0% found this document useful (0 votes)
17 views14 pages

BasicVSR: Essential Components in VSR

Uploaded by

matin fazel
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
17 views14 pages

BasicVSR: Essential Components in VSR

Uploaded by

matin fazel
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

BasicVSR: The Search for Essential Components in

Video Super-Resolution and Beyond

Kelvin C.K. Chan1 Xintao Wang2 Ke Yu3 Chao Dong4,5 Chen Change Loy1 *
1
S-Lab, Nanyang Technological University 2 Applied Research Center, Tencent PCG
3
CUHK – SenseTime Joint Lab, The Chinese University of Hong Kong
4
Shenzhen Key Lab of Computer Vision and Pattern Recognition, SIAT-SenseTime Joint Lab,
arXiv:2012.02181v2 [[Link]] 7 Apr 2021

Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences


5
SIAT Branch, Shenzhen Institute of Artificial Intelligence and Robotics for Society
{chan0899, ccloy}@[Link] [Link]@[Link]
yk017@[Link] [Link]@[Link]

Abstract
IconVSR
Video super-resolution (VSR) approaches tend to have BasicVSR (ours)
(ours) EDVR
more components than the image counterparts as they need (CVPRW19)
to exploit the additional temporal dimension. Complex de-
RSDN RBPN
signs are not uncommon. In this study, we wish to untangle (ECCV20) (CVPR19)
the knots and reconsider some most essential components
DUF
for VSR guided by four basic functionalities, i.e., Propaga- PFNL (CVPR18)
RLSP (ICCV19)
tion, Alignment, Aggregation, and Upsampling. By reusing (ICCVW19) # Params
some existing components added with minimal redesigns,
we show a succinct pipeline, BasicVSR, that achieves ap- FRVSR
pealing improvements in terms of speed and restoration (CVPR18) 5M 10M 15M 20M

quality in comparison to many state-of-the-art algorithms.


We conduct systematic analysis to explain how such gain
can be obtained and discuss the pitfalls. We further show Figure 1. Speed and performance comparison. Without bells
the extensibility of BasicVSR by presenting an information- and whistles, BasicVSR outperforms state-of-the-art methods with
refill mechanism and a coupled propagation scheme to fa- high efficiency. Built upon BasicVSR, IconVSR further im-
cilitate information aggregation. The BasicVSR and its ex- proves the performance. Comparisons are performed on UDM10
dataset [34].
tension, IconVSR, can serve as strong baselines for future
VSR approaches.
from different frames. In RBPN [9], multiple projection
modules are used to sequentially aggregate features from
1. Introduction multiple frames. Such designs are effective but inevitably
Compared to single-image super-resolution, which fo- increase the runtime and model complexity (see Fig. 1).
cuses on the intrinsic properties of a single image for the In addition, unlike SISR, the potentially complex and dis-
upscaling task, video super-resolution (VSR) poses an ex- similar designs of VSR methods pose difficulties in imple-
tra challenge as it involves aggregating information from menting and extending existing approaches, hampering re-
multiple highly-related but misaligned frames in video se- producibility and fair comparisons.
quences. There is a need to step back and reconsider the di-
Various approaches have been proposed to address the verse designs of VSR models, with the aim to search for
challenge. Some designs can be highly complex. For in- a more generic, efficient, and easy-to-implement baseline
stance, in the representative method EDVR [32], a multi- for VSR. We start our search by decomposing popular VSR
scale deformable alignment module and multiple attention approaches into submodules based on functionalities. As
layers are adopted for aligning and integrating the features summarized in Table 1, most existing methods entail four
inter-related components, namely, propagation, alignment,
* Corresponding author aggregation, and upsampling. Such a decomposition allows
Table 1. Components in existing VSR methods. We categorize components based on their functionalities: i) Propagation refers to the way
in which features are propagated temporally, ii) Alignment concerns on the spatial transformation applied to misaligned images/features, iii)
Aggregation defines the steps to combine aligned features, and iv) Upsampling describes the method to transform the aggregated features
to the final output image. Bolded texts correspond to designs that were reported to achieve better performance in the literature.

Sliding-Window Recurrent
EDVR [31] MuCAN [20] TDAN [30] BRCN [10, 11] FRVSR [25] RSDN [12] BasicVSR IconVSR
Propagation Local Local Local Bidirectional Unidirectional Unidirectional Bidirectional Bidirectional (coupled)
Alignment Yes (DCN) Yes (correlation) Yes (DCN) No Yes (flow) No Yes (flow) Yes (flow)
Aggregation Concatenate + TSA Concatenate Concatenate Concatenate Concatenate Concatenate Concatenate Concatenate + Refill
Upsampling Pixel-Shuffle Pixel-Shuffle Pixel-Shuffle Pixel-Shuffle Pixel-Shuffle Pixel-Shuffle Pixel-Shuffle Pixel-Shuffle

us to systematically study various options under each com- present an efficient baseline for VSR. We show that sim-
ponent and understand their pros and cons. ple components, when integrated properly, would synergize
Through extensive experiments, we find that with mini- and lead to state-of-the-art performance. We further present
mal redesigns of existing options, one could already reach a an example of extending BasicVSR with two novel modules
strong yet efficient baseline for VSR without bells and whis- to refine the propagation and aggregation components.
tles. In this paper, we highlight one of such possibilities,
named BasicVSR. We observe that, among the four afore- 2. Related Work
mentioned components, the choices of propagation and
Existing VSR approaches [10, 21, 28, 34, 20, 12, 13] can
alignment components could lead to a big swing in terms
be mainly divided into two frameworks – sliding-window
of performance and efficiency. Our experiments suggest the
and recurrent. Earlier methods [1, 29, 33] in the sliding-
use of bidirectional propagation scheme to maximize infor-
window framework predict the optical flow between low-
mation gathering, and an optical flow-based method to esti-
resolution (LR) frames and perform spatial warping for
mate the correspondence between two neighboring frames
alignment. Later approaches resort to a more sophisticated
for feature alignment. By simply streamlining these prop-
approach of implicit alignment. For example, TDAN [30]
agation and alignment components with the commonly-
adopts deformable convolutions (DCNs) [5, 37] to align
adopted designs for aggregation (i.e. feature concatenation)
different frames at the feature level. EDVR [32] further
and upsampling (i.e. pixel-shuffle [27]), BasicVSR outper-
uses DCNs in a multi-scale fashion for more accurate align-
forms existing state of the arts [9, 12, 32] in both perfor-
ment. DUF [16] leverages dynamic upsampling filters to
mance (up to 0.61 dB) and efficiency (up to 24× speedup).
handle motions implicitly. Some approaches take a recur-
Thanks to its simplicity and versatility, BasicVSR pro- rent framework. RSDN [12] proposes a recurrent detail-
vides a viable starting point for extending to more elabo- structural block and a hidden state adaptation module to
rated networks. By using BasicVSR as a foundation, we enhance the robustness to appearance change and error ac-
present IconVSR that comprises two novel extensions to cumulation. RRN [14] adopts a residual mapping between
improve the aggregation and the propagation components. layers with identity skip connections to ensure a fluent in-
The first extension is named information-refill. This mech- formation flow and preserve the texture information over
anism leverages an additional module to extract features long periods. The aforementioned studies have led to many
from sparsely selected frames (keyframes), and the fea- new and sophisticated components to address the propaga-
tures are then inserted into the main network for feature re- tion and alignment problems in VSR. Here, we reinvestigate
finement. The second extension is a coupled propagation some of the components and find that bidirectional propaga-
scheme, which facilitates information exchange between tion coupled with a simple optical flow-based feature align-
the forward and backward propagation branches. The two ment suffice to outperform many state-of-the-art methods.
modules not only reduce error accumulation during propa- The information-refill mechanism in IconVSR is remi-
gation due to occlusions and image boundaries, but also al- niscent of the concept of interval-based processing [4, 15,
low the propagation to access complete information in a se- 26, 35, 36, 38, 39]. These methods divide video frames
quence for generating high-quality features. With these two into independent intervals characterized by keyframes and
new designs, IconVSR surpasses BasicVSR with a PSNR non-keyframes. The keyframes and non-keyframes are then
improvement of up to 0.31 dB. processed by different pipelines. For instance, FAST [35]
We believe that our work is timely, given the increasing applies SRCNN [6, 7] to super-resolve the keyframes. Non-
number of approaches centered around the research of VSR. keyframes are then restored using the upscaled keyframes
A strong, simple yet extensible baseline is needed. Guided and the motion vectors stored in the compressed video
by the main functionalities in VSR approaches, we recon- codec. IconVSR inherits the concept of keyframes, but
sider some essential components in existing pipelines and unlike existing methods that process the intervals indepen-
··· <latexit sha1_base64="C/ThE3LO0bsU2xdfe5suNCnJ9S4=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69LBbBU0lE0GPRi8cK9gPaUDabTbt2kw27E6GE/gcvHhTx6v/x5r9x2+agrQ8GHu/NMDMvSKUw6LrfTmltfWNzq7xd2dnd2z+oHh61jco04y2mpNLdgBouRcJbKFDybqo5jQPJO8H4duZ3nrg2QiUPOEm5H9NhIiLBKFqp3WehQjOo1ty6OwdZJV5BalCgOah+9UPFspgnyCQ1pue5Kfo51SiY5NNKPzM8pWxMh7xnaUJjbvx8fu2UnFklJJHSthIkc/X3RE5jYyZxYDtjiiOz7M3E/7xehtG1n4skzZAnbLEoyiRBRWavk1BozlBOLKFMC3srYSOqKUMbUMWG4C2/vEraF3XPrXv3l7XGTRFHGU7gFM7BgytowB00oQUMHuEZXuHNUc6L8+58LFpLTjFzDH/gfP4ArlePLw==</latexit>

···
<latexit sha1_base64="C/ThE3LO0bsU2xdfe5suNCnJ9S4=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69LBbBU0lE0GPRi8cK9gPaUDabTbt2kw27E6GE/gcvHhTx6v/x5r9x2+agrQ8GHu/NMDMvSKUw6LrfTmltfWNzq7xd2dnd2z+oHh61jco04y2mpNLdgBouRcJbKFDybqo5jQPJO8H4duZ3nrg2QiUPOEm5H9NhIiLBKFqp3WehQjOo1ty6OwdZJV5BalCgOah+9UPFspgnyCQ1pue5Kfo51SiY5NNKPzM8pWxMh7xnaUJjbvx8fu2UnFklJJHSthIkc/X3RE5jYyZxYDtjiiOz7M3E/7xehtG1n4skzZAnbLEoyiRBRWavk1BozlBOLKFMC3srYSOqKUMbUMWG4C2/vEraF3XPrXv3l7XGTRFHGU7gFM7BgytowB00oQUMHuEZXuHNUc6L8+58LFpLTjFzDH/gfP4ArlePLw==</latexit>
xi
<latexit sha1_base64="96Omxy49SPGED3hOl6/e89unjWM=">AAAB6nicbVBNS8NAEJ3Ur1q/qh69LBbBU0lEqMeiF48V7Qe0oWy2k3bpZhN2N2IJ/QlePCji1V/kzX/jts1BWx8MPN6bYWZekAiujet+O4W19Y3NreJ2aWd3b/+gfHjU0nGqGDZZLGLVCahGwSU2DTcCO4lCGgUC28H4Zua3H1FpHssHM0nQj+hQ8pAzaqx0/9Tn/XLFrbpzkFXi5aQCORr98ldvELM0QmmYoFp3PTcxfkaV4UzgtNRLNSaUjekQu5ZKGqH2s/mpU3JmlQEJY2VLGjJXf09kNNJ6EgW2M6JmpJe9mfif101NeOVnXCapQckWi8JUEBOT2d9kwBUyIyaWUKa4vZWwEVWUGZtOyYbgLb+8SloXVc+teneXlfp1HkcRTuAUzsGDGtThFhrQBAZDeIZXeHOE8+K8Ox+L1oKTzxzDHzifP2CQjdg=</latexit>
xi
<latexit sha1_base64="96Omxy49SPGED3hOl6/e89unjWM=">AAAB6nicbVBNS8NAEJ3Ur1q/qh69LBbBU0lEqMeiF48V7Qe0oWy2k3bpZhN2N2IJ/QlePCji1V/kzX/jts1BWx8MPN6bYWZekAiujet+O4W19Y3NreJ2aWd3b/+gfHjU0nGqGDZZLGLVCahGwSU2DTcCO4lCGgUC28H4Zua3H1FpHssHM0nQj+hQ8pAzaqx0/9Tn/XLFrbpzkFXi5aQCORr98ldvELM0QmmYoFp3PTcxfkaV4UzgtNRLNSaUjekQu5ZKGqH2s/mpU3JmlQEJY2VLGjJXf09kNNJ6EgW2M6JmpJe9mfif101NeOVnXCapQckWi8JUEBOT2d9kwBUyIyaWUKa4vZWwEVWUGZtOyYbgLb+8SloXVc+teneXlfp1HkcRTuAUzsGDGtThFhrQBAZDeIZXeHOE8+K8Ox+L1oKTzxzDHzifP2CQjdg=</latexit>

Fb Ff Fb
U xi xi+1
xi+1
<latexit sha1_base64="iGci6Aoi0MlEX9iTlkpWbFwtl7A=">AAAB7nicbVBNS8NAEJ3Ur1q/qh69LBZBEEoiBT0WvXisYD+gDWWznbRLN5uwuxFL6I/w4kERr/4eb/4bt20O2vpg4PHeDDPzgkRwbVz32ymsrW9sbhW3Szu7e/sH5cOjlo5TxbDJYhGrTkA1Ci6xabgR2EkU0igQ2A7GtzO//YhK81g+mEmCfkSHkoecUWOl9lM/4xfetF+uuFV3DrJKvJxUIEejX/7qDWKWRigNE1Trrucmxs+oMpwJnJZ6qcaEsjEdYtdSSSPUfjY/d0rOrDIgYaxsSUPm6u+JjEZaT6LAdkbUjPSyNxP/87qpCa/9jMskNSjZYlGYCmJiMvudDLhCZsTEEsoUt7cSNqKKMmMTKtkQvOWXV0nrsuq5Ve++Vqnf5HEU4QRO4Rw8uII63EEDmsBgDM/wCm9O4rw4787HorXg5DPH8AfO5w/+tI9U</latexit>
Ff <latexit sha1_base64="YiaPGNfyhnPSoaBVjh+5KKPfOig=">AAAB7nicbVBNS8NAEJ3Ur1q/qh69LBbBiyWRgh6LXjxWsB/QhrLZTtqlm03Y3Ygl9Ed48aCIV3+PN/+N2zYHbX0w8Hhvhpl5QSK4Nq777RTW1jc2t4rbpZ3dvf2D8uFRS8epYthksYhVJ6AaBZfYNNwI7CQKaRQIbAfj25nffkSleSwfzCRBP6JDyUPOqLFS+6mf8Qtv2i9X3Ko7B1klXk4qkKPRL3/1BjFLI5SGCap113MT42dUGc4ETku9VGNC2ZgOsWuppBFqP5ufOyVnVhmQMFa2pCFz9fdERiOtJ1FgOyNqRnrZm4n/ed3UhNd+xmWSGpRssShMBTExmf1OBlwhM2JiCWWK21sJG1FFmbEJlWwI3vLLq6R1WfXcqndfq9Rv8jiKcAKncA4eXEEd7qABTWAwhmd4hTcncV6cd+dj0Vpw8plj+APn8wcBz49W</latexit>
1
S S
<latexit sha1_base64="iGci6Aoi0MlEX9iTlkpWbFwtl7A=">AAAB7nicbVBNS8NAEJ3Ur1q/qh69LBZBEEoiBT0WvXisYD+gDWWznbRLN5uwuxFL6I/w4kERr/4eb/4bt20O2vpg4PHeDDPzgkRwbVz32ymsrW9sbhW3Szu7e/sH5cOjlo5TxbDJYhGrTkA1Ci6xabgR2EkU0igQ2A7GtzO//YhK81g+mEmCfkSHkoecUWOl9lM/4xfetF+uuFV3DrJKvJxUIEejX/7qDWKWRigNE1Trrucmxs+oMpwJnJZ6qcaEsjEdYtdSSSPUfjY/d0rOrDIgYaxsSUPm6u+JjEZaT6LAdkbUjPSyNxP/87qpCa/9jMskNSjZYlGYCmJiMvudDLhCZsTEEsoUt7cSNqKKMmMTKtkQvOWXV0nrsuq5Ve++Vqnf5HEU4QRO4Rw8uII63EEDmsBgDM/wCm9O4rw4787HorXg5DPH8AfO5w/+tI9U</latexit>

sfi sbi
Fb <latexit sha1_base64="xVz7wWYHU0XVC0yW7ELAVxcS3eQ=">AAAB7HicbVBNS8NAEJ3Ur1q/qh69LBbBU0lE0GPRi8cKpi20sWy2m3bpZhN2J0IJ/Q1ePCji1R/kzX/jts1BWx8MPN6bYWZemEph0HW/ndLa+sbmVnm7srO7t39QPTxqmSTTjPsskYnuhNRwKRT3UaDknVRzGoeSt8Px7cxvP3FtRKIecJLyIKZDJSLBKFrJN33xGPWrNbfuzkFWiVeQGhRo9qtfvUHCspgrZJIa0/XcFIOcahRM8mmllxmeUjamQ961VNGYmyCfHzslZ1YZkCjRthSSufp7IqexMZM4tJ0xxZFZ9mbif143w+g6yIVKM+SKLRZFmSSYkNnnZCA0ZygnllCmhb2VsBHVlKHNp2JD8JZfXiWti7rn1r37y1rjpoijDCdwCufgwRU04A6a4AMDAc/wCm+Ocl6cd+dj0Vpyiplj+APn8wfOt46r</latexit> <latexit sha1_base64="9l9wzYx4mEos9MpTVwY07IVVAR4=">AAAB7HicbVBNS8NAEJ3Ur1q/qh69LBbBU0lE0GPRi8cKpi20sWy2m3bpZhN2J0IJ/Q1ePCji1R/kzX/jts1BWx8MPN6bYWZemEph0HW/ndLa+sbmVnm7srO7t39QPTxqmSTTjPsskYnuhNRwKRT3UaDknVRzGoeSt8Px7cxvP3FtRKIecJLyIKZDJSLBKFrJN33xGParNbfuzkFWiVeQGhRo9qtfvUHCspgrZJIa0/XcFIOcahRM8mmllxmeUjamQ961VNGYmyCfHzslZ1YZkCjRthSSufp7IqexMZM4tJ0xxZFZ9mbif143w+g6yIVKM+SKLRZFmSSYkNnnZCA0ZygnllCmhb2VsBHVlKHNp2JD8JZfXiWti7rn1r37y1rjpoijDCdwCufgwRU04A6a4AMDAc/wCm+Ocl6cd+dj0Vpyiplj+APn8wfIp46n</latexit>

U hfi 1 hbi+1
Ff
<latexit sha1_base64="M0CBOL9CS95kZtltHMy/DSAYIAE=">AAAB8HicbVDLSgNBEOyNrxhfUY9eBoPgxbArgh6DXjxGMA9J1jA7mU2GzGOZmRXCkq/w4kERr36ON//GSbIHTSxoKKq66e6KEs6M9f1vr7Cyura+UdwsbW3v7O6V9w+aRqWa0AZRXOl2hA3lTNKGZZbTdqIpFhGnrWh0M/VbT1QbpuS9HSc0FHggWcwItk56GPYydhZMHuNeueJX/RnQMglyUoEc9V75q9tXJBVUWsKxMZ3AT2yYYW0Z4XRS6qaGJpiM8IB2HJVYUBNms4Mn6MQpfRQr7UpaNFN/T2RYGDMWkesU2A7NojcV//M6qY2vwozJJLVUkvmiOOXIKjT9HvWZpsTysSOYaOZuRWSINSbWZVRyIQSLLy+T5nk18KvB3UWldp3HUYQjOIZTCOASanALdWgAAQHP8ApvnvZevHfvY95a8PKZQ/gD7/MHYZqQHg==</latexit>

W W <latexit sha1_base64="VOvmnYJWcsnmkidN/kgox3pYJbQ=">AAAB8HicbVDLSgNBEOyNrxhfUY9eBoMgCGFXBD0GvXiMYB6SrGF2MpsMmccyMyuEJV/hxYMiXv0cb/6Nk2QPmljQUFR1090VJZwZ6/vfXmFldW19o7hZ2tre2d0r7x80jUo1oQ2iuNLtCBvKmaQNyyyn7URTLCJOW9HoZuq3nqg2TMl7O05oKPBAspgRbJ30MOxl7CyYPEa9csWv+jOgZRLkpAI56r3yV7evSCqotIRjYzqBn9gww9oywumk1E0NTTAZ4QHtOCqxoCbMZgdP0IlT+ihW2pW0aKb+nsiwMGYsItcpsB2aRW8q/ud1UhtfhRmTSWqpJPNFccqRVWj6PeozTYnlY0cw0czdisgQa0ysy6jkQggWX14mzfNq4FeDu4tK7TqPowhHcAynEMAl1OAW6tAAAgKe4RXePO29eO/ex7y14OUzh/AH3ucPWHqQGA==</latexit>

xi
<latexit sha1_base64="96Omxy49SPGED3hOl6/e89unjWM=">AAAB6nicbVBNS8NAEJ3Ur1q/qh69LBbBU0lEqMeiF48V7Qe0oWy2k3bpZhN2N2IJ/QlePCji1V/kzX/jts1BWx8MPN6bYWZekAiujet+O4W19Y3NreJ2aWd3b/+gfHjU0nGqGDZZLGLVCahGwSU2DTcCO4lCGgUC28H4Zua3H1FpHssHM0nQj+hQ8pAzaqx0/9Tn/XLFrbpzkFXi5aQCORr98ldvELM0QmmYoFp3PTcxfkaV4UzgtNRLNSaUjekQu5ZKGqH2s/mpU3JmlQEJY2VLGjJXf09kNNJ6EgW2M6JmpJe9mfif101NeOVnXCapQckWi8JUEBOT2d9kwBUyIyaWUKa4vZWwEVWUGZtOyYbgLb+8SloXVc+teneXlfp1HkcRTuAUzsGDGtThFhrQBAZDeIZXeHOE8+K8Ox+L1oKTzxzDHzifP2CQjdg=</latexit>

h̄fi
<latexit sha1_base64="CMJrtgivqIqUOueRy+KmM83HAJo=">AAAB8nicbVBNS8NAEJ3Ur1q/qh69LBbBU0lE0GPRi8cK9gPSWDbbTbt0swm7E6GE/gwvHhTx6q/x5r9x2+agrQ8GHu/NMDMvTKUw6LrfTmltfWNzq7xd2dnd2z+oHh61TZJpxlsskYnuhtRwKRRvoUDJu6nmNA4l74Tj25nfeeLaiEQ94CTlQUyHSkSCUbSS3wupzkfTvniM+tWaW3fnIKvEK0gNCjT71a/eIGFZzBUySY3xPTfFIKcaBZN8WullhqeUjemQ+5YqGnMT5POTp+TMKgMSJdqWQjJXf0/kNDZmEoe2M6Y4MsveTPzP8zOMroNcqDRDrthiUZRJggmZ/U8GQnOGcmIJZVrYWwkbUU0Z2pQqNgRv+eVV0r6oe27du7+sNW6KOMpwAqdwDh5cQQPuoAktYJDAM7zCm4POi/PufCxaS04xcwx/4Hz+AIXikWU=</latexit>
h̄bi
<latexit sha1_base64="LNzLyFwqCE+LB1+mo9QuOk9HMlI=">AAAB8nicbVBNS8NAEJ3Ur1q/qh69LBbBU0lE0GPRi8cK9gPSWDbbTbt0swm7E6GE/gwvHhTx6q/x5r9x2+agrQ8GHu/NMDMvTKUw6LrfTmltfWNzq7xd2dnd2z+oHh61TZJpxlsskYnuhtRwKRRvoUDJu6nmNA4l74Tj25nfeeLaiEQ94CTlQUyHSkSCUbSS3wupzkfTvngM+9WaW3fnIKvEK0gNCjT71a/eIGFZzBUySY3xPTfFIKcaBZN8WullhqeUjemQ+5YqGnMT5POTp+TMKgMSJdqWQjJXf0/kNDZmEoe2M6Y4MsveTPzP8zOMroNcqDRDrthiUZRJggmZ/U8GQnOGcmIJZVrYWwkbUU0Z2pQqNgRv+eVV0r6oe27du7+sNW6KOMpwAqdwDh5cQQPuoAktYJDAM7zCm4POi/PufCxaS04xcwx/4Hz+AH/SkWE=</latexit>

Fb Rf Rb
U hfi hbi
xi
<latexit sha1_base64="YiaPGNfyhnPSoaBVjh+5KKPfOig=">AAAB7nicbVBNS8NAEJ3Ur1q/qh69LBbBiyWRgh6LXjxWsB/QhrLZTtqlm03Y3Ygl9Ed48aCIV3+PN/+N2zYHbX0w8Hhvhpl5QSK4Nq777RTW1jc2t4rbpZ3dvf2D8uFRS8epYthksYhVJ6AaBZfYNNwI7CQKaRQIbAfj25nffkSleSwfzCRBP6JDyUPOqLFS+6mf8Qtv2i9X3Ko7B1klXk4qkKPRL3/1BjFLI5SGCap113MT42dUGc4ETku9VGNC2ZgOsWuppBFqP5ufOyVnVhmQMFa2pCFz9fdERiOtJ1FgOyNqRnrZm4n/ed3UhNd+xmWSGpRssShMBTExmf1OBlwhM2JiCWWK21sJG1FFmbEJlWwI3vLLq6R1WfXcqndfq9Rv8jiKcAKncA4eXEEd7qABTWAwhmd4hTcncV6cd+dj0Vpw8plj+APn8wcBz49W</latexit>
1 Ff <latexit sha1_base64="Rsw+ZNw1R0mD4pP5NhcE2HnoO2A=">AAAB7HicbVBNS8NAEJ3Ur1q/qh69LBbBU0lE0GPRi8cKpi20sWy2k3bpZhN2N0IJ/Q1ePCji1R/kzX/jts1BWx8MPN6bYWZemAqujet+O6W19Y3NrfJ2ZWd3b/+genjU0kmmGPosEYnqhFSj4BJ9w43ATqqQxqHAdji+nfntJ1SaJ/LBTFIMYjqUPOKMGiv5oz5/jPrVmlt35yCrxCtIDQo0+9Wv3iBhWYzSMEG17npuaoKcKsOZwGmll2lMKRvTIXYtlTRGHeTzY6fkzCoDEiXKljRkrv6eyGms9SQObWdMzUgvezPxP6+bmeg6yLlMM4OSLRZFmSAmIbPPyYArZEZMLKFMcXsrYSOqKDM2n4oNwVt+eZW0LuqeW/fuL2uNmyKOMpzAKZyDB1fQgDtogg8MODzDK7w50nlx3p2PRWvJKWaO4Q+czx+9346g</latexit> <latexit sha1_base64="hrlW+5FT5NXZogvKrTRYWvsPeH0=">AAAB7nicbVBNS8NAEJ3Ur1q/qh69LBbBU0lE0GPRi8cK9gPaWDbbTbt0swm7E6GE/ggvHhTx6u/x5r9x0+agrQ8GHu/NMDMvSKQw6LrfTmltfWNzq7xd2dnd2z+oHh61TZxqxlsslrHuBtRwKRRvoUDJu4nmNAok7wST29zvPHFtRKwecJpwP6IjJULBKFqpM34MBpmYDao1t+7OQVaJV5AaFGgOql/9YczSiCtkkhrT89wE/YxqFEzyWaWfGp5QNqEj3rNU0YgbP5ufOyNnVhmSMNa2FJK5+nsio5Ex0yiwnRHFsVn2cvE/r5dieO1nQiUpcsUWi8JUEoxJ/jsZCs0ZyqkllGlhbyVsTDVlaBOq2BC85ZdXSfui7rl17/6y1rgp4ijDCZzCOXhwBQ24gya0gMEEnuEV3pzEeXHenY9Fa8kpZo7hD5zPH34Ej6g=</latexit>

hfi hbi
··· <latexit sha1_base64="C/ThE3LO0bsU2xdfe5suNCnJ9S4=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69LBbBU0lE0GPRi8cK9gPaUDabTbt2kw27E6GE/gcvHhTx6v/x5r9x2+agrQ8GHu/NMDMvSKUw6LrfTmltfWNzq7xd2dnd2z+oHh61jco04y2mpNLdgBouRcJbKFDybqo5jQPJO8H4duZ3nrg2QiUPOEm5H9NhIiLBKFqp3WehQjOo1ty6OwdZJV5BalCgOah+9UPFspgnyCQ1pue5Kfo51SiY5NNKPzM8pWxMh7xnaUJjbvx8fu2UnFklJJHSthIkc/X3RE5jYyZxYDtjiiOz7M3E/7xehtG1n4skzZAnbLEoyiRBRWavk1BozlBOLKFMC3srYSOqKUMbUMWG4C2/vEraF3XPrXv3l7XGTRFHGU7gFM7BgytowB00oQUMHuEZXuHNUc6L8+58LFpLTjFzDH/gfP4ArlePLw==</latexit>

···
<latexit sha1_base64="C/ThE3LO0bsU2xdfe5suNCnJ9S4=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69LBbBU0lE0GPRi8cK9gPaUDabTbt2kw27E6GE/gcvHhTx6v/x5r9x2+agrQ8GHu/NMDMvSKUw6LrfTmltfWNzq7xd2dnd2z+oHh61jco04y2mpNLdgBouRcJbKFDybqo5jQPJO8H4duZ3nrg2QiUPOEm5H9NhIiLBKFqp3WehQjOo1ty6OwdZJV5BalCgOah+9UPFspgnyCQ1pue5Kfo51SiY5NNKPzM8pWxMh7xnaUJjbvx8fu2UnFklJJHSthIkc/X3RE5jYyZxYDtjiiOz7M3E/7xehtG1n4skzZAnbLEoyiRBRWavk1BozlBOLKFMC3srYSOqKUMbUMWG4C2/vEraF3XPrXv3l7XGTRFHGU7gFM7BgytowB00oQUMHuEZXuHNUc6L8+58LFpLTjFzDH/gfP4ArlePLw==</latexit>
<latexit sha1_base64="Rsw+ZNw1R0mD4pP5NhcE2HnoO2A=">AAAB7HicbVBNS8NAEJ3Ur1q/qh69LBbBU0lE0GPRi8cKpi20sWy2k3bpZhN2N0IJ/Q1ePCji1R/kzX/jts1BWx8MPN6bYWZemAqujet+O6W19Y3NrfJ2ZWd3b/+genjU0kmmGPosEYnqhFSj4BJ9w43ATqqQxqHAdji+nfntJ1SaJ/LBTFIMYjqUPOKMGiv5oz5/jPrVmlt35yCrxCtIDQo0+9Wv3iBhWYzSMEG17npuaoKcKsOZwGmll2lMKRvTIXYtlTRGHeTzY6fkzCoDEiXKljRkrv6eyGms9SQObWdMzUgvezPxP6+bmeg6yLlMM4OSLRZFmSAmIbPPyYArZEZMLKFMcXsrYSOqKDM2n4oNwVt+eZW0LuqeW/fuL2uNmyKOMpzAKZyDB1fQgDtogg8MODzDK7w50nlx3p2PRWvJKWaO4Q+czx+9346g</latexit>
<latexit sha1_base64="hrlW+5FT5NXZogvKrTRYWvsPeH0=">AAAB7nicbVBNS8NAEJ3Ur1q/qh69LBbBU0lE0GPRi8cK9gPaWDbbTbt0swm7E6GE/ggvHhTx6u/x5r9x0+agrQ8GHu/NMDMvSKQw6LrfTmltfWNzq7xd2dnd2z+oHh61TZxqxlsslrHuBtRwKRRvoUDJu4nmNAok7wST29zvPHFtRKwecJpwP6IjJULBKFqpM34MBpmYDao1t+7OQVaJV5AaFGgOql/9YczSiCtkkhrT89wE/YxqFEzyWaWfGp5QNqEj3rNU0YgbP5ufOyNnVhmSMNa2FJK5+nsio5Ex0yiwnRHFsVn2cvE/r5dieO1nQiUpcsUWi8JUEoxJ/jsZCs0ZyqkllGlhbyVsTDVlaBOq2BC85ZdXSfui7rl17/6y1rgp4ijDCZzCOXhwBQ24gya0gMEEnuEV3pzEeXHenY9Fa8kpZo7hD5zPH34Ej6g=</latexit>

(a) BasicVSR architecture (b) Forward and backward propagation branches

Figure 2. An overview of BasicVSR. BasicVSR is a generic and efficient baseline for VSR. With minimal redesigns of existing components
including optical flow and residual blocks, it outperforms existing state of the arts with high efficiency. (a) BasicVSR adopts a typical
bidirectional recurrent network. The upsampling module U contains multiple pixel-shuffle and convolutions. The red and blue colors
represent the backward and forward propagations, respectively. (b) The propagation branches contain only generic components. S, W ,
and R refer to the flow estimation module, spatial warping module, and residual blocks, respectively.

dently, we make one advancement by connecting the inter- and bidirectional propagations. In what follows, we discuss
vals through the propagation branches. With this design, the weaknesses of the former two to motivate our choice of
long-term information can be propagated across the inter- bidirectional propagation in BasicVSR.
connected intervals, further improving the effectiveness.
• Local Propagation. The sliding-window methods [9,
3. Methodology 13, 32] take the LR images within a local window as
inputs and employ the local information for restora-
Video super-resolution, by nature, involves a long and tion. In this design, the accessible information is re-
complex processing pipeline since it needs to aggregate in- stricted in a local neighborhood. The omittance of
formation from not only the spatial dimension but also the distant frames inevitably limits the potential of the
temporal dimension. Existing studies typically focus on sliding-window methods. To verify our claim, we start
one aspect of the functionalities to make advancement and with a global receptive field (in the temporal dimen-
may not collectively consider the synergy of various com- sion) and gradually reduce the receptive field. We sep-
ponents. There is an urge to revisit various components arate the test sequences into K segments and use our
macroscopically and uncover a generic baseline that inher- BasicVSR to restore each segment independently. The
its the strengths of existing approaches. In this work, we PSNR difference to the case K=1 (global propaga-
conduct extensive analysis and present a simple, strong and tion) is depicted in Fig. 3.
versatile baseline, BasicVSR, which can serve as a back- First, the difference in PSNR is reduced (i.e. better
bone with abundant flexibilities in design. performance) when the number of segments decreases
(i.e. temporal receptive field increases). This suggests
3.1. BasicVSR
that the information in distant frames is beneficial to
Aiming at discovering generic frameworks for facilitat- the restoration and should not be neglected. Second,
ing analysis and development of VSR methods, we confine the difference in PSNR is the largest at the two ends
our search to commonly-adopted elements such as optical of each segment, indicating the necessity of adopting
flow and residual blocks. An overview of BasicVSR is de- long sequences to accumulate long-term information.
picted in Fig. 2.
Propagation. Propagation is one of the most influential • Unidirectional Propagation. The aforementioned
components in VSR. It specifies how the information in a problem can be resolved by adopting a unidirectional
video sequence is leveraged. Existing propagation schemes propagation [8, 12, 14, 25], where the information is
can be divided into three main groups: local, unidirectional sequentially propagated from the first frame to the last
Figure 3. Local vs Global propagation. When the number of seg- Figure 4. Unidirectional vs Bidirectional. In unidirectional prop-
ments K is reduced, the increased temporal receptive field leads agation, earlier timesteps receive less information, leading to infe-
to higher PSNR. This demonstrates the importance of aggregat- rior performance. Values smaller than zero (dotted line) indicates
ing long-term information. Values smaller than zero (dotted line) a lower PSNR than the bidirectional counterpart. Note that the
indicate a lower PSNR than the case K=1. unidirectional model outperforms the bidirectional model only for
the last frame, owing to the zero feature initialization in the bidi-
rectional model.
frame. However, in this setting, the information re-
ceived by different frames is imbalanced. Specifically,
the first frame receives no information from the video aligned images/features for subsequent aggregation. Main-
sequence except itself, whereas the last frame receives stream works can be divided into three categories: without
information from the whole sequence. Hence, subop- alignment, image alignment, and feature alignment. In this
timal results are expected for the earlier frames. section, we conduct experiments to analyze each of the cat-
egories and to validate our choice of feature alignment.
To demonstrate the effects, we compare BasicVSR (us-
ing bidirectional propagation) with its unidirectional
variant (with comparable network complexity). From • Without Alignment. Existing recurrent methods [8,
Fig. 4, we see that the unidirectional model obtains 10, 11, 12, 14] generally do not perform alignment dur-
a significantly lower PSNR than bidirectional propa- ing propagation. The non-aligned features/images im-
gation at early timesteps, and the difference gradually pede aggregation and eventually lead to substandard
reduces as more information is aggregated with the in- performance. This suboptimality can be reflected by
crease in the number of frames. Moreover, a consistent our experiment, where we remove the spatial align-
performance drop of 0.5 dB is observed with only par- ment module in BasicVSR. In this case, we directly
tial information employed. These observations reveal concatenate the non-aligned features for restoration.
the suboptimality of unidirectional propagation. One Without proper alignment, the propagated features are
can improve the output quality by propagating infor- not spatially aligned with the input image. As a re-
mation back from the last frame of the sequence. sult, the local operations such as convolutions, which
have relatively small receptive fields, are inefficient in
• Bidirectional Propagation. The above two problems aggregating the information from corresponding loca-
can be simultaneously addressed by bidirectional prop- tions. A drop of 1.19 dB of PSNR is observed. This
agation, in which the features are propagated forward result suggests that it is pivotal to adopt operations that
and backward in time independently. Motivated by have a large enough receptive field to aggregate infor-
this, BasicVSR adopts a typical bidirectional propa- mation from distant spatial locations.
gation scheme. Given an LR image xi , its neighboring
frames xi−1 and xi+1 , and the corresponding features • Image Alignment. Earlier works [17, 33] perform
propagated from its neighbors, denoted as hfi−1 and alignment by computing the optical flow and warping
hbi+1 , we have the images before restoration. Recently, Chan et al. [2]
show that moving the spatial alignment from the image
hbi = Fb (xi , xi+1 , hbi+1 ), level to the feature level yields a marked improvement.
(1) In this work, we further conduct experiments to verify
hfi = Ff (xi , xi−1 , hfi−1 ),
their claim. We compare image warping and feature
where Fb and Ff denote the backward and forward warping1 on a variant of BasicVSR. Resulting from
propagation branches, respectively. the inaccuracy of optical flow estimation, the warped
Alignment. Spatial alignment plays an important role in 1 We compute optical flow from the images and use the optical flow for
VSR as it is responsible to align highly related but mis- feature warping.
images inevitably suffer from blurriness and incorrect- xi
<latexit sha1_base64="96Omxy49SPGED3hOl6/e89unjWM=">AAAB6nicbVBNS8NAEJ3Ur1q/qh69LBbBU0lEqMeiF48V7Qe0oWy2k3bpZhN2N2IJ/QlePCji1V/kzX/jts1BWx8MPN6bYWZekAiujet+O4W19Y3NreJ2aWd3b/+gfHjU0nGqGDZZLGLVCahGwSU2DTcCO4lCGgUC28H4Zua3H1FpHssHM0nQj+hQ8pAzaqx0/9Tn/XLFrbpzkFXi5aQCORr98ldvELM0QmmYoFp3PTcxfkaV4UzgtNRLNSaUjekQu5ZKGqH2s/mpU3JmlQEJY2VLGjJXf09kNNJ6EgW2M6JmpJe9mfif101NeOVnXCapQckWi8JUEBOT2d9kwBUyIyaWUKa4vZWwEVWUGZtOyYbgLb+8SloXVc+teneXlfp1HkcRTuAUzsGDGtThFhrQBAZDeIZXeHOE8+K8Ox+L1oKTzxzDHzifP2CQjdg=</latexit>
xi
<latexit sha1_base64="96Omxy49SPGED3hOl6/e89unjWM=">AAAB6nicbVBNS8NAEJ3Ur1q/qh69LBbBU0lEqMeiF48V7Qe0oWy2k3bpZhN2N2IJ/QlePCji1V/kzX/jts1BWx8MPN6bYWZekAiujet+O4W19Y3NreJ2aWd3b/+gfHjU0nGqGDZZLGLVCahGwSU2DTcCO4lCGgUC28H4Zua3H1FpHssHM0nQj+hQ8pAzaqx0/9Tn/XLFrbpzkFXi5aQCORr98ldvELM0QmmYoFp3PTcxfkaV4UzgtNRLNSaUjekQu5ZKGqH2s/mpU3JmlQEJY2VLGjJXf09kNNJ6EgW2M6JmpJe9mfif101NeOVnXCapQckWi8JUEBOT2d9kwBUyIyaWUKa4vZWwEVWUGZtOyYbgLb+8SloXVc+teneXlfp1HkcRTuAUzsGDGtThFhrQBAZDeIZXeHOE8+K8Ox+L1oKTzxzDHzifP2CQjdg=</latexit>

ness. The loss of details eventually leads to degraded


outputs. In our experiments, a drop of 0.17 dB is ob- S Fb
served when adopting image alignment. This observa-
{xi 1 , xi , xi+1 } i 2 Ikey
tion confirms the necessity of shifting the spatial align- <latexit sha1_base64="zqgHW+9Ej46HIdh1QNXPHOHUjUA=">AAACAnicbVDLSgMxFM3UV62vUVfiJlgEQS0TEXRZdOOygn1AZxgyaaYNZjJDkpGWYXDjr7hxoYhbv8Kdf2Om7UJbD9zL4Zx7Se4JEs6Udpxvq7SwuLS8Ul6trK1vbG7Z2zstFaeS0CaJeSw7AVaUM0GbmmlOO4mkOAo4bQf314XffqBSsVjc6VFCvQj3BQsZwdpIvr3nZkM/Y6coP4FDnxUtY8cod3Pfrjo1Zww4T9CUVMEUDd/+cnsxSSMqNOFYqS5yEu1lWGpGOM0rbqpogsk97tOuoQJHVHnZ+IQcHhqlB8NYmhIajtXfGxmOlBpFgZmMsB6oWa8Q//O6qQ4vvYyJJNVUkMlDYcqhjmGRB+wxSYnmI0Mwkcz8FZIBlphok1rFhIBmT54nrbMacmro9rxav5rGUQb74AAcAQQuQB3cgAZoAgIewTN4BW/Wk/VivVsfk9GSNd3ZBX9gff4AuRWWVw==</latexit>

<latexit sha1_base64="KSy8L+a7XDqlY6cs3brldgj5amg=">AAAB83icbVBNS8NAEJ3Ur1q/qh69LBbBU0lE0GPRi94q2A9oQtlsp+3SzSbsboQQ+je8eFDEq3/Gm//GbZuDtj4YeLw3w8y8MBFcG9f9dkpr6xubW+Xtys7u3v5B9fCoreNUMWyxWMSqG1KNgktsGW4EdhOFNAoFdsLJ7czvPKHSPJaPJkswiOhI8iFn1FjJ5z6X5L6fTzCb9qs1t+7OQVaJV5AaFGj2q1/+IGZphNIwQbXueW5igpwqw5nAacVPNSaUTegIe5ZKGqEO8vnNU3JmlQEZxsqWNGSu/p7IaaR1FoW2M6JmrJe9mfif10vN8DrIuUxSg5ItFg1TQUxMZgGQAVfIjMgsoUxxeythY6ooMzamig3BW355lbQv6p5b9x4ua42bIo4ynMApnIMHV9CAO2hCCxgk8Ayv8Oakzovz7nwsWktOMXMMf+B8/gDpPJGX</latexit>

W
ment to the feature level. Ff
i2
/ Ikey
• Feature Alignment. The inferior performance of re- E C <latexit sha1_base64="Kh3h3vHpbdv0q3qzdNQOfAAAwdk=">AAAB+HicbVBNS8NAEJ3Ur1o/GvXoZbEInkoigh6LXvRWwX5AG8Jmu22XbnbD7kaIob/EiwdFvPpTvPlv3LY5aOuDgcd7M8zMixLOtPG8b6e0tr6xuVXeruzs7u1X3YPDtpapIrRFJJeqG2FNORO0ZZjhtJsoiuOI0040uZn5nUeqNJPiwWQJDWI8EmzICDZWCt0q6wtpmEB3YT6h2TR0a17dmwOtEr8gNSjQDN2v/kCSNKbCEI617vleYoIcK8MIp9NKP9U0wWSCR7RnqcAx1UE+P3yKTq0yQEOpbAmD5urviRzHWmdxZDtjbMZ62ZuJ/3m91AyvgpyJJDVUkMWiYcqRkWiWAhowRYnhmSWYKGZvRWSMFSbGZlWxIfjLL6+S9nnd9+r+/UWtcV3EUYZjOIEz8OESGnALTWgBgRSe4RXenCfnxXl3PhatJaeYOYI/cD5/AN+akzc=</latexit>

moving/image alignment motivates us to resort to fea- U


ture alignment. Similar to flow-based methods [17, 25, R
33], BasicVSR adopts optical flow for spatial align-
ment. But instead of warping the images as in previous yi
<latexit sha1_base64="fYHWTPnicP3K/a8xK5ZBBgxUcng=">AAAB6nicbVBNS8NAEJ3Ur1q/qh69LBbBU0lE0GPRi8eK9gPaUDbbTbt0swm7EyGE/gQvHhTx6i/y5r9x2+agrQ8GHu/NMDMvSKQw6LrfTmltfWNzq7xd2dnd2z+oHh61TZxqxlsslrHuBtRwKRRvoUDJu4nmNAok7wST25nfeeLaiFg9YpZwP6IjJULBKFrpIRuIQbXm1t05yCrxClKDAs1B9as/jFkacYVMUmN6npugn1ONgkk+rfRTwxPKJnTEe5YqGnHj5/NTp+TMKkMSxtqWQjJXf0/kNDImiwLbGVEcm2VvJv7n9VIMr/1cqCRFrthiUZhKgjGZ/U2GQnOGMrOEMi3srYSNqaYMbToVG4K3/PIqaV/UPbfu3V/WGjdFHGU4gVM4Bw+uoAF30IQWMBjBM7zCmyOdF+fd+Vi0lpxi5hj+wPn8AWIWjdk=</latexit>

works, we perform warping on the features for better (a) Information-Refill (b) Coupled Propagation
performance. The aligned features are then passed to
multiple residual blocks for refinement. Formally, we Figure 5. (a) An additional feature extractor is used for feature re-
have finement, alleviating the error accumulation during propagation.
Ikey denotes the set of indices of the selected keyframes. E and
{b,f }
si = S(xi , xi±1 ), C denote the feature extractor and convolution, respectively. (b)
{b,f } {b,f } {b,f } The inter-connected propagation branches facilitate the informa-
h̄i = W (hi±1 , si ), (2) tion exchange by passing the outputs of the backward branches to
{b,f } {b,f } the forward branches. The proposed components are colored in
hi = R{b,f } (xi , h̄i ),
purple.
and F{b,f } = R{b,f } ◦ W ◦ S with a slight abuse of
notations. Here S and W denote the flow estimation Information-Refill. Inaccurate alignment in occluded re-
and spatial warping modules, respectively, and R{b,f } gions and on image boundaries is a prominent challenge that
denotes a stack of residual blocks. can lead to error accumulation, especially if we adopt long-
term propagation in our framework. To alleviate undesir-
Aggregation and Upsampling. BasicVSR adopts ba-
able effects brought by such erroneous features, we propose
sic components for aggregation and upsampling. Specif-
{b,f } an information-refill mechanism for feature refinement.
ically, given the intermediate features hi , an upsam-
As shown in Fig. 5(a), an additional feature extractor is
pling module composed of multiple convolutions and pixel-
used to extract deep features from a subset of input frames
shuffle [27] is used to generate the output HR images:
(keyframes) and their respective neighbors. The extracted
yi = U (hfi , hbi ), (3) features are then fused with the aligned features h̄i (Eq. 2)
by a convolution:
where U denotes the upsampling module.
ei = E(xi−1 , xi , xi+1 ),
Summary of BasicVSR. The analysis above motivates   
{b,f }
the design choice of BasicVSR. For propagation, Ba- 
C e i , h̄i if i ∈ Ikey ,
{b,f }
 (4)
sicVSR has chosen bidirectional propagation with empha- ĥi =
sis on long-term and global propagation. For alignment, h̄ {b,f }


BasicVSR adopts a simple flow-based alignment but tak- i otherwise,
ing place at feature level. For aggregation and upsam-
where E and C correspond to the feature extractor and con-
pling, popular choices on feature concatenation and pixel-
volution, respectively. Ikey denotes the set of indices of the
shuffle suffice. Despite being a simple and succinct method,
selected keyframes. The refined features are then passed to
BasicVSR achieves great performance in both restoration
the residual blocks for further refinement:
quality and efficiency. BasicVSR is also highly versatile as
it can readily accommodate additional components to han- {b,f } {b,f }
hi = R{b,f } (xi , ĥi ). (5)
dle more challenging scenarios, as we show next.
It is noteworthy that the feature extractor and feature fusion
3.2. From BasicVSR to IconVSR
are applied to the sparsely-selected keyframes only. Hence,
Using BasicVSR as a backbone, we introduce two novel the computational burden brought by the information-refill
components – Information-refill mechanism and coupled mechanism is insignificant.
propagation (IconVSR), to mitigate error accumulation While information-refill inherits the idea of keyframes,
during propagation and to facilitate information aggrega- we remark here that unlike existing interval-based meth-
tion. ods [15, 35] that isolate the intervals for independent
processing, the intervals (separated by the keyframes) in 4.1. Comparisons with State-of-the-Art Methods
IconVSR are connected to maintain a global information
We conduct comprehensive experiments by comparing
propagation.
BasicVSR and IconVSR with 14 models: VESPCN [1],
Coupled Propagation. In bidirectional settings, features SPMC [29], TOFlow [33], FRVSR [25], DUF [16],
are typically propagated in two opposite directions inde- RBPN [9], EDVR-M [32], EDVR [32], MuCAN [20],
pendently. In this design, the features in each propagation PFNL [34], RLSP [8], TGA [13], RSDN [12], and
branch are computed based on partial information, from ei- RRN [14]. The quantitative results are summarized in Ta-
ther previous frames or future frames. To exploit the in- ble 2 and the speed and performance comparison is pro-
formation in the sequences, we propose a coupled propa- vided in Fig. 1. Note that the parameters of BasicVSR and
gation scheme, where the propagation modules are inter- IconVSR are inclusive of that in the optical flow network,
connected. As depicted in Fig. 5(b), in coupled propagation, SPyNet. So the comparison is fair.
the features propagated backward hbi are taken as inputs in
the forward propagation module (c.f . Eq. 1, 3): BasicVSR. BasicVSR outperforms existing state of the
arts on various datasets, including REDS4, UDM10, and
hbi = Fb (xi , xi+1 , hbi+1 ), Vid4. BasicVSR also demonstrates high efficiency in ad-
dition to improvements in restoration quality. As shown
hfi = Ff (xi , xi−1 , hbi , hfi−1 ), (6) in Fig. 1, BasicVSR surpasses RSDN [12] by 0.61 dB
yi = U (hfi ). on UDM10 while having a similar number of parameters.
When compared with EDVR [32], which has a significantly
With coupled propagation, the forward propagation branch larger complexity, BasicVSR obtains a marked improve-
receives information from both past and future frames, lead- ment of 0.33 dB on REDS4 and competitive performances
ing to features of higher quality and hence better outputs. on Vimeo-90K-T and Vid4. We note that the performance
More importantly, since coupled propagation requires only of BasicVSR on Vimeo-90K-T is slightly lower than that
changes of the branch connections, the performance gain achieved by sliding-window methods such as EDVR [32]
can be obtained without introducing computational over- and TGA [13]. This is expected since Vimeo-90K-T con-
head. tains sequences with only seven frames, while the success
of BasicVSR partially comes from the aggregation of long-
4. Experiments term information (which is a realistic assumption).
Datasets and Settings We consider two widely-used IconVSR. IconVSR further improves the performance by
datasets for training: REDS [23] and Vimeo-90K [33]. For up to 0.31 dB over BasicVSR with slightly longer runtime.
REDS, following [32], we use the REDS4 dataset2 as our The performance gain is especially obvious in Vimeo-90K-
test set. We additionally define REDSval43 as our valida- T and REDS4, showing that our proposed coupled propa-
tion set. The remaining clips are used for training. We use gation and information-refill mechanisms are beneficial in
Vid4 [21], UDM10 [34], and Vimeo-90K-T [33] as test sets videos (1) lacking long-term information (Vimeo-90K-T)
along with Vimeo-90K. We test our models with 4× down- and (2) containing large and complicated motions (REDS4).
sampling using two degradations – Bicubic (BI) and Blur Overall, both BasicVSR and IconVSR are able to achieve
Downsampling (BD). remarkable performance while being faster than most state
We use pre-trained SPyNet [24] and EDVR-M4 [32] as of the arts.
our flow estimation module and feature extractor, respec- Qualitative comparisons are shown in Figures 11 and 12.
tively. We adopt Adam optimizer [18] and Cosine Anneal- BasicVSR and IconVSR are able to recover finer details and
ing scheme [22]. The initial learning rates of the feature ex- sharper edges. For instance, only BasicVSR and IconVSR
tractor and flow estimator are set to 1×10−4 and 2.5×10−5 , successfully recover clear square patterns in Fig. 11 and the
respectively. The learning rate for all other modules is set vertical strip patterns in Fig. 12. With the proposed compo-
to 2×10−4 . The total number of iterations is 300K, and the nents, IconVSR is able to reconstruct images with sharper
weights of the feature extractor and flow estimator are fixed edges. More examples are provided in the appendix.
during the first 5,000 iterations. The batch size is 8 and the
patch size of input LR frames is 64×64. We use Charbon- 5. Ablation Studies
nier loss [3] since it better handles outliers and improves the
performance over the conventional `2 loss [19]. Detailed
5.1. From BasicVSR to IconVSR
experimental settings are provided in the appendix. Information-Refill. We qualitatively visualize the features
2 Clips 000, 011, 015, 020 of REDS training set.
before and after information-refill to gain insights into the
3 Clips 000, 001, 006, 017 of REDS validation set. mechanism. As shown in Fig. 8(a), before information-
4 A lightweight version of EDVR. refill, the boundary pixels in the warped feature essentially
Table 2. Quantitative comparison (PSNR/SSIM). All results are calculated on Y-channel except REDS4 [23] (RGB-channel). Red and
blue colors indicate the best and the second-best performance, respectively. Blanked entries correspond to results unable to be reported.
The runtime is computed on an LR size of 180×320.

BI degradation BD degradation
Params (M) Runtime (ms) REDS4 [23] Vimeo-90K-T [33] Vid4 [21] UDM10 [34] Vimeo-90K-T [33] Vid4 [21]
Bicubic - - 26.14/0.7292 31.32/0.8684 23.78/0.6347 28.47/0.8253 31.30/0.8687 21.80/0.5246
VESPCN [1] - - - - 25.35/0.7557 - - -
SPMC [29] - - - - 25.88/0.7752 - - -
TOFlow [33] - - 27.98/0.7990 33.08/0.9054 25.89/0.7651 36.26/0.9438 34.62/0.9212 -
FRVSR [25] 5.1 137 - - - 37.09/0.9522 35.64/0.9319 26.69/0.8103
DUF [16] 5.8 974 28.63/0.8251 - - 38.48/0.9605 36.87/0.9447 27.38/0.8329
RBPN [9] 12.2 1507 30.09/0.8590 37.07/0.9435 27.12/0.8180 38.66/0.9596 37.20/0.9458 -
EDVR-M [32] 3.3 118 30.53/0.8699 37.09/0.9446 27.10/0.8186 39.40/0.9663 37.33/0.9484 27.45/0.8406
EDVR [32] 20.6 378 31.09/0.8800 37.61/0.9489 27.35/0.8264 39.89/0.9686 37.81/0.9523 27.85/0.8503
PFNL [34] 3.0 295 29.63/0.8502 36.14/0.9363 26.73/0.8029 38.74/0.9627 - 27.16/0.8355
MuCAN [20] - - 30.88/0.8750 37.32/0.9465 - - - -
TGA [13] 5.8 - - - - - 37.59/0.9516 27.63/0.8423
RLSP [8] 4.2 49 - - - 38.48/0.9606 36.49/0.9403 27.48/0.8388
RSDN [12] 6.2 94 - - - 39.35/0.9653 37.23/0.9471 27.92/0.8505
RRN [14] 3.4 45 - - - 38.96/0.9644 - 27.69/0.8488
BasicVSR (ours) 6.3 63 31.42/0.8909 37.18/0.9450 27.24/0.8251 39.96/0.9694 37.53/0.9498 27.96/0.8553
IconVSR (ours) 8.7 70 31.67/0.8948 37.47/0.9476 27.39/0.8279 40.03/0.9694 37.84/0.9524 28.04/0.8570

Frame 075, Clip 020 Bicubic/25.52 dB PFNL/28.59 dB RBPN/28.97 dB EDVR-M/29.79 dB

EDVR/30.21 dB BasicVSR/30.68 dB IconVSR/30.80 dB GT/PSNR


Figure 6. Qualitative comparison on REDS4 [23]. BasicVSR and IconVSR restores clearer square patterns. IconVSR restores sharper
edges. (Zoom-in for best view)

Sequence 0238, Clip 002 Bicubic/21.08 dB PFNL/23.31 dB RBPN/27.18 dB EDVR-M/25.29 dB

EDVR/26.58 dB BasicVSR/27.23 dB IconVSR/27.75 dB GT/PSNR


Figure 7. Qualitative comparison on Vimeo-90K-T [33]. Only BasicVSR and IconVSR are able to recover the vertical strip patterns.
IconVSR restores sharper edges. (Zoom-in for best view)

become zero due to non-existing correspondences. The lost lost information in regions where the features are poorly
information inevitably worsens the feature quality, leading aligned. The retrieved information can then be employed
to degraded outputs. With our information-refill mecha- for the subsequent feature refinement and propagation.
nism, the additional features can be used to “refill” the The above effect is especially obvious in regions with
Table 3. Evaluations of IconVSR components. The two compo-
nents bring an improvement of up to 0.28 dB over BasicVSR. The
PSNR is computed on REDS4/REDSval4.

BasicVSR IconVSR (w/o refill) IconVSR


Info-Refill 7 7 3
Coupled-Prop 7 3 3

coupled coupled
Without
Before

PSNR 31.42/30.17 31.60/30.38 31.67/30.45


refill

N = 21

With
After
refill

N = 11

(a) Information-Refill (b) Coupled Propagation


N =6
Figure 8. (a) Information lost during spatial warping can be com- N =4
pensated by the additional features. (b) With more effective use
N =3
of the backward-propagated features, coupled propagation leads
to clearer details and finer edges, especially in regions that are N =0 N =2
occluded in previous frames and regions that exist in the whole
sequence. (Zoom-in for best view)
Figure 10. Tradeoff in IconVSR. One can reduce the number
Without refill With refill
of keyframes for faster inference. The PSNR is positively corre-
lated with the number of keyframes N , indicating the effectiveness
of the information-refill mechanism. The PSNR is calculated on
REDSval4. The total number of frames in each clip is 100, and the
keyframes are evenly spaced.

Figure 9. Effect of Information-Refill. The contribution of is depicted in Fig. 10, where we see that the PSNR is pos-
information-refill is more obvious in regions with fine details, itively correlated with the number of keyframes, verifying
where alignment is error-prone. The information from the addi- the contributions of the information-refill mechanism. In an
tional feature extractor leads to marked improvements. extreme case when there is no keyframe, IconVSR degener-
ates to a recurrent network. Nevertheless, it still achieves a
fine details. In those regions, information from neighboring PSNR of 30.38 dB on REDSval4, which is 0.21 dB higher
frames cannot be effectively aggregated due to alignment than BasicVSR. This demonstrates the effectiveness of our
error, often resulting in inferior quality. With information- coupled propagation scheme, which can be used without in-
refill, the additional features assist in the restoration of the troducing additional computational overhead.
details, leading to in improved quality. For example, as
shown in Fig. 9, the license plate number can be recon-
6. Conclusion
structed more clearly with the refill mechanism. This work devotes attention to the search of generic and
Coupled Propagation. To ablate the coupled propagation efficient VSR baselines to ease the analysis and extension
scheme, we disable the information-refill mechanism and of VSR approaches. Through decomposing and analyzing
compare IconVSR with BasicVSR. In Fig. 8(b), the yellow existing elements, we propose BasicVSR, a simple yet ef-
box represents a region occluded in previous frames, and the fective network that outperforms existing state of the arts
forward propagation branch in BasicVSR could not receive with high efficiency. We build upon BasicVSR and propose
information of that region. The red box denotes a region IconVSR with two novel components to further improve the
that exists in all frames of the sequence, and hence abun- performance. BasicVSR and IconVSR can serve as strong
dant “snapshots” of the region can be found in latter frames. baselines for future works, and the discovery on the archi-
With coupled propagation, the backward-propagated fea- tecture designs could potentially be extended to other low-
tures are employed more effectively, and hence more details level vision tasks, such as video deblurring, denoising and
and finer edges can be reconstructed. The PSNR improve- colorization.
ment over BasicVSR is summarized in Table 3. Acknowledgement. This research was conducted in col-
laboration with SenseTime and supported by the Singapore
5.2. Tradeoff in IconVSR
Government through the Industry Alignment Fund - Indus-
Although IconVSR is trained with a fixed keyframe in- try Collaboration Projects Grant. It is also partially sup-
terval, one can reduce the number of keyframes for faster ported by Singapore MOE AcRF Tier 1 (2018-T1-002-056)
inference. The PSNR using different numbers of keyframes and NTU SUG.
References [18] Diederik Kingma and Jimmy Ba. Adam: A method for
stochastic optimization. In ICLR, 2015. 6, 10
[1] Jose Caballero, Christian Ledig, Aitken Andrew, Acosta Ale- [19] Wei-Sheng Lai, Jia-Bin Huang, Narendra Ahuja, and Ming-
jandro, Johannes Totz, Zehan Wang, and Wenzhe Shi. Real- Hsuan Yang. Deep laplacian pyramid networks for fast and
time video super-resolution with spatio-temporal networks accurate super-resolution. In CVPR, pages 5835–5843, 2017.
and motion compensation. In CVPR, 2017. 2, 6, 7 6, 10
[2] Kelvin CK Chan, Xintao Wang, Ke Yu, Chao Dong, and [20] Wenbo Li, Xin Tao, Taian Guo, Lu Qi, Jiangbo Lu, and Jiaya
Chen Change Loy. Understanding deformable alignment in Jia. MuCAN: Multi-correspondence aggregation network for
video super-resolution. arXiv preprint arXiv:2009.07265. 4 video super-resolution. In ECCV, 2020. 2, 6, 7
[3] Pierre Charbonnier, Laure Blanc-Feraud, Gilles Aubert, and [21] Ce Liu and Deqing Sun. On bayesian adaptive video super
Michel Barlaud. Two deterministic half-quadratic regular- resolution. TPAMI, 2014. 2, 6, 7, 10, 13
ization algorithms for computed imaging. In ICIP, 1994. 6, [22] Ilya Loshchilov and Frank Hutter. SGDR: Stochas-
10 tic gradient descent with warm restarts. arXiv preprint
[4] Kai Chen, Jiaqi Wang, Shuo Yang, Xingcheng Zhang, Yuan- arXiv:1608.03983, 2016. 6, 10
jun Xiong, Chen Change Loy, and Dahua Lin. Optimizing [23] Seungjun Nah, Sungyong Baik, Seokil Hong, Gyeongsik
video object detection via a scale-time lattice. In CVPR, Moon, Sanghyun Son, Radu Timofte, and Kyoung Mu Lee.
2018. 2 NTIRE 2019 challenge on video deblurring and super-
[5] Jifeng Dai, Haozhi Qi, Yuwen Xiong, Yi Li, Guodong resolution: Dataset and study. In CVPRW, 2019. 6, 7, 10,
Zhang, Han Hu, and Yichen Wei. Deformable convolutional 11
networks. In ICCV, 2017. 2 [24] Anurag Ranjan and Michael J Black. Optical flow estimation
[6] Chao Dong, Chen Change Loy, Kaiming He, and Xiaoou using a spatial pyramid network. In CVPR, 2017. 6, 10
Tang. Learning a deep convolutional network for image [25] Mehdi S M Sajjadi, Raviteja Vemulapalli, and Matthew
super-resolution. In ECCV, 2014. 2 Brown. Frame-recurrent video super-resolution. In CVPR,
[7] Chao Dong, Chen Change Loy, Kaiming He, and Xiaoou 2018. 2, 3, 5, 6, 7, 10
Tang. Image super-resolution using deep convolutional net- [26] Evan Shelhamer, Kate Rakelly, Judy Hoffman, and Trevor
works. TPAMI, 38(2):295–307, 2016. 2 Darrell. Clockwork convnets for video semantic segmenta-
[8] Dario Fuoli, Shuhang Gu, and Radu Timofte. Efficient video tion. In ECCV, 2016. 2
super-resolution through recurrent latent space propagation. [27] Wenzhe Shi, Jose Caballero, Ferenc Huszár, Johannes Totz,
In ICCVW, 2019. 3, 4, 6, 7 Andrew P Aitken, Rob Bishop, Daniel Rueckert, and Zehan
[9] Muhammad Haris, Greg Shakhnarovich, and Norimichi Wang. Real-time single image and video super-resolution
Ukita. Recurrent back-projection network for video super- using an efficient sub-pixel convolutional neural network. In
resolution. In CVPR, 2019. 1, 2, 3, 6, 7 CVPR, pages 1874–1883, 2016. 2, 5
[10] Yan Huang, Wei Wang, and Liang Wang. Bidirectional [28] Hiroyuki Takeda, Peyman Milanfar, Matan Protter, and
recurrent convolutional networks for multi-frame super- Michael Elad. Super-resolution without explicit subpixel
resolution. In NeurIPS, 2015. 2, 4 motion estimation. TIP, 18(9):1958–1975, 2009. 2
[11] Yan Huang, Wei Wang, and Liang Wang. Video super- [29] Xin Tao, Hongyun Gao, Renjie Liao, Jue Wang, and Jiaya
resolution via bidirectional recurrent convolutional net- Jia. Detail-revealing deep video super-resolution. In CVPR,
works. TPAMI, 2018. 2, 4 2017. 2, 6, 7
[12] Takashi Isobe, Xu Jia, Shuhang Gu, Songjiang Li, Shengjin [30] Yapeng Tian, Yulun Zhang, Yun Fu, and Chenliang Xu.
Wang, and Qi Tian. Video super-resolution with recurrent TDAN: Temporally deformable alignment network for video
structure-detail network. In ECCV, 2020. 2, 3, 4, 6, 7, 10 super-resolution. In CVPR, 2020. 2
[13] Takashi Isobe, Songjiang Li, Xu Jia, Shanxin Yuan, Gregory [31] Hua Wang, Dewei Su, Chuangchuang Liu, Longcun Jin, Xi-
Slabaugh, Chunjing Xu, Ya-Li Li, Shengjin Wang, and Qi anfang Sun, and Xinyi Peng. Deformable non-local network
Tian. Video super-resolution with temporal group attention. for video super-resolution. IEEE Access, 2019. 2
In CVPR, 2020. 2, 3, 6, 7 [32] Xintao Wang, Kelvin C.K. Chan, Ke Yu, Chao Dong, and
[14] Takashi Isobe, Fang Zhu, and Shengjin Wang. Revisiting Chen Change Loy. EDVR: Video restoration with enhanced
temporal modeling for video super-resolution. In BMVC, deformable convolutional networks. In CVPRW, 2019. 1, 2,
2020. 2, 3, 4, 6, 7 3, 6, 7, 10
[15] Samvit Jain, Xin Wang, and Joseph E Gonzalez. Accel: A [33] Tianfan Xue, Baian Chen, Jiajun Wu, Donglai Wei, and
corrective fusion network for efficient semantic segmenta- William T Freeman. Video enhancement with task-oriented
tion on video. In CVPR, 2019. 2, 5 flow. IJCV, 2019. 2, 4, 5, 6, 7, 10, 12
[16] Younghyun Jo, Seoung Wug Oh, Jaeyeon Kang, and Seon [34] Peng Yi, Zhongyuan Wang, Kui Jiang, Junjun Jiang, and
Joo Kim. Deep video super-resolution network using dy- Jiayi Ma. Progressive fusion video super-resolution net-
namic upsampling filters without explicit motion compensa- work via exploiting non-local spatio-temporal correlations.
tion. In CVPR, 2018. 2, 6, 7 In ICCV, 2019. 1, 2, 6, 7, 10, 13
[17] Tae Hyun Kim, Mehdi S M Sajjadi, Michael Hirsch, and [35] Zhengdong Zhang and Vivienne Sze. FAST: A framework
Bernhard Schölkopf. Spatio-temporal transformer network to accelerate super-resolution processing on compressed
for video restoration. In ECCV, 2018. 4, 5 videos. In CVPRW, 2017. 2, 5
[36] Xizhou Zhu, Jifeng Dai, Xingchi Zhu, Yichen Wei, and Lu input sequence to allow longer propagation. In other words,
Yuan. Towards high performance video object detection for we train with a sequence of 14 frames. During inference,
mobiles. arXiv preprint arXiv:1804.05830, 2018. 2 we take the whole video sequence as input.
[37] Xizhou Zhu, Han Hu, Stephen Lin, and Jifeng Dai. De- We adopt Adam optimizer [18] and Cosine Annealing
formable convnets v2: More deformable, better results. In scheme [22]. The initial learning rates of the feature extrac-
CVPR, 2019. 2 tor and flow estimator are set to 1×10−4 and 2.5×10−5 ,
[38] Xizhou Zhu, Yujie Wang, Jifeng Dai, Lu Yuan, and Yichen respectively. The learning rate for all other modules is set
Wei. Flow-guided feature aggregation for video object de- to 2×10−4 . The total number of iterations is 300K, and the
tection. In ICCV, 2017. 2
weights of the feature extractor and flow estimator are fixed
[39] Xizhou Zhu, Yuwen Xiong, Jifeng Dai, Lu Yuan, and Yichen
during the first 5,000 iterations. The batch size is 8 and the
Wei. Deep feature flow for video recognition. In CVPR,
2017. 2
patch size of input LR frames is 64×64.
Loss Function. We use Charbonnier loss [3] since it bet-
Appendix ter handles outliers and improves the performance over the
conventional `2 loss [19]:
A. Architecture and Experimental Settings
N
1 X
Architecture. In all our models, we adopt SPyNet [24] as L= ρ(yi − zi ), (7)
our flow estimator because of its simplicity and efficiency. N i=0
We use 30 residual blocks in each propagation branch. The √
feature channel is set to 64. In IconVSR, we adopt EDVR- where ρ(x) = x2 + 2 , =1×10−8 , zi denotes the
M5 [32] as the additional feature extractor since it main- ground-truth HR frame, and N denotes to the number of
tains a good balance between efficiency and quality. The pixels.
complexity of the components are summarized in Table 4. Degradations. We train and test our models with 4× down-
BasicVSR and IconVSR share the same flow estimator and sampling using two degradations – Bicubic (BI) and Blur
main network. The main network is a lightweight network, Downsampling (BD) [12, 25]. For BI, we use the MAT-
consisting of only 4.9M parameters. The flow estimator LAB function imresize for downsampling. For BD, we
blur the ground-truths by a Gaussian filter with σ=1.6, fol-
Table 4. Model complexity of BasicVSR and IconVSR.
lowed by a subsampling every four pixels.
BasicVSR IconVSR
Flow Estimator 1.4M 1.4M Implementation. We implement our models with PyTorch
Main Network 4.9M 4.9M and train the models using two NVIDIA Tesla V100 GPUs.
Feature Extractor - 2.4M Codes will be made publicly available.
Total 6.3M 8.7M
B. Qualitative Results
and feature extractor are fine-tuned together with the main B.1. Comparison with State of the Arts
network. In all our experiments, every five frames are se-
lected as keyframes. Note that the feature extractor is ap- In this section, we provide additional qualitative com-
plied to keyframes only. Therefore, the computational bur- parisons on REDS4 [23], Vimeo-90K [33], Vid4 [21], and
den brought by it is insignificant. UDM10 [34]. In Fig. 11 to Fig. 14, it is observed that
BasicVSR and IconVSR successfully produce outputs with
Datasets. We consider two widely-used datasets for train-
finer details and sharper edges. Furthermore, with the pro-
ing: REDS [23] and Vimeo-90K [33]. For REDS, follow-
posed information-refill and coupled propagation, IconVSR
ing [32], we use the REDS4 dataset6 as our test set. We
further improves the quality of the outputs.
additionally define REDSval47 as our validation set. The
remaining clips are used for training. We use Vid4 [21], B.2. BasicVSR vs IconVSR
UDM10 [34], and Vimeo-90K-T [33] as test sets along with
Vimeo-90K. In Fig. 15, we provide additional visual comparison of
BasicVSR and IconVSR to demonstrate the effectiveness
Experimental Settings. When training on REDS, we use a of our proposed components. We see that (1) information-
sequence of 15 frames as inputs, and loss is computed for refill improves the output quality on the fine regions, where
the 15 output images. When training on Vimeo-90K, we alignment is error-prone, and (2) coupled propagation leads
temporally augment the sequence by flipping the original to sharper edges by better employing the long-term infor-
5A lightweight version of EDVR.
mation in the sequence.
6 Clips 000, 011, 015, 020 of REDS training set.
7 Clips 000, 001, 006, 017 of REDS validation set.
Frame 074, Clip 000 Bicubic/24.42 dB PFNL/27.49 dB RBPN/27.51 dB EDVR-M/27.72 dB

EDVR/27.95 dB BasicVSR/28.29 dB IconVSR/28.42 dB GT/PSNR

Frame 070, Clip 020 Bicubic/25.87 dB PFNL/28.62 dB RBPN/28.83 dB EDVR-M/29.66 dB

EDVR/30.14 dB BasicVSR/30.60 dB IconVSR/30.75 dB GT/PSNR

Figure 11. Qualitative comparison on REDS [23].


Sequence 837, Clip 001 Bicubic/22.99 dB PFNL/26.52 dB RBPN/27.37 dB EDVR-M/27.34 dB

EDVR/27.71 dB BasicVSR/27.40 dB IconVSR/27.94 dB GT/PSNR

Sequence 943, Clip 010 Bicubic/35.65 dB PFNL/38.26 dB RBPN/40.22 dB EDVR-M/40.01 dB

EDVR/40.78 dB BasicVSR/41.20 dB IconVSR/41.41 dB GT/PSNR

Figure 12. Qualitative comparison on Vimeo-90K [33].


Frame 012, Clip Calendar Bicubic/18.83 dB PFNL/21.74 dB RBPN/22.11 dB EDVR-M/22.17 dB

EDVR/22.31 dB BasicVSR/22.27 dB IconVSR/22.42 dB GT/PSNR

Frame 018, Clip Foliage Bicubic/22.24 dB PFNL/24.48 dB RBPN/24.78 dB EDVR-M/24.74 dB

EDVR/24.93 dB BasicVSR/25.30 dB IconVSR/25.32 dB GT/PSNR

Figure 13. Qualitative comparison on Vid4 [21].

Frame 029, Clip Archpeople Bicubic/29.78 dB EDVR-M/36.14 dB EDVR/36.51 dB

BasicVSR/36.49 dB IconVSR/36.84 dB GT/PSNR

Figure 14. Qualitative comparison on UDM10 [34].


Bicubic w/o refill w/ refill GT Bicubic w/o coupled w/ coupled GT

Bicubic w/o refill w/ refill GT Bicubic w/o coupled w/ coupled GT

(a) Information-Refill (b) Coupled Propagation

Figure 15. Ablation of IconVSR. With information-refill and coupled propagation, IconVSR produces outputs with details and sharper
edges.

You might also like