0% found this document useful (0 votes)
3 views6 pages

Convolutional Transformer Image Compression

This paper introduces SwinNPE, a novel transformer-based architecture for end-to-end image compression that integrates convolutional operations within the multi-head attention mechanism, eliminating the need for positional encoding. The proposed framework demonstrates superior performance in the rate-distortion trade-off compared to state-of-the-art CNN-based architectures while maintaining lower computational complexity. SwinNPE achieves comparable results to existing transformer-based methods with fewer parameters, highlighting the effectiveness of combining convolutions and transformers for image compression.

Uploaded by

manel zouaoui
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views6 pages

Convolutional Transformer Image Compression

This paper introduces SwinNPE, a novel transformer-based architecture for end-to-end image compression that integrates convolutional operations within the multi-head attention mechanism, eliminating the need for positional encoding. The proposed framework demonstrates superior performance in the rate-distortion trade-off compared to state-of-the-art CNN-based architectures while maintaining lower computational complexity. SwinNPE achieves comparable results to existing transformer-based methods with fewer parameters, highlighting the effectiveness of combining convolutions and transformers for image compression.

Uploaded by

manel zouaoui
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Convolutional Transformer-Based Image

Compression
Bouzid Arezki Fangchen Feng Anissa Mokraoui
L2TI Laboratory L2TI Laboratory L2TI Laboratory
University Sorbonne Paris Nord University Sorbonne Paris Nord University Sorbonne Paris Nord
Villetaneuse, France Villetaneuse, France Villetaneuse, France
[Link]@[Link] [Link]@[Link] [Link]@[Link]
arXiv:2409.04118v1 [[Link]] 6 Sep 2024

Abstract—In this paper, we present a novel transformer-based token sequence simultaneously. Therefore, positional encoding
architecture for end-to-end image compression. Our architecture plays a crucial role in preserving the sequence order and
incorporates blocks that effectively capture local dependencies various approaches have been proposed to better model the
between tokens, eliminating the need for positional encoding
by integrating convolutional operations within the multi-head positional information and maintain local context [11]–[13].
attention mechanism. We demonstrate through experiments that In the context of image compression, positional encoding
our proposed framework surpasses state-of-the-art CNN-based has demonstrated benefits in terms of Rate-Distortion (RD)
architectures in terms of the trade-off between bit-rate and performance, as shown in works such as [8], [10].
distortion and achieves comparable results to transformer-based Despite its advantages, employing positional encoding in
methods while maintaining lower computational complexity.
Index Terms—Image Compression,Rate-Distortion, Trans- transformers can increase the dimensionality of embeddings,
former, Transform Coding, Attention Mechanism. leading to higher computational costs during training and
limiting model flexibility. Recently, the authors of [14] demon-
I. I NTRODUCTION strated that positional encoding can be omitted in the attention
Transform coding is a widely employed method for com- module for image classification without any performance drop.
pressing images and serves as the foundation for several pop- They achieved this by introducing convolution in the tokeniza-
ular coding standards like JPEG. Codecs that utilize transform tion process of patches and the self-attention block to preserve
coding typically consist of three components for lossy com- local spatial information. This combination of convolution
pression: transform, quantization, and entropy coding. These and the attention mechanism leverages the advantages of both
components have seen advancements through the application convolutional neural networks and transformers.
of deep neural networks and end-to-end training, as evidenced In this paper, we introduce a novel image compression
by various studies [1]–[6]. framework called SwinNPE. It’s built upon our proposed
Among the early works, the authors of [1] introduced convolutional Swin block, which integrates patch convolution
a CNN-based two-level hierarchical variational autoencoder and shift window-based attention in Swin, eliminating the
with a hyper-prior serving as the entropy model. This archi- requirement for positional encoding.
tecture comprises two sets of encoders/decoders one for the We believe that this framework excels in capturing spatial
generative model and another for the hyper-prior model. contextual information more effectively. Preliminary experi-
Recently, transformers [7] have demonstrated remarkable ments demonstrate that SwinNPE achieves comparable results
success in computer vision, including neural image compres- to the SwinT architecture [10] while eliminating the need for
sion. The authors of [8] incorporated the attention mecha- positional encoding and utilizing fewer parameters. Some of
nism into the image compression framework by introducing the results of this paper have been presented at [15].
self-attention in the hyper-prior model. Additionally, More
sophisticated Swin block [9] in both the generative and II. P ROPOSED FRAMEWORK
hyper-prior models [10], utilizing shift window-based attention The proposed SwinNPE uses the same architecture as
to confine attention to local windows. Unlike convolutional in [10], which is shown in Figure 1. Specifically, the input
neural networks, transformers possess the ability to adapt image x is first encoded by the generative encoder y = ga (x),
their receptive field based on the task, with the attention and the hyper-latent z = ha (y) is obtained. The quantized
mechanism’s capacity to handle global context. This enhanced version of the hyper-latent ẑ is modeled and entropy-coded
understanding of global information enables the capture of with a learned factorized prior to passe through hs (ẑ) to
long-range dependencies in image compression applications. obtain µ and σ which are the parameters of a factorized
Positional encoding holds significant importance in trans- Gaussian distribution P (y|ẑ) = N (µ, diag(σ)) to model y.
formers. In the original ViT transformer [7], images are The quantized latent ŷ = Q(y−µ)+µ is finally entropy-coded
divided into non-overlapping patches, each mapped to a to- (Arithmetic encoding/decoding AE/AD) and sent to x̂ = ga (ŷ)
ken. The standard transformer layers then process the entire to reconstruct the image x̂. We use the classical strategy of
adding uniform noise to simulate the quantization operation Reconstruction Input Image
(Q) which makes the operation differentiable. The channel-
wise autoregressive block [2], [3] is designed to learn the
auto-regressive prior which factorizes the distribution of the Patch Split Patch Merge

latent as a product of conditional distributions incorporating


prediction from the causal context of the latents [4]–[6]. Convolutional window size Convolutional
Swin Block Swin Block
The generative and the hyper-prior encoder, ga and ha , are
built with the patch merge block and the convolutional Swin
Patch Split Patch Merge
block. The patch merge block contains the Depth-to-Space
operation [10] for down-sampling, a normalization layer, and Convolutional Convolutional
window size
a linear layer to project the input to a certain depth Ci . In Swin Block Swin Block

ga , the depth Ci of the latent representation increases as


Patch Split Patch Merge
the network gets deeper which allows for getting a more
abstract representation of the image. The size of the latent
Convolutional Convolutional
representation decreases accordingly. In each stage, we down- Swin Block
window size
Swin Block
sample the input feature by a factor of 2.
The convolutional Swin block proposed in this work is Patch Split Patch Merge
an extension of the Swin cell [9], as illustrated in Fig. 3.
Instead of employing position-wise linear projections, we Convolutional window size Convolutional
utilize convolutions to project the K, Q, and V matrices Swin Block Swin Block

within the multi-head attention block. Rather than relying on


hand-crafted positional encoding, we leverage the convolution
layer to capture positional information. More specifically, AD AE Q
following [14], we reshape the tokens into 2D dimensions and
apply 2D-convolution and flattening operation to get tokens
more sensitive to spatial context as illustrated in Fig. 2:
Channel-wise
AutoRegressive Model
K, Q, V = Flatten(Conv2d(Reshape2D(x))) (1)
To achieve parameter efficiency, we employ depth-wise sep-
Patch Split Patch Merge
arable convolution [14]. Specifically, the depth-wise separable
convolution performs a 2D convolution independently in each
Convolutional window size Convolutional
feature channel. The results are then concatenated and passed Swin Block Swin Block
through another convolution layer. This approach reduces
the number of parameters and computations while enhancing Patch Split Patch Merge
representational efficiency, as it operates not only on spatial
dimensions but also on the depth dimension. Importantly, it Convolutional window size Convolutional
Swin Block Swin Block
should be noted that the proposed block is not limited to
convolution operations; different forms of convolution are
possible [16], [17], making the convolutional Swin block
AD AE Q
highly adaptable. Unlike the convolutional attention block
mentioned in [14], we retain the shift window structure, which
facilitates cross-window connections.
Factorized
The generative and hyper-prior decoders, denoted as gs Model

and hs respectively, are constructed using the patch split


block and the convolutional Swin block. In the patch split Fig. 1. Network architecture of our proposed SwinNPE.
block, we reverse the merging sequence and Space-to-Depth
operation [10] for up-sampling.
III. E XPERIMENT AND A NALYSIS The SwinNPE’s performance was evaluated on the Kodak
and JPEG-AI test dataset [18], [19] and we center-cropped
A. Experiment configuration all images to multiples of 256 to avoid padding. We choose
This section presents an assessment of the SwinNPE ar- the following loss function to optimize the trade-off between
chitecture and a comparison of its image compression results the bit-rate R and the quality of reconstruction D which
against state-of-the-art approaches. The SwinNPE was trained corresponds to the Mean Squared Error (MSE) in RGB color
on the CLIC2020 training set for 3.3 million steps. During space:
training, each batch consisted of eight randomly cropped
images with a size of 256 × 256 pixels. L = D + βR, (2)
with β ∈ {0.003, 0.001, 0.0003, 0.0001}. presents a crop of one original image (K24) from the Kodak
The schedule learning rate starts at 10−4 and the hyper- dataset [18]. Figure 7 (b) compares the reconstructed cropped
parameters of the architecture shown in Fig. 1 are as fol- image using JPEG2000 with our method, showing that Swin-
lows: (d1 , d2 , d3 , d4 , d5 , d6 ) = (2, 2, 6, 2, 5, 1), (wg , wg ) = NPE preserves more details even if the bit-rate is smaller than
(8, 8), (wh , wh ) = (4, 4), and (C1 , C2 , C3 , C4 , C5 , C6 ) = those of JPEG2000.
(128, 192, 256, 320, 192, 192). For the autoregressive model,
IV. C ONCLUSION
we use the model proposed in [6] with 10 slices. The kernel
size in all convolutional Swin blocks for depth-wise separable This paper introduces SwinNPE, an image compression
convolution is set to 3. model based on transformers that utilizes convolutional Swin
blocks instead of positional encoding. SwinNPE achieves com-
B. Analysis parable performance to state-of-the-art methods while employ-
We compare our proposed SwinNPE model with the results ing fewer parameters and surpassing CNN-based architectures.
of two transformers-based architectures [8], [10] and some of The convolutional Swin block proposed in this work enables
the most used CNN-based image compression architectures enhanced utilization of spatial context without relying on
and standard codecs. The rate-distortion curves of different positional encoding, thereby providing increased flexibility
methods are shown in Figure 4 and Figure 5 on the Ko- and reducing the number of parameters.
dak dataset [18] and JPEG-AI test-set [19] respectively. In In future research, it would be interesting to explore the
these two figures, the PSNR and the rate shown are the utilization of diverse convolution operations and sizes within
average values across all images of the respective datasets. the SwinNPE model. This exploration could enable more
We summarize the number of parameters and GMACs of the precise modeling of complex spatial relationships and patterns,
tested transformer-based architectures in Table I where we ultimately leading to improved image compression perfor-
also illustrate the Bijønteguard metric [20] using the SwinT- mance.
CHARM as the reference for the Kodak dataset [18]. Furthermore, integrating the convolution operation into the
From Figure 4, we can clearly see that the SwinNPE patch merge/split module could leverage the advantages of
outperforms all of the tested CNN-based architectures in terms CNNs. The incorporation of convolutional Swin blocks in the
of the bit-rate/distortion tradeoff. It is particularly interesting SwinNPE model offers a promising avenue for developing
to notice that our proposed approach obtains almost the same efficient and effective transformer-based models for image
results as Entroformer [8] (orange dashed line in Figure 4) compression.
with much fewer model parameters (see Table I). Specifically,
the saving bit-rate of SwinNPE is 5.46% less than SwinT-
R EFERENCES
CHARM (optimal saving bit-rate) which is at the same level
as Entroformer with 4.33% more bit-rate saving compare to [1] Johannes Ballé, David Minnen, Saurabh Singh, Sung Jin Hwang, and
Nick Johnston. Variational image compression with a scale hyperprior.
SwinT-CHARM. We argue that it is due to the fact that 6th International Conference on Learning Representations (ICLR), 2018.
the convolutional layer in the proposed convolutional Swin [2] Mu Li, Wangmeng Zuo, Shuhang Gu, Debin Zhao, and David Zhang.
block can capture the local contextual information. From Learning convolutional networks for content weighted image compres-
sion. IEEE/CVF Conference on Computer Vision and Pattern Recogni-
Figure 4 and Figure 5, we can see that with fewer parameters, tion, pages 3214–3223, 2018.
the proposed SwinNPE has results comparable to SwinT- [3] Fabian Mentzer, Eirikur Agustsson, Michael Tschannen, Radu Timofte,
CHARM on both datasets. We emphasize that our proposed and Luc Van Gool. Conditional probability models for deep image
compression. IEEE/CVF Conference on Computer Vision and Pattern
architecture is particularly advantageous compared to SwinT- Recognition, pages 4394–4402, 2018.
based architecture without positional encoding 1 validating the [4] David Minnen, Johannes Ballé, and George D Toderici. Joint autoregres-
advantages of combining convolutions and transformers for sive and hierarchical priors for learned image compression. Advances
in Neural Information Processing Systems, 2018.
image compression. [5] Jooyoung Lee, Seunghyun Cho, and SeungKwon Beack. Context-
Figure 6 provides the Rate-Distortion (RD) curves on each adaptive entropy model for end-to- end optimized image compression.
image of the Kodak dataset [18]. Each curve presents the International Conference on Learning Representations, 2019.
[6] David Minnen and Saurabh Singh. Channel-wise autoregressive entropy
evaluation of an image with different versions of β value models for learned image compression. IEEE International Conference
(i.e. different SwinNPE models). As expected, experiments on Image Processing (ICIP), pages 3339–3343, 2020.
with the same β values reveal variability in PSNR and rate [7] Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weis-
senborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani,
values across different images. From the curves, a signif- Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit,
icant difference can be seen between images, highlighting and Neil Houlsby. An image is worth 16x16 words: Transformers
the substantial discrepancy in results when using the same for image recognition at scale. International Conference on Learning
Representations, 2021.
approach on different images. This observation highlights the [8] Yichen Qian, Ming Lin, Xiuyu Sun, Tan Zhiyu, and Rong Jin. Entro-
dependence of the results on the individual characteristics of former: A transformer-based entropy model for learned image compres-
the images in the dataset. sion. International Conference on Learning Representations (ICLR), 02
2022.
We inspect the visual quality in Figure 7 of the proposed [9] Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B.
approach with a standard codec as a reference. Figure 7 (a) Guo. Swin transformer: Hierarchical vision transformer using shifted
windows. IEEE/CVF International Conference on Computer Vision
1 The results are shown in the ablation studies in [10]. (ICCV), pages 9992–10002, Los Alamitos, CA, USA, oct 2021.
(a) (b)

MLP MLP

Norm Norm
....

....
Multi-Head Attention
Reshape Flatten Multi-Head Attention
with Shifted window
pad

....
Query Key Value Query Key Value

Convolutional Projection Convolutional Projection

....
Token Query (Q)

Key(K)

Value(V)

Fig. 2. (a) Convolutional Projection. (b) Convolutional Swin Block.

TABLE I
P ERFORMANCE COMPARISON USING B IJØNTEGUARD METRIC [20] FOR KODAK DATASET [18] WHERE ∆PSNR MEASURES THE AVERAGE PSNR
DIFFERENCE AND % ∆ RATE THE AVERAGE RATE SAVING IN PERCENTAGE BETWEEN S WIN T-CHARM [10] ( SELECTED AS THE REFERENCE NETWORK )
AND ANOTHER GIVEN NETWORK . -* GMAC S OF THE CORRESPONDING MODEL ARE NOT PROVIDED

Network #Param. (M) GMACs Positional Encoding Bijønteguard Metric


∆ PSNR %∆ rate
SwinT-CHARM [10] 32 223 Positional Relative Encoding 2D 0 0%
Entroformer [8] 142.7 -* Positional Relative Encoding 2D + Diamond -0.228 4.33%
SwinNPE (Ours) 27 178 - -0.311 5.46%

(a) (b)
Fig. 3. (a) The attention mechanism scheme for multi-head attention (b)
The attention mechanism scheme for convolutional Swin. DS conv means
depthwise separable convolution.

[10] Yinhao Zhu, Yang Yang, and Taco Cohen. Transformer-based transform
coding. International Conference on Learning Representations, 2022. Fig. 4. SwinNPE achieves nearly the same results as Entroformer [8] and
[11] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, L slion SwinT-CHARM [10] that relying on Positional encoding and better RD
Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. Attention performance than CNNs-based methods Factorized [21], Scale [1], Mean-
is all you need. I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Scale [4], Joint hyperprior [4] and standard codecs on the Kodak [18] image
Fergus, S. Vishwanathan, and R. Garnett, éditeurs, Advances in Neural set.
Information Processing Systems, volume 30. Curran Associates, Inc.,
2017.
[12] Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. Self attention with
relative position representations. Proceedings of the Conference. of Shen. Conditional positional encodings for vision transformers. The
the North American Chapter of the Association for Computational Eleventh Inter. Conf. on Learning Representations, 2023.
Linguistics : Human Language Technologies, Volume 2 (Short Papers), [14] Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyang Dai,
pages 464–468. Association for Computational Linguistics, Juin 2018. Lu Yuan, and Lei Zhang. Cvt: Introducing convolutions to vision
[13] Xiangxiang Chu, Zhi Tian, Bo Zhang, Xinlong Wang, and Chunhua transformers. Pro ceedings of the IEEE/CVF International Conference.
[8] Yichen Qian, Ming Lin, Xiuyu Sun, Tan Zhiyu, and Rong Jin, “Entro-
former: A transformer-based entropy model for learned image compres-
sion,” 02 2022, Inter. Conf. on Learning Representations (ICLR).
[9] Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and
B. Guo, “Swin transformer: Hierarchical vision transformer using shifted
windows,” in 2021 IEEE/CVF International Conference on Computer
Vision (ICCV), Los Alamitos, CA, USA, oct 2021, pp. 9992–10002,
IEEE Computer Society.
[10] Yinhao Zhu, Yang Yang, and Taco Cohen, “Transformer-based transform
coding,” in Inter. Conf. on Learning Representations, 2022.
[11] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion
Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin, “Attention
is all you need,” in Advances in Neural Information Processing
Systems, I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus,
S. Vishwanathan, and R. Garnett, Eds. 2017, vol. 30, Curran Associates,
Inc.
[12] Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani, “Self-attention with
relative position representations,” in Proceedings of the 2018 Confe.
of the North American Chapter of the Association for Computational
Linguistics: Human Language Technologies, Volume 2 (Short Papers).
June 2018, pp. 464–468, Association for Computational Linguistics.
[13] Xiangxiang Chu, Zhi Tian, Bo Zhang, Xinlong Wang, and Chunhua
Shen, “Conditional positional encodings for vision transformers,” in
Fig. 5. SwinNPE achieves nearly the same results as SwinT-CHARM [10] The Eleventh Inter. Conf. on Learning Representations, 2023.
and better RD performance than standard codecs on the JPEG-AI test-set [19]. [14] Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyang Dai,
Lu Yuan, and Lei Zhang, “Cvt: Introducing convolutions to vision
transformers,” in Proceedings of the IEEE/CVF Inter. Conf. on Computer
on Computer Vision (ICCV), pages 22–31, October 2021. Vision (ICCV), October 2021, pp. 22–31.
[15] Bouzid Arezki, Fangchen Feng, and Anissa Mokraoui, Transformer- [15] Fangchen Feng Bouzid Arezki and Anissa Mokraoui, “Transformer-
Based Image Compression Without Positional Encoding. Compression based image compression without positional encoding,” in 22ème
et Représentation des Signaux Audiovisuels, CORESA, 7-9 juin, 2023, édition de la conférence Compression et Représentation des Signaux
Lille, France. Audiovisuels (CORESA), 2023.
[16] Jifeng Dai, Haozhi Qi, Yuwen Xiong, Yi Li, Guodong Zhang, Han Hu, [16] Jifeng Dai, Haozhi Qi, Yuwen Xiong, Yi Li, Guodong Zhang, Han Hu,
and Yichen Wei. Deformable convolutional networks. Proceedings of and Yichen Wei, “Deformable convolutional networks,” in Proceedings
the IEEE International Conference. on computer vision, pages 764–773, of the IEEE inter. conf. on computer vision, 2017, pp. 764–773.
2017. [17] Lu Chi, Borui Jiang, and Yadong Mu, “Fast fourier convolution,”
[17] Lu Chi, Borui Jiang, and Yadong Mu. Fast Fourier convolution. Ad- Advances in Neural Information Processing Systems, vol. 33, pp. 4479–
vances in Neural Information Processing Systems, 33: 4479–4488, 2020. 4488, 2020.
[18] Kodak, “Kodak test images,” [Link] 1999.
[18] Kodak test images. [Link] graphics/kodak/, 1999.
[19] JPEG-AI, “Jpeg-ai test images,” [Link] images/,
[19] Johannes Ballé, Valero Laparra, and Eero P Simoncelli. End-to-end
2020.
optimized image compression. 5th International Conference on Learning
[20] Gisle Bjøntegaard, “Calculation of average psnr differences between
Representations (ICLR), 2017.
rd-curves,” 2001.
[20] JPEG-AI Test Images. [Link] images/, 2020.
[21] Johannes Ballé, Valero Laparra, and Eero P Simoncelli, “End-to-
[21] Gisle Bjøntegaard. Calculation of average PSNR differences between
end optimized image compression,” 5th Inter. Conf. on Learning
rd-curves. 2001.
Representations (ICLR), 2017.

R EFERENCES
[1] Johannes Ballé, David Minnen, Saurabh Singh, Sung Jin Hwang, and
Nick Johnston, “Variational image compression with a scale hyperprior,”
in 6th Inter. Conf. on Learning Representations (ICLR), 2018.
[2] Mu Li, Wangmeng Zuo, Shuhang Gu, Debin Zhao, and David Zhang,
“Learning convolutional networks for content-weighted image compres-
sion,” in 2018 IEEE/CVF Conference on Computer Vision and Pattern
Recognition, 2018, pp. 3214–3223.
[3] Fabian Mentzer, Eirikur Agustsson, Michael Tschannen, Radu Timofte,
and Luc Van Gool, “Conditional probability models for deep image
compression,” in 2018 IEEE/CVF Conf. on Computer Vision and Pattern
Recognition, 2018, pp. 4394–4402.
[4] David Minnen, Johannes Ballé, and George D Toderici, “Joint autore-
gressive and hierarchical priors for learned image compression,” in
Advances in Neural Information Processing Systems, 2018.
[5] Jooyoung Lee, Seunghyun Cho, and Seung-Kwon Beack, “Context-
adaptive entropy model for end-to-end optimized image compression,”
in International Conference on Learning Representations, 2019.
[6] David Minnen and Saurabh Singh, “Channel-wise autoregressive entropy
models for learned image compression,” in 2020 IEEE International
Conference on Image Processing (ICIP), 2020, pp. 3339–3343.
[7] Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weis-
senborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani,
Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and
Neil Houlsby, “An image is worth 16x16 words: Transformers for image
recognition at scale,” in Inter. Conf. on Learning Representations, 2021.
Fig. 6. Rate-Distortion (RD) performance of SwinNPE on each image of Kodak [18] dataset.

(a) (b)
Fig. 7. (a) One example (K24) of the original image from Kodak dataset [18]. (b) Left: image compressed with JPEG2000 (0.9257 bpp and PSNR: 30.6395
dB). Right: image compressed with the proposed SwinNPE for β = 0.001 (0.6987 bpp and PSNR: 32.9314 dB).

You might also like