Pyramid Vision Transformer Overview
Pyramid Vision Transformer Overview
without Convolutions
[Link]
!× !! ×
Conv 1 TF-E Transformer TF-E 1 Transformer
Block Block
Shrink
(a) CNNs: VGG [54], ResNet [22], etc. (b) Vision Transformer [13] (c) Pyramid Vision Transformer (ours)
Figure 1: Comparisons of different architectures, where “Conv” and “TF-E” stand for “convolution” and “Transformer
encoder”, respectively. (a) Many CNN backbones use a pyramid structure for dense prediction tasks such as object detection
(DET), instance and semantic segmentation (SEG). (b) The recently proposed Vision Transformer (ViT) [13] is a “columnar”
structure specifically designed for image classification (CLS). (c) By incorporating the pyramid structure from CNNs, we
present the Pyramid Vision Transformer (PVT), which can be used as a versatile backbone for many computer vision tasks,
broadening the scope and impact of ViT. Moreover, our experiments also show that PVT can easily be combined with
DETR [6] to build an end-to-end object detection system without convolutions.
1
44 shorter edge of 800 pixels in the COCO benchmark [40]).
PVT-L
PVT-M
42 To address the above limitations, this work proposes a
PVT-S pure Transformer backbone, termed Pyramid Vision Trans-
40 X101-64x4d
COCO BBox AP (%)
Patch Emb
Patch Emb
Patch Emb
Encoder
Encoder
Encoder
Encoder
Stage i
SRA
Multi-Head
Attention
Forward
Linear
Norm
Norm
Norm
Reshape Reshape
Feed
Reduction
Spacial
Position Embedding
𝐻!"#𝑊!"#
×𝐶!
𝑃!$
Element-wise Add Patch 𝐻!"# 𝑊!"#
×
𝐻!"# 𝑊!"# 𝑃! 𝑃!
× Embedding ×𝐶!
𝑃! 𝑃!
Feature Map ×(𝑃!$𝐶!"#) Transformer Encoder (𝐿% ×)
Figure 3: Overall architecture of Pyramid Vision Transformer (PVT). The entire model is divided into four stages, each
of which is comprised of a patch embedding layer and a Li -layer Transformer encoder. Following a pyramid structure, the
output resolution of the four stages progressively shrinks from high (4-stride) to low (32-stride).
backbone for dense prediction tasks, rather than a task- downstream tasks, including image classification, object de-
specific head or an image classification model. tection, and semantic segmentation.
Table 1: Detailed settings of PVT series. The design follows the two rules of ResNet [22]: (1) with the growth of network
depth, the hidden dimension gradually increases, and the output resolution progressively shrinks; (2) the major computation
resource is concentrated in Stage 3.
tively. Detailed hyper-parameter settings of the PVT series be an arbitrary shape, the position embeddings pre-trained
are provided in the supplementary material (SM). on ImageNet may no longer be meaningful. Therefore, we
For image classification, we follow ViT [13] and perform bilinear interpolation on the pre-trained position
DeiT [63] to append a learnable classification token to the embeddings according to the input resolution.
input of the last stage, and then employ a fully connected
(FC) layer to conduct classification on top of the token. 5. Experiments
We compare PVT with the two most representative CNN
4.2. Pixel-Level Dense Prediction
backbones, i.e., ResNet [22] and ResNeXt [73], which are
In addition to image-level prediction, dense prediction widely used in the benchmarks of many downstream tasks.
that requires pixel-level classification or regression to be
performed on the feature map, is also often seen in down- 5.1. Image Classification
stream tasks. Here, we discuss two typical tasks, namely Settings. Image classification experiments are performed
object detection, and semantic segmentation. on the ImageNet 2012 dataset [51], which comprises 1.28
We apply our PVT models to three representative dense million training images and 50K validation images from
prediction methods, namely RetinaNet [39], Mask R- 1,000 categories. For fair comparison, all models are
CNN [21], and Semantic FPN [32]. RetinaNet is a widely trained on the training set, and report the top-1 error on the
used single-stage detector, Mask R-CNN is the most pop- validation set. We follow DeiT [63] and apply random crop-
ular two-stage instance segmentation framework, and Se- ping, random horizontal flipping [59], label-smoothing reg-
mantic FPN is a vanilla semantic segmentation method ularization [60], mixup [78], CutMix [76], and random eras-
without special operations (e.g., dilated convolution). Us- ing [82] as data augmentations. During training, we employ
ing these methods as baselines enables us to adequately ex- AdamW [46] with a momentum of 0.9, a mini-batch size of
amine the effectiveness of different backbones. 128, and a weight decay of 5 × 10−2 to optimize models.
The implementation details are as follows: (1) Like The initial learning rate is set to 1 × 10−3 and decreases fol-
ResNet, we initialize the PVT backbone with the weights lowing the cosine schedule [45]. All models are trained for
pre-trained on ImageNet; (2) We use the output feature 300 epochs from scratch on 8 V100 GPUs. To benchmark,
pyramid {F1 , F2 , F3 , F4 } as the input of FPN [38], and we apply a center crop on the validation set, where a 224×
then the refined feature maps are fed to the follow-up de- 224 patch is cropped to evaluate the classification accuracy.
tection/segmentation head; (3) When training the detec- Results. In Table 2, we see that our PVT models are supe-
tion/segmentation model, none of the layers in PVT are rior to conventional CNN backbones under similar parame-
frozen; (4) Since the input for detection/segmentation can ter numbers and computational budgets. For example, when
Method #Param (M) GFLOPs Top-1 Err (%) ers. Our models are trained with a batch size of 16 on 8
ResNet18* [22] 11.7 1.8 30.2
ResNet18 [22] 11.7 1.8 31.5
V100 GPUs and optimized by AdamW [46] with an ini-
DeiT-Tiny/16 [63] 5.7 1.3 27.8 tial learning rate of 1 × 10−4 . Following common prac-
PVT-Tiny (ours) 13.2 1.9 24.9 tices [39, 21, 7], we adopt 1× or 3× training schedule (i.e.,
ResNet50* [22] 25.6 4.1 23.9 12 or 36 epochs) to train all detection models. The training
ResNet50 [22] 25.6 4.1 21.5
ResNeXt50-32x4d* [73] 25.0 4.3 22.4
image is resized to have a shorter side of 800 pixels, while
ResNeXt50-32x4d [73] 25.0 4.3 20.5 the longer side does not exceed 1,333 pixels. When using
T2T-ViTt -14 [75] 22.0 6.1 19.3 the 3× training schedule, we randomly resize the shorter
TNT-S [19] 23.8 5.2 18.7 side of the input image within the range of [640, 800]. In
DeiT-Small/16 [63] 22.1 4.6 20.1
PVT-Small (ours) 24.5 3.8 20.2 the testing phase, the shorter side of the input image is fixed
ResNet101* [22] 44.7 7.9 22.6 to 800 pixels.
ResNet101 [22] 44.7 7.9 20.2 Results. As shown in Table 3, when using RetinaNet for
ResNeXt101-32x4d* [73] 44.2 8.0 21.2
object detection, we find that under comparable number of
ResNeXt101-32x4d [73] 44.2 8.0 19.4
T2T-ViTt -19 [75] 39.0 9.8 18.6 parameters, the PVT-based models significantly surpasses
ViT-Small/16 [13] 48.8 9.9 19.2 their counterparts. For example, with the 1× training sched-
PVT-Medium (ours) 44.2 6.7 18.8 ule, the AP of PVT-Tiny is 4.9 points better than that of
ResNeXt101-64x4d* [73] 83.5 15.6 20.4
ResNet18 (36.7 vs. 31.8). Moreover, with the 3× training
ResNeXt101-64x4d [73] 83.5 15.6 18.5
ViT-Base/16 [13] 86.6 17.6 18.2 schedule and multi-scale training, PVT-Large archive the
T2T-ViTt -24 [75] 64.0 15.0 17.8 best AP of 43.4, surpassing ResNeXt101-64x4d (43.4 vs.
TNT-B [19] 66.0 14.1 17.2 41.8), while our parameter number is 30% fewer. These re-
DeiT-Base/16 [63] 86.6 17.6 18.2
PVT-Large (ours) 61.4 9.8 18.3
sults indicate that our PVT can be a good alternative to the
CNN backbone for object detection.
Table 2: Image classification performance on the Ima- Similar results are found in instance segmentation exper-
geNet validation set. “#Param” refers to the number of iments based on Mask R-CNN, as shown in Table 4. With
parameters. “GFLOPs” is calculated under the input scale the 1× training schedule, PVT-Tiny achieves 35.1 mask AP
of 224 × 224. “*” indicates the performance of the method (APm ), which is 3.9 points better than ResNet18 (35.1 vs.
trained under the strategy of its original paper. 31.2) and even 0.7 points higher than ResNet50 (35.1 vs.
34.4). The best APm obtained by PVT-Large is 40.7, which
is 1.0 points higher than ResNeXt101-64x4d (40.7 vs. 39.7),
the GFLOPs are roughly similar, the top-1 error of PVT- with 20% fewer parameters.
Small reaches 20.2, which is 1.3 points higher than that of
ResNet50 [22] (20.2 vs. 21.5). Meanwhile, under similar or 5.3. Semantic Segmentation
lower complexity, PVT models archive performances com- Settings. We choose ADE20K [83], a challenging scene
parable to the recently proposed Transformer-based mod- parsing dataset, to benchmark the performance of semantic
els, such as ViT [13] and DeiT [63] (PVT-Large: 18.3 vs. segmentation. ADE20K contains 150 fine-grained semantic
ViT(DeiT)-Base/16: 18.3). Here, we clarify that these re- categories, with 20,210, 2,000, and 3,352 images for train-
sults are within our expectations, because the pyramid struc- ing, validation, and testing, respectively. We evaluate our
ture is beneficial to dense prediction tasks, but brings little PVT backbones on the basis of Semantic FPN [32], a sim-
improvements to image classification. ple segmentation method without dilated convolutions [74].
Note that ViT and DeiT have limitations as they are In the training phase, the backbone is initialized with the
specifically designed for classification tasks, and thus are weights pre-trained on ImageNet [12], and other newly
not suitable for dense prediction tasks, which usually re- added layers are initialized with Xavier [18]. We optimize
quire effective feature pyramids. our models using AdamW [46] with an initial learning rate
of 1e-4. Following common practices [32, 8], we train our
5.2. Object Detection
models for 80k iterations with a batch size of 16 on 4 V100
Settings. Object detection experiments are conducted on GPUs. The learning rate is decayed following the polyno-
the challenging COCO benchmark [40]. All models are mial decay schedule with a power of 0.9. We randomly
trained on COCO train2017 (118k images) and evalu- resize and crop the image to 512 × 512 for training, and
ated on val2017 (5k images). We verify the effectiveness rescale to have a shorter side of 512 pixels during testing.
of PVT backbones on top of two standard detectors, namely Results. As shown in Table 5, when using Seman-
RetinaNet [39] and Mask R-CNN [21]. Before training, we tic FPN [32] for semantic segmentation, PVT-based
use the weights pre-trained on ImageNet to initialize the models consistently outperforms the models based on
backbone and Xavier [18] to initialize the newly added lay- ResNet [22] or ResNeXt [73]. For example, with al-
#Param RetinaNet 1x RetinaNet 3x + MS
Backbone
(M) AP AP50 AP75 APS APM APL AP AP50 AP75 APS APM APL
ResNet18 [22] 21.3 31.8 49.6 33.6 16.3 34.3 43.2 35.4 53.9 37.6 19.5 38.2 46.8
PVT-Tiny (ours) 23.0 36.7(+4.9) 56.9 38.9 22.6 38.8 50.0 39.4(+4.0) 59.8 42.0 25.5 42.0 52.1
ResNet50 [22] 37.7 36.3 55.3 38.6 19.3 40.0 48.8 39.0 58.4 41.8 22.4 42.8 51.6
PVT-Small (ours) 34.2 40.4(+4.1) 61.3 43.0 25.0 42.9 55.7 42.2(+3.2) 62.7 45.0 26.2 45.2 57.2
ResNet101 [22] 56.7 38.5 57.8 41.2 21.4 42.6 51.1 40.9 60.1 44.0 23.7 45.0 53.8
ResNeXt101-32x4d [73] 56.4 39.9(+1.4) 59.6 42.7 22.3 44.2 52.5 41.4(+0.5) 61.0 44.3 23.9 45.5 53.7
PVT-Medium (ours) 53.9 41.9(+3.4) 63.1 44.3 25.0 44.9 57.6 43.2(+2.3) 63.8 46.1 27.3 46.3 58.9
ResNeXt101-64x4d [73] 95.5 41.0 60.9 44.0 23.9 45.2 54.0 41.8 61.5 44.4 25.2 45.4 54.6
PVT-Large (ours) 71.1 42.6(+1.6) 63.7 45.4 25.8 46.0 58.4 43.4(+1.6) 63.6 46.1 26.1 46.0 59.5
Table 3: Object detection performance on COCO val2017. “MS” means that multi-scale training [39, 21] is used.
Table 4: Object detection and instance segmentation performance on COCO val2017. APb and APm denote bounding
box AP and mask AP, respectively.
.
pixels per patch) as input, leading to poor detection perfor- 2, 4, 6, and 8 of ViT-Small/32, and interpolate them to different scales.
#Param Mask R-CNN 1x Time RetinaNet 1x
Method GFLOPs Method Scale GFLOPs
(M) APm APm 50 APm75 (ms) AP AP50 AP75
ResNet50+GC r4 [5] 54.2 279.6 36.2 58.7 38.3 ResNet50 [22] 800 239.3 55.9 36.3 55.3 38.6
PVT-Small (ours) 44.1 304.4 37.8 60.1 40.3 640 157.2 51.7 38.7 59.3 40.8
PVT-Small (ours)
800 285.8 76.9 40.4 61.3 43.0
m
Table 10: PVT vs. CNN w/ non-local. AP denotes
mask AP. Under similar parameter nubmer and GFLOPs, Table 11: Latency and AP under different input scales.
our PVT outperform the CNN backbone w/ Non-Local “Scale” and “Time” denote the input scale and time cost
(ResNet50+GC r4) by 1.6 APm (37.8 vs. 36.2). per image. When the shorter side is 640 pixels, the PVT-
Small+RetinaNet has a lower GFLOPs and time cost (on a
V100 GPU) than ResNet50+RetinaNet, while obtaining 2.4
problem in our PVT. For fair comparisons, we multiply points better AP (38.7 vs. 36.3).
the hidden dimensions {C1 , C2 , C3 , C4 } of PVT-Small by
a scale factor 1.4 to make it have an equivalent parameter
equipped with non-local blocks (e.g., GCNet).
number to the deep model (i.e., PVT-Medium). As shown
(2) Regular convolutions can be deemed as special in-
in Table 9, the deep model (i.e., PVT-Medium) consistently
stantiations of spatial attention mechanisms [84]. In other
works better than the wide model (i.e., PVT-Small-Wide) on
words, the format of MHA is more flexible than the regular
both ImageNet and COCO. Therefore, going deeper is more
convolution. For example, for different inputs, the weights
effective than going wider in the design of PVT. Based on
of the convolution are fixed, but the attention weights of
this observation, in Table 1, we develop PVT models with
MHA change dynamically with the input. Thus, the features
different scales by increasing the model depth.
learned by the pure Transformer backbone full of MHA lay-
Pre-trained Weights. Most dense prediction models (e.g., ers, could be more flexible and expressive.
RetinaNet [39]) rely on the backbone whose weights are
Computation Overhead. With increasing input scale, the
pre-trained on ImageNet. We also discuss this problem in
growth rate of the GFLOPs of our PVT is greater than
our PVT. In the top of Figure 5, we plot the validation AP
ResNet [22], but lower than ViT [13], as shown in Figure
curves of RetinaNet-PVT-Small w/ (red curves) and w/o
6. However, when the input scale does not exceed 640×640
(blue curves) pre-trained weights. We find that the model
pixels, the GFLOPs of PVT-Small and ResNet50 are simi-
w/ pre-trained weights converges better than the one w/o
lar. This means that our PVT is more suitable for tasks with
pre-trained weights, and the gap between their final AP
medium-resolution input.
reaches 13.8 under the 1× training schedule and 8.4 under
On COCO, the shorter side of the input image is 800
the 3× training schedule and multi-scale training. There-
pixels. Under this condition, the inference speed of Reti-
fore, like CNN-based models, pre-training weights can also
naNet based on PVT-Small is slower than the ResNet50-
help PVT-based models converge faster and better. More-
based model, as reported in Table 11. (1) A direct solution
over, in the bottom of Figure 5, we also see that the con-
for this problem is to reduce the input scale. When reduc-
vergence speed of PVT-based models (red curves) is faster
ing the shorter side of the input image to 640 pixels, the
than that of ResNet-based models (green curves).
model based on PVT-Small runs faster than the ResNet50-
PVT vs. “CNN w/ Non-Local” To obtain a global recep- based model (51.7ms vs., 55.9ms), with 2.4 higher AP
tive field, some well-engineered CNN backbones, such as (38.7 vs. 36.3). 2) Another solution is to develop a self-
GCNet [5], integrate the non-local block in the CNN frame- attention layer with lower computational complexity. This
work. Here, we compare the performance of our PVT (pure is a worth exploring direction, we recently propose a solu-
Transformer) and GCNet (CNN w/ non-local), using Mask tion PVTv2 [67].
R-CNN for instance segmentation. As reported in Table Detection & Segmentation Results. In Figure 7, we also
10, we find that our PVT-Small outperforms ResNet50+GC present some qualitative object detection and instance seg-
r4 [5] by 1.6 points in APm (37.8 vs. 36.2), and 2.0 points in mentation results on COCO val2017 [40], and semantic
APm 75 (38.3 vs. 40.3), under comparable parameter number segmentation results on ADE20K [83]. These results indi-
and GFLOPs. There are two possible reasons for this result: cate that a pure Transformer backbone (i.e., PVT) without
(1) Although a single global attention layer (e.g., non- convolutions can also be easily plugged in dense prediction
local [70] or multi-head attention (MHA) [64]) can ac- models (e.g., RetinaNet [39], Mask R-CNN [21], and Se-
quire global-receptive-field features, the model perfor- mantic FPN [32]), and obtain high-quality results.
mance keeps improving as the model deepens. This indi-
cates that stacking multiple MHAs can further enhance the 6. Conclusions and Future Work
representation capabilities of features. Therefore, as a pure
Transformer backbone with more global attention layers, We introduce PVT, a pure Transformer backbone for
our PVT tends to perform better than the CNN backbone dense prediction tasks, such as object detection and seman-
300 References
ViT-Small/16
250 ViT-Small/32 [1] Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hin-
PVT-Small (ours) ton. Layer normalization. arXiv preprint arXiv:1607.06450,
200 2016. 5
ResNet50
GFLOPs
pe rson :0 .97
ze bra:1 .00
ze bra:1 .00
ze bra:1 .00
pe rson :0 .85
ca ke:0 .99
pe rson :0 .99
pe rson :0 .94
pe rson :1 .00
tr uck:0 .84
ca r:1 .00
pe rson :0 .94 pe rson :0 .99
be nch:1 .00
Figure 7: Qualitative results of object detection and instance segmentation on COCO val2017 [40], and semantic
segmentation on ADE20K [83]. The results (from left to right) are generated by PVT-Small-based RetinaNet [39], Mask
R-CNN [21], and Semantic FPN [32], respectively.
[15] Deng-Ping Fan, Ge-Peng Ji, Ming-Ming Cheng, and Ling arXiv:2103.00112, 2021. 7
Shao. Concealed object detection. IEEE Trans. Pattern Anal. [20] Song Han, Huizi Mao, and William J Dally. Deep com-
Mach. Intell., 2021. 11 pression: Compressing deep neural networks with pruning,
[16] Deng-Ping Fan, Ge-Peng Ji, Tao Zhou, Geng Chen, Huazhu trained quantization and huffman coding. arXiv preprint
Fu, Jianbing Shen, and Ling Shao. Pranet: Parallel reverse arXiv:1510.00149, 2015. 11
attention network for polyp segmentation. In International [21] Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Gir-
Conference on Medical Image Computing and Computer- shick. Mask r-cnn. In Proc. IEEE Int. Conf. Comp. Vis.,
Assisted Intervention, 2020. 11 2017. 1, 2, 3, 6, 7, 8, 10, 12
[17] Shanghua Gao, Ming-Ming Cheng, Kai Zhao, Xin-Yu [22] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun.
Zhang, Ming-Hsuan Yang, and Philip HS Torr. Res2net: A Deep residual learning for image recognition. In Proc. IEEE
new multi-scale backbone architecture. IEEE Trans. Pattern Conf. Comp. Vis. Patt. Recogn., 2016. 1, 2, 3, 4, 5, 6, 7, 8, 9,
Anal. Mach. Intell., 2019. 11 10, 11
[18] Xavier Glorot and Yoshua Bengio. Understanding the diffi- [23] Jie Hu, Li Shen, and Gang Sun. Squeeze-and-excitation net-
culty of training deep feedforward neural networks. In Proc. works. In Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2018.
Int. Conf. Artificial Intell. & Stat., 2010. 7 11
[19] Kai Han, An Xiao, Enhua Wu, Jianyuan Guo, Chunjing Xu, [24] Ronghang Hu and Amanpreet Singh. Transformer is all you
and Yunhe Wang. Transformer in transformer. arXiv preprint need: Multimodal multitask learning with a unified trans-
former. arXiv preprint arXiv:2102.10772, 2211. 2 [40] Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays,
[25] Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kil- Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence
ian Q Weinberger. Densely connected convolutional net- Zitnick. Microsoft coco: Common objects in context. In
works. In Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2017. Proc. Eur. Conf. Comp. Vis., 2014. 2, 7, 9, 10, 12
3 [41] Chenxi Liu, Liang-Chieh Chen, Florian Schroff, Hartwig
[26] Zilong Huang, Xinggang Wang, Lichao Huang, Chang Adam, Wei Hua, Alan L Yuille, and Li Fei-Fei. Auto-
Huang, Yunchao Wei, and Wenyu Liu. Ccnet: Criss-cross at- deeplab: Hierarchical neural architecture search for semantic
tention for semantic segmentation. In Proc. IEEE Int. Conf. image segmentation. In Proc. IEEE Conf. Comp. Vis. Patt.
Comp. Vis., 2019. 3 Recogn., 2019. 3
[27] Le Hui, Mingmei Cheng, Jin Xie, and Jian Yang. Efficient [42] Nian Liu, Ni Zhang, Kaiyuan Wan, Ling Shao, and Junwei
3D point cloud feature learning for large-scale place recog- Han. Visual saliency transformer. In Proc. IEEE Int. Conf.
nition. arXiv preprint arXiv:2101.02374, 2021. 11 Comp. Vis., 2021. 2
[28] Le Hui, Rui Xu, Jin Xie, Jianjun Qian, and Jian Yang. Pro- [43] Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian
gressive point cloud deconvolution generation network. In Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander C
Proc. Eur. Conf. Comp. Vis., 2020. 11 Berg. Ssd: Single shot multibox detector. In Proc. Eur. Conf.
Comp. Vis., 2016. 3
[29] Ge-Peng Ji, Yu-Cheng Chou, Deng-Ping Fan, Geng Chen,
Huazhu Fu, Debesh Jha, and Ling Shao. Progressively nor- [44] Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully
malized self-attention network for video polyp segmentation. convolutional networks for semantic segmentation. In Proc.
In International Conference on Medical Image Computing IEEE Conf. Comp. Vis. Patt. Recogn., 2015. 3
and Computer-Assisted Intervention, 2021. 11 [45] Ilya Loshchilov and Frank Hutter. SGDR: stochastic gradi-
ent descent with warm restarts. In Proc. Int. Conf. Learn.
[30] Xu Jia, Bert De Brabandere, Tinne Tuytelaars, and Luc V
Representations, 2017. 6
Gool. Dynamic filter networks. In Proc. Advances in Neural
Inf. Process. Syst., 2016. 3 [46] Ilya Loshchilov and Frank Hutter. Decoupled weight decay
regularization. In Proc. Int. Conf. Learn. Representations,
[31] Asifullah Khan, Anabia Sohail, Umme Zahoora, and
2019. 6, 7
Aqsa Saeed Qureshi. A survey of the recent architectures of
[47] Hyeonwoo Noh, Seunghoon Hong, and Bohyung Han.
deep convolutional neural networks. Artificial Intelligence
Learning deconvolution network for semantic segmentation.
Review, 53(8):5455–5516, 2020. 3
In Proc. IEEE Int. Conf. Comp. Vis., 2015. 3
[32] Alexander Kirillov, Ross Girshick, Kaiming He, and Piotr
[48] Niki Parmar, Prajit Ramachandran, Ashish Vaswani, Irwan
Dollár. Panoptic feature pyramid networks. In Proc. IEEE
Bello, Anselm Levskaya, and Jon Shlens. Stand-alone
Conf. Comp. Vis. Patt. Recogn., 2019. 1, 3, 6, 7, 10, 12
self-attention in vision models. In Hanna M. Wallach,
[33] Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Hugo Larochelle, Alina Beygelzimer, Florence d’Alché-
Imagenet classification with deep convolutional neural net- Buc, Emily B. Fox, and Roman Garnett, editors, Proc. Ad-
works. Proc. Advances in Neural Inf. Process. Syst., 2012. vances in Neural Inf. Process. Syst., 2019. 2, 3
3
[49] Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun.
[34] Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Faster r-cnn: Towards real-time object detection with region
Haffner. Gradient-based learning applied to document recog- proposal networks. In Proc. Advances in Neural Inf. Process.
nition. 1998. 3 Syst., 2015. 1, 3
[35] Xiang Li, Wenhai Wang, Xiaolin Hu, Jun Li, Jinhui Tang, [50] Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-
and Jian Yang. Generalized focal loss v2: Learning reliable net: Convolutional networks for biomedical image segmen-
localization quality estimation for dense object detection. In tation. In International Conference on Medical image com-
Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2021. 3 puting and computer-assisted intervention, 2015. 3
[36] Xiang Li, Wenhai Wang, Xiaolin Hu, and Jian Yang. Selec- [51] Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, San-
tive kernel networks. In Proc. IEEE Conf. Comp. Vis. Patt. jeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy,
Recogn., 2019. 3, 11 Aditya Khosla, Michael Bernstein, et al. Imagenet large
[37] Xiang Li, Wenhai Wang, Lijun Wu, Shuo Chen, Xiaolin Hu, scale visual recognition challenge. Int. J. Comput. Vision,
Jun Li, Jinhui Tang, and Jian Yang. Generalized focal loss: 2015. 3, 6
Learning qualified and distributed bounding boxes for dense [52] Suyash Shetty. Application of convolutional neural net-
object detection. In Proc. Advances in Neural Inf. Process. work for image classification on pascal voc challenge 2012
Syst., 2020. 3 dataset. arXiv preprint arXiv:1607.03785, 2016. 3
[38] Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, [53] Connor Shorten and Taghi M Khoshgoftaar. A survey on
Bharath Hariharan, and Serge Belongie. Feature pyramid image data augmentation for deep learning. Journal of Big
networks for object detection. In Proc. IEEE Conf. Comp. Data, 6(1):1–48, 2019. 3
Vis. Patt. Recogn., 2017. 3, 6 [54] Karen Simonyan and Andrew Zisserman. Very deep con-
[39] Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and volutional networks for large-scale image recognition. In
Piotr Dollár. Focal loss for dense object detection. In Proc. Yoshua Bengio and Yann LeCun, editors, Proc. Int. Conf.
IEEE Int. Conf. Comp. Vis., 2017. 1, 2, 3, 6, 7, 8, 10, 12 Learn. Representations, 2015. 1, 3, 4
[55] Peize Sun, Yi Jiang, Enze Xie, Zehuan Yuan, Changhu efficient and accurate end-to-end spotting of arbitrarily-
Wang, and Ping Luo. Onenet: Towards end-to-end one-stage shaped text. IEEE Trans. Pattern Anal. Mach. Intell., 2021.
object detection. arXiv preprint arXiv:2012.05780, 2020. 3 11
[56] Peize Sun, Yi Jiang, Rufeng Zhang, Enze Xie, Jinkun Cao, [70] Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaim-
Xinting Hu, Tao Kong, Zehuan Yuan, Changhu Wang, and ing He. Non-local neural networks. In Proc. IEEE Conf.
Ping Luo. Transtrack: Multiple-object tracking with trans- Comp. Vis. Patt. Recogn., 2018. 2, 3, 10
former. arXiv preprint arXiv:2012.15460, 2020. 2 [71] Enze Xie, Peize Sun, Xiaoge Song, Wenhai Wang, Xuebo
[57] Peize Sun, Rufeng Zhang, Yi Jiang, Tao Kong, Chenfeng Liu, Ding Liang, Chunhua Shen, and Ping Luo. Polarmask:
Xu, Wei Zhan, Masayoshi Tomizuka, Lei Li, Zehuan Yuan, Single shot instance segmentation with polar representation.
Changhu Wang, et al. Sparse r-cnn: End-to-end object de- In Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2020. 3
tection with learnable proposals. In Proc. IEEE Conf. Comp. [72] Enze Xie, Wenjia Wang, Wenhai Wang, Peize Sun, Hang Xu,
Vis. Patt. Recogn., 2021. 3 Ding Liang, and Ping Luo. Segmenting transparent object in
[58] Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and the wild with transformer. In Proc. Int. Joint Conf. Artificial
Alexander Alemi. Inception-v4, inception-resnet and the im- Intell., 2021. 2, 9
pact of residual connections on learning. In Proc. AAAI Conf. [73] Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and
Artificial Intell., 2017. 3 Kaiming He. Aggregated residual transformations for deep
[59] Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, neural networks. In Proc. IEEE Conf. Comp. Vis. Patt.
Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Recogn., 2017. 1, 2, 3, 6, 7, 8
Vanhoucke, and Andrew Rabinovich. Going deeper with
[74] Fisher Yu and Vladlen Koltun. Multi-scale context aggrega-
convolutions. In Proc. IEEE Conf. Comp. Vis. Patt. Recogn.,
tion by dilated convolutions. In Yoshua Bengio and Yann Le-
2015. 3, 6
Cun, editors, Proc. Int. Conf. Learn. Representations, 2016.
[60] Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon 7, 11
Shlens, and Zbigniew Wojna. Rethinking the inception ar-
[75] Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi,
chitecture for computer vision. In Proc. IEEE Conf. Comp.
Zihang Jiang, Francis EH Tay, Jiashi Feng, and Shuicheng
Vis. Patt. Recogn., 2016. 3, 6
Yan. Tokens-to-token vit: Training vision transformers
[61] Mingxing Tan and Quoc Le. Efficientnet: Rethinking model
from scratch on imagenet. arXiv preprint arXiv:2101.11986,
scaling for convolutional neural networks. In Proc. Int. Conf.
2021. 7
Mach. Learn., 2019. 11
[76] Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk
[62] Zhi Tian, Chunhua Shen, Hao Chen, and Tong He. Fcos:
Chun, Junsuk Choe, and Youngjoon Yoo. Cutmix: Regular-
Fully convolutional one-stage object detection. In Proc.
ization strategy to train strong classifiers with localizable fea-
IEEE Int. Conf. Comp. Vis., 2019. 3
tures. In Proc. IEEE Conf. Comp. Vis. Patt. Recogn., pages
[63] Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco 6023–6032, 2019. 6
Massa, Alexandre Sablayrolles, and Hervé Jégou. Training
[77] Erwan Zerhouni, Dávid Lányi, Matheus Viana, and Maria
data-efficient image transformers & distillation through at-
Gabrani. Wide residual networks for mitosis detection.
tention. In Proc. Int. Conf. Mach. Learn., 2021. 3, 6, 7
In IEEE International Symposium on Biomedical Imaging,
[64] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko-
2017. 9
reit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia
Polosukhin. Attention is all you need. In Proc. Advances in [78] Hongyi Zhang, Moustapha Cissé, Yann N. Dauphin, and
Neural Inf. Process. Syst., 2017. 2, 3, 4, 5, 10 David Lopez-Paz. mixup: Beyond empirical risk minimiza-
tion. In Proc. Int. Conf. Learn. Representations, 2018. 6
[65] Wenhai Wang, Xiang Li, Jian Yang, and Tong Lu. Mixed
link networks. Proc. Int. Joint Conf. Artificial Intell., 2018. [79] Hang Zhang, Chongruo Wu, Zhongyue Zhang, Yi Zhu, Zhi
3 Zhang, Haibin Lin, Yue Sun, Tong He, Jonas Mueller, R
[66] Wenhai Wang, Xuebo Liu, Xiaozhong Ji, Enze Xie, Ding Manmatha, et al. Resnest: Split-attention networks. arXiv
Liang, ZhiBo Yang, Tong Lu, Chunhua Shen, and Ping Luo. preprint arXiv:2004.08955, 2020. 11
Ae textspotter: Learning visual and linguistic representation [80] Hengshuang Zhao, Jiaya Jia, and Vladlen Koltun. Exploring
for ambiguous text spotting. In Proc. Eur. Conf. Comp. Vis., self-attention for image recognition. In Proc. IEEE Conf.
2020. 11 Comp. Vis. Patt. Recogn., 2020. 2
[67] Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao [81] Hengshuang Zhao, Jianping Shi, Xiaojuan Qi, Xiaogang
Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Wang, and Jiaya Jia. Pyramid scene parsing network. In
Pvtv2: Improved baselines with pyramid vision transformer. Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2017. 3
arXiv preprint arXiv:2106.13797, 2021. 10 [82] Zhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, and
[68] Wenhai Wang, Enze Xie, Xiang Li, Wenbo Hou, Tong Lu, Yi Yang. Random erasing data augmentation. In Proc. AAAI
Gang Yu, and Shuai Shao. Shape robust text detection with Conf. Artificial Intell., 2020. 6
progressive scale expansion network. In Proc. IEEE Conf. [83] Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela
Comp. Vis. Patt. Recogn., 2019. 11 Barriuso, and Antonio Torralba. Scene parsing through
[69] Wenhai Wang, Enze Xie, Xiang Li, Xuebo Liu, Ding Liang, ade20k dataset. In Proc. IEEE Conf. Comp. Vis. Patt.
Yang Zhibo, Tong Lu, and Chunhua Shen. Pan++: Towards Recogn., 2017. 2, 7, 9, 10, 12
[84] Xizhou Zhu, Dazhi Cheng, Zheng Zhang, Stephen Lin, and
Jifeng Dai. An empirical study of spatial attention mech-
anisms in deep networks. In Proc. IEEE Int. Conf. Comp.
Vis., 2019. 10
[85] Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang,
and Jifeng Dai. Deformable DETR: deformable transformers
for end-to-end object detection. In Proc. Int. Conf. Learn.
Representations, 2021. 2, 3