Image Segmentation Using Deep Learning
Image Segmentation Using Deep Learning
fully edited. Content may change prior to final publication. Citation information: DOI 10.1109/TPAMI.2021.3059968, IEEE
Transactions on Pattern Analysis and Machine Intelligence
TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. ??, NO.??, ?? 2021 1
Abstract—Image segmentation is a key task in computer vision and image processing with important applications such as scene
understanding, medical image analysis, robotic perception, video surveillance, augmented reality, and image compression, among others,
and numerous segmentation algorithms are found in the literature. Against this backdrop, the broad success of Deep Learning (DL) has
prompted the development of new image segmentation approaches leveraging DL models. We provide a comprehensive review of this
recent literature, covering the spectrum of pioneering efforts in semantic and instance segmentation, including convolutional pixel-labeling
networks, encoder-decoder architectures, multiscale and pyramid-based approaches, recurrent networks, visual attention models, and
generative models in adversarial settings. We investigate the relationships, strengths, and challenges of these DL-based segmentation
models, examine the widely used datasets, compare performances, and discuss promising research directions.
Index Terms—Image segmentation, deep learning, convolutional neural networks, encoder-decoder models, recurrent models,
generative models, semantic segmentation, instance segmentation, panoptic segmentation, medical image segmentation.
1 I NTRODUCTION
0162-8828 (c) 2021 IEEE. Personal use is permitted, but republication/redistribution requires IEEE permission. See [Link] for more information.
Authorized licensed use limited to: Carleton University. Downloaded on May 28,2021 at 14:10:11 UTC from IEEE Xplore. Restrictions apply.
This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI 10.1109/TPAMI.2021.3059968, IEEE
Transactions on Pattern Analysis and Machine Intelligence
2 TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. ??, NO.??, ?? 2021
Output
0162-8828 (c) 2021 IEEE. Personal use is permitted, but republication/redistribution requires IEEE permission. See [Link] for more information.
Authorized licensed use limited to: Carleton University. Downloaded on May 28,2021 at 14:10:11 UTC from IEEE Xplore. Restrictions apply.
This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI 10.1109/TPAMI.2021.3059968, IEEE
Transactions on Pattern Analysis and Machine Intelligence
S. MINAEE et al.: IMAGE SEGMENTATION USING DEEP LEARNING: A SURVEY 3
Discriminator Network
Generator Fake
Network Images FCNs have been applied to a variety of segmentation
Predicted
problems, such as brain tumor segmentation [31], instance-
Real
Labels
aware semantic segmentation [32], skin lesion segmenta-
Images
tion [33], and iris segmentation [34]. While demonstrating
that DNNs can be trained to perform semantic segmentation
in an end-to-end manner on variable-sized images, the
Fig. 6. Architecture of a GAN. Courtesy of Ian Goodfellow.
conventional FCN model has some limitations—it is too
computationally expensive for real-time inference, it does
not account for global context information in an efficient
manner, and it is not easily generalizable to 3D images.
3 DL-BASED I MAGE S EGMENTATION M ODELS Several researchers have attempted to overcome some of the
limitations of the FCN. For example, Liu et al. [35] proposed
This section is a survey of numerous learning-based seg- ParseNet (Fig. 9), which adds global context to FCNs by using
mentation methods, grouped into 10 categories based on the average feature for a layer to augment the features at
their model architectures. Several architectural features are each location. The feature map for a layer is pooled over the
common among many of these methods, such as encoders whole image, resulting in a context vector. The context vector
and decoders, skip-connections, multiscale architectures, and is normalized and unpooled to produce new feature maps of
more recently the use of dilated convolutions. It is convenient the same size as the initial ones, which are then concatenated,
to group models based on their architectural contributions which amounts to an FCN whose convolutional layers are
over prior models. replaced by the described module (Fig. 9e).
0162-8828 (c) 2021 IEEE. Personal use is permitted, but republication/redistribution requires IEEE permission. See [Link] for more information.
Authorized licensed use limited to: Carleton University. Downloaded on May 28,2021 at 14:10:11 UTC from IEEE Xplore. Restrictions apply.
This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI 10.1109/TPAMI.2021.3059968, IEEE
Transactions on Pattern Analysis and Machine Intelligence
4 TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. ??, NO.??, ?? 2021
0162-8828 (c) 2021 IEEE. Personal use is permitted, but republication/redistribution requires IEEE permission. See [Link] for more information.
Authorized licensed use limited to: Carleton University. Downloaded on May 28,2021 at 14:10:11 UTC from IEEE Xplore. Restrictions apply.
This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI 10.1109/TPAMI.2021.3059968, IEEE
Transactions on Pattern Analysis and Machine Intelligence
S. MINAEE et al.: IMAGE SEGMENTATION USING DEEP LEARNING: A SURVEY 5
channel conv.
maps unit
strided
conv. upsample
Fig. 15. The V-Net model for 3D image segmentation. From [51].
0162-8828 (c) 2021 IEEE. Personal use is permitted, but republication/redistribution requires IEEE permission. See [Link] for more information.
Authorized licensed use limited to: Carleton University. Downloaded on May 28,2021 at 14:10:11 UTC from IEEE Xplore. Restrictions apply.
This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI 10.1109/TPAMI.2021.3059968, IEEE
Transactions on Pattern Analysis and Machine Intelligence
6 TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. ??, NO.??, ?? 2021
3.5 R-CNN Based Models Fig. 19. Mask R-CNN instance segmentation results. From [62].
The Regional CNN (R-CNN) and its extensions have proven
successful in object detection applications. In particular, the
Faster R-CNN [61] architecture (Fig. 17) uses a region pro- improving the propagation of lower-layer features. Each
posal network (RPN) that proposes bounding box candidates. stage of this third pathway takes as input the feature maps
The RPN extracts a Region of Interest (RoI), and an RoIPool of the previous stage and processes them with a 3 × 3
layer computes features from these proposals to infer the convolutional layer. A lateral connection adds the output
bounding box coordinates and class of the object. Some to the same-stage feature maps of the top-down pathway
extensions of R-CNN have been used to address the instance and these feed the next stage.
segmentation problem; i.e., the task of simultaneously per-
forming object detection and semantic segmentation.
Fig. 20. The Path Aggregation Network. (a) FPN backbone. (b) Bottom-
up path augmentation. (c) Adaptive feature pooling. (d) Box branch. (e)
Fully-connected fusion. From [63].
0162-8828 (c) 2021 IEEE. Personal use is permitted, but republication/redistribution requires IEEE permission. See [Link] for more information.
Authorized licensed use limited to: Carleton University. Downloaded on May 28,2021 at 14:10:11 UTC from IEEE Xplore. Restrictions apply.
This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI 10.1109/TPAMI.2021.3059968, IEEE
Transactions on Pattern Analysis and Machine Intelligence
S. MINAEE et al.: IMAGE SEGMENTATION USING DEEP LEARNING: A SURVEY 7
0162-8828 (c) 2021 IEEE. Personal use is permitted, but republication/redistribution requires IEEE permission. See [Link] for more information.
Authorized licensed use limited to: Carleton University. Downloaded on May 28,2021 at 14:10:11 UTC from IEEE Xplore. Restrictions apply.
This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI 10.1109/TPAMI.2021.3059968, IEEE
Transactions on Pattern Analysis and Machine Intelligence
8 TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. ??, NO.??, ?? 2021
Visin et al. [82] proposed an RNN-based model for propose an end-to-end trainable recurrent and convolutional
semantic segmentation called ReSeg (Fig. 25). This model model that jointly learns to process visual and linguistic
is mainly based on ReNet [83], which was developed for information (Fig. 28). This differs from traditional semantic
image classification. Each ReNet layer is composed of four segmentation over a predefined set of semantic classes; i.e.,
RNNs that sweep the image horizontally and vertically in the phrase “two men sitting on the right bench” requires
both directions, encoding patches/activations, and providing segmenting only the two people on the right bench and no
relevant global information. To perform image segmentation others sitting on another bench or standing. Fig. 29 shows
with the ReSeg model, ReNet layers are stacked atop pre- an example segmentation result by the model.
trained VGG-16 convolutional layers, which extract generic
local features, and are then followed by up-sampling layers to
recover the original image resolution in the final predictions.
Byeon et al. [84] performed per-pixel segmentation and
classification of images of natural scenes using 2D LSTM
networks, which learn textures and the complex spatial
dependencies of labels in a single model that carries out
classification, segmentation, and context integration.
Fig. 26. The graph-LSTM model for semantic segmentation. From [85].
Xiang and Fox [86] proposed Data Associated Recurrent Chen et al. [88] proposed an attention mechanism that
Neural Networks (DA-RNNs) for joint 3D scene mapping learns to softly weight multiscale features at each pixel
and semantic labeling. DA-RNNs use a new recurrent neural location. They adapt a powerful semantic segmentation
network architecture for semantic labeling on RGB-D videos. model and jointly train it with multiscale images and the
The output of the network is integrated with mapping attention model In Fig. 30, the model assigns large weights
techniques such as Kinect-Fusion in order to inject semantic to the person (green dashed circle) in the background for
information into the reconstructed 3D scene. features from scale 1.0 as well as on the large child (magenta
Hu et al. [87] developed a semantic segmentation al- dashed circle) for features from scale 0.5. The attention
gorithm that combines a CNN to encode the image and mechanism enables the model to assess the importance of
an LSTM to encode its linguistic description. To produce features at different positions and scales, and it outperforms
pixel-wise image segmentations from language inputs, they average and max pooling.
0162-8828 (c) 2021 IEEE. Personal use is permitted, but republication/redistribution requires IEEE permission. See [Link] for more information.
Authorized licensed use limited to: Carleton University. Downloaded on May 28,2021 at 14:10:11 UTC from IEEE Xplore. Restrictions apply.
This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI 10.1109/TPAMI.2021.3059968, IEEE
Transactions on Pattern Analysis and Machine Intelligence
S. MINAEE et al.: IMAGE SEGMENTATION USING DEEP LEARNING: A SURVEY 9
0162-8828 (c) 2021 IEEE. Personal use is permitted, but republication/redistribution requires IEEE permission. See [Link] for more information.
Authorized licensed use limited to: Carleton University. Downloaded on May 28,2021 at 14:10:11 UTC from IEEE Xplore. Restrictions apply.
This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI 10.1109/TPAMI.2021.3059968, IEEE
Transactions on Pattern Analysis and Machine Intelligence
10 TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. ??, NO.??, ?? 2021
Cheng et al. [114] proposed the Deep Active Ray Network 4 DATASETS
(DarNet), which is similar to DSAC, but with a different
In this section we survey the image datasets most commonly
explicit ACM formulation based on polar coordinates to
used to train and test DL image segmentation models,
prevent contour self-intersection.
grouping them into 3 categories—2D (pixel) images, 2.5D
A truly end-to-end backpropagation trainable, fully-
RGB-D (color+depth) images, and 3D (voxel) images—and
integrated FCN-ACM combination was recently introduced
provide details about the characteristics of each dataset.
by Hatamizadeh et al. [115], dubbed Trainable Deep Active
Data augmentation is often used to increase the number
Contours (TDAC). Going beyond [112], they implemented
of labeled samples, especially for small datasets such as
the locally-parameterized level-set ACM in the form of
those in the medical imaging domain, thus improving the
additional convolutional layers following the layers of the
performance of DL segmentation models. A set of trans-
backbone FCN, exploiting Tensorflow’s automatic differen-
formations is applied either in the data space, or feature
tiation mechanism to backpropagate training error gradi-
space, or both (i.e., both the image and the segmentation
ents throughout the entire DCAC framework. The fully-
map). Typical transformations include translation, reflection,
automated model requires no intervention either during
rotation, warping, scaling, color space shifting, cropping, and
training or segmentation, can naturally segment multiple
projections onto principal components. Data augmentation
instances of objects of interest, and deal with arbitrary object
can also benefit by yielding faster convergence, decreasing
shape including sharp corners.
the chance of over-fitting, and enhancing generalization. For
some small datasets, data augmentation has been shown to
3.11 Other Models boost model performance by more than 20%.
Other popular DL architectures for image segmentation
include the following:
4.1 2D Image Datasets
Context Encoding Network (EncNet) [116] uses a basic
feature extractor and feeds the feature maps into a context The bulk of image segmentation research has focused on 2D
encoding module. RefineNet [117] is a multipath refinement images; therefore, many 2D image segmentation datasets are
network that explicitly exploits all the information available available. The following are some of the most popular:
along the down-sampling process to enable high-resolution PASCAL Visual Object Classes (VOC) [150] is a highly
prediction using long-range residual connections. Seed- popular dataset in computer vision, with annotated images
net [118] introduced an automatic seed generation technique available for 5 tasks—classification, segmentation, detection,
with deep reinforcement learning that learns to solve the in- action recognition, and person layout. For the segmentation
teractive segmentation problem. Object-Contextual Represen- task, there are 21 labeled object classes and pixels are labeled
tations (OCR) [42] learns object regions and the relation be- as background if they do not belong to any of these classes.
tween each pixel and each object region, augmenting the rep- The dataset is divided into two sets, training and validation,
resentation pixels with the object-contextual representation. with 1,464 and 1,449 images, respectively, and a private test
Additional models and methods include BoxSup [119], Graph set for the actual challenge. Fig. 34 shows an example image
Convolutional Networks (GCN) [120], Wide ResNet [121], and its pixel-wise label.
Exfuse [122] (enhancing low-level and high-level features PASCAL Context [152] is an extension of the PASCAL
fusion), Feedforward-Net [123], saliency-aware models for VOC 2010 detection challenge. It includes pixel-wise labels
geodesic video segmentation [124], Dual Image Segmentation for all the training images. It contains more than 400 classes
(DIS) [125], FoveaNet [126] (perspective-aware scene pars- (including the original 20 classes plus backgrounds from
ing), Ladder DenseNet [127], Bilateral Segmentation Network PASCAL VOC segmentation), in three categories (objects,
(BiSeNet) [128], Semantic Prediction Guidance for Scene stuff, and hybrids). Many of the object categories of this
Parsing (SPGNet) [129], gated shape CNNs [130], Adaptive dataset are too sparse and; therefore, a subset of 59 classes is
Context Network (AC-Net) [131], Dynamic-Structured Se- usually selected for use.
mantic Propagation Network (DSSPN) [132], Symbolic Graph Microsoft Common Objects in Context (MS
Reasoning (SGR) [133], CascadeNet [134], Scale-Adaptive COCO) [153] is a large-scale object detection, segmentation,
Convolutions (SAC) [135], Unified Perceptual parsing Net- and captioning dataset. COCO includes images of complex
work (UperNet) [136], segmentation by re-training and everyday scenes, containing common objects in their natural
self-training [137], densely connected neural architecture contexts. This dataset contains photos of 91 object types,
search [138], hierarchical multiscale attention [139], Efficient with a total of 2.5 million labeled instances in 328K images.
RGB-D Semantic Segmentation (ESA-Net) [140], Iterative Fig. 35 compares MS-COCO labels with those of previous
Pyramid Contexts [141], and Learning Dynamic Routing for datasets for a sample image.
Semantic Segmentation [142]. Cityscapes [154] is a large database with a focus on
Panoptic segmentation [143] is growing in popularity. semantic understanding of urban street scenes. It contains
Efforts in this direction include Panoptic Feature Pyramid a diverse set of stereo video sequences recorded in street
Network (PFPN) [144], attention-guided network for panop- scenes from 50 cities, with high quality pixel-level annotation
tic segmentation [145], seamless scene segmentation [146], of 5K frames, in addition to a set of 20K weakly annotated
panoptic Deeplab [147], unified panoptic segmentation net- frames. It includes semantic and dense pixel annotations of
work [148], and efficient panoptic segmentation [149]. 30 classes, grouped into 8 categories—flat surfaces, humans,
Fig. 33 provides a timeline of some of the most represen- vehicles, constructions, objects, nature, sky, and void. Fig. 36
tative DL image segmentation models since 2014. shows sample segmentation maps from this dataset.
0162-8828 (c) 2021 IEEE. Personal use is permitted, but republication/redistribution requires IEEE permission. See [Link] for more information.
Authorized licensed use limited to: Carleton University. Downloaded on May 28,2021 at 14:10:11 UTC from IEEE Xplore. Restrictions apply.
This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI 10.1109/TPAMI.2021.3059968, IEEE
Transactions on Pattern Analysis and Machine Intelligence
S. MINAEE et al.: IMAGE SEGMENTATION USING DEEP LEARNING: A SURVEY 11
Fig. 33. Timeline of representative DL-based image segmentation algorithms. Orange, green, and yellow blocks indicate semantic, instance, and
panoptic segmentation algorithms, respectively.
0162-8828 (c) 2021 IEEE. Personal use is permitted, but republication/redistribution requires IEEE permission. See [Link] for more information.
Authorized licensed use limited to: Carleton University. Downloaded on May 28,2021 at 14:10:11 UTC from IEEE Xplore. Restrictions apply.
This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI 10.1109/TPAMI.2021.3059968, IEEE
Transactions on Pattern Analysis and Machine Intelligence
12 TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. ??, NO.??, ?? 2021
0162-8828 (c) 2021 IEEE. Personal use is permitted, but republication/redistribution requires IEEE permission. See [Link] for more information.
Authorized licensed use limited to: Carleton University. Downloaded on May 28,2021 at 14:10:11 UTC from IEEE Xplore. Restrictions apply.
This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI 10.1109/TPAMI.2021.3059968, IEEE
Transactions on Pattern Analysis and Machine Intelligence
S. MINAEE et al.: IMAGE SEGMENTATION USING DEEP LEARNING: A SURVEY 13
TABLE 1 TABLE 2
Accuracies of segmentation models on the PASCAL VOC test set Accuracies of segmentation models on the Cityscapes dataset
0162-8828 (c) 2021 IEEE. Personal use is permitted, but republication/redistribution requires IEEE permission. See [Link] for more information.
Authorized licensed use limited to: Carleton University. Downloaded on May 28,2021 at 14:10:11 UTC from IEEE Xplore. Restrictions apply.
This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI 10.1109/TPAMI.2021.3059968, IEEE
Transactions on Pattern Analysis and Machine Intelligence
14 TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. ??, NO.??, ?? 2021
TABLE 4 TABLE 7
Accuracies of segmentation models on the ADE20k validation dataset Segmentation model performance on the NYUD-v2 and SUN-RGBD
0162-8828 (c) 2021 IEEE. Personal use is permitted, but republication/redistribution requires IEEE permission. See [Link] for more information.
Authorized licensed use limited to: Carleton University. Downloaded on May 28,2021 at 14:10:11 UTC from IEEE Xplore. Restrictions apply.
This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI 10.1109/TPAMI.2021.3059968, IEEE
Transactions on Pattern Analysis and Machine Intelligence
S. MINAEE et al.: IMAGE SEGMENTATION USING DEEP LEARNING: A SURVEY 15
6.4 Weakly-Supervised and Unsupervised Learning algorithms. Similarly, DL-based segmentation techniques in
Weakly-supervised (a.k.a. few shot) learning [186] and un- the evaluation of construction materials [194] face challenges
supervised learning [187] are becoming very active research related to the massive volume of the related image data and
areas. These techniques promise to be specially valuable for the limited reference information for validation purposes.
image segmentation, as collecting pixel-accurately labeled Last but not least, an important application field for DL-
training images is problematic in many application domains, based segmentation has been biomedical imaging [195]. Here,
particularly so in medical image analysis. The transfer an opportunity is to design standardized image databases
learning approach is to train a generic image segmentation useful in evaluating new infectious diseases and tracking
model on a large set of labeled samples (perhaps from pandemics [196].
a public benchmark) and then fine-tune that model on a
few samples from some specific target application. Self-
7 C ONCLUSIONS
supervised learning is another promising direction that is
attracting much attraction in various fields. With the help We have surveyed image segmentation algorithms based
of self-supervised learning, many details in images can be on deep learning models, which have achieved impres-
captured in order to train segmentation models with far sive performance in various image segmentation tasks and
fewer training samples. Models based on reinforcement benchmarks, grouped into architectural categories such as:
learning could also be another potential future direction, as CNN and FCN, RNN, R-CNN, dilated CNN, attention-
they have scarcely received attention for image segmentation. based models, generative and adversarial models, among
For example, MOREL [188] introduced a deep reinforcement others. We have summarized the quantitative performance
learning approach for moving object segmentation in videos. of these models on some popular benchmarks, such as
the PASCAL VOC, MS COCO, Cityscapes, and ADE20k
datasets. Finally, we discussed some of the open challenges
6.5 Real-time Models for Various Applications
and promising research directions for deep-learning-based
In many applications, accuracy is the most important factor; image segmentation in the coming years.
however, there are applications in which it is also critical
to have segmentation models that can run in near real-time,
or at common camera frame rates (at least 25 frames per ACKNOWLEDGMENTS
second). This is useful for computer vision systems that are, We thank Tsung-Yi Lin of Google Brain as well as Jingdong
for example, deployed in autonomous vehicles. Most of the Wang and Yuhui Yuan of Microsoft Research Asia for
current models are far from this frame-rate; e.g., FCN-8 takes providing helpful comments that improved the manuscript.
roughly 100 ms to process a low-resolution image. Models
based on dilated convolution help to increase the speed of
segmentation models to some extent, but there is still plenty R EFERENCES
of room for improvement.
[1] A. Rosenfeld and A. C. Kak, Digital Picture Processing. Academic
Press, 1976.
6.6 Memory Efficient Models [2] R. Szeliski, Computer Vision: Algorithms and Applications. Springer,
2010.
Many modern segmentation models require a significant [3] D. Forsyth and J. Ponce, Computer Vision: A Modern Approach.
amount of memory even during the inference stage. So Prentice Hall, 2002.
far, much effort has been directed towards improving the [4] N. Otsu, “A threshold selection method from gray-level his-
tograms,” IEEE Transactions on Systems, Man, and Cybernetics, vol. 9,
accuracy of such models, but in order to fit them into no. 1, pp. 62–66, 1979.
specific devices, such as mobile phones, the networks must be [5] R. Nock and F. Nielsen, “Statistical region merging,” IEEE
simplified. This can be done either by using simpler models, Transactions on Pattern Analysis and Machine Intelligence, vol. 26,
no. 11, pp. 1452–1458, 2004.
or by using model compression techniques, or even by [6] N. Dhanachandra, K. Manglem, and Y. J. Chanu, “Image seg-
training a complex model and using knowledge distillation mentation using K-means clustering algorithm and subtractive
techniques to compress it into a smaller, memory efficient clustering algorithm,” Procedia Computer Science, vol. 54, pp. 764–
network that mimics the complex model. 771, 2015.
[7] L. Najman and M. Schmitt, “Watershed of a continuous function,”
Signal Processing, vol. 38, no. 1, pp. 99–112, 1994.
6.7 Applications [8] M. Kass, A. Witkin, and D. Terzopoulos, “Snakes: Active contour
models,” International Journal of Computer Vision, vol. 1, no. 4, pp.
DL-based segmentation methods have been successfully 321–331, 1988.
applied to satellite images in remote sensing [189], such as to [9] Y. Boykov, O. Veksler, and R. Zabih, “Fast approximate energy
minimization via graph cuts,” IEEE Transactions on Pattern Analysis
support urban planning [190] and precision agriculture [191]. and Machine Intelligence, vol. 23, no. 11, pp. 1222–1239, 2001.
Images collected by airborne platforms [192] and drones [193] [10] N. Plath, M. Toussaint, and S. Nakajima, “Multi-class image
have also been segmented using DL-based segmentation segmentation using conditional random fields and global classi-
methods in order to address important environmental prob- fication,” in International Conference on Machine Learning. ACM,
2009, pp. 817–824.
lems including ones related to climate change. The main [11] J.-L. Starck, M. Elad, and D. L. Donoho, “Image decomposition
challenges of the remote sensing domain stem from the via the combination of sparse representations and a variational
typically formidable size of the imagery (often collected by approach,” IEEE Transactions on Image Processing, vol. 14, no. 10,
pp. 1570–1582, 2005.
imaging spectrometers with hundreds or even thousands of
[12] S. Minaee and Y. Wang, “An ADMM approach to masked signal
spectral bands) and the limited ground-truth information decomposition using subspace representation,” IEEE Transactions
necessary to evaluate the accuracy of the segmentation on Image Processing, vol. 28, no. 7, pp. 3192–3204, 2019.
0162-8828 (c) 2021 IEEE. Personal use is permitted, but republication/redistribution requires IEEE permission. See [Link] for more information.
Authorized licensed use limited to: Carleton University. Downloaded on May 28,2021 at 14:10:11 UTC from IEEE Xplore. Restrictions apply.
This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI 10.1109/TPAMI.2021.3059968, IEEE
Transactions on Pattern Analysis and Machine Intelligence
16 TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. ??, NO.??, ?? 2021
[13] L.-C. Chen, G. Papandreou, F. Schroff, and H. Adam, “Rethinking recurrent neural networks,” in IEEE International Conference on
atrous convolution for semantic image segmentation,” arXiv Computer Vision, 2015, pp. 1529–1537.
preprint arXiv:1706.05587, 2017. [39] G. Lin, C. Shen, A. Van Den Hengel, and I. Reid, “Efficient
[14] S. Minaee, Y. Boykov, F. Porikli, A. Plaza, N. Kehtarnavaz, and piecewise training of deep structured models for semantic seg-
D. Terzopoulos, “Image segmentation using deep learning: A mentation,” in IEEE Conference on Computer Vision and Pattern
survey,” arXiv preprint arXiv:2001.05566, 2020. Recognition, 2016, pp. 3194–3203.
[15] Y. LeCun, L. Bottou, Y. Bengio, P. Haffner et al., “Gradient-based [40] Z. Liu, X. Li, P. Luo, C.-C. Loy, and X. Tang, “Semantic image
learning applied to document recognition,” Proceedings of the IEEE, segmentation via deep parsing network,” in IEEE International
vol. 86, no. 11, pp. 2278–2324, 1998. Conference on Computer Vision, 2015, pp. 1377–1385.
[16] K. Fukushima, “Neocognitron: A self-organizing neural network [41] H. Noh, S. Hong, and B. Han, “Learning deconvolution network
model for a mechanism of pattern recognition unaffected by shift for semantic segmentation,” in IEEE International Conference on
in position,” Biological Cybernetics, vol. 36, no. 4, pp. 193–202, 1980. Computer Vision, 2015, pp. 1520–1528.
[17] A. Waibel, T. Hanazawa, G. Hinton, K. Shikano, and K. J. Lang, [42] Y. Yuan, X. Chen, and J. Wang, “Object-contextual representations
“Phoneme recognition using time-delay neural networks,” IEEE for semantic segmentation,” arXiv preprint arXiv:1909.11065, 2019.
Transactions on Acoustics, Speech, and Signal Processing, vol. 37, no. 3, [43] J. Fu, J. Liu, Y. Wang, J. Zhou, C. Wang, and H. Lu, “Stacked decon-
pp. 328–339, 1989. volutional network for semantic segmentation,” IEEE Transactions
[18] A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classifi- on Image Processing, 2019.
cation with deep convolutional neural networks,” in Advances in [44] A. Chaurasia and E. Culurciello, “LinkNet: Exploiting encoder
Neural Information Processing Systems, 2012, pp. 1097–1105. representations for efficient semantic segmentation,” in IEEE Inter-
[19] K. Simonyan and A. Zisserman, “Very deep convolutional national Conference on Visual Communications and Image Processing.
networks for large-scale image recognition,” arXiv preprint IEEE, 2017, pp. 1–4.
arXiv:1409.1556, 2014.
[45] X. Xia and B. Kulis, “W-Net: A deep model for fully unsupervised
[20] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image segmentation,” arXiv preprint arXiv:1711.08506, 2017.
image recognition,” in IEEE Conference on Computer Vision and
Pattern Recognition, 2016, pp. 770–778. [46] Y. Cheng, R. Cai, Z. Li, X. Zhao, and K. Huang, “Locality-sensitive
deconvolution networks with gated fusion for RGB-D indoor
[21] [Link]
semantic segmentation,” in IEEE Conference on Computer Vision
[22] D. E. Rumelhart, G. E. Hinton, R. J. Williams et al., “Learning and Pattern Recognition, 2017, pp. 3029–3037.
representations by back-propagating errors,” Cognitive Modeling,
[47] O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional
vol. 5, no. 3, p. 1, 1988.
networks for biomedical image segmentation,” in International
[23] S. Hochreiter and J. Schmidhuber, “Long short-term memory,”
Conference on Medical Image Computing and Computer-Assisted
Neural Computation, vol. 9, no. 8, pp. 1735–1780, 1997.
Intervention. Springer, 2015, pp. 234–241.
[24] I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning. MIT
[48] Z. Zhou, M. M. R. Siddiquee, N. Tajbakhsh, and J. Liang, “UNet++:
Press, 2016.
A nested U-Net architecture for medical image segmentation,” in
[25] V. Badrinarayanan, A. Kendall, and R. Cipolla, “SegNet: A deep
Deep Learning in Medical Image Analysis and Multimodal Learning
convolutional encoder-decoder architecture for image segmenta-
for Clinical Decision Support. Springer, 2018, pp. 3–11.
tion,” IEEE Transactions on Pattern Analysis and Machine Intelligence,
vol. 39, no. 12, pp. 2481–2495, 2017. [49] Z. Zhang, Q. Liu, and Y. Wang, “Road extraction by deep residual
U-Net,” IEEE Geoscience and Remote Sensing Letters, vol. 15, no. 5,
[26] I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley,
pp. 749–753, 2018.
S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial
nets,” in Advances in Neural Information Processing Systems, 2014, [50] Ö. Çiçek, A. Abdulkadir, S. S. Lienkamp, T. Brox, and O. Ron-
pp. 2672–2680. neberger, “3D U-Net: learning dense volumetric segmentation
[27] A. Radford, L. Metz, and S. Chintala, “Unsupervised represen- from sparse annotation,” in International Conference on Medical
tation learning with deep convolutional generative adversarial Image Computing and Computer-Assisted Intervention. Springer,
networks,” arXiv preprint arXiv:1511.06434, 2015. 2016, pp. 424–432.
[28] M. Mirza and S. Osindero, “Conditional generative adversarial [51] F. Milletari, N. Navab, and S.-A. Ahmadi, “V-Net: Fully convo-
nets,” arXiv preprint arXiv:1411.1784, 2014. lutional neural networks for volumetric medical image segmen-
[29] M. Arjovsky, S. Chintala, and L. Bottou, “Wasserstein GAN,” arXiv tation,” in International Conference on 3D Vision. IEEE, 2016, pp.
preprint arXiv:1701.07875, 2017. 565–571.
[30] J. Long, E. Shelhamer, and T. Darrell, “Fully convolutional [52] T. Brosch, L. Y. Tang, Y. Yoo, D. K. Li, A. Traboulsee, and R. Tam,
networks for semantic segmentation,” in IEEE Conference on “Deep 3D convolutional encoder networks with shortcuts for
Computer Vision and Pattern Recognition, 2015, pp. 3431–3440. multiscale feature integration applied to multiple sclerosis lesion
[31] G. Wang, W. Li, S. Ourselin, and T. Vercauteren, “Automatic brain segmentation,” IEEE Transactions on Medical Imaging, vol. 35, no. 5,
tumor segmentation using cascaded anisotropic convolutional pp. 1229–1239, 2016.
neural networks,” in International MICCAI Brainlesion Workshop. [53] T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Be-
Springer, 2017, pp. 178–190. longie, “Feature pyramid networks for object detection,” in IEEE
[32] Y. Li, H. Qi, J. Dai, X. Ji, and Y. Wei, “Fully convolutional instance- Conference on Computer Vision and Pattern Recognition, 2017, pp.
aware semantic segmentation,” in IEEE Conference on Computer 2117–2125.
Vision and Pattern Recognition, 2017, pp. 2359–2367. [54] H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia, “Pyramid scene parsing
[33] Y. Yuan, M. Chao, and Y.-C. Lo, “Automatic skin lesion seg- network,” in IEEE Conference on Computer Vision and Pattern
mentation using deep fully convolutional networks with Jaccard Recognition, 2017, pp. 2881–2890.
distance,” IEEE Transactions on Medical Imaging, vol. 36, no. 9, pp. [55] G. Ghiasi and C. C. Fowlkes, “Laplacian pyramid reconstruction
1876–1886, 2017. and refinement for semantic segmentation,” in European Conference
[34] N. Liu, H. Li, M. Zhang, J. Liu, Z. Sun, and T. Tan, “Accurate on Computer Vision. Springer, 2016, pp. 519–534.
iris segmentation in non-cooperative environments using fully [56] J. He, Z. Deng, and Y. Qiao, “Dynamic multi-scale filters for se-
convolutional networks,” in International Conference on Biometrics. mantic segmentation,” in IEEE International Conference on Computer
IEEE, 2016, pp. 1–8. Vision, 2019, pp. 3562–3572.
[35] W. Liu, A. Rabinovich, and A. C. Berg, “ParseNet: Looking wider [57] H. Ding, X. Jiang, B. Shuai, A. Qun Liu, and G. Wang, “Context
to see better,” arXiv preprint arXiv:1506.04579, 2015. contrasted feature and gated multi-scale aggregation for scene
[36] L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. segmentation,” in IEEE Conference on Computer Vision and Pattern
Yuille, “Semantic image segmentation with deep convolutional Recognition, 2018, pp. 2393–2402.
nets and fully connected CRFs,” arXiv preprint arXiv:1412.7062, [58] J. He, Z. Deng, L. Zhou, Y. Wang, and Y. Qiao, “Adaptive pyramid
2014. context network for semantic segmentation,” in Conference on
[37] A. G. Schwing and R. Urtasun, “Fully connected deep structured Computer Vision and Pattern Recognition, 2019, pp. 7519–7528.
networks,” arXiv preprint arXiv:1503.02351, 2015. [59] D. Lin, Y. Ji, D. Lischinski, D. Cohen-Or, and H. Huang, “Multi-
[38] S. Zheng, S. Jayasumana, B. Romera-Paredes, V. Vineet, Z. Su, scale context intertwining for semantic segmentation,” in European
D. Du, C. Huang, and P. H. Torr, “Conditional random fields as Conference on Computer Vision, 2018, pp. 603–619.
0162-8828 (c) 2021 IEEE. Personal use is permitted, but republication/redistribution requires IEEE permission. See [Link] for more information.
Authorized licensed use limited to: Carleton University. Downloaded on May 28,2021 at 14:10:11 UTC from IEEE Xplore. Restrictions apply.
This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI 10.1109/TPAMI.2021.3059968, IEEE
Transactions on Pattern Analysis and Machine Intelligence
S. MINAEE et al.: IMAGE SEGMENTATION USING DEEP LEARNING: A SURVEY 17
[60] G. Li, Y. Xie, L. Lin, and Y. Yu, “Instance-level salient object [83] F. Visin, K. Kastner, K. Cho, M. Matteucci, A. Courville, and
segmentation,” in IEEE Conference on Computer Vision and Pattern Y. Bengio, “ReNet: A recurrent neural network based alternative
Recognition, 2017, pp. 2386–2395. to convolutional networks,” arXiv preprint arXiv:1505.00393, 2015.
[61] S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards [84] W. Byeon, T. M. Breuel, F. Raue, and M. Liwicki, “Scene labeling
real-time object detection with region proposal networks,” in with LSTM recurrent neural networks,” in IEEE Conference on
Advances in Neural Information Processing Systems, 2015, pp. 91–99. Computer Vision and Pattern Recognition, 2015, pp. 3547–3555.
[62] K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask R-CNN,” in [85] X. Liang, X. Shen, J. Feng, L. Lin, and S. Yan, “Semantic object
IEEE International Conference on Computer Vision, 2017, pp. 2961– parsing with graph LSTM,” in European Conference on Computer
2969. Vision. Springer, 2016, pp. 125–143.
[63] S. Liu, L. Qi, H. Qin, J. Shi, and J. Jia, “Path aggregation network [86] Y. Xiang and D. Fox, “DA-RNN: Semantic mapping with data
for instance segmentation,” in IEEE Conference on Computer Vision associated recurrent neural networks,” arXiv:1703.03098, 2017.
and Pattern Recognition, 2018, pp. 8759–8768. [87] R. Hu, M. Rohrbach, and T. Darrell, “Segmentation from natural
[64] J. Dai, K. He, and J. Sun, “Instance-aware semantic segmentation language expressions,” in European Conference on Computer Vision.
via multi-task network cascades,” in IEEE Conference on Computer Springer, 2016, pp. 108–124.
Vision and Pattern Recognition, 2016, pp. 3150–3158. [88] L.-C. Chen, Y. Yang, J. Wang, W. Xu, and A. L. Yuille, “Attention
[65] R. Hu, P. Dollár, K. He, T. Darrell, and R. Girshick, “Learning to to scale: Scale-aware semantic image segmentation,” in IEEE
segment every thing,” in IEEE Conference on Computer Vision and Conference on Computer Vision and Pattern Recognition, 2016, pp.
Pattern Recognition, 2018, pp. 4233–4241. 3640–3649.
[66] L.-C. Chen, A. Hermans, G. Papandreou, F. Schroff, P. Wang, [89] Q. Huang, C. Xia, C. Wu, S. Li, Y. Wang, Y. Song, and C.-C. J. Kuo,
and H. Adam, “Masklab: Instance segmentation by refining object “Semantic segmentation with reverse attention,” arXiv preprint
detection with semantic and direction features,” in IEEE Conference arXiv:1707.06426, 2017.
on Computer Vision and Pattern Recognition, 2018, pp. 4013–4022. [90] H. Li, P. Xiong, J. An, and L. Wang, “Pyramid attention network
[67] X. Chen, R. Girshick, K. He, and P. Dollár, “Tensormask: for semantic segmentation,” arXiv preprint arXiv:1805.10180, 2018.
A foundation for dense object segmentation,” arXiv preprint [91] J. Fu, J. Liu, H. Tian, Y. Li, Y. Bao, Z. Fang, and H. Lu, “Dual
arXiv:1903.12174, 2019. attention network for scene segmentation,” in IEEE Conference on
[68] J. Dai, Y. Li, K. He, and J. Sun, “R-FCN: Object detection via Computer Vision and Pattern Recognition, 2019, pp. 3146–3154.
region-based fully convolutional networks,” in Advances in Neural [92] Y. Yuan and J. Wang, “OCNet: Object context network for scene
Information Processing Systems, 2016, pp. 379–387. parsing,” arXiv preprint arXiv:1809.00916, 2018.
[69] P. O. Pinheiro, R. Collobert, and P. Dollár, “Learning to segment [93] H. Zhang, C. Wu, Z. Zhang, Y. Zhu, Z. Zhang, H. Lin, Y. Sun,
object candidates,” in Advances in Neural Information Processing T. He, J. Mueller, R. Manmatha et al., “Resnest: Split-attention
Systems, 2015, pp. 1990–1998. networks,” arXiv preprint arXiv:2004.08955, 2020.
[94] S. Choi, J. T. Kim, and J. Choo, “Cars can’t fly up in the sky:
[70] E. Xie, P. Sun, X. Song, W. Wang, X. Liu, D. Liang, C. Shen, and
Improving urban-scene segmentation via height-driven attention
P. Luo, “PolarMask: Single shot instance segmentation with polar
networks,” in Proceedings of the IEEE/CVF Conference on Computer
representation,” arXiv preprint arXiv:1909.13226, 2019.
Vision and Pattern Recognition, 2020, pp. 9373–9383.
[71] Z. Hayder, X. He, and M. Salzmann, “Boundary-aware instance
[95] X. Li, Z. Zhong, J. Wu, Y. Yang, Z. Lin, and H. Liu, “Expectation-
segmentation,” in IEEE Conference on Computer Vision and Pattern
maximization attention networks for semantic segmentation,” in
Recognition, 2017, pp. 5696–5704.
IEEE International Conference on Computer Vision, 2019, pp. 9167–
[72] Y. Lee and J. Park, “CenterMask: Real-time anchor-free instance 9176.
segmentation,” in IEEE Conference on Computer Vision and Pattern
[96] Z. Huang, X. Wang, L. Huang, C. Huang, Y. Wei, and W. Liu,
Recognition, 2020, pp. 13 906–13 915.
“CCNet: Criss-cross attention for semantic segmentation,” in IEEE
[73] M. Bai and R. Urtasun, “Deep watershed transform for instance International Conference on Computer Vision, 2019, pp. 603–612.
segmentation,” in IEEE Conference on Computer Vision and Pattern [97] M. Ren and R. S. Zemel, “End-to-end instance segmentation with
Recognition, 2017, pp. 5221–5229. recurrent attention,” in IEEE Conference on Computer Vision and
[74] D. Bolya, C. Zhou, F. Xiao, and Y. J. Lee, “YOLACT: Real- Pattern Recognition, 2017, pp. 6656–6664.
time instance segmentation,” in IEEE International Conference on [98] H. Zhao, Y. Zhang, S. Liu, J. Shi, C. Change Loy, D. Lin, and J. Jia,
Computer Vision, 2019, pp. 9157–9166. “PSANet: Point-wise spatial attention network for scene parsing,”
[75] A. Fathi, Z. Wojna, V. Rathod, P. Wang, H. O. Song, S. Guadarrama, in European Conference on Computer Vision, 2018, pp. 267–283.
and K. P. Murphy, “Semantic instance segmentation via deep [99] C. Yu, J. Wang, C. Peng, C. Gao, G. Yu, and N. Sang, “Learning
metric learning,” arXiv preprint arXiv:1703.10277, 2017. a discriminative feature network for semantic segmentation,” in
[76] L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. IEEE Conference on Computer Vision and Pattern Recognition, 2018,
Yuille, “DeepLab: Semantic image segmentation with deep convo- pp. 1857–1866.
lutional nets, atrous convolution, and fully connected crfs,” IEEE [100] P. Luc, C. Couprie, S. Chintala, and J. Verbeek, “Semantic segmen-
Transactions on Pattern Analysis and Machine Intelligence, vol. 40, tation using adversarial networks,” arXiv preprint arXiv:1611.08408,
no. 4, pp. 834–848, 2017. 2016.
[77] F. Yu and V. Koltun, “Multi-scale context aggregation by dilated [101] N. Souly, C. Spampinato, and M. Shah, “Semi supervised semantic
convolutions,” arXiv preprint arXiv:1511.07122, 2015. segmentation using generative adversarial network,” in IEEE
[78] P. Wang, P. Chen, Y. Yuan, D. Liu, Z. Huang, X. Hou, and G. Cot- International Conference on Computer Vision, 2017, pp. 5688–5696.
trell, “Understanding convolution for semantic segmentation,” in [102] W.-C. Hung, Y.-H. Tsai, Y.-T. Liou, Y.-Y. Lin, and M.-H. Yang,
IEEE Winter Conference on Applications of Computer Vision, 2018, pp. “Adversarial learning for semi-supervised semantic segmentation,”
1451–1460. arXiv preprint arXiv:1802.07934, 2018.
[79] M. Yang, K. Yu, C. Zhang, Z. Li, and K. Yang, “DenseASPP for [103] Y. Xue, T. Xu, H. Zhang, L. R. Long, and X. Huang, “SegAN:
semantic segmentation in street scenes,” in IEEE Conference on Adversarial network with multi-scale L1 loss for medical image
Computer Vision and Pattern Recognition, 2018, pp. 3684–3692. segmentation,” Neuroinformatics, vol. 16, no. 3-4, pp. 383–392, 2018.
[80] A. Paszke, A. Chaurasia, S. Kim, and E. Culurciello, “ENet: A deep [104] M. Majurski, P. Manescu, S. Padi, N. Schaub, N. Hotaling, C. Si-
neural network architecture for real-time semantic segmentation,” mon Jr, and P. Bajcsy, “Cell image segmentation using generative
arXiv preprint arXiv:1606.02147, 2016. adversarial networks, transfer learning, and augmentations,”
[81] L.-C. Chen, Y. Zhu, G. Papandreou, F. Schroff, and H. Adam, in IEEE Conference on Computer Vision and Pattern Recognition
“Encoder-decoder with atrous separable convolution for semantic Workshops, 2019, pp. 0–0.
image segmentation,” in European Conference on Computer Vision, [105] K. Ehsani, R. Mottaghi, and A. Farhadi, “SegAN: Segmenting and
2018, pp. 801–818. generating the invisible,” in IEEE Conference on Computer Vision
[82] F. Visin, M. Ciccone, A. Romero, K. Kastner, K. Cho, Y. Bengio, and Pattern Recognition, 2018, pp. 6144–6153.
M. Matteucci, and A. Courville, “ReSeg: A recurrent neural [106] T. F. Chan and L. A. Vese, “Active contours without edges,” IEEE
network-based model for semantic segmentation,” in IEEE Con- Transactions on Image Processing, vol. 10, no. 2, pp. 266–277, 2001.
ference on Computer Vision and Pattern Recognition Workshops, 2016, [107] X. Chen, B. M. Williams, S. R. Vallabhaneni, G. Czanner,
pp. 41–48. R. Williams, and Y. Zheng, “Learning active contour models for
0162-8828 (c) 2021 IEEE. Personal use is permitted, but republication/redistribution requires IEEE permission. See [Link] for more information.
Authorized licensed use limited to: Carleton University. Downloaded on May 28,2021 at 14:10:11 UTC from IEEE Xplore. Restrictions apply.
This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI 10.1109/TPAMI.2021.3059968, IEEE
Transactions on Pattern Analysis and Machine Intelligence
18 TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. ??, NO.??, ?? 2021
medical image segmentation,” in IEEE Conference on Computer [128] C. Yu, J. Wang, C. Peng, C. Gao, G. Yu, and N. Sang, “BiSeNet:
Vision and Pattern Recognition, 2019, pp. 11 632–11 640. Bilateral segmentation network for real-time semantic segmenta-
[108] S. Gur, L. Wolf, L. Golgher, and P. Blinder, “Unsupervised tion,” in European Conference on Computer Vision, 2018, pp. 325–341.
microvascular image segmentation using an active contours [129] B. Cheng, L.-C. Chen, Y. Wei, Y. Zhu, Z. Huang, J. Xiong, T. S.
mimicking neural network,” in IEEE International Conference on Huang, W.-M. Hwu, and H. Shi, “SPGNet: Semantic prediction
Computer Vision, 2019, pp. 10 722–10 731. guidance for scene parsing,” in IEEE International Conference on
[109] P. Marquez-Neila, L. Baumela, and L. Alvarez, “A morphological Computer Vision, 2019, pp. 5218–5228.
approach to curvature-based evolution of curves and surfaces,” [130] T. Takikawa, D. Acuna, V. Jampani, and S. Fidler, “Gated-SCNN:
IEEE Transactions on Pattern Analysis and Machine Intelligence, Gated shape cnns for semantic segmentation,” in IEEE International
vol. 36, no. 1, pp. 2–17, 2014. Conference on Computer Vision, 2019, pp. 5229–5238.
[110] T. H. N. Le, K. G. Quach, K. Luu, C. N. Duong, and M. Savvides, [131] J. Fu, J. Liu, Y. Wang, Y. Li, Y. Bao, J. Tang, and H. Lu, “Adaptive
“Reformulating level sets as deep recurrent neural network context network for scene parsing,” in IEEE International Conference
approach to semantic segmentation,” IEEE Transactions on Image on Computer Vision, 2019, pp. 6748–6757.
Processing, vol. 27, no. 5, pp. 2393–2407, 2018. [132] X. Liang, H. Zhou, and E. Xing, “Dynamic-structured semantic
[111] C. Rupprecht, E. Huaroc, M. Baust, and N. Navab, “Deep active propagation network,” in IEEE Conference on Computer Vision and
contours,” arXiv preprint arXiv:1607.05074, 2016. Pattern Recognition, 2018, pp. 752–761.
[112] A. Hatamizadeh, A. Hoogi, D. Sengupta, W. Lu, B. Wilcox, [133] X. Liang, Z. Hu, H. Zhang, L. Lin, and E. P. Xing, “Symbolic graph
D. Rubin, and D. Terzopoulos, “Deep active lesion segmentation,” reasoning meets convolutions,” in Advances in Neural Information
in International Workshop on Machine Learning in Medical Imaging, Processing Systems, 2018, pp. 1853–1863.
ser. Lecture Notes in Computer Science, vol. 11861. Springer, [134] B. Zhou, H. Zhao, X. Puig, S. Fidler, A. Barriuso, and A. Torralba,
2019, pp. 98–105. “Scene parsing through ADE20K dataset,” in IEEE Conference on
[113] D. Marcos, D. Tuia, B. Kellenberger, L. Zhang, M. Bai, R. Liao, Computer Vision and Pattern Recognition, 2017.
and R. Urtasun, “Learning deep structured active contours end-to- [135] R. Zhang, S. Tang, Y. Zhang, J. Li, and S. Yan, “Scale-adaptive
end,” in IEEE Conference on Computer Vision and Pattern Recognition, convolutions for scene parsing,” in IEEE International Conference
2018, pp. 8877–8885. on Computer Vision, 2017, pp. 2031–2039.
[114] D. Cheng, R. Liao, S. Fidler, and R. Urtasun, “DARNet: Deep [136] T. Xiao, Y. Liu, B. Zhou, Y. Jiang, and J. Sun, “Unified perceptual
active ray network for building segmentation,” in IEEE Conference parsing for scene understanding,” in European Conference on
on Computer Vision and Pattern Recognition, 2019, pp. 7431–7439. Computer Vision, 2018, pp. 418–434.
[137] B. Zoph, G. Ghiasi, T.-Y. Lin, Y. Cui, H. Liu, E. D. Cubuk, and
[115] A. Hatamizadeh, D. Sengupta, and D. Terzopoulos, “End-to-end
Q. V. Le, “Rethinking pre-training and self-training,” arXiv preprint
trainable deep active contour models for automated image seg-
arXiv:2006.06882, 2020.
mentation: Delineating buildings in aerial imagery,” in European
Conference on Computer Vision, 2020, pp. 730–746. [138] X. Zhang, H. Xu, H. Mo, J. Tan, C. Yang, and W. Ren, “DCNAS:
Densely connected neural architecture search for semantic image
[116] H. Zhang, K. Dana, J. Shi, Z. Zhang, X. Wang, A. Tyagi, and
segmentation,” arXiv preprint arXiv:2003.11883, 2020.
A. Agrawal, “Context encoding for semantic segmentation,” in
IEEE Conference on Computer Vision and Pattern Recognition, 2018, [139] A. Tao, K. Sapra, and B. Catanzaro, “Hierarchical multi-scale atten-
pp. 7151–7160. tion for semantic segmentation,” arXiv preprint arXiv:2005.10821,
2020.
[117] G. Lin, A. Milan, C. Shen, and I. Reid, “RefineNet: Multi-path
[140] D. Seichter, M. Köhler, B. Lewandowski, T. Wengefeld, and H.-M.
refinement networks for high-resolution semantic segmentation,”
Gross, “Efficient rgb-d semantic segmentation for indoor scene
in IEEE Conference on Computer Vision and Pattern Recognition, 2017,
analysis,” arXiv preprint arXiv:2011.06961, 2020.
pp. 1925–1934.
[141] M. Zhen, J. Wang, L. Zhou, S. Li, T. Shen, J. Shang, T. Fang, and
[118] G. Song, H. Myeong, and K. Mu Lee, “SeedNet: Automatic seed L. Quan, “Joint semantic segmentation and boundary detection
generation with deep reinforcement learning for robust interactive using iterative pyramid contexts,” in Proceedings of the IEEE/CVF
segmentation,” in IEEE Conference on Computer Vision and Pattern Conference on Computer Vision and Pattern Recognition, 2020, pp.
Recognition, 2018, pp. 1760–1768. 13 666–13 675.
[119] J. Dai, K. He, and J. Sun, “BoxSup: Exploiting bounding boxes to [142] Y. Li, L. Song, Y. Chen, Z. Li, X. Zhang, X. Wang, and J. Sun,
supervise convolutional networks for semantic segmentation,” in “Learning dynamic routing for semantic segmentation,” in Pro-
IEEE International Conference on Computer Vision, 2015, pp. 1635– ceedings of the IEEE/CVF Conference on Computer Vision and Pattern
1643. Recognition, 2020, pp. 8553–8562.
[120] C. Peng, X. Zhang, G. Yu, G. Luo, and J. Sun, “Large kernel [143] A. Kirillov, K. He, R. Girshick, C. Rother, and P. Dollár, “Panoptic
matters — improve semantic segmentation by global convolu- segmentation,” in IEEE Conference on Computer Vision and Pattern
tional network,” in IEEE Conference on Computer Vision and Pattern Recognition, 2019, pp. 9404–9413.
Recognition, 2017, pp. 4353–4361. [144] A. Kirillov, R. Girshick, K. He, and P. Dollar, “Panoptic feature
[121] Z. Wu, C. Shen, and A. Van Den Hengel, “Wider or deeper: pyramid networks,” in IEEE Conference on Computer Vision and
Revisiting the resnet model for visual recognition,” Pattern Pattern Recognition, 2019, pp. 6399–6408.
Recognition, vol. 90, pp. 119–133, 2019. [145] Y. Li, X. Chen, Z. Zhu, L. Xie, G. Huang, D. Du, and X. Wang,
[122] Z. Zhang, X. Zhang, C. Peng, X. Xue, and J. Sun, “ExFuse: “Attention-guided unified network for panoptic segmentation,” in
Enhancing feature fusion for semantic segmentation,” in European IEEE Conference on Computer Vision and Pattern Recognition, 2019.
Conference on Computer Vision, 2018, pp. 269–284. [146] L. Porzi, S. R. Bulo, A. Colovic, and P. Kontschieder, “Seamless
[123] M. Mostajabi, P. Yadollahpour, and G. Shakhnarovich, “Feedfor- scene segmentation,” in IEEE Conference on Computer Vision and
ward semantic segmentation with zoom-out features,” in IEEE Pattern Recognition, 2019, pp. 8277–8286.
Conference on Computer Vision and Pattern Recognition, 2015, pp. [147] B. Cheng, M. D. Collins, Y. Zhu, T. Liu, T. S. Huang, H. Adam, and
3376–3385. L.-C. Chen, “Panoptic-DeepLab,” arXiv preprint arXiv:1910.04751,
[124] W. Wang, J. Shen, and F. Porikli, “Saliency-aware geodesic video 2019.
object segmentation,” in IEEE Conference on Computer Vision and [148] Y. Xiong, R. Liao, H. Zhao, R. Hu, M. Bai, E. Yumer, and R. Urtasun,
Pattern Recognition, 2015, pp. 3395–3402. “UPSNet: A unified panoptic segmentation network,” in IEEE
[125] P. Luo, G. Wang, L. Lin, and X. Wang, “Deep dual learning for Conference on Computer Vision and Pattern Recognition, 2019, pp.
semantic image segmentation,” in IEEE International Conference on 8818–8826.
Computer Vision, 2017, pp. 2718–2726. [149] R. Mohan and A. Valada, “EfficientPS: Efficient panoptic segmen-
[126] X. Li, Z. Jie, W. Wang, C. Liu, J. Yang, X. Shen, Z. Lin, Q. Chen, tation,” arXiv preprint arXiv:2004.02307, 2020.
S. Yan, and J. Feng, “FoveaNet: Perspective-aware urban scene [150] M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zis-
parsing,” in IEEE International Conference on Computer Vision, 2017, serman, “The PASCAL visual object classes (VOC) challenge,”
pp. 784–792. International Journal of Computer Vision, vol. 88, pp. 303–338, 2010.
[127] I. Kreso, S. Segvic, and J. Krapac, “Ladder-style densenets for se- [151] [Link]
mantic segmentation of large natural images,” in IEEE International [152] R. Mottaghi, X. Chen, X. Liu, N.-G. Cho, S.-W. Lee, S. Fidler,
Conference on Computer Vision, 2017, pp. 238–245. R. Urtasun, and A. Yuille, “The role of context for object detection
0162-8828 (c) 2021 IEEE. Personal use is permitted, but republication/redistribution requires IEEE permission. See [Link] for more information.
Authorized licensed use limited to: Carleton University. Downloaded on May 28,2021 at 14:10:11 UTC from IEEE Xplore. Restrictions apply.
This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI 10.1109/TPAMI.2021.3059968, IEEE
Transactions on Pattern Analysis and Machine Intelligence
S. MINAEE et al.: IMAGE SEGMENTATION USING DEEP LEARNING: A SURVEY 19
and semantic segmentation in the wild,” in IEEE Conference on Conference on Computer Vision and Pattern Recognition, 2019, pp.
Computer Vision and Pattern Recognition, 2014, pp. 891–898. 6172–6181.
[153] T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, [175] K. Sofiiuk, O. Barinova, and A. Konushin, “AdaptIS: Adaptive
P. Dollár, and C. L. Zitnick, “Microsoft COCO: Common objects instance selection network,” in Proceedings of the IEEE International
in context,” in European Conference on Computer Vision. Springer, Conference on Computer Vision, 2019, pp. 7355–7363.
2014. [176] J. Lazarow, K. Lee, K. Shi, and Z. Tu, “Learning instance occlusion
[154] M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Be- for panoptic segmentation,” in IEEE Conference on Computer Vision
nenson, U. Franke, S. Roth, and B. Schiele, “The cityscapes dataset and Pattern Recognition, 2020, pp. 10 720–10 729.
for semantic urban scene understanding,” in IEEE Conference on [177] Z. Deng, S. Todorovic, and L. Jan Latecki, “Semantic segmentation
Computer Vision and Pattern Recognition, 2016, pp. 3213–3223. of RGBD images with mutex constraints,” in IEEE International
[155] C. Liu, J. Yuen, and A. Torralba, “Nonparametric scene parsing: Conference on Computer Vision, 2015, pp. 1733–1741.
Label transfer via dense scene alignment,” in IEEE Conference on [178] D. Eigen and R. Fergus, “Predicting depth, surface normals
Computer Vision and Pattern Recognition, 2009. and semantic labels with a common multi-scale convolutional
[156] S. Gould, R. Fulton, and D. Koller, “Decomposing a scene into architecture,” in IEEE International Conference on Computer Vision,
geometric and semantically consistent regions,” in International 2015, pp. 2650–2658.
Conference on Computer Vision. IEEE, 2009, pp. 1–8. [179] A. Mousavian, H. Pirsiavash, and J. Kosecka, “Joint semantic
[157] D. Martin, C. Fowlkes, D. Tal, and J. Malik, “A database of segmentation and depth estimation with deep convolutional
human segmented natural images and its application to evaluating networks,” in International Conference on 3D Vision. IEEE, 2016.
segmentation algorithms and measuring ecological statistics,” in [180] A. Kendall, V. Badrinarayanan, and R. Cipolla, “Bayesian SegNet:
International Conference on Computer Vision, vol. 2, July 2001, pp. Model uncertainty in deep convolutional encoder-decoder archi-
416–423. tectures for scene understanding,” arXiv preprint arXiv:1511.02680,
2015.
[158] A. Prest, C. Leistner, J. Civera, C. Schmid, and V. Ferrari, “Learning
[181] X. Qi, R. Liao, J. Jia, S. Fidler, and R. Urtasun, “3D graph neural
object class detectors from weakly annotated video,” in IEEE
networks for RGBD semantic segmentation,” in IEEE International
Conference on Computer Vision and Pattern Recognition. IEEE, 2012,
Conference on Computer Vision, 2017, pp. 5199–5208.
pp. 3282–3289.
[182] W. Wang and U. Neumann, “Depth-aware CNN for RGB-D
[159] S. D. Jain and K. Grauman, “Supervoxel-consistent foreground segmentation,” in European Conference on Computer Vision, 2018,
propagation in video,” in European Conference on Computer Vision. pp. 135–150.
Springer, 2014, pp. 656–671. [183] S. Vandenhende, S. Georgoulis, and L. Van Gool, “Mti-net: Multi-
[160] A. Geiger, P. Lenz, C. Stiller, and R. Urtasun, “Vision meets scale task interaction networks for multi-task learning,” arXiv
robotics: The KITTI dataset,” The International Journal of Robotics preprint arXiv:2001.06902, 2020.
Research, vol. 32, no. 11, pp. 1231–1237, 2013. [184] S.-J. Park, K.-S. Hong, and S. Lee, “RDFNet: RGB-D multi-level
[161] J. M. Alvarez, T. Gevers, Y. LeCun, and A. M. Lopez, “Road scene residual feature fusion for indoor semantic segmentation,” in IEEE
segmentation from a single image,” in European Conference on International Conference on Computer Vision, 2017, pp. 4980–4989.
Computer Vision. Springer, 2012, pp. 376–389. [185] J. Jiao, Y. Wei, Z. Jie, H. Shi, R. W. Lau, and T. S. Huang, “Geometry-
[162] B. Hariharan, P. Arbeláez, L. Bourdev, S. Maji, and J. Malik, “Se- aware distillation for indoor semantic segmentation,” in IEEE
mantic contours from inverse detectors,” in International Conference Conference on Computer Vision and Pattern Recognition, 2019, pp.
on Computer Vision. IEEE, 2011, pp. 991–998. 2869–2878.
[163] X. Chen, R. Mottaghi, X. Liu, S. Fidler, R. Urtasun, and A. Yuille, [186] Z.-H. Zhou, “A brief introduction to weakly supervised learning,”
“Detect what you can: Detecting and representing objects using National Science Review, vol. 5, no. 1, pp. 44–53, 2018.
holistic models and body parts,” in IEEE Conference on Computer [187] L. Jing and Y. Tian, “Self-supervised visual feature learning with
Vision and Pattern Recognition, 2014, pp. 1971–1978. deep neural networks: A survey,” IEEE Transactions on Pattern
[164] G. Ros, L. Sellart, J. Materzynska, D. Vazquez, and A. M. Lopez, Analysis and Machine Intelligence, 2020.
“The SYNTHIA dataset: A large collection of synthetic images for [188] V. Goel, J. Weng, and P. Poupart, “Unsupervised video object
semantic segmentation of urban scenes,” in IEEE Conference on segmentation for deep reinforcement learning,” in Advances in
Computer Vision and Pattern Recognition, 2016, pp. 3234–3243. Neural Information Processing Systems, 2018, pp. 5683–5694.
[165] X. Shen, A. Hertzmann, J. Jia, S. Paris, B. Price, E. Shechtman, and [189] L. Ma, Y. Liu, X. Zhang, Y. Ye, G. Yin, and B. A. Johnson, “Deep
I. Sachs, “Automatic portrait segmentation for image stylization,” learning in remote sensing applications: A meta-analysis and
in Computer Graphics Forum, vol. 35, no. 2. Wiley Online Library, review,” ISPRS Journal of Photogrammetry and Remote Sensing, vol.
2016, pp. 93–102. 152, pp. 166 – 177, 2019.
[166] N. Silberman, D. Hoiem, P. Kohli, and R. Fergus, “Indoor segmen- [190] L. Gao, Y. Zhang, F. Zou, J. Shao, and J. Lai, “Unsupervised urban
tation and support inference from RGBD images,” in European scene segmentation via domain adaptation,” Neurocomputing, vol.
Conference on Computer Vision. Springer, 2012, pp. 746–760. 406, pp. 295 – 301, 2020.
[167] J. Xiao, A. Owens, and A. Torralba, “Sun3D: A database of [191] M. Paoletti, J. Haut, J. Plaza, and A. Plaza, “Deep learning
big spaces reconstructed using SFM and object labels,” in IEEE classifiers for hyperspectral imaging: A review,” ISPRS Journal of
International Conference on Computer Vision, 2013, pp. 1625–1632. Photogrammetry and Remote Sensing, vol. 158, pp. 279 – 317, 2019.
[168] S. Song, S. P. Lichtenberg, and J. Xiao, “Sun RGB-D: A RGB-D [192] J. F. Abrams, A. Vashishtha, S. T. Wong, A. Nguyen, A. Mo-
scene understanding benchmark suite,” in IEEE Conference on hamed, S. Wieser, A. Kuijper, A. Wilting, and A. Mukhopadhyay,
Computer Vision and Pattern Recognition, 2015, pp. 567–576. “Habitat-Net: Segmentation of habitat images using deep learning,”
Ecological Informatics, vol. 51, pp. 121 – 128, 2019.
[169] A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser, and
[193] M. Kerkech, A. Hafiane, and R. Canals, “Vine disease detection in
M. Nießner, “ScanNet: Richly-annotated 3D reconstructions of
UAV multispectral images using optimized image registration and
indoor scenes,” in IEEE Conference on Computer Vision and Pattern
deep learning segmentation approach,” Computers and Electronics
Recognition, 2017, pp. 5828–5839.
in Agriculture, vol. 174, p. 105446, 2020.
[170] I. Armeni, A. Sax, A. Zamir, and S. Savarese, “Joint 2D-3D- [194] Y. Song, Z. Huang, C. Shen, H. Shi, and D. A. Lange, “Deep
Semantic Data for Indoor Scene Understanding,” ArXiv e-prints, learning-based automated image segmentation for concrete petro-
Feb. 2017. graphic analysis,” Cement and Concrete Research, vol. 135, p. 106118,
[171] K. Lai, L. Bo, X. Ren, and D. Fox, “A large-scale hierarchical multi- 2020.
view RGB-D object dataset,” in IEEE International Conference on [195] N. Tajbakhsh, L. Jeyaseelan, Q. Li, J. N. Chiang, Z. Wu, and
Robotics and Automation. IEEE, 2011, pp. 1817–1824. X. Ding, “Embracing imperfect datasets: A review of deep learning
[172] C.-Y. Fu, M. Shvets, and A. C. Berg, “RetinaMask: Learning to solutions for medical image segmentation,” Medical Image Analysis,
predict masks improves state-of-the-art single-shot detection for vol. 63, p. 101693, 2020.
free,” arXiv preprint arXiv:1901.03353, 2019. [196] A. Amyar, R. Modzelewski, H. Li, and S. Ruan, “Multi-task
[173] P. O. Pinheiro, T.-Y. Lin, R. Collobert, and P. Dollár, “Learning to deep learning based CT imaging analysis for COVID-19 pneumo-
refine object segments,” in European Conference on Computer Vision. nia: Classification and segmentation,” Computers in Biology and
Springer, 2016, pp. 75–91. Medicine, vol. 126, p. 104037, 2020.
[174] H. Liu, C. Peng, C. Yu, J. Wang, X. Liu, G. Yu, and W. Jiang,
“An end-to-end network for panoptic segmentation,” in IEEE
0162-8828 (c) 2021 IEEE. Personal use is permitted, but republication/redistribution requires IEEE permission. See [Link] for more information.
Authorized licensed use limited to: Carleton University. Downloaded on May 28,2021 at 14:10:11 UTC from IEEE Xplore. Restrictions apply.
This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI 10.1109/TPAMI.2021.3059968, IEEE
Transactions on Pattern Analysis and Machine Intelligence
20 TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. ??, NO.??, ?? 2021
0162-8828 (c) 2021 IEEE. Personal use is permitted, but republication/redistribution requires IEEE permission. See [Link] for more information.
Authorized licensed use limited to: Carleton University. Downloaded on May 28,2021 at 14:10:11 UTC from IEEE Xplore. Restrictions apply.