0% found this document useful (0 votes)
9 views7 pages

Cross-View Semantic Segmentation for Robots

The document introduces a novel task called Cross-view Semantic Segmentation, aimed at enabling robots to sense their surroundings by converting first-view observations into a top-down-view semantic map. A framework named View Parsing Network (VPN) is proposed, which utilizes domain adaptation techniques to train on synthetic data and apply the model to real-world scenarios. Experimental results demonstrate the effectiveness of the VPN in understanding spatial information from various views and modalities, facilitating improved robot navigation and perception capabilities.

Uploaded by

keyfanshehefu
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
9 views7 pages

Cross-View Semantic Segmentation for Robots

The document introduces a novel task called Cross-view Semantic Segmentation, aimed at enabling robots to sense their surroundings by converting first-view observations into a top-down-view semantic map. A framework named View Parsing Network (VPN) is proposed, which utilizes domain adaptation techniques to train on synthetic data and apply the model to real-world scenarios. Experimental results demonstrate the effectiveness of the VPN in understanding spatial information from various views and modalities, facilitating improved robot navigation and perception capabilities.

Uploaded by

keyfanshehefu
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

IEEE ROBOTICS AND AUTOMATION LETTERS. PREPRINT VERSION.

ACCEPTED JUNE, 2020 1

Cross-view Semantic Segmentation for Sensing


Surroundings
Bowen Pan1,∗ , Jiankai Sun2,∗ , Ho Yin Tiga Leung2 , Alex Andonian1 , and Bolei Zhou2

Abstract—Sensing surroundings plays a crucial role in human Environment Top-down-view Semantic Map
spatial perception, as it extracts the spatial configuration of Pillar
Wall
objects as well as the free space from the observations. To Chair
facilitate the robot perception with such a surrounding sensing
arXiv:1906.03560v3 [[Link]] 18 Jun 2020

capability, we introduce a novel visual task called Cross-view Sofa


Semantic Segmentation as well as a framework named View Mobile robot Glass
Parsing Network (VPN) to address it. In the cross-view semantic Mobile robot
segmentation task, the agent is trained to parse the first-view
Chair
observations into a top-down-view semantic map indicating the First-view Observations Floor
Wall
spatial location of all the objects at pixel-level. The main issue of Cross-view
this task is that we lack the real-world annotations of top-down- Chair
Segmentation
Chair & Human
view data. To mitigate this, we train the VPN in 3D graphics Human
environment and utilize the domain adaptation technique to Cabinet
transfer it to handle real-world data. We evaluate our VPN on
both synthetic and real-world agents. The experimental results
show that our model can effectively make use of the information Fig. 1: Top-down-view semantics is predicted from the first-view real-
from different views and multi-modalities to understanding world observations in the cross-view semantic segmentation. Input
spatial information. Our further experiment on a LoCoBot robot observations from multiple angles are fused. Notice that the result in
shows that our model enables the surrounding sensing capability this figure is generated without training on real-world data.
from 2D image input. Code and demo videos can be found at
[Link]
Index Terms—Semantic Scene Understanding, Deep Learning information of the surrounding environment. Based on the
for Visual Perception, Visual Learning, Visual-Based Navigation, top-down-view semantic map we can then infer the position
Computer Vision for Other Robotic Applications coordinates and functional properties of surrounding regions
and objects.
I. I NTRODUCTION To enable machines to capture the spatial structure of the
surroundings from 2D images, we explore a new image-based
R ECENT progress in semantic understanding enables
machine perception to segment a scene precisely into
meaningful regions and objects [1], [2]. These semantic seg-
scene understanding task, Cross-View Semantic Segmentation.
Different from the standard semantic segmentation predicting
mentation techniques have benefited many automation appli- the labels of each pixel in the input image, the cross-view
cations, like autonomous driving [3]. Though the semantic semantic segmentation aims at predicting the top-down-view
segmentation network can recognize semantic content in a semantic map from a set of first-view observations (see Fig. 1).
static image, it is still far from enough to facilitate robots to The resulting top-down-view semantic map, as a 2.5D spatial
sense in an unknown environment and navigate freely there. representation of the surrounding, indicates the spatial layout
One important reason is that the parsed first-view semantic of the discrete objects such as chair and human, as well as
mask is still at pure image-level without providing any spa- the stuff classes floor and wall. Note that although there is a
tial information about the surroundings. To perceive spatial huge literature of 3D methods to reconstruct environments [4],
configuration from pure image input, an intuitive approach our method has its unique advantages. For example, robot
is to explicitly train networks to infer the top-down-view perception systems based on 3D sensors involve expensive
semantic map which directly contains the spatial configuration cost not only in sensor setup but also in computational
power. Instead, the top-down-view map from the cross-view
* indicates equal contribution. semantic segmentation can facilitate the robot to understand
Manuscript received: February 24, 2020; Revised May 13, 2020; Accepted its surroundings in a lightweight and efficient way. In many
June 4, 2020.
This paper was recommended for publication by Editor Tamim Asfour situations such as free space exploration for mobile robots
upon evaluation of the Associate Editor and Reviewers’ comments. This work where the height information is not that essential, the 2D top-
was supported by CUHK FoE Direct Grant and Facebook PyRobot Research down-view semantic map would be sufficient to provide spatial
Award.
1 B. Pan and A. Andonian are with the Computer Science and Artificial information with much less computation cost.
Intelligence Laboratory, Massachusetts Institute of Technology, USA. One challenge in cross-view semantic segmentation is the
2 J. Sun, H. Y. T. Leung and B. Zhou are with the Department of
difficulty of collecting the top-down-view semantic annota-
Information Engineering, The Chinese University of Hong Kong, Hong Kong,
China. Corresponding email: bzhou@[Link] tions. Recently, simulation environments such as House3D [5]
Digital Object Identifier (DOI): see top of this page. and CARLA [6] have been proposed for training navigation
2 IEEE ROBOTICS AND AUTOMATION LETTERS. PREPRINT VERSION. ACCEPTED JUNE, 2020

agents. In these environments, cameras can be placed at pulled from simulation engines (i.e., for visual navigation
any location in the simulated scene while the observations models [16]). Several techniques have been proposed to ad-
in multiple modalities can be extracted. Thus, we leverage dress the domain adaptation issue when models trained with
the simulation environments to acquire cross-view annotated simulated images are transferred to real scenes [17]. Rather
data. To reduce the domain gap between the synthetic scenes than working on the task of visual navigation directly, our
and the real-world scenes, we transfer the models trained in work aims at parsing the top-down-view semantic map from
the simulation environment to the real-world scenes through the first-view observations. The resulting top-down-view map
domain adaptation. will further facilitate visual navigation.
In this work, we propose a novel framework with View
Parsing Network (VPN) for cross-view semantic segmentation III. CROSS-VIEW SEMANTIC SEGMENTATION
using simulation environments and then transfer them to real- A. Problem Formulation
world environments. In VPN, a view transformer module is
The objective of cross-view semantic segmentation is as
designed to aggregate the information from multiple first-view
follows: given the first-view observations as input, the algo-
observations with different angles and different modalities. It
rithm must generate the top-down-view semantic map. The
outputs the top-down-view semantic map with a spatial layout
top-down-view semantic map is a map captured by a camera
of objects. We evaluate the proposed models on the indoor
at a certain height from the top-down view with the annotations
scene of the House3D environment [5] and the outdoor driving
of the semantic label of each pixel. The input first-view
scene of the CARLA environment [6]. Furthermore, to show
observations are a set of images with different modalities. They
the cross-view semantic task helps visual navigation, we have
are captured at N different angles by the robot’s camera (with
demonstrations of real robot.
360/N degrees apart).
Our main contributions are as follows: (1) We introduce
a novel task named cross-view semantic segmentation to
facilitate robots to flexibly sense the surrounding environment. B. Framework of the View Parsing Network
(2) We propose a framework with View Parsing Network Fig. 2 illustrates two stages of our framework. In the first
which effectively learns and aggregates features across first- stage, we propose View Parsing Network (VPN) to learn and
view observations with multiple angles and modalities. (3) We aggregate features from multiple first-view observations in the
further apply the domain adaptation technique to transferring simulation environment. In VPN, first-view observations are
our model so that it can work in real-world data while without first fed into the encoder to extract first-view feature maps. For
any extra annotations. each modality, VPN has a corresponding encoder to process
it. All of these first-view feature maps from different angles
II. RELATED WORK and different modalities are transformed and then aggregated
A. Semantic Segmentation and Semantic Mapping into one top-down-view feature map in the View Transformer
Deep learning networks for semantic segmentation [7] are Module. Then the aggregated feature map is decoded into a
designed to segment the image pixel-wise within one-view. top-down-view semantic map. Details of how to transform
Image datasets with pixel-wise annotations such as CityScapes and aggregate these first-view feature maps can be found
[3] are used for the training of semantic segmentation net- in Sec. III-D. In the second stage of our framework, we
works. There is also a huge literature about semantic mapping transfer the knowledge which VPN learns from the simulation
in robotics domain [1], [2], [8], [9], which provides the seman- environment to the real-world data. We slightly modified the
tic abstraction of the environment and a way to communicate domain adaptation algorithm proposed by [17] to fit our cross-
with robots. view semantic segmentation task and our VPN architecture.
More details of this part will be revealed in Section III-C.
Pipeline. As shown in Fig. 2, from one spatial position in a
B. Layout estimation and view synthesis
3D environment, we first sample N ×M first-view observations
Estimating layout has been an active topic of research from N angles and M modalities (here N = 6, M = 2 in Fig. 2)
(i.e. room layout estimation [10], free space estimation [11], in even angles so that all-around information is captured.
and road layout estimation [12], [13]). Most of the previous The first-view observations are encoded by M encoders for
methods use annotations of the layout or geometric constraints M corresponding modalities respectively. These CNN-based
for the estimation, while our proposed framework estimates encoders extract N × M spatial feature maps for their first-
the top-down-view map directly from the image, without the view input. Then all of these feature maps are fed into the
intermediate step of estimating the 3D structure of the scene. View Transformer Module (VTM). VTM transforms these
On the other hand, view synthesis has been explored in many view feature maps from first-view space into the top-down-
works [14], [15]. They focus on generating realistic cross- view feature space and fuses them to get one final feature map
view images while cross-view segmentation aims at parsing which already contains sufficient spatial information. Finally,
semantics across different views. we decode it to predict the top-down-view semantic map using
a convolutional decoder.
C. Learning in Simulation Environments View Transformer Module. Although the encoder-decoder
Given that current graphics simulation engines can render structure gets huge success in the classical semantic seg-
realistic scenes, recognition algorithms can be trained on data mentation area [7], our experiment (cf. Table III) shows that
PAN et al.: CROSS-VIEW SEMANTIC SEGMENTATION FOR SENSING SURROUNDINGS 3

First-view Observations
Depth inputs View Transformer Module
top-down-view semantic map
VRM
(simulation)
Depth


encoder car
VRM
full top-down-view
First-view feature map road mark
Sensor road car
unknown
Fusion Decoder
RGB inputs car
sidewalk
car
VRM

RGB Simulation environment


encoder
VRM
Real world Discriminator
(real or simulation?)

Target domain
semantic
segmentation
Real RGB

Generator road
Segmentation car
& mapping
Source Semantic car
domain VPN unknown
sidewalk
car
Synthetic Real masks with top-down-view
masks synthetic style semantic map (real)

Fig. 2: Framework of the View Parsing Network for cross-view semantic segmentation. The simulation part shows the architecture and
training scheme of our VPN, while the real-world part demonstrates the domain adaptation process for transferring VPN to the real world.

it performs poorly in the cross-view semantic segmentation spectively along the flattened dimension, and Ri models the
task. We conjecture that it is because in standard semantic relations between the ith pixel on top-down-view feature map
segmentation architecture the receptive field of the output and every pixel on first-view feature map. Here we simply use
spatial feature map is roughly aligned with the input spatial multilayer perceptron (MLP) in our view relation module R.
feature map. However, in cross-view semantic segmentation, After that, the top-down-view feature map is reshaped back to
each pixel on the top-down-view map should consider all input H ×W ×C. Notice that each first-view input has its own VRM
first-view feature maps, not just a local receptive field region. to get the top-down-view feature map t i ∈ RH×W ×C based on
After thinking about the flaws of the current semantic its own observations. To aggregate the information from all
segmentation structure, we design the View Transformer Mod- observation inputs, we fuse these top-down-view feature map
ule (VTM) to learn the dependencies across all the spatial t i by using VFM. More details of VFM and VRM will be
locations between the first-view feature map and the top-down- introduced in Sec. III-D.
view feature map. VTM will not change the shape of input
feature map, so it can be plugged into any existing encoder- C. Sim-to-real Adaptation
decoder type of network architecture for classical semantic To generalize our VPN to real-world data without the real-
segmentation. It consists of two parts: View Relation Module world ground truth, we implement the sim-to-real domain
(VRM) and View Fusion Module (VFM). The diagram at the adaptation scheme shown in Fig. 2 to narrow the gap. This
central of Figure 2 illustrates the whole process: The first- scheme contains the following pixel-level adaptation and out-
view feature map is first flattened while the channel dimension put space adaptation.
remains unchanged. Then we use a view relation module R Pixel-level adaptation. To mitigate the domain shift, we
to learn the relations between the any two pixel positions in adopt the pixel-level adaptation on the real-world inputs to
flattened first-view feature map and flattened top-down-view make them look more like the style of the simulation data.
feature map. That is: Semantic mask is an ideal mid-level representation without
ft [i] = Ri ( f [1], ..., f [ j], ..., f [HW ]), (1) texture gap while including sufficient information and it is
easy to transfer. This process can be formulated as follows:
where i, j ∈ [0, HW ) are the indices of top-down-view feature
map t ∈ RHW ×C and first-view feature map f ∈ RHW ×C re- {IS } = MReal→Synthetic (PRGB→Mask ({IR })), (2)
4 IEEE ROBOTICS AND AUTOMATION LETTERS. PREPRINT VERSION. ACCEPTED JUNE, 2020

TABLE I: Results on House3D cross-view dataset with different modalities and view numbers.

3D Geometric Baseline X-Fork(RGB) in [15] RGB VPN Semantic VPN Depth VPN
Networks PA mIoU PA mIoU PA mIoU PA mIoU PA mIoU
1-view model 31.3% 2.4% 38.0% 1.5% 55.8% 6.5% 59.6% 13.2% 56.9% 7.6%
2-view model 46.8% 7.2% 40.0% 1.9% 70.1% 14.8% 75.7% 25.9% 70.2% 15.6%
4-view model 63.2% 22.8% 39.5% 2.0% 80.3% 27.2% 85.0% 40.6% 77.3% 22.0%
8-view model 67.6% 27.1% 43.9% 1.8% 81.2% 28.5% 84.7% 41.0% 82.1% 29.9%

TABLE II: Results of cross-modality learning for VPN. Here we TABLE III: Ablation study of View Transformer Module.
compare the results with the inputs from 4 views.
Modality VPN w/o VTM VPN
Method Pixel Accuracy mIoU 1-view Pix. Acc. mIoU Pix. Acc. mIoU
RGB VPN 80.3% 27.2% RGB 53.9% 6.3% 55.8% 6.5%
Depth VPN 77.3% 22.0% Depth 55.7% 6.5% 56.9% 7.6%
Semantic VPN 85.0% 40.6% Semantic 57.4% 10.0% 59.6% 13.2%
R+D (late fusion) 81.2% 27.3% 8-view Pix. Acc. mIoU Pix. Acc. mIoU
R-D VPN 82.8% 31.2%
RGB 60.5% 8.7% 81.2% 28.5%
D+S (late fusion) 83.5% 33.9%
Depth 43.8% 2.5% 82.1% 29.9%
D-S VPN 86.3% 43.2%
Semantic 47.6% 6.5% 84.7% 41.0%
S+R (late fusion) 84.3% 35.7%
S-R VPN 85.1% 42.3%
D+S+R (late fusion) 84.3% 29.4%
D-S-R VPN 86.2% 43.6% last Residual Block and the Average Pool layer so that the
resolution of the encoding feature map remains large, which
better preserves the details of the view. We employ the pyramid
where IR , IS are the real RGB image and synthetic-style se- pooling module used in [7] as the decoder.
mantic mask respectively, PRGB→Mask is the existing semantic View Transformer Module. For each view relation module,
segmentation model which parses the real-world RGB into we simply use the two-layer MLP. We choose this because
semantic mask, and MReal→Synthetic is the semantic category two-layer MLP doesn’t bring too much extra computation so
mapping process where we construct the concept mappings that we can keep our model following the lightweight-and-
between the real world and the simulation environment. efficient rationale. Input and output dimensions of the VRM
Output space adaptation. Beyond the pixel-level transfer are both HI WI , where HI and WI are respectively the height
on input data, we also devise an adversarial training scheme and width of the intermediate feature map. As for the view
in structured output space based on the method proposed fusion module, we just add all the features up to keep the
in [17]. Here the generator G is a view parsing network shape consistent.
generating the top-down-view prediction P, which is initialized Sim-to-real. For the generator G , we use the architecture
by the weights of a VPN trained on the semantic data in the of the 4-view VPN. For the discriminator D, we adopt the
simulation environment as we illustrated before. During the same architecture in [17]. It has 5 convolution layers, each
training phase, we first forward a group of input images from of which is followed by a leaky ReLU with the parameter
the source domain {Is } to G and optimize it with a normal 0.2 (except the last layer). We use HRNet [18] pretrained on
segmentation loss Lseg . Then we use G to extract the feature CityScapes dataset [3] to extract the semantic mask from real-
map Fi (after the softmax layer) of the images from the target world images.
domain {It } and use discriminator to distinguish whether Ft
is from the source domain. The loss function to optimize G IV. EXPERIMENTS
can be written as follows: We first go through the overview of the cross-view segmen-
tation datasets in Section IV-A. Then we show the performance
L ({Is }, {It }) = Lseg ({Is }) + λadv Ladv ({It }), (3)
of VPN on synthetic data of the House3D and CARLA
where Lseg is the cross-entropy loss for semantic segmenta- environment in Section IV-B. Finally in Section IV-C, we
tion, Ladv is designed to train the G and fool the discriminator demonstrate the real-world performance of our VPN which
D. The loss function for the discriminator Ld is a cross- is trained in the simulation environment.
entropy loss for binary source & target classification.
A. Benchmarks
D. Network configuration Here we introduce two synthetic cross-view datasets,
View encoder and decoder. To balance efficiency and per- House3D cross-view dataset and Carla cross-view dataset, and
formance, we use ResNet-18 as the encoder. We remove the one real-world cross-view dataset, nuScenes dataset.
PAN et al.: CROSS-VIEW SEMANTIC SEGMENTATION FOR SENSING SURROUNDINGS 5

House3D cross-view dataset. Each data pair contains 8 first- the late-fusion baseline to compare with our multi-modalities
view input images captured from 8 different orientations with VPN, which simply averages the softmax outputs of each
45 degrees apart. Additionally, each data pair comes with single-modality VPN to obtain the final results. We find that
the top-down-view semantic mask captured in the ceiling- the Depth-Semantic VPN achieves the best performance and
level height. To be complete, we store the input image with makes a great improvement. This may be because semantic
multiple modalities including the RGB images, depth maps, mask and depth map are two complementary information.
and semantic masks. The training set contains 143k data pairs However, the Semantic-RGB combination does not bring too
from 342 scenes while the validation set contains 20k data much improvement. The reason can be that, for this cross-view
pairs from 68 scenes. semantic segmentation task, semantic input contains most of
NuScenes dataset. Each data sample contains in the useful information in the RGB.
NuScenes[19] first-view RGB images from 6 directions Importance of View Transformer Module. We further eval-
(Front, Front-right, Back-right, Back, Back-left, Front-left) in uate our model in Table III to show the importance of the
different modalities. We select 919 data samples without the view transformer module. The baseline network is a classic
top-down-view mask for unsupervised training and 515 data encoder-decoder architecture used in the standard semantic
samples with the binary top-down-view mask for evaluation. segmentation, in which the encoder and the decoder are the
CARLA cross-view dataset. To build the synthetic source same as our VPN. It simply sums up the feature maps from
domain dataset, we extract 28, 000 data pairs with top-down- different views and then feeds it to the decoder. Our VPN
view annotations and different input modalities from 14 driv- easily outperforms the baseline and, in some multi-view cases,
ing episodes in CARLA. Each data pair contains 6 first-view the baseline model does even worse than single-view one due
input image sets captured from the same 6 directions. to the bad fusion strategy.
Comparing with baseline. Table I shows that our VPN can
easily outperform the 3D geometric method. 3D Geometric
B. Evaluation
method is very easy to fail when there are obstacles. In
We present VPN performances on the synthetic data of Fig. 4, we can see that the 3D geometric method is unable
House3D cross-view and CARLA cross-view datasets. to reconstruct the objects which can not be directly observed,
Metrics. We report the results of cross-view semantic seg- even after filling the holes, such as the desk behind the chairs
mentation using two commonly used metrics in semantic shown in the figure. As for X-Fork, we can see that the
segmentation: P IXEL ACCURACY (PA) which characterizes original generator performs badly in our cross-view semantic
the proportion of correctly classified pixels, and M EAN I O U segmentation task. This is because X-Fork doesn’t have a
( M I O U) which indicates the intersection-and-union between necessary module to transform the first-view feature map into
the predicted and ground truth pixels. the top-down-view space. The ablation study in Table III
Baselines. Two methods are included as the comparison shows a similar issue that there is a significant performance
baselines: (1) 3D geometric method. With the observed depth drop when VPN doesn’t contain the VTM.
and RGB images, we can reconstruct the 3D points cloud
with the voxel-level semantic label. (2) Cross-view synthesis. C. Results of sim-to-real adaptation
We also compare with the architecture used in cross-view After we train and test our VPNs in the simulation environ-
image synthesis literature [15], which adopts a conditional ment, we transfer our model to the real-world data. We first
GAN called X-Fork to generate aerial images from street-view train a 6-view semantic VPN model on the predicted semantic
images. masks in CARLA simulator and then transfer it to nuScenes
1) Results of VPNs: We present the results of our VPN dataset by using an unsupervised domain adaptation process
for cross-view semantic segmentation in House3D, including as depicted in Section III-C. We provide the qualitative results
the ones of single-modality and multi-modalities VPN respec- in Fig. 3, from which we can see that our VPN can roughly
tively. To better evaluate our VPN, we impose an upper bound segment various road shapes like crossroads and also sketch
that we perform segmentation using top-view RGB images the relative locations of surrounding objects such as cars and
directly as inputs, where we get the performance of 91.4% buildings. As shown in Table IV, we evaluate the quantitative
pixel acc. and 41.2% mIoU. We also show the comparison results of real-world performance by using binary drivable-
with the geometric baseline and the ablation study of View area ground truth.
Transformer Module.
Single-modality VPN. We show the House3D results of TABLE IV: Results in real world.
single-modality VPN with different modalities and different Method Pix. Acc. Mean Acc. mIoU
numbers of views in Table I. We can see that as VPN receives
more views, the segmentation results improve rapidly. We also Before Adaptation 72.6% 61.4% 28.0%
plot some qualitative results by our VPNs in Fig. 4. On Carla After Adaptation 78.8% 65.2% 31.9%
dataset, we achieve the performance of 84.7% pixel acc. and
33.2% mIoU with a 6-view RGB-input model.
Multi-modalities VPN. We demonstrate the results of multi- V. EXPLORATION WITH TOP-DOWN-VIEW MAP
modalities VPN in Table II to show that our VPN can effec- When exploring an unknown space our humans head to the
tively synthesize information from multiple modalities. We set regions which they have not visited. This intuition reflects that
6 IEEE ROBOTICS AND AUTOMATION LETTERS. PREPRINT VERSION. ACCEPTED JUNE, 2020

Real observations Predicted Ground-truth


Before adaptation After adaptation binary mask binary mask

Fig. 3: Qualitative results of sim-to-real adaptation. The results of source prediction before and after domain adaptation, drivable area
prediciton after adaptation and the groud-truth drivable area map.

First-view Projected 3D Geometric View Parsing Algorithm 1 Exploration decision policy at time t
observations points method Network Ground truth Input: A top-down-view free-space map Tt and a state map
St at time step t, where Tt , St ∈ {0, 1}L×L .
Output: Policy action at , where at ∈ {Forward, Back, Left-
forward, Right-forward, Done}.
1: Ut ← Tt ¬St ; at ← Done; ds ← +∞
T

Fig. 4: Qualitative results of 3D geometric method and our VPN. 2: Dt ← computeDistMap(Ut )


Considering that the geometric method requires semantic mask and 3: for a in {Forward, Back, Left-forward, Right-forward} do
depth map, we use the 4-view Depth-Semantic VPN to predict the
4: d =execute(a)
top-down-view semantic map to fairly compare these two methods.
5: if ds > d then
6: ds ← d; at ← a
exploration requires the agent to identify free space as well as 7: end if
remember which areas it has not visited yet. To achieve this 8: end for
goal, we make the agent able to identify the free space by 9: return at
training it to predict the top-down-view free-space map. computeDistMap(): Compute the shortest distance of each map
Top-down-view free-space map. We train the VPN to predict pixel to the unvisited free-space region.
top-down-view free-space map. Different from the semantic execute(): Return the shortest distance of the pixel to which
map, free-space map has only two categories, obstacle and the agent transit if execute the action a.
free space, which are denoted by 0 1 respectively.
State map. Due to the ideal assumption made above, by
memorizing the previous actions it has executed, the agent exploration trajectories given the first-view observations. The
can easily build the state map which contains the information trajectories are generated by the baseline above with a top-
of the already-visited positions. We label the unvisited pixels down-view ground truth map. Network inputs are 4 first-view
as 0 and the already-visited pixels as 1 on the state map. depth images. We also input the state map to indicate the
Exploration algorithm. We detail the navigation policy already-visited area. We extract 729 trajectories for the training
decision algorithm in Algorithm 1. At each time step t, we set and 121 trajectories for the validation set to train the
make the action at and update the agent with the next top- navigation agent. Each trajectory contains 150 states which
down-view free-space map Tt+1 and state map St+1 . In both the are all labeled with expert policy.
top-down-view free-space map and the state map, we assume
TABLE V: Comparison on exploration.
that the agent is always at the center of the map.
Method Coverage Area
A. Result and comparison Random walk 260.3 ± 82.7
To demonstrate that VPN can help navigation, we com- IL w/o top-down-view 443.8 ± 340.6
pare it with the following baselines for exploration. Random
walk: Random walk agent randomly chooses one action from Top-down-view navigation 673.8 ± 349.8
Forward, Back, Right-forward and Left-forward, at each time Top-down-view navigation with GT 1070.8±326.2
step. Top-down-view navigation with ground truth (GT): By
planning on the ground truth top-down-view free-space map We run the algorithm directly on our predicted top-down-
with Algorithm 1, we can obtain the upper-bound performance view map. For testing all the methods, we start the episode
of our method. The difference is that in our case the top- by initializing the state maps from zero, indicating that all
down-view free-space map is predicted by VPN, rather than free space is yet to be visited. Coverage Area is defined to
the ground truth. Imitation learning (IL) without top-down- measure exploration performance. We randomly choose 100
view: A reactive CNN network learns to imitate the expert starting points on a scene map. For each starting point, we
PAN et al.: CROSS-VIEW SEMANTIC SEGMENTATION FOR SENSING SURROUNDINGS 7

Start point End point

Coverage area: 326 Coverage area: 632 Coverage area: 991 Coverage area: 1163
on the experimental results, we demonstrate that VPN can
be applied to mobile robots to facilitate the surrounding
awareness through a lightweight and efficient top-down-view
(a) Random walk (b) IL w/o top-view (c) Top-view navigation (d) Top-view navigation semantic map. In many situations where the height information
with GT
of objects is not essential, VPN could be a good alternative
Fig. 5: Examples of the surrounding exploration. Start point and end compared to the traditional 3D-based methods which are costly
point are marked as red point and green point respectively on both data memory and computation.

(a) (b) (c) R EFERENCES


[1] I. Kostavelis and A. Gasteratos, “Semantic mapping for mobile robotics
target target tasks: A survey,” Robotics and Autonomous Systems, vol. 66, pp. 86–
centroid 103, 2015.
[2] N. Sünderhauf, F. Dayoub, S. McMahon, B. Talbot, R. Schulz, P. Corke,
G. Wyeth, B. Upcroft, and M. Milford, “Place categorization and
semantic mapping on a mobile robot,” in 2016 IEEE international
conference on robotics and automation (ICRA). IEEE, 2016, pp. 5729–
start point 5736.
[3] M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benen-
“Go to the bench” son, U. Franke, S. Roth, and B. Schiele, “The cityscapes dataset for
semantic urban scene understanding,” in Proc. CVPR, 2016.
LoCoBot Test environment Top-down-view semantic map [4] A. Dai, A. X. Chang, M. Savva, M. Halber, T. A. Funkhouser, and
M. Nießner, “Scannet: Richly-annotated 3d reconstructions of indoor
scenes.” in Proc. CVPR, vol. 2, 2017, p. 10.
Fig. 6: Our experiments are conducted on a LoCoBot mobile robot [5] Y. Wu, Y. Wu, G. Gkioxari, and Y. Tian, “Building generalizable
in the PyRobot platform. (a) We show a picture of the LoCoBot agents with a realistic and rich 3d environment,” arXiv preprint
robot. (b) We show our test environment and the target specified by arXiv:1801.02209, 2018.
a semantic token (e.g. bench). (c) We show the process that parses [6] A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, and V. Koltun,
the instruction and calculate the target coordinates. “CARLA: An open urban driving simulator,” in Proceedings of the 1st
Annual Conference on Robot Learning, 2017, pp. 1–16.
[7] H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia, “Pyramid scene parsing
network,” in Proc. CVPR, 2017.
let the agent explore the space for 300 steps and compute the [8] Y. Katsumata, A. Taniguchi, Y. Hagiwara, and T. Taniguchi, “Semantic
coverage area. Then final results are obtained by averaging mapping based on spatial concepts for grounding words related to places
in daily environments,” Frontiers in Robotics and AI, vol. 6, p. 31, 2019.
the coverage area of these 100 episodes. Table V plots the [9] K. Zheng, A. Pronobis, and R. P. Rao, “Learning graph-structured sum-
exploration result for different methods and Fig. 5 shows product networks for probabilistic semantic maps,” in Thirty-Second
some sample trajectories. We can see that equipped with the AAAI Conference on Artificial Intelligence, 2018.
[10] C. Zou, A. Colburn, Q. Shan, and D. Hoiem, “Layoutnet: Reconstructing
predicted top-down-view map from our VPN, the agent can the 3d room layout from a single rgb image,” in Proc. CVPR, 2018.
efficiently explore the environment. [11] V. Hedau, D. Hoiem, and D. Forsyth, “Recovering free space of indoor
scenes from a single image,” in Proc. CVPR, 2012.
[12] S. Schulter, M. Zhai, N. Jacobs, and M. Chandraker, “Learning to
VI. REAL ROBOT EXPERIMENT look around objects for top-view representations of outdoor scenes,” in
Proceedings of the European Conference on Computer Vision (ECCV),
To verify the performance of our model in the real-world 2018, pp. 787–802.
robotic environment, we conduct a semantic navigation exper- [13] Z. Wang, B. Liu, S. Schulter, and M. Chandraker, “A parametric top-
iment by using a LoCoBot mobile robot [20] (cf. Fig. 6a). view representation of complex road scenes,” in Proceedings of the
IEEE Conference on Computer Vision and Pattern Recognition, 2019,
In this task, the robot is required to identify and reach the pp. 10 325–10 333.
target specified by a semantic token. For instance, given the [14] M. Zhai, Z. Bessinger, S. Workman, and N. Jacobs, “Predicting ground-
instruction “go to the bench”, the robot has to move to the level scene layout from aerial imagery,” in Proc. CVPR, vol. 3, 2017.
[15] K. Regmi and A. Borji, “Cross-view image synthesis using conditional
target which is shown in Fig. 6b. Similar settings are also gans,” in Proc. CVPR, 2018, pp. 3501–3510.
used in [16]. At the initial location, the robot takes 8 RGB [16] W. B. Shen, D. Xu, Y. Zhu, L. J. Guibas, L. Fei-Fei, and S. Savarese,
images (45 degrees apart) using its head camera. Then it uses “Situational fusion of visual representation for visual navigation,” in
Proceedings of the IEEE International Conference on Computer Vision,
the existing semantic segmentation technique to obtain the 2019, pp. 2881–2890.
semantic mask of each RGB image. After that, it predicts [17] Y.-H. Tsai, W.-C. Hung, S. Schulter, K. Sohn, M.-H. Yang, and
the top-down-view semantic map with our VPN. Finally, it M. Chandraker, “Learning to adapt structured output space for semantic
segmentation,” in Proceedings of the IEEE Conference on Computer
parses the instruction and calculates the centroid coordinates Vision and Pattern Recognition, 2018, pp. 7472–7481.
of all “bench” pixels (cf. Fig. 6c). Our real robot experiment [18] K. Sun, Y. Zhao, B. Jiang, T. Cheng, B. Xiao, D. Liu, Y. Mu, X. Wang,
shows that though the model is trained in a simulator it W. Liu, and J. Wang, “High-resolution representations for labeling pixels
and regions,” arXiv preprint arXiv:1904.04514, 2019.
exhibits reasonable robustness when we randomly set the [19] H. Caesar, V. Bankiti, A. H. Lang, S. Vora, V. E. Liong, Q. Xu, A. Kr-
initial location and change the layout of surrounding objects. ishnan, Y. Pan, G. Baldan, and O. Beijbom, “nuscenes: A multimodal
dataset for autonomous driving,” arXiv preprint arXiv:1903.11027, 2019.
[20] A. Murali, T. Chen, K. V. Alwala, D. Gandhi, L. Pinto, S. Gupta, and
VII. CONCLUSION A. Gupta, “Pyrobot: An open-source robotics framework for research
and benchmarking,” arXiv preprint arXiv:1906.08236, 2019.
In this work, we propose the cross-view semantic segmen-
tation task to sense the environment and a neural architecture
design View Parsing Network (VPN) to address that. Based

You might also like