Urban Building Classification (UBC) – A Dataset for Individual Building
Detection and Classification from Satellite Imagery
Xingliang Huang1,2 *, Libo Ren1,2 *, Chenglong Liu1,2 , Yixuan Wang3,5 , Hongfeng Yu1,2
Michael Schmitt5 , Ronny Hänsch4 , Xian Sun1,2†, Hai Huang5†, Helmut Mayer5
1
Aerospace Information Research Institute, Chinese Academy of Sciences, China
2
University of Chinese Academy of Sciences, China 3 Technische Universität München, Germany
4
German Aerospace Center, Germany 5 Universität der Bundeswehr München, Germany
{huangxingliang20, renlibo20, liuchenglong20}@[Link],
ge35qej@[Link], hfyu@[Link], [Link]@[Link], sunxian@[Link],
{[Link], [Link], [Link]}@[Link]
Abstract 1. Introduction
We present a dataset for building detection and classifi- Buildings are one of the most important components of
cation from very high-resolution satellite imagery with the urban areas. The investigation of buildings plays an essen-
focus on object-level interpretation of individual buildings. tial role in urban planning, city administration, emergency
It is meant to provide not only a flexible test platform for ob- management, tourism, etc. With the advent of deep learn-
ject detection algorithms but also a solid basis for the com- ing techniques, the performance of building detection and
parison of city morphologies and the investigation of ur- classification in remote sensing data has been significantly
ban planning. In most current open datasets, buildings are improved. One key driver are the ever increasing remote
treated either as a class of landcover in the form of masks or sensing datasets [14], of which the building-related datasets
as simple objects defined by separate contours (footprints). are summarized in Section 2.
Our dataset, instead, represents individual buildings using We propose a novel dataset with a specific focus on the
in-depth object-level descriptions concerning geometry as object-level interpretation of individual buildings, which are
well as functionality. Buildings are treated as objects with represented with in-depth descriptions concerning both ge-
individual ID and boundary. Adjacent building blocks are ometry as well as functionality. Based on this, the dataset
also separated according to house numbers making a subse- provides the possibility to quantitatively compare different
quent high-level classification of individual buildings possi- cities with regard to statistics and morphology. As shown in
ble. The buildings are classified into predefined roof types, Figure 1, Beijing’s buildings (rows 1 and 2) typically show
such as flat, gable and hipped roof as well as functional pur- a neat arrangement and a modern steel/concrete style, while
poses, i.e., residential, commercial, industrial, public, and buildings in Munich (rows 3 and 4) are mostly distributed
their sub-classes, e.g., single-family house, office building along the historical streets and are of lower height.
and school. In the first version of the dataset we provide Although we can readily identify geometrical informa-
selected urban areas from two cities: Beijing in China and tion of buildings, e.g., contour and roof shape, in satellite
Munich in Germany. It, therefore, (1) allows to verify algo- imagery, it is usually substantially difficult to accurately
rithms that are not only valid for specific regions but also identify the function of buildings. So we label the function
work robustly in spite of the diversity of cities on different using additional map information.
continents with various land forms and styles of architecture Data sources for cities such as OSM (OpenStreetMap)
and at the same time (2) provides the possibility to quanti- and Google Maps provide the basis for a large number
tatively compare the statistics and morphology of different of statistics on buildings worldwide. But different data
cities. It is planned to extend the dataset by a continuous sources have different definitions for building attributes,
integration of various urban areas worldwide. so it is difficult to combine multiple data sources to
automate the labeling of buildings for remote sensing
* Equal contribution. images. Moreover, many building attributes are missing
† Corresponding author. in these data sources and they do not provide up-to-date,
1413
grained categories concerning (1) building geometry
with 25 roof types as well as (2) two levels of func-
tional purposes with the five main classes: residential,
commercial, industrial, public and other, and 36 sub-
classes, e.g., single-family house, office building and
school.
• Flexible Structure: This dataset consists of explic-
itly separated subsets corresponding to different cities.
The subsets can be combined to train general detectors
and to test their robustness or employed separately to
investigate and compare characteristics of different ur-
ban areas. The categories are given on various levels of
both roof type and function. They can be used to define
different setups for competition with varying amounts
of data. Please note that it is planned to extend this
dataset. We expect that along with the increasing size
of the urban areas and the extended coverage of differ-
ent classes, the advantage of our multi-level category
definition will become even more apparent.
2. Related Work
Suitable datasets are critical for the development and
evaluation of object detection and classification algorithms,
especially deep neural network models. An increasing num-
ber of remote sensing datasets has been introduced in recent
Image Image + GT GT
years with various data sources as well as target objects.
The DOTA dataset [8] contains over 1.7 million instances of
Figure 1. Example UBC data of Beijing (rows 1 and 2) and Mu- 18 classes with oriented bounding box annotations collected
nich (rows 3 and 4) with input images (left), ground truth (GT, from 11,268 aerial images. It has, thus, greatly contributed
right), and overlaid views (center). Colors of ground truth indicate to the development of detection algorithms for rotated ob-
various classes of roof type. jects in remote sensing data. The FAIR1M dataset [17] also
consists of over one million instances of fine-grained ob-
jects in high-resolution remote sensing imagery, providing
complete and accurate statistics of building locations and the community data with 5 categories and 37 sub-categories
attributes. Further information on urban buildings can be of ground targets. In the ISPRS Urban Modelling and Se-
provided by high resolution and low-cost remote sensing mantic Labeling Benchmark [13] multispectral imagery and
images. To learn large-scale statistics of urban buildings airborne laserscanner data of Vaihingen and Potsdam, Ger-
in remote sensing images with deep learning techniques, many, as well as Toronto, Canada, are meant for the de-
a building dataset with consistent attribute standards and tection of urban objects, such as buildings, roads, trees,
global diversity is urgently needed. as well as for 3D building reconstruction. The TorontoC-
ity dataset [21] provides aerial imagery with about 10 cm
Particularly, the main contributions of our work are: ground resolution depicting around 400 thousands build-
• In-depth annotation: Instead of visual annotations, ings. SpaceNet consists of a series of remote sensing
in this dataset boundaries as well as functions of in- datasets with various basic data including multispectral im-
dividual buildings are labeled according to OSM and agery and synthetic aperture radar (SAR) data and purposes
Google Maps. This means that adjacent buildings in a such as building and road network extraction as well as clas-
building block as for instance shown in Figure 4 can sification. The SpaceNet 2 Challenge [20] contains 302,701
be correctly separated into different objects. The func- building footprints in 24,586 scenes, while SpaceNet 6 [16]
tions of buildings can often be more accurately derived is a multi-sensor all-weather mapping dataset, consisting
from their attributes in the maps than from images. of both optical and SAR imagery, aiming to map build-
ing footprints using multi-modal data. SpaceNet 7 Multi-
• Fine-grained categories: We provide novel fine- Temporal Urban Development Challenge [19] is based on
1414
Dataset Classes Instance Quantity Modality Resolution
SpaceNet 2 [20] 1 500k RGB MSI 0.3 m
SpaceNet 6 [16] 1 4.8k RGB SAR 0.5 m
Toronto City [21] 1 400k RGB 0.05-0.1 m
GaoFen-3 Building [24] 1 (semantic segmentation) RGB SAR 1m
INRIA [11] 2 (semantic segmentation) RGB 0.1-0.3 m
DSTL [15] 5 2k RGB MSI 0.3 m
SemCity Toulouse [12] 6 9k PAN 0.5 m
UBC 61 41k RGB 0.5-0.8 m
Table 1. Comparison of building datasets
imagery collected by Planet Labs’ Dove Satellites and con- tained 4-band images (red, green, blue and near-infrared)
tains around 500,000 buildings tracked over time. at 0.5 m and 0.8 m by pan-sharpening. In the current ver-
The INRIA aerial image labeling benchmark [11] con- sion of the dataset only the three visible bands, i.e., RGB,
sists of precisely registered cadastral records as well as are used. Possible multispectral as well as SAR data are
15 cm or 30 cm orthorectified imagery. It considers the scheduled for the further extension of the dataset (cf. Sec-
two classes building and non-building, i.e., trees and roads. tion 5). The whole dataset consists of 800 tiles with 600
The DSTL Satellite Imagery Feature Detection dataset [15] × 600 pixels and 200 pixels overlap for adjacent tiles. The
provides multispectral satellite imagery in RGB as well as information about data coverage and instances is shown in
16-bands with a resolution of 0.3 m. It employs a coarse Figure 2 as well as Table 2.
classification of buildings, including residential and non-
residential building, fuel storage facility, and fortified build-
ing. It comprises about two thousand building instances.
The SemCity Toulouse benchmark [12] focuses on build-
ing instance segmentation. It provides multi-class seman-
tic segmentation annotation including residential and office
building, shop, department store, discount store, shopping
center, as well as industrial building.
Related datasets include also the GaoFen-3 SAR
dataset [24] for semantic segmentation of buildings. It is
acquired in spotlight (SL) mode with high-resolution (1 m)
(a) Beijing (b) Munich
and a wide swath (10 km) and covers urban as well as rural
areas in, e.g., Hongkong, Berlin, and Shanghai. A com- Figure 2. Data coverage in Beijing and Munich: Tiles from Super-
prehensive comparison of our dataset with selected remote view (blue) and Gaofen-2 (red)
sensing datasets containing buildings is given in Table 1.
3. Dataset City Source Coverage (km2 ) Building Instances
3.1. Satellite Imagery SuperView 20 10,105
Beijing
Gaofen-2 12.8 4,824
The SuperView (or “GaoJing” in Chinese) satellites are SuperView 20 19,352
commercial very high-resolution earth observation space- Munich
Gaofen-2 12.8 7305
crafts operated by Beijing Space View Tech Co Ltd. They
are equipped with sensors collecting both panchromatic
(0.5 m, Ground Sampling Distance – GSD) and multispec- Table 2. Data coverage and building instances for Beijing and
tral (2 m) GSD imagery with a maximum scene size of 60 Munich
km × 70 km [5]. The Gaofen-2 high-resolution imaging
satellites from the China National Space Administration
3.2. Definition of Classes
(CNSA) are capable of collecting images with a GSD of
0.81 m in the panchromatic and 3.24 m in the multispectral Current datasets focus on the location, footprint extrac-
bands with a swath width of 45 km [4]. Here, we chose tion and segmentation of buildings. Classification, espe-
panchromatic and multispectral data from SuperView and cially fine-grained classification of buildings is rare yet. The
Gaofen-2 for urban areas of Beijing and Munich and ob- latter is of great interest for applications in city planning and
1415
urban development analysis. We are aware that the defini- For the building functions, as shown in Table 4, we de-
tion of classes has a huge influence on the performance of fine five coarse classes, i.e., “residential”, “commercial”,
the classification. To derive plausible classes which con- “industrial”, “public” and “other” which can be split into 36
form to common understanding as well as to what can be fine-grained classes. Also inside each coarse class there is
seen from remote sensing data, we first summarize the cat- a fine-grained class named “other”, e.g., “public other”. In
egories of the main popular data sources and standards: contrast to the “other” in the coarse classes, it is employed
OpenStreetMap [23], CityGML (partially open) as well as to label the instances which are not listed in the fine-grained
Google Maps and derive our own classes taking the char- classes, but still can be determined as belonging to one of
acteristics and limits of satellite imagery into account. One the coarse classes. This happens quite often as the function
obvious advantage is that the existing geometrical and func- classes hardly cover all possibilities.
tional attributes of buildings in the above mentioned data
sources can be easily mapped to our classes. This makes Coarse Fine-grained
the current manual annotation easier and will allow for a single-family house,
(semi-) automatic annotation (cf. Section 5) in the future. multi-family house,
residential row house, apartment high,
The definition the roof classes is given in Table 3. For the
apartment block, villa, garage,
roof type, we use 25 fine-grained classes (including “other”)
residential other
based on the geometry of the roof. The fine classes are
grouped into nine coarse classes based on geometrical sim- office building, retail and mall,
ilarities. The coarse classes are especially useful when the commercial hotel, parking house,
instances of the fine-grained classes are not enough for a restaurant, commercial other
stable training. power plant, warehouse,
industrial manufacturing, water treatment,
industrial other
Coarse Fine-grained administration, gas station,
flat roof education, stadium, sports hall,
flat flat roof HVAC* transportation, theatre,
flat roof complex public fire station, police station,
shed shed roof military, church, mosque,
gable roof temple, airport building,
gable roof asymm* hangar, public other
gable other other
gambrel roof
butterfly roof*
Table 4. Building functionality classes
row roof shed*
row row roof gable*
row roof arched*
3.3. Annotation
multiple eave roof*
hipped roof v1 The footprint of buildings are annotated with polygons.
hipped roof v2 The footprints from OSM are used as basis and are man-
hipped
half hipped roof* ually refined/corrected according to the input images. If
mansard roof available, the roof type and function information is taken
pinnacle roof from OSM as well as Google Maps and is mapped to the
arched roof predefined classes. Figure 4 shows one example.
arched
half arched roof* We are aware of the heterogeneous quality of OSM [6]
dome* and noticed that the availability and quality of OSM data
revolved cone* for Munich is substantially better than for Beijing. For the
cupola* selected areas in relation to the corresponding final ground
freeshape surface* truth, the footprints of 89.7% of the buildings have been
freeshape
freeshape poly* provided in the OSM dataset. 39.9% of the buildings in Mu-
other other nich have the function attribute and 22.2% roof type infor-
mation. Yet, in Beijing, only for 27.6% of the building foot-
Table 3. Roof type classes. * indicates the fine-grained classes prints have been included and not all of them are correctly
with few instances, which are merged into the class “other” in the located. Only 4.2% of the buildings have function informa-
following experiments tion and there is no roof type information in the OSM data
1416
for Beijing. A summary is given in Figure 3. Therefore, 3.5. Dataset Splits
we have referred to information from Google Maps to man-
Our dataset is designed to be used as a whole set as well
ually improve the annotation of roof types and functions.
as two separate subsets: Beijing and Munich. Each sub-
I.e., the annotators visually check the roof shape as well as
set contains the same amount of data: 400 tiles of satellite
(heterogeneous) labels of buildings and manually interpret
images with a size of 600×600 pixels selected from their
their roof types and functions according to the predefined
urban areas. Please note that we also ensured that the data
categories in UBC. For difficult instances, particularly the
partitions of SuperView (80%) and Gaofen-2 (20%) on each
buildings with multiple labels (cf. Section 3.4), additional
subset is constant, so that this ratio is also kept in the whole
Google Street View data (for Munich only) are optionally
set. In the experiments, the dataset is divided (on a random
employed to assistant the (visual) interpretation. To ensure
sampling basis) into training, validation as well as test sets
the quality of annotation, the results from annotators are
with partitions of 7:2:1.
examined by more experienced inspectors in two rounds:
One complete check and one random check. Controversial
labels are determined by consensus of multiple annotators
and inspectors. The consistency of annotations are also en-
sured for the buildings in overlapping areas. As shown in
Table 2, the UBC dataset provides altogether 41,586 build-
ing instances, including 14,929 for Beijing and 26,657 for
Munich, with complete footprints and annotations for both
the roof type as well as the function. Additionally, 4,790 for
Munich and 210 for Beijing building instances have multi-
label annotations.
(a) (b)
3.4. Multi-label Annotation
In a realistic urban scenario, individual buildings often
do not have a single function. For example, for many tall
buildings and large structures in the city centers, the lower
floors are typically shopping areas, while the upper floors
serve as office space or for habitation. It is, therefore, not
reasonable to label these buildings with just one function
class. To make the annotation of the dataset more accurate,
we, thus, introduce multi-label annotation. Yet, in order not
to add too much extra complexity to the annotation process,
(c) (d)
we restrict the multi-label annotation to only two reasonable
situations: (1) “Apartment block” as well as “commercial Figure 4. An annotation example: Input image (a), building foot-
other” and (2) “apartment high”, “office building” as well prints (b, green polygons), roof types (c) and functions (d, coarse
as “retail and mall”. For the latter, multiple choice selection classes)
of up to three labels at the same time is allowed.
4. Experiments
High building densities in urban areas, various sizes,
fine-grained classes as well as multi-label annotations pose
great challenges to instance segmentation methods on the
UBC dataset. To quantitatively measure the state-of-the-art
under these circumstances, we propose two instance seg-
mentation tasks as well as corresponding evaluation metrics
and we evaluate general methods on the dataset. In the first
task, we predict roof types for all buildings and use pixel-
level masks to localize them. Fine- and coarse-grained roof
Figure 3. Comparison of the availability of OSM building labels are set as two sub-tasks . The class of each instance
footprints and attributes for Munich and Beijing. The missing in the second task is defined by the building function. Due
(“empty”) data are supplemented by manual annotation. to the difficulty of reasoning about the function from visual
features alone, we only conducted baseline experiments on
1417
coarse function classes. for the instance segmentation algorithms presented in Sec-
tion 4.1. The results in Table 5 show that the Cascade Mask
4.1. Baseline Models R-CNN model outperforms the other models concerning av-
We selected Mask-RCNN [9], Cascade Mask RCNN [1], erage precision with an IoU of at least 0.5 (AP50 ) and the
SOLOv2 [22] and QueryInst [7] as our baseline mod- mean average precision (mAP ) calculated from all differ-
els. The backbone network for all models is ResNet-50- ent IoUs. The cascade architecture has shown its advan-
FPN. The implementation of these methods is based on tage in the detection of densely distributed buildings with
the MMDetection library [2] [18]. Specifically, Mask- varying size by means of the multi-stage refinement of the
RCNN and Cascade Mask RCNN are classic two-stage IoU threshold. Following the basic concept of “segmenting
models. SOLOv2 is a novel effective single-stage approach. objects by location” fitting the distribution of building in-
QueryInst is an end-to-end query-based framework that stances, SOLOv2 demonstrated a competitive performance.
achieves state-of-the-art performance on the COCO dataset. It, however, also showed its limit for irregularly distributed
and objects with varying scale because of the fixed size
4.2. Evaluation Metrics grid cells. QueryInst performs better for COCO than on
For evaluation, we use the standard COCO metrics [10]: the UBC dataset, probably because our dataset mainly cov-
APmask (averaged over IoU threshold), AP50 , AP75 , APS , ers urban areas with building instances, for which Queryinst
APM and APL . Due to the presence of buildings in high- doesn’t seem to be suitable. The average precision for small
density areas but also rather small buildings, we adjusted buildings (smaller than 10m×10m) APs is quite low for all
parts of the metrics. The area of objects referred to by S, M, models, showing that the detection of small buildings is a
and L is redefined for small as area less than 400 pixels, for challenge for recent models.
medium as area 400-1600 pixels, and for large as area 1600 Class-wise results for different roof types are shown in
pixels and above. Table 6. Some classes, e.g., flat and hipped roof, contain
To train the CNN-based models, we employ the dataset a large number of building instances and have obvious fea-
splits given in Section 3.5 and initialize the network with tures for classification. Thus, they could be better detected
ResNet50 pre-trained on ImageNet [3]. All models are and classified by the models. Other classes such as arched,
trained on 4 GPUs for 100 epochs with 4 tiles per GPU. flat complex and shed roofs consist of small numbers of in-
The learning rate is initialized with 0.02 and then reduced stances and also do not have distinctive features, adding to
by a factor of 0.1 at epochs 60 and 90. Multi-scale training the difficulty and challenge of detection and classification.
and random flipping are used as data augmentation during This implies the necessity to use few-shot learning or fea-
training. Instances with multiple function labels are corre- ture augmentation to improve the instance segmentation re-
spondingly employed multiple times for training and testing sults. Figure 5 shows segmentation results using different
of the different classes. The rest of the hyper-parameters models. The Cascade Mask R-CNN model has the best per-
of the model are set to the same values as in the original formance for building detection. Precise segmentation of
MMDetection [2] setup. individual buildings and roof type classification are chal-
lenging in complex scenes.
4.3. Experimental Setups
4.3.2 Baselines with Coarse-grained Function
4.3.1 Baselines with Fine-grained Roof Type
Compared with roof type classification, it is difficult to de-
Method AP AP50 AP75 APS APM APL termine the function of buildings from the visual features in
Mask R-CNN [9] 13.4 23.1 14.7 3.8 17.2 17.2 just a single image tile, especially in urban areas. In this pa-
C Mask R-CNN [1] 15.3 25.2 17.5 3.4 19.7 19.3 per, therefore, only the results with coarse-grained classes
SOLOv2 [22] 13.5 23.1 15.3 3.0 20.4 15.8 of building functions are demonstrated (Table 7). Different
QueryInst [7] 13.1 21.9 14.9 3.7 18.4 16.4 from the experimental results for roof type, the Mask R-
CNN model achieves the best performance with respect to
Table 5. Instance segmentation results using maskAP on UBC mAP , while the SOLOv2 model performs best concerning
test set with fine-grained roof categories. “C Mask R-CNN” de- AP50 . However, one has to admit that the overall perfor-
notes Cascade Mask R-CNN [1]. mance of all models is relatively poor. On one hand, it is
hard to classify the function based only on the visual ge-
As the geometry type of the roof is reasonably deter- ometry features. On the other hand, since the size of each
minable from visual features, we employed the fine-grained image tile is small, it is often not possible to determine the
classes of roof types to compare the current state-of-the-art relationships between nearby buildings with different func-
models. Table 5 shows the average precision (AP) results tions and, thus, discover the complex information necessary
also with different intersection over union (IoU) thresholds to discriminate the various functional areas of a city.
1418
Method AP AP50 FL FC SH GA GM H1 H2 MA PI AR OT
Mask R-CNN [9] 13.4 23.1 26.0 2.4 3.9 21.1 13.5 14.8 40.1 7.8 3.0 7.3 7.1
Cascade Mask R-CNN [1] 15.3 25.2 27.3 2.3 2.8 21.5 13.4 16.3 40.2 10.9 20.8 6.6 6.7
SOLOv2 [22] 13.5 23.1 26.1 3.3 2.9 20.5 14.1 15.2 33.9 10.4 10.4 6.5 5.9
QueryInst [7] 13.1 21.9 21.5 5.8 2.1 18.4 14.6 15.6 30.4 17.8 8.6 5.5 4.7
Table 6. Class-wise instance segmentation results on UBC test set with fine-grained roof classes: FL-flat, FC-flat complex, SH-shed,
GA-gable, GM-gambrel, H1-hipped V1, H2-hipped V2, MA-mansard, PI-pinnacle, AR-arched, OT-other
(a) Ground Truth (b) Mask R-CNN (c) Cascade Mask R-CNN (d) SOLOv2 (e) QueryInst
Figure 5. Example results of the selected models in Beijing (rows 1 and 2) and Munich (rows 3 and 4). RGB images converted to gray-scale
for better visualization.
4.3.3 Comparisons between Beijing and Munich of buildings’ functions in urban areas. To compare the dif-
ferences between the two cities and the influence on classi-
fications, we produced separate datasets for each city.
Individual cities have unique characteristics influenced by
their different culture and history. Particularly, Beijing and We compared these two datasets using the Mask R-CNN
Munich have different architectural styles and distributions model. The results of the experiments are shown in Tables 8
1419
Method AP AP50 AP75 APS APM APL Function
Method AP AP50
Mask R-CNN [9] 14.4 24.9 15.4 5.6 16.3 22.0 RE CO IN PU OT
C Mask R-CNN [1] 13.6 24.0 14.4 5.1 14.7 21.4 Beijing 14.0 23.5 40.5 13.2 0.0 9.4 7.0
SOLOv2 [22] 14.1 25.4 14.4 4.9 18.0 21.1 Munich 13.5 25.3 28.1 10.9 8.2 8.6 11.6
QueryInst [7] 10.4 20.4 10.8 4.3 12.8 16.9
Table 9. Results of separate experiments for dividing function
Table 7. Instance segmentation results with coarse-grained func- data in Beijing and Munich, respectively. RE-residential, CO-
tion classes using mask AP on the UBC test set. “C Mask R-CNN” commercial, IN-industrial, PU-public, OT-other
denotes Cascade Mask R-CNN.
5. Conclusion
and 9. The segmentation accuracy in Munich is lower than We have presented a novel remote sensing dataset with
in Beijing. This is mainly because Munich, on one hand, a specific focus on individual buildings and fine-grained
has a large number of connected closed loops of buildings classification concerning both, building geometry, i.e., roof
(apartment building blocks) and, on the other hand, many type, as well as functionality, i.e., occupation/usage. For
tiny independent buildings. For the former it is difficult classification, predefined classes are given on both a coarse-
to distinguish the individual building instances inside the and a fine-grained level. They can be employed according
building blocks, while the latter are hard to detect. The APs to different purposes or available numbers of instances. Se-
for both roof type and function for Beijing are, therefore, lected typical urban areas of Beijing and Munich are pro-
higher than for Munich. vided for combined as well as separate investigation. Exper-
iments with multiple state-of-the-art object detection mod-
Furthermore, the amount of available data and different
els have been conducted as baseline for further research and
characteristics of buildings lead to varying results in differ-
possible competition. The experiments demonstrate that the
ent cities. E.g., there are more flat roofs but fewer gable
detection and classification of individual buildings in dense
roofs in Beijing than in Munich. Therefore, the flat roof
urban areas are challenging. The instances show a large
accuracy in Beijing is much higher than in Munich while
diversity concerning size, shape, texture and relationship
the gable roof accuracy in Munich is higher than in Bei-
with neighboring objects, along with influences from his-
jing. The different characteristics and styles of buildings
tory, culture, climate, as well as density of habitation. The
in different cities also are decisive for the final results of
function of buildings is a latent feature which can only par-
classification for both, roof type and function. There are
tially (sometimes not at all) be derived from the appearance.
many high-rise buildings in Beijing but few in Munich. On
Even for the manual annotation often geo-information from
the other hand, there are many more churches in the city
different sources is needed for a reliable decision. We ex-
center of Munich than in Beijing. Commercial buildings
pect it to be a great challenge for classification to find more
in Beijing are most likely large shopping malls and look
high-level evidence by integrating, e.g., geometrical fea-
different from residential buildings. Opposed to this, most
tures like roof types and the structure of the neighborhood.
commercial buildings in Munich look similar to residential
In this first version of the dataset selected urban areas of
ones, therefore, they are difficult to distinguish. Overall, for
Beijing, China and Munich, Germany, are provided and the
all the results for Beijing and Munich, the AP of residen-
input data are solely RGB satellite imagery. In the future, it
tial buildings is the highest. Industrial buildings have the
is planned to extend the dataset in multiple dimensions: (1)
worst classification results concerning their function, prob-
Coverage of additional urban areas of interest worldwide,
ably because the amount of data is small and some indus-
(2) Use of multispectral as well as SAR imagery, and (3)
trial buildings are easily wrongly classified as residential or
Extension with multi-temporal data. Please note that with
commercial buildings.
an increasing amount of instances, those classes “merged”
in the current experiments can be “split” again as individual
Roof Type classes. Additionally, new classes can be added if neces-
Method AP AP50
FL GA HI AR OT sary.
Beijing 22.8 36.6 33.6 18.2 49.0 5.9 7.3
Munich 15.9 28.1 16.3 26.3 30.9 0.0 6.0 6. Acknowledgements
Table 8. Results of separate experiments for dividing roof type This work was supported by the National Key R&D Pro-
data in Beijing and Munich, respectively. FL-flat, GA-gable, HI- gram of China (Grant No. 2021YFB3900504) and Na-
hipped, AR-arched, OT-other tional Natural Science Foundation of China (Grant No.
62171436).
1420
References [13] Franz Rottensteiner, Gunho Sohn, Jaewook Jung, Markus
Gerke, Caroline Baillard, Sebastien Benitez, and Uwe Bre-
[1] Zhaowei Cai and Nuno Vasconcelos. Cascade R-CNN: high itkopf. The ISPRS benchmark on urban object classification
quality object detection and instance segmentation. IEEE and 3D building reconstruction. ISPRS Annals of the Pho-
Transactions on Pattern Analysis and Machine Intelligence, togrammetry, Remote Sensing and Spatial Information Sci-
43(5):1483–1498, 2019. 6, 7, 8 ences I-3 (2012), Nr. 1, 1(1):293–298, 2012. 2
[2] Kai Chen, Jiaqi Wang, Jiangmiao Pang, Yuhang Cao, Yu [14] Michael Schmitt, Seyed Ali Ahmadi, and Ronny Hänsch.
Xiong, Xiaoxiao Li, Shuyang Sun, Wansen Feng, Ziwei Liu, There is no data like more data - current status of machine
Jiarui Xu, Zheng Zhang, Dazhi Cheng, Chenchen Zhu, Tian- learning datasets in remote sensing. In 2021 IEEE Interna-
heng Cheng, Qijie Zhao, Buyu Li, Xin Lu, Rui Zhu, Yue Wu, tional Geoscience and Remote Sensing Symposium IGARSS,
Jifeng Dai, Jingdong Wang, Jianping Shi, Wanli Ouyang, pages 1206–1209, 2021. 1
Chen Change Loy, and Dahua Lin. MMDetection: Open
[15] Defence Science and Technology Laboratory. DSTL satellite
MMLab detection toolbox and benchmark. arXiv preprint
imagery feature detection. Available at [Link]
arXiv:1906.07155, 2019. 6
kaggle . com / c / dstl - satellite - imagery -
[3] Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, feature-detection/data. 3
and Li Fei-Fei. Imagenet: A large-scale hierarchical image
[16] Jacob Shermeyer, Daniel Hogan, Jason Brown, Adam
database. In 2009 IEEE Conference on Computer Vision and
Van Etten, Nicholas Weir, Fabio Pacifici, Ronny Hansch,
Pattern Recognition, pages 248–255. Ieee, 2009. 6
Alexei Bastidas, Scott Soenen, Todd Bacastow, and Ryan
[4] edPortal. Gaofen-2. Available at [Link]
Lewis. Spacenet 6: Multi-sensor all weather mapping
eoportal . org / web / eoportal / satellite -
dataset. In Proceedings of the IEEE/CVF Conference on
missions/g/gaofen-2. 3
Computer Vision and Pattern Recognition (CVPR) Work-
[5] Inc. EOS Data Analytics. Superview1. Available at https: shops, June 2020. 2, 3
//[Link]/find-satellite/superview-1/. 3
[17] Xian Sun, Peijin Wang, Zhiyuan Yan, Feng Xu, Ruiping
[6] Hongchao Fan, Alexander Zipf, Qing Fu, and Pascal Neis. Wang, Wenhui Diao, Jin Chen, Jihao Li, Yingchao Feng,
Quality assessment for building footprints data on Open- Tao Xu, et al. Fair1m: A benchmark dataset for fine-
StreetMap. International Journal of Geographical Informa- grained object recognition in high-resolution remote sens-
tion Science, 28(4):700–719, 2014. 4 ing imagery. ISPRS Journal of Photogrammetry and Remote
[7] Yuxin Fang, Shusheng Yang, Xinggang Wang, Yu Li, Chen Sensing, 184:116–130, 2022. 2
Fang, Ying Shan, Bin Feng, and Wenyu Liu. Instances as [18] Zhi Tian, Hao Chen, Xinlong Wang, Yuliang Liu, and Chun-
queries. In Proceedings of the IEEE/CVF International Con- hua Shen. AdelaiDet: A toolbox for instance-level recogni-
ference on Computer Vision, pages 6910–6919, 2021. 6, 7, tion tasks. [Link] 2019. 6
8
[19] Topcoder. Spacenet challenge 7: Multi-temporal urban
[8] Hirsh Goldberg, Myron Brown, and Sean Wang. A bench-
development challenge. Available at https : / / go .
mark for building footprint classification using orthorectified
[Link]/spacenet/. 2
RGB imagery and digital surface models from commercial
[20] Adam Van Etten, Dave Lindenbaum, and Todd M Bacastow.
satellites. In 2017 IEEE Applied Imagery Pattern Recogni-
Spacenet: A remote sensing dataset and challenge series.
tion Workshop (AIPR), pages 1–7. IEEE, 2017. 2
arXiv preprint arXiv:1807.01232, 2018. 2, 3
[9] Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Gir-
[21] Shenlong Wang, Min Bai, Gellert Mattyus, Hang Chu, Wen-
shick. Mask R-CNN. In Proceedings of the IEEE Interna-
jie Luo, Bin Yang, Justin Liang, Joel Cheverie, Sanja Fidler,
tional Conference on Computer Vision (ICCV), Oct 2017. 6,
and Raquel Urtasun. TorontoCity: Seeing the world with a
7, 8
million eyes. arXiv preprint arXiv:1612.00423, 2016. 2, 3
[10] Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays,
[22] Xinlong Wang, Rufeng Zhang, Tao Kong, Lei Li, and Chun-
Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence
hua Shen. Solov2: Dynamic and fast instance segmenta-
Zitnick. Microsoft coco: Common objects in context. In
tion. Advances in Neural Information Processing Systems,
David Fleet, Tomas Pajdla, Bernt Schiele, and Tinne Tuyte-
33:17721–17732, 2020. 6, 7, 8
laars, editors, Computer Vision – ECCV 2014, pages 740–
755, Cham, 2014. Springer International Publishing. 6 [23] OpenStreetMap Wiki. Openstreetbrowser/category list.
Available at [Link]
[11] Emmanuel Maggiori, Yuliya Tarabalka, Guillaume Charpiat,
wiki/OpenStreetBrowser/Category_list. 4
and Pierre Alliez. Can semantic labeling methods generalize
to any city? the INRIA aerial image labeling benchmark. [24] Junshi Xia, Naoto Yokoya, Bruno Adriano, Lianchong
In 2017 IEEE International Geoscience and Remote Sensing Zhang, Guoqing Li, and Zhigang Wang. A benchmark high-
Symposium (IGARSS), pages 3226–3229, 2017. 3 resolution GaoFen-3 SAR dataset for building semantic seg-
[12] Ribana Roscher, Michele Volpi, Clément Mallet, Lukas mentation. IEEE Journal of Selected Topics in Applied Earth
Drees, and Jan Dirk Wegner. SemCity Toulouse: A bench- Observations and Remote Sensing, 14:5950–5963, 2021. 3
mark for building instance segmentation in satellite images.
In ISPRS Annals of Photogrammetry, Remote Sensing and
Spatial Information Sciences, 2020, volume V-5-2020, pages
109–116, Aug. 2020. 3
1421