Building Classification
Building Classification
Keywords: Several studies have focused on generating seismic vulnerability maps for earthquake-prone areas, particularly
Building characterization in Indonesia. Building typologies are a key factor in determining vulnerability to earthquakes. However,
Street-level imagery conducting large-scale field surveys to determine the spatial distribution of building typologies in a city is
Deep convolutional neural networks
uneconomical. This paper explores the use of a convolutional neural network (CNN) to automatically detect
Image classification
building typologies from diverse regions in Indonesia, utilizing both conventional and automated building
Seismic risk assessment
image acquisition processes. In this study, datasets from three distinct image acquisition methods are trained
with four unique CNN architectures to identify the best-performing model to classify building typologies. The
sample size effect on CNN performance is also investigated. The results showed that randomly sampled Google
Street View (GSV) images are the most effective dataset for the CNN model, achieving an f1-score of 84.33%.
Among the network architectures tested, MobileNet demonstrated superior performance on the majority of
evaluated datasets. As the sample size increases by about 350% in the dataset, there is a positive correlation
with up to 2.3% f1-score improvement. Using the best-performing CNN model, two building vulnerability
models were employed to assess the spatial distribution of building damage in the urban area of Bandung,
considering a hypothetical scenario of an M7 earthquake. Incorporating local construction data, one of the
generated maps estimated that approximately 55% of buildings in Bandung would experience moderate to
severe structural damage. This study showcases the potential of CNN models in automating regional seismic
assessments and providing valuable insights for comprehensive seismic mitigation strategies.
1. Introduction To reduce seismic risk, many studies have produced a global risk
map due to earthquakes (Silva et al., 2018; Pittore et al., 2017) which
Earthquakes are classified as highly destructive natural phenomena. can greatly benefit decision-makers in understanding their popula-
Even though their incidence accounted for merely 8% of global disaster tion’s vulnerability and help them address problems before the disaster
events during the period from 2000 to 2019, they were responsible for event. However, the problem with these risk estimations is the inade-
a staggering 58% of total fatalities within that timeframe (CRED and quate building information (typology, construction time, etc.) of low to
UNDRR, 2020). However, these casualties are not only caused by the middle-income countries (Moroni and Ghomez, 2002; Charleson et al.,
seismic hazard itself but are also deeply intertwined with the building 2017). Also, the available global risk maps commonly assume the same
quality and the social readiness to face disaster. While seismic hazard
building type for wide regions (sometimes up to about 1000 km2 ) (Silva
is a natural phenomenon, the other two components are oftentimes de-
et al., 2018), which is less than ideal.
pendent on the strength of the economy. Consequently, the unfortunate
Ideally, to obtain better risk prediction, large-scale building surveys
reality emerges where approximately 90% of earthquake-related deaths
take place in countries with low to middle-income levels (World Health must be performed to assess building vulnerability, where the determi-
Organization, 2018). nation of the lateral load-resisting system and its constituent elements
✩ This document is the results of the research project funded by P2MI Institut Teknologi Bandung, Indonesia.
∗ Corresponding author.
E-mail address: [Link]@[Link] (P.W. Sarli).
[Link]
Received 30 July 2023; Received in revised form 19 December 2023; Accepted 26 December 2023
Available online 9 January 2024
0952-1976/© 2024 The Authors. Published by Elsevier Ltd. This is an open access article under the CC BY-NC-ND license ([Link]
nc-nd/4.0/).
H.R. Firmansyah et al. Engineering Applications of Artificial Intelligence 131 (2024) 107824
Table 1
Image sources used in different studies to obtain building information of certain regions or buildings either through top view or side view of buildings.
Data category Machine learning methodology Reference
Top View of Buildings
Unmanned Aerial Vehicle (UAV) CNN, deep convolutional encoder–decoder Meng et al. (2022), Zhang (2020), Wang et al. (2023a)
Aerial Photography bag-of-visual-words (BoW model); CNN Naito et al. (2020)
Remote Sensing Images (RSIs) CNN, Swin Transformer Wang et al. (2023b, 2022), Cui et al. (2022)
Side View of Buildings
Google Street View (GSV) CNN, DCNN Gonzalez et al. (2020), Pelizari et al. (2021)
is crucial. Nonetheless, this information can only be achieved through to detect earthquakes (Perol et al., 2018) and also seismic damage
examination of the building blueprints or expert observations. How- identification of post-earthquake buildings (Xu et al., 2019). While
ever, access to building blueprints is frequently limited, and vital infor- with GSV images, CNN has been used in street number recognition,
mation may also be lacking for structures built using self-construction traffic sign recognition, number of building stories, estimation of the
methods. As a result, expert surveys are the preferred method for cre- sky, tree, and building view factors, classification of building instances,
ating extensive building inventories, particularly in terms of building and estimation of neighborhood demographic composition (Gong et al.,
typology, which aids in vulnerability assessment. This entails visiting 2018). A particular study using GSV to classify building typologies has
multiple buildings and conducting visual inspections or interviews with already been conducted (Gonzalez et al., 2020), although the ground-
the building owners to gather general specifications (Harirchian et al., image acquisition is still manual using screenshots and is limited by
2020), which can be achieved through census data (Stachl et al., 2020) GSV coverage and capabilities. Another study (Pelizari et al., 2021) ex-
or a partial survey of the entire structural population (Salgado-Gálvez plores the automated inference of seismic building structural types, the
et al., 2014; Acevedo et al., 2017). Then, expert opinions are necessary material of the lateral-load resisting system, and building height with a
to verify and process these data to meaningful information for risk or collection of GSV imagery that yields exceptional results. However, this
vulnerability mapping. Nevertheless, in the case of expansive urban study has yet to demonstrate an instance of applying automated build-
areas, conducting surveys for each asset may prove impractical, costly, ing vulnerability mapping, taking into account a building vulnerability
and sometimes infeasible, especially for the low to middle-income model that is contingent upon the building type, construction locality,
country setting. and a specific earthquake scenario. In conclusion, the technology’s po-
Recent advances have opened up opportunities to apply the automa- tential is undeniable. Its application, particularly for local government,
tion concept to the building vulnerability visual assessment using image enables wide-scale earthquake risk assessment to be conducted with
[Link] researchers have started to use either top view images significant time and cost reductions.
of buildings, whether using Unmanned Aerial Vehicle (UAV) or remote Due to Indonesia’s location in a tectonically active region, the
sensing images (RSIs), or side view images such as Google Street view majority of the country is susceptible to earthquakes and associated
secondary hazards like tsunamis, landslides, liquefaction, floods, and
(GSV) images to replace costly surveys with virtual tours to acquire
fires. Consequently, there is an urgent and significant need to un-
needed information remotely, i.e. facade imagery, lateral load-resisting
derstand earthquake risks in Indonesia. In the last 30 years alone, 9
systems, and materials (see Table 1 for list of recent references). This
major earthquake disasters have been recorded and have caused a total
visual information then is coupled with the rapid development of
economic loss of more than 160 trillion IDR (±$10.7b as of 2023)
remote sensing technologies since the early 2000s, which has allowed
and more than 200.000 casualties (Meilano et al., 2019). The 2020–
for the measurement of variables such as plan-built areas, building
2024 National Disaster Management Plan (BNPB Indonesia, 2020), on
height, type of roof, building classifications, and building age, therefore
the one hand, aims to enhance sustainable development and minimize
reducing the cost of obtaining ground data for large assessments.
economic losses caused by disasters that adversely affect the country’s
Although both the top view and side view of buildings can be useful in
GDP. In pursuit of this goal, the plan emphasizes the importance of mi-
estimating possible structural information, from a structural perspective
cro zonation mapping of geologically vulnerable areas as a significant
the side view can give more information on the typology and building
indicator.
construction data of the building, hence data such as GSV can be highly
What amplifies Indonesia’s seismic risk is its reportedly significant
beneficial for seismic vulnerability [Link], in cases of
number of non-engineered residential houses, even in the capital city
large and detailed implementation of this building typology assessment, of Jakarta (Struyk et al., 1990). Non-engineered buildings are sponta-
which frequently requires thousands of building data images for each neously and informally constructed without any or little intervention
city, the process continues to incur substantial expenses associated with by professional engineers (Arya et al., 2014). Non-engineered build-
data acquisition due to its predominantly manual nature. Regardless of ings can vary depending on locally available materials or construction
whether it is accomplished through expert visual analysis or automated methodologies that have permeated a particular region. In this defi-
techniques, the accurate identification of building typologies remains a nition, non-engineered buildings can be wooden buildings, masonry,
challenging task. reinforced concrete (RC) frames, adobe structures, etc. However, in
Concurrently, there has been a notable utilization of deep learn- the context of Southeast Asia, especially Indonesia, the most prevalent
ing techniques across various perceptual tasks. These methodologies building techniques involve unconfined masonry or RC frames with
circumvent the necessity for the manual design of specialized image infilled masonry (Watanabe et al., 2013). Non-engineered buildings
feature detectors by seeking a comprehensive set of transformations also typically do not follow national standards or international design
directly from the available data. As a consequence, remarkable achieve- practices and, therefore, might have limited detailing at the connec-
ments have been attained, particularly in the realm of computer vision tions or between elements (Arya et al., 2014). In the examination
challenges like object classification. Convolutional Neural Network of construction quality, a study (Maqsood et al., 2013) undertook an
(CNN) is a specialized artificial neural network architecture intended evaluation of the susceptibility of the building inventory across the
for signal processing in general and 2D images (O’Shea and Nash, Asia-Pacific region, with a particular focus on Indonesia. The findings
2015). In CNN, images represented by a 3D matrix pass through succes- reveal a prevalent issue concerning substandard construction materials
sive layers organized in a hierarchical pattern recognition system that and practices which might result in higher risks in the event of earth-
learns the most useful features to classify images. [Link], the development of a CNN model that can result in
CNNs are starting to be used in seismic risk assessment in various better prediction of building typology within the Indonesian context,
ways. For example, CNN was used to process ground velocity records that can also help estimate seismic risk is crucial.
2
H.R. Firmansyah et al. Engineering Applications of Artificial Intelligence 131 (2024) 107824
This study aims to facilitate automated earthquake vulnerability Each of these building stock or typologies represents different
mapping down to individual building levels that address unique build- earthquake-resisting systems, and hence will impact its vulnerability.
ing typologies in Indonesia by using ground-level image processing Here the definition of each typology and their visual characteristics are
and utilizing a convolutional neural network (CNN) to identify the explained.
building typologies. The study will also predict population vulnera-
bility through the developed method. The remainder of this paper 2.1. Unconfined Masonry (UC)
is organized as follows. Section 2 describes an overview of common
residential building typologies in Indonesia. Section 3 describes the UC structure utilizes bricks as its load-bearing element without any
data acquisition process, the assembling of the data sets, and the details vertical wall binding elements such as concrete columns. During an
of the experiment. The results and subsequent discussion are presented earthquake, these structures perform poorly, resulting in cracking and
in Section 4. Finally, Section 5 presents the main conclusions of this failure. This structure typology has the lowest lateral load capacity
study. of any other masonry structure and could be improved by adding a
confinement element (Shukla et al., 2021). Low-rise UC structures have
2. Overview of common residential building typologies in Indone- low survival rates, even for low-to-moderate earthquakes.
sia According to insights derived from experts in the field, the visual
attributes associated with UC structures suggest that the identification
There are four common residential building typologies in Indone- of these structures becomes relatively challenging when their exterior
sia: confined masonry (CM) systems, reinforced concrete frames with components have been covered with plaster or paint.
infilled masonry walls (RC) systems, timber structure (TB) systems,
and unconfined masonry (UC) systems (see Fig. 1). CM and RC struc- 2.2. RC infilled masonry (RC)
tures are the most common household-dwelling structures in Indone-
sia (Watanabe et al., 2013), while TB and UC structures represent RC typology is a structure characterized by its concrete frame
typologies more common in rural areas. structure, which is considered to be the main load-bearing element, and
3
H.R. Firmansyah et al. Engineering Applications of Artificial Intelligence 131 (2024) 107824
the brick walls are considered non-structural. During the construction objectivity in choosing a given label for the data set. A total of 4563
process, the RC structure was built with concrete columns and beams images from various ground-image sources were then labeled with this
first, followed by the construction of the brick walls. decision tree. The result of the labeling was compared with expert
For RC typology, the difference in thickness between the column opinions on building datasets.
and wall provides visual characteristics that may appear in the form In Phase II, a combination of five datasets, four CNN architec-
of column protrusions that are visible from the side of the building. In tures, and three resampled distributions were trained independently to
cases where the side of the building cannot be seen due to equipment discover the best-performing CNN model to predict building typologies.
limitations or building arrangement, identification of this typology In Phase III, the best-performing CNN model will be used to identify
can be estimated by considering socio-economic factors. Considering building typology on a city level, where then using the fragility curve
that construction with a concrete frame requires a relatively large estimation for the given typology and the size of the possible hazard
cost (Sazedj et al., 2013), this RC structure is generally only built by scenario predicted, the full vulnerability of the city is estimated. Each
the upper-class economy. This socio-economic level can be estimated of these phases will be explored in more detail in the subsections below.
by the appearance of the building facade which is visibly luxurious and
modern, and also if the number of stories of the building is more than 3.2. Phase I-Data labeling
2.
Labeling is a process that is time-consuming and expensive in
2.3. Confined Masonry (CM) artificial intelligence research (Dimitrakakis and Savu-Krohn, 2008).
This is because the algorithm requires human guidance to provide
CM structures use the common materials of concrete and brick case examples and solutions, including deep learning cases in image
which are relatively inexpensive and widely used for low-rise build- recognition. However, the presence of ambiguity in the classification
ings (Brzev and Mitra, 2007). The principal distinction between CM of building typology necessitates the exploration of a clear decision
(concrete masonry) and RC (reinforced concrete) construction lies in tree. Therefore in Phase I, to provide objectivity to the labeling process
the sequence of building stages, which affects the load-bearing mech- firstly a decision tree was prepared with the assistance of experts.
anism. In CM structures, brick walls are erected partly before filling The resultant labels obtained through this objective decision tree will
the concrete columns between them. Conversely, in RC structures, the subsequently be reviewed by experts. This decision tree will serve as
columns and beams of the portal frame structure are erected initially, a tool for categorizing the complete datasets acquired from multiple
with the brick walls being filled in subsequently. sources in the study. The building typology decision tree and datasets
Due to the difference in construction, for CM structures masonry used in the study will be further explored below.
walls are considered as lateral resisting loads, and the reinforced con-
crete is used as wall binders or confining elements. This confining 3.2.1. Building typology decision tree
element serves to increase the stability, integrity, and ductility of the A typology decision tree allows for a more efficient and cost-
brick wall against lateral and horizontal loads (EERI, 2019). On the effective labeling process. Rather than using professional services to
other hand, in RC structures, the concrete frame supports the entire assess thousands of building image samples, undergraduate students
load, and the brick walls are considered non-structural components. could be employed, as was done in this research. This decision tree was
Differentiating between CM and UC structures becomes challenging first created as a collaboration with Indonesian structural engineers, ar-
when their exterior components have been covered with plaster or chitects, and contractors who are familiar with local construction prac-
paint. But differentiation between CM and RC is still possible, if there tices and surveyors who have done ground truth surveys of structural
are visible protruding columns from the walls. classification before.
The building typology decision tree implemented in this study is
2.4. Timber (TB) proposed in Fig. 3, where the scheme is strictly based on visually in-
ferable indicators (visual structural criteria) and represents Indonesian
TB structure is defined as a structure with a load-bearing element in structural typology which is aligned with EMS taxonomy and existing
the form of a wooden frame, where the TB structure transmits lateral vulnerability models (Grunthal, 1998).
loads and gravity to the foundation. In contrast to previous typologies The decision tree is systematically structured starting from the
considered, the TB structure is relatively easy to identify, particularly more prominent and unique building features that are the main dis-
with visible wood texture on the walls or frame structure. tinguisher between building types, progressing to the bottom decision
tree with features that are more general and common in Indonesia.
3. Data and methodology Firstly, only residential buildings below 4 stories were examined, as
structures above 4 stories have different typological features. Then
This section is organized as follows: firstly, the general methodology as the TB structure has the most unique feature, that is the wood
is presented, then a more detailed explanation for each of the phases texture, it is used as the first decision element. Then, the general
within the method is given. features in buildings such as supporting columns and plaster walls will
be observed. If the structure has protruding columns, then it is an RC
3.1. General methodology structure. But if the structure has protruding columns but is only a
one-story building with a width below 5 m, then the structure is a
This study was performed in three phases. The first phase focuses on CM structure. If there are no protruding columns but the structure has
the structural engineering part of the problem, where a labeling process a premium or modern building facade then it is considered RC. This
on how to identify a building’s typology through visual information decision is made as interviews with contractors and practitioners have
was created (Phase-I). Then the second phase focuses on the creation of stated that RC structures have only become more common in recent
the machine learning model via a convolutional neural network (CNN) decades, due to the better literacy of earthquake-resistant construction
using the labeled datasets from earlier phases (Phase II). Lastly, in the practices in Indonesia. Therefore more premium and modern structures
third phase, the potential application of the CNN model to estimate a with the building style of the 2000s and later will be considered RC.
city’s vulnerability is explored (Phase III). See Fig. 2 for clarity. Then, the next step is to the walls, whether it is plastered or not. If
In Phase I, a labeling typology decision tree is created through plastered, the structure will be considered CM, if not depending on
expert interviews. As visual identification of building typology can be the invisibility of supporting columns, it can be UC. Therefore, it is
subjective, the decision tree was created to ensure consistency and important to emphasize that through the decision tree, in this study
4
H.R. Firmansyah et al. Engineering Applications of Artificial Intelligence 131 (2024) 107824
Fig. 2. An illustration of the general methodology adopted in this study. Overall, the methodology can be divided into three phases: (1) Data Labeling; (2) Development of the
CNN model; and (3) Estimation of the seismic vulnerability map with the aid of the CNN model to predict building typology.
UC type structures are only identified as such if the masonry walls are from Bandung fragility assessment study (Sharon, 2022) and personal
exposed and there are no visible bonding columns as in CM structures. archives. This method was included as some areas, especially in high-
However, in this study, it is argued that even if a UC structure is density regions such as slums cannot be obtained through GSV images,
misidentified as CM because of its wall plastering, plastered walls can as GSV images typically employ a car that has limited access to these
increase lateral capacity by about 35% compared to plastered walls, as areas. GSV virtual tour images implemented virtual tours in Google
well as an increase in the ductility of plastered walls (Rildova et al., Maps and took building facades with screenshots. This method ensures
2012). Therefore the potential misidentification of UC structures to be better framing of the building as the pictures were taken manually.
CM is considered acceptable as CM structures do have an improved While randomly sampled images were collected using GSV API in
fragility behavior. However, in general, these results indicate that the predetermined areas across Indonesia. For both the GSV virtual tour
labeling process can be streamlined both in time and cost with the images and the randomly sampled GSV images, the precise coordinate
implementation of a building typology decision tree. From several key data source from which the visual view of the front of the building
components that aid the decision-making process of building typology was accessed can be obtained and later used in the vulnerability pre-
labels, the labeling process can be done more efficiently. diction of the building. The differences between these types of images
are described in detail in Table 2. Especially for randomly sampled
3.2.2. The dataset images, we queried the street-level imagery available and its metadata
The dataset used in this study is divided into 3 main sources in Indonesia.
of images: (1) camera images, (2) GSV virtual tour images, and (3) The metadata includes image location (in 𝑙𝑎𝑡𝑖 and 𝑙𝑜𝑛𝑖 ), the heading
randomly sampled images from GSV. The camera image was captured which indicates the compass heading of the camera, radius which sets
5
H.R. Firmansyah et al. Engineering Applications of Artificial Intelligence 131 (2024) 107824
the radius to search for a street-level image (taken at the default value 3.3. Phase II-convolutional neural network
of 50 m), and size that specifies the output size of images in pixels. In
this study, the imagery is acquired with a size of 256 𝑥 256 pixels which In total, we conducted 60 experiments with three resampled dis-
is suitable for the default image size available in Tensorflow, device tributions, four pre-trained network architectures, and five datasets
limitations, and visibility of structural details in the facade imagery. with varying domain sources and the number of samples. The total
Overview of the image acquisition method using GSV API is illustrated optimization and training times required between 40 to 50 h and
in Fig. 4. The pseudocode of randomly sampled method utilizing GSV were executed on a laptop equipped with the Keras/Tensorflow deep
API in this study is depicted in the Algorithm 1. learning framework and an RTX 3060 NVIDIA GPU.
Algorithm 1 GSV API Image Collection
1: boundary = {𝑙𝑜𝑛1 , 𝑙𝑜𝑛2 , 𝑙𝑎𝑡1 , 𝑙𝑎𝑡2 } 3.3.1. Resampled distributions
2: radius = 50 In the experimental workflow of deep learning, data serves three
3: model = BuildingImageClassifier() distinct purposes: (1) facilitating the training of models; (2) assessing
4: building_conf_threshold = 80% model performance using data not employed in any preceding stages;
5: for 𝑁 = 1, 2, 3, … num_iteration do and (3) fine-tuning model hyperparameters to identify the optimal set
6: coord_x=random_uniform(𝑙𝑜𝑛1 , 𝑙𝑜𝑛2 ) yielding the best performance. Therefore, the dataset is randomly split
7: coord_y=random_uniform(𝑙𝑎𝑡1 , 𝑙𝑎𝑡2 ) into three subsets: training, validation, and testing, which are used in
8: heading=random in{0, 45, 90, 135, 180, 225, 270, 315} each of the stages mentioned above. Our splits involved using 70% of
9: images = query_GSV_API(coord_x, coord_y, heading, radius) the data for training, 15% for validation, and 15% for testing, which
10: if ([Link](images) > building_conf_threshold) then is commonly used in deep learning research (Kasapbaşi et al., 2022;
11: save(images) Prashanth et al., 2020).
12: else We also repeated each random train–validation–test split three
13: delete(images) times to ensure consistency and minimize experimental bias due to
14: end if the train–validation–test split with a limited amount of data, and all
15: end for the model performance for analysis will be averaged with these three
splits. With these operations, we can obtain an unbiased estimate of the
A similar study (Pelizari et al., 2021) uses administrative boundaries
model performance, allowing us to predict the performance observed
in the study area to pick GSV panorama shooting points randomly.
when the models are applied to real-world data.
Nevertheless, it has been observed through field studies that a signif-
Knowing our sample distribution for each class, it is found that CM
icant number of buildings in Indonesia are inaccessible through the
building typology is the supermajority for all the data sources. This
use of GSV due to their locations in narrow alleys and remote regions.
Consequently, it is necessary to explore alternative data sources, such causes a heavily imbalanced dataset in our experiment. Our imbalanced
as camera images, to develop a specialized model capable of effectively datasets may cause the model to gravitate to predicting the majority
addressing these particular scenarios. class, to minimize wrong predictions (losses). Therefore, we employed
This virtual tour survey required 45 working hours over one month an oversampling strategy with data augmentation to prevent model
by a master’s student and a final-year civil engineering student. ran- bias.
domly sampled GSV images required about 5 h to acquire building
images with decent quality and another 20 h to label the images. A total 3.3.2. Data augmentation
of 4563 images comprised of 685 camera images, 1321 GSV virtual Data augmentation is a method in deep learning to modify input
tour images, and 2557 random-sampled GSV images were considered variables, with certain variation characteristics to enrich input combi-
for this deep learning experiment process. The Table 3 presents the nations or resolve data limitations. The model that learns images of
building typology distribution of the defined dataset. buildings with various conditions is expected to be more robust and
6
H.R. Firmansyah et al. Engineering Applications of Artificial Intelligence 131 (2024) 107824
Table 2
Dataset characteristics.
Characteristics Camera image GSV virtual tour Randomly
sampled GSV
Image source Camera Screenshot GSV API
Automated sampling No No Yes
Data acquisition costs High Moderate Low
Remote area accessibility Yes No No
Proper angle of view of building facades Yes Yes No
able to have good generalization capability, as indicated by higher on model performance, consequently, the number of samples has to be
validation and testing performance. Augmentation performed on the kept constant. Thus, 2 new datasets will be created from the subset of
dataset should not change the final meaning of the transformed images. the VT,L, and RS,L dataset which have the same number of samples and
For example, the augmentation of flipping a photo vertically will typology class distribution as the CI dataset, which in this study will be
change the meaning of the building, while flipping a photo horizontally abbreviated as VT,S and RS,S.
does not change the meaning.
In this study, data augmentation will be implemented with two 3.3.3. Network architectures
main objectives (1) Enriching the variety of building images to improve For this study, we select four state-of-the-art CNN architectures
the generalization capability of the CNN model, and (2) balancing the that are commonly used with pre-trained weights on ImageNet: Incep-
number of samples in each class to prevent majority bias in the model. tionV3, Xception, MobileNet V3L, and EfficientNet B0.
Regarding objective (1), the random transformation combination
A unique feature of InceptionV3 and other Inception architectures
that will be given to the training images are as follows: rotation, hor-
is a convolution layer group known as inception block (Szegedy et al.,
izontal flipping, zoom in /out, dropout pixels, cropping with Gaussian
2015), the main principle of this inception module is that each CNN
noise, salt & pepper noise, Gaussian blur, gamma contrast enhance-
layer will learn the detailed features from the 1 × 1 convolution to the
ment, image saturation level, and color sharpness. These transfor-
coarse features of the 5 × 5 convolution, and the salient features of the
mations are commonly used in several convolutional neural network
max-pooling layer from the previous layer. The results of the detailed
studies (Shorten and Khoshgoftaar, 2019; Khalifa et al., 2022; Hao
and coarse feature learning in this image will be combined with the
et al., 2021; Barai and Heikkinen, 2017; Saikia et al., 2021; Rahman
concatenate function which will be forwarded to the next inception
et al., 2021; Afifi and Brown, 2019).
module for further processing.
Perez and Wang (2017) shows that an adequate augmentation
process can increase the accuracy of the validation set and test set by Xception itself is an acronym for extreme inception as a further
6%–7%. The optimal combination of augmentation functions given may innovation from InceptionV3. Spatial separable convolution in the
be dependent on the problem to be solved at hand. inception module is replaced with depthwise separable convolution.
Regarding objective (2), the balancing process of the class sample This discovery reduces the matrix multiplication operation by up to
on the dataset with data augmentation will be carried out with an 83% (Chollet, 2017).
oversampling strategy to even out the number of samples for each class MobileNet V3L is a further development from the previous vari-
in the dataset. Augmentation will only be carried out on the training ant of the MobileNet architecture which is optimized for the use of
sample, which is 70% of the sample in the dataset. For the majority computer vision technology for smartphone devices that have lower
class (CM) sample, it was determined in this study that 2 augmentation computing capabilities than personal computers (Howard et al., 2019).
images would be generated for each original image. The other class MobileNet V3L uses depthwise separable convolution and residual
samples will be augmented/generated proportionally until the total inverse structure. Each layer is enhanced with a hard-swish nonlinear
sample in each class is equal to the majority class sample. activation function.
The dataset distribution after the application of this augmentation The EfficientNet model is a family of CNN architectures that have
strategy is presented in Table 3. It should be noted for this study, due 7 variants. This variant can be adjusted according to performance
to one of the factors to be identified is the effect of the image sources requirements with trade-offs against the number of parameters that
7
H.R. Firmansyah et al. Engineering Applications of Artificial Intelligence 131 (2024) 107824
Table 3
The distribution of the building typology datasets post-augmentation.
Typology class Total training % Training Augmented Total sample
sample sample (70%) sample post-augmentation
Camera Images (CI)
Confined Masonry 415 60.6% 291 579 870
RC Infilled Masonry 152 22.2% 106 764 870
Timber Structure 80 11.7% 56 814 870
Unconfined Masonry 38 5.5% 27 843 870
Total 685 100.0% 480 3000 3480
GSV Virtual Tour (VT,L)
Confined Masonry 761 57.6% 532 1064 1596
RC Infilled Masonry 275 20.8% 192 1404 1596
Timber Structure 174 13.2% 121 1475 1596
Unconfined Masonry 111 8.4% 77 1519 1596
Total 1321 100.0% 922 5462 6384
Randomly sampled GSV (RS, L)
Confined Masonry 1544 60.4% 1080 2160 3240
RC Infilled Masonry 565 22.1% 395 2845 3240
Timber Structure 326 12.7% 228 3012 3240
Unconfined Masonry 122 4.8% 85 3155 3240
Total 2557 100.0% 1788 11 172 12 960
Table 4
Hyperparameter optimization results.
Model Inception Xception MobileNet EfficientNet
LR (%) Batch Val loss LR (%) Batch Val loss LR (%) Batch Val loss LR (%) Batch Val loss
CI 0.010 16 1.286 0.020 16 1.14269 0.083 8 1.2173 0.167 32 1.1458
VT-S 0.010 8 1.007 0.033 16 0.97445 0.067 16 1.31 0.083 8 0.99759
RS-S 0.013 16 0.826 0.027 16 0.788 0.010 32 1.27518 0.100 16 0.7833
affect memory requirements. The main inspiration for this CNN model Algorithm 2 Multi-stage Hyperparameter Optimization
family is to achieve optimal accuracy and efficiency when a CNN
for dataset = {CI, VT-S, RS-S} do
architecture is enlarged to complete more complex tasks (Tan and Le,
for architecture = {Inception, Xception, MobileNet, EfficientNet}
2019). For this study, EfficientNet with the B0 variant, requiring the
do
minimum amount of memory is used due to device memory limitation.
for trial = 1, 2, … num_trial_coarse do ⊳ Coarse search
hyperparams = (𝑖𝑛𝑖𝑡_𝑣𝑎𝑙_𝑙𝑜𝑠𝑠, 0, 0)
3.3.4. Hyperparameter optimization, fine-tuning, and training 𝑙𝑟1 = random in {10−2 , 10−3 , 10−4 , 10−5 }
For each CNN architecture, several hyperparameters can be tuned to 𝑏𝑎𝑡𝑐ℎ1 = random in {8, 16, 32}
achieve better performance. Hyperparameters are parameters that con- 𝑣𝑎𝑙_𝑙𝑜𝑠𝑠1 = calculate loss(𝑙𝑟1 , 𝑏𝑎𝑡𝑐ℎ1 )
trol the learning process of deep learning systems and in the end, will if 𝑣𝑎𝑙_𝑙𝑜𝑠𝑠1 < hyperparams[0] then
indirectly determine the parameters of the learning model (Wu et al., hyperparams = (𝑣𝑎𝑙_𝑙𝑜𝑠𝑠1 , 𝑙𝑟1 , 𝑏𝑎𝑡𝑐ℎ1 )
2019). The optimal hyperparameters will maximize the performance of end if
the deep learning system. However, this optimal hyperparameter value end for
is specific to each problem encountered. for trial = 1, 2, … num_trial_fine do ⊳ Fine search
In this study, two hyperparameters (learning rate (LR) and batch hyperparams = (𝑣𝑎𝑙_𝑙𝑜𝑠𝑠1 , 𝑙𝑟1 , 𝑏𝑎𝑡𝑐ℎ1 )
size) for CI, VT-S, and RS-S datasets will be optimized using a multi- 𝑙𝑟2 = random in {𝑙𝑟2 ∕2, 𝑙𝑟2 ∗ 2}
stage random search. For VT-L and RS-L datasets, the optimized hyper- 𝑣𝑎𝑙_𝑙𝑜𝑠𝑠2 = calculate loss(𝑙𝑟2 , 𝑏𝑎𝑡𝑐ℎ1 )
parameter from its corresponding VT-S and RS-S will be used due to if 𝑣𝑎𝑙_𝑙𝑜𝑠𝑠2 < hyperparams[0] then
device and time limitations. The pseudocode of this multi-stage random hyperparams = (𝑣𝑎𝑙_𝑙𝑜𝑠𝑠2 , 𝑙𝑟2 , 𝑏𝑎𝑡𝑐ℎ1 )
search method is depicted in Algorithm 2: end if
The random search uses several predetermined values that are end for
randomly chosen for each hyperparameter (coarse search), then the for trial = 1, 2, … num_trial_vfine do ⊳ Very fine search
model is trained with that hyperparameter combination only for a few hyperparams = (𝑣𝑎𝑙_𝑙𝑜𝑠𝑠2 , 𝑙𝑟2 , 𝑏𝑎𝑡𝑐ℎ1 )
iterations (epochs), and the validation loss is calculated and compared 𝑙𝑟3 = random in {𝑙𝑟2 ∕3 , 𝑙𝑟2 ∗ 2∕3 , 𝑙𝑟2 ∗ 4∕3 , 𝑙𝑟2 ∗ 5∕3}
to the current hyperparameter values that yield minimum validation 𝑣𝑎𝑙_𝑙𝑜𝑠𝑠3 = calculate loss(𝑙𝑟3 , 𝑏𝑎𝑡𝑐ℎ1 )
loss. Bergstra and Bengio (2012). The results of this hyperparameter if 𝑣𝑎𝑙_𝑙𝑜𝑠𝑠3 < hyperparams[0] then
optimization process are depicted in Table 4 hyperparams = (𝑣𝑎𝑙_𝑙𝑜𝑠𝑠3 , 𝑙𝑟3 , 𝑏𝑎𝑡𝑐ℎ1 )
This process resulted in each model trained on a specific dataset end if
source and CNN architecture having the most optimal hyperparameters end for
for their actual training stages. It ensures each model is trained to end for
reach the global optimum, avoiding the local minima present in the end for
loss function.
Following the hyperparameter optimization process, these CNN ar-
chitectures with pre-trained weights from ImageNet allow us to use
a technique called fine-tuning. Fine-tuning is the process of slowly
8
H.R. Firmansyah et al. Engineering Applications of Artificial Intelligence 131 (2024) 107824
Box I.
changing the parameters that have been studied in the model in solving perfect misclassification and perfect classification, respectively, while
the previous problem (ImageNet) (Chu et al., 2016), and reused to solve MCC = 0 is the expected value for the random classifier.
the specific problem presented in this study, which is the classification 𝑇𝑃 × 𝑇𝑁 − 𝐹𝑃 × 𝐹𝑁
of building typology. MCC = √ (5)
(𝑇 𝑃 + 𝐹 𝑃 )(𝑇 𝑃 + 𝐹 𝑁)(𝑇 𝑁 + 𝐹 𝑃 )(𝑇 𝑁 + 𝐹 𝑁)
Fine-tuning is carried out starting from the top convolution and
fully-connected components of the CNN system, by freezing the param- Therefore, in this study, f1-score and MCC will be used as the
eters in the previous layers (Ferlitsch, 2021). The purpose of this freeze primary metrics for comparing performances across distinct datasets
operation is the assumption that simple features from the initial layers trained with various CNN models, and accuracy will be used as sup-
of the pre-trained model are adequate to learn simple features common porting metrics.
to every object in nature such as lines and angles.
Then, progressively from the top layer, each prior layer will be
3.3.6. Visual evaluation
unfrozen to allow the model to gradually change the more funda-
mental parameters (Castro et al., 2018); if the model simultaneously Convolutional Neural Networks are known for their black box sys-
unfreezes all the layers during fine-tuning, there is a great risk that tem. It is due to difficulty in visualizing how the system processes the
the model will overfit too quickly, where the training accuracy has images to generate prediction results. Sole utilization of numerical met-
reached nearly 100% while the validation accuracy is very low, leading rics, in this case, may not be enough to fully understand the behavior
to generalization issues. of this model. Furthermore, even if the metrics are showing exceptional
While the process of unfreezing CNN layer by layer is executed results, they may not guarantee similar performance when the model is
gradually, the model will be retrained with an incrementally smaller deployed in the actual production. However, recent advances in deep
learning rate, with the aim that the model does not change low-level learning have discovered several algorithms to aid in elucidating how
parameters (textures, lines, angles, etc.) quickly and increasing the risk CNN works. In addition to performance metrics, for this study, we will
of overfitting. evaluate the model’s ability through the use of the GradCAM algorithm.
GradCAM (Gradient-weighted Class Activation Mapping) is a CNN
3.3.5. Performance metrics visualization technique published in a study from Selvaraju et al.
There are several methods to measure performance in classification (2019) to allow us to have better knowledge of CNN’s decision-making
tasks in deep learning. The confusion matrix is commonly used to steps in determining building typology from images. Using GradCAM
measure the performance of classification problems solved by machine simplifies the CNN performance evaluation process. This is due to the
learning (Provost et al., 1998). The columns in the confusion matrix transparency of the model which has high complexity. It has been
show the predicted classes, while the rows in the confusion matrix show described previously that the bottom layer of CNN detects simple
the actual classes. features such as lines and curves, while the top layer detects more
Accuracy measures how often the model is right by dividing the complex features such as the predicted type of building object.
number of correct predictions by the total number of predictions.
However, accuracy metrics may be less useful when faced with cases
of class imbalances within the dataset (see Box I). 3.3.7. Model selection
The issue of dataset class imbalance can be solved by using the f1- For this study, each model will be trained on 3 different resampled
score as the alternative performance metric. This measure requires both distributions, aiming to reach maximum validation performance. Then,
precision and recall to be high to achieve a better performance. Huč each model will be tested with the test samples of the respective
et al. (2021). resampled distribution to achieve model unbiased performance. Fi-
2 (Precision × Recall) nally, for every model, their f1-score performance for each class, and
𝑓1 -score = (2) weighted f1-score performance will be compared with models trained
Precision + Recall
with different sample sizes, data sources, and CNN architecture. This
Where precision measures the model’s ability to correctly predict
the target building typologies. Achieving the perfect precision of 100% comparison will yield insights into how several datasets and model
means the model is always correct when predicting the target building characteristics can affect performance. In addition to f1-score metrics,
typology: we will compare the model’s ability to differentiate building typology
features through GradCAM visualization to gain a more comprehensive
𝑇𝑃
Precision = (3) analysis of the model performance.
𝑇𝑃 + 𝐹𝑃
and recall measures the model’s ability to find all the relevant building
typologies within the dataset : 3.4. Phase III-Seismic vulnerability mapping
𝑇𝑃
Recall = (4)
𝑇𝑃 + 𝐹𝑁 After identifying the best-performing CNN model in Phase II, the
On the other hand, the Matthew Correlation Coefficient (MCC) are model’s capability to predict the typology of buildings in certain areas
additional performance measure that only produces a high score if the will be tested. First, GSV Images and their coordinate will be acquired,
prediction obtained excellent results in all of the four confusion matrix then the typology will be predicted. With the given coordinates, the
categories (true positives, false negatives, true negatives, and false peak ground acceleration (PGA) for a specific scenario is estimated.
positives) proportionally both to the number of positive and negative Then, using the vulnerability class associated with the predicted typol-
elements in the dataset (Chicco and Jurman, 2020). MCC ranges in the ogy, the most possible damage grade experienced in each building will
interval [−1, +1], with extreme values −1 and +1 reached in case of be estimated and mapped.
9
H.R. Firmansyah et al. Engineering Applications of Artificial Intelligence 131 (2024) 107824
Fig. 6. The distribution of the generated coordinates with available GSV images to be reviewed within the rectangular boundary condition and the location of Lembang Fault
which will be the earthquake scenario observed within the study.
3.4.1. Study area There are 1538 generated coordinates of which available GSV im-
In this study, the selected city which will be observed is the greater ages exist that also show building facades. The distribution of these
Bandung area, one of the largest and most populated cities in Indonesia building coordinates can be seen in Fig. 6.
close to a major active fault. Bandung is located on the island of
3.4.2. Peak ground acceleration (PGA) and modified Mercalli intensity
Java, Indonesia (see Fig. 5). Then, a rectangular boundary will be
(MMI) calculation
made (North: 6.8684S, South: 6.9709S, West: 107.55587E and East:
For this study, the possibility of a major earthquake due to the
107.67844E; see Fig. 6 for illustration), inside which 1538 coordinates
Lembang Fault movement is explored. The Lembang Fault is a major
will be generated randomly with a uniform distribution. Then, for each
fault in western Java that skirts the northern edge of Bandung, just
generated coordinate, the closest GSV image will be mined with a south of the active Tangkuban Perahu volcano. The fault is accom-
search radius of 50 m. If there is a street view that can be accessed by modating a parallel trench slip that is a result of a slight obliquity
the GSV system in that radius, images will be taken and their typology in plate convergence at the Java Trench. With a length of 29 km,
will be predicted using this deep learning system. The overview of this has suggested that the Lembang Fault could produce a Mw 6.5–
the street-view image acquisition method applied in this study area is 7.0 earthquake with a recurrence time of 170–670 years (Daryono
illustrated in Fig. 4 et al., 2019). Therefore, in this study, the earthquake scenario that will
10
H.R. Firmansyah et al. Engineering Applications of Artificial Intelligence 131 (2024) 107824
Table 5
Building types according to EMS-98 and their qualitative probabilistic correlations to the vulnerability classes for (i)
original EMS-98, and (ii) modified EMS-98 from Maqsood et al. (2013). Numerical probabilistic assignments for the EMS-98
vulnerability classes are derived from fuzzy sets developed by Bernardini et al. (2010).
be considered is a magnitude 7 Mw scenario plausible for the given from Bernardini et al. (2010) to randomly generate the vulnerability
Lembang Fault. class of the building. There are 6 vulnerability classes present in EMS-
The earthquake attenuation function of the Lembang fault uses the 98, denoted alphabetically from A (highest vulnerability) to F (lowest
Boore attenuation function (Boore et al., 1997) with the equation as vulnerability), to describe the ability of different types of structures to
follows: withstand seismic loads. The EMS-98 building typology classification
𝑉 gives not only the most probable vulnerability class but a range of
ln 𝑌 = 𝑏1 + 𝑏2 (𝑀 − 6) + 𝑏3 (𝑀 − 6)2 + 𝑏5 ln 𝑟 + 𝑏𝑣 ln 𝑆 (6)
𝑉𝐴 classes for most types of building types. Next, fuzzy sets are used to esti-
√
mate the probability of occurrence of a certain level of building damage
𝑟 = 𝑟2𝑐𝑙 + ℎ2 (7)
due to a specific earthquake MMI intensity. Fuzzy sets are commonly
Where 𝑌 is the ground motion parameter (peak horizontal acceleration used to represent some form of uncertainty and are widely used in other
(PGA) in g); M is (moment) magnitude; 𝑟𝑐𝑙 is the closest distance from technological advances such as home environment control Wozniak
the station to a site of interest in km; 𝑉𝑠 is the shear wave velocity et al. (2021). For this study, these fuzzy sets are used to measure the
(taken here at 1070 m/s (Irsyam et al., 2000)); while 𝑏1 , 𝑏2 , 𝑏3 , 𝑏5 , linguistic frequencies (e.g., ‘‘few’’, ‘‘many’’, ‘‘most’’) of damage grade
ℎ, 𝑏𝑣 and 𝑉𝐴 are parameters that are dependent on the earthquake distribution corresponding to each vulnerability class in the EMS-98
mechanism and assumptions. As the Lembang fault is classified as a scale, developed by Bernardini et al. (2010). This probability of damage
strike-slip earthquake, the above parameters are taken as 𝑏1 = −0.313, grade level occurrence is presented in a matrix known as the damage
𝑏2 = 0.527, 𝑏3 = 0, 𝑏5 −0.778, ℎ = 5.57 km, 𝑏𝑣 and 𝑉𝐴 = 1396 m∕s (Irsyam probability matrix (DPM).
et al., 2000). Two vulnerability class definitions are applied. The first defini-
After obtaining the 𝑌 value, the site coefficients can be calculated tion aligns with the original EMS-98 framework, which was initially
which amplifies earthquake waves from bedrock to the ground sur- devised based on the building characteristics prevalent in Europe. Con-
face by linear interpolation Table 10 from Indonesia’s Building Code versely, the second definition utilizes a modified EMS-98 vulnerability
Standards (Badan Standardisasi Nasional, 2019). definition that more effectively captures local construction practices
𝑌𝑠 = 𝐶𝑌 × 𝑌 (8) specifically tailored to Indonesia (see the comparison of the two models
in Table 5). This modification is derived from a comprehensive exam-
Where 𝑌𝑠 is the ground motion at surface level in g, and 𝐶𝑌 is ination of numerous photographs depicting the damage in residential
a multiplier factor dependant on the soil class (Badan Standardisasi buildings during the 2009 Sumatra earthquakes. Maqsood et al. (2013).
Nasional, 2019). In this study, the soil class in Bandung was assumed to
be soft soil (SE). Furthermore, after obtaining PGA at the surface level,
4. Results and discussion
the Modified Mercalli Intensity (MMI) felt by the building at the point
under consideration is calculated, with the equation (Panjamani et al.,
2016): 4.1. Labeling
MMI = 0.1417 + 3.2335 ∗ 𝑙𝑜𝑔10 𝑌𝑠 (9) Following the labeling of the dataset, a straightforward evaluation
This MMI is then used to estimate the damage grade for a given was conducted by comparing a subset of the dataset utilizing the
typology. suggested building typology decision tree with the ensemble label
predictions generated by a panel of five building experts. A total of
3.4.3. Building damage grade mapping based on predicted building typology 46 unlabeled building images were presented to the experts, who were
After obtaining the typology of a building using CNN at a particular tasked with selecting the most appropriate building typology corre-
coordinate and knowing the corresponding MMI, the level of damage sponding to each image. The analysis revealed that 43 out of the 46
to the building could be estimated in the study area. Firstly, the undergraduate student predictions concurred with the expert ensemble
numerical probabilities of the EMS-98 vulnerability classes are taken predictions, resulting in an accuracy rate of 93.48%.
11
H.R. Firmansyah et al. Engineering Applications of Artificial Intelligence 131 (2024) 107824
Fig. 7. Examples of randomly sampled image resulting in misidentification from the experts, due to (a) multiple building facades appearing in one image, or (b) obscure structural
details.
12
H.R. Firmansyah et al. Engineering Applications of Artificial Intelligence 131 (2024) 107824
13
H.R. Firmansyah et al. Engineering Applications of Artificial Intelligence 131 (2024) 107824
Table 6
Performance summary, averaged from 3 resampled distributions for (a) F1-score, (b) Matthew Correlation Coefficient (MCC), and (c) Accuracy.
this photo is UC. Fig. 11(b)(ii) shows the decision-making process of the wall pattern in the photo resulting in a building typology decision that
CNN model that correctly predicts the typology of this building as UC. is UC, and discarding the vehicles as the main differentiator of building
The identification process starts from the stem convolution layer which typology. From Fig. 11(c)(i), it can be seen that the original photo of the
recognizes simple objects such as vehicle edges. Then the next layers building has a wooden wall, so the ground-truth typology of this photo
gradually recognize walls and bigger parts of vehicles based on the is the TB structure. Fig. 11(c)(ii) shows the decision-making process of
results of the identification of the previous layers. In the end, the last the CNN model which correctly predicts the typology of this building
layer (top convolution) is making concrete decisions based on the brick as TB structure.
14
H.R. Firmansyah et al. Engineering Applications of Artificial Intelligence 131 (2024) 107824
Fig. 12. Confusion matrix results for images trained with (a) MobileNetV3L (MBL) architecture, a second set of resampled data distribution, and various image sources (b) a full
version of randomly sampled GSV images, first set of resampled data distribution, and various CNN architecture.
Fig. 13. Model incorrectly predicted unconfined masonry structures as other typology class.
4.2.6. Confusion matrix evaluation in experiment standard deviation can be attributed to several factors.
In this subsection, we will assess the experimental outcomes through Firstly, the extensive environmental variations present in UC pho-
the examination of the generated confusion matrices. Fig. 12(a) il- tos pose a challenge in accurately identifying the absence of column
lustrates confusion matrices of MBL CNN architectures trained with structures. Additionally, the random sampling method employed for
various image sources, while Fig. 12(b) showcases the confusion ma- photo selection may yield suboptimal angles that further exacerbate
trices of randomly sampled GSV images trained across different CNN this issue. Conversely, the GSV virtual tour method ensures screenshots
architectures. Both sets of confusion matrices indicate that the most are captured from angles that optimize the visibility of UC building
prevalent misclassification occurs when the model wrongly identifies features. Consequently, dataset VT-L exhibits slightly better average
the UC class as CM. This occurrence can be attributed to the similarities UC typology performance compared to dataset RS-L, although it has
between UC and CM, where buildings with exposed bricks may exhibit a smaller size of the training sample. To mitigate the issue related to
visible columns that resemble CM structures. However, in cases where these UC images, it is advisable to provide the model with more UC
walls feature a mixture of brick types, the model may incorrectly samples during training and to incorporate UC images captured from
interpret the variations in brick types as columns, as depicted in more diverse angles.
Fig. 13(a).
Another interesting observation is that buildings surrounded by 4.3. Seismic vulnerability mapping
abundant vegetation are more likely to be predicted as TB, as shown
in Fig. 13(b). This phenomenon may arise from the fact that the The presented map depicts the mapping results of 1538 build-
majority of TB samples in the datasets were obtained from rural areas in ing images in Bandung, obtained using the randomly sampled GSV
Indonesia, such as East Nusa Tenggara, due to the challenging nature method. The building typology for each image was determined using
of acquiring TB images in urban settings. This pattern within the TB the most accurate CNN model trained specifically with randomly sam-
datasets leads the CNN model to associate trees and bushes surrounding pled GSV images. The distribution of building typologies can be seen in
buildings with TB typology. One potential solution to address this Fig. 14(a), where the majority of the structures identified are CM and
issue is to collect more TB building images in urban areas with fewer RC.
trees and bushes in the vicinity or to incorporate additional images of Then, after predicting the building typology, the following vul-
different typology classes obtained from similar rural regions. nerability map is obtained for the city of Bandung with the most
Considering the unique case presented in Fig. 10, the decrease in possible damage grade levels ranging from damage grade 0 to 5, where
the average performance of UC typology coupled with an increase damage grade 0 corresponds to no damage to buildings, and damage
15
H.R. Firmansyah et al. Engineering Applications of Artificial Intelligence 131 (2024) 107824
grade 5 represents destruction to affected buildings. Each point size fragility curve estimation, although the used fragility curve employs
corresponds to the relative probability magnitude of the most possible results based on the Padang Earthquake in Sumatera, Indonesia, it still
damage grade experienced in the particular building. only showcases the behavior of structures in that area. It is plausible
For the first case, when considering the vulnerability of buildings that different building construction practices in other areas and islands
according to the EMS-98 standard, which was originally devised based will result in different fragility estimations. Additionally, other types of
on building typologies prevalent in Europe, the mapping results are image acquisition methods could be conducted to cover the rural area
depicted in Fig. 14(b). By analyzing Fig. 14(b), it becomes evident that is currently inaccessible to Google Street View, producing a more
that the southern region of Bandung City shows a large proportion representative seismic vulnerability map.
of buildings remaining undamaged in the event of a magnitude 7
earthquake scenario originating from the Lembang fault. Only 2.4% of 4.4. Limitations
the buildings in Bandung City possess the potential for experiencing
severe structural damage, while a mere 0.07% of the buildings are The previous subsection demonstrates that CNN automated building
projected to face a complete devastation level of damage. In contrast, typology detection can be used to estimate seismic vulnerability in wide
an estimated 54% of the buildings in the southern region of Bandung regions quickly. There are still some challenges and limitations, such as:
are expected to remain unharmed.
• The decision tree, which concentrates on visible characteristics
Given the origin of EMS-98 in Europe, which primarily accounts
in building images, may lead to misclassification stemming from
for the building morphology prevalent in that region, it is reasonable
variations in building facades and the actual structural system.
to observe that CM and RC building types are presumed to exhibit
For instance, a building with RC structures featuring wooden fa-
adequate earthquake resistance, with the majority falling into class D.
cades, while may be uncommon, may be misclassified as TB struc-
However, in the context of Indonesia, despite being prone to earth-
tures. Similarly, buildings with traditional aesthetics are prone to
quakes, construction practices often deviate from optimal standards.
being predicted as a typology other than RC. This highlights the
Common shortcomings include the absence of reinforcing bars at joints,
need to update the models with relevant images along with the
inadequate spacing on stirrups, and insufficient anchorage on reinforce-
development of building architectural trends in Indonesia over
ment bars. Hence, Fig. 14(b) might underestimate the hazards resulting
time.
from this earthquake scenario in the dense city of Bandung.
• The dataset from camera images (CI) is limited, mainly collected
Next, a subsequent mapping was conducted utilizing the EMS-
in Bandung, West Java, Indonesia, due to time constraints. In
98 definition, incorporating modifications to the vulnerability classes
contrast, virtual tour and randomly sampled GSV images are
tailored for the unique building conditions in Indonesia as shown in
captured from various locations across Indonesia accessible with
Table 5. As a result of these adaptations, the majority of CM and
GSV. This discrepancy in data collection may introduce bias in
RC buildings were assigned to vulnerability class C, with vulnerability
terms of dataset diversity when comparing the performance of the
class B as the next most probable class. This modification highlights
models with each other. Future research with extensive camera
the lower resilience of buildings in Indonesia when compared to their
images gathered around Indonesia would be highly beneficial.
European counterparts. Analysis of Fig. 14(c) reveals that nearly all
• The limitations of the GSV system, constrained by vehicle restric-
buildings located on the northern side of Bandung possess the potential
tions, prevent access to slums and rural areas. Consequently, the
to experience moderate to severe levels of structural damage. Merely a
variety of building types represented in the study is limited.
meager 6% of the buildings are anticipated to experience minimal to no
damage. Overall, it is estimated that a minimum of 55% of the buildings However, this is a significant step towards achieving a better and
in Bandung would encounter moderate to severe structural damage. more precise data-driven assessment system that can aid the govern-
Considering the existing construction practices in Indonesia, Fig. 14(c) ment in preparing and mitigating the potential economic damage and
portrays a more sensible outcome in the event of an M7 earthquake loss of life in the next earthquake event.
with an epicenter along the Lembang fault. It provides a conservative Considering Indonesia’s unique status as an archipelagic country
representation to stakeholders, offering insight into the scale of impact with diverse cultures, architectural styles, and building construction
expected from such an earthquake scenario. However the distinct risk methods, further research development should be pursued to identify
prediction using different fragility curves confirms the need for better vulnerabilities specific to these distinctive building types. Utilizing
16
H.R. Firmansyah et al. Engineering Applications of Artificial Intelligence 131 (2024) 107824
recent advancements in technology, Sarli et al. (2022) demonstrates object localization and detection. This refinement could reduce the
the rapid generation of vulnerability curves and prediction of building need to filter out GSV images with just one building. Introducing
collapse probabilities based on experimental outcomes. Gathering a background class, especially when the randomly sampled method did
limited amount of field data about structural dimensions is sufficient not yield building images could also be applied to improve the city-
to establish a vulnerability curve model. This approach enables the scale seismic vulnerability assessment speed. Alternative approaches
development of advanced image processing models, capable of accom- such as combining building imagery from aerial view and street view
modating diverse typologies found in different regions across Indonesia. could also be used to obtain more information to detect the building
Another study by Sandoli et al. (2023) had recently developed an Urban typologies. Using semantic segmentation techniques to extract rele-
Fragility Matrix that represents the vulnerability of selected urban vant earthquake-resisting elements in buildings could also be useful
sectors, potentially creating better population seismic risk estimates. to support the building typology detection algorithms. The automation
Incorporating these advances in fragility estimations has the potential of data mining and typology prediction offers significant time-saving
to produce more comprehensive and representative mappings, thus benefits for conducting vulnerability assessments at a regional scale.
enhancing the overall earthquake risk assessment process.
CRediT authorship contribution statement
5. Conclusions and future directions
Hafidz R. Firmansyah: Conceptualization, Methodology, Software,
This study used a combination of camera, GSV virtual tour, and Validation, Formal analysis, Investigation, Data curation, Writing –
randomly sampled GSV images, with a manually annotated dataset original draft, Visualization. Prasanti Widyasih Sarli: Conceptualiza-
totaling more than 4500 images at the street level captured in various tion, Methodology, Writing – review & editing, Supervision, Project
areas in Indonesia to discover the prospective benefit of using CNN to administration, Funding acquisition. Andru Putra Twinanda: Method-
classify building typologies. ology, Supervision. Devin Santoso: Data curation. Iswandi Imran:
Varying sample size to test the performance improvement shown Conceptualization, Supervision, Funding acquisition.
that doubling the sample size brings f1-score performance increase by
1.5% on average, with GSV virtual tour images, whereas quadrupling
Declaration of competing interest
the sample size leads to similar performance improvement by up to
2.34%, with randomly sampled GSV images. The best sources of im-
The authors declare the following financial interests/personal rela-
ages to allow CNN to learn distinct features in building typology are
tionships which may be considered as potential competing interests:
randomly sampled images, with the f1-score performance of 83.99%,
Iswandi Imran reports financial support was provided by Bandung
keeping the constant sample size. The best results of one of the devel-
Institute of Technology.
oped models achieved a weighted f1-score of 88.2% when identifying
building typologies. From testing of 4 distinct network architectures.
Among the four CNN architectures trained in this study, MobileNet V3L Data availability
showed the best performance because this network achieved the top-
weighted f1-score when trained and tested in 3 of 5 datasets in this Data will be made available on request.
study.
Declaration of Generative AI and AI-assisted technologies in the
For the individual building typology class, CM consistently is the
writing process
easiest typology to predict, with a minimum average performance of
74.67% up to 90.58%, this may be caused by CM being the superma-
jority class, which comprises about 60% of images in the unaugmented During the preparation of this work the author(s) used ChatGPT
dataset. On the other side, UC typology, due to its limited samples, to improve readability. After using this tool/service, the author(s) re-
only comprises about 5%–8% of images in the dataset, and consistently viewed and edited the content as needed and take(s) full responsibility
have the lowest performance compared to other typology class, with for the content of the publication.
a minimum average performance of 31.17% up to 57.5%. This result
shows that obtaining more images, especially from the minority class Acknowledgments
is crucial to achieving better performance.
Based on the findings of the building vulnerability mapping con- This research was funded through the Program Penelitian, Pengab-
ducted in Bandung, it became evident that variations in fragility curve dian kepada Masyarakat dan Inovasi ITB (P2MI) administered by Insti-
models for different building typologies led to notable disparities in tut Teknologi Bandung, Indonesia.
the resulting mappings. The original EMS-98 model, developed with
European buildings as a reference, assumed acceptable construction References
qualities for CM and RC structures, thereby creating the impression
of a city that is resilient to significant earthquake events. Conversely, Acevedo, A.B., Jaramillo, J.D., Yepes, C., Silva, V., Osorio, F.A., Villar, M., 2017.
Evaluation of the seismic risk of the unreinforced masonry building stock in
the modified EMS-98 model, which incorporates vulnerability classes
Antioquia, Colombia. Nat. Hazards 86, 31–54.
adjusted to reflect the construction quality in Indonesia, produced Afifi, M., Brown, M.S., 2019. What else can fool deep learning? Addressing color
more sensible outcomes. In this model, the majority of buildings were constancy errors on deep neural network performance. In: Proceedings of the
observed to experience moderate to severe structural damage, aligning IEEE/CVF International Conference on Computer Vision. pp. 243–252.
with the actual scenario expected during such events. Arif, Z.H., Mahmoud, M.A., Abdulkareem, K.H., Kadry, S., Mohammed, M.A., Al-
Mhiqani, M.N., Al-Waisy, A.S., Nedoma, J., 2022. Adaptive deep learning detection
Future research endeavors can be directed towards several im-
model for multi-foggy images. Int. J. Interact. Multimedia Artif. Intell. 7 (7).
portant areas. Firstly, it is crucial to explore the applicability of the Arya, A.S., Boen, T., Ishiyama, Y., 2014. Guidelines for Earthquake Resistant
proposed approach in additional regions, encompassing building ty- Non-Engineered Construction. UNESCO.
pologies that were not covered in the current study, such as steel Badan Standardisasi Nasional, 2019. SNI 1726-2019: Tata Cara Perencanaan Ketahanan
structures and high-rise buildings with moderate to high levels of Gempa Untuk Struktur Bangunan Gedung Dan Nongedung. Badan Standardisasi
Nasional (BSN).
earthquake-resistant design (ERD). Furthermore, given the existing ef-
Barai, M., Heikkinen, A., 2017. Impact of data augmentations when training the
fectiveness issues associated with the data mining process conducted inception model for image classification.
on the GSV API, there is room for further refinement of the data Bergstra, J., Bengio, Y., 2012. Random search for hyper-parameter optimization. J.
mining algorithm and usage of more advanced tasks in CNN such as Mach. Learn. Res. 13 (2).
17
H.R. Firmansyah et al. Engineering Applications of Artificial Intelligence 131 (2024) 107824
Bernardini, A., Lagomarsino, S., Mannella, A., Martinelli, A., Milano, L., Parodi, S., Naito, S., Tomozawa, H., Mori, Y., Nagata, T., Monma, N., Nakamura, H., Fujiwara, H.,
2010. Forecasting seismic damage scenarios of residential buildings from rough Shoji, G., 2020. Building-damage detection method based on machine learning
inventories: a case-study in the Abruzzo Region (Italy). Proc. Inst. Mech. Eng. O utilizing aerial photographs of the Kumamoto earthquake. Earthq. Spectra 36 (3),
224 (4), 279–296. 1166–1187.
BNPB Indonesia, 2020. Rencana Nasional Penanggulangan Bencana 2020-2024. O’Shea, K., Nash, R., 2015. An introduction to convolutional neural networks. arXiv:
Boore, D.M., Joyner, W.B., Fumal, T.E., 1997. Equations for estimating horizontal 1511.08458.
response spectra and peak acceleration from western North American earthquakes: Panjamani, Bajaj, Moustafa, Al-Arifi, 2016. Relationship between intensity and recorded
A summary of recent work. Seismol. Res. Lett. 68 (1), 128–153. ground-motion and spectral parameters for the himalayan region. Bull. Seismol.
Brzev, S., Mitra, K., 2007. Earthquake-Resistant Confined Masonry Construction. Soc. Am. 106 (4).
National Information Center of Earthquake Engineering (NICEE), India. Pelizari, P.., Geiß, C., Aguirre, P., Maria, H., Pena, Y., Taubenbock, H., 2021. Automated
Castro, F.M., Marín-Jiménez, M.J., Guil, N., Schmid, C., Alahari, K., 2018. End-to-end building characterization for seismic risk assessment using street-level imagery and
incremental learning. In: Proceedings of the European Conference on Computer deep learning. ISPRS J. Photogramm. Remote Sens. 180, 370–386.
Vision. ECCV, pp. 233–248. Perez, L., Wang, J., 2017. The effectiveness of data augmentation in image classification
Charleson, A., Brzev, S., Jaiswal, K., Greene, M., 2017. Improving housing seismic using deep learning. arXiv:1712.04621.
safety in developing countries: The world housing encyclopedia. In: Proc. 16th Perol, T., Gharbi, M., Denolle, M., 2018. Convolutional neural network for earthquake
World Conference on Earthquake Engineering. Santiago, Chile. detection and location. Sci. Adv. 4 (2), e1700578.
Pittore, M., Wieland, M., Fleming, K., 2017. Perspectives on global dynamic exposure
Chicco, D., Jurman, G., 2020. The advantages of the Matthews correlation coefficient
modelling for geo-risk assessment. Nat. Hazards 86 (1), 7–30.
(MCC) over F1 score and accuracy in binary classification evaluation. BMC
Prashanth, D.S., Mehta, R.V.K., Sharma, N., 2020. Classification of handwritten devana-
Genomics 21 (6).
gari number–an analysis of pattern recognition tool using neural network and CNN.
Chollet, F., 2017. Xception: Deep learning with depthwise separable convolutions. In:
Procedia Comput. Sci. 167, 2445–2457.
Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.
Provost, F.J., Fawcett, T., Kohavi, R., et al., 1998. The case against accuracy estimation
pp. 1251–1258.
for comparing induction algorithms. In: ICML, Vol. 98. pp. 445–453.
Chu, B., Madhavan, V., Beijbom, O., Hoffman, J., Darrell, T., 2016. Best practices
Rahman, T., Khandakar, A., Qiblawey, Y., Tahir, A., Kiranyaz, S., Kashem, S.B.A.,
for fine-tuning visual classifiers to new domains. In: Computer Vision–ECCV
Islam, M.T., Al Maadeed, S., Zughaier, S.M., Khan, M.S., et al., 2021. Exploring
2016 Workshops: Amsterdam, the Netherlands, October 8-10 and 15-16, 2016,
the effect of image enhancement techniques on COVID-19 detection using chest
Proceedings, Part III 14. Springer, pp. 435–442.
X-ray images. Comput. Biol. Med. 132, 104319.
CRED, UNDRR, 2020. Human cost of disasters. An overview of the last 20 years: Rildova, D., Suarjana, D., Pribadi, K., 2012. Experimental study on the behaviour of
2000–2019. plastered confined masonry wall under lateral cyclic load. In: 15th World Confrence
Cui, L., Jing, X., Wang, Y., Huan, Y., Xu, Y., Zhang, Q., 2022. Improved swin of Earthquake Engineering.
transformer-based semantic segmentation of postearthquake dense buildings in Saikia, T., Schmid, C., Brox, T., 2021. Improving robustness against common corrup-
urban areas using remote sensing images. IEEE J. Sel. Top. Appl. Earth Obs. Remote tions with frequency biased models. In: Proceedings of the IEEE/CVF International
Sens. 16, 369–385. Conference on Computer Vision. pp. 10211–10220.
Daryono, M.R., Natawidjaja, D.H., Sapiie, B., Cummins, P., 2019. Earthquake geology Salgado-Gálvez, M.A., Zuloaga-Romero, D., Bernal, G.A., Mora, M.G., Cardona, O.D.,
of the lembang fault, West Java, Indonesia. Tectonophysics 751, 180–191. 2014. Fully probabilistic seismic risk assessment considering local site effects for
Dimitrakakis, C., Savu-Krohn, C., 2008. Cost-minimising strategies for data labelling: op- the portfolio of buildings in Medellín, Colombia. Bull. Earthq. Eng. 12, 671–695.
timal stopping and active learning. In: Foundations of Information and Knowledge Sandoli, A., Brandonisio, G., Lignola, G., Prota, A., Fabbrocino, G., 2023. Seismic
Systems: 5th International Symposium, FoIKS 2008, Pisa, Italy, February 11-15, fragility matrices for large scale probabilistic structural safety assessment. Soil Dyn.
2008. Proceedings 5. Springer, pp. 96–111. Earthq. Eng. 171, 107963.
EERI, 2019. EERI Policy White Paper. Technical Report, Earthquake Engineering Sarli, P., Palar, P., Azhari, Y., Setiawan, A., Sanjaya, Y., S.C., S., Imran, I.,
Research Institute (EERI), pp. 2–6. 2022. Gaussian process regression for seismic fragility assessment: Application to
Ferlitsch, A., 2021. Deep Learning Patterns and Practices. Manning. non-engineered residential buildings in Indonesia. Buildings 13, 59.
Gong, F.Y., Zeng, Z.C., Zhang, F., Li, X., Ng, E., Norford, L.K., 2018. Mapping sky, tree, Sazedj, S., Morais, A., Jalali, S., 2013. Comparison of Costs of Brick Construction and
and building view factors of street canyons in a high-density urban environment. Concrete Structure Based on Functional Units. Universidade de Minho, Tecminho,
Build. Environ. 134, 155–167. Luis Bragança, Manuel Pinheiro, Ricardo . . . .
Gonzalez, D., Rueda-Plata, D., Acevedo, A.B., Duque, J.C., Ramos-Pollan, R., Betan- Selvaraju, R.R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., Batra, D., 2019. Grad-
court, A., Garcia, S., 2020. Automatic detection of building typology using deep CAM: Visual explanations from deep networks via gradient-based localization. Int.
learning methods on street level images. Build. Environ. 177, 106805. J. Comput. Vis. 128 (2), 336–359.
Grunthal, G., 1998. Chaiers du Centre Européen de Géodynamique et de Séismologie: Sharon, S., 2022. Development of a Simple Vulnerability Assessment Method for
Volume 15—European Macroseismic Scale 1998. European Center for Geodynamics Confined Masonry-type Residential Houses for Indonesian Big Data (Case Study:
and Seismology, Luxembourg. Cikahuripan Village, West Bandung Regency) (Master’s thesis). Bandung Institute
Hao, R., Namdar, K., Liu, L., Haider, M.A., Khalvati, F., 2021. A comprehensive study of Technology (in Indonesian).
of data augmentation strategies for prostate cancer detection in diffusion-weighted Shorten, C., Khoshgoftaar, T.M., 2019. A survey on image data augmentation for deep
MRI using convolutional neural networks. J. Digit. Imaging 34, 862–876. learning. J. Big Data 6 (1), 1–48.
Harirchian, E., Lahmer, T., Buddhiraju, S., Mohammad, K., Mosavi, A., 2020. Earth- Shukla, A.K., Kumar, S., Maiti, P., 2021. Failure analysis of unconfined brick masonry
quake safety assessment of buildings through rapid visual screening. Buildings 10 with experimental verification. J. Fail. Anal. Prev. 21 (2), 419–428.
Silva, V., Amo-Oduro, D., Calderon, A., Dabbeek, J., Despotaki, V., Martins, L., Rao, A.,
(3), 51.
Simionato, M., Viganò, D., Yepes, C., et al., 2018. Global earthquake model (GEM)
Howard, A., Sandler, M., Chu, G., Chen, L.C., Chen, B., Tan, M., Wang, W., Zhu, Y.,
seismic risk map (version 2018.1).
Pang, R., Vasudevan, V., et al., 2019. Searching for mobilenetv3. In: Proceedings
Stachl, C., Pargent, F., Hilbert, S., Harari, G.M., Schoedel, R., Vaid, S., Gosling, S.D.,
of the IEEE/CVF International Conference on Computer Vision. pp. 1314–1324.
Bühner, M., 2020. Personality research and assessment in the era of machine
Huč, A., Šalej, J., Trebar, M., 2021. Analysis of machine learning algorithms for
learning. Eur. J. Pers. 34 (5), 613–631.
anomaly detection on edge devices. Sensors 21 (14), 4946.
Struyk, R.J., Hoffman, M.L., Katsura, H.M., 1990. The Market for Shelter in Indonesian
Irsyam, M., Himawan, A., Subki, B.A., Suntoko, H., 2000. Analisis seismisitas untuk
Cities. The Urban Insitute.
semenanjung muria. Jurnal Pengembangan Energi Nuklir 2 (2).
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Van-
Kasapbaşi, A., Elbushra, A.E.A., Omar, A.H., Yilmaz, A., 2022. DeepASLR: A CNN houcke, V., Rabinovich, A., 2015. Going deeper with convolutions. In: Proceedings
based human computer interface for American Sign Language recognition for of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 1–9.
hearing-impaired individuals. Comput. Methods Programs Biomed. Update 2, Tan, M., Le, Q., 2019. Efficientnet: Rethinking model scaling for convolutional neural
100048. networks. In: International Conference on Machine Learning. PMLR, pp. 6105–6114.
Khalifa, N.E., Loey, M., Mirjalili, S., 2022. A comprehensive survey of recent trends in Wang, Y., Cui, L., Zhang, C., Chen, W., Xu, Y., Zhang, Q., 2022. A two-stage seismic
deep learning for digital images augmentation. Artif. Intell. Rev. 1–27. damage assessment method for small, dense, and imbalanced buildings in remote
Maqsood, T., Schwarz, J., Edwards, M., 2013. Application of European macroseismic sensing images. Remote Sens. 14 (4), 1012.
scale -1998 in the Asia-Pacific region. pp. 28–30. Wang, Y., Jing, X., Cui, L., Zhang, C., Xu, Y., Yuan, J., Zhang, Q., 2023a. Geometric con-
Meilano, I., Virtriana, R., Hanifa, N., et al., 2019. Basis Data Historis Bencana sistency enhanced deep convolutional encoder-decoder for urban seismic damage
Dan Estimasi Potensi Kerusakan Serta Kerugian Akibat Bencana Di Indonesia (in assessment by UAV images. Eng. Struct. 286, 116132.
Indonesian). Wang, Y., Jing, X., Xu, Y., Cui, L., Zhang, Q., Li, H., 2023b. Geometry-guided semantic
Meng, C., Song, Y., Ji, J., Jia, Z., Zhou, Z., Gao, P., Liu, S., 2022. Automatic segmentation for post-earthquake buildings using optical remote sensing images.
classification of rural building characteristics using deep learning methods on Earthq. Eng. Struct. Dyn. 52 (11), 3392–3413.
oblique photography. In: Build. Simul.. Springer, pp. 1–14. Watanabe, S., Shima, N., Fujita, K., 2013. Research on non-engineered housing
Moroni, O., Ghomez, C., 2002. World Housing Encyclopedia Report. International construction based on a field investigation in Jakarta. J. Asian Archit. Build. Eng.
Association of Earthquake Engineering. 12 (1), 33–40.
18
H.R. Firmansyah et al. Engineering Applications of Artificial Intelligence 131 (2024) 107824
World Health Organization, 2018. World Health Statistics 2018: Monitoring Health for Xu, Y., Wei, S., Bao, Y., Li, H., 2019. Automatic seismic damage identification of
the SDGs, Sustainable Development Goals. World Health Organization. reinforced concrete columns from images by a region-based deep convolutional
Wozniak, M., Zielonka, A., Sikora, A., Piran, M., 2021. 6G-enabled IoT home neural network. Struct. Control Health Monit. 26 (3), e2313.
environment control using fuzzy rules. IEEE Internet Things J. 8, 5442. Zhang, X., 2020. Village-level homestead and Building Floor Area estimates based on
Wu, J., Chen, X.Y., Zhang, H., Xiong, L.D., Lei, H., Deng, S.H., 2019. Hyperparameter UAV imagery and U-net algorithm. ISPRS Int. J. Geo-Inf. 9 (6), 403.
optimization for machine learning models based on Bayesian optimization. J.
Electron. Sci. Technol. 17 (1), 26–40.
19