0% found this document useful (0 votes)
15 views10 pages

Self-Supervised Multisensor Change Detection

This document presents a self-supervised learning method for multisensor change detection (CD) that utilizes bitemporal images captured by optical and synthetic aperture radar (SAR) sensors without requiring labeled data. The proposed approach leverages deep clustering and contrastive learning to effectively identify changes in complex urban scenes, demonstrating its versatility and efficacy through experimental evaluations on various datasets. The method addresses the challenges posed by the significant differences in image characteristics between optical and SAR sensors, making it suitable for real-world applications such as disaster management and urban monitoring.

Uploaded by

marcos
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
15 views10 pages

Self-Supervised Multisensor Change Detection

This document presents a self-supervised learning method for multisensor change detection (CD) that utilizes bitemporal images captured by optical and synthetic aperture radar (SAR) sensors without requiring labeled data. The proposed approach leverages deep clustering and contrastive learning to effectively identify changes in complex urban scenes, demonstrating its versatility and efficacy through experimental evaluations on various datasets. The method addresses the challenges posed by the significant differences in image characteristics between optical and SAR sensors, making it suitable for real-world applications such as disaster management and urban monitoring.

Uploaded by

marcos
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING, VOL.

60, 2022 4405710

Self-Supervised Multisensor Change Detection


Sudipan Saha , Member, IEEE, Patrick Ebel, and Xiao Xiang Zhu , Fellow, IEEE

Abstract— Most change detection (CD) methods assume that a crucial step for several applications, including disaster man-
prechange and postchange images are acquired by the same agement, urban monitoring, forestry, glacier monitoring, and
sensor. However, in many real-life scenarios, e.g., natural dis- precision agriculture. Considering the variation of applications,
asters, it is more practical to use the latest available images
before and after the occurrence of incidence, which may be rarity of occurrences of some change-inducing incidents (e.g.,
acquired using different sensors. In particular, we are interested natural disasters), and large geographic variation, it is impru-
in the combination of the images acquired by optical and dent to assume that large-scale training datasets corresponding
synthetic aperture radar (SAR) sensors. SAR images appear to all such tasks can be ever collected. Thus, there is a
vastly different from the optical images even when capturing significant inclination in the CD literature toward methods that
the same scene. Adding to this, CD methods are often con-
strained to use only target image-pair, no labeled data, and can process the target bitemporal region-of-interest without
no additional unlabeled data. Such constraints limit the scope using any training label or any additional pool of unlabeled
of traditional supervised machine learning and unsupervised images. Motivated by its excellent performance in computer
generative approaches for multisensor CD. The recent rapid vision, researchers have applied deep learning to satellite
development of self-supervised learning methods has shown that image CD [8]. To exploit the potential of deep learning while
some of them can even work with only few images. Motivated
by this, in this work, we propose a method for multisensor CD not using any training label or additional unlabeled images,
using only the unlabeled target bitemporal images that are used transfer learning-based CD methods are popular, which reuse
for training a network in a self-supervised fashion by using a pretrained network for bitemporal feature extraction and
deep clustering and contrastive learning. The proposed method comparison [1].
is evaluated on four multimodal bitemporal scenes showing A striking feature of satellite data is its variability, in terms
change, and the benefits of our self-supervised approach are
demonstrated. Code is available at [Link] of different sensors. Images captured using a passive optical
/tree/main/sarOpticalMultisensorTgrs2021. sensor are quite similar to the natural images studied in com-
puter vision. However, images captured by the active sensors,
Index Terms— Change detection (CD), deep learning, multi-
sensor analysis, self-supervised learning. e.g., synthetic aperture radar (SAR), are remarkably different
from the optical images [9]–[11]. While optical sensors use
I. I NTRODUCTION wavelengths near visible light (approx. 1 μm), SAR uses a
wavelength of 1 cm to 1 m. Moreover, optical sensors rely
O UR earth is rapidly changing, both due to natural and
man-made causes. Satellite image-based change detec-
tion (CD) is generally used to monitor the temporal evolution
upon the natural illumination (e.g., sun) to create the bright-
ness observed by the sensor, while the SAR sensors carry their
of the dynamic earth [1]–[7]. CD ingests bitemporal images as own illumination source, in the form of radio waves transmit-
input and segregates all pixels as changed/unchanged. CD is ted by an antenna. Moreover, satellite images are captured with
a different number of spectral bands (one to a few hundred),
Manuscript received May 10, 2021; revised June 23, 2021 and July 21, different spatial resolutions (few cm/pixel to Km/pixel), and
2021; accepted July 25, 2021. Date of publication September 15, 2021; date different polarizations. While this vast variation provides an
of current version January 21, 2022. This work was supported in part by
the European Research Council (ERC) under the European Union’s Horizon opportunity for detailed earth observation, it is not trivial to
2020 Research and Innovation Programme (grant agreement No. [ERC-2016- use the same set of methods for images from different sensors.
StG-714087], Project acronym: So2Sat), in part by the Helmholtz Association Due to this reason, most existing CD methods assume that the
through the Framework of Helmholtz Artificial Intelligence (AI)—Local Unit
“Munich Unit at Aeronautics, Space and Transport (MASTr)” under Grant prechange and postchange images are acquired using the same
ZT-I-PF-5-01, in part by the Helmholtz Excellent Professorship “Data Science sensor. The temporal frequency at which the same sensor can
in Earth Observation—Big Data Fusion for Urban Research” under Grant image the same place depends on the revisit period of the
W2-W3-100, and in part by the German Federal Ministry of Education
and Research (BMBF) in the framework of the international future AI satellite on which the sensor is mounted. However, the better
Laboratory “AI4EO—Artificial Intelligence for Earth Observation: Reasoning, the spatial resolution, the more close the satellite is to the
Uncertainties, Ethics and Beyond” under Grant 01DD20001. (Corresponding earth, and the more time it takes to revisit the same place. This
author: Xiao Xiang Zhu.)
Sudipan Saha and Patrick Ebel are with the Department of Aerospace is a hindrance in the use of same-sensor CD in time-bound
and Geodesy, Data Science in Earth Observation, Technical University applications, e.g., fast response for disaster management and
of Munich, 85521 Ottobrunn, Germany (e-mail: [Link]@[Link]; precision agriculture. Using different sensors may allow us
[Link]@[Link]).
Xiao Xiang Zhu is with the Remote Sensing Technology Institute, German to obtain temporal sequences with better temporal frequency
Aerospace Center (DLR), 82234 Weßling, Germany, and also with the without sacrificing spatial resolution. However, it is not trivial
Department of Aerospace and Geodesy, Data Science in Earth Observa- to process multisensor bitemporal images as they are affected
tion, Technical University of Munich, 85521 Ottobrunn, Germany (e-mail:
[Link]@[Link]). by the spectral characteristics of the sensors. Moreover, dif-
Digital Object Identifier 10.1109/TGRS.2021.3109957 ferent sensors capture a different type of information, making
This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
4405710 IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING, VOL. 60, 2022

their comparison often challenging [12]. The difficulty of


this problem is further accentuated by the fact that we are
interested to detect change without using any labeled training
data or any abundant pool of unlabeled data.
The emergence of deep learning has seen many such
problems solved that were thought to be very challenging
in the past [13], [14]. Self-supervised learning has shown
remarkable success recently, even when only few images are
available [15]. Intrigued by this, in this article, we explore
the challenging problem of CD between optical and SAR
images, the disparity between which is evident in Fig. 1. Fig. 1. Visual contrast for Las Vegas between (a) optical image (prechange)
We exploit recent developments in the self-supervised learning and (b) SAR image (postchange). Optical and SAR images emphasize differ-
and deep clustering to propose a method for challenging ent properties of the target area, thus performing CD on them is challenging.
SAR-optical CD where one of the bitemporal images is
acquired by an optical sensor, while the other is acquired by an 3) We experimentally show the efficacy of the proposed
SAR sensor. method on four different bitemporal multisensor scenes.
The proposed method requires only the bitemporal target The rest of this article is organized as follows. Related
scene (where change is to be detected), no training label, and works are briefly discussed in Section II. Section III outlines
no additional unlabeled data. The target bitemporal scene is the proposed method. Datasets and experimental results are
typically large, few hundred pixels by few hundred pixels. detailed in Section IV. Finally, we conclude this article in
Smaller bitemporal patches (e.g., 64 × 64) are extracted Section V.
from it to train a two-branch network, similar to the Siamese
II. R ELATED W ORK
network [16]. Each branch of the network has a projection
module and a predictor. Projection modules learn features In this section, we briefly discuss existing works on unsu-
unique to optical and SAR data without sharing weights, pervised CD (with a focus on the multisensor CD) and self-
while predictors share the weight. The output of the predictors supervised learning.
is used to estimate deep clustering loss for both images
A. Change Detection
separately. Moreover, considering that the prior probability
of changed pixels is much less than the unchanged ones, Prior to the emergence of deep learning, most unsupervised
a temporal consistency loss is proposed, which ensures that CD methods used the concept of pixelwise image differencing,
pixels in the same location at two different times tend to i.e., change vector analysis (CVA) [17]. A number of super-
get the same label. To ensure that this does not lead the pixels and spatial neighborhood-based variants of CVA have
network to learn a trivial solution, a contrastive loss is used. been proposed, e.g., parcel change vector analysis (PCVA) [18]
By the combination of these losses, the proposed method and robust change vector analysis (RCVA) [19]. Most deep
learns useful semantic features from the multisensor (SAR- learning-based unsupervised CD methods use transfer learn-
optical) bitemporal target scene, and after training, the network ing. Reference [1] proposed deep change vector analysis
predictions can be compared for CD. (DCVA), a CD framework that combines ideas from CVA
The contributions of this article are given as follows. with feature extraction based on pretrained neural networks.
In nutshell, a deep model that has been trained for some
1) We propose a self-supervised learning method for CD other task is reused to obtain pixelwise bitemporal deep
in a bitemporal scene where one image is captured by features from the target scene. Bitemporal deep features are
the optical sensor and the other by the SAR sensor. The then compared to obtain deep change hypervectors for each
proposed method, only exploiting the available target pixel in the scene, which is analyzed based on magnitude (2
unlabeled scene, effectively absorbs several concepts norm) to identify the changed pixels. While [20] shows that
from the recent self-supervised learning literature, e.g., sensor-specific pretrained network is more suitable for transfer
deep clustering, augmented view, Siamese network, and learning, [5] advocates models trained on ImageNet [21] for
contrastive learning. By effectively exploiting these con- transfer learning in CD. There is another class of unsupervised
cepts and modifying them appropriately for the target CD methods that preclassifies some pixels with high confi-
multisensor bitemporal data, the proposed method is able dence as changed/unchanged using some traditional approach
to train a network that is further used for bitemporal and further uses those confident samples for training a CD
comparison and CD. model [22].
2) We show the versatility of self-supervised learning on It is not trivial to process multisensor bitemporal images
spatiotemporal satellite data that are very different from as they are affected by differences in spatial resolution and
typical computer vision images. Even though some form differences in the spectral characteristics of the sensors.
of aerial images (e.g., drone images) is often studied in Due to this, there are very few works that can work in
computer vision, we stress that our satellite data (both the setting where prechange and postchange images have
optical and SAR) are significantly different from the different spatial resolution [23], [24] or bands with differ-
typical aerial images. ent spectral characteristics [25]. Moreover, those works deal
SAHA et al.: SELF-SUPERVISED MULTISENSOR CD 4405710

with only minor variations in spatial or spectral character- III. P ROPOSED M ETHOD
istics. Saha et al. [23] proposed a cycle-consistent genera- Let X 1 and Z 2 be two images of size R × C taken over
tive adversarial network-based method to learn transcoding the same geographical region at times t1 and t2 , respectively.
between multisensor multitemporal domain. However, their Without loss of generality, we assume that the prechange
work assumes that a large (unlabeled) area corresponding to image X 1 is acquired by an optical sensor (RGB), and the
both sensors is available as training data. Liu et al. [26] used a postchange image Z 2 is acquired by the SAR sensor. Since
symmetric convolutional coupling network (SCCN), and [27] SAR image is grayscale, the same channel is replicated thrice
used denoising autoencoder (DAE) for CD in multisensor to make it three-channel like the optical input. We aim to
images. Though those works considered optical-SAR images, detect changes from the images X 1 and Z 2 in an unsuper-
they applied their methods to scenes with limited spatial vised manner, i.e., without using any training labels and any
complexity. While our work is strongly motivated by the additional unlabeled data pool. Our goal is to divide the set
existing works on multisensor CD [23], [24], it takes them of all pixels  into two subsets c and ωnc corresponding
a step further by considering the challenging scenario of to changed and unchanged pixels, respectively. Like most
optical-SAR CD in complex urban scenes and, furthermore, existing unsupervised CD methods [1], we assume that the
by integrating recent developments in self-supervised learning. prior probability of occurrence of change is less compared to
no change [38].
B. Self-Supervised Learning We can extract a set of bitemporal patches of size R  × C 
(R  < R and C  < C) from the images X 1 and Z 2 .
Considering the difficulty of collecting labeled data and the In practice, one training iteration involves only a batch of
abundance of unlabeled data, machine learning researchers B patches from X 1 , denoted as X = {x 11 , . . . , x 1B }, and
have focused on developing unsupervised and self-supervised corresponding patches from Z 2 , denoted as Z = {z 21 , . . . , z 2B }.
deep learning methods in the recent past. Gidaris et al. [28] x 1b and z 2b are processed separately with deep clustering loss,
used image rotation as a pretext task to learn unsuper- as detailed in Section III-C. Furthermore, considering that x 1b
vised semantic feature. Several other pretext tasks have been and z 2b represent same location at two different times and prior
explored in the literature, e.g., relative patch prediction [29] probability of change is less, a temporal consistency loss (see
and image inpainting [30]. Deep clustering, i.e., joint learn- Section III-D) is formulated using each such pair. Furthermore,
ing of the parameters of the deep network and the cluster Z is shuffled to form negative samples Z  , and a contrastive
assignment of the resulting features, has also been shown loss is used between pairs from X and Z  , as outlined in
to be effective for unsupervised representation learning [31]. Section III-E. The proposed method is outlined in Fig. 2.
Remarkably, [15] has shown that the abovementioned unsu-
pervised methods learn useful semantic features even with a
single-image input. Contrastive methods function by bring- A. Bitemporal Patches are Multiple Views of the Same
ing the representation of different views of the same image Location
(“positive pairs”) closer while spreading representations of We recall from Section II-B that many self-supervised
different images (“negative pairs”) apart [32]–[34]. Boostrap learning approaches build upon the concept of bringing closer
your own latent [35] and its variant SiamSiam [16] eliminate the representation of the multiple views of the same image.
the requirement of negative pair by using multiple views of the Different views of the same image are generally obtained
same image. In more detail, SiamSiam [16] ingests as input by different augmentation techniques, e.g., random crops.
two randomly augmented views of an image and processes it We argue that multisensor bitemporal patches x 1b and z 2b
through a Siamese architecture. Each Siamese branch consists can be similarly thought to be multiple views of the same
of an encoder and a prediction head. The encoders share location. They represent augmentation of the same place,
weight between two views. where the augmentation transformation is naturally caused by
The proposed method is strongly inspired from the above multisensor differences and other factors, including weather
self-supervised methods. Like deep clustering [31], the pro- conditions. Considering that the prior probability of change
posed method uses the concept of simultaneous representation is less [38], most of the time, such a pair of patches x 1b and
learning and cluster/label assignment. The bitemporal images z 2b represent the same information but from the eyes of two
can be considered to be views of the same scene, such as different viewers (sensors).
SiamSiam [16]. Like the contrastive methods, the proposed
method uses the idea of bringing closer the representation of
B. Siamese Representation
positive pairs and spreading apart the negative pairs. Like [15],
the proposed method works on a single scene (a pair of images Since bitemporal patches can be seen as multiple views of
capturing the same location at two different times). the same location, we argue that semantic information can
Multitemporal satellite image processing researchers have be captured from them by using a Siamese-like architecture.
also proposed self-supervised representation learning methods, Similar to [16], both branches of the two-branch network
e.g., deep clustering for multitemporal segmentation [36] have projection modules fopt and f sar for the optical and
and learning by rearranging randomly shuffled time-series SAR branch, respectively. In addition, both branches have
images [37]. The proposed method is related to them, using prediction modules h opt and h sar for the optical and SAR
the concept of deep clustering as in [36]. branches, respectively. However, unlike [16], the projection
4405710 IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING, VOL. 60, 2022

Fig. 2. Proposed unsupervised multisensor (optical-SAR) CD framework. The left-hand side denotes the self-supervised training process, while the right-hand
side shows the CD process using already trained model.

Fig. 3. Network simplified architecture with L 1 = 4 and L 2 = 1. Optical and SAR inputs are processed separately and subsequently fed to a common
prediction layer.

modules fopt and f sar do not share weight. This is because the resulting features [31]. Deep clustering helps the network
SAR and optical images are significantly different processed to learn discriminative features that can identify different
by two different projection modules using different sets of classes/clusters in the images. Considering the processing of
weights. However, the prediction modules h opt and h sar share the two images as an independent process, deep clustering can
weights and, henceforth, simply denoted as h. be performed for each of them. The output obtained by the
The projection and the prediction networks consist of L 1 network for a paired input patches x 1b and z 2b is
and L 2 (generally L 2 = 1) convolutional layers, respectively,   
where L = L 1 + L 2 . The two projections compute a projected y1b = h fopt x 1b (1)
  b 
representation from the optical and SAR images and project y2 = h fsar z 2 .
b
(2)
them to a common domain. In the ideal scenario, where the
projectors have perfectly learned to project optical and SAR y1b has same spatial dimension R  × C  as x 1b and has kernel
images into a common domain and the bitemporal images do number (or, feature dimension) K . The deep clustering process
not show any change, the output generated for an input pair is is performed over the pixels, i.e., each pixel is assigned to a
expected to be identical. However, practically even in absence cluster. Without loss of generality, we, henceforth, explain the
b
of any change, there are differences caused by multisensor deep clustering process in reference to a generic pixel y1,n
b b
acquisition and other factors that are not trivial for projection from y1 . The dimension of y1,n is K that can be converted to
b
modules to mitigate. 1-D label c1,n by argmax classification. This is achieved by
b
All but the last convolution layers are followed by the ReLU selecting the kernel/feature in y1,n that has maximum value.
activation function. They are further followed by the batch
b b
If the kth feature of y1,n is represented by y1,n (k), then label
b
normalization layer. We do not use any pooling layer; hence, c1,n is obtained as follows:
the size of the input is preserved in the output. While filters b
c1,n = arg max y1,n
b
(k). (3)
of spatial size 3 × 3 are used for all convolution layers for k∈K
projection, the prediction module uses 1 × 1 filter. The kernel
The rationale behind finding the highest activation of an
number of the final layer is K and can be thought of as K
input pixel is that the pixels that obtain the highest activation
different clusters/classes. Each pixel can be assigned to one of
in the same feature are likely to have similar semantics, thus
these K clusters (as detailed in Section III-C). The network
belonging to the same group. While there are several possible
architecture is shown in Fig. 3.
ways to define the pseudolabel, our approach more closely
follows the ones based on argmax classification of the final
C. Deep Clustering layer [39], [40]. Once the pixels are assigned to the K clusters,
The deep clustering process involves the joint learning of the parameters of the deep network can be updated by using a
b b
parameters of the deep network and the cluster assignment of loss between the feature y1,n and the cluster c1,n . We use
SAHA et al.: SELF-SUPERVISED MULTISENSOR CD 4405710

cross-entropy loss as Algorithm 1 Self-Supervised Training for Multisensor CD


 b 
b1,n = crossentropy y1,n , c1,n
b
. (4) 1: Initialize W1 , . . . , W L
2: for i ← 1 to I do
In practice, the loss term L1 is computed by taking mean 3: Sample B patches from X 1 , denoted as X =
of b1,n over all pixels in x 1b and all patches in the batch (b = {x 11 , . . . , x 1B }
1, . . . , B). L1 is used to adjust the weights of h and f opt . 4: Obtain corresponding B patches from Z 2 , denoted as
Similarly, L2 is computed from z 2b (b = 1, . . . , B) and used Z = {z 21 , . . . , z 2B }
to modulate the weights of h and f sar . 5: Obtain Z  as random shuffling of Z
While deep clustering helps to learn representation for each 6: for j ← 1 to J do
sensor separately, they do not ensure that the independently 7: for b ∈ B do
learned features are aligned with each other. 8: y1b = h( fopt (x 1b ))
9: y2b = h( fsar (z 2b ))
 
D. Temporal Consistency 10: y2b = h( f sar (z 2b ))
Recalling from Section III-B, multisensor bitemporal 11: end for
patches x 1b and z 2b are multiple views of the same location 12: Calculate deep clustering losses L1 , L2
in the absence of any change. In other words, in coregistered 13: Calculate temporal consistency loss L1,2
bitemporal images, pixels in the same spatial location gener- 14: Calculate contrastive loss L1,2
ally tend to belong to the same object as changes have a low 15: if i ≤ I1 then
prior probability than the unchanged class. Thus, the features 16: Use loss (L1 + L2 )/2 to modulate W1 , . . . , W L
computed for the bitemporal paired patches x 1b and z 2b should 17: else
b
be similar in most cases. For each input pixel x 1,n b
and z 2,n , 18: For each 3 consecutive iterations j , use L1 , L1,2 ,
we compute absolute error (AE) loss as and L1,2 , respectively, to modulate W1 , . . . , W L
 b  end if
b12,n =  y1,n b  19:
− y2,n . (5)
1 20: end for
A loss term L1,2 is computed by taking the mean of b12,n 21: end for
over all considered pixels for all patches in the batch. The
proposed temporal consistency only ensures that the pixels at 
b b
the same location, however, at two different times, tend to have it approaches 0, i.e., y1,n and y2,n become too similar. This
the same label. This may lead to a degenerate solution where is achieved by computing the loss term L1,2 as mean of

all pixels simply have the same prediction for both times. exponentials of b12,n over all considered pixels for all patches
Moreover, some bitemporal pairs x 1b and z 2b may be indeed in the batch.
changed and, however, penalized for producing dissimilar
output in this step. F. Overall Loss and Network Refinement
The initialization process [41] is used to initialize all the
E. Contrastive Learning trainable weights of the network W1 , . . . , W L , corresponding
While Section III-D encourages the features computed for to L layers. For updating of weights, we exploit stochastic
paired patch x 1b and z 2b to be similar, in this section, we encour- gradient descent (SGD) mechanism with momentum [42].
age the network to produce a dissimilar feature for different The training process is executed in two different steps of I1
inputs by employing concepts inspired by contrastive learning. and I2 epochs (summing to I). For each batch of data, J
While we do not have negative samples under the unsupervised iterations are performed. For the first I1 epochs, only the sum
setting in which our work is based on, we simply shuffle the of deep clustering losses (L1 + L2 ) is used to modulate the
batch of patches Z to Z  . Recall that X and Z have location- network weights. For subsequent I2 epochs, in one training
wise paired patches. This implies that X and Z  have unpaired iteration, L1 is used as loss function; in the following iteration,
patches. Thus, there should be more dissimilar in comparison L1,2 is used; and in the following iteration, L1,2 is used.
to the paired patches in Section III-D. We encourage features The combination of three loss functions yields a balanced

computed for x 1b and z 2b to be dissimilar. This is achieved training process taking into account coherent cluster formation,
b
by computing (negative) AE loss for each input pixel x 1,n temporal feature consistency, and feature dissimilarity for
b
and z 2,n unpaired patches. Alternatively, sum of L1 , L1,2 , and L1,2 can
 b  also be used as aggregated loss function. The self-supervised
b12,n = − y1,n b 

− y2,n 1
. (6) mechanism for network training is shown in Algorithm 1.
 
b12,n has negative value. Ideally, b12,n should be encouraged to
be more and more negative. However, in practice, we note that G. Change Detection
simply shuffling Z to Z  does not always ensure that X and Z  Once the network is trained, it can be used to detect change
have semantically different patches. Even after shuffling, they between X 1 and Z 2 . Since the network is fully convolutional,
may have the semantically paired patches, however penalized it enables us to obtain a pixelwise feature vector of dimension
in this step for producing similar features. Thus, to control K from X 1 and Z 2 . Similar to [1], the pixelwise change

its impact, we penalize the network with b12,n only when information is captured by taking the magnitude (2 norm)
4405710 IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING, VOL. 60, 2022

Fig. 5. CD results for the Chongqing. CD maps: (a) reference, (b) proposed,
(c) FCC between reference and proposed (the correctly detected region are in
black, false alarms are in green, and missed alarms are in pink), (d) PCVA,
(e) DCVA, and (f) SCCN.

Fig. 4. CD results for Las Vegas. CD maps: (a) reference, (b) proposed,
(c) false color composite (FCC) between reference and proposed (the correctly
detected region are in black, false alarms are in green, and missed alarms are
in pink), (d) CVA, (e) RCVA, (f) PCVA, (g) DCVA, (h) encoder–decoder, and
(i) SCCN.

of difference of the feature vectors computed from prechange


and postchange pixels. Changed pixels (c ) generate a higher
difference magnitude in comparison to the unchanged ones
ωnc , and they can be distinguished by using any suitable
threshold determination scheme [43]. Fig. 6. CD results for the Abu Dhabi. CD maps: (a) reference, (b) proposed,
(c) FCC between reference and proposed (the correctly detected region are in
black, false alarms are in green, and missed alarms are in pink), (d) PCVA,
IV. E XPERIMENTAL VALIDATION (e) DCVA, and (f) SCCN.

A. Datasets
We use four paired optical (prechange)–SAR (postchange)
images to validate the proposed method. Optical images are
acquired by the Sentinel-2 sensor and are taken from the
Onera Satellite Change Detection (OSCD) dataset [44]. They
show 10-m/pixel spatial resolution. The OSCD dataset is
originally a single-sensor dataset consisting of only Sentinel-
2 images. Recalling the importance of multisensor CD (see
Section I), we extend this dataset by collecting the postchange
SAR Sentinel-1 images for the nearest available date as the
postchange image in the original OSCD dataset. Both Sentinel-
2 and Sentinel-1 sensors are part of the European Space
Agency’s Copernicus program.
Fig. 7. Qualitative CD results for the Montpellier. CD maps: (a) reference,
The four scenes are collected over Las Vegas in United (b) proposed, (c) FCC between reference and proposed (the correctly detected
States (824 × 716 pixels) (see Fig. 4), Chongqing in China region are in black, false alarms are in green, and missed alarms are in pink),
(730 × 544 pixels) (see Fig. 5), Abu Dhabi (799 × 785 pixels) (d) PCVA, (e) DCVA, and (f) SCCN.
(see Fig. 6), and Montpellier in France (426 × 451 pixels) (see
Fig. 7). Thus, this provides us an opportunity to validate the B. Compared Methods
proposed method on geographically distributed complex urban To verify the effectiveness of the proposed method, we com-
scenes with large variation. pare it to related unsupervised CD methods.
SAHA et al.: SELF-SUPERVISED MULTISENSOR CD 4405710

TABLE I TABLE II
S TRUCTURE OF THE N ETWORK FOR P ROCESSING C OMPARISON OF D IFFERENT M ETHODS ON L AS V EGAS
O NE OF THE T WO I NPUTS

1) CVA [17], [45], a classical difference-based unsuper-


vised model for CD. TABLE III
2) RCVA [19] that modifies CVA by taking into account VARIATION OF R ESULT FOR L AS V EGAS AS I IS VARIED
pixel neighborhood effects.
3) PCVA [18] that incoporates notion of the object (super-
pixels) in CVA.
4) DCVA [1] that detects change by comparing bitempo-
ral deep features extracted using a pretrained network.
We used the second convolution layer of pretrained
VGGNet [46] for feature extraction.
5) Image-to-image transfer model based on an TABLE IV
encoder–decoder network architecture that projects VARIATION OF R ESULT FOR L AS V EGAS AS K IS VARIED
prechange optical images into postchange SAR
image [47]. The CD map can be obtained by the
difference between the simulated prechange SAR image
(obtained as the projection of prechange optical image)
and the original postchange SAR image.
6) DAE-based joint feature extraction [27].
7) SCCN [26] that first identifies some unchanged pixels
and uses them to learn a coupled network. detail, given true positive (TP), true negative (TN), false
While methods 1–3 are not deep learning-based, the fol- positive (FP), and false negative (FN), sensitivity is TP/(TP +
lowing ones are deep learning-based. Methods 1–4 do not FN), and specificity is TN/(TN + FP).
have any explicit adaptation for multisensor input, while D. Results
methods 5–7 have.
1) Las Vegas: The reference CD map (ground truth) for Las
Vegas is shown in Fig. 4(a). Fig. 4(b) shows the result obtained
C. Experimental Settings by the proposed method. For better visualization, a false color
The proposed method and compared methods are fed with composition between the reference map and the obtained result
preprocessed images and postprocessed similarly. For the is shown in Fig. 4(c). The proposed method can detect most of
proposed method, we use I = 5 (I1 = 1, I2 = 4), the changed objects with fewer false alarms in comparison to
J = 50 K = 4, L 1 = 4, and L 2 = 1. We show the architecture the compared methods. In many cases, the proposed method
of the network in Table I. A relatively simple architecture is partly detects the changed object, thus missing some objects
used considering that the number of patches available to us is only partially [shown in pink in Fig. 4(c)]. CVA [see Fig. 4(d)]
very few compared to the images in typical computer vision performs poorly and incorrectly detects most urban areas as
datasets. Moreover, our target image has a coarse resolution changed. The result obtained by RCVA [see Fig. 4(e)] is
(10 m/pixel) compared to natural images in computer vision. similar to CVA. While PCVA [see Fig. 4(f)], DCVA [see
Spatial complexity in such coarse images can be handled by Fig. 4(g)], encoder–decoder [see Fig. 4(h)], DAE, and SCCN
simpler architecture compared to those in computer vision. [see Fig. 4(i)] improve the result over CVA, the proposed
64 × 64 patches are used to train the model, and patches are method still outperforms them by large margin. Quantitative
extracted from the bitemporal scene with a stride of 32. The evaluation (see Table II) clearly shows the superiority of the
actual number of training patches for a scene depends on the proposed method over state-of-the-art unsupervised methods.
size of the particular scene. For example, for the Las Vegas This can be attributed to the superior capability of the proposed
scene (824 × 716 pixels), the number of patches extracted is method to ingest multisensor multitemporal images.
504. For optimization, the SGD method is used with a learning Further studies are conducted by varying different parame-
rate set to 0.001. ters on the Las Vegas image pair.
We show the result in terms of sensitivity (accuracy in per- Training epochs I are varied with different values, as tab-
centage computed over reference changed pixels) and speci- ulated in Table III, while setting K = 4. We observe clear
ficity (computed over reference unchanged pixels). In more improvement in performance from I = 1 to 2. Recalling
4405710 IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING, VOL. 60, 2022

TABLE V
VARIATION OF R ESULT FOR L AS V EGAS AS T HRESHOLD
D ETERMINATION S CHEME IS VARIED

Fig. 8. Evolution of the loss over training iterations for Las Vegas: (a) deep
TABLE VI clustering loss L1 and (b) temporal consistency loss L1,2 and contrastive
C OMPARISON OF D IFFERENT M ETHODS ON C HONGQING loss L1,2 .

TABLE VIII
C OMPARISON OF D IFFERENT M ETHODS ON M ONTPELLIER

TABLE VII
C OMPARISON OF D IFFERENT M ETHODS ON A BU D HABI

L1,2 and L1,2 are introduced to the training process. L1,2 and
L1,2 balance each other, as shown in Fig. 8(b).
Projection layers f opt and f sar need to be modeled inde-
pendently by not sharing weights between them to capture
the different semantic properties of optical and SAR patches,
as hypothesized in Section III-B. Here, we test this hypothesis
by instead sharing the weights between f opt and f sar . For I = 5
and K = 4, the proposed method fails to detect most of the
changes. This shows that it is crucial to model the optical and
from Section III-F that, for first I1 = 1 iterations, only SAR patches differently.
deep clustering loss is used, this shows that bitemporal deep The computation time requirement is not high. We tested
clustering itself is not sufficient to learn the correspondence our code on a machine equipped with a Quadro T2000 GPU,
between two images, and the other losses (L1,2 and L1,2 ) are which is a low-end GPU. For processing the Las Vegas dataset
required. From I = 2 onward, we observe an increment in (training process over five epochs), it takes approx. 460 s.
performance initially followed by performance getting satu- The Las Vegas scene is 824 × 716 pixels with the 10-m/pixel
rated/dropping. Despite variation in performance, the proposed resolution, and thus, processing it is equivalent to processing
method outperforms all compared methods for I = 3, 5, 10. an approximate area of 8 ∗ 7 = 56 km2 in terms of geography.
The kernel number of the last layer (K ) is varied from The same sensor bitemporal input can be ingested by
2 to 16 in multiplicative steps of 2 while fixing the I = 5. the proposed method, though designed for multisensor CD.
The variation in performance is shown in Table IV. While For Las Vegas prechange optical–postchange optical input,
performance improves from K = 2 to K = 4, a gradual fall the proposed method can obtain a sensitivity of 64.74% and
in performance is observed henceforth. The increasing value specificity of 97.89%. However, we note that some character-
of K is equivalent to allowing the scene to be partitioned into istics of the proposed method (e.g., temporal consistency loss)
more classes. Since the spatial area of the scene is fixed and are designed to reduce the representation gap of multisensor
not too large (only few hundred pixels by few hundred pixels), input, which is less relevant in single-sensor input. Thus,
a large number of classes potentially leads the model to learn the proposed method may not be the most suitable choice
irrelevant classes, impacting CD performance. for single-sensor scenarios as there are numerous existing
Thresholding is done using Otsu’s method [43], as it is CD techniques particularly designed for the same-sensor sce-
popular in unsupervised CD methods [19], [48]. However, nario [1].
any other suitable method can be used, as shown in Table V. 2) Chongqing and Abu Dhabi: Reference CD map (ground
Results obtained by the ISODATA method [49], [50] and the truth) for Chongqing is shown in Fig. 5(a). Fig. 5(b) and
adaptive method [1] are similar to Otsu’s method [43]. (c) shows the result obtained by the proposed method and
Loss plot visualization in Fig. 8 shows the interplay between false color composition between the reference map and
different components of loss. L1 consistently decreases [see the obtained result, respectively. The proposed method out-
Fig. 8(a)] except that it rises for a while after epoch 1 when performs all compared methods, as can be observed in
SAHA et al.: SELF-SUPERVISED MULTISENSOR CD 4405710

quantitative results in Table VI). Similar result is obtained for [8] J. E. Ball, D. T. Anderson, and C. S. Chan, “Comprehensive survey
Abu Dhabi (see Fig. 6 and Table VII). of deep learning in remote sensing: Theories, tools, and challenges
for the community,” J. Appl. Remote Sens., vol. 11, no. 4, 2017,
3) Montpellier: Reference CD map (ground truth) for Mont- Art. no. 042609.
pellier is shown in Fig. 7(a). The proposed method [see [9] M. Hirschmugl, J. Deutscher, C. Sobe, A. Bouvet, S. Mermoz,
Fig. 7(b)] outperforms most of the state-of-the-art methods, and M. Schardt, “Use of SAR and optical time series for tropical
forest disturbance mapping,” Remote Sens., vol. 12, no. 4, p. 727,
including PCVA [see Fig. 7(d)] and DCVA [see Fig. 7(e)], Feb. 2020.
as shown in Table VIII. However, SCCN [see Fig. 7(f)] [10] N. Zhou, X. Li, Z. Shen, T. Wu, and J. Luo, “Geo-parcel-based change
outperforms the proposed method. The performance of the detection using optical and SAR images in cloudy and rainy areas,” IEEE
proposed method is relatively poor for Montpellier, which can J. Sel. Topics Appl. Earth Observ. Remote Sens., vol. 14, pp. 1326–1332,
2021.
be possibly explained by: 1) smaller size of Montpellier scene, [11] U. I. Ahmed, B. Rabus, and M. F. Beg, “SAR and optical image
which implies fewer data to learn proposed self-supervised fusion for urban infrastructure detection and monitoring,” Proc. SPIE,
network and 2) uniform (showing mostly urban areas) geospa- vol. 11535, Sep. 2020, Art. no. 115350M.
tial characteristics of Montpellier scene in comparison to [12] M. Schmitt and X. X. Zhu, “Data fusion and remote sensing: An ever-
growing relationship,” IEEE Geosci. Remote Sens. Mag., vol. 4, no. 4,
Las Vegas and Chongqing that show complex distribution pp. 6–23, Dec. 2016.
consisting of both urban and nonurban areas. [13] X. X. Zhu et al., “Deep learning in remote sensing: A comprehensive
review and list of resources,” IEEE Geosci. Remote Sens. Mag., vol. 5,
no. 4, pp. 8–36, Dec. 2017.
V. C ONCLUSION [14] G. Camps-Valls, D. Tuia, X. X. Zhu, and M. Reichstein, Deep Learning
This article proposed a self-supervised learning-based for the Earth Sciences: A Comprehensive Approach to Remote Sensing,
Climate Science and Geosciences. Hoboken, NJ, USA: Wiley, 2021.
method for CD in multisensor bitemporal images where one of [15] Y. M. Asano, C. Rupprecht, and A. Vedaldi, “A critical analysis of
the images is acquired by an optical sensor and the other one self-supervision, or what we can learn from a single image,” 2019,
is captured by an SAR sensor. The proposed method effec- arXiv:1904.13132. [Online]. Available: [Link]
tively utilizes several concepts from self-supervised learning, [16] X. Chen and K. He, “Exploring simple Siamese representation
learning,” 2020, arXiv:2011.10566. [Online]. Available: [Link]
e.g., deep clustering, Siamese network, multiple views, and org/abs/2011.10566
contrastive learning, and operates under severe constraints, [17] W. A. Malila, “Change vector analysis: An approach for detecting forest
i.e., nothing except that the target scene is used, and no labeled changes with Landsat,” in Proc. LARS Symp., 1980, p. 385.
data or additional unlabeled image is used. Despite the strong [18] F. Bovolo, “A multilevel parcel-based approach to change detection in
very high resolution multitemporal images,” IEEE Geosci. Remote Sens.
difference in the input modalities and operating under stringent Lett., vol. 6, no. 1, pp. 33–37, Jan. 2009.
constraints, it can identify a large fraction of the changed [19] F. Thonfeld, H. Feilhauer, M. Braun, and G. Menz, “Robust change
pixels. Comparisons with the existing methods working under vector analysis (RCVA) for multi-sensor very high resolution optical
satellite data,” Int. J. Appl. Earth Observ. Geoinf., vol. 50, pp. 131–140,
unsupervised scenarios show that the proposed method brings Aug. 2016.
significant improvement, especially when the target scene is [20] S. Saha, F. Bovolo, and L. Bruzzone, “Building change detection in VHR
large. Potential improvement of the proposed method may SAR images via unsupervised deep transcoding,” IEEE Trans. Geosci.
Remote Sens., vol. 59, no. 3, pp. 1917–1929, Mar. 2021.
be achieved by prior learning of clusters on the unrelated
[21] J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet:
domains/sensors and transferring them to target sensors on the A large-scale hierarchical image database,” in Proc. IEEE Conf. Comput.
fly [51]. In addition, our future work will focus on extending Vis. Pattern Recognit., Jun. 2009, pp. 248–255.
the method to other application domains, e.g., the comparison [22] H. M. Keshk and X.-C. Yin, “Change detection in SAR images based on
deep learning,” Int. J. Aeronaut. Space Sci., vol. 21, no. 2, pp. 549–559,
of biomedical images. 2020.
[23] S. Saha, F. Bovolo, and L. Bruzzone, “Unsupervised multiple-change
R EFERENCES detection in VHR multisensor images via deep-learning based adap-
tation,” in Proc. IEEE Int. Geosci. Remote Sens. Symp. (IGARSS),
[1] S. Saha, F. Bovolo, and L. Bruzzone, “Unsupervised deep change vector Jul. 2019, pp. 5033–5036.
analysis for multiple-change detection in VHR images,” IEEE Trans. [24] P. Zhang, M. Gong, L. Su, J. Liu, and Z. Li, “Change detection
Geosci. Remote Sens., vol. 57, no. 6, pp. 3677–3693, Jun. 2019. based on deep feature representation and mapping transformation for
[2] A. Appice, N. Di Mauro, F. Lomuscio, and D. Malerba, “Empowering multi-spatial-resolution remote sensing images,” ISPRS J. Photogramm.
change vector analysis with autoencoding in bi-temporal hyperspectral Remote Sens., vol. 116, pp. 24–41, Jun. 2016.
images,” in Proc. CEUR Workshop, vol. 2466, 2019, pp. 1–10. [25] M. Volpi, G. Camps-Valls, and D. Tuia, “Spectral alignment of multi-
[3] H. Chen, C. Wu, B. Du, L. Zhang, and L. Wang, “Change detection temporal cross-sensor images with automated kernel canonical correla-
in multisource VHR images via deep Siamese convolutional multiple- tion analysis,” ISPRS J. Photogramm. Remote Sens., vol. 107, pp. 50–63,
layers recurrent neural network,” IEEE Trans. Geosci. Remote Sens., Sep. 2015.
vol. 58, no. 4, pp. 2848–2864, Apr. 2020.
[4] F. Rahman, B. Vasu, J. V. Cor, J. Kerekes, and A. Savakis, “Siamese [26] J. Liu, M. Gong, K. Qin, and P. Zhang, “A deep convolutional coupling
network with multi-level features for patch-based change detection in network for change detection based on heterogeneous optical and
satellite imagery,” in Proc. IEEE Global Conf. Signal Inf. Process. radar images,” IEEE Trans. Neural Netw. Learn. Syst., vol. 29, no. 3,
(GlobalSIP), Nov. 2018, pp. 958–962. pp. 545–559, Mar. 2018.
[5] A. Pomente, M. Picchiani, and F. Del Frate, “Sentinel-2 change detection [27] T. Zhan, M. Gong, X. Jiang, and S. Li, “Log-based transforma-
based on deep features,” in Proc. IEEE Int. Geosci. Remote Sens. Symp. tion feature learning for change detection in heterogeneous images,”
(IGARSS), Jul. 2018, pp. 6859–6862. IEEE Geosci. Remote Sens. Lett., vol. 15, no. 9, pp. 1352–1356,
[6] S. T. Seydi and M. Hasanlou, “A new land-cover match-based change Sep. 2018.
detection for hyperspectral imagery,” Eur. J. Remote Sens., vol. 50, no. 1, [28] S. Gidaris, P. Singh, and N. Komodakis, “Unsupervised representa-
pp. 517–533, Jan. 2017. tion learning by predicting image rotations,” 2018, arXiv:1803.07728.
[7] M. Puhm, J. Deutscher, M. Hirschmugl, A. Wimmer, U. Schmitt, and [Online]. Available: [Link]
M. Schardt, “A near real-time method for forest change detection based [29] C. Doersch, A. Gupta, and A. A. Efros, “Unsupervised visual represen-
on a structural time series model and the Kalman filter,” Remote Sens., tation learning by context prediction,” in Proc. IEEE Int. Conf. Comput.
vol. 12, no. 19, p. 3135, Sep. 2020. Vis., Dec. 2015, pp. 1422–1430.
4405710 IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING, VOL. 60, 2022

[30] D. Pathak, P. Krahenbuhl, J. Donahue, T. Darrell, and A. A. Efros, Sudipan Saha (Member, IEEE) received the
“Context encoders: Feature learning by inpainting,” in Proc. IEEE Conf. [Link]. degree in electrical engineering from IIT
Comput. Vis. Pattern Recognit., Jun. 2016, pp. 2536–2544. Bombay, Mumbai, India, in 2014, and the Ph.D.
[31] M. Caron, P. Bojanowski, A. Joulin, and M. Douze, “Deep clustering for degree in information and communication technolo-
unsupervised learning of visual features,” in Proc. Eur. Conf. Comput. gies from the University of Trento, Trento, Italy, and
Vis. (ECCV), 2018, pp. 132–149. Fondazione Bruno Kessler, Trento, in 2020.
[32] T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple He is currently a Post-Doctoral Researcher with
framework for contrastive learning of visual representations,” 2020, the Technical University of Munich (TUM), Munich,
arXiv:2002.05709. [Online]. Available: [Link] Germany. Previously, he worked as an Engineer with
[33] X. Chen, H. Fan, R. Girshick, and K. He, “Improved baselines with TSMC Ltd., Hsinchu, Taiwan, from 2015 to 2016.
momentum contrastive learning,” 2020, arXiv:2003.04297. [Online]. In 2019, he was a Guest Researcher with TUM.
Available: [Link] His research interests include multitemporal remote sensing image analysis,
[34] Y. Tian, C. Sun, B. Poole, D. Krishnan, C. Schmid, and P. Isola, domain adaptation, time-series analysis, image segmentation, deep learning,
“What makes for good views for contrastive learning?” 2020, image processing, and pattern recognition.
arXiv:2005.10243. [Online]. Available: [Link] Dr. Saha is also a reviewer for several international journals. He was
[35] J.-B. Grill et al., “Bootstrap your own latent: A new approach to a recipient of the Fondazione Bruno Kessler Best Student Award 2020.
self-supervised learning,” 2020, arXiv:2006.07733. [Online]. Available: He has served as a Guest Editor for Remote Sensing (MDPI) Special Issue
[Link] on “Advanced Artificial Intelligence for Remote Sensing: Methodology and
[36] S. Saha, L. Mou, C. Qiu, X. X. Zhu, F. Bovolo, and L. Bruzzone, Application.”
“Unsupervised deep joint segmentation of multitemporal high-resolution
images,” IEEE Trans. Geosci. Remote Sens., vol. 58, no. 12,
pp. 8780–8792, Dec. 2020. Patrick Ebel received the [Link]. degree in cog-
[37] S. Saha, F. Bovolo, and L. Bruzzone, “Change detection in nitive science from the University of Osnabrück,
image time-series using unsupervised LSTM,” IEEE Geosci. Remote Osnabrück, Germany, in 2015, and the dual [Link].
Sens. Lett., early access, Dec. 24, 2020, doi: 10.1109/LGRS.2020. degree in cognitive neuroscience and artificial intel-
3043822. ligence from Radboud University, Nijmegen, The
[38] F. Bovolo and L. Bruzzone, “The time variable in data fusion: A change Netherlands, in 2018. He is currently pursuing the
detection perspective,” IEEE Geosci. Remote Sens. Mag., vol. 3, no. 3, Ph.D. degree with the Department of Aerospace and
pp. 8–26, Sep. 2015. Geodesy, Technical University of Munich, Munich,
[39] A. Kanezaki, “Unsupervised image segmentation by backpropagation,” Germany.
in Proc. IEEE Int. Conf. Acoust., Speech Signal Process. (ICASSP), His research interests include deep learning and
Apr. 2018, pp. 1543–1547. its applications in computer vision and to remote
[40] S. Saha, S. Sudhakaran, B. Banerjee, and S. Pendurkar, “Semantic sensing data.
guided deep unsupervised image segmentation,” in Proc. Int. Conf.
Image Anal. Process. Cham, Switzerland: Springer, 2019, pp. 499–510.
[41] K. He, X. Zhang, S. Ren, and J. Sun, “Delving deep into rectifiers: Xiao Xiang Zhu (Fellow, IEEE) received the
Surpassing human-level performance on ImageNet classification,” in [Link]., Dr. Ing., and “Habilitation” degrees in signal
Proc. IEEE Int. Conf. Comput. Vis., Dec. 2015, pp. 1026–1034. processing from the Technical University of Munich
[42] S. Ruder, “An overview of gradient descent optimization algo- (TUM), Munich, Germany, in 2008, 2011, and 2013,
rithms,” 2016, arXiv:1609.04747. [Online]. Available: [Link] respectively.
org/abs/1609.04747 She was a Guest Scientist or a Visiting Professor
[43] N. Otsu, “A threshold selection method from gray-level histograms,” with the Italian National Research Council (CNR-
IEEE Trans. Syst., Man, Cybern., vol. SMC-9, no. 1, pp. 62–66, IREA), Naples, Italy, Fudan University, Shanghai,
Jan. 1979. China, The University of Tokyo, Tokyo, Japan, and
[44] R. C. Daudt, B. L. Saux, A. Boulch, and Y. Gousseau, “Urban change the University of California at Los Angeles, Los
detection for multispectral Earth observation using convolutional neural Angeles, CA, USA, in 2009, 2014, 2015, and 2016,
networks,” in Proc. IEEE Int. Geosci. Remote Sens. Symp. (IGARSS), respectively. Since 2019, she has been a Co-Coordinator of Munich Data
Jul. 2018, pp. 2115–2118. Science Research School ([Link]). Since 2019, she has also been the
[45] L. Bruzzone and D. F. Prieto, “Automatic analysis of the difference Head of the Helmholtz Artificial Intelligence–Research Field “Aeronautics,
image for unsupervised change detection,” IEEE Trans. Geosci. Remote Space and Transport.” Since May 2020, she has also been the Director of the
Sens., vol. 38, no. 3, pp. 1171–1182, May 2000. International Future AI Laboratory “AI4EO–Artificial Intelligence for Earth
[46] K. Simonyan and A. Zisserman, “Very deep convolutional networks Observation: Reasoning, Uncertainties, Ethics and Beyond,” Munich. Since
for large-scale image recognition,” 2014, arXiv:1409.1556. [Online]. October 2020, she has been serving as the Co-Director for Munich Data
Available: [Link] Science Institute (MDSI), TUM. She is currently a Professor of data science
[47] Y. Xu, S. Xiang, C. Huo, and C. Pan, “Change detection based on auto- in earth observation (former: signal processing in earth observation) with
encoder model for VHR images,” Proc. SPIE, vol. 8919, Oct. 2013, TUM and the Head of the Department “EO Data Science,” Remote Sensing
Art. no. 891902. Technology Institute, German Aerospace Center (DLR), Weßling, Germany.
[48] S. Saha, Y. T. Solano-Correa, F. Bovolo, and L. Bruzzone, “Unsuper- She is also a Visiting AI Professor with ESA’s Phi-Lab. Her research interests
vised deep transfer learning-based change detection for HR multispectral include remote sensing and Earth observation, signal processing, machine
images,” IEEE Geosci. Remote Sens. Lett., vol. 18, no. 5, pp. 856–860, learning, and data science, with a special application focus on global urban
May 2021. mapping.
[49] T. W. Ridler and S. Calvard, “Picture thresholding using an iterative Dr. Zhu is also a member of Young Academy (Junge Akademie/Junges
selection method,” IEEE Trans. Syst., Man, Cybern., vol. SMC-8, no. 8, Kolleg) with the Berlin-Brandenburg Academy of Sciences and Humanities,
pp. 630–632, Aug. 1978. the German National Academy of Sciences Leopoldina, and the Bavarian
[50] M. Sezgin and B. Sankur, “Survey over image thresholding techniques Academy of Sciences and Humanities. She serves on the scientific advisory
and quantitative performance evaluation,” J. Electron. Imag., vol. 13, board in several research organizations, among others the German Research
no. 1, pp. 146–165, 2004. Center for Geosciences (GFZ) and the Potsdam Institute for Climate Impact
[51] W. Menapace, S. Lathuilière, and E. Ricci, “Learning to cluster Research (PIK). She is also an Associate Editor of IEEE T RANSACTIONS
under domain shift,” 2020, arXiv:2008.04646. [Online]. Available: ON G EOSCIENCE AND R EMOTE S ENSING . She also serves as the Area Editor
[Link] responsible for Special Issue of IEEE Signal Processing Magazine.

You might also like