0% found this document useful (0 votes)
18 views13 pages

Semd Net

The document presents SemDNet, a novel semantic-guided despeckling network designed to enhance the quality of Synthetic Aperture Radar (SAR) images affected by speckle noise. By integrating high-level semantic information through a coarse despeckling module, a semantic segmentation module, and a semantic-guided despeckling module, SemDNet effectively balances noise suppression and detail preservation. Experimental results demonstrate significant performance improvements over existing methods, with notable gains in PSNR, SSIM, and perceptual quality metrics on real SAR images.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
18 views13 pages

Semd Net

The document presents SemDNet, a novel semantic-guided despeckling network designed to enhance the quality of Synthetic Aperture Radar (SAR) images affected by speckle noise. By integrating high-level semantic information through a coarse despeckling module, a semantic segmentation module, and a semantic-guided despeckling module, SemDNet effectively balances noise suppression and detail preservation. Experimental results demonstrate significant performance improvements over existing methods, with notable gains in PSNR, SSIM, and perceptual quality metrics on real SAR images.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Expert Systems With Applications 296 (2026) 129200

Contents lists available at ScienceDirect

Expert Systems With Applications


journal homepage: [Link]/locate/eswa

SemDNet: Semantic-guided despeckling network for SAR images ⋆,⋆⋆


Fuyu Bo , Yi Jin, Xiaole Ma∗, Yigang Cen∗, Shaohai Hu, Yidong Li
State Key Laboratory of Advanced Rail Autonomous Operation, School of Computer Science and Technology, Visual Intellgence +X International Cooperation Joint
Laboratory of MOE, Beijing Jiaotong University, Beijing, 100044, China

a r t i c l e i n f o a b s t r a c t

Keywords: Synthetic Aperture Radar (SAR) is highly valued for its all-weather and all-day imaging capabilities. However,
SAR despeckling SAR images are often severely degraded by speckle noise, which poses significant challenges to accurate image
Semantic segmentation interpretation. Recently, deep learning-based methods have shown great promise in despeckling tasks. However,
Deep learning
most existing approaches process images at a global level, failing to account for the semantic differences between
Semantic guided
regions. Ideally, a despeckling algorithm should suppress speckle noise effectively in homogeneous areas while
preserving fine texture details in heterogeneous regions. To address this limitation, we propose a novel semantic-
guided despeckling network (SemDNet), which leverages high-level semantic information to reconstruct high-
quality despeckled images. SemDNet consists of three core components: a coarse despeckling module (CDM),
a semantic segmentation module (SSM), and a semantic-guided despeckling module (SGDM). The CDM first
generates a coarse despeckled image, which is then processed by the SSM to extract semantic segmentation
features. These features, together with the coarse despeckled image, are further refined by the SGDM, enabling an
effective interaction between low-level pixel information and high-level semantic features, resulting in enhanced
despeckling performance. Additionally, the proposed SGDM can function as a plug-and-play module that can be
integrated into existing despeckling networks to improve their performance. Extensive experiments demonstrate
that SemDNet achieves performance gains of up to +0.98 dB in PSNR and +2.1 % in SSIM compared to baseline
methods. In real SAR images, it also shows an improvement of 13 % in the NIQE score and a reduction of 19 %
in the BRISQUE score, confirming its superior ability to balance noise suppression with detail preservation. The
source code of our proposed method is available at [Link]

1. Introduction target detection (Pan et al., 2023), classification (Zhang et al., 2024b),
and change detection (Norouzi et al., 2024).
Synthetic aperture radar (SAR), as a coherent imaging system, has In recent decades, numerous SAR despeckling methods have been
been widely applied in various fields, including disaster monitoring, developed (Rasti et al., 2021). Traditional spatial domain filtering tech-
environmental classification, urban planning, and mapping (Bao et al., niques include local and non-local filters. Local filters reduce noise by
2023; Baraha & Sahoo, 2023; Ponmani & Saravanan, 2021). Its ability computing a weighted average of the pixel values within a local win-
to operate in all weather conditions and provide day-and-night imaging dow (Wang et al., 2023). However, they often oversmooth the image,
capabilities makes SAR an indispensable tool for remote sensing appli- resulting in the loss of details and textures. Non-local filters, on the
cations (Saleh et al., 2024). However, SAR images are inherently con- other hand, search for similar patches across a broader area or the
taminated with speckle noise due to the coherent nature of radar signal entire image, averaging these patches to reduce noise. For instance,
processing. This noise manifests as granular interference, degrading im- Xu et al. (2024), proposed a non-local iterative triangular filtering al-
age quality and hindering accurate image interpretation. Speckle noise gorithm. Non-local filtering can also be combined with methods like
poses significant challenges in extracting reliable information, as it ob- sparse representation (Zhang et al., 2021), total variation (Sun et al.,
scures fine details and distorts critical features. Consequently, effective 2020), and low-rank estimation (Bo et al., 2022, 2024). Despite their
speckle reduction is a fundamental preprocessing step for tasks such as effectiveness, these methods frequently suffer from complex parameter


This document is the results of the research project funded by the National Science Foundation.
⋆⋆
The second title footnote which is a longer text matter to fill through the whole text width and overflow into another line in the footnotes area of the first page.
∗ Corresponding authors.

E-mail addresses: 23111084@[Link] (F. Bo), yjin@[Link] (Y. Jin), maxiaole@[Link] (X. Ma), ygcen@[Link] (Y. Cen), shhu@[Link]
(S. Hu), ydli@[Link] (Y. Li).

[Link]
Received 12 April 2025; Received in revised form 7 July 2025; Accepted 27 July 2025
Available online 5 August 2025
0957-4174/© 2025 Elsevier Ltd. All rights are reserved, including those for text and data mining, AI training, and similar technologies.
F. Bo et al. Expert Systems With Applications 296 (2026) 129200

selection, high computational cost, and slow processing speed. Another noise suppression. Additionally, we incorporate semantic loss that
popular strategy is transform domain despeckling. For example, SAR- implicitly leverages semantic priors. This dual-level integration en-
BM3D (Parrilli et al., 2011) integrates non-local filtering with wavelet hances both despeckling performance and the retention of essential
shrinkage for speckle noise reduction, while Penna and Mascarenhas image details.
(2019), applied Haar wavelet domain despeckling within a non-local 3) Extensive experiments demonstrate that the proposed method out-
framework. Aranda-Bojorges et al. (2022) developed a sparse represen- performs previous state-of-the-art methods qualitatively and quanti-
tation method using discrete wavelet transforms and maximize a pos- tatively.
teriori estimation. However, these methods may introduce artifacts and
distort image structures due to the inherent limitations of the transform The remainder of this paper is organized as follows: Section 2 re-
domain. views related work. The proposed SemDNet framework is detailed in
Recently, deep learning-based methods have achieved significant ad- Section 3. Section 4 presents the experimental results and analysis. Fi-
vances in image processing (Tsokas et al., 2022; Zhang et al., 2023a,b). nally, Section 5 concludes the paper and discusses future directions.
Chierchia et al. (2017), were the first to apply convolutional neural net-
work (CNN) to SAR image despeckling. Perera et al. (2023), proposed a 2. Related works
despeckling approach based on probabilistic diffusion denoising models.
Saha et al. (2023), introduced a network utilizing parallel convolutional In recent years, deep learning-based methods for SAR image despeck-
encoders, and Liu et al. (2023), developed a multi-scale feature adap- ling have demonstrated significant potential. Cheng et al. (2023), pro-
tive enhancement network. Additionally, Xiao et al. (2024), combined posed a two-stream CNN employing a hybrid truncated loss function
Transformers with non-local means for despeckling. Ma et al. (2024), to improve despeckling performance. Deng et al. (2024), introduced the
proposed a conditional diffusion model for despeckling. To address the Sublook2Sublook method, which utilizes phase information from single-
challenge of the unavailability of noise-free SAR images, several un- look complex (SLC) images through sublook decomposition and gener-
supervised methods have also been proposed (Bo et al., 2025). Molini ates speckle image pairs for training using selected sublook SAR images.
et al. (2021), proposed a blind-spot denoising network using a self- Liu et al. (2024b), developed a method that combined CNN and Trans-
supervised Bayesian approach, while Dalsasso et al. (2022), presented former architectures to efficiently extract local and global features.
a self-supervised despeckling algorithm separating real and imaginary However, these methods fail to address a crucial aspect of SAR image
parts of single-look complex SAR images. Lin et al. (2023), developed an despeckling: the diverse nature of SAR image regions, such as buildings,
unsupervised method that synthesizes pseudo-SAR images using CNN to soil, and water. Performing despeckling on a global scale often over-
extract speckle noise. looks the distinct characteristics of these regions, leading to blurred or
However, most existing despeckling methods operate globally, treat- lost structural and edge details. Earlier studies attempted to mitigate this
ing all regions uniformly. This global strategy often leads to indiscrimi- issue by incorporating handcrafted features to differentiate between re-
nate smoothing, which can obscure important textures and distort struc- gion types and enhance despeckling results (Gragnaniello et al., 2016;
tural details, particularly in heterogeneous regions. Such limitations Wang et al., 2022). However, handcrafted features struggle to capture
compromise the interpretability of objects within the scene, necessitat- the intricate textures and complex variations in SAR images, particularly
ing methods that adaptively balance noise suppression and detail preser- in heterogeneous scenes.
vation. In this context, semantic information offers valuable high-level More recently, semantic priors have shown great promise in vari-
contextual knowledge that can help address this challenge. Specifically, ous image processing tasks (Kan et al., 2022; Liu et al., 2020, 2017).
semantic segmentation delivers explicit categorical priors through ad- Wu et al. (2023), developed a framework utilizing pre-trained segmen-
vanced feature representations, enabling the differentiation of distinct tation networks as semantic knowledge bases, extracting multi-scale se-
region types (e.g., urban areas, water bodies, vegetation). By lever- mantic features. Their method incorporates semantic guidance through
aging these priors, a semantically aware despeckling framework can color histogram loss and adversarial loss while maintaining efficiency by
adaptively tailor denoising strategies for each class: preserving intricate freezing the parameters of the semantic model and employing a modular
structures in textured regions while aggressively suppressing noise in ho- design. Wei et al. (2022), introduced a semantic-guided interactive net-
mogeneous areas. To address this issue, we propose SemDNet, a novel work that facilitates comprehensive interaction between image derain-
semantic-guided method for SAR image despeckling. The proposed ap- ing and semantic segmentation tasks. Similarly, Zhang et al. (2024a),
proach consists of three key stages. First, a coarse despeckling module introduced a generalized framework for transferring semantic priors
(CDM) generates an initial despeckled image, providing a foundational from the Segment Anything Model to lightweight image restoration
estimate for further refinement. Next, a semantic segmentation module models. By leveraging distillation techniques and parameter freezing,
(SSM), pre-trained on clean remote sensing images, extracts high-level their approach reduces computational burdens during inference. These
semantic features from the coarse despeckled image. Finally, a semantic- semantic-guided approaches can be broadly categorized into two types:
guided despeckling module (SGDM) leverages these semantic features to loss-level and feature-level methods. Loss-level methods employ seman-
refine despeckling through an effective interaction between low-level tic segmentation loss as an additional constraint, implicitly guiding the
and high-level tasks. Moreover, a semantic loss is introduced to enforce restoration process. Aakerberg et al. (2022), proposed a framework
implicit semantic constraints, enabling an effective integration of se- where the loss function of an semantic segmentation network guides
mantic knowledge. Extensive experiments show that our method con- super-resolution learning, utilizing semantic information to suppress
sistently outperforms baseline methods, with average improvements of noise and refine reconstruction. The segmentation network is decou-
0.66–0.98 dB in PSNR and 2.1 % in SSIM across different datasets. The pled during inference to minimize runtime. Zheng et al. (2022), intro-
perceptual quality metrics (NIQE and BRISQUE) also show significant duced a semantic-aware single image deraining method that freezes the
enhancements of 13–19 % on real SAR images. The main contributions parameters of the segmentation network during training to reduce com-
of this paper are as follows. putational burden. Yuan et al. (2023), proposed a self-supervised SAR
despeckling method that incorporates semantic segmentation. However,
1) We propose SemDNet, a novel SAR image despeckling framework this approach only uses segmentation loss and is limited to binary-class
that incorporates high-level semantic information to guide fine-level images, which restricts its ability to handle more complex and diverse
despeckling. SemDNet effectively addresses the limitations of tradi- scenes. In contrast, feature-level methods explicitly extract semantic
tional global approaches. features through a segmentation network and integrate these with im-
2) We propose SGDM, which explicitly integrates low-level pixel infor- age features to enhance restoration quality (Li et al., 2024b, 2022).
mation with high-level semantic features to achieve more effective Notably, semantic priors have been applied to various types of image

2
F. Bo et al. Expert Systems With Applications 296 (2026) 129200

restoration tasks. For example, Liu et al. (2024a), developed a semantic- semantic-guided fine despeckling stage. Second, this approach allows
guided despeckling method for sonar images using blind-spot networks, us to isolate and evaluate the contribution of the semantic-guided fine
combining spatial and semantic features through semantic encoders despeckling process to overall performance improvements.
while employing knowledge distillation for efficiency. Huang et al. Given a noisy SAR image, the CDM generates a coarse despeckled
(2022), proposed a low-dose CT denoising approach that first trains a result:
segmentation network to generate semantic labels, then fixes its param-
𝐼𝐶 = 𝑓𝐶𝐷𝑀 (𝐼𝑁 ) (1)
eters and feeds the labels as priors into the denoising network. Chen
et al. (2024), designed a dual-path Transformer for ultrasound image where 𝐼𝑁 is the noisy input image and 𝐼𝐶 is the coarse despeckling
despeckling, where one path extracts global structural semantics and image by CDM. The loss function for this stage is defined as:
the other preserves local details, with cross-attention dynamically fusing
both features. Li et al. (2024a), combined contrastive learning with se- 𝐷1 = ‖𝐼𝐶 − 𝐼𝐺𝑇 ‖1 (2)
mantic guidance for underwater image enhancement, maintaining com- where 𝐼𝐺𝑇 is the clean image and || ⋅ ||1 denotes the L1 loss.
putational efficiency through frozen segmentation network parameters. In the next stage, the preprocessed output 𝐼𝐶 from the CDM is then
Inspired by these methods, we propose SemDNet, which combines both passed through the SSM. For the SSM, we adopt DeepLabV3+ (Chen
semantic loss and semantic feature integration to improve SAR image et al., 2018) with a ResNet101 backbone, pre-trained on noiseless re-
despeckling. mote sensing datasets. The SSM extracts high-level semantic features 𝐹𝐶 ,
which encapsulate contextual and structural knowledge of the scene. To
3. Proposed method ensure stability and consistency, the weights of the SSM are frozen dur-
ing the training of the entire despeckling network. This configuration
In this section, we present SemDNet, a novel framework that inte- allows the SSM to act as a robust semantic encoder, providing domain-
grates semantic segmentation into SAR image despeckling to enhance relevant guidance for the subsequent fine despeckling stage. The feature
performance through effective interaction between noise suppression extraction process can be expressed as:
and semantic understanding. The architecture of SemDNet is illustrated
in Fig. 1. The following subsections provide a detailed description of 𝐹𝐶 = 𝑓𝑆𝑆𝑀 (𝐼𝐶 ) (3)
SemDNet’s components, network architecture, and loss functions.
where 𝐹𝐶 is the semantic features obtained by SSM. To optimize the
SSM, we define a hybrid loss function:
3.1. CDM and SSM
𝐿𝑆𝑒𝑔 = 𝛼𝐿𝐶𝐸 + (1 − 𝛼)𝐿𝐷𝑖𝑐𝑒 (4)
Speckle noise severely disrupts image details and structural features,
which can lead to significant misclassification if noisy images are di- where 𝛼 is the balance parameter empirically set to 0.5. 𝐿𝐶𝐸 and 𝐿𝐷𝑖𝑐𝑒
rectly fed into the SSM. To mitigate this issue, we introduce CDM as denote cross entropy (CE) loss and Dice loss, respectively. The CE loss
a preprocessing step. The primary objective of CDM is to suppress the is defined as:

1 ∑∑
majority of the noise while preserving critical structural and textural 𝑁 𝐶

details in the image. This step provides a cleaner input for the sub- 𝐿𝐶𝐸 = − 𝑦 log(𝑝𝑖,𝑐 ), (5)
𝑁 𝑖=1 𝑐=1 𝑖,𝑐
sequent semantic segmentation process, thereby improving segmenta-
tion accuracy. Instead of designing a new network architecture for the where 𝑁 is the total number of pixels, 𝐶 is the number of classes, 𝑦𝑖,𝑐 is
CDM, we adopt existing denoising networks to perform this task. In this the ground truth label for class 𝑐 for the 𝑖th pixel (1 if the sample belongs
study, we consider RIDNet (Anwar & Barnes, 2019) and SADNet (Chang to class 𝑐, 0 otherwise), and 𝑝𝑖,𝑐 is the predicted probability for class 𝑐.
et al., 2020) as candidate networks for the CDM. This strategy is moti- While CE loss ensures accurate pixel-wise classification, it may struggle
vated by two key considerations. First, using proven denoising architec- with imbalanced datasets, as it weights all classes equally. To address
tures ensures reliable performance, creating a strong foundation for the the limitations of CE loss, we incorporate Dice loss, which focuses on

Fig. 1. Overall framework of our proposed SemDNet. SemDNet consists of three key modules: the coarse despeckling module (CDM), the semantic segmentation
module (SSM), and the semantic guided despeckling module (SGDM). The primary goal of the CDM is to suppress the majority of the noise and provide a cleaner
input for the subsequent semantic segmentation process. The SSM is responsible for extracting high-level semantic features, while the SGDM combines low-level
pixel information with high-level semantic features to achieve more effective noise suppression.

3
F. Bo et al. Expert Systems With Applications 296 (2026) 129200

Fig. 2. Network architecture of SGDM. The SGDM consists of two parallel input branches: the upper branch processes high-level semantic features, while the lower
branch handles low-level image features. After feature extraction through the ResBlock, the outputs from both branches are merged in the SF block.

the overlap between predicted and ground truth regions. It is defined


as:
∑ ∑𝐶
2 𝑁 𝑦𝑖,𝑐 𝑝𝑖,𝑐
𝐿𝐷𝑖𝑐𝑒 = 1 − ∑𝑁 ∑𝐶 𝑖=1 𝑐=1 ∑𝑁 ∑𝐶 , (6)
𝑖=1 𝑦
𝑐=1 𝑖,𝑐 + 𝑖=1 𝑐=1 𝑝𝑖,𝑐

Unlike CE loss, Dice loss is robust to class imbalance, making it


particularly effective in scenarios where certain regions dominate the
dataset. This hybrid loss formulation ensures both pixel-level precision
and reliable region-based segmentation, enabling the SSM to provide
accurate and meaningful semantic guidance.

3.2. Semantic guided despeckling module

As discussed earlier, traditional pixel-based despeckling methods Fig. 3. Network architecture of semantic fusion block. Image features are di-
face inherent limitations, particularly their inability to capture contex- vided into components based on the number of semantic segmentation classes.
The semantic features are then unfolded to generate corresponding class-specific
tual information and their challenges in balancing noise suppression
attention masks, which are used for spatial multiplication with the image fea-
with detail preservation. To address these shortcomings, the SGDM is
tures to emphasize or suppress specific regions.
proposed in this paper, as illustrated in Fig. 2, which incorporates high-
level semantic features into the despeckling pipeline, enabling effective
noise reduction while preserving essential structural and contextual de- This design ensures a rich representation that combines detailed spatial
tails. The SGDM begins with two parallel input branches. The upper information with robust semantic guidance. The feature fusion process
branch processes the semantic features generated by the SSM, while the is repeated three times, with each iteration comprising a series of Res-
lower branch takes a concatenation of the coarse despeckled image and Blocks and SF Blocks. Each repetition refines the feature representations,
the original noisy image as input. The overall operation of the SGDM enabling the network to achieve a balance between noise suppression
can be expressed as: and detail preservation. The final despeckled image is obtained as the
output of this module. The loss function for the SGDM is defined as:
𝐼𝐷 = 𝑓𝑆𝐺𝐷𝑀 [𝑐𝑎𝑡(𝐼𝐶 , 𝐼𝑁 ), 𝐹𝐶 ] (7)
𝐷2 = ‖𝐼𝐷 − 𝐼𝐺𝑇 ‖1 (8)
where 𝐼𝐷 is the final despeckling image, and 𝑐𝑎𝑡 denotes concatenation
operation. This architecture integrates semantic features at multiple stages,
The inputs of both branches are initially processed through Res- bridging the gap between low-level and high-level processing. This is
Blocks to extract features. Each ResBlock consists of two identical con- fundamentally different from simple feature concatenation or conven-
volutional layers followed by ReLU activation functions. After feature tional attention mechanisms such as SE or CBAM. While SE and CBAM
extraction, the outputs of the two branches are merged into the SF Block. focus on internal channel or spatial attention by learning weights from
The SF Block integrates features from both branches, combining com- the feature tensor itself, they do not incorporate any external semantic
plementary information to enhance the overall feature representation. information (e.g., class labels, semantic priors, or outputs from exter-
The detailed structure of the SF Block is shown in Fig. 3. The image fea- nal networks). In contrast, the SF Block introduces class-wise semantic
tures are first split into 𝐶 components, where 𝐶 represents the number modulation by leveraging a SSM to provide region-specific semantic pri-
of semantic segmentation classes. Simultaneously, the semantic features ors. It enables the network to distinguish various region types, adopt-
are spatially unfolded to align with the spatial dimensions of the image ing a divide-and-conquer approach for more effective despeckling. In
features. Each unfolded semantic feature is then multiplied element- essence, the SF Block retains the advantages of attention-based modu-
wise with its corresponding image feature component (𝑇1 to 𝑇𝐶 ) through lation while incorporating domain-specific knowledge, improving both
spatial-wise multiplication. This process acts as an attention mechanism, interpretability and despeckling performance.
with semantic information emphasizing or suppressing specific regions
of the image features based on their contextual relevance, such as build- 3.3. Implementation details
ings, vegetation, or roads. The products of these multiplications are con-
catenated to form the semantic-fused features, effectively integrating The total loss for SemDNet is defined as:
multi-scale spatial details with semantic context. A convolutional layer
is then applied to the concatenated features and generates the output. 𝐿 = 𝜆1 𝐿𝐷1 + 𝜆2 𝐿𝐷2 + 𝜆3 𝐿𝑆𝑒𝑔 (9)

4
F. Bo et al. Expert Systems With Applications 296 (2026) 129200

Table 1
The average objective indices of each method on synthetic SAR images. Red and (Green) denotes the improvement (reduction) of performance. The bold indicates
the best values and underline indicates the second best values.

where 𝜆1 , 𝜆2 , and 𝜆3 are weight parameters. In this paper, we set them 4.1. Experiments on synthetic SAR images
to 0.5, 2, and 0.5, respectively.
The proposed SemDNet framework was implemented using the Py- We evaluated the proposed SemDNet on various test datasets, includ-
Torch platform and trained on an NVIDIA 2080Ti GPU with 11 GB of ing WHDLD-test, SAR-9,3 and SSAR.4 The evaluation utilized objective
memory. The training process was conducted in two stages. First, the metrics such as peak signal-to-noise ratio (PSNR), structural similarity
CDM and SSM were independently trained for 100 epochs with a learn- index (SSIM) (Wang et al., 2004), and feature similarity index (FSIM)
ing rate of 5 × 10−5 . Subsequently, the CDM and SGDM were fine-tuned (Zhang et al., 2011). These metrics were chosen to comprehensively
together using a reduced learning rate of 1 × 10−5 . assess the quality of despeckled images in terms of noise suppression,
structural preservation, and perceptual similarity.
4. Experiments Table 1 summarizes the quantitative results across the three test
datasets, where SemDNet-R/S represents the proposed improvements
The training process utilized two publicly available datasets: the based on the CDM modules. As shown in Table 1, AGSDNet and MRD-
WHDLD dataset,1 which contains six-class semantic labels (buildings, DANet exhibit poor performance. Similarly, G-MONet, LGDBNet, and
roads, pavements, vegetation, bare soil, and water) annotated by ex- SIFSDNet fail to clearly distinguish fine image structures from noise,
perts from Wuhan University (Shao et al.), and the UCM dataset2 , man- leading to inadequate noise removal and inferior performance. In con-
ually labeled by Yang and Newsam at UC Merced. To adapt these optical trast, our proposed method consistently outperforms other approaches,
datasets for SAR domain training, we first converted all RGB images to including the highly competitive MFAENet. Our framework demon-
grayscale and generated synthetic SAR images by contaminating clean strates superior performance compared to the baseline models, achiev-
versions with multiplicative single-look speckle noise. To ensure the va- ing significant improvements across multiple datasets in terms of objec-
lidity and reliability of the training and evaluation data, we conducted tive metrics. Specifically, SemDNet-R and SemDNet-S achieve average
thorough visual inspections of all image samples and excluded any with PSNR improvements of 0.73 dB and 0.45 dB, respectively, over the base-
ambiguous semantics or poor visual quality. The WHDLD dataset was di- line models across the three datasets. Both SSIM and FSIM also show sig-
vided into training, validation, and testing subsets (7:2:1 ratio), where nificant enhancements, indicating improved structural preservation and
its clean grayscale images trained the SSM, while the synthetic noisy perceptual similarity. Notably, SemDNet-S achieves the highest PSNR
versions of both WHDLD and UCM datasets trained the CDM. The noisy and SSIM values, while SemDNet-R achieves the highest FSIM. This vari-
version of WHDLD was then used to train the full SemDNet framework. ation can be attributed to the robust semantic guidance in SemDNet,
Random horizontal and vertical flips were applied during the training which enhances structural preservation and contextual understanding,
process. particularly in complex scenes.
The proposed SemDNet was compared with several despeckling al- Visual inspection also confirms the analysis of objective metrics. As
gorithms, including AGSDNet (Thakur & Maji, 2022a), G-MONet (Vi- shown in Fig. 4, AGSDNet fails to adequately suppress speckle noise,
tale et al., 2023), LGDBNet (Liu et al., 2024b), MFAENet (Liu et al., leaving residual noise in the despeckled images. Conversely, G-MONet
2023), MRDDANet (Liu et al., 2022), SIFSDNet (Thakur & Maji, 2022b), suffers from excessive smoothing, leading to the loss of critical struc-
SAR-NNFN(Bo et al., 2024), MSANN(Guo et al., 2024), and the baseline tural details. While LGDBNet, MSANN, and SIFSDNet manage to retain
models (Anwar & Barnes, 2019; Chang et al., 2020). Additionally, to some structural information, they introduce artificial artifacts that de-
evaluate the performance on real SAR images, we included advanced grade the overall image quality. Similarly, MFAENet, RIDNet, and SAD-
self-supervised despeckling algorithms, such as SAR2SAR (Dalsasso Net achieve effective noise suppression but struggle to preserve finer
et al., 2021), MERLIN (Dalsasso et al., 2022), and Speckle2Void (Molini details, as highlighted in the red zoomed-in regions. In contrast, the pro-
et al., 2021). posed SemDNet excels at balancing noise suppression and detail preser-

1 3
[Link] [Link]
2 [Link] 4 [Link]

5
F. Bo et al. Expert Systems With Applications 296 (2026) 129200

Fig. 4. Visual comparison of despeckling results on WHDLD-Test dataset. (a) Clean image. (b) AGSDNet. (c) SAR-NNFN. (d) G-MONet. (e) LGDBNet. (f) MFAENet.
(g) MSANN. (h) SIFSDNet. (i) RIDNet. (j) SemDNet-R. (k) SADNet. (l) SemDNet-S.

Fig. 5. Visual comparison of despeckling results on SAR-9 dataset. (a) Clean image. (b) AGSDNet. (c) SAR-NNFN. (d) G-MONet. (e) LGDBNet. (f) MFAENet. (g)
MSANN. (h) SIFSDNet. (i) RIDNet. (j) SemDNet-R. (k) SADNet. (l) SemDNet-S.

vation. It effectively reduces noise while maintaining intricate structural & Saravanan, 2021). Lower NIQE and BRISQUE values indicate better
details, ensuring that even subtle image features remain intact without image quality, reflecting reduced noise and enhanced naturalness of the
introducing artifacts. Comparable results are presented in Fig. 5, which despeckled image. Conversely, higher ENL and TCR values signify im-
further confirms the robustness of SemDNet. These findings highlight proved despeckling performance. ENL is widely used in SAR despeckling
the effectiveness of leveraging semantic information as a guiding prior, to assess the uniformity of homogeneous regions, while TCR evaluates
enhancing the network’s ability to differentiate noise and meaningful the clarity of target structures relative to surrounding clutter. These met-
image structures. This capability underscores the framework’s improve- rics provide a comprehensive assessment of despeckling performance in
ment in addressing the challenges of SAR image despeckling. the absence of reference images. ENL and TCR are defined as:
( )2
4.2. Experiments on real SAR images ̂𝑅𝑜𝐼 )
𝑚𝑒𝑎𝑛(𝑋
𝐸𝑁𝐿 = 20𝑙𝑜𝑔10 ( )2 (10)
To evaluate the performance of the proposed SemDNet on real-world ̂𝑅𝑜𝐼 )
𝑠𝑡𝑑(𝑋
SAR images, we selected five images from diverse scenes,5 as shown in
𝑚𝑎𝑥(𝑋̂𝑅𝑜𝐼 )
Fig. 6. Since clean ground truth images are unavailable for real SAR data, 𝑇 𝐶𝑅 = 20𝑙𝑜𝑔10 (11)
several no-reference evaluation metrics were employed: natural image ̂𝑅𝑜𝐼 )
𝑚𝑒𝑎𝑛(𝑋
quality evaluator (NIQE) (Mittal et al., 2012b), blind/referenceless im-
where 𝑚𝑒𝑎𝑛, 𝑠𝑡𝑑, and 𝑚𝑎𝑥 denote mean, standard deviation, and maxi-
age spatial quality evaluator (BRISQUE) (Mittal et al., 2012a), equiva-
mum operations, respectively. 𝑅𝑜𝐼 represents the region of interest.
lent number of looks (ENL), and target-to-clutter ratio (TCR) (Ponmani
Visual inspection serves as a crucial qualitative method to assess the
effectiveness of despeckling methods. Fig. 7 presents the despeckling
5 [Link] results for R1. It is evident that AGSDNet and MRDDANet exhibit lim-

6
F. Bo et al. Expert Systems With Applications 296 (2026) 129200

Fig. 6. Real SAR images. (a) R1. (b) R2. (c) R3. (d) R4. (e) R5 The green and blue boxes are selected ROI to calculate TCR and ENL, respectively.

Fig. 7. Visual comparison of despeckling results for R3. (a) AGSDNet. (b) G-MONet. (c) LGDBNet. (d) MFAENet. (e) MRDDANet. (f) SIFSDNet. (g) SAR-NNFN. (h)
MSANN. (i) SAR2SAR. (j) Speckle2Void. (k) MERLIN. (l) RIDNet. (m) SemDNet-R. (n) SADNet. (o) SemDNet-S.

ited speckle suppression capabilities, leaving significant residual noise ear structures, as highlighted by the red boxes in Fig. 8. The find-
in the despeckled images. Methods such as LGDBNet, SIFSDNet, SAR- ings highlight the robustness of SemDNet in leveraging semantic in-
NNFN, and MERLIN remove some noise but introduce noticeable ar- formation to distinguish noise from meaningful features, leading to
tifacts or retain granular noise in homogeneous regions. G-MONet, clearer, artifact-free despeckled images. Additionally, Fig. 9 shows the
SAR2SAR, and Speckle2Void tend to over-smooth the images, leading to ratio images (the original SAR image divided by the despeckled image)
blurred edge structures and a loss of crucial details. Similarly, MSANN, for R3.
MFAENet, RIDNet, and SADNet, which perform despeckling at a same Since noise in SAR images is multiplicative, an ideal ratio image af-
scale, fail to effectively preserve texture details while suppressing noise. ter despeckling should contain only noise without any texture or struc-
In contrast, the proposed SemDNet demonstrates superior performance tural features. As shown in Fig. 9, G-MONet, SAR-NNFN, SAR2SAR,
by achieving a balanced trade-off between noise reduction and detail Speckle2Void, and MERLIN exhibit noticeable texture leakage, indicat-
preservation. ing their inability to preserve textures effectively during despeckling.
The despeckling results for R3 are presented in Fig. 8, where similar LGDBNet, and SIFSDNet exhibit less leakage, their performance in pre-
observations can be obtained. We can see that the proposed SemDNet serving structural details remains suboptimal. In contrast, the sparse
outperforms baseline models, delivering images with significantly en- ratio images of AGSDNet and MRDDANet reflect inadequate speckle
hanced perceptual quality. Notably, SemDNet effectively restores lin- suppression. The remaining methods, including the proposed SemD-

7
F. Bo et al. Expert Systems With Applications 296 (2026) 129200

Fig. 8. Visual comparison of despeckling results for R3. (a) AGSDNet. (b) G-MONet. (c) LGDBNet. (d) MFAENet. (e) MRDDANet. (f) SIFSDNet. (g) SAR-NNFN. (h)
MSANN. (i) SAR2SAR. (j) Speckle2Void. (k) MERLIN. (l) RIDNet. (m) SemDNet-R. (n) SADNet. (o) SemDNet-S.

Fig. 9. Ratio images for R3 obtained by different despeckling methods. (a) AGSDNet. (b) G-MONet. (c) LGDBNet. (d) MFAENet. (e) MRDDANet. (f) SIFSDNet. (g)
SAR-NNFN. (h) SAR2SAR. (i) Speckle2Void. (j) MERLIN. (k) RIDNet. (l) SemDNet-R. (m) SADNet. (n) SemDNet-S.

Net, show minimal texture leakage, demonstrating their effectiveness in noise suppression, they fail to adequately retain the fine structures in the
suppressing speckle noise while retaining underlying texture and struc- image. These observations are further corroborated by the ratio images.
tural features. In contrast, SemDNet achieves a remarkable balance between noise sup-
Another case is shown in Figs. 10 and 11. G-MONet, SAR2SAR, and pression and detail preservation. It effectively retains subtle structural
Speckle2Void continue to exhibit poor performance in preserving im- features and textures, demonstrating its robustness and superior perfor-
age details. AGSDNet and MRDDANet overly retain noise, leading to mance across diverse scenarios.
unsatisfactory despeckling results. While MFAENet, RIDNet, and SAD- The quantitative evaluation results are presented in Table 2. It
Net perform better than LGDBNet, SIFSDNet, MSANN, and MERLIN in can be observed that the proposed SemDNet consistently improves the

8
F. Bo et al. Expert Systems With Applications 296 (2026) 129200

Fig. 10. Visual comparison of despeckling results for R3. (a) AGSDNet. (b) G-MONet. (c) LGDBNet. (d) MFAENet. (e) MRDDANet. (f) SIFSDNet. (g) SAR-NNFN. (h)
MSANN. (i) SAR2SAR. (j) Speckle2Void. (k) MERLIN. (l) RIDNet. (m) SemDNet-R. (n) SADNet. (o) SemDNet-S.

Fig. 11. Ratio images for R3 obtained by different despeckling methods. (a) AGSDNet. (b) G-MONet. (c) LGDBNet. (d) MFAENet. (e) MRDDANet. (f) SIFSDNet. (g)
SAR-NNFN. (h) SAR2SAR. (i) Speckle2Void. (j) MERLIN. (k) RIDNet. (l) SemDNet-R. (m) SADNet. (n) SemDNet-S.

performance of the baseline models, with the exception of runtime. highest ENL score along with a relatively high TCR value, reflecting
SemDNet-S achieves the best NIQE score and the second-best BRISQUE its ability to balance noise suppression with detail preservation. The
score, demonstrating its effectiveness in enhancing image quality. It is final column of Table 2 presents the runtime of all methods, while
important to note that ENL and TCR are mutually opposing metrics. Table 3 provides details on the number of parameters and FLOPs for
Excessive smoothing tends to increase ENL but decreases TCR, and each SemDNet module. Although the proposed SGDM has the smallest
vice versa. For instance, MERLIN and MRDDANet achieve high TCR number of parameters and moderate FLOPs, the integration of seman-
scores but low ENL, indicating insufficient noise suppression. Con- tic information introduces additional computational complexity, result-
versely, G-MONet and SAR2SAR achieve high ENL but at the ex- ing in a longer runtime compared to baselines. However, the significant
pense of low TCR, reflecting their tendency to oversmooth and fail improvement in despeckling performance justifies the added computa-
in preserving structural details. In contrast, SemDNet achieves the tional cost.

9
F. Bo et al. Expert Systems With Applications 296 (2026) 129200

Table 2
The average objective indices of each method on real SAR images. Red and (Green) denotes the improvement (reduction) of performance.
The bold indicates the best values and underline indicates the second best values.

Table 3 Table 4
Comparison of Parameters (Unit: M) and FLOPs (Unit: Results of ablation studies on synthetic and real SAR images.
G) of Different Modules in SemDNet.
SAR-9 Real SAR
SADNet RIDNet DeepLabV3+ SGDM
SGDM 𝐿𝑆𝑒𝑔 PSNR SSIM NIQE BRISQUE
Para. 3.45 5.89 59.34 1.91 × × 25.80 0.7534 4.91 37.37
Flops 18.85 391.03 22.24 124.95 ✓ × 26.01 0.7594 4.52 34.68
× ✓ 25.92 0.7568 4.67 35.26
✓ ✓ 26.10 0.7612 4.29 30.26

4.3. Ablation study

To evaluate the effectiveness of each component in the proposed


SemDNet, we conducted an ablation study using SADNet as the base-
line model on the SAR-9 dataset. As shown in Table 3, incorporating
the SGDM significantly improved both PSNR and SSIM, demonstrating
its ability to enhance despeckling performance by effectively integrat-
ing semantic guidance with low-level image features. Similarly, the in-
clusion of the semantic loss 𝐿𝑆𝑒𝑔 further boosted performance, as evi-
denced by higher evaluation metrics. This indicates that the semantic
loss contributes to the preservation of meaningful structural details and
reduces noise artifacts. In addition to synthetic datasets, we also con-
ducted ablation experiments on real SAR images to validate the gen-
eralization capability of each component. The perceptual quality was
evaluated using no-reference image quality assessment metrics, includ-
ing NIQE and BRISQUE. The results revealed consistent improvements
with each module. These findings suggest that the combination of SGDM
and 𝐿𝑆𝑒𝑔 plays a pivotal role in the success of SemDNet, enabling the net-
work to achieve a robust balance between noise suppression and detail
preservation (Table 4).
We also conduct ablation experiments to investigate the impact of
hyperparameters 𝜆1 , 𝜆2 , and 𝜆3 in balancing the losses 𝐿𝐷1 , 𝐿𝐷2 , and
𝐿𝑆𝑒𝑔 , respectively. We fixed two hyperparameters while systematically
Fig. 12. Ablation study on loss weights of our framework.
varying the third. The results of these experiments are summarized in
Fig. 12. Based on the experimental findings, we empirically set 𝜆1 = 0.5,
𝜆2 = 2, and 𝜆3 = 0.5 as they yield the best performance. While we ac-
knowledge that more exhaustive parameter exploration could theoret- current configuration provides robust performance while maintaining
ically yield marginally better results, our experiments suggest that the reasonable computational requirements.

10
F. Bo et al. Expert Systems With Applications 296 (2026) 129200

5. Discussion

5.1. Sensitivity analysis

The accuracy of semantic segmentation has a certain impact on


the denoising performance of the proposed SemDNet. To better un-
derstand how segmentation errors impact despeckling performance, we
conducted a sensitivity analysis under two representative scenarios:

(1) Semantically consistent misclassification: This scenario refers to


cases where the segmentation network mislabels a region as a se-
mantically similar class. For example, as shown in Fig. 13, a flat
grassland area may be misclassified as water. Our experimental Fig. 15. Visual assessment results by experts for baseline models and SemDNet.
analysis indicates that the proposed method exhibits significant
robustness against semantically consistent misclassifications, with
the performance gain from semantic guidance largely unaffected. maintains effective performance as demonstrated in our comprehen-
This robustness can be attributed to the fact that semantic guidance sive evaluations.
operates based on generalized feature representations rather than
being strictly dependent on categorical boundaries. Consequently,
when the underlying semantic context remains consistent, minor 5.2. Expert evaluation
discrepancies in segmentation labels do not significantly influence
the denoising quality. This result highlights the advantage of using While quantitative metrics are standard in SAR despeckling, expert
a more flexible, context-aware approach in semantic guidance. validation would further enhance the reliability and practical relevance
(2) Semantically inconsistent misclassification: In contrast, semanti- of the evaluation. To this end, we conducted a blind visual evaluation
cally inconsistent misclassifications present a greater challenge. As involving five experienced remote sensing experts. Each expert indepen-
illustrated in Fig. 14, when a textured soil region is misclassified as dently scored the despeckling results from the baseline models and our
grassland, the semantic prior becomes misleading. Such errors can SemDNet on five synthetic and five real SAR images. The evaluation cov-
arise from excessive smoothing in the coarse despeckling stage, seg- ered four aspects: sharpness, despeckling, structural fidelity, and overall
mentation inaccuracies, or the influence of residual noise. This mis- visual quality, with each criterion rated on a 1-to-5 scale. As shown in
match in semantic context leads to a degradation in performance, Fig. 15, SemDNet-R showed the most notable improvement over RID-
particularly in regions requiring precise edge preservation and Net, highlighting the benefit of semantic guidance. Similarly, SemDNet-
structured pattern recognition. The misdirected semantic guidance S received higher scores than SADNet from the majority of experts, in-
results in over-smoothing or incorrect interpretation of complex dicating its advantages in preserving fine details and improving overall
textures, which compromises the overall denoising quality. These perceptual quality. These findings further support the practical effective-
findings emphasize the importance of accurate semantic segmenta- ness of our method beyond quantitative indicators. Notably, achieving
tion, particularly in areas where fine structural details are critical. It large-scale and consistent expert annotations for SAR despeckling is in-
should be noted that such cases represent rare exceptions rather than herently challenging. Such annotation efforts require extensive domain
the norm. In broader operational scenarios, the proposed method knowledge and significant human resources, limiting their feasibility in
practice.

6. Conclusion

In this paper, we proposed SemDNet, a novel framework that com-


bines SAR image despeckling with semantic segmentation to enhance
the performance of existing despeckling methods. The core innovation
of this framework lies in the design of the SGDM, which explicitly
fuses low-level pixel-based information with high-level semantic fea-
tures. This integration ensures effective noise suppression while preserv-
ing critical structural and textural details in SAR images. The inclusion
of the semantic loss function as semantic constraints further strengthens
Fig. 13. Semantically consistent misclassification. (a) SADNet (PSNR:30.86,
the network’s ability to retain meaningful information and accurately re-
SSIM:0.8934, FSIM:0.8856). (b) Segmentation map. (c) SemDNet-S
construct fine image details. By incorporating semantic guidance, SemD-
(PSNR:31.60, SSIM:0.9004, FSIM:0.8929).
Net bridges the gap between noise reduction and detail preservation, ad-
dressing the limitations of traditional despeckling methods. Experiments
conducted on both synthetic and real-world SAR images demonstrate
that SemDNet significantly outperforms existing state-of-the-art meth-
ods in terms of both quantitative metrics and visual quality. Moreover,
it achieves substantial performance improvements over baseline models.
However, the incorporation of semantic segmentation introduces addi-
tional computational complexity and a heightened sensitivity to seg-
mentation accuracy. Although the model exhibits robustness to seman-
tically consistent misclassifications, its performance tends to degrade in
the presence of inconsistent semantic errors. This limitation arises from
Fig. 14. Semantically inconsistent misclassification. (a) SADNet (PSNR:28.11, the inherent reliance on segmentation accuracy. To address this chal-
SSIM:0.7802, FSIM:0.8641). (b) Segmentation map. (c) SemDNet-S lenge, future work could explore the adoption of more sophisticated
(PSNR:28.12, SSIM:0.7781, FSIM:0.8629). segmentation models, while employing techniques such as knowledge

11
F. Bo et al. Expert Systems With Applications 296 (2026) 129200

distillation to reduce computational complexity without compromising Dalsasso, E., Denis, L., & Tupin, F. (2021). SAR2SAR: A semi-supervised despeckling
performance. algorithm for SAR images. IEEE Journal of Selected Topics in Applied Earth Observations
and Remote Sensing, 14, 4321–4329. [Link]
Dalsasso, E., Denis, L., & Tupin, F. (2022). As if by magic: Self-supervised training of
CRediT authorship contribution statement deep despeckling networks with MERLIN. IEEE Transactions on Geoscience and Remote
Sensing, 60, 1–13. [Link]
Deng, J.-W., Li, M.-D., & Chen, S.-W. (2024). Sublook2Sublook: A self-supervised speckle
Fuyu Bo: Writing - original draft, Formal analysis, Methodology; Yi filtering framework for single SAR images. IEEE Transactions on Geoscience and Remote
Jin: Data curation, Investigation; Xiaole Ma: Supervision, Data cura- Sensing, 62, 1–13.
Gragnaniello, D., Poggi, G., Scarpa, G., & Verdoliva, L. (2016). Sar image despeckling
tion; Yigang Cen: Writing - review & editing; Shaohai Hu: Supervision,
by soft classification. IEEE Journal of Selected Topics in Applied Earth Observations and
Funding acquisition; Yidong Li: Validation, Data curation. Remote Sensing, 9(6), 2118–2130.
Guo, Y., Lu, Y., Liu, R. W., & Zhu, F. (2024). Blind image despeckling using a mul-
tiscale attention-guided neural network. IEEE Transactions on Artificial Intelligence,
Uncited Refrerences 5(1), 205–216.
Huang, Z., Liu, Z., He, P., Ren, Y., Li, S., Lei, Y., Luo, D., Liang, D., Shao, D., Hu, Z. et al.
Zheng et al. (2022) (2022). Segmentation-guided denoising network for low-dose CT imaging. Computer
Methods and Programs in Biomedicine, 227, 107199.
Kan, S., Cen, Y., Li, Y., Vladimir, M., & He, Z. (2022). Local semantic correlation modeling
Data availability over graph neural networks for deep feature embedding and image retrieval. IEEE
Transactions on Image Processing, 31, 2988–3003.
Li, F., Zheng, J., Wang, L., & Wang, S. (2024a). Integrating cross-domain feature rep-
Data will be made available on request. resentation and semantic guidance for underwater image enhancement. IEEE Signal
Processing Letters, 31, 1511–1515. [Link]
Li, S., Liu, M., Zhang, Y., Chen, S., Li, H., Dou, Z., & Chen, H. (2024b). SAM-Deblur: Let
Declaration of competing interest segment anything boost image deblurring. In Proceedings of the international conference
on acoustics, speech and signal processing (ICASSP) (pp. 2445–2449).
Li, Y., Chang, Y., Yu, C., & Yan, L. (2022). Close the loop: A unified bottom-up and
The authors declare that they have no known competing financial
top-down paradigm for joint image deraining and segmentation. In Proceedings of
interests or personal relationships that could have appeared to influence the association for the advancement of artificial intelligence (AAAI) (pp. 1438–1446).
the work reported in this paper. (vol. 36).
Lin, H., Zhuang, Y., Huang, Y., & Ding, X. (2023). Unpaired speckle extraction for SAR
despeckling. IEEE Transactions on Geoscience and Remote Sensing, 61, 1–14. https:
Acknowledgments //[Link]/10.1109/TGRS.2022.3233892
Liu, D., Wen, B., Jiao, J., Liu, X., Wang, Z., & Huang, T. S. (2020). Connecting image
denoising and high-level vision tasks via deep learning. IEEE Transactions on Image
This work was supported in part by the Fundamental Research Funds Processing, 29, 3695–3706.
for the Central Universities, China (2025YJS054), in part by the National Liu, D., Wen, B., Liu, X., Wang, Z., & Huang, T. S. (2017). When image denoising meets
Key Research and Development Program, China (2022YFB3103500), high-level vision tasks: A deep learning approach. arXiv preprint arXiv:1706.04284.
Liu, S., Lei, Y., Zhang, L., Li, B., Hu, W., & Zhang, Y.-D. (2022). MRDDANet: A multi-
in part by the National Natural Science Foundation, China (62473033, scale residual dense dual attention network for SAR image denoising. IEEE Transac-
62172030 and 62202036), in part by the Beijing Natural Science Foun- tions on Geoscience and Remote Sensing, 60, 1–13. [Link]
dation under Grant L231012. 3106764
Liu, S., Lu, J., Dou, H., Li, J., & Deng, Y. (2024a). SEGSID: A semantic-guided framework
for sonar image despeckling. IEEE Transactions on Image Processing, 34, 652–666.
References Liu, S., Tian, S., Zhao, Y., Hu, Q., Li, B., & Zhang, Y.-D. (2024b). LG-DBNet: Local and
global dual-branch network for SAR image denoising. IEEE Transactions on Geoscience
Aakerberg, A., Johansen, A. S., Nasrollahi, K., & Moeslund, T. B. (2022). Semantic seg- and Remote Sensing, 62, 1–15. [Link]
mentation guided real-world super-resolution. In Proceedings of the IEEE/CVF winter Liu, S., Zhang, L., Tian, S., Hu, Q., Li, B., & Zhang, Y. (2023). MFAENet: A multi-scale
conference on applications of computer vision (WACV) (pp. 449–458). feature adaptive enhancement network for SAR image despeckling. IEEE Journal of
Anwar, S., & Barnes, N. (2019). Real image denoising with feature attention. In Proceed- Selected Topics in Applied Earth Observations and Remote Sensing, 16, 10420–10433.
ings of the IEEE/CVF international conference on computer vision (ICCV) (pp. 3155–3164). Ma, Y., Ke, P., Aghababaei, H., Chang, L., & Wei, J. (2024). Despeckling SAR images
Aranda-Bojorges, G., Ponomaryov, V., Reyes-Reyes, R., Sadovnychiy, S., & Cruz-Ramos, with Log-Yeo-Johnson transformation and conditional diffusion models. IEEE Trans-
C. (2022). Clustering-based 3-d-MAP despeckling of SAR images using sparse wavelet actions on Geoscience and Remote Sensing, 62, 1–17. [Link]
representation. IEEE Geoscience and Remote Sensing Letters, 19, 1–5. [Link] 2024.3419083
10.1109/LGRS.2021.3108774 Mittal, A., Moorthy, A. K., & Bovik, A. C. (2012a). No-reference image quality assessment
Bao, X., Zhang, R., Lv, J., Wu, R., Zhang, H., Chen, J., Zhang, B., Ouyang, X., & Liu, G. in the spatial domain. IEEE Transactions on Image Processing, 21(12), 4695–4708.
(2023). Vegetation descriptors from sentinel-1 SAR data for crop growth monitoring. Mittal, A., Soundararajan, R., & Bovik, A. C. (2012b). Making a “completely blind” image
ISPRS Journal of Photogrammetry and Remote Sensing, 203, 86–114. quality analyzer. IEEE Signal Processing Letters, 20(3), 209–212.
Baraha, S., & Sahoo, A. K. (2023). Synthetic aperture radar image and its despeckling Molini, A. B., Valsesia, D., Fracastoro, G., & Magli, E. (2021). Speckle2Void: Deep self-
using variational methods: A review of recent trends. Signal Process., 212, 109156. supervised SAR despeckling with blind-spot convolutional neural networks. IEEE
Bo, F., Lu, W., Wang, G., Zhou, M., Wang, Q., & Fang, J. (2022). A blind SAR image Transactions on Geoscience and Remote Sensing, 60, 1–17.
despeckling method based on improved weighted nuclear norm minimization. IEEE Norouzi, J., Helfroush, M. S., Liaghat, A., & Danyali, H. (2024). A deep-based approach
Geoscience and Remote Sensing Letters, 19, 1–5. for multi-descriptor feature extraction: Applications on SAR image registration. Expert
Bo, F., Ma, X., Cen, Y., & Hu, S. (2024). SAR image speckle reduction based on nuclear Systems with Applications, 254, 124291.
norm minus frobenius norm regularization. IEEE Transactions on Geoscience and Remote Pan, Y., Khan, I. A., & Meng, H. (2023). SAR-to-optical image translation using multi-
Sensing, 62, 1–15. [Link] stream deep ResCNN of information reconstruction. Expert Systems with Applications,
Bo, F., Ma, X., Hu, S., An, G., Li, Y., & Cen, Y. (2025). Speckle-driven unsupervised 224, 120040.
despeckling for SAR images. IEEE Journal of Selected Topics in Applied Earth Obser- Parrilli, S., Poderico, M., Angelino, C. V., & Verdoliva, L. (2011). A nonlocal SAR im-
vations and Remote Sensing, 18, 13023–13034. [Link] age denoising algorithm based on LLMMSE wavelet shrinkage. IEEE Transactions on
3568854 Geoscience and Remote Sensing, 50(2), 606–616.
Chang, M., Li, Q., Feng, H., & Xu, Z. (2020). Spatial-adaptive network for single image Penna, P. A. A., & Mascarenhas, N. D. A. (2019). Sar speckle nonlocal filtering with statis-
denoising. In Proceedings of the european conference on computer vision (ECCV) (pp. tical modeling of haar wavelet coefficients and stochastic distances. IEEE Transactions
171–187). on Geoscience and Remote Sensing, 57(9), 7194–7208.
Chen, L.-C., Zhu, Y., Papandreou, G., Schroff, F., & Adam, H. (2018). Encoder-decoder Perera, M. V., Nair, N. G., Bandara, W. G. C., & Patel, V. M. (2023). SAR despeckling
with atrous separable convolution for semantic image segmentation. In Proceedings of using a denoising diffusion probabilistic model. IEEE Geoscience and Remote Sensing
the european conference on computer vision (ECCV) (pp. 801–818). Letters, 20, 1–5. [Link]
Chen, Y., Guo, Z., Yuan, J., Li, X., & Yu, H. (2024). Dual-TranSpeckle: Dual-pathway Ponmani, E., & Saravanan, P. (2021). Image denoising and despeckling methods for SAR
transformer based encoder-decoder network for medical ultrasound image despeck- images to improve image enhancement performance: a survey. Multimedia Tools and
ling. Computers in Biology and Medicine, 173, 108313. Applications, 80(17), 26547–26569.
Cheng, L., Guo, Z., Li, Y., & Xing, Y. (2023). Two-stream multiplicative heavy-tail noise Rasti, B., Chang, Y., Dalsasso, E., Denis, L., & Ghamisi, P. (2021). Image restoration for
despeckling network with truncation loss. IEEE Transactions on Geoscience and Remote remote sensing: Overview and toolbox. IEEE Geoscience and Remote Sensing Magazine,
Sensing, , 61, 1–17. 10(2), 201–230.
Chierchia, G., Cozzolino, D., Poggi, G., & Verdoliva, L. (2017). SAR image despeckling Saha, A., Arihant, K. R., & Maji, S. K. (2023). PCENet: Deep SAR despeckling network
through convolutional neural networks. In Proceedings of the IEEE international geo- utilizing parallel convolutional encoding modules. IEEE Geoscience and Remote Sensing
science and remote sensing symposium (IGARSS) (pp. 5438–5441). IEEE. Letters, 21, 1–5.

12
F. Bo et al. Expert Systems With Applications 296 (2026) 129200

Saleh, T., Weng, X., Holail, S., Hao, C., & Xia, G.-S. (2024). DAM-Net: Flood detection Xiao, S., Zhang, S., Huang, L., & Wang, W.-Q. (2024). Trans-NLM network for SAR image
from SAR imagery using differential attention metric-based vision transformers. ISPRS despeckling. IEEE Transactions on Geoscience and Remote Sensing, 62, 1–12. https:
Journal of Photogrammetry and Remote Sensing, 212, 440–453. //[Link]/10.1109/TGRS.2024.3397325
Sun, Y., Lei, L., Guan, D., Li, X., & Kuang, G. (2020). SAR image speckle reduction based on Xu, L., Liu, P., & Jin, Y.-Q. (2024). A new nonlocal iterative trilateral filter for SAR
nonconvex hybrid total variation model. IEEE Transactions on Geoscience and Remote images despeckling. IEEE Transactions on Geoscience and Remote Sensing, 62, 1–19.
Sensing, 59(2), 1231–1249. [Link]
Thakur, R. K., & Maji, S. K. (2022a). AGSDNet: Attention and gradient-based SAR denois- Yuan, Y., Wu, Y., Feng, P., Fu, Y., & Wu, Y. (2023). Segmentation-guided semantic-aware
ing network. IEEE Geoscience and Remote Sensing Letters, 19, 1–5. [Link] self-supervised denoising for SAR image. IEEE Transactions on Geoscience and Remote
1109/LGRS.2022.3166565 Sensing, 61, 1–16. [Link]
Thakur, R. K., & Maji, S. K. (2022b). SIFSDNet: Sharp image feature based SAR denoising Zhang, J., Chen, J., Yu, H., Yang, D., Xu, X., & Xing, M. (2021). Learning an SAR image
network. In Proceedings of the IEEE international geoscience and remote sensing symposium despeckling model via weighted sparse representation. IEEE Journal of Selected Topics
(IGARSS) (pp. 3428–3431). [Link] in Applied Earth Observations and Remote Sensing, 14, 7148–7158. [Link]
Tsokas, A., Rysz, M., Pardalos, P. M., & Dipple, K. (2022). SAR data applications in earth 1109/JSTARS.2021.3097119
observation: An overview. Expert Systems with Applications, 205, 117342. Zhang, L., Zhang, L., Mou, X., & Zhang, D. (2011). FSIM: A feature similarity index
Vitale, S., Ferraioli, G., Frery, A. C., Pascazio, V., Yue, D.-X., & Xu, F. (2023). SAR de- for image quality assessment. IEEE Transactions on Image Processing, 20(8), 2378
speckling using multiobjective neural network trained with generic statistical samples. –2386.
IEEE Transactions on Geoscience and Remote Sensing, 61, 1–12. [Link] Zhang, N., Fang, J., Bo, F., Mao, T., Song, Y., Zhao, Y., & Gao, J. (2023a). Grid-guided
TGRS.2023.3314857 localization network based on the spatial attention mechanism for synthetic aperture
Wang, C., Guo, B., & He, F. (2023). A novel SAR image despeckling method based on radar ship detection. Journal of Applied Remote Sensing, 17(2), 024505.
local filter with nonlocal preprocessing. IEEE Journal of Selected Topics in Applied Earth Zhang, Q., Liu, X., Li, W., Chen, H., Liu, J., Hu, J., Xiong, Z., Yuan, C., & Wang, Y. (2024a).
Observations and Remote Sensing, 16, 2915–2930. Distilling semantic priors from SAM to efficient image restoration models. In Proceed-
Wang, G., Bo, F., Chen, X., Lu, W., Hu, S., & Fang, J. (2022). A collaborative despeckling ings of the IEEE/CVF conference on computer vision and pattern recognition (CVPR) (pp.
method for SAR images based on texture classification. Remote Sensing, 14(6), 1465. 25409–25419).
Wang, Z., Bovik, A. C., Sheikh, H. R., & Simoncelli, E. P. (2004). Image quality assessment: Zhang, Y., Zeng, G.-Q., Chen, M.-R., Geng, G.-G., Weng, J., & Lu, K.-D. (2024b). DoFA:
from error visibility to structural similarity. IEEE Transactions on Image Processing, Adversarial examples detection for SAR images by dual-objective feature attribution.
13(4), 600–612. [Link] Expert Systems with Applications, 255, 124705.
Wei, Y., Zhang, Z., Zheng, H., Hong, R., Yang, Y., & Wang, M. (2022). SGINet: Toward Zhang, Y., Zhang, F., Jin, Y., Cen, Y., Voronin, V., & Wan, S. (2023b). Local correla-
sufficient interaction between single image deraining and semantic segmentation. In tion ensemble with GCN based on attention features for cross-domain person re-ID.
Proceedings of the 30th ACM international conference on multimedia (ACM MM) (pp. ACM Transactions on Multimedia Computing, Communications, and Applications, 19(2), 1
6202–6210). –22.
Wu, Y., Pan, C., Wang, G., Yang, Y., Wei, J., Li, C., & Shen, H. T. (2023). Learn- Zheng, S., Lu, C., Wu, Y., & Gupta, G. (2022). SAPNet: Segmentation-aware progressive
ing semantic-aware knowledge guidance for low-light image enhancement. In Pro- network for perceptual contrastive deraining. In Proceedings of the IEEE/CVF winter
ceedings of the IEEE/CVF conference on computer vision and pattern recognition (CVPR) conference on applications of computer vision (pp. 52–62).
(pp. 1662–1671).

13

You might also like