0% found this document useful (0 votes)
11 views532 pages

Transformer-Based Radar Signal Processing

This document contains abstracts of various research articles focusing on advanced technologies such as millimeter-wave radar, transformers in autonomous driving, and deep learning applications in radar-based sensing and forecasting. Key themes include the evaluation of existing methodologies, the introduction of novel algorithms, and the exploration of challenges and future directions in these fields. The articles collectively contribute to the understanding and development of intelligent sensing, communication strategies, and environmental monitoring technologies.

Uploaded by

zh.assanova98
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
11 views532 pages

Transformer-Based Radar Signal Processing

This document contains abstracts of various research articles focusing on advanced technologies such as millimeter-wave radar, transformers in autonomous driving, and deep learning applications in radar-based sensing and forecasting. Key themes include the evaluation of existing methodologies, the introduction of novel algorithms, and the exploration of challenges and future directions in these fields. The articles collectively contribute to the understanding and development of intelligent sensing, communication strategies, and environmental monitoring technologies.

Uploaded by

zh.assanova98
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Yash Soni, Malhaar Goswami, Nishit Prabhakar Shetty, Dhiraj,

Millimeter-wave radar for intelligent sensing: A comprehensive review of techniques,


applications, and challenges,
Computers and Electrical Engineering,
Volume 128, Part A,
2025,
110696,
ISSN 0045-7906,
[Link]
([Link]
Abstract: Millimeter-wave (mmWave) radar sensing has established itself as a robust
technology across diverse applications, such as automotive, healthcare, security, and smart
homes. Its exceptional capacity to function effectively in varying environmental conditions,
detect concealed objects, sense physiological signals, and facilitate precise target detection
positions it as a pivotal enabler for next-generation sensing solutions. The survey employs
bibliometric analysis to critically evaluate the existing literature surrounding mmWave radar,
highlighting key research trends, notable publications, and the challenges faced within the
field. This work presents a comprehensive examination of mmWave radar-based sensing,
detailing its fundamental operating principles, signal processing methodologies,
advancements in hardware, and the latest developments in machine learning applications. It
also addreses the key challenges in signal processing, including resolution enhancement,
environmental adaptability, and data fusion with complementary sensors such as LiDAR and
cameras. Furthermore, explored the potential of deep learning techniques to enhance target
classification, activity recognition, gesture identification, and healthcare applications while
addressing concerns related to accuracy and precision. This survey also sheds light on
emerging trends by assessing the strengths, limitations, and prospects of mmWave radar
technology. This review aims to provide insightful guidance for researchers and practitioners
committed to advancing radar-based sensing and its real-world implementations.
Keywords: mmWave radar; FMCW radar; Wireless sensing; Machine learning; Deep learning;
Application taxonomy

Fulin Chu, Haoyu Li, Lili Xie, Jingyuan Zhao,


A survey of transformer architectures for autonomous driving,
Expert Systems with Applications,
Volume 299, Part D,
2026,
130338,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Transformers have emerged as a foundational paradigm in autonomous driving,
enabling high-capacity modeling of complex, multimodal, and dynamic environments. Their
self-attention mechanisms, scalability, and sequence modeling capabilities support spatial–
temporal reasoning and long-range dependency capture. Although increasingly adopted in
core modules—such as perception, trajectory prediction, decision-making, and anomaly
detection—a system-level survey of their architectural evolution and deployment challenges
remains lacking. This paper presents a structured survey of Transformer models in
autonomous driving, introducing a task-oriented taxonomy that spans object detection,
sensor fusion, trajectory forecasting, motion planning, and intent prediction. We compare
Transformer architectures with traditional deep learning models (e.g., CNNs, RNNs),
highlighting advantages in global context modeling, multimodal alignment, and unified
representations across heterogeneous inputs (camera, LiDAR, radar, HD maps). Beyond
current applications, this work examines large-scale, end-to-end Transformer systems and
their potential as foundation models for autonomous driving. We analyze design patterns
involving chain-of-thought reasoning, neuro-symbolic integration, federated learning, and
privacy-preserving edge deployment. Case studies from industry leaders (e.g., Tesla, Baidu,
NVIDIA, Aurora) illustrate practical trade-offs and architectural adaptations. Despite
progress, challenges remain in achieving real-time efficiency, robustness in open-world
scenarios, interpretability in sequential decision-making, and integration with cost-sensitive
sensors. We identify research gaps and propose directions in scalable Transformer design,
explainable AI, and policy-aware planning. This survey aims to guide researchers and
practitioners at the intersection of AI and intelligent transportation, supporting the
development of interpretable, efficient, and generalizable Transformer-based autonomous
driving systems.
Keywords: Transformer; Autonomous driving; Perception; Decision-making; Multimodal

Yiming Zhang, Jiaqi Li, Zheng Tong, Weiguang Zhang, Xiyuan Shen,
A direction-aware and expert-inspired network for internal crack size detection using on-site
ground penetrating radar data,
Engineering Applications of Artificial Intelligence,
Volume 165, Part A,
2026,
113414,
ISSN 0952-1976,
[Link]
([Link]
Abstract: The size of internal cracks is a key basis for determining maintenance measures.
Existing methods primarily utilize Ground Penetrating Radar (GPR) signals or images to
detect crack size, but still face two challenges: limited robustness in interpreting on-site data
based on GPR signals and the inability to directly characterize the crack size based on raw
GPR image features. To address these issues, this study has proposed a novel internal crack
size detection network, which was trained by a dataset with on-site GPR B-Scans and
interpreted crack size labels. In the proposed model, a deformable Cross Stage Partial (CSP)
block is first used to extract the irregular hyperbolic features of crack reflected waves from
B-Scans. Then, a directional fusion attention module is designed to construct direction-
aware channel attention and generate spatial interaction weights. Finally, a bipartite graph
matching detection head is proposed to emulate the expert behavior to analyze crack
reflected waves from a global B-Scan perspective, outputting trapezoidal-sized boxes to
detect internal crack size. The experimental results demonstrate that the proposed model
exceeds other state-of-the-art models on the tasks thanks to the channel-spatial weight
aggregation and global output strategy via bipartite graph matching. Additionally, the model
exhibits good stability across various antenna frequencies and pavement structures. The on-
site testing indicates that the predicted crack sizes sufficiently meet engineering
requirements in most scenarios, though challenges remain in detecting the bottom width of
small and water-saturated cracks.
Keywords: Asphalt pavement; Non-destructive testing; Ground penetrating radar; Internal
crack; Size detection

Long He, Kun Zheng, Huihua Ruan, Shuo Yang, Jinbiao Zhang, Cong Luo, Siyu Tang, Yunlei Yi,
Yugang Tian, Jianmei Cheng,
A spatiotemporal mixed-enhanced generative adversarial network for radar-based
precipitation nowcasting,
Computers & Geosciences,
Volume 200,
2025,
105919,
ISSN 0098-3004,
[Link]
([Link]
Abstract: Skillful precipitation nowcasting with high resolution and detailed information
holds promise for providing reliable alerts about severe weather events to society. Radar
echo extrapolation is an essential method for precipitation nowcasting, but traditional
methods struggle to capture rapidly changing regions. Deep learning (DL)-based methods
exhibit superior performance. However, existing DL-based methods face challenges such as
low accuracy, particularly in producing clear forecasts over longer lead times and accurately
forecasting moderate to heavy rainfall events. To address these challenges, we developed a
novel radar-based precipitation nowcasting model, STMixGAN, which can be described as a
nonlinear proximity forecasting model. This model effectively aggregates global-to-local
information and imposes constraints to represent the complex evolution of rainfall
efficiently. Consequently, STMixGAN produces realistic and spatiotemporally consistent
predictions. Using radar observations from South China, STMixGAN successfully forecasted
radar maps for the next 1 h using 24 min of input data. Two traditional methods (Persistence
and Optical flow) and five DL-based methods (ConvLSTM, Rainformer, IAM4VP, REMNet, and
GAN-argcPredNet) were employed as benchmarks to validate STMixGAN’s forecasting
capabilities. The experimental results demonstrate STMixGAN’s superior performance and
provide valuable insights for enhancing heavy rainfall forecasting.
Keywords: Precipitation nowcasting; Spatiotemporal mixed enhancement; Generative
adversarial networks; Self-attention

Wenju Zhao, Shijie Xing, Futao Ni, Yongding Tian, Qiang Liu,
FMCW radar-based high-precision range estimation with generalized eigenvalue
decomposition algorithm,
Measurement,
Volume 255,
2025,
117957,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Frequency-modulated Continuous Wave (FMCW) radar has been widely applied in
defense systems, automotive collision avoidance, intelligent traffic monitoring, and precision
level measurement due to its excellent long-range detection capabilities and improved range
resolution. However, traditional FMCW radar techniques are inherently prone to spectral
leakage artifacts and picket fence effects caused by non-integer period sampling and limited
frequency resolution of discrete Fourier transform (DFT) processing, which seriously
compromise ranging accuracy. To overcome these challenges, this paper proposes a novel
high-precision range estimation algorithm based on generalized eigenvalue decomposition
(GEVD). The key contributions of this study are: (1) the derivation of the analytical
relationship between generalized eigenvalues and beat frequencies through formulating a
generalized eigenvalue equation that includes both the beat frequency signal and its first-
order derivative; and (2) the effective reduction of discretization errors and noise
interference using frequency-shifting techniques combined with singular value
decomposition (SVD)-based signal enhancement. A thorough parametric analysis has been
performed to evaluate the impact of sampling frequency, matrix dimension, signal-to-noise
ratio, and the number of targets on range precision. Extensive numerical simulations and
controlled laboratory experiments validate the theoretical framework and operational
effectiveness of the proposed methodology. Comparative results demonstrate that the
GEVD-based method achieves superior resolution compared to conventional techniques,
even in noisy environments. Field validation using single-target measurement trials confirms
exceptional measurement stability, with empirical data showing maximum range deviation
within 0.2 mm at sampling frequencies exceeding 400 kHz.
Keywords: Frequency-modulated continuous wave radar; Range estimation; Generalized
eigenvalue decomposition; Civil engineering

Wenjun Hou, Hu Jin, Chuang Peng, Li Jiang,


A cognitive communication jamming strategy based on Transformer and Deep
Reinforcement Learning,
Computers and Electrical Engineering,
Volume 120, Part A,
2024,
109610,
ISSN 0045-7906,
[Link]
([Link]
Abstract: The advent of sophisticated communication technologies, such as cognitive radio
and anti-jamming techniques, has significantly elevated the challenge of disrupting enemy
communications. Nevertheless, the inherent openness of wireless communications remains
a vulnerability that can be exploited to interfere with them. Some contemporary
Reinforcement Learning (RL)-based jamming strategies examine methods for rapidly
identifying the optimal jamming strategy for a specific modulated signal. However, such
algorithms lack the flexibility and responsiveness required to effectively counter the enemy’s
evolving communication strategies. To address this issue, we propose a Transformer and
Deep Reinforcement Learning (DRL)-based jamming strategy that can be trained to identify
jamming methods for multiple digital and analog signals. In particular, the Transformer
Encoder is employed as a network for DRL to process the state information pertaining to the
enemy communication. Subsequently, the decision module of the Double Deep Q Network
(DDQN) is utilized to select the jamming action based on the processed information.
Furthermore, we have devised a reward function and constructed an invalid jamming list,
with the objective of selecting an action that requires low power consumption and enhances
the convergence speed of the algorithm. The experimental results demonstrate that the
algorithm proposed in this paper exhibits notable performance advantages in comparison to
other networks and DRL algorithms.
Keywords: Cognitive interference; Deep Reinforcement Learning; Transformer model;
Double Deep Q Network

Peifeng Ma, Zherong Wu, Zhengjia Zhang, Francis T.K. Au,


SAR-Transformer-based decomposition and geophysical interpretation of InSAR time-series
deformations for the Hong Kong-Zhuhai-Macao Bridge,
Remote Sensing of Environment,
Volume 302,
2024,
113962,
ISSN 0034-4257,
[Link]
([Link]
Abstract: Time-series interferometric synthetic aperture radar (InSAR) provides a unique tool
for measuring large-scale and long-term land surface deformation. Under the assumption of
a single linear deformation model in conventional InSAR, it is difficult to quantify and
interpret the impacts of multiple environmental factors that presumably induce nonlinear
deformations. In this paper, we propose a SAR-Transformer method to decompose InSAR
time-series signals into various physics-related components and apply the method to
evaluate the deformation of the world's longest cross-sea bridge, the Hong Kong-Zhuhai-
Macao Bridge (HZMB). We first developed an improved bridge geometry-based InSAR
network to monitor the deformation of the HZMB using Sentinel-1 and COSMO-SkyMed
images from 2019 to 2022, which were validated using the leveling and GPS data. The SAR-
Transformer model was trained using synthetic InSAR time-series samples and applied to
decompose the monitored InSAR measurements. Compared with that of conventional curve-
fitting and seasonal-trend decomposition using LOESS, SAR-Transformer reduced the mean
absolute error at least by 58.32% and mean absolute percentage error at least by 8.84% for
time-series signal reconstruction. We evaluated the decomposed patterns according to the
geotechnical, meteorological, and marine processes, and found that: 1) Seasonal thermal
expansion owing to temperature changes was significant in all parts of the bridge, and
deflection due to concrete shrinkage and creep was observed on cable-stayed bridges. 2)
The artificial islands experienced evident ground subsidence with a decelerating trend. In
particular, the newly adopted non-dredged reclamation method resulted in a lower
decelerated settlement than that of fully-dredged reclamation areas. 3) The seawall showed
linear horizontal movement from the outward stretching of the reclaimed soil consolidation
and periodic displacement related to sea tidal loading. Furthermore, typhoons and coastal
earthquakes had limited effects on the permanent movement of the bridge. These results
improve the understanding of the interactions between artificial super-infrastructures and
environmental factors, and provide valuable guidelines for the maintenance and
management of the HZMB.
Keywords: Cross-sea bridge; InSAR time-series signals; Reclamation; SAR-transformer; Tidal
loading

Wei Quan, Wenjing Cheng, Yike Yang, Haiquan Zhao, Zhaoyu Chen, Yunfan Luo,
A signal fingerprint feature extraction method based on decomposition and fusion for radar
emitter individual identification,
Digital Signal Processing,
Volume 164,
2025,
105257,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Radar emitter individual identification is one of the key technologies of modern
electronic countermeasure reconnaissance and electronic intelligence. With the
advancement of radar technology and the increasingly complex electromagnetic
environment, existing methods for identifying emitter are gradually becoming unable to
meet the performance requirements of modern radar individual identification. Aiming at
improving the adaptability of feature extraction for non-cooperative radar emitter signals
and the robustness of individual identification in the complex modern electronic warfare
environment, a signal fingerprint feature extraction method based on decomposition and
fusion is proposed. It firstly integrates signal decomposition and scattering convolution
networks (SCN) to adaptively extract the multi-scale intra-pulse feature of the signal, while
removing the potential noise of the redundant component by energy proportion. And then a
deep feature fusion model based on multi-head self-attention and residual connection is
proposed to fuse the multi-scale features and the time domain features to further extract
signal fingerprint of radar emitter. Experimental results based on the real radar emitter
signals demonstrate that the identification method proposed in this paper can more
effectively extract signal fingerprint features and the identification accuracy reaches 96.45%,
which outperforms other existing identification methods.
Keywords: Radar emitter individual identification; Signal fingerprint feature; Signal
decomposition; Scattering convolution networks (SCN); Fusion

Haoyuan Ding, Tingsong Zhang, Zhangting Wang, Guangran Bai, Liujun Han, Yujia Dai, Ziyuan
Liu,
Classification and quantification of sodium metabisulfite in goji berry powder: Applications
of hyperspectral technology and transformer-based hybrid models,
Spectrochimica Acta Part A: Molecular and Biomolecular Spectroscopy,
Volume 349,
2026,
127379,
ISSN 1386-1425,
[Link]
([Link]
Abstract: Sodium metabisulfite (Na2S2O5) is widely used as an antioxidant and preservative
in food products, but excessive residues pose health risks and are strictly regulated. Rapid
and non-destructive detection methods are therefore essential, particularly for goji berry
powder where sulfite addition is common during processing. In this study, near-infrared
hyperspectral imaging (HSI, 900–1700 nm) was employed in combination with advanced
deep learning models to classify and quantify sodium metabisulfite concentrations. A total
of 360 samples (nine concentration levels including control, 40 replicates each) were
prepared, with 70 % allocated for training and 30 % for testing. For classification, a hybrid
ResLocalformer model that integrates local attention and residual paths within a
Transformer framework achieved an accuracy of 97.22 %. For regression, the Resformer
model, combining Transformer's global attention with dense layers, yielded an R2 of 0.9945
and RMSE of 0.0343. Spectral preprocessing using first-derivative Gaussian smoothing
(1DER-GS) significantly enhanced feature quality, improving overall model performance by
more than 40 %. Compared with conventional approaches such as CNN and LSTM, the
proposed models demonstrated superior robustness and predictive accuracy. These results
indicate that HSI combined with transformer-based hybrid models provides an effective and
non-destructive approach for monitoring sodium metabisulfite in goji berry powder, with
strong potential for extension to the detection of other food additives and matrices.
Keywords: Hyperspectral imaging; Deep learning; Classification; Quantification; Goji berry;
Transformer-Based Hybrid Models.

Shuai Xu, Lutao Liu, Zhongkai Zhao,


Unsupervised recognition of radar signals combining multi-block TFR with subspace
clustering,
Digital Signal Processing,
Volume 151,
2024,
104552,
ISSN 1051-2004,
[Link]
([Link]
Abstract: In the realm of radar systems, the proliferation of new modulation techniques
introduces an increased level of complexity in the identification of both existing and
potentially novel modulation schemes. Conventional supervised recognition methodologies,
which rely heavily on labeled datasets, exhibit limitations in distinguishing amongst a
multitude of unlabeled signals. This paper introduces a novel framework that synergizes
Subspace Clustering (SC) with Multi-Block Time-Frequency Representations (TFRs, denoted
as mT), referred to as SC-mT. This framework leverages subspace clustering for the
unsupervised categorization of radar modulations and employs an innovative multi-block
strategy to augment classification precision. The process commences with the generation of
TFR datasets for radar signals, followed by the construction of two distinct multi-block
models, Model-A and Model-B. These models segment the radar TFRs into overlapping
multi-block sets, utilizing the concept of random receptive fields. The subspace clustering
algorithm is then applied to each block within the sub-TFR sets to procure an affinity matrix.
The culmination of this process involves the aggregation of the affinity matrices from all
blocks, facilitating the derivation of classification outcomes via spectral clustering. Empirical
analyses affirm that the sequential integration of the SC-mT algorithm surpasses the
classification efficacy of various conventional and cutting-edge algorithms. Notably, this
algorithm attains a classification accuracy surpassing 90% for ten unlabeled signals, even in
scenarios where the signal-to-noise ratio (SNR) is as low as -4 dB.
Keywords: Radar modulation; Unsupervised classification; Multi-block; Subspace clustering

Siyu Chen, Xiaoyan Zhang, Hongjun Xue, Xiang Fang, Xueren Li,
Knowledge-guided graph transformer for gaze-based pilot operation recognition in dynamic
flight operations,
Aerospace Science and Technology,
Volume 171,
2026,
111610,
ISSN 1270-9638,
[Link]
([Link]
Abstract: Real-time and accurate recognition of pilots’ operational intent in dynamic flight
tasks is critical for enhancing aviation safety, preventing human error, enabling intelligent
decision support, and supporting online human-reliability assessment. Eye-tracking data, an
objective physiological signal obtainable in real time in the cockpit, provides crucial evidence
for inferring intent. However, the volatility of time-series gaze data and its tight, nonlinear
coupling with complex operational responses limit the accuracy of approaches that rely on
gaze features alone. Methods based on physical parameters capture action outcomes rather
than intent, while purely data-driven models often generalize poorly and offer limited
interpretability due to the absence of task-logic guidance. This study introduces a
Knowledge-enhanced Graph Transformer Network (KGTN) that integrates structured domain
knowledge as a computable guidance signal with dynamic eye-movement behavior, and
explicitly models their nonlinear interactions. A task knowledge graph grounded in a
cognitive-activity taxonomy is constructed and encoded by a Relational Graph Neural
Network, alongside a multi-scale Patch-Transformer and a Hierarchical Task-Context Aware
Operation Recognition Module (HTCA-ORM) for knowledge-guided fusion and precise
operation recognition. On a high-fidelity flight-simulation dataset, KGTN attains 87.63%
accuracy and 87.48% weighted F1, outperforming a range of baselines. Ablation and
interpretability analyses indicate strong potential for high-precision, reliable aviation human-
factors analytics.
Keywords: Human-machine systems; Intelligent cockpit; Aviation human factors; Eye
tracking; Cognitive activity modeling

Asim Saleem, Guoyun Lv, Safa Hussein Mohammed,


Deep Learning-Based Radar Fingerprinting for Open-Set Generalization Using Dynamic
Thresholding and Embedding Rejection,
Knowledge-Based Systems,
Volume 333,
2026,
115047,
ISSN 0950-7051,
[Link]
([Link]
Abstract: This paper presents a comprehensive framework for radar-specific emitter
identification (SEI), starting with the simulation of a large-scale radar signal dataset designed
to mimic real hardware impairments. By incorporating diverse distortions–such as phase
noise, frequency jitter, amplitude nonlinearity, and multipath reflections–for multiple radar
types and signal-to-noise ratio (SNR) conditions, we generate a realistic and challenging
dataset, SimRF-14, suitable for learning-based signal analysis. We utilized this dataset to
develop RAFNet, a hybrid deep learning model specifically designed for closed-set and open-
set radar emitter classification. The proposed architecture combines convolutional,
recurrent, and attention-based components to capture spatial, temporal, and contextual
features from normalized I/Q waveforms. For open-set recognition, we integrate the
OpenMax algorithm enhanced with Extreme Value Theory (EVT), where class-wise Weibull
modeling of embedding distances enables outlier detection. In addition, an SNR-adaptive
thresholding mechanism improves open-set reliability under varying noise conditions. The
proposed method achieved 97.43% unknown rejection at -20 dB SNR and maintained low
false positives (<4%) at high SNRs, validating its effectiveness and reliability for practical SEI
scenarios.
Keywords: Radar Specific Emitter Identification; Open-set Recognition; RF Fingerprinting;
SNR-Adaptive Classification; Deep Learning for SEI

Tiantian Wang, Nan Yan, Chaosan Yang, Zeliang An, Gongjing Zhang, Yuqing Xu,
Electromagnetic signal recognition using multimodal tri-branch semantic fusion network in
the UAV-assist integrated sensing and communication systems,
Digital Signal Processing,
Volume 171,
2026,
105820,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Driven by the proliferation of integrated sensing and communication (ISAC)
systems, the accurate recognition of unauthorized unmanned aerial vehicle (UAV) signals in
dynamic electromagnetic environments has emerged as a critical challenge for spectrum
security and cognitive radio applications. Conventional automatic modulation recognition
(AMR) frameworks suffer from significant performance degradation in low signal-to-noise
ratio (SNR) regimes and exhibit limited adaptability to resource-constrained edge computing
platforms. To address these limitations, we propose a novel Multimodal Tri-branch Fusion
Network (MTF-Net) architecture that synergistically integrates time-frequency analysis with
statistical feature learning. The framework systematically processes binarized time-
frequency images (B-TFIs) and higher-order cumulant vectors through three collaboratively
operating branches: (1) A primary temporal feature extractor employing dilated convolution-
residual blocks (DCRBlocks) with hierarchical dilatation factors, incorporating channel
attention mechanisms to dynamically emphasize discriminative temporal patterns; (2) Dual
auxiliary branches based on Edge-Transformer modules (ETFormers), which achieve efficient
spatial-structural learning through depthwise separable convolutions (DSC) while capturing
long-range spectral dependencies via additive attention mechanisms with linear complexity;
(3) A hierarchical fusion module implementing cross-branch feature recalibration through
learnable parameter matrices. Extensive Monte Carlo experiments demonstrate that our
MTF-Net significantly outperforms traditional methods in recognition accuracy for radar and
communication signals under low SNR conditions, establishing a new benchmark for
lightweight AMR solutions in ISAC systems.
Keywords: Multi-modal feature fusion; Unmanned aerial vehicle(UAV); Integrated sensing
and communication (ISAC); Lightweight neural network; Transformer

Sidra Ghayour Bhatti, Imtiaz Ahmad Taj, Mohsin Ullah, Aamer Iqbal Bhatti,
Transformer-based models for intrapulse modulation recognition of radar waveforms,
Engineering Applications of Artificial Intelligence,
Volume 136, Part B,
2024,
108989,
ISSN 0952-1976,
[Link]
([Link]
Abstract: The increasing prevalence of low probability of intercept (LPI) radars in electronic
warfare (EW) systems highlights the need to effectively recognize phase-coded radar
waveforms intercepted at radar warning receivers (RWRs) from various threat emitters. The
complexities of the electromagnetic (EM) spectrum necessitate the implementation of an
automatic modulation recognition system (AMRS) within the RWR. However, a major
challenge is accurately identifying phase-coded waveforms with high accuracy at low signal-
to-noise ratios (SNRs). This research addresses the challenge by exploring three artificial
intelligence (AI)-driven AMRS architectures for identifying phase-coded waveforms using
short-time Fourier transform (STFT): vision transformer (ViT), vicinity vision transformer
(VViT), and deep convolutional neural network (DCNN). Unlike recent methods focusing on
amplitude spectra, our research delves into the phase spectra for the feature extraction of
phase-coded waveforms. We leverage phase-based features extracted from intercepted
phase-coded waveforms to classify six types of phase-coded signals using these AMRS
architectures across SNR levels ranging from −16 dB to 8 dB. The simulation experiments
show that these methods are effective at an SNR of −16 dB, with VViT and ViT achieving
recognition accuracies of 93% and 92.7%, respectively. Both outperform the DCNN, which
achieves an RA of 89% at the same SNR. This approach promises to enhance situational
awareness and decision-making in EW operations by improving phase-coded radar
waveform recognition and enabling appropriate countermeasure deployment.
Keywords: Automatic modulation recognition system; Feature extraction; Low probability of
intercept; Short time Fourier transform

Shuai Guo, Ting Chen, Penghui Wang, Jun Ding, Junkun Yan, Hongwei Liu,
Knowledge embedding fusion based on language model for enhanced radar target
recognition,
Signal Processing,
Volume 238,
2026,
110199,
ISSN 0165-1684,
[Link]
([Link]
Abstract: Traditional radar target recognition methods typically model only single echoes,
neglecting the crucial information that domain knowledge can provide for understanding
data. In this paper, we propose a knowledge embedding fusion (KEF) method for enhanced
high-resolution range profile (HRRP) recognition, which utilizes the target state descriptions
available during radar detection. KEF leverages a language model (LM) to integrate textual
knowledge with echo features for fusion recognition. It consists of three components: HRRP
feature extraction, measurement-based knowledge construction, and knowledge embedding
fusion module. First, we perform feature extraction on the HRRP to obtain echo tokens.
Next, in the knowledge construction module, the measurement statuses are standardized to
a natural language format, and the LM is utilized to extract semantic information, resulting in
text tokens. Finally, in the knowledge embedding fusion module, a cross-attention HRRP-text
fusion strategy is employed to facilitate interaction between echo tokens and textual tokens.
We also design a combination of HRRP-text matching loss and fusion classification loss to
guide model training. Experiments are conducted on a real measured dataset, and the
results indicate that KEF effectively enhances recognition performance across multiple
scenarios compared with approaches that only utilize echoes.
Keywords: High-resolution range profile (HRRP); Knowledge embedding fusion; Language
model (LM); Radar target recognition

Mingyue Lu, Menglong Wang, Qian Zhang, Manzhu Yu, Caifen He, Yadong Zhang, Yuchen Li,
A vision transformer for lightning intensity estimation using 3D weather radar,
Science of The Total Environment,
Volume 853,
2022,
158496,
ISSN 0048-9697,
[Link]
([Link]
Abstract: Lightning has strong destructive powers; its blast wave, high temperature, and high
voltage can pose a great threat to human production, life, and personal safety. The
destructive power of high-intensity lightning is much greater than that of low-intensity
lightning. The estimation of lightning intensity can provide an important reference for
determining the lightning protection level and lightning disaster risk assessment. Lightning is
a type of small-scale severe convective weather phenomenon. Weather radar is one of the
best monitoring systems that can frequently sample the detailed three-dimensional (3D)
structures of convective storms, with a small spatial scale and short lifetime at high temporal
and spatial resolutions. Therefore, it is possible to extract the 3D spatial feature strongly
correlated with lightning from 3D weather radar for estimating lightning intensity. This paper
proposes a Vision Transformer model for lightning intensity estimation that can
automatically estimate lightning intensity from 3D weather radar data. In an experiment, we
transferred the task of estimating lightning intensity into a multicategory classification task.
A framework was designed to produce lightning feature samples for model input from 3D
weather radar and lightning location data. Then, the Synthetic Minority Over-Sampling
Technique (SMOTE) algorithm was used to balance and optimize the sample distribution.
Finally, samples were input into the proposed lightning intensity estimation model based on
Vision Transformer for training and evaluation. Experimental results show that the proposed
model based on Vision Transformers performs well with lightning intensity estimation.
Keywords: Lightning intensity estimation; 3D weather radar; Vision transformer; SMOTE;
Multicategory classification
Shenghua Lv, Xiaowei Zhang, Xuan Zhao, Meng Li, Jianghao Zhang, Chen Lin, Jian Wen,
Rapid and accurate assessment of filed scale soil moisture using ground-penetrating radar
deep learning-based inversion,
Measurement,
2026,
120594,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Accurate quantification of soil moisture content (SMC) is essential for sustaining
plant growth and maintaining ecosystem stability. Current SMC monitoring approaches are
subject to several constraints: satellite-based remote sensing frequently suffers from
inadequate spatial resolution for large-scale precision, while point-scale techniques are
incapable of effectively capturing soil moisture spatiotemporal variations and are unsuitable
for large-area monitoring. Ground Penetrating Radar (GPR), as an efficient and non-
destructive subsurface detection technique, has been widely employed for estimating soil
water content. However, existing methods based on full-waveform inversion exhibit strong
dependency and involve computationally expensive processes, resulting in insufficient
efficiency when handling large-scale GPR data. To overcome these challenges, this study
proposes a GPR deep learning-based inversion framework for rapid and precise SMC
estimation. A 900 MHz GPR system was deployed to survey an experimental site equipped
with pre-installed moisture sensors. The results demonstrated strong agreement between
GPR-derived SMC values and measurements obtained via Time Domain Reflectometry (TDR).
Furthermore, continuous high-temporal-resolution data collection verified the capability of
the method to characterize the spatiotemporal dynamics of SMC. Notably, for regional soil
moisture content (RSMC) estimation, the proposed method achieved a substantially lower
error (0.0019 m3) compared to conventional point-based measurements (0.0146 m3) and
stratified estimation (0.0122 m3). This methodology provides a robust technical foundation
for accurate field-scale SMC monitoring and exhibits significant potential for use in ecological
surveillance, precision agriculture irrigation, and sustainable water resource management.
Keywords: Ground-penetrating radar; Non-destructive testing; Soil moisture content; Deep
learning-based inversion; Spatio-temporal evolution

Zhuoning Hao, Shaojuan Luo, Wei Meng, Huapan Xiao, Heng Wu, Chunhua He,
SAR ship detection in complex coastal areas: A multi-scale residual fusion transformer with
spatial-channel attention for high detection accuracy,
Measurement,
Volume 264,
2026,
120257,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Synthetic Aperture Radar (SAR) provides critical all-weather measurement
capabilities for microwave imaging in remote sensing. Accurate SAR ship detection is pivotal
for measurement-driven applications such as maritime surveillance, where reliable
identification and localization of ships are essential. Current SAR target detection methods
enable accurate ship detection in open waters but face challenges in detecting ships near
complex coastlines and small targets. To address these problems, we propose a SAR ship
detection network based on a multi-scale query refinement transformer and spatial-channel
attention fusion module, MRFSA-Net. Specifically, a lightweight backbone network with a
SAR enhancement module is designed to reduce model complexity while preserving feature
extraction capabilities, simultaneously enhancing multi-scale feature extraction and feature
fusion. A spatial-channel attention fusion module is developed to strengthen contextual
awareness and suppress background interference. Cross-scale interaction branches are
employed to achieve effective fusion of information across different receptive fields.
Additionally, a transformer with biased attention and query optimization mechanisms is
constructed to enhance the robustness of multi-scale object detection. Quantitative and
qualitative experimental results show that MRFSA-Net achieves competitive performance
compared to the previous state-of-the-art object detection methods, significantly improving
detection accuracy in complex coastal environments.
Keywords: Synthetic Aperture Radar; Deep Learning; Object Detection

Kangle Song, Jingbin Li, Yang Li, Jing Nie, Yuntao Sun, Wujun Zhang, Xiaojie Hou, Qiang
Wang, Pengxiang Song,
Long-distance soil moisture monitoring via Helmholtz resonator–enhanced acoustic
transmission and Swin-Transformer modeling,
Computers and Electronics in Agriculture,
Volume 242,
2026,
111344,
ISSN 0168-1699,
[Link]
([Link]
Abstract: To address the severe energy attenuation of acoustic waves during propagation
through soil, which restricts signal reception and modeling in large scale soil moisture
monitoring, this study designs a resonance enhancement structure to overcome the short
detection range of conventional acoustic methods and proposes a large-scale transmission
method for soil moisture detection based on a Helmholtz resonator. The response
characteristics and consistency of the Helmholtz resonator under identical excitation
conditions were first evaluated. The enhancement of acoustic signal reception and the
extension of detection capability were then verified by comparing system responses over
various propagation distances. Furthermore, the influence of excitation periodicity on
resonance behavior was analyzed, leading to optimized excitation parameters for improved
system performance. A mapping model was subsequently constructed between acoustic
time–frequency spectrograms and soil water content through feature extraction and model
development, with a visualization mechanism introduced to interpret model decision
making processes. System validation was conducted under real field conditions.
Experimental results demonstrated that the Helmholtz resonator exhibited highly consistent
response characteristics and strong frequency selectivity. The effective detection range of
the system was extended from 4 m to 60 m, significantly enhancing the reception of acoustic
signals in large-scale soil environments. In the test dataset, the Swin-Transformer–based
regression model achieved a mean absolute error (MAE) of 0.207 %, root mean square error
(RMSE) of 0.244 %, and coefficient of determination (R2) of 0.992. In field trials, the
corresponding metrics were 1.046 %, 0.851 %, and 0.735, respectively, confirming the
effectiveness of the feature extraction and modeling approach. This study demonstrates that
the proposed Helmholtz resonator–based method enables effective large-scale soil moisture
detection and offers a novel, low cost, transmission-based solution for wide-area soil
moisture monitoring and water management.
Keywords: Soil moisture monitoring; Helmholtz resonator; Acoustic wave; EMD; Swin-
transformer

Xiaosong Tang, Feng Yang, Xu Qiao, Jialin Liu, Haitao Zuo, Liang Gao, Jianshe Zhao, Suping
Peng,
GPR-HIDiff: A diffusion-based model for horizontal interference suppression in urban
underground detection radar profiles,
Underground Space,
Volume 26,
2026,
Pages 458-478,
ISSN 2467-9674,
[Link]
([Link]
Abstract: Automated subsurface utility detection systems in construction rely heavily on the
quality of ground-penetrating radar (GPR) profiles, which are often degraded by high-
amplitude horizontal interference. Existing low-rank decomposition methods lack the
intelligence and flexibility required for multi-site data processing and involve labor-intensive
parameter tuning, impeding their integration into intelligent construction workflows. To
address these challenges, this paper proposes a horizontal interference suppression
algorithm based on a diffusion model, termed GPR-HIDiff. The proposed model replaces
conventional sequential convolutional operators with ResBlocks throughout the encoder,
intermediate layer, and decoder of the UNet architecture, enhancing training stability.
Lightweight agent attention modules are embedded between ResBlocks at each level to
improve global information modeling capability. A spatial attention mechanism is deployed
between the encoder and decoder to achieve adaptive spatial feature optimization.
Furthermore, the forward diffusion phase adopts a cosθ schedule-based strategy to ensure a
smooth temporal variation of noise variance. A standardized dataset comprising real-world
measured samples and finite difference time domain simulation samples of urban road
models has also been constructed. The effectiveness of the hybrid dataset, the introduced
modules, the robustness analysis, and the cosθ schedule is validated through training with
single/mixed datasets, ablation studies, evaluation of metric variations before and after the
introduction of different noise levels, and comparative experiments with constant, linear,
and cosθ schedules. Experimental results demonstrate that GPR-HIDiff significantly
outperforms both traditional methods and state-of-the-art deep learning models on both
simulated and real-world test samples. It effectively suppresses horizontal artifacts,
preserves target hyperbolic contours, and avoids excessive reduction of target scattering,
showcasing its exceptional performance. This method provides a powerful algorithmic
foundation for high-resolution GPR imaging and target detection.
Keywords: Ground-penetrating radar; Horizontal interference; Diffusion model; Agent
attention module; Spatial attention; Hybrid dataset
Zhiyan Lin, Minming Gu, Keyu Pan, Wei-Ping Zhu,
Adaptive temporal convolutional network with multi-head EMA-gated attention for
continuous radar-based human activity recognition,
Biomedical Signal Processing and Control,
Volume 117,
2026,
109667,
ISSN 1746-8094,
[Link]
([Link]
Abstract: Continuous human activity recognition (HAR) using radar signals offers strong
potential for privacy-preserving clinical health monitoring. However, its performance is
limited by challenges such as multi-scale temporal variation, signal noise, and unstable
activity transitions. To address these issues, this study introduces a radar-based HAR
framework with three tailored components. First, an adaptive temporal convolutional
network (ATCN) uses learnable dilation rates and sampling offsets to flexibly capture both
abrupt and periodic motion patterns over time. Second, an exponential moving average
(EMA)-gated attention (EDGA) module integrates linear attention with exponential moving
average smoothing through a dynamic gating mechanism, effectively suppressing noise
while preserving temporal continuity. Third, an attention-guided multi-stage refinement
(AMSR) module refines coarse predictions using global attention-driven residual corrections,
thereby reducing segmentation noise and improving boundary precision. Experiments on a
77 GHz frequency modulated continuous wave (FMCW) radar dataset show that the
proposed model achieves 96.09% accuracy, demonstrating its strong potential for
continuous and unobtrusive activity monitoring in healthcare applications.
Keywords: ATCN; EDGA; AMSR; Continuous HAR; FMCW radar

Yiming Xiao, Ali Mostafavi,


DamageCAT: A deep learning transformer framework for typology-based post-disaster
building damage categorization,
International Journal of Disaster Risk Reduction,
Volume 128,
2025,
105704,
ISSN 2212-4209,
[Link]
([Link]
Abstract: Rapid, accurate, and descriptive building damage assessment is critical for directing
post-disaster resources, yet current automated methods typically provide only binary
(damaged/ undamaged) or ordinal severity scales. This paper introduces DamageCAT, a
framework that advances damage assessment through typology-based categorical
classifications. We contribute: (1) the BD-TypoSAT dataset containing satellite image triplets
from Hurricane Ida with four damage categories – partial roof damage, total roof damage,
partial structural collapse, and total structural collapse – and (2) a hierarchical U-Net-based
transformer architecture for processing pre- and post-disaster image pairs. Our model
achieves 0.737 IoU and 0.846 F1-score overall, with cross-event evaluation demonstrating
transferability across Hurricane Harvey, Florence, and Michael data. While performance
varies across damage categories due to class imbalance, the framework shows that typology-
based classification can provide more actionable damage assessments than traditional
severity-based approaches, enabling targeted emergency response and resource allocation.
Keywords: Damage assessment; Satellite imagery; Transformers; Damage description

Lijie Yang, Yu Wang, Zhaohui Yang, Tongkai Xu, Lizhi Dang, Zhongyue Chen, Weipeng Mao,
mmPPT: Hierarchical-serialization-enhanced point transformer for mmWave pedestrian
reconstruction,
Information Fusion,
Volume 127, Part B,
2026,
103835,
ISSN 1566-2535,
[Link]
([Link]
Abstract: In autonomous driving, redundant perception is critical for robust environmental
understanding, especially under adverse conditions where traditional sensors like RGB-D fail.
This paper introduces a novel Point Transformer for mmWave-radar-based pedestrian 3D
reconstruction (mmPPT), addressing the challenges of sparse, noisy 4D mmWave radar point
clouds. mmPPT employs a novel Hierarchical Serialization Strategy that combines Hybrid
Local Serialization to preserve fine-grained body details via adaptive space-filling curves and
Anchor-based Global Serialization to model long-range structural relationships. Additionally,
it integrates temporal information and Doppler velocity features to leverage radar’s unique
capabilities, while a Bone Length Loss is adopted to enforce geometric constraints on joint
position prediction. Experiments on 4D radar dataset demonstrate that mmPPT outperforms
state-of-the-art methods, achieving the lowest total reconstruction error (12.76 cm with
multi-frame input) and superior robustness in harsh environments like smoke, occlusion, and
poor lighting scenarios where RGB-D sensors fail. Ablation studies validate the effectiveness
of the Hybrid Local Serialization Strategy and the Anchor-based Global Serialization Strategy,
highlighting mmPPT’s potential to enhance redundant perception in autonomous driving.
The source code will be available at [Link]
Keywords: Redundant perception; 4D mmWave radar; 3D reconstruction; Hierarchical
serialization strategy; Hybrid local serialization; Anchor-based global serialization; Temporal
and velocity embedding; Bone length loss

Chaofeng Huang, Xiaowo Xu, Fan Fan, Shunjun Wei, Xiaoling Zhang, Dongmei Liu, Min Gu,
A low-SNR-adaptive temporal network with smart mask attention for radar signal
modulation recognition,
Digital Signal Processing,
Volume 168, Part D,
2026,
105640,
ISSN 1051-2004,
[Link]
([Link]
Abstract: The automatic modulation recognition of radar signals is a key technology in
electronic warfare and communication systems. However, traditional handcrafted features
often struggle to achieve high recognition accuracy under low signal-to-noise ratio (SNR)
conditions. With the rapid development of artificial intelligence technologies, deep learning-
based approaches have emerged as a promising alternative for modulation recognition. In
this article, a low-SNR-adaptive network architecture is proposed, which integrates a
bidirectional temporal convolutional network (Bi-TCN) and dual-channel smart mask
attention (DSMA) modules. The DSMA adaptively highlights informative features and
suppresses noise through complementary attention masks, enhancing robustness in low-SNR
conditions. Experimental results demonstrate that the autocorrelation domain outperforms
both time and frequency domains, with recognition accuracy improvements of 13.33 % and
14.71 %, respectively. Compared to state-of-the-art models, the proposed network achieves
63 % accuracy at -20 dB and more than 99 % accuracy at -6 dB, significantly enhancing radar
signal modulation recognition.
Keywords: Modulation recognition; Deep learning; Radar signal analysis,

Jiaquan Wan, Junchao Wang, Wei Zhang, Hao Song, Congyi Nai, Fengchang Xue, Tao Yang,
Chunxiang Shi, Quan J. Wang, Baoxiang Pan,
RadarDiT: An advanced radar echo extrapolation model for three gorges reservoir area via
diffusion transformer,
Journal of Hydrology: Regional Studies,
Volume 61,
2025,
102703,
ISSN 2214-5818,
[Link]
([Link]
Abstract: Study region
The Three Gorges Reservoir Area (TGRA)
Study focus
TGRA faces increasing vulnerability to extreme precipitation events driven by complex
convective weather systems. Radar echo extrapolation—predicting future precipitation
patterns from current radar data—is essential for early warning systems but faces significant
challenges in this topographically complex region. While data-driven approaches have
advanced the field, current convolutional neural network-based diffusion models struggle
with the TGRA's dynamic meteorological conditions due to their reliance on translational
invariance, which often fails to capture rapid weather transitions in complex terrain.
New hydrogeological insights from the region
To address these limitations, we introduce RadarDiT, a Vision Transformer-based diffusion
model specifically engineered for radar extrapolation in the TGRA. First, we develop a five-
year radar dataset capturing diverse convective weather phenomena unique to this region.
Then, leveraging this dataset, RadarDiT employs multi-layer Vision Transformers that
effectively model global dependencies and complex spatial relationships, enabling accurate
prediction of convective cell evolution. Our model demonstrates superior performance in
maintaining strong echo and spatial coherence over longer forecast horizons. Quantitative
evaluations across multiple metrics and thresholds confirm RadarDiT's enhanced skill in
forecasting heavy precipitation events, with particular improvements in Critical Success
Index at higher radar echo values. This work establishes a foundation for more reliable
nowcasting systems in regions with complex terrain and dynamic weather patterns, directly
supporting enhanced disaster preparedness and response strategies.
Keywords: Radar Echo Extrapolation; Three Gorges Reservoir Area; Diffusion Model; Vision
Transformer; Nowcasting

Xiaole Han, Jintao Liu, Jian Ye, Zihe Wang, Pengfei Wu, Hai Yang,
Deep Learning-Based GPR interpretation of soil thickness in headwater hillslopes,
Geoderma,
Volume 462,
2025,
117530,
ISSN 0016-7061,
[Link]
([Link]
Abstract: Soil thickness strongly influences eco-hydrological and geomorphic processes, yet
conventional measurements such as auger drilling are invasive, labor-intensive, and
unsuitable for large-scale surveys. Ground-penetrating radar (GPR) provides a non-invasive
alternative, but its manual interpretation remains slow and prone to observer bias. To
address this challenge, we developed a fully automated framework that couples a hybrid
CNN-Transformer deep learning architecture with optimized signal filtering to predict soil
thickness directly from GPR profiles. The convolutional layers extract local waveform
features, while the attention mechanism captures long-range dependencies. Using field data
from a steep headwater hillslope (H1) in the Taihu Basin, China, we compared five filtering
strategies—median, Savitzky-Golay, Gaussian, moving average, and none—and found that
median filtering yielded the most accurate results (R2 up to 0.92, CCC of 0.96, RMSE near
10 cm). We further identified optimal filter window sizes (61–101 samples) and a training
duration threshold (≥500 epochs) that ensured stable and accurate predictions. Cross-site
validation on an independent hillslope (H2) without retraining showed that the pretrained
CNN-Transformer model achieved the highest R2 (0.80), CCC (0.89), and lowest RMSE
(11.3 cm), outperforming traditional machine learning models (CNN, MLP, RF, SVM) in
transferability. These findings demonstrate that integrating CNN-Transformer architectures
with appropriate signal filtering enables scalable, accurate, and objective soil thickness
mapping in complex terrain. The proposed approach also holds promise for broader GPR-
based subsurface applications, including soil horizon delineation and root system detection.
Keywords: Ground-penetrating radar; Soil thickness; Transformer; Headwater hillslopes;
Median filtering

Teng Li, Liwen Zhang, Youcheng Zhang, Qingmin Liao,


AXFL: Axial prior-guided cross-view fusion learning for radar semantic segmentation,
Expert Systems with Applications,
Volume 303,
2026,
130552,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Multiple 2D View spectrogram-based Radar Semantic Segmentation (MVRSS)
simultaneously leverages different radar-view spectrograms to capture comprehensive
spatial and velocity information of targets. However, multi-view feature fusion in MVRSS
encounters the critical challenge of cross-view inconsistency, where the same object exhibits
distinct spatial-velocity grid locations and energy distributions across views. Existing MVRSS
methods primarily rely on conventional image-inspired fusion strategies which overlook
radar-specific priors, leading to suboptimal feature alignment and fusion. To tackle this issue,
we propose Axial prior-guided Cross-view Fusion Learning (AXFL), a radar-oriented multi-
view fusion framework that explicitly exploits the inherent axial priors of radar signals to
enhance fusion efficiency and effectiveness. Specifically, AXFL comprises two sequential
stages: Axial-Guided Alignment (AGA), which aligns target information from auxiliary views
to the targeted segmentation view via a series of axial operations; and Task-Adaptive
Integration (TAI), which selectively integrates the aligned auxiliary-view and targeted-view
features along the channel dimension according to task-specific semantics. Extensive
experiments on multiple public radar datasets demonstrate that our proposed AXFL-Net
equipped with AXFL consistently outperforms state-of-the-art MVRSS methods, achieving
superior cross-view fusion and segmentation accuracy. The source code will be available at
[Link]
Keywords: Radar semantic segmentation; Deep learning; Autonomous driving; Feature
fusion; Radar prior

Ruizhe Feng, Shuzhao Zhu, Ruixin Jiang, Xin Cai, Rui Lin,
Degradation prediction of the low-Pt loading proton exchange membrane fuel cell based on
spatio-temporal Transformer network,
Energy,
Volume 342,
2026,
139625,
ISSN 0360-5442,
[Link]
([Link]
Abstract: Owing to the high energy density and low pollutant emissions, proton exchange
membrane fuel cells (PEMFCs) have emerged as promising solutions for sustainable energy
applications. Low-Pt loading proton exchange membrane fuel cells (low-Pt PEMFCs) hold
great promise in the field of sustainable energy for reducing the use of precious metals and
enhancing cost-efficiency. However, the long-term operational stability of low-Pt PEMFCs
remains a major barrier to large-scale commercialization. Accurate prediction of the
degradation process is essential for achieving system health management and ensuring
operational reliability. Under dynamic operating conditions, the increased nonlinearity and
uncertainty in the degradation process of low-Pt PEMFCs make accurate prediction more
difficult. To address these issues, a novel Transformer model named Adaptive Cross-
Dimensional Transformer (ACD-Transformer) is proposed. The model integrates a Dual
Deformation Attention Block (DDAB) and an Adaptive Filtering Block (AFB), effectively
extracting spatio-temporal dependencies in time-series data through deformable attention
mechanisms and learnable filtering thresholds. To evaluate the effectiveness and
generalizability of the proposed model, two degradation tests were designed for low-Pt
PEMFCs: Steady-State Conditions (SSC) and Multi-Temperature Region Conditions (MTRC),
the latter simulating real-world onboard temperature distributions. The results show that
the ACD-Transformer achieves superior performance on low-Pt PEMFCs compared to the
baseline models. Compared to the LSTM model, MAE is reduced by 63.64 % and 47.62 %
across two datasets. This work contributes to the development of prognostic methodologies,
facilitating effective health monitoring and durability enhancement of low-Pt PEMFCs.
Keywords: PEMFCs; Low-Pt loading; Transformer; Degradation prediction; Deep learning

Jackson S. Zaunegger, Paul G. Singerman, Ram M. Narayanan, Muralidhar Rangaswamy,


RadarTD: A Radar Text Dataset for multi-parameter optimization,
Natural Language Processing Journal,
Volume 12,
2025,
100178,
ISSN 2949-7191,
[Link]
([Link]
Abstract: This paper introduces the radar text dataset (RadarTD) for technical language
modeling. This dataset is comprised of sentences containing radar parameters, values, and
units determined from published radar literature. Additionally, each statement is assigned a
sentiment, goal priority, and goal direction label. In this work, we show how RadarTD may be
used to train simple Natural Language Processing (NLP) models to identify the attributes of
each sentence listed in RadarTD. Once the NLP models have identified these attributes from
text, we can use this information to develop Language Based Cost Functions (LBCF). Our
study shows that the proposed text classification model achieves a classification accuracy
between 96.7% and 97.8%, while the proposed named entity recognition model achieves an
F1 score of 99.7. These findings suggest that the developed models are capable of achieving
good performance for both text classification and named entity recognition for autonomous
radar applications. We then illustrate an example of how these models could be used with
Language Based Cost Functions to develop multi-parameter radar optimization schemes. We
also provide a method of providing scalarization weights for each parameter, to improve the
results of the optimization process.
Keywords: Text classification; Named entity recognition; Language modeling; Language-
based cost functions; Multi-parameter optimization; Cognitive radar

Changlong Wang, Jiawei Jiang, Chong Han, Hengyi Ren, Lijuan Sun, Jian Guo,
Through-Wall Multihuman Activity Recognition Based on MIMO Radar,
Computers, Materials and Continua,
Volume 83, Issue 3,
2025,
Pages 4537-4550,
ISSN 1546-2218,
[Link]
([Link]
Abstract: Existing through-wall human activity recognition methods often rely on Doppler
information or reflective signal characteristics of the human body. However, static
individuals, lacking prominent motion features, do not generate Doppler information.
Moreover, radar signals experience significant attenuation due to absorption and scattering
effects as they penetrate walls, limiting recognition performance. To address these
challenges, this study proposes a novel through-wall human activity recognition method
based on MIMO radar. Utilizing a MIMO radar operating at 1–2 GHz, we capture activity data
of individuals through walls and process it into range-angle maps to represent activity
features. To tackle the issue of minimal variation in reflection areas caused by static
individuals, a multi-scale activity feature extraction module is designed, capable of extracting
effective features from radar signals across multiple scales. Simultaneously, a temporal
attention mechanism is employed to extract keyframe information from sequential signals,
focusing on critical moments of activity. Furthermore, this study introduces an activity
recognition network based on a Deformable Transformer, which efficiently extracts both
global and local features from radar signals, delivering precise human posture and activity
sequences. In experimental scenarios involving 24 cm-thick brick walls, the proposed
method achieves an impressive 97.1% accuracy in activity recognition classification.
Keywords: MIMO radar; human activity; Transformer; through-wall

Lixing Shi, Xueling Liang, Wenchao Chen, Yaoqiang Liu, Tong Ding, Kun Qin, Bo Chen,
Hongwei Liu,
Masked variational transformer for complex clutter modeling and target detection,
Signal Processing,
Volume 239,
2026,
110236,
ISSN 0165-1684,
[Link]
([Link]
Abstract: Weak target detection commonly encounters intense clutter interference, which
overshadows weak signals and complicates the task. Taking advantage of the powerful data
mining capability of neural networks, more and more deep learning-based methods are
applied to radar target detection. Among the approaches, those founded upon unsupervised
learning methodologies exhibit remarkable merit because they dispense with the
requirement for target samples within the training step, making them highly applicable in
practical target detecting scenarios. However, existing methods suffer from limitations in
leveraging the range-Doppler (R-D) two-dimensional correlation and finely modeling in
multiple clutter scenarios. In this paper, an unsupervised Transformer-based detector (TrDet)
is proposed to break through the boundary of modeling capability. First, with the designed
two-dimensional position embedding (2-DPE) and global query embedding (GQE)
techniques, an unsupervised training strategy for R-D spectrum based on Transformer
framework is utilized to achieve refined clutter modeling. Then, radar target detection is
formulated as an out-of-distribution (OOD) detection task to mitigate clutter interference.
Moreover, the masked variational Transformer-based detector (MVTrDet) is further
proposed to prevent target information leakage when the target is in close proximity to the
clutter in Doppler domain. Compared with several relative algorithms, our proposed
methods are better suited for radar target detection in complex clutter environments. The
experimental results derived from both measured data and simulated data verify the
effectiveness of our proposed methods.
Keywords: Radar target detection; Clutter modeling; Range-Doppler (R-D) spectrum;
Unsupervised learning; Out-of-distribution detection; Transformer
Han Zhang, Shengheng Liu, Hao Chi Zhang, Le Peng Zhang, Tong Chen,
Globally fused hierarchical transformer for nonuniform frequency diverse arrays,
Digital Signal Processing,
Volume 159,
2025,
105009,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Nonuniform frequency diverse array (FDA) enables significant advantages in joint
range-angle measurement for radar localization tasks. In this article, we propose a global-
context two-stage transformer (GC-TSformer) specifically tailored for high-resolution target
localization using a nonuniform FDA transmitter and a single-channel receiver. The global
context network aggregates learnable array-wide features, which enhances resilience to
amplitude-phase errors and improves frequency accuracy. Built on the Transformer
backbone, the hierarchical model facilitates high-order estimation through multi-scale
interactions among array elements. This enables effective processing within the size
constraints of the covariance matrix. Additionally, the hybrid model refines element-specific
attention weights to ensure balanced representation. Extensive simulations and field
experiments verify that GC-TSformer achieves sub-meter localization accuracy, with
attention map analysis confirming robustness across varied sensor configurations.
Keywords: Target localization; Frequency diverse array; Nonuniform frequency offsets;
Parameter estimation; Transformer architecture

Jinyang Xie, Kanghui Zhou, Lei Han, Liang Guan, Maoyu Wang, Yongguang Zheng, Hongjin
Chen, Jiaqi Mao,
Enhancing multi-task learning-based Tornado identification using spatial and temporal
information from weather radar images,
Applied Soft Computing,
Volume 184, Part B,
2025,
113834,
ISSN 1568-4946,
[Link]
([Link]
Abstract: Tornadoes, as dynamic weather phenomena, exhibit unique spatial and temporal
evolution characteristics that reflect their formation and development. Existing tornado
detection algorithms often struggle with high false alarm rates, primarily due to insufficient
capture of temporal correlations in tornado development. As an improvement, we propose a
multi-task tornado identification network with three-dimensional temporal and spatial
information (TS-MTINet). Taking continuous three-frame radar data as input, the Multi-
frame Temporal Interaction Block (MTIB) utilizes multi-head attention to model the dynamic
interaction information between the radar data, thus exploring in-depth the temporal
features during tornado development. Further, we design a Spatial-Temporal Enhancement
Module (STEM), which analyzes the difference information between continuous data to
extract local and global spatial and temporal feature variations about tornadoes. Based on
this architecture, TS-MTINet incorporates a multi-task learning framework to perform
tornado detection and number estimation tasks simultaneously, thus extracting
comprehensive information related to tornadoes. To validate the performance of the
proposed model, we construct the first Chinese tornado identification dataset with fine
radar features. The experimental results show that the proposed method shows significant
advantages in several evaluation metrics, especially in reducing false alarms. In practical case
studies, compared to the traditional TVS method, TS-MTINet achieves an increase in POD of
approximately 30% and a decrease in FAR of about 20% in several typical tornado events.
Particularly in environments with strong interference, TS-MTINet demonstrates higher
detection accuracy, reflecting greater robustness and practical value.
Keywords: Deep learning; Multi-task learning; Tornado identification; Weather radar;
Attention mechanisms

Hao Huang, Xueli Hao, Lili Pei, Jiangang Ding, Yujiao Hu, Wei Li,
Automated detection of through-cracks in pavement using three-instantaneous attributes
fusion and Swin Transformer network,
Automation in Construction,
Volume 158,
2024,
105179,
ISSN 0926-5805,
[Link]
([Link]
Abstract: To improve the performance of existing through-crack detection networks by
solving the problem in which through-cracks are misidentified as simple surface cracks due
to limited feature extraction, this study proposes an automated detection method based on
the fusion of three instantaneous attributes and the Swin Transformer network. First, a
900MHZ ground-coupled radar system was used to collect data and construct the original
dataset (Origin). Then, the Hilbert-Huang transform was used to extract three instantaneous
attributes (instantaneous amplitude (IA), instantaneous phase (IP) and instantaneous
frequency (IF)). Second, four fusion-feature datasets, i.e., IA + IP, IA + IF, IP + IF and
IA + IP + IF, were constructed using the spectral weighting of individual features. Finally, the
Swin Transformer network was proposed to detect through-cracks. The results show that the
IA + IF dataset exhibited the best performance. The improved network achieved a 5.4%
increase in the mean average precision compared with the initial network, reaching 87.78%.
Keywords: Through-cracks detection; Ground penetrating radar (GPR); Hilbert-Huang
Transform; Three instantaneous attribute; Feature fusion; Swin Transformer

Zhigao Huang, Musheng Chen, Shiyan Zheng,


Dynamic spectral weighting in CausalSelfAttention: Enhancing transformer performance
through frequency-based head modulation,
Neurocomputing,
Volume 670,
2026,
132562,
ISSN 0925-2312,
[Link]
([Link]
Abstract: Transformer-based models are foundational to natural language processing, yet
optimizing their attention mechanisms remains challenging. This paper introduces Dynamic
Spectral Weighting (DSW), which enhances Transformer performance by modulating
attention heads based on their frequency-domain characteristics. Our hybrid static-dynamic
approach combines learned weights with spectral metrics through an adaptive gating
mechanism, significantly improving model generalization across various datasets.
Experiments demonstrate substantial validation loss reductions, with the most dramatic
improvements on larger datasets. Direct comparison with AdaAttention, a state-of-the-art
dynamic attention method, shows DSW’s consistent superiority across all datasets, achieving
up to 8.76% relative improvement on complex datasets, validating the effectiveness of
frequency-domain analysis over content-based dynamic attention. While introducing
moderate computational overhead, the performance gains validate the effectiveness of
incorporating spectral analysis into attention mechanisms. This work offers new insights into
attention head specialization and opens promising directions for integrating signal
processing concepts into deep learning architectures, potentially transforming how we
understand and design attention mechanisms for sequence modeling tasks.
Keywords: Transformer; Attention mechanism; Spectral analysis; Dynamic weighting;
Language modeling

Fulin Song, Hong Zhao, Yixin Zhao,


Three-dimensional inversion method of buried pipeline defects based on information fusion
technology and MSCPO-Transformer-BiGRU model,
Measurement,
Volume 264,
2026,
120250,
ISSN 0263-2241,
[Link]
([Link]
Abstract: As the primary mode of transportation in the oil and gas industry, pipeline
transportation offers significant advantages in terms of efficiency and cost-effectiveness.
However, due to the unique working environment of buried pipelines, internal defects are
prone to accumulate, leading to leakage accidents. To address these problems, firstly, an
information fusion technology that integrates the features of Transient electromagnetic
method (TEM) and weak magnetic field detection method (WMDM) is proposed for
detecting defects in buried pipelines. Secondly, recognizing the challenge of collecting a
substantial and diverse dataset of defects under actual working conditions, a defect data set
of buried pipelines is constructed by combining simulation and experiments. The importance
of the corresponding features of the three-dimensional information of buried pipeline
defects is analyzed by the Grey Relation Analysis (GRA). Then, in view of the current poor
performance of the three-dimensional defect prediction model for buried pipelines, A hybrid
model combining the Multi-strategy improvement of Crested Porcupine Optimization
Bidirectional Gated Recurrent neural network and Transformer mechanism (MSCPO-
Transformer-BiGRU) is proposed. To validate the superiority of proposed model, the
performance of MSCPO algorithm is evaluated by nine benchmark test functions. The
effectiveness of the MSCPO-Transformer-BiGRU model is further assessed through
comparative experiments as well as ablation experiments. The results demonstrate that
proposed model exhibits exceptional accuracy and stability when predicting three-
dimensional defect information in buried pipelines. Notably, when predicting the defect
length information of buried pipelines, its R2, MAE, MAPE, MSE and RMSE reached 0.98, 5.7,
0.07, 44.01 and 6.63 respectively.
Keywords: Buried pipeline; Information fusion technology; Defect detection; Defect three-
dimensional information; MSCPO-Transformer-BiGRU

Xuning Wang, Yuan Chen, Fuhao Wang, Kai Zheng, Yi Zong,


mPCT-LSTM: A lightweight human activity recognition model for 3D point clouds in
millimeter-wave radar,
Digital Signal Processing,
Volume 164,
2025,
105263,
ISSN 1051-2004,
[Link]
([Link]
Abstract: The millimeter-wave radar-based human activity recognition technology shows
significant potential in various domains, such as health monitoring, sports analysis, and
smart homes. Traditional methods rely on 2D feature spectrograms, which struggle to
balance low computational complexity with high recognition accuracy. Transforming radar
echoes into 3D point clouds and designing appropriate neural network models can mitigate
this conflict. However, most existing point cloud processing models are intended for dense
point clouds generated by LiDAR or depth sensors, leaving a gap in effective algorithms for
sparse mmWave radar point clouds. To address this need, we propose a lightweight model
for human activity recognition using sparse mmWave radar point clouds: mPCT-LSTM. This
model extracts spatial features from point clouds of human activity through a Point Cloud
Transformer (PCT) module, which consists of an embedding layer and four stacked offset-
attention layers. These spatial features are fed into a Long Short-Term Memory (LSTM)
module to capture temporal relationships between point cloud frames. Experimental results
demonstrate that the mPCT-LSTM model achieves an average recognition accuracy of
97.26% across three public datasets, outperforming the state-of-the-art by 1.41%.
Additionally, the model’s computational complexity is only 0.09 GFLOPS, a reduction of 70%
compared to current solutions.
Keywords: Millimeter-wave radar; Human activity recognition; 3D point cloud; Deep
learning; Long short-term memory (LSTM)

Seyed Alireza Khoshnevis, Abdollah Amirkhani,


Tracking with attention: A review of transformer-based object tracking,
Engineering Science and Technology, an International Journal,
Volume 73,
2026,
102263,
ISSN 2215-0986,
[Link]
([Link]
Abstract: Traditional object tracking methods are often based on convolutional neural
networks and handcrafted feature extraction techniques where they have seen remarkable
success. However, these methods still face limitations in capturing global dependencies and
contextual relationships in complex scenarios. Transformers, which were initially introduced
to the field of natural language processing, have transfigured vision tasks by leveraging the
self-attention mechanisms and global feature modeling capabilities. One of the tasks that
has been most affected by the use of transformers is the object tracking task. This review
explores the transformative impact of attention-based architectures in object tracking, and
provides a comprehensive analysis of the current frameworks and their core principles. The
ability of the attention mechanism to capture local and global dependencies and to associate
queries between frames, has helped transformer-based models to achieve state-of-the-art
performance. The utilization of transformers in object tracking has drastically increased over
the past few years, initiating the new “tracking-by-attention” paradigm. This work focuses on
different applications of transformer architecture in both single and multi-object tracking
where each task is divided further by methodology. End-to-end approaches and hybrid
fusion models that leverage additional data for tracking are also discussed. The models that
are discussed, are categorized by their main approaches and transformer usage, and
challenges such as computational cost and scalability are outlined, along with future
research opportunities informed by successful methods. By examining recent advancements,
this review is intended to advance understanding of transformer-based tracking capabilities
and to promote continued innovation in this rapidly evolving field.
Keywords: Single-object tracking (SOT); Multi-object tracking (MOT); Object detection;
Transformer model; Attention module; Data fusion

Hongxi Zhao, Yiran Shi, Wenchao He, Hewei Sun, Haoran Wang, Jiahao Liu, Lin Gui,
Novel graph neural network and GNN-C-Transformer model construction for direction of
arrival estimation,
Digital Signal Processing,
Volume 168, Part D,
2026,
105619,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Direction of Arrival (DOA) estimation is essential in radar, sonar, wireless
communications, and speech processing. Traditional methods like MUSIC and ESPRIT provide
high resolution but suffer from high computational complexity and poor performance in low
signal-to-noise ratio (SNR) environments. Recent advances in neural networks, particularly
Convolutional Neural Networks (CNN), improve accuracy and robustness; however, CNNs’
ability to reduce time complexity and improving robustness under low SNR conditions
remains insufficient. This paper presents a novel framework for DOA estimation in sparse
arrays based on Graph Neural Networks (GNN) and proposes an entirely new array-based
graph connectivity structure. By modeling the array geometry as a graph, our GNN approach
captures spatial relationships effectively, addressing the challenges of time complexity and
low SNR. We further integrate Transformer layers to capture both spatial and temporal
dependencies, enhancing the model’s performance. Experimental results demonstrate that,
at SNRs ≤5dB, our GNN-based framework and the GNN-C-Transformer model developed
thereon achieve superior accuracy compared to existing methods, while exhibiting lower
computational complexity than all other algorithms except ESPRIT. This work advances the
application of GNN-based DOA estimation by providing a scalable solution for large-scale,
multi-dimensional signal processing in both dense and sparse array configurations.
Keywords: Direction of arrival estimation; Deep learning; Parameter estimation; Graph
neural networks; GNN-C-Transformer; Sparse array

Yongsheng Yao, Chen Liu, Jue Li, Jinliang Wu,


Transformer-based data generation and lightweight robust detection network for complex
pavement defects,
Measurement,
Volume 256, Part C,
2025,
118213,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Ground-penetrating radar (GPR) accurate detection of road surface hidden defects
is very important for road maintenance. However, due to the complexity of disease
waveforms and the scarcity of sample data, existing models face problems such as
insufficient accuracy and have weak generalization ability in engineering scenarios. To this
end, this paper proposes a two-stage method for improving the accuracy of complex
pavement diseases with small-scale data. Firstly, a B-scan Latent Generative Adversarial
Network (BL-GAN) was proposed to effectively expand the small-scale dataset by combining
the self-attention mechanism and hierarchical feature fusion to synthesize realistic defect
patterns. The t-SNE visualization confirms the high feature agreement between the synthetic
and real defect images. Secondly, by systematically integrating MobileViTv2 and the channel
attention mechanism, a lightweight transformer-based detector ECA-MobileViTv2
YOLOv5s(EV2-YOLOv5s) is designed to achieve accurate localization of multi-scale defects
while maintaining computational efficiency. Experimental results show that the proposed
method achieves 96.5% mAP@0.5 with 144.1 FPS (GPU) inference speed, 19.7 FPS (CPU)
inference speed, and 7.4 MB model size, which is significantly better than traditional
methods. Finally, the contributions of Multi-head Attention (MHA) and Linearly Separable
Self-Attention (LSA) in global feature aggregation and local feature extraction were analyzed
to support the subsequent optimization of Transformer-based defect detection algorithms.
This work provides a practical solution for the detection of hidden road surface diseases in
resource-constrained scenarios.
Keywords: Asphalt pavement; Ground penetrating radar; Transformer; Object detection;
Data augmentation

Bixuan Gao, Riwei Zhang, Xiangyu Kong, Gaohua Liu, Kaijie Fang, Meimei Duan,
A Novel Carbon Emission Calculation Method for Power System Based on Personalized
Transformer with Two-Stage Training,
Engineering,
2026,
,
ISSN 2095-8099,
[Link]
([Link]
Abstract: Accurately and comprehensively calculating carbon emissions in the power system
is a fundamental prerequisite for achieving low-carbon energy transitions. Existing carbon
flow theory-based methods primarily concentrate on emissions from the grid and load sides,
while the methods for generation-side emissions often rely on costly continuous monitoring
systems or imprecise default emission factors. However, in practice, most power generation
units cannot achieve real-time and precise carbon emission measurement, leading to
generation side data deviations that affect overall computational accuracy. To address the
limitations, this paper proposes a novel carbon emission measurement method based on
heterogeneous data and personalized Transformer, applicable to various power generation
units. This method has several key innovations: ① extends traditional total electricity
production based models are extended to a multi-feature framework to capture similarities
and differences in time-series data among various generator units; ② an improved
Transformer is designed that integrates short-term relationship extraction, long-term
differential identification and long-term and short-term feature fusion modules to enhance
the multi-feature based emission mapping process, and ③ a two stage training protocol is
adopted, with self-supervised pretraining of feature extractors followed by fine tuning, to
accelerate convergence and improve accuracy. Experiments on real-world generation-unit
data show that the proposed method reduces average RMSE by 22.3% relative to a standard
Transformer and by 15.9% relative to Informer. Further validation utilizing the Institute of
Electrical and Electronics Engineers (IEEE) 30-bus test case confirms the effectiveness and
applicability of the model for carbon emission measurement across all segments of the
power system.
Keywords: Power system carbon emissions; Carbon emission calculation; Transformer; Multi-
frequency feature mapping

Walter Brescia, Pedro Gomes, Laura Toni, Saverio Mascolo, Luca De Cicco,
GT-MilliNoise: Graph transformer for point-wise denoising of indoor millimetre-wave point
clouds,
Signal Processing: Image Communication,
Volume 142,
2026,
117453,
ISSN 0923-5965,
[Link]
([Link]
Abstract: Millimetre-wave (mmWave) radars are gaining popularity thanks to their low cost
and robustness in low-visibility conditions. However, the 3D point clouds they produce are
sparser and noisier than those from LiDARs and depth cameras. These differences create
challenges when applying existing methods, originally designed for dense point clouds, to
mmWave data. Specifically, there is a gap in point-level precision tasks, such as full point
cloud denoising for mmWave data, partly due to the lack of fully annotated datasets. In this
work, we employ the MilliNoise dataset, a fully annotated indoor mmWave point clouds
dataset, to advance the understanding of mmWave point clouds denoising via two main
steps: (i) we carry out an experimental analysis of the most common point cloud processing
approaches and show their limitations in exploring the local-to-global structures in sparse
and noisy point clouds; (ii) in light of the identified limitations, we propose a graph-based
transformer architecture, denoted as GT-MilliNoise, composed of two main blocks to
effectively leverage both the temporal and geometric structures of the data: a Temporal
block leverages the sparsity of data to learn the dynamic behaviour of the points; a
Geometric block, uses a point-wise attention mechanism to form representative
neighbourhoods for feature extraction. The experimental results obtained in the MilliNoise
dataset show that our proposed GT-MilliNoise architecture outperforms the state-of-the-art
both qualitatively and quantitatively. Specifically, it achieves 75% accuracy (5% gain
compared to the state-of-the-art), and a significantly low Earth Mover’s distance value of
0.193.
Keywords: Point cloud; mmWave; Denoising; Deep learning

Adil Ali Saleem, Hafeez Ur Rehman Siddiqui, Muhammad Amjad Raza, Sandra Dudley, Julio
César Martínez Espinosa, Luis Alonso Dzul López, Isabel de la Torre Díez,
Ultra Wideband radar-based gait analysis for gender classification using artificial intelligence,
Array,
Volume 27,
2025,
100477,
ISSN 2590-0056,
[Link]
([Link]
Abstract: Gender classification plays a vital role in various applications, particularly in
security and healthcare. While several biometric methods such as facial recognition, voice
analysis, activity monitoring, and gait recognition are commonly used, their accuracy and
reliability often suffer due to challenges like body part occlusion, high computational costs,
and recognition errors. This study investigates gender classification using gait data captured
by Ultra-Wideband radar, offering a non-intrusive and occlusion-resilient alternative to
traditional biometric methods. A dataset comprising 163 participants was collected, and the
radar signals underwent preprocessing, including clutter suppression and peak detection, to
isolate meaningful gait cycles. Spectral features extracted from these cycles were
transformed using a novel integration of Feedforward Artificial Neural Networks and
Random Forests , enhancing discriminative power. Among the models evaluated, the
Random Forest classifier demonstrated superior performance, achieving 94.68% accuracy
and a cross-validation score of 0.93. The study highlights the effectiveness of Ultra-wideband
radar and the proposed transformation framework in advancing robust gender classification.
Keywords: Gait; Ultra-wide band radar; Gender classification; Spectral features; Feed
forward artificial neural network; Ridge classifier; Hist gradient boosting

Jiahao Liu, Yiming Zhang, Liang Song, Zheng Tong,


SCB-ADAE: An attention-based deep autoencoder for ground penetrating radar signal
denoising,
Engineering Applications of Artificial Intelligence,
Volume 160, Part A,
2025,
111902,
ISSN 0952-1976,
[Link]
([Link]
Abstract: In buried object detection, recorded signals of a ground penetrating radar (GPR)
inevitably include noise interference owing to complex underground environments. Existing
rule- and data-driven denoising methods struggle to handle non-Gaussian and real-world
noise because the rule-driven ones rely on the assumptions of simplified noise
characteristics and the data-driven ones cannot capture fine- and global-scale features of a
GPR signal well. To address the problem, this study proposes an attention-based denoising
model called the Swin-Conv Block with Attention Denoising Autoencoder (SCB-ADAE). The
model first feeds a GPR signal into a SCB module, which extracts a tensor with the fine-scale
features in the signal, such as sharp reflective interfaces and abrupt amplitude variations.
The feature tensor then passes through an ADAE module that uses encoder-decoder
structure with the self-attention to enhances the representation of the global-scale signal
features. Finally, the feature tensor from the ADAE module is decoded by another SCB
module to generate a denoised GPR signal, where the tensor includes the fine-scale and
global features of the raw signal. An experiment with three types of GPR signals
demonstrates the effectiveness of the proposed model: radar signals with Gaussian noise,
radar signals with inhomogeneous-material noise, and real-world signals. radar signals with
Gaussian noise, radar signals with inhomogeneous-material noise, and real-world signals.
Experimental results demonstrate that the proposed model outperforms other state-of-the-
art denoising methods on denosing the three types of GPR signals, where the signal-to-noise
ratio, peak signal-to-noise ratio, and structural similarity index are improved to 20.64, 14.59,
and 0.366, respectively.
Keywords: Ground penetrating radar; Denoising; Attention-based model; Autoencoder

Prasshanth Chennai Viswanathan, Ahaan Banerjee, Naveen Venkatesh Sridharan, Ganjikunta


Chakrapani, Sugumaran Vaithiyanathan,
Advancing automobile dry clutch fault diagnosis through innovative imaging techniques and
Vision transformer integration,
Measurement,
Volume 242, Part B,
2025,
115975,
ISSN 0263-2241,
[Link]
([Link]
Abstract: The study investigates the significance of clutch condition monitoring in
automotive transmissions to preempt mechanical failures, enhance efficiency, and mitigate
risks to human safety and maintenance costs. It explores the integration of Vision
Transformer (ViT) with imaging techniques, such as scalograms, spectrograms, polar plots,
radar plots, and Hilbert-Huang transforms, to diagnose faults in dry friction clutches. By
transforming vibration signals into image representations and utilizing ViT for fault
classification, the study aims to identify the most effective imaging technique and optimal
hyperparameters for accurate fault diagnosis. Experimental studies on a test rig with varying
fault conditions demonstrate the effectiveness of ViT in diagnosing clutch faults when
coupled with different image conversion techniques. The results highlight the potential of
integrating spectrogram image processing with ViT, achieving a 100% accuracy in fault
diagnosis for clutch systems, thus advancing the analysis of faults in clutch systems.
Keywords: Vision transformer; Condition monitoring; Dry clutch; Imaging technique

Md. Jalil Piran, Xiaoding Wang, Ho Jun Kim, Hyun Han Kwon,
Precipitation nowcasting using transformer-based generative models and transfer learning
for improved disaster preparedness,
International Journal of Applied Earth Observation and Geoinformation,
Volume 132,
2024,
103962,
ISSN 1569-8432,
[Link]
([Link]
Abstract: Due to the rapidly changing climate conditions, precipitation nowcasting poses a
daunting challenge because it is impossible to make accurate short-term forecasts due to the
rapid fluctuations in weather conditions. There are limitations to traditional methods of
forecasting precipitation, such as the use of numerical models and radar extrapolation, when
it comes to providing highly detailed and timely forecasts. With the help of contemporary
machine learning (ML) models, including deep neural networks, transformers and generative
models, complex precipitation nowcasting tasks can be performed in an efficient way. To
address this critical task and enhance proactive emergency disaster management, we
propose an innovative method based on transformer-based generative models for
precipitation nowcasting. Our study area is the Soyang Dam basin in South Korea, located
upstream of the Han River, characterized by a monsoon climate with approximately 1200
mm of annual precipitation. To develop a precipitation nowcasting model, radar composite
data from 10 weather radars across South Korea is used. By utilizing radar reflective data in
order to train our model, we are able to effectively predict future precipitation patterns, thus
mitigating the risk of catastrophic weather conditions caused by heavy rainfalls. This dataset
covers reflectivity data from 2018 to 2022, with a spatial resolution of 1km over a 960 ×
1200 grid. Normalization using the min–max scaler method is applied to this reflectivity
data, which is then transformed into grayscale images for uniform comparison. We enhance
performance effectively by employing transfer learning with pre-trained Transformer
models. Initially, we train the model using a comprehensive dataset. Subsequently, we fine-
tune it for precipitation nowcasting using radar reflective data. This adaptation improves the
accuracy of rainfall forecasting by capturing crucial features. Leveraging prior task knowledge
through transfer learning not only enhances prediction accuracy but also increases overall
efficiency. In terms of predictive accuracy, extensive experimental results demonstrate that
our transformer-based nowcasting model outperforms related approaches, including
conditional generative adversarial networks (cGANs), U-Net, convolutional long short-term
memory (ConvLSTM), pySTEP. As a result of this research, disaster preparedness and
response will be greatly improved through improved weather prediction.
Keywords: Precipitation nowcasting; Transformer-based generative model; Radar reflective
data; ConvLSTM; cGAN; U-net

Dikun Hu, Weidong Gao, Kai Keng Ang, Mengjiao Hu, Rong Huang, Yingying Shao,
FECT-OSA: A transformer-enhanced multimodal system for non-contact sleep apnea
monitoring,
Alexandria Engineering Journal,
Volume 128,
2025,
Pages 628-641,
ISSN 1110-0168,
[Link]
([Link]
Abstract: Obstructive sleep apnea (OSA) is a common sleep disorder linked to an increased
risk of cardiovascular and neurocognitive disorders. While portable wearables provide a low-
cost alternative for OSA detection compared to polysomnography (PSG), their prolonged
wear burden, limited accuracy, and susceptibility to motion artifacts hinder practical
application. This study introduces FECT-OSA, a non-contact framework that extracts reliable
OSA indicators from piezoelectric physiological signals (PPS). FECT-OSA combines three key
components: adaptive signal processing, TransUnet, and ConvTransLSTM. Adaptive signal
processing utilizes entropy-based selection to determine the most stable input, improving
signal robustness. TransUnet improves feature extraction by integrating U-Net’s localization
with Transformer-based global modeling. It segments critical spectrogram regions, mitigating
motion artifacts and sidelobe interference. ConvTransLSTM integrates CNNs for local feature
extraction, a Multimodal Transformer for global alignment, and LSTM for temporal
modeling, enhancing OSA detection accuracy. In a study of 34 patients, FECT-OSA achieved
84.79% accuracy, 74.80% sensitivity, and 88.53% specificity, surpassing state-of-the-art
portable devices. For severe OSA patients (AHI≥30), it attains 94.1% accuracy and 90.9%
sensitivity/specificity, exceeding traditional methods by 5%–15% in accuracy and 5%–22% in
sensitivity. FECT-OSA’s non-contact design minimizes patient discomfort while ensuring high
accuracy across varying OSA severities, providing a practical home-based alternative.
Keywords: Obstructive sleep apnea; Smart sleep monitoring; Non-contact detection;
Enhanced feature extraction; Cross-attention fusion; Time-frequency analysis

Kai Yue, Zemeng Huang, Yubing Li, Yujia Chen, Tao Tan, Tao He, Yu Wang, Peng Ke, Xiuping Li,
A parallel Class-E power amplifier with doubly-tuned transformer-based load network and
high-efficiency cascode in 110-nm CMOS,
AEU - International Journal of Electronics and Communications,
Volume 205,
2026,
156117,
ISSN 1434-8411,
[Link]
([Link]
Abstract: This article presents a doubly tuned (DT) transformer-based parallel Class-E power
amplifier (PA). A compact parallel Class-E load network consisting of only one DT transformer
and a pair of capacitors is proposed to enhance output power and efficiency. Compared with
the traditional DT transformer-based series Class-E load, the proposed DT transformer-based
parallel Class-E load can further mitigate the constraints placed on device size and reduce
the impedance transformation ratio of the load. Besides, a cascode structure with
neutralization and charging acceleration capacitor (CX) is used as active core to enhance gain
and efficiency. The gain and stability of the active core are quantitatively analyzed based on
the transistor small-signal model, and it can be concluded that the gain of the active core
exhibits an increasing trend with the growth of CX while ensuring stability. As a proof of the
design, a 12 GHz Class-E PA is fabricated using 110-nm CMOS process. The measurement
results show that the proposed PA realizes a peak power-added-efficiency (PAE) of 30.9%, a
maximum saturated output power (Psat) of 18.2 dBm and a peak gain of 19.0 dB. The core
area of the circuit is only 990 μm ×260μm.
Keywords: Class-E power amplifier; Doubly-tuned transformer; Harmonic impedance;
Charging acceleration capacitor; Stability

Gang Xiong, Wenyu Huang, Tao Zhen, Shuning Zhang,


Fractal-domain transformer based on learnable multifractal spectrum for chaotic systems
classification,
Physica A: Statistical Mechanics and its Applications,
Volume 658,
2025,
130276,
ISSN 0378-4371,
[Link]
([Link]
Abstract: Conventional deep learning in the spatiotemporal-frequency domain frequently
encounter challenges in terms of slow convergence rates and limited generalization,
particularly for classification of chaotic systems. To address these limitations, this paper
introduces a novel fractal-inspired deep network model, specifically, the Multifractal
Spectrum Transformer (MFS-Transformer), grounded in learnable multifractal analysis.
Initially, we put forward the conceptual framework of fractal learning, compared with
traditional fractal signal processing methodologies and spatial-temporal domain learning
paradigms. Subsequently, a Learnable Multifractal Spectrum (LMFS) derived from 3D spatial
gridding, coupled with fractal-domain filtering, is proposed to construct the iterative
learning process within the fractal domain. Further, we formulate the MFS-Transformer, an
innovative architecture that integrates multi-channel embedding, LMFS, fractal-domain
filtering, residual fusion mechanisms, a mixer module, and a classifier, tailored for chaotic
system classification. Ultimately, we evaluate the efficacy of our model in classifying 3D
chaotic systems under stringent conditions of short-term sequences and low Signal-to-Noise
Ratio (SNR). Experimental outcomes underscore the substantial performance gains achieved
by the MFS-Transformer, with classification accuracy enhancements of 13.34 % and 5.00 %
over existing Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs),
respectively, under SNR= 0 dB and 32-sample sequences. These findings validate the
superiority of the MFS-Transformer in addressing the complexities of chaotic system
classification under complex scenarios. This research not only advances the frontier of
fractal deep learning but also presents a novel perspective and methodology for tackling
intricate spatiotemporal classification problems.
Keywords: Fractal deep learning; Learnable multifractal spectrum; Multifractal Transformer;
Classification

Sachin Kishanrao Bhingikar, Rishi Raj Sharma, Ram Bilas Pachori,


Radar-based non-contact heart rate monitoring: A comprehensive review,
Digital Signal Processing,
Volume 168, Part D,
2026,
105630,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Heart rate (HR) measurement is an essential physiological parameter used in
various fields, such as healthcare and human-computer interaction. This article presents a
comparative evaluation of contact and non-contact approaches for heart rate monitoring,
with a specific emphasis on radar-based non-contact methods. It begins by examining
conventional contact-based techniques like electrocardiography and photoplethysmography,
outlining their operational principles, applications, and limitations. While contact methods
ensure high accuracy and reliability, they can be inconvenient for long-term monitoring in
ambulatory environments owing to their inconvenience. The discussion then shifts to non-
contact approaches with a focus on radar-based techniques, which offer advantages such as
being non-intrusive, capable of remote operation, and having the potential for continuous
monitoring without physical contact. The paper provides insights into the underlying
principles of radar-based heart rate detection along with its varied applications across
domains like healthcare, automotive systems, and smart environments while also addressing
implementation challenges. Additionally, it discusses future directions pertaining to radar-
based non-contact heart rate detection and concludes by highlighting the strengths and
limitations.
Keywords: Radar sensing; Heart rate (HR) detection; Non-contact sensing; Vital sign
monitoring; Radar signal processing

Mohammad Hossein Shirazi, Sira Yongchareon, Anuradha Singh, Jing Ma,


A survey on machine learning approaches for vital sign monitoring using radar,
Measurement,
Volume 253, Part D,
2025,
117707,
ISSN 0263-2241,
[Link]
([Link]
Abstract: The integration of machine learning methodologies with radar-based vital sign
monitoring represents a significant advancement in non-contact healthcare surveillance
systems. This systematic literature review synthesizes and critically analyzes research from
2020 to 2025, addressing substantive theoretical and methodological gaps in extant
literature. Our comprehensive taxonomic classification of machine learning paradigms
employed in this domain elucidates the progressive refinement from conventional
algorithmic approaches to sophisticated deep learning architectures, with particular
emphasis on hybrid neural network configurations optimized for physiological signal
extraction in non-stationary environments. Methodologically, this survey contributes a
rigorous evaluation framework comprising standardized assessment protocols, quantifiable
performance metrics, and cross-validation methodologies—elements conspicuously absent
in previous reviews. Empirical analysis demonstrates substantial correlations between
dataset demographic characteristics and algorithmic generalizability, with heterogeneous
participant cohorts yielding markedly enhanced performance across cardiac, respiratory, and
hemodynamic parameter estimation tasks. The review delineates four distinct
developmental phases in the field’s chronological evolution and provides analytical insight
into persistent technical challenges: motion artifact compensation, multi-subject
disambiguation, and the translation of laboratory efficacy to clinical utility. This
comprehensive examination of computational approaches for radar-based vital sign
monitoring establishes a theoretical foundation and methodological framework to guide
future research towards physiologically robust and clinically viable implementations.
Keywords: Non-intrusive vital sign monitoring; Machine learning; Radar

Aurora Polo-Rodríguez, Miguel Ángel Anguita-Molina, Ignacio Rojas-Ruiz, Javier Medina-


Quero,
Multi-occupant tracking with radar and wearable devices for enhanced accuracy in indoor
environments,
Engineering Applications of Artificial Intelligence,
Volume 154,
2025,
110872,
ISSN 0952-1976,
[Link]
([Link]
Abstract: This work explores the integration of millimetre-wave (mmWave) radar and a
minimal configuration of ultra-wideband (UWB) devices for enhanced multi-occupant
tracking in real domestic environments. Using a low-cost, non-intrusive, and rapidly
deployable device setup, our approach addresses key challenges in multi-occupant tracking,
including individual identification and ease of installation. While mmWave radar precisely
detects occupant presence, it lacks individual recognition and exhibits limited sensitivity.
This limitations are addressed by incorporating a minimal configuration of UWB (wearable
tags and ambient anchors), enabling individual identification through signal strength
measurements. Several data autoencoder models, such as long short-term memories
(LSTMs), Convolutional Neural Networks (CNN) or Transformers, were evaluated.
Experiments conducted in two real-world domestic settings, each with up to three
inhabitants, demonstrate the effectiveness of combining mmWave and UWB technologies
for indoor multi-occupant tracking. Our results show that ConvLSTM achieves the best
performance with a mean squared error (MSE) between 0,0142 and 0,0433 in single and
multi-occupation, respectively. These findings suggest promising applications for accurate
inhabitant tracking in ambient assisted living and other smart environment contexts.
Keywords: Multi-tracking; Ultra-wideband; Wave radar; Autoencoder models

Yuankang Ye, Feng Gao, Shaoqing Zhang, Chang Liu,


Improving precipitation nowcasting via multiphysical parameter fusion in radar echo
extrapolation,
Journal of Hydrology,
Volume 668,
2026,
134947,
ISSN 0022-1694,
[Link]
([Link]
Abstract: Radar-based precipitation nowcasting plays a vital role in short-term
hydrometeorological forecasting and water resource management. Existing modeling
methodologies typically simplify precipitation nowcasting to a task of spatiotemporal
sequence prediction based on radar echo reflectivity data. However, the reliance on
unimodal reflectivity data including intensity-only information restricts the model’s ability to
characterize the phase evolution and dynamic processes of hydrometeor particles,
ultimately leading to insufficient extrapolation accuracy. This study breaks through the
conventional unimodal data paradigm, aiming to capture the complex dynamic evolutionary
features of hydrometeor particles. We integrate radar echo reflectivity and four additional
physical parameters of hydrometeor particles into a deep learning framework and propose a
novel Physics-Informed Multimodal Echo Extrapolation neural network (PIEE). Furthermore,
we systematically investigate the individual contributions of each physical parameter to the
accuracy of radar echo extrapolation. Specifically, PIEE adopts a three-stage structure. First, a
multimodal encoder with a dual-branch attention-based fusion strategy is used to capture
diverse physical signals. Second, a novel gated spatiotemporal self-attention module is
designed for deep feature extraction. Finally, the decoding stage generates the extrapolated
radar echoes. Experimental results on a real multimodal radar echo dataset show that the
proposed model demonstrates superior performance in two aspects. First, under a unimodal
baseline architecture, the PIEE model clearly outperforms the comparison model. Second,
after fusing multiple physical parameters, the PIEE achieves significant improvements in all
the evaluated metrics, especially in the CSI and HSS metrics for the high echo intensity
region (≥ 40 dBZ), with improvements of up to 24.2% and 20.3%, respectively. Furthermore,
systematic ablation experiments on physical parameters quantify the effects of different
combination methods on extrapolation accuracy, highlighting the potential of physics-
informed, multimodal deep learning approaches in improving short-term hydrological
prediction accuracy, with implications for flood forecasting, early warning systems, and
hydrometeorological risk management at catchment scales.
Keywords: Deep learning; Precipitation nowcasting; Radar echo extrapolation;
Hydrometeorological forecasting

Mohammad Alikhani, Amirhossein Nikoofard,


Long-term error compensation in inertial IEKF-based localization with transformer-adaptive
noise tuning and IMU missing data handling,
Measurement,
Volume 257, Part E,
2026,
118928,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Reliable inertial localization requires accurate state estimation over extended
durations, particularly in challenging environments where sensor noise and missing data can
significantly degrade performance. In this work, we present an inertial localization
framework based on invariant extended Kalman filter (IEKF) that incorporates a transformer-
based adaptive noise covariance tuner. This tuner becomes effective after accumulating
inertial measurement unit (IMU) data over time, enabling long-term error correction. The
proposed approach is designed to complement other localization frameworks by gradually
correcting accumulated localization errors using historical sensor readings. Additionally, we
introduce a missing IMU data handling mechanism based on a long short-term memory
(LSTM) network that reconstructs incomplete inertial measurements, ensuring the
robustness of the framework under real-world conditions. Experimental results demonstrate
that our method effectively reduces drift over time and significantly enhances localization
performance. Compared to other inertial approaches, our method achieves an average
reduction of 0.89% in relative translational error across all sequences. Notably, for
sequences with missing values, it yields improvements of 5.09%, 1.21%, and 1.22%,
respectively.
Keywords: Localization; Kinematic state estimation; Inertial navigation; Intelligent Kalman
filter; Invariant extended Kalman filter

Kitoshi Kawai, Bungo Konishi, Ryo Natsuaki, Akira Hirose,


Quaternion reservoir computing for spatiotemporal analysis in polarimetric synthetic
aperture radar,
Neurocomputing,
Volume 658,
2025,
131633,
ISSN 0925-2312,
[Link]
([Link]
Abstract: Quaternion neural networks possess high generalization ability in three-
dimensional (3D) information space by representing every 3D data point as a single
quaternion entity. In polarimetric synthetic aperture radar (PolSAR) applications such as land
surface classification, they are expected to deal with 3D Poincare parameters as inseparable
physical entities. With the increasing acquisition frequency, there is a growing demand also
for efficient and robust techniques to monitor temporal or spatiotemporal changes.
Reservoir computing (RC) is a variation of recurrent neural networks (RNNs) capable of
detecting changes in series data with low computational cost. In this context, we propose
quaternion reservoir computing (QRC) for spatiotemporal analysis in PolSAR. First, in a
benchmark prediction task for chaotic time-series derived from the 3D Lorenz equations, we
demonstrate that QRC achieves higher prediction accuracy than real-valued RC and
conventional RNNs. Secondly, we conduct spatiotemporal anomalous change detection for
actual PolSAR data of (1) rice fields having seasonal changes in Japan and (2) Amazon
rainforest suffering from deforestation in Brazil. Compared with real-valued RC, RNNs, one-
dimensional convolutional neural networks, Transformer, and non-adaptive methods based
on complex Wishart and Pauli RGB, QRC shows a larger area under the curve (AUC) score,
demonstrating its high efficacy in capturing spatiotemporal anomalous variations in PolSAR
data. Besides such high performance, QRC shows low training cost, which is very suitable for
real-time processing in edge computing including highly frequent satellite observations.
These experimental results indicate that combining quaternion representation with RC is a
promising approach for analyzing the ever-increasing volume of PolSAR data.
Keywords: Quaternion; Reservoir computing; Polarimetric synthetic aperture radar (PolSAR);
Spatiotemporal analysis

Nan Xia, Siqi Wang, Weijia Lu,


Automotive radar co-channel interference mitigation and target detection based on third-
order cumulant and neural network,
Measurement,
Volume 266,
2026,
120494,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Automotive millimeter-wave (mmWave) radar systems face significant challenges
from noise and co-channel interference, which can obscure weak target reflections. This
paper proposes a novel target localization algorithm that combines a third-order cumulant
based interference suppression technique with a dual-stream deep neural network to
enhance detection performance. The third-order cumulant operation exploits higher-order
statistics to suppress co-channel interference among radars. A dual-stream convolutional
neural network then extracts image-based features from heatmaps and edge maps, enabling
robust detection and localization of weak and overlapping targets in complex multi-target
scenarios. The proposed method is evaluated on both simulated radar data and real
measurements from an automotive mmWave radar. Experimental results demonstrate that
the proposed method effectively mitigates interference and significantly improves target
detection accuracy and localization precision compared to baseline methods. These results
indicate that the integration of higher-order statistical filtering with advanced neural
networks can greatly enhance radar target localization in challenging interference
environments.
Keywords: Automotive radar; Co-channel interference mitigation; The third-order cumulant;
Deep learning; Target localization

Zhipeng Qing, Kecheng Ge, Shunsheng Zhang, Jing Yang, Zhijin Wen, Youlei Pu,
An inverse synthetic aperture radar imaging framework based on multi-layer networks and
heat conduction attention,
Engineering Applications of Artificial Intelligence,
Volume 167, Part 1,
2026,
113708,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Accurate compensation is essential for achieving high-resolution inverse synthetic
aperture radar (ISAR) imaging. Traditional parametric methods usually rely on iterative
optimization of the objective function to compensate for target motion in radar echoes.
However, the iteration process is often computationally intensive, difficult to integrate into
deep learning frameworks, and may discard sufficiently acceptable intermediate solutions.
To address these challenges, this study proposes a deep unfolding-based translational
compensation network that combines unsupervised learning with gradient back-
propagation. A prototype network is incorporated to monitor the imaging process, enabling
early termination of iterations. Moreover, a U-shaped network architecture based on a heat
conduction attention is employed to enhance ISAR image resolution and focusing
performance. To solve the problem of offset or splitting in the imaging results caused by
residual motion errors, a learnable affine transformation is employed for automatic
centering. These modules are integrated into an echo-to-image ISAR imaging framework.
Experimental results on both simulated and real radar data demonstrate the framework’s
effectiveness and robustness.
Keywords: Inverse synthetic aperture radar imaging; Translational compensation; Heat
conduction; Deep unfolding network; Affine transformation

Tamer Saleh, Shimaa Holail, Xiongwu Xiao, Gui-Song Xia,


High-precision flood detection and mapping via multi-temporal SAR change analysis with
semantic token-based transformer,
International Journal of Applied Earth Observation and Geoinformation,
Volume 131,
2024,
103991,
ISSN 1569-8432,
[Link]
([Link]
Abstract: Flood detection in crisis and disaster management is significantly facilitated by the
analysis of synthetic aperture radar (SAR) imagery. Traditional flood detection techniques
focus more on SAR image pairs than on the optical level. However, the distinctive
characteristics of SAR images, characterized by limited visual information, pervasive speckle
noise, and analogous backscatter signals, present formidable obstacles to accurately
identifying water bodies and extracting change features. Consequently, the performance of
existing methods remains unsatisfactory. This paper addresses this challenge by focusing on
disparities between SAR image pairs and introducing a pioneering semantic token-based
transformer network, denoted as SemT-Former, to enhance flood detection accuracy. SemT-
Former operates by prioritizing changes of interest rather than fully comprehending the
entire image scene. This is achieved through the integration of temporal-wise feature
representation and the introduction of a class token to capture high-level segmentation
associated with changes in water bodies. These innovations augment the model’s capacity to
discriminate between genuine changes in water bodies and spurious changes induced by
similar signals or speckle noise. The effectiveness of SemT-Former is evaluated through a
case study in Khartoum, Sudan, focusing on flood detection and the estimation of damaged
farmland near river confluences. Experimental results demonstrate that SemT-Former
outperforms existing methods, exhibiting a 90.6% improvement in F1-score and an 88.5%
enhancement in IoU. This underscores SemT-Former as a promising solution for precise and
effective flood mapping from SAR images.
Keywords: Flood extraction; Sentinel-1 SAR; CNN; Transformer; Khartoum–Sudan

Xinwen Yi, Dongyu He, Jiachang Liu, Xiaoling Zhu, Zhifang Pan,
HaFeiT: A fetal hypoxia diagnosis model using health status and fetal heart rate based on
vision transformer,
Biomedical Signal Processing and Control,
Volume 116,
2026,
109503,
ISSN 1746-8094,
[Link]
([Link]
Abstract: Electronic fetal monitoring is widely employed during pregnancy and labor periods
to detect fetal hypoxia. Due to internal observer differences, visual inspection of
cardiotocography based on clinical guidelines exhibits a high false-positive rate. Therefore,
AI-based cardiotocography classification using fetal heart rate signals is essential and
challenging, as it assists clinicians in objectively and accurately assessing fetal health status.
Most existing methods either focus on single-modal cardiotocography classification based on
signals or employ multimodal modeling with signals combined with natural language or
other statistical features. However, these methods do not consider the health status of the
gravida or fetus. This study is the first to incorporate maternal and fetal health status into
the fetal heart rate classification task. The fetal heart rate is transformed into a two-
dimensional image, and health status is extracted from the database. Both of them are fed
to the network we propose for classification. To address the lack of generalization and
robustness in existing methods, we propose a fetal hypoxia diagnosis model based on the
vision transformer. Compared to the related works, the proposed method demonstrates
strong generalization, achieving 95.582% accuracy with a 97.222% AUC on the public
database, and 97.658% accuracy with an 79.910% AUC on the private database. Compared
to related works, our proposed model demonstrates the most balanced performance across
all metrics. Moreover, experiments conducted in various scenarios demonstrate our model’s
strong robustness.
Keywords: Cardiotocography; Fetal hypoxia; Feature fusion network; Vision transformer

Farhana Ahmed Chowdhury, Md Kamal Hosain, Md Sakib Bin Islam, Md Shafayet Hossain,
Promit Basak, Sakib Mahmud, M. Murugappan, Muhammad E.H. Chowdhury,
ECG waveform generation from radar signals: A deep learning perspective,
Computers in Biology and Medicine,
Volume 176,
2024,
108555,
ISSN 0010-4825,
[Link]
([Link]
Abstract: Cardiovascular diagnostics relies heavily on the ECG (ECG), which reveals significant
information about heart rhythm and function. Despite their significance, traditional ECG
measures employing electrodes have limitations. As a result of extended electrode
attachments, patients may experience skin irritation or pain, and motion artifacts may
interfere with signal accuracy. Additionally, ECG monitoring usually requires highly trained
professionals and specialized equipment, which increases the treatment's complexity and
cost. In critical care scenarios, such as continuous monitoring of hospitalized patients,
wearable sensors for collecting ECG data may be difficult to use. Although there are issues
with ECG, it remains a valuable tool for diagnosing and monitoring cardiac disorders due to
its non-invasive nature and the detailed information it provides about the heart. The goal of
this study is to present an innovative method for generating continuous ECG waveforms
from non-contact radar data by using Deep Learning. The method can eliminate the need for
invasive or wearable biosensors and expensive equipment to collect ECGs. In this paper, we
propose the MultiResLinkNet, a one-dimensional convolutional neural network (1D CNN)
model for generating ECG signals from radar waveforms. With the help of a publicly
accessible radar benchmark dataset, an end-to-end DL architecture is trained and assessed.
There are six ports of raw radar data in this dataset, along with ground truth physiological
signals collected from 30 participants in five distinct scenarios: Resting, Valsalva, Apnea, Tilt-
up, and Tilt-down. By using strong temporal and spectral measurements, we assessed our
proposed framework's ability to convert ECG data from Radar signals in three distinct
scenarios, namely Resting, Valsalva, and Apnea (RVA). ECG segmentation performed better
by MultiResLinkNet than by state-of-the-art networks in both combined and individual
cases. As a result of the simulations, the resting, valsalva, and RVA scenarios showed the
highest average temporal values, respectively: 66.09523 ± 19.33, 60.13625 ± 21.92, and
61.86265 ± 21.37. In addition, it exhibited the highest spectral correlation values
(82.4388 ± 18.42 (Resting), 77.05186 ± 23.26 (Valsalva), 74.65785 ± 23.17 (Apnea), and
79.96201 ± 20.82 (RVA)), along with minimal temporal and spectral errors in almost every
case. The qualitative evaluation revealed strong similarities between generated and actual
ECG waveforms. As a result of our method of forecasting ECG patterns from remote radar
data, we can monitor high-risk patients, especially those undergoing surgery.
Keywords: ECG; Raw radar data; MultiResLinkNet; CNN; Deep learning

Tao Wang, Ye Xu, Yu Qin, Xu Wang, Feifan Zheng, Wei Li,


Short-term PV forecasting of multiple scenarios based on multi-dimensional clustering and
hybrid transformer-BiLSTM with ECPO,
Energy,
Volume 334,
2025,
137654,
ISSN 0360-5442,
[Link]
([Link]
Abstract: Accurate and effective photovoltaic output forecasting is critical for the high-
proportion integration of solar power generation into the power grid. To address the issues
that existing prediction methods often ignore long-term dependencies in photovoltaic
power sequences and tend to produce suboptimal hyperparameters combinations, a hybrid
photovoltaic power prediction model composed of GIKM, ECPO, VMD and Transformer-
BiLSTM is established. Firstly, a brand-new integrated method for selecting multi-
dimensional similar days based on GRA and an improved K-medoids is proposed to identify
the strongly correlated historical days with meteorological conditions similar to those of the
predicted day. Secondly, an innovative combination of Logistic-Tent composite chaotic
mapping and Gaussian mutation is used to enhance the optimal performance of the CPO,
which overcomes the issues of numerous suboptimal solutions and premature convergence
to local optima that afflict traditional CPO algorithms. Thirdly, ECPO is used for the first time
to determine the optimal parameter combination of VMD method (i.e. number of
decomposed subsequences and penalty factor), which provides high-quality training
samples. Next, a novel hybrid forecasting model combining BiLSTM and Transformer is
proposed, where BiLSTM extracts temporal dynamics and Transformer captures the
interdependencies among multivariate energy-related time series, thereby enhancing the
generalization ability and robustness of this combined model during the power forecasting
process. Accurate prediction is achieved by optimizing the epoch, number of hidden unit and
learning rate of the Transformer-BiLSTM model based on ECPO. The performance of the
proposed method is evaluated through four sets of comparative experiments and three
evaluation metrics for two distinct PV stations located in two provinces (i.e. Yunnan and
Gansu), China. The proposed ensemble model significantly outperforms other baseline
models, with average MAE, RMSE, and MSE of 0.2328 MW, 0.2778 MW and 0.0809 MW2 in
Yunnan, respectively, and 0.7231 MW, 1.0186 MW and 1.0583 MW2 in Gansu, respectively.
Parallel cross-experiments at two photovoltaic stations with distinct meteorological
conditions and installed capacities further verify the proposed model's technical superiority
and robust performance, demonstrating strong potential for real-world application and
promotion.
Keywords: Short-term PV power forecasting; Similar day selection; Enhanced crested
porcupine optimization; Transformer-BiLSTM; Cross validation

Cries Avian, Jenq-Shiou Leu, Hang Song, Jun-ichi Takada, Nur Achmad Sulistyo Putro,
Muhammad Izzuddin Mahali, Setya Widyawan Prakosa,
RCTrans-Net: A spatiotemporal model for fast-time human detection behind walls using
ultrawideband radar,
Computers and Electrical Engineering,
Volume 120, Part C,
2024,
109873,
ISSN 0045-7906,
[Link]
([Link]
Abstract: Ultrawideband (UWB) radar systems are becoming increasingly popular for
detecting human presence, even through walls. Recent advancements in signal processing
use deep learning techniques, which are known for their accuracy. While earlier methods
focused on spatial information using Convolutional Neural Networks (CNNs), newer research
highlights the importance of temporal information, such as how data peaks shift over time.
This study introduces RCTrans-Net, a deep-learning architecture that combines RCNet (a
Residual CNN) for spatial features with TransNet (a Transformer) for temporal features. This
fusion improves human presence classification in fast-time signal processing. Tested under
various conditions—different materials, body orientations, ranges, and radar heights—
RCTrans-Net achieved high performance with F1-scores of 0.997±0.000 for static,
0.967±0.004 for dynamic, and 0.978±0.001 for combined scenarios. The architecture
outperforms previous methods and offers real-time processing with an inference time of
about one millisecond.
Keywords: Human presence behind the wall; Residual network; Spatiotemporal'
Transformer; Ultrawideband radar system

Do-Soo Kwon, Chungkuk Jin, MooHyun Kim, Sung-Jae Kim,


Transformers and neural networks for estimation of parameters of multi-directional waves
from rich statistics of FPSO motion signals,
Applied Ocean Research,
Volume 166,
2026,
104927,
ISSN 0141-1187,
[Link]
([Link]
Abstract: This study presents a machine learning (ML) framework for inverse estimation of
parameters of multi-directional waves from moored FPSO (floating production storage
offloading) motion-sensor synthetic data. A time-domain hull/mooring/riser coupled
dynamics numerical simulation program was used to generate realistic vessel-response time
series under varying wind–wave–current conditions, from which motion-statistical features
up to 122 were extracted. These statistical features served as inputs to two different ML
models, artificial neural networks (ANNs) and transformer-based ensemble (TBE) model.
Then, different combinations of motion-statistical features were selected as inputs to the ML
models to estimate key spectral parameters of multi-directional-waves including significant
wave height, peak period, mean wave direction, spectral enhancement (peakedness) factor,
and directional spreading, and the results were systematically compared. A more advanced
ML method, the transformer architecture, combined with an ensemble approach,
demonstrated improved robustness and generality across complex sea states. The systematic
comparisons of ML performances with measured wave parameters versus artificially-
generated wave parameters provided insights into how the hidden intrinsic correlations
among wave parameters can improve the overall performance of the ML-based inverse wave
estimation. The results highlight the potential of FPSOs as near-real-time wave-sensing
devices. The estimated parameters can serve as crucial inputs for optimizing dynamic
positioning (DP) systems and other active controls, as well as for digital-twin and smart-
ship/platform applications, reducing reliance on external measurement systems.
Keywords: Directional wave spectrum; Inverse wave estimation; Artificial neural network;
Transformer; Synthetic data; Digital Twin; Richer statistical motion inputs

Lele Qu, Jinpeng Tao, Tianhong Yang,


Human sleep posture recognition and vital sign monitoring method using FWCW radar,
Measurement,
Volume 258, Part D,
2026,
119349,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Radar technology has the great potential for recognizing human sleep postures and
tracking respiration rate (RR) and heart rate (HR) during sleep. In this paper, we propose a
novel integrated framework that simultaneously performs sleep posture recognition and
vital sign monitoring using dual frequency-modulated continuous-wave (FMCW) radar
modules. A distinctive aspect of the proposed framework is its dynamic radar switching
strategy, where sleep posture is first classified by applying histogram of oriented gradients
(HOG) and support vector machine (SVM) to the merged range-time map (RTM). Based on
the identified posture, either the top or side radar is selectively engaged to capture
physiological signals from the most effective orientation. To obtain the robust and accurate
estimation of RR and HR, we further propose the GA-MVMD algorithm that integrates
genetic algorithm (GA) and multivariate variational mode decomposition (MVMD) to jointly
decompose chest wall displacement signals across multiple range bins. Experimental results
demonstrate that the proposed method can effectively enhance the recognition accuracy
with an average accuracy rate of 96.7 % for four typical sleep postures and the proposed GA-
MVMD algorithm can provide more accurate RR and HR estimation results.
Keywords: Frequency modulated continuous wave (FMCW) radar; Sleep posture recognition;
Vital sign monitoring

Son Minh Nguyen, Duc Viet Le, Paul J.M. Havinga,


Seeing the world from its words: All-embracing Transformers for fingerprint-based indoor
localization,
Pervasive and Mobile Computing,
Volume 100,
2024,
101912,
ISSN 1574-1192,
[Link]
([Link]
Abstract: In this paper, we present all-embracing Transformers (AaTs) that are capable of
deftly manipulating attention mechanism for Received Signal Strength (RSS) fingerprints in
order to invigorate localizing performance. Since most machine learning models applied to
the RSS modality do not possess any attention mechanism, they can merely capture
superficial representations. Moreover, compared to textual and visual modalities, the RSS
modality is inherently notorious for its sensitivity to environmental dynamics. Such
adversities inhibit their access to subtle but distinct representations that characterize the
corresponding location, ultimately resulting in significant degradation in the testing phase. In
contrast, a major appeal of AaTs is the ability to focus exclusively on relevant anchors in RSS
sequences, allowing full rein to the exploitation of subtle and distinct representations for
specific locations. This also facilitates disregarding redundant clues formed by noisy ambient
conditions, thus enhancing accuracy in localization. Apart from that, explicitly resolving the
representation collapse (i.e., none-informative or homogeneous features, and gradient
vanishing) can further invigorate the self-attention process in transformer blocks, by which
subtle but distinct representations to specific locations are radically captured with ease. For
that purpose, we first enhance our proposed model with two sub-constraints, namely
covariance and variance losses at the Anchor2Vec. The proposed constraints are
automatically mediated with the primary task towards a novel multi-task learning manner. In
an advanced manner, we present further the ultimate in design with a few simple tweaks
carefully crafted for transformer encoder blocks. This effort aims to promote representation
augmentation via stabilizing the inflow of gradients to these blocks. Thus, the problems of
representation collapse in regular Transformers can be tackled. To evaluate our AaTs, we
compare the models with the state-of-the-art (SoTA) methods on three benchmark indoor
localization datasets. The experimental results confirm our hypothesis and show that our
proposed models could deliver much higher and more stable accuracy.
Keywords: RSS fingerprints; Indoor localization; Deep learning; Transformers

Ziwei Zhang, Mengtao Zhu, Yunjie Li, Yan Li, Shafei Wang,
Joint recognition and parameter estimation of cognitive radar work modes with LSTM-
transformer,
Digital Signal Processing,
Volume 140,
2023,
104081,
ISSN 1051-2004,
[Link]
([Link]
Abstract: The recent developed cognitive radars can implement flexible work modes with
programmable modulation types and optimized modulating values for each mode definition
parameter. Automatic analysis of these work modes is a significant challenge for modern
electromagnetic reconnaissance receivers. In this paper, a Multi-Output Multi-Structure
(MOMS) learning-based processing framework is proposed for Joint inter-pulse automatic
Modulation Recognition and Parameter Estimation (JMRPE-MOMS). We propose a label
construction method as a feature interpretation method of the network to facilitate MOMS
learning and utilize the correlations between labels for performance gain. Moreover, an
LSTM-Transformer is designed to mine deep time-series characteristics, which can model
local and global relationships and reduce quantization loss. The proposed framework can
perform joint modulation recognition and parameter estimation (JMRPE) tasks
simultaneously with flexible output structures including scalar output and vector output
with fixed or variable sizes. Extensive simulations are performed based on the simulated
radar work modes defined with pulse repetition interval (PRI) sequences. The simulation
results validate the effectiveness and superiority of the proposed method especially under
non-ideal electromagnetic environments.
Keywords: Radar work mode; Automatic modulation recognition; Modulation parameter
estimation; Multi-output learning; Transformer

Wentao He, Jianfeng Ren, Ruibin Bai, Xudong Jiang,


Radar gait recognition using Dual-branch Swin Transformer with Asymmetric Attention
Fusion,
Pattern Recognition,
Volume 159,
2025,
111101,
ISSN 0031-3203,
[Link]
([Link]
Abstract: Video-based gait recognition suffers from potential privacy issues and performance
degradation due to dim environments, partial occlusions, or camera view changes. Radar has
recently become increasingly popular and overcome various challenges presented by vision
sensors. To capture tiny differences in radar gait signatures of different people, a dual-branch
Swin Transformer is proposed, where one branch captures the time variations of the radar
micro-Doppler signature and the other captures the repetitive frequency patterns in the
spectrogram. Unlike natural images where objects can be translated, rotated, or scaled, the
spatial coordinates of spectrograms and CVDs have unique physical meanings, and there is
no affine transformation for radar targets in these synthetic images. The patch splitting
mechanism in Vision Transformer makes it ideal to extract discriminant information from
patches, and learn the attentive information across patches, as each patch carries some
unique physical properties of radar targets. Swin Transformer consists of a set of cascaded
Swin blocks to extract semantic features from shallow to deep representations, further
improving the classification performance. Lastly, to highlight the branch with larger
discriminant power, an Asymmetric Attention Fusion is proposed to optimally fuse the
discriminant features from the two branches. To enrich the research on radar gait
recognition, a large-scale NTU-RGR dataset is constructed, containing 45,768 radar frames of
98 subjects. The proposed method is evaluated on the NTU-RGR dataset and the MMRGait-
1.0 database. It consistently and significantly outperforms all the compared methods on
both datasets. The codes are available at: [Link]
Keywords: Micro-Doppler signature; Radar gait recognition; Spectrogram; Cadence velocity
diagram; Asymmetric Attention Fusion

Huan Wang, Zi-Hao Ren, Junyu Qi,


TFD-Trans: Time-frequency hierarchical decomposition transformer for mechanical fault
diagnosis,
Advanced Engineering Informatics,
Volume 71, Part B,
2026,
104303,
ISSN 1474-0346,
[Link]
([Link]
Abstract: Rolling bearings are critical components in rotating machinery, and their health
directly affects the safety and stability of industrial equipment. To improve fault
identification under strong noise and complex operating conditions, this paper proposes a
Time–Frequency Decomposition Transformer (TFD-Trans) for robust rolling bearing fault
diagnosis. TFD-Trans integrates wavelet transform and Fourier transform within a signal-
processing-informed deep architecture. A wavelet-driven frequency decomposition module
hierarchically partitions vibration signals into multiple sub-bands, enabling fine-grained
multi-scale feature extraction. A Fourier-driven autocorrelation encoding module then
models global periodic dependencies in the frequency domain and enhances fault-sensitive
representations. Extensive experiments on two real-world rolling bearing datasets show that
TFD-Trans consistently achieves higher accuracy and stronger noise robustness than existing
mainstream methods across multiple operating and signal-to-noise conditions.
Keywords: Mechanical Fault Diagnosis; Transformer; Deep Learning; Health monitoring

Abid Hussain, Yueshan Chen, Arif Ullah, Sihai Zhang,


WiSigPro: Transformer for elevating CSI-based human activity recognition through attention
mechanisms,
Expert Systems with Applications,
Volume 258,
2024,
124976,
ISSN 0957-4174,
[Link]
([Link]
Abstract: The utilization of Channel State Information (CSI) in Wi-Fi-based passive sensing
has become popular due to its cost-effectiveness and broad applicability. This technology is
highly valued for its ability to gather data without requiring active user interaction, making it
versatile in various contexts. However, existing methods face significant challenges, including
low sensing accuracy, high processing requirements, and system instability. To address these
issues, we introduce the WiSigPro Transformer, an advanced CSI-based Wi-Fi passive sensing
model specifically designed for Human Activity Recognition. This model employs multi-head
attention mechanisms and positional encoding to effectively capture complex
spatiotemporal patterns, thereby enhancing both robustness and accuracy. Our approach
integrates advanced signal processing techniques to improve signal quality and feature
extraction. These techniques include wavelet denoising to reduce noise, median filtering to
smooth the signal, and Power Spectral Density analysis using Welch’s method to capture
frequency domain features. Additionally, we use normalization to standardize amplitude and
phase data, and feature engineering methods to extract comprehensive signal
characteristics, such as skewness and kurtosis. To address data imbalance, we apply the
Synthetic Minority Over-sampling Technique and data augmentation strategies, ensuring
balanced representation and improved model generalization. Through comprehensive
simulations, the WiSigPro Transformer demonstrates superior performance across key
metrics, including recognition accuracy, precision, recall, and F1-score. Achieving an
impressive 98% accuracy, it outperforms conventional neural networks such as CNN, RNN,
BiLSTM, LSTM, and ABLSTM. These results underscore the transformative potential of the
WiSigPro Transformer in Wi-Fi-based passive sensing and activity recognition applications,
making it a powerful tool for accurately capturing and analyzing spatiotemporal data.
Keywords: Channel state information; Passive WiFi sensing; Human activity recognition; IoT
monitoring systems; Deep learning; Multi-head attention mechanisms; Positional encoding

Yuanzhi Su, Huiying Cynthia Hou, Chun Zhao, Zhuojun Nan,


PoseGraphNet: Pose prior and graph structure for 3D human pose estimation using
mmWave radar,
Measurement,
Volume 257, Part C,
2026,
118851,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Human pose estimation (HPE) is a crucial task in computer vision with extensive
applications in healthcare, surveillance, and human–computer interaction. Traditional HPE
research primarily utilizes RGB cameras, which may suffer from poor performance under
varying lighting conditions and raise privacy concerns. Recently, millimeter-wave (mmWave)
radar technology has emerged as a promising alternative, providing a non-invasive and
privacy-preserving solution for HPE. However, the progress in mmWave-based HPE is
hindered by the limited availability of high-quality datasets that encompass a diverse range
of poses and provide accurate data annotations. Current mmWave-based datasets for HPE
often feature only basic poses or rely on imprecise annotations, typically derived from pre-
trained image-based HPE models using synchronized RGB images, which can limit the
potential of derived models. This study introduces a pioneering approach to HPE by
synergizing wearable motion capture sensors with mmWave radar technology to create a
comprehensive and precise dataset tailored for enhancing HPE with mmWave radar.
Leveraging this dataset, we develop an innovative deep learning framework specifically
designed to explore the unique properties of radar signals for HPE. The performance of our
proposed model is evaluated and compared with several well-known deep learning models.
Extensive experimental results affirm the robustness of the dataset, establishing it as a
rigorous benchmark for mmWave radar-based HPE. The proposed methodology
demonstrates exceptional accuracy in estimating human poses from radar data, setting the
stage for its application in environments where privacy and complexity are critical concerns.
Keywords: Human pose estimation; 3D point cloud; Millimeter radar; Deep learning

Binyue Cao, Mi He, Meiyun Zhao, Qinwen Ping, Chang He, Xiangyu Zhou, Yushun Gong,
Non-contact detection of electrocardiogram fiducial points with millimeter-wave radar,
Biomedical Signal Processing and Control,
Volume 111,
2026,
108273,
ISSN 1746-8094,
[Link]
([Link]
Abstract: Purpose:
Accurate detection of fiducial points in electrocardiograms (ECG) is crucial for diagnosing
heart diseases. However, traditional ECG devices require direct contact with the skin, which
can cause discomfort for patients. Therefore, there is an increasing need for non-contact
detection of ECG fiducial points, while research in this area remains limited.
Methods:
In response to this need, we proposed a millimeter-wave radar-based method for non-
contact detection of ECG fiducial points. Initially, fiducial points from synchronously acquired
ECG signals were annotated to serve as labels for network training. Subsequently, a radar
cardiogram (RCG) was derived from radar radio frequency signals using a second-order
differentiator. To enhance the detection process, we developed a one-dimensional semantic
segmentation model by integrating a deep residual shrinkage network block (DRSN) into the
U-Net architecture, which we term DRSN-Unet.
Results:
The results from the testing dataset demonstrate that the proposed method accurately
detects the five fiducial points, achieving an average precision of 0.986, an average recall of
0.986, and an average F1-score of 0.986. Notably, the network’s parameter count is only
2.042 Mega, with a computation amount of 0.086 GFLOPs.
Conclusion:
This method effectively detects the onsets of the P-wave and QRS complex, the peak of the
R-wave, and the offsets of the QRS complex and T-wave without direct contact with the
human body, laying a solid foundation for the automated diagnosis of complex cardiac
diseases.
Keywords: Millimeter-wave radar; Non-contact; Electrocardiogram; Fiducial point detection;
Semantic segmentation

Lindong Wang, Hongya Tuo, Yu Yuan, Henry Leung, Zhongliang Jing,


RCMixer: Radar-camera fusion based on vision transformer for robust object detection,
Journal of Visual Communication and Image Representation,
Volume 107,
2025,
104367,
ISSN 1047-3203,
[Link]
([Link]
Abstract: In real-world object detection applications, the camera would be affected by poor
lighting conditions, resulting in a deteriorate performance. Millimeter-wave radar and
camera have complementary advantages, radar point cloud can help detecting small objects
under low light. In this study, we focus on feature-level fusion and propose a novel end-to-
end detection network RCMixer. RCMixer mainly includes depth pillar expansion(DPE),
hierarchical vision transformer and radar spatial attention (RSA) module. DPE enhances
radar projection image according to perspective principle and invariance assumption of
adjacent depth; The hierarchical vision transformer backbone alternates the feature
extraction of spatial dimension and channel dimension; RSA extracts the radar attention,
then it fuses radar and camera features at the late stage. The experiment results on
nuScenes dataset show that the accuracy of RCMixer exceeds all comparison networks and
its detection ability of small objects in dark light is better than the camera-only method. In
addition, the ablation study demonstrates the effectiveness of our method.
Keywords: Sensor fusion; Object detection; Feature level fusion; Vision transformer; Neural
network

Yongqiang Cui, Yiyang Zhang, Di Bai, Yi Diao, Yulei Wang,


3D map and mmWave radar-based self-localization for UAVs in GNSS-denied environments,
Vehicular Communications,
Volume 57,
2026,
100986,
ISSN 2214-2096,
[Link]
([Link]
Abstract: Reliable self-localization of unmanned aerial vehicles (UAVs) in dense urban
environments remains a major challenge due to the frequent unavailability or degradation of
Global Navigation Satellite Systems (GNSS) and other radio signals. This paper presents a
robust and cost-effective method for UAV self-localization by using vision and millimeter-
wave (mmWave) radar data in GNSS-denied environments. The approach generates an initial
dense point cloud through depth estimation and semantic segmentation, which is then
geometrically refined using sparse mmWave radar point cloud. A semantic-guided clustering
method is applied to the mmWave radar point cloud to remove noise and extract key
structural elements such as walls, which are later fused with vision-based depth information.
For positioning, image matching algorithm provides coarse localization, followed by fine
registration that leverages geometric features of windows to enhance precision.
Experimental results demonstrate that the proposed method can achieve self-localization
accuracy within 0.4 m, while maintaining low system complexity and deployment cost,
offering a practical solution for UAV self-localization in GNSS-denied urban scenarios.
Keywords: Millimeter-wave (mmWave) radar; Data fusion; 3D reconstruction; Point cloud
registration; Self-localization

Yueming Su, Qiusheng Lian, Dan Zhang, Baoshun Shi,


Transformer based Douglas-Rachford unrolling network for compressed sensing,
Signal Processing: Image Communication,
Volume 127,
2024,
117153,
ISSN 0923-5965,
[Link]
([Link]
Abstract: Compressed sensing (CS) with the binary sampling matrix is hardware-friendly and
memory-saving in the signal processing field. Existing Convolutional Neural Network (CNN)-
based CS methods show potential restrictions in exploiting non-local similarity and lack
interpretability. In parallel, the emerging Transformer architecture performs well in
modelling long-range correlations. To further improve the CS reconstruction quality from
highly under-sampled CS measurements, a Transformer based deep unrolling reconstruction
network abbreviated as DR-TransNet is proposed, whose design is inspired by the traditional
iterative Douglas-Rachford algorithm. It combines the merits of structure insights of
optimization-based methods and the speed of the network-based ones. Therein, a U-type
Transformer based proximal sub-network is elaborated to reconstruct images in the wavelet
domain and the spatial domain as an auxiliary mode, which aims to explore local informative
details and global long-term interaction of the images. Specially, a flexible single model is
trained to address the CS reconstruction with different binary CS sampling ratios. Compared
with the state-of-the-art CS reconstruction methods with the binary sampling matrix, the
proposed method can achieve appealing improvements in terms of Peak Signal to Noise
Ratio (PSNR), Structural Similarity Index Measure (SSIM) and visual metrics. Codes are
available at [Link]
Keywords: Compressed sensing; Transformer; Binary sampling; Douglas-Rachford algorithm;
Deep learning

Jiayuan Liang, Yu Cheng, Jiafeng He,


Transformer-based flexible sampling ratio compressed ghost imaging,
Engineering Analysis with Boundary Elements,
Volume 170,
2025,
106050,
ISSN 0955-7997,
[Link]
([Link]
Abstract: Recently, deep learning has been tried to improve the efficiency of compressed
ghost imaging. However, these current learning-based ghost imaging methods have to
modify and retrain the learning model to cope with different sampling ratios. This will
consume a lot of computing resources and energy. In this paper, we propose a deep
learning-based compressed ghost imaging method that can adapt to arbitrary sampling
ratios without tailoring and retraining model. By simultaneously optimizing the weights of
both the speckle patterns and the transformer model, we achieve a network for ghost
imaging at arbitrary sampling ratios. The feasibility and effectiveness of the proposed
method were validated through numerical simulations. The results indicate that the
proposed method, requiring only a single training session, is capable of reconstructing high-
quality images under varying sampling ratios. Furthermore, the performance of the
proposed method surpasses that of currently widely employed deep learning ghost imaging
methods. At a sampling ratio of 5%, the proposed method achieves an increase of 1.87 dB in
Peak Signal-to-Noise Ratio (PSNR) and 0.171 in Structural Similarity Index (SSIM).
Keywords: Compressed ghost imaging; Deep learning; Transformer; Ghost imaging

Yanwen Bai, Jibin Zheng, Hanxing Shao, Hongwei Liu,


A dual-driven hybrid tracking architecture for radar targets based on innovation,
Information Fusion,
Volume 129,
2026,
104056,
ISSN 1566-2535,
[Link]
([Link]
Abstract: Targets such as hypersonic missiles and stealth aircraft are characterized by
complex motion patterns, strong maneuverability, and anomalous radar measurement
statistics. Although model-driven radar target tracking methods offer physical
interpretability, they suffer from their dependence on explicit prior assumptions. Data-
driven methods can theoretically approximate arbitrarily complex motions through
nonlinear mappings, but suffer from poor interpretability, vulnerability to noise during
feature extraction, and loss of low-frequency maneuvering features due to sample
imbalance. Therefore, this paper proposes a Dual-Driven Hybrid Tracking Architecture Based
on Innovation (DDHTA), which fuses the advantages of both model-driven and data-driven
approaches. First, a model-driven approach is adopted for basic state estimation, and a Dual
Condition Judgment Adjustment (DCJA) method is proposed to adaptively adjust the
measurement error variance, thereby providing a high-quality baseline estimate for the
data-driven layer and reducing the interference of anomalous noise on feature extraction.
Further, in the data-driven layer, a Dual-Scale Temporal Network (DSTNet) is designed. By
learning the mapping from the innovation to the estimation errors, it combines the
strengths of causal dilated convolution and multi-head self-attention to provide dynamic
compensation, which corrects the estimation errors of the model-driven method. Numerical
simulation results demonstrate that the proposed method enhances the algorithm’s ability
to handle target maneuvers in complex environments, achieving higher tracking accuracy
and robustness.
Keywords: Radar tracking; Deep learning; Self-attention mechanism; Causal convolution

Li Qiusheng, Zhu Huajuan,


Target classification with low-resolution radars based on cyclic bispectrum and improved
ACGAN,
Measurement,
Volume 259, Part B,
2026,
119715,
ISSN 0263-2241,
[Link]
([Link]
Abstract: To address the challenges of insufficient generalization and high noise sensitivity in
low-resolution radar target recognition under limited-sample conditions, this paper
proposes a joint optimization framework integrating cyclic bispectral analysis and an
improved Auxiliary Classifier Generative Adversarial Network (ACGAN). First, a third-order
cyclic cumulant spectral model is designed to extract modulation-specific signatures of
aircraft targets in the cyclostationary domain, effectively suppressing both Gaussian and
non-Gaussian noise while preserving discriminative features that are robust to low SNR
conditions (maintaining 92.7 % accuracy at 0 dB). Second, an enhanced ACGAN architecture
is developed by incorporating self-attention mechanisms and Wasserstein distance
optimization with gradient penalty, with spectral normalization and dynamic gradient
penalties introduced to stabilize training dynamics and improve synthetic sample fidelity.
Extensive experiments on a real-world dataset collected by a certain Chinese-made VHF-
band radar demonstrate that the proposed method achieves state-of-the-art performance,
with average recognition accuracies of 98.46 % and 98.52 % for approaching and departing
targets in complex noise environments, respectively, alongside a Kappa coefficient exceeding
0.97. Comprehensive comparisons with traditional methods (wavelet, HOS, FrFT), GAN
variants (WGAN-GP, AFGAN + ResNet, Diffusion-GAN), VAE-based approaches, few-shot
learning models (ProtoNet, MatchingNet), and a modern Vision Transformer (ViT) baseline
show consistent improvements of 1.38–12.19 % in accuracy. Ablation studies validate the
contributions of key components, where the self-attention module and Wasserstein
optimization improve accuracy by 1.27 % and 0.97 %, respectively. Furthermore, embedded
platform tests confirm the framework’s feasibility for real-time deployment (inference
time < 15 ms/sample), offering a robust solution for resource-constrained radar systems.
This work highlights the efficacy of unifying physics-inspired feature extraction with
stabilized deep generative models for practical radar recognition.
Keywords: Radar target recognition; Cyclic bispectral analysis; Auxiliary classifier generative
adversarial network (ACGAN); Self-attention mechanism; Few-shot learning

Wentao Zhao, Guoxiang Tong,


A deep learning based heart rate estimation method for millimeter wave radar,
Measurement,
Volume 255,
2025,
117923,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Contact-based vital sign detection technology has been widely used in the medical
field. However, contact-based devices may cause discomfort to users and suffer from user
dependency issues. Frequency Modulated Continuous Wave (FMCW) millimeter-wave radar
provides an efficient and accurate solution for heart rate and respiratory rate monitoring.
Nevertheless, due to the small amplitude of heartbeat micro-motion signals, they are
susceptible to noise interference such as respiratory harmonics, making accurate
measurement challenging. Aiming at the problem that traditional methods are difficult to
adapt to different environmental noises, we propose a heart rate estimation method based
on a convolutional neural network. The heart rate estimation accuracy is significantly
improved by recognizing the phase change patterns in radar signals. We reduce the
interference of respiratory harmonics on the heart rate micromotion signal by decomposing
the extracted phase signal into multiple frequency components using the Empirical Wavelet
Transform (EWT) algorithm. The proposed deep learning model is used by means of depth
convolution in order to balance the model size and accuracy. Additionally, we introduce
time–frequency channels into the convolutional neural network to further enhance its
feature extraction capability. Comparisons with various related works on heart rate
distribution, sensing range, and subjects demonstrate the lightweight nature and higher
accuracy of the proposed model. Extensive experiments are conducted on a dataset based
on the TI AWR1642BOOST radar. The results show that the proposed method achieves
outstanding performance, with an accuracy of ±5 BPM in heart rate monitoring.
Keywords: Deep learning; Heart rate (HR); Millimeter-wave (mmW) radar; Vital signs
monitoring; Moving target indication (MTI)

Philipp Reitz, Tobias Veihelmann, Norman Franchi, Maximilian Lübke,


Dual radar vision: A feature fusion approach for advanced object detection in IoT radar
networks,
Machine Learning with Applications,
Volume 21,
2025,
100703,
ISSN 2666-8270,
[Link]
([Link]
Abstract: 60GHz radar technology is one of the most promising movement detector
solutions for Internet of Things (IoT) applications. However, challenges remain in accurately
classifying different objects and detecting small objects in a multi-target scenario. This work
investigates whether sensor fusion between multiple radars can enhance object detection
and classification performance. A one-stage detection architecture, designed based on the
features of the latest YOLO generations, is used to perform fusion based on range-Doppler
(RD) maps of two non-coherent spatially separated radars. A complete physical 3D
propagation simulation using ray tracing evaluates the fusion methods. This approach
enables precise ground truth, as all unprocessed signal components are known, and
guarantees a consistent, error-free reference. Results demonstrate that dynamic, attention-
based fusion significantly improves detection and classification compared to static fusion in
homogeneous and heterogeneous radar setups.
Keywords: Data fusion; Deep learning; FMCW radar; IoT; Radar networks; Ray tracing; YOLO
Saeed Sotoudeh, Mohammad Sadegh Ayubirad, Francesco Benedetto, Fabio Tosti,
Ground-Based Interferometric Radar for Bridge Dynamic System Identification: A Multi-
Domain Operational Modal Analysis and Integrated Method Selection Framework,
NDT & E International,
2026,
103649,
ISSN 0963-8695,
[Link]
([Link]
Abstract: This study proposes a cohesive multi-domain analytical framework for Ground-
Based Interferometric Radar (GBIR) dynamic system identification (DSI), designed to
overcome inherent spatial resolution limits and range-bin averaging in non-contact bridge
monitoring. While GBIR enables high-precision displacement monitoring, the extraction of
reliable modal parameters under operational conditions, independent of auxiliary contact
sensors, remains a significant technical bottleneck. The proposed framework addresses this
by integrating seven systematically derived advanced computational methods into a multi-
domain operational modal analysis (OMA) and integrated method-selection pipeline. This
architecture is structured around three analytical pillars: (i) spatial-multivariate
identification, (ii) frequency-domain benchmarking, and (iii) non-stationary time-frequency
tracking. The framework was validated using field data from the Dong-Yi cable-stayed bridge
in South Korea, and internally cross-validated by verifying agreement in modal estimates
across the three analytical pillars. By utilising the systematic agreement across the three
pillars as a mechanism for internal consistency and methodological reliability, the
methodology demonstrates that multi-domain approaches, specifically Covariance-driven
Stochastic Subspace Identification (Cov-SSI) and Enhanced Frequency Domain
Decomposition (EFDD), exploit spatial correlations across radar range-bins to achieve
superior modal separation. Results show that the framework successfully isolates three
dominant natural frequencies (0.33 Hz, 0.46 Hz, and 0.63 Hz) and resolves significant
frequency discrepancies associated with univariate processing. This study demonstrates that
an integrated, multi-domain analytical strategy provides the necessary methodological
rigour to establish GBIR as a viable independent solution for the dynamic characterisation
and long-term health assessment of complex bridge infrastructure.
Keywords: Ground-Based Interferometric Radar (GBIR); Multi-Domain Operational Modal
Analysis; Closely Spaced Modes; Bridge Dynamic System Identification (DSI); Integrated
Method Selection; Cable-Stayed Bridge; Internal Consistency Validation; Structural Health
Monitoring (SHM); Remote Sensing

Li Zha, Chen Gong, Kunfeng Lv,


Real-time localization and navigation method for autonomous vehicles based on multi-
modal data fusion by integrating memory transformer and DDQN,
Image and Vision Computing,
Volume 156,
2025,
105484,
ISSN 0262-8856,
[Link]
([Link]
Abstract: In the field of autonomous driving, real-time localization and navigation are the
core technologies that ensure vehicle safety and precise operation. With advancements in
sensor technology and computing power, multi-modal data fusion has become a key method
for enhancing the environmental perception capabilities of autonomous vehicles. This study
aims to explore a novel visual-language navigation technology to achieve precise navigation
of autonomous cars in complex environments. By integrating information from radar, sonar,
5G networks, Wi-Fi, Bluetooth, and a 360-degree visual information collection device
mounted on the vehicle's roof, the model fully exploits rich multi-source data. The model
uses the Memory Transformer for efficient data encoding and a data fusion strategy with a
self-attention network, ensuring a balance between feature integrity and algorithm real-time
performance. Furthermore, the encoded data is input into a DDQN vehicle navigation
algorithm based on an automatically growing environmental target knowledge graph and
large-scale scene maps, enabling continuous learning and optimization in real-world
environments. Comparative experiments show that the proposed model outperforms
existing SOTA models, particularly in terms of macro-spatial reference from large-scale scene
maps, background knowledge support from the automatically growing knowledge graph,
and the experience-optimized navigation strategies of the DDQN algorithm. In the
comparative experiments with the SOTA models, the proposed model achieved scores of
3.99, 0.65, 0.67, 0.65, 0.63, and 0.63 on the six metrics NE, SR, OSR, SPL, CLS, and DTW,
respectively. All of these results significantly enhance the intelligent positioning and
navigation capabilities of autonomous driving vehicles.
Keywords: DDQN; Memory transformer; Self-attention network; Autonomous vehicles;
Knowledge graph; Navigation

Jiale Chang, Yanhui Wang, Siya Mi, Yu Zhang,


MVL-Tra: Multi-view LFM signal source classification using Transformer,
Computers and Electrical Engineering,
Volume 111, Part B,
2023,
108967,
ISSN 0045-7906,
[Link]
([Link]
Abstract: Classifying the sources of linear frequency modulation (LFM) signals has practical
significance, as these signals are widely applied in various scenarios. However, the signals
emitted by the same configured sources exhibit insignificant differences and may suffer
severe degradation in low signal-to-noise ratio (SNR) conditions. Previous research failed to
achieve satisfactory classification results in these two conditions. To address such challenges,
we propose a signal source classification method for LFM signals using multi-view
representations, based on the Transformer network, named MVL-Tra. Multi-view learning
improves robustness by utilizing multiple feature sets and multi-dimensional analysis on
these features. Transformer, a emerging deep neural network, enhances the distinctiveness
of features between signals through capturing long-range dependencies when dealing with
large-scale time-series data. In the experiment, we conducted classification on a dataset
comprising 15 sets of low-SNR LFM signals. Notably, Even under the challenging condition of
SNR = −5, our method achieved an accuracy of 93.78%. Th results confirm the effectiveness
of using multi-view representations and Transformer for LFM signal source classification.
Keywords: LFM signal source classification; Transformer; Multi-view learning; Deep neural
network; Multi-dimensional analysis

Zhigang Cheng, Zhizhou He, Peng Pan,


3D reconstruction of subsurface pipes and cavities using ground penetrating radar based on
deep learning,
NDT & E International,
Volume 158,
2026,
103579,
ISSN 0963-8695,
[Link]
([Link]
Abstract: Detecting subsurface pipes and cavities is important in urban infrastructure
management, but existing methods struggle to accurately reconstruct the 3D shapes of deep
subsurface objects. This study pioneers a new paradigm for this task by reformulating the ill-
posed permittivity regression problem as a 3D semantic segmentation problem. A novel
neural network, 3DReconNet, to predict the material type of each subsurface voxel from
ground penetrating radar (GPR) data was proposed. This approach leverages the intrinsic
relationship between material composition and reflected signal intensity to simultaneously
recover both geometry and material properties. A dataset of 3150 synthetic cases was
generated using full-scale simulation models and a Markov model-based algorithm to
simulate irregular cavities. The 3DReconNet adopts a U-shaped architecture and
incorporates residual connections to reduce information loss. The network is trained using
the Dice Loss function regularized with total variation (TV) constraints, which enhances
geometric consistency and reconstruction accuracy. The proposed method was validated
using both simulated and experimental data, and the qualitative as well as quantitative
results confirmed its effectiveness, robustness, and generalizability.
Keywords: Ground penetrating radar (GPR); Subsurface pipes; Subsurface cavities; Deep
learning; 3D reconstruction

Shuchen Wang, Qizhi Xu, Shunpeng Zhu, Biao Wang,


Making transformer hear better: Adaptive feature enhancement based multi-level
supervised acoustic signal fault diagnosis,
Expert Systems with Applications,
Volume 264,
2025,
125736,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Acoustic signal fault diagnosis has been receiving increasing attention in the field of
engine health management due to its effectiveness and non-invasiveness. Despite the
progress made in fault diagnosis models, challenges still exist due to the complexity of
acoustic signals and environmental factors. (1) End-to-end deep networks for fault diagnosis
are at risk of underperformance or overfitting due to complex models and imbalanced data.
(2) The complex acoustic environment within the vehicle power compartment poses
obstacles to extracting subtle fault features. (3) Time–Frequency (TF) analysis has been
proven to be an effective tool for characterizing the nonlinear features of fault signals, but it
falls short in achieving ideal fidelity and resolution. To address these issues, an engine
acoustic signal fault diagnosis method based on multi-level supervised learning and time–
frequency transformation was proposed. First, adopting a multi-level supervised learning
paradigm decomposes the fault diagnosis task into three stages: feature enhancement, fault
detection, and fault identification, thereby incorporating additional experiential knowledge
to mitigate overfitting. Second, an adaptive fault feature band extraction algorithm based on
the fusion of multiple time–frequency analyses is proposed, specifically for extracting unique
features from different vehicle datasets. Finally, a frequency band attention module was
designed to focus on the frequency range most relevant to the characteristics of engine
fault. The proposed method was validated on various audio signal fault datasets, and the
results indicated its superior performance compared to other state-of-art fault detection and
identification methods.
Keywords: Time–Frequency analysis; Fault diagnosis; Multi-level supervised learning;
Acoustic signal

Hao Yang, Shirong Zhou, Liyan Liu, Zhong Zhou,


A fast and accurate detection model of internal defects in tunnel lining for ground
penetrating radar image data,
Advanced Engineering Informatics,
Volume 68, Part C,
2025,
103812,
ISSN 1474-0346,
[Link]
([Link]
Abstract: Defects in tunnel linings accelerate structural deterioration, reduce service life, and
pose serious safety risks. Existing algorithms for detecting defect signals in ground-
penetrating radar (GPR) images often struggle to balance accuracy and efficiency, with
limited capacity to extract meaningful features. To address these limitations, this paper
proposes a lightweight algorithm, MGD-DETR, for accurate recognition of internal tunnel
lining defects, using RT-DETR as the base model. First, a Multi-HGNet backbone feature
extraction network is introduced to reduce model size (MS) and enhance dynamic fusion and
interaction between feature layers, thereby improving feature extraction. Second, the
lightweight convolution module GSConv replaces standard convolution operations to reduce
the parameter count. Third, a dual attention module (DAM) is integrated to dynamically
adjust spatial and channel feature weights, improving the model’s generalization
performance. Five models—RT-DETR, YOLO-LD, YOLOv10, YOLOv11, and SSD—were used for
comparative evaluation. Experimental results show that MGD-DETR outperforms the other
models across all metrics, achieving a mean average precision (mAP) of 0.834, mean F1
score (mF1) of 0.818, MS of 26.9 M, and frames per second (FPS) of 91.2f/s, enabling fast
and accurate recognition of defect signals and facilitate subsequent deployment into tunnel
detection mobile devices.
Keywords: Tunnel engineering; Lining defects; Deep learning; Ground-penetrating radar
images
Xinrui Zhao, Zeng Liu, Qi Hu, Jianglong Sun, Xiaoyan Yang,
Significant wave height estimation and prediction from synthetic X-band radar data by
spatio-temporal deep neural networks,
Ocean Engineering,
Volume 339, Part 1,
2025,
122061,
ISSN 0029-8018,
[Link]
([Link]
Abstract: Estimation and prediction of real-time significant wave height (SWH) is a
fundamental requirement for the safety of offshore activities. This study utilizes a spatio-
temporal deep neural network model to effectively estimate and predict the SWH by
extracting key spatial and temporal features from synthetic X-band radar [Link] study
considers three deep neural networks: InceptionV3, ResNet and Vision Transformer (ViT) to
extract multi-scale spatial features from radar images for SWH estimation. Subsequently, a
gated recurrent unit (GRU) is employed on these spatial features to perform the time-series
SWH prediction. Irregular waves with sea state ranging from 4 to 6 were considered based
on the synthetic radar data. Results indicate that the ResNet model with deep residual
structure performs the best in both the estimation and prediction tasks, demonstrating
excellent generalization ability and adaptability to complex sea conditions. The ViT model
shows outstanding performance in scenarios without out-of-distribution data. While the
InceptionV3 model is inferior, it exhibits significant improvement for the SWH prediction
when the GRU is incorporated.
Keywords: Significant wave height(SWH); X-band marine radar; InceptionV3; ResNet; Vision
Transformer(ViT); Gated recurrent unit(GRU)

Jun Li, Yihui Wang, Qinghong Sheng, Zhaocong Wu, Bo Wang, Xiao Ling, Xiang Liu, Yang Du,
Fan Gao, Gustau Camps-Valls, Matthieu Molinier,
CloudRuler: Rule-based transformer for cloud removal in Landsat images,
Remote Sensing of Environment,
Volume 328,
2025,
114913,
ISSN 0034-4257,
[Link]
([Link]
Abstract: Clouds are a key factor influencing transmission of the radiance signal in optical
remote sensing images. For mapping or monitoring the Earth's surface, it is inevitable to
mask or remove clouds before applying optical remote sensing images. Nowadays, deep
learning (DL) based thin cloud removal methods far outperform traditional methods. Yet
these DL-based methods often overlook position information or the physical cloud model in
thermal bands. Moreover, most existing cloud physical models for cloud removal overlook
the down-transmittance of the cloud in optical bands and do not account for the radiance of
thermal bands. This work proposes a novel transformer network, CloudRuler, coupled with
three rules in remote sensing domain for cloud removal. The proposed CloudRuler can
distinguish the semantic meanings between similar features in different pixel positions by
utilizing the Half-Spherical Coordinate System, aggregating features from local neighborhood
windows with remote sensing mosaicking, and solving the parameters of the cloud physical
model without limitations. Experimental results on 20 paired Landsat 8 and 9 images
demonstrate that CloudRuler outperforms seven baseline methods, based on GAN, CNN,
and transformer, both visually and quantitatively. Ablation experiments demonstrate that
the proposed rule-based modules are highly effective in improving CloudRuler's
performance for thin cloud removal. This work demonstrates that the joint use of Landsat 8
and 9 images for cloud removal is effective, producing more reliable data for downstream
applications than methods that utilize only one satellite with a longer revisit period. For
future research of the field, the code and dataset for reproducing the reported results are
available on: [Link]
Keywords: Cloud removal; Transformer; Cloud physical model; Landsat imagery; Deep
learning

Jiale Ren, Hengyi Li, Aihui Wang, Kenshi Saho, Lin Meng,
Radar-based gait analysis by Transformer-liked network for dementia diagnosis,
Biomedical Signal Processing and Control,
Volume 91,
2024,
105986,
ISSN 1746-8094,
[Link]
([Link]
Abstract: Providing reliable diagnostic evidence to doctors while minimizing financial and
physical burdens on patients is a prominent focus of current research. Gait features show
potential as a clinical marker for dementia diagnosis. Radar is capable of efficient,
contactless collecting human motion. This paper proposes a novel radar-based gait analysis
strategy with a Transformer-liked network for dementia diagnosis. The gait data are collected
by Micro-Doppler radar. Then Welch’s power spectral density estimation is adopted to
obtain the frequency features while unifying and compressing data size. The network is
designed to explore the relationship between dementia and gait features. In the network,
1D convolution is crucial in extracting local features and encoding features into deeper
dimensions. The attention-based module, inspired by the encoder of the Transformer,
possesses an edge in capturing long-sequence dependencies. The dual-stage gating
mechanism enhances the discriminative power of the learned representations by fine-tuning
the weights of extracted features. To validate the effectiveness of the proposed strategy,
comparative experiments are performed with prevailing networks in both time and
frequency domains. Experimental results demonstrate the superiority of the frequency
domain processing method over the time domain processing method and fusion time-
frequency processing method. Notably, the proposed model outperforms others, achieving
the highest accuracy of 94.93% in frequency domain processing-based experiments — 5.91%
and 4.25% higher than the highest accuracies in time domain processing-based experiments
and fusion time-frequency processing-based experiments respectively. The overall findings
illustrate that our proposal can provide a reliable reference for dementia diagnosis
effectively.
Keywords: Dementia diagnosis; Gait analysis; Power spectral density estimation; Attention
mechanism; Convolutional neural network; Gating mechanism
Yangxiaoyue Liu, Yuan Tian, Ying Xin, Yizhuo Yang, Jiangyuan Zeng, Min Feng, Chunqiao Song,
Transformer-based soil moisture simulation for understanding future drying trend globally,
Journal of Hydrology,
Volume 665,
2026,
134709,
ISSN 0022-1694,
[Link]
([Link]
Abstract: As a crucial element of the terrestrial water cycle, multiple future scenario soil
moisture (SM) datasets are widely applied to investigating Earth surface processes using
ensemble averages. However, they may run the risk of vague variation trend resulted from
averaging multiple models, which are characterized by different land surface models on
hydrological process simulation. To improve spatiotemporal pattern reliability, this study
innovatively designs a Transformer SM Simulation Net (TSMSNet), to conduct global SM
simulation of SSP1-2.6, SSP2-4.5, and SSP5-8.5 during 2016–2099. Nine qualified future SM
datasets, along with their spatial distribution of error parameters, and geographic data are
selected as model inputs. The learning target is calculated through merging merits from Soil
Moisture Active Passive and European Centre for Medium-Range Weather Forecast
Reanalysis v5-Land SM. The TSMSNet SM (R = 0.68, ubRMSE = 0.045 m3/m3) achieves good
matching degree against in situ measurements compared to the Convolutional Long Short
Term Memory (CSMSNet) simulated SM (R = 0.65, ubRMSE = 0.047 m3/m3). The TSMSNet
SM could favorably match the long-term trend of learning target, which exhibits advantage
over ensemble averages and CSMSNet SM. TSMSNet SM presents an overwhelming drying
trend. The decline magnitude rises accompanied by SSP changing from sustainable pathway
to fossil-fueled development. In terms of land cover types, evident drying trends are found
in cropland and forest. SM shows faster descent rate at habitable areas than inhabitable
areas. This paper develops a reliable TSMSNet SM dataset, which is expected to be a
valuable reference for understanding future SM variations.
Keywords: Soil moisture; Simulation; Global scale; Transformer; Trend analysis

Muhammad Yasir, Liu Shanwei, Xu Mingming, Wan Jianhua, Shah Nazir, Qamar Ul Islam, Kinh
Bac Dang,
SwinYOLOv7: Robust ship detection in complex synthetic aperture radar images,
Applied Soft Computing,
Volume 160,
2024,
111704,
ISSN 1568-4946,
[Link]
([Link]
Abstract: Using satellite-based SAR (Synthetic Aperture Radar) imagery to detect and track
ships is a formidable challenge. However, accurate analysis is hampered by inherent
difficulties such as obscured edges, multiple targets, varying dimensions and complex
backgrounds. These factors contribute to a poor signal-to-noise ratio. Various artificial
intelligence models have been developed to improve these problems, especially with YOLO-
based models. To overcome the above challenges and achieve robust performance, this
study presents a model called SwinYOLOv7, a novel fusion of the YOLOv7 framework with
the Features Pyramid Network and the Swin Transformer using ground-breaking anchor-free
detection algorithms. This innovative method aims to increase the accuracy of vessel
detection while reducing the impact of background clutter. The proposed model improved
by YOLOv7 examined three different datasets in detail: SRSDD-v1.0, HRSID, and SSDD to
optimize the performance of the model. The training process was consistently verified using
superior recall, precision, and F1-score values, which can be easily compared with previous
studies. The results show that the model using the Swin Transformer attention mechanism
and using an image size of 640×640 achieves the highest accuracy of 96.59%. Alternative
attention mechanisms, including the Squeeze-and-Excitation Network (SEnet), the
Convolutional Block Attention Module (CBAM), Channel Attention (CA) and Efficient Channel
Attention (ECA), deliver poorer accuracy rates. The combination of YOLOv7 and Swin
Transformer yielded encouraging results that enabled the proposed model to outperform
the current benchmarking models. Therefore, the proposed model provides a compelling
solution to ensure accurate vessel identification in complex search and rescue scenarios.
Keywords: Robust Ship detection; YOLO; Swin transformer; Complex synthetic aperture radar
images

Wenxu Zhang, Kang Luo, Fuli Sun, Zhongkai Zhao, Yunxiao Fu, Feiran Liu,
Radar working mode recognition for small samples based on the DBA-CIB-IMP method,
Physical Communication,
Volume 72,
2025,
102718,
ISSN 1874-4907,
[Link]
([Link]
Abstract: To address the issue that the reliability of radar working mode recognition
decreases as detected radar pulses decrease, a Dynamic Time Warping (DTW) Barycenter
Averaging-Circularly Integrated Bispectrum-Informer Multilayer Perceptron (DBA-CIB-IMP)
recognition method is proposed. This method uses CIB feature extraction to reduce the
complexity of the input without losing the feature information, and computes the global
feature similarity by DTW Barycenter Averaging (DBA). Generating samples based on the
original data set and expanding the data by generating a supplementary database through a
weighted average algorithm. The supplementary database is then fused with the original
database to complete the radar working mode recognition work by intelligent network
model. Higher recognition accuracies are achieved in scenarios with a limited training
samples. Recognition accuracy exceeds 90% at 0 dB SNR with low time spent.
Keywords: Radar working mode recognition; Multi-functional radar; Informer; Circularly
integrated bispectrum

Xinyue Huang, Yongtao Ma, Zedong Yu, Haibo Zhao,


RCDformer: Transformer-based dense depth estimation by sparse radar and camera,
Neurocomputing,
Volume 589,
2024,
127668,
ISSN 0925-2312,
[Link]
([Link]
Abstract: Accurate depth cues are crucial for 3D perception tasks, and monocular depth
estimation networks are no longer sufficient for realistic scenarios. Currently, the most
effective approaches are to introduce depth information from other modalities into the
image. Radar has become a popular sensor for fusion with cameras due to its low price and
all-weather working characteristics. This paper aims to explore how to more effectively
integrate the heterogeneous data of radar point clouds and RGB images to improve the
performance of depth estimation. Most of the previous works have not fully exploited the
potential of integrating these two modalities, so we propose RCDformer, a novel network
based on the transformer architecture that fuses radar-camera for dense depth estimation.
Without reducing the receptive field, our approach can fully model the contextual
relationships between sensors to reduce the impact of radar noise on overall performance.
With the proposed Radar-guided Multi-scale Depth Fusion (RGDF) module, the prior spatial
information mapped by the Radar Feature Extractor (RFE) is embedded into a set of multi-
scale hierarchical features output by Image Feature Extractor (IFE) via the modified
deformable cross-attention, which aims to guide the depth prediction of images.
Furthermore, we discover that incorporating the Radar Cross Section (RCS) attribute as an
extended channel for the radar map is beneficial for dense depth estimation, which
improves the overall performance of our model. We evaluate the proposed method on the
nuScenes dataset, and the experiment results show that our method still achieves significant
advantages in most metrics compared to the state-of-the-art models.
Keywords: Dense depth estimation; Transformer; Radar-camera fusion; Radar Cross Section

Linqi Zhao, Pedro Cheong,


Open-set pedestrian identification via transformer-based neural architecture with extreme
value theory,
Engineering Applications of Artificial Intelligence,
Volume 164, Part A,
2026,
113253,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Real-world radio frequency (RF) sensing deployments frequently encounter
identities not observed during training. However, most existing methods are designed under
closed-set assumptions. To address this limitation, we introduce an open-set pedestrian
identification system that leverages an attention neural network (ANN) together with an
extreme value theory (EVT)-based algorithm. Initially, pedestrian features are captured via a
radio frequency identification (RFID) system and processed by a meticulously crafted ANN
with a transformer architecture to synthesize and analyze features from two-channel RFID
signals. To counteract the final logit layer’s tendency to generate low-dimensional
embeddings overly specific to known classes, we employ a decision-driven EVT approach on
intermediate layers to establish robust class-specific acceptance zones. Experimental results
demonstrate a closed-set identification accuracy of 99%, while maintaining over 90%
accuracy even when half of the test identities are unknown. These findings highlight a
practical RF-sensing solution for large-scale access control and privacy-preserving, ambient-
assisted Internet of Things (IoT) applications.
Keywords: Attention mechanism; Extreme value theory; Open-set recognition; Pedestrian
identification; Radio frequency identification

Zheng Tong, Yiming Zhang, Tao Mao,


Guiding GPT models for specific one-for-all tasks in ground penetrating radar,
Automation in Construction,
Volume 171,
2025,
105979,
ISSN 0926-5805,
[Link]
([Link]
Abstract: Buried object detection using ground penetrating radar (GPR) benefits from deep
neural networks but still faces the problem of condition- and question-limited outputs. This
paper presents an approach to conduct “one-for-all” (OFA) tasks in GPR data processing. In
the approach, a generative pre-trained transformer (GPT) generates the prompts based on
input GPR data and an open-ended question. The question, prompts, and GPR data are fed
into a GPT-based large language model to obtain a general-purpose answer. Finally, another
GPT model summarizes the answer into a GPR-purpose one. An experiment with 10k GPR
samples indicates that the proposed approach exceeds the other OFA models with the AP of
77.35% on visual grounding, the BLEU@4 of 65.72 on grounded captioning, the Acc1 of
83.09% on visual question answer, and the Acc1 of 83.64% on object-text matching. Besides,
the proposed approach can handle GPR data with different frequencies and civil structures.
Keywords: Ground penetrating radar; Generative pre-trained transformer; Image and signal
captioning; Natural language processing; Content generation

Hui Wang, Qinghua Liu, Lijun Zhou,


Underground target localization method for ground penetrating radar based on deep
learning,
Measurement,
Volume 253, Part B,
2025,
117647,
ISSN 0263-2241,
[Link]
([Link]
Abstract: To tackle the challenge of subsurface target localization under interference in field
scenarios, a novel two-level cascade network referred to as dual cascade is proposed. The
first level, Cascade-1, is a deep feature extraction network designed to extract and eliminate
direct wave interference signals. On this basis, Cascade-2 is developed using domain
knowledge from ground penetrating radar as prior information, and it incorporates an
attention mechanism along with a feature fusion strategy to enhance the accuracy of target
feature hyperbola detection. Subsequently, the least squares method is employed to fit the
feature hyperbola, and location estimation is performed based on geometric equations. The
proposed cascade network model has demonstrated superior performance compared to
other algorithms, such as column-connection clustering algorithm, YOLOv9, and Faster R-
CNN, in terms of the composite metric F1, which validates the model’s effectiveness in
extracting the feature hyperbola. Additionally, the proposed localization method has
exhibited greater accuracy than the conventional full waveform inversion algorithm.
Keywords: Ground Penetrating Radar; Buried Target Location; Domain Knowledge; Deep
Learning; Cascade Network

Bin Zhang, En-Cheng Liou, Yi-Chih Tung, Muhammad Usman, Chiung-An Chen, Chao-Shun
Yang,
Cross-Site Map-Free Indoor Localization for 6G ISAC Systems Using Low-Frequency Radio and
Transformer Networks,
CMES - Computer Modeling in Engineering and Sciences,
Volume 145, Issue 2,
2025,
Pages 2551-2571,
ISSN 1526-1492,
[Link]
([Link]
Abstract: Indoor localization is a fundamental requirement for future 6G Intelligent Sensing
and Communication (ISAC) systems, enabling precise navigation in environments where
Global Positioning System (GPS) signals are unavailable. Existing methods, such as map-
based navigation or site-specific fingerprinting, often require intensive data collection and
lack generalization capability across different buildings, thereby limiting scalability. This
study proposes a cross-site, map-free indoor localization framework that uses low-frequency
sub-1 GHz radio signals and a Transformer-based neural network for robust positioning
without prior environmental knowledge. The Transformer’s self-attention mechanisms allow
it to capture spatial correlations among anchor nodes, facilitating accurate localization in
unseen environments. Evaluation across two validation sites demonstrates the framework’s
effectiveness. In cross-site testing (Site-A), the Transformer achieved a mean localization
error of 9.44 m, outperforming the Deep Neural Network (DNN) (10.76 m) and
Convolutional Neural Network (CNN) (12.02 m) baselines. In a real-time deployment (Site-B)
spanning three floors, the Transformer maintained an overall mean error of 9.81 m,
compared with 13.45 m for DNN, 12.88 m for CNN, and 53.08 m for conventional
trilateration. For vertical positioning, the Transformer delivered a mean error of 4.52 m,
exceeding the performance of DNN (4.59 m), CNN (4.87 m), and trilateration (>45 m). The
results confirm that the Transformer-based framework generalizes across heterogeneous
indoor environments without requiring site-specific calibration, providing stable, sub-12 m
horizontal accuracy and reliable vertical estimation. This capability makes the framework
suitable for real-time applications in smart buildings, emergency response, and autonomous
systems. By utilizing multipath reflections as an informative structure rather than treating
them as noise, this work advances artificial intelligence (AI)-native indoor localization as a
scalable and efficient component of future 6G ISAC networks.
Keywords: Indoor localization; 6G; ISAC; transformer; deep learning; map-free; cross-site;
wireless sensing

Jingpeng Gao, Sisi Jiang, Xiangyu Ji, Chen Shen,


Cross-domain prototype similarity correction for few-shot radar modulation signal
recognition,
Signal Processing,
Volume 223,
2024,
109575,
ISSN 0165-1684,
[Link]
([Link]
Abstract: The new classes of radar signals are increasingly difficult to acquire under non-
cooperative environments, which makes it difficult to support convolutional neural network
training with limited labeled samples. The few-shot learning (FSL) methods have shown
great performance in classification with limited labeled samples, but the FSL methods ignore
that the class distributions between the new and original tasks are significantly different,
resulting in a massive challenge in identifying new radar signals. To solve this problem, a
few-shot radar modulation signal recognition method based on cross-domain prototype
similarity correction (CDPSC) is proposed. Specifically, a residual feature tokenizer
transformer (RFTT) model embedded with a pooling token generation block is designed to
focus on the important features and improve the ability to represent samples. Meanwhile,
the proposed domain prototype similarity mapping (DPSM) strategy adaptively learns the
class mapping, reduces the inter-domain difference through feature distribution alignment,
and effectively corrects the target domain prototypes. In addition, we introduce a sample
prototype embedding (SPE) strategy in the training phase, which can reduce the intra-class
distance and increase the inter-class distance. Experimental results demonstrate that the
CDPSC method is superior to typical FSL methods in recognition accuracy under different
sample numbers.
Keywords: Cross-domain; Few-shot learning; Prototype similarity correction; Radar
modulation signal recognition

André Mariano, Bernardo Leite, Thierry Taris, Jean Baptiste Bégueret,


Co-design of a wideband double-balanced active mixer and transformer-based baluns for
77GHz radar applications,
Microelectronics Journal,
Volume 45, Issue 11,
2014,
Pages 1566-1574,
ISSN 1879-2391,
[Link]
([Link]
Abstract: This paper presents a 130-nm BiCMOS active mixer dedicated to 77GHz
automotive radar applications. The architecture is based on a double-balanced Gilbert cell
with integrated transformer-based baluns. Interconnections between devices, capacitor
accesses and Tee-junctions are modeled using EM software in order to improve the
simulation accuracy. Focusing on wideband operation, the transformer-based baluns are
considered as part of the input matching network. Sizing of the transformer is detailed along
with its amplitude and phase balance performances. The design of the input matching circuit
integrating the transformer is presented, providing a 12-GHz bandwidth. Measured noise
figure, conversion gain and compression point of the mixer are displayed and compared to
the state of the art.
Keywords: mm-waves; Mixer; Gilbert cell; Transformer; Balun

Xinyi Fu, Zhengchun Zhou, Hua Meng, Shuting Li,


A synthetic aperture radar small ship detector based on transformers and multi-dimensional
parallel feature extraction,
Engineering Applications of Artificial Intelligence,
Volume 137, Part A,
2024,
109049,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Deep learning-based target detection methods have been widely used for
Synthetic aperture radar (SAR) ship detection in recent years. However, the large amount of
clutter noise and land interference information in SAR images poses challenges to small ship
target detection. Existing methods are still inadequate for solving the problem of noise-
induced feature loss of small targets, and feature extraction of small targets is inadequate,
leading to poor detection accuracy. In response to these problems, we propose a new small
ship detector-you only look once (SSD-YOLO). First, to solve the problem that noise
interference leads to the loss of ship body features, a small target feature enhancement
module (STFEM) based on multi-dimensional parallel feature extraction is introduced.
STFEM uses attention mechanisms to capture the main features of small ships from multiple
dimensions, effectively enhancing the boundary texture information of the ship body. In
addition, a transformer-based skip connection path aggregation network (Tr-PANet) is
introduced to adequately extract the contextual information of small targets. Tr-PANet uses
self-attention to model the non-local contextual features of small targets in bottom-up
feature fusion and uses skip connection to maximally maintain the global structural
information of small targets. Since SSD-YOLO is based on a lightweight model, the number of
parameters and the computational amount are lower than most of the comparative models.
Moreover, SSD-YOLO shows the best detection performance on three real datasets.
Keywords: Feature enhancement; Transformer; Small ship detection; Noise interference;
Synthetic aperture radar

Bingzhe Fu, Wei Wang, Yihuan Li, Guorui Ren, Kang Li,
A frequency loss function based dynamic convolutional transformer model with data
denoising for short-term wind speed forecasting,
Expert Systems with Applications,
Volume 299, Part B,
2026,
130166,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Accurate and efficient wind speed forecasting is essential for power grid
management in balancing supply and demand, and support cost-effective transition to
renewable energy. This study develops a novel wind power forecasting system designed to
reduce deployment training computational time and support periodic updates in practical
applications. The system first applies Stagewise Orthogonal Matching Pursuit (StOMP) to
suppress noise in the raw wind speed sequences and performs correlation-based feature
selection to determine the most informative historical inputs. Subsequently, DCFDTFDD, the
forecasting model is constructed, in which multi-layer Dynamic Convolution (DC) modules
are employed for multi-scale feature fusion, and a frequency debiased Transformer is
introduced to directly learn multiple frequency components of wind speed sequences
through frequency-domain modeling and patching operations, without the need for prior
signal decomposition. Additionally, a Frequency-Domain Decorrelation (FDD) loss function is
incorporated to mitigate the inherent autocorrelation of label sequences in Transformer.
Experiments conducted on three real-world wind speed datasets demonstrate that the
proposed system delivers both accurate and efficient predictions, reducing training
computational time by 36.20 %, 34.24 % and 41.46 % compared with state-of-the-art
decomposition-based hybrid methods, while maintaining accuracy within 0.54 m/s RMSE.
These results indicate that the developed system offers a practical solution for wind farm
applications considering engineering time.
Keywords: Wind speed forecasting; Stagewise orthogonal matching pursuit denoising;
Transformer; Frequency-domain decorrelation loss function

Hua Wang, Qiangyu Zeng, Hao Wang, Jianxin He, Tiantian Yu, Guangpu Liu,
Temporal super-resolution reconstruction of weather radar echoes using a deep learning
approach,
Expert Systems with Applications,
Volume 300,
2026,
130189,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Severe convective weather events are characterised by rapid evolution and high
destructive potential, requiring weather radars to provide observations with high temporal
resolution. However, current S-band weather radar systems, constrained by their volumetric
scanning strategies, often fail to capture the rapidly changing features of these systems
promptly. To address this limitation, we propose EMAIRA-VFI, a deep learning–based
method for temporal super-resolution reconstruction of radar echoes, which enhances the
temporal resolution of radar data to meet the demands of severe convective weather
monitoring. By introducing an inter-frame attention mechanism, the proposed method
effectively fuses spatiotemporal features from sequential radar echoes, enabling accurate
modelling of dynamic weather evolution and the generation of continuous, high-temporal-
resolution radar echoes. Compared with conventional temporal interpolation methods,
EMAIRA-VFI demonstrates significant improvements in both interpolation accuracy and the
preservation of fine-scale meteorological structures. Experimental results show that the
model not only enhances the capability of S-band radars in monitoring rapidly evolving
weather events but also provides a new perspective for spatiotemporal fusion and the
intelligent application of radar data. We have open-sourced the code for this work at
[Link]
Keywords: Temporal super-resolution; Radar echo; Inter-frame attention mechanism

Yunlong Zhou, Chen Zhao, Fanfan Ji, Renlong Hang, Qingshan Liu, Xiao-Tong Yuan,
More realistic and accurate precipitation nowcasting with Conditional Rectified Flow
Transformers,
Engineering Applications of Artificial Intelligence,
Volume 165, Part A,
2026,
113402,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Precipitation nowcasting plays a critical role in disaster prevention and daily life but
remains challenging due to the intricate spatiotemporal dynamics of atmospheric processes.
In response to these challenges, recent research has shown that diffusion models can
generate visually realistic precipitation results. However, challenges such as accurately
predicting precipitation positions and improving inference speed remain unresolved. To
address these issues, we propose a novel Conditional Rectified Flow Transformers (CRFT)
architecture to improve precipitation nowcasting, which is designed to deliver both
predictive accuracy and visual realism. At its core, CRFT features an efficient latent space
predictor powered by OmniFormer blocks, which integrate spatial, temporal, and
spatiotemporal Transformers to holistically capture the atmosphere dynamics. We explore
five variants of spatiotemporal dynamic information interactions for OmniFormer and
demonstrate that integrating triple Transformers achieves the best performance.
Additionally, we significantly reduce inference time by employing a rectified flow approach,
achieving a reduction in inference steps by 98.4% compared to existing methods, enabling
high-quality 20-frame predictions within 2 s. Evaluated on three benchmark datasets, CRFT
outperforms state-of-the-art (SOTA) models in both accuracy and quality across multiple
metrics, offering an accurate and efficient solution for real-world nowcasting. The code is
publicly available at [Link]
Keywords: Precipitation nowcasting; Rectified flow; Transformer; Variational autoencoder

Wen Su, Peter Xiaoping Liu, Jun Yu,


Boosting monocular depth estimation with semantic diffusion guided convolutional
transformer model,
Neurocomputing,
Volume 664,
2026,
132113,
ISSN 0925-2312,
[Link]
([Link]
Abstract: Monocular depth estimation has garnered significant attention in the past decade
due to its convenient data acquisition and wide range of applications. While fully supervised
estimation methods have traditionally been considered the gold standard in this field, the
disparity between the accuracy of estimating fine-grained objects within a scene and
maintaining consistency in global scene structure has constrained research to autonomous
driving applications. This paper systematically analyzes the factors that may impact global
consistency and fine-grained target estimation under full supervision. A semantic diffusion
guided convolutional transformer network that takes dual-resolution images of the same
scene as input is proposed. Specifically, it divides the supervised depth into base depth and
residual depth, smoothing fine-grained depth perturbations while preserving boundary
depth variations. The base branch provides a globally consistent depth estimate by encoding
wide-area context. In contrast, the residual branch is specifically designed for fine-grained
depth estimation, utilizing multiscale CNN features and transformer context to address the
task of refining depth at boundaries and recovering local geometric details. Additionally, a
semantic-guided anisotropic diffusion model is employed to optimize depth at boundaries.
Benchmark experiments demonstrate that the proposed semantic diffusion guided
convolutional transformer network consistently outperforms several typical and state-of-the-
art models on challenging indoor and outdoor public benchmarks, particularly in capturing
contours in both simple and complex scenes. The source code is available at:
[Link]
Keywords: Depth estimation; Monocular; Convolutional neural network; Transformer;
Diffusion model

Liuyu Yang, Yuan An, Gang Zhang, Tuo Xie, Mengxin Liu,
Day-ahead electricity price forecasting method integrating multi-scale hypergraph features
and dual-layer transformer,
Applied Energy,
Volume 407,
2026,
127396,
ISSN 0306-2619,
[Link]
([Link]
Abstract: Accurate forecasting of spot electricity prices is critical yet challenging due to the
multi-scale temporal coupling, nonlinear volatility, and complex spatial dependencies
influenced by supply-demand fluctuations, extreme weather, and transmission topology.
This study proposes a novel day-ahead price forecasting model integrating multi-scale
hypergraph features with a dual-layer Transformer. A hypergraph is constructed based on
price trend similarity to capture spatial dependencies at local, global, and full-fusion levels.
High-relevance exogenous variables are selected using the maximum information coefficient
(MIC), and a two-tier Transformer separately models temporal and spatial dynamics.
Spectral hypergraph convolution is introduced to generate dynamic spatial representations.
The model is evaluated on real-world data from the Guangdong electricity market using both
single-day and rolling forecast tasks. Compared with the second-best model, RMSE, MAE,
and MAPE are reduced by 9.23%, 12.00%, and 21.74%, respectively, with R2 improved by
2.25%. Additionally, SHAP analysis quantifies feature contributions, forming a closed-loop
feature selection and validation process with MIC. The results demonstrate that
incorporating multi-scale dynamic modeling and spatiotemporal feature fusion can
significantly enhance forecasting accuracy.
Keywords: Multi-scale analysis; Spatial-temporal feature; Hypergraph; Transformer;
Electricity price forecasting; SHAP
Yunlin Ma, Tengfei Bao, Yangtao Li, Mengfan Zhao, Zhenhao Wu, Chengbo Fan,
GANFormerNet: A UAV-based Concrete Crack Segmentation Model for Water-related
Structures Using Vision Transformer and Graph Attention Network,
Advanced Engineering Informatics,
Volume 68, Part B,
2025,
103725,
ISSN 1474-0346,
[Link]
([Link]
Abstract: In this study, we propose a novel concrete crack segmentation model for water-
related structures, GANFormerNet, which realizes multi-feature fusion, dynamic attention
mechanism, and lightweight design. Specifically, the combination of ViT and Graph Attention
Network improves the crack recognition effect through topological modeling; the innovative
design of spatial-channel synergetic attention (SCSA) and semantic interaction module
(MSAF) work in concert to achieve multi-scale feature extraction and dynamic focusing; the
use of Depth-Separable Convolution reduces the amount of parameters, and the
introduction of the Focal Tversky loss function is introduced to solve the crack imbalance
problem at the boundary. The experiments show that the mIoU, Recall, Precision, and F1-
score of the GANFormerNet model reach 0.92178, 0.94375, 0.93531, and 0.94289,
respectively, which are significantly better than those of existing methods. The final model
was integrated into a GUI system to realize offline real-time detection of dam cracks.
Keywords: Water-related Structures; UAV; Crack Segmentation; GANFormerNet; Vision
Transformer; Graph Attention Network

Zhu Li, Xu Wanru, Zhu Chunqiang, Gao Jingkai, Mi Lugema, Deng Fan, Qu Jinqi,
AFMT:Adaptive frequency decomposition and multi-scale transformer for time series
forecasting,
Information Sciences,
Volume 726,
2026,
122735,
ISSN 0020-0255,
[Link]
([Link]
Abstract: Time series forecasting is essential in various fields, including power systems,
transportation, and meteorology. Although many existing methods improve predictive
performance by decomposing sequences into trend and seasonal components, traditional
techniques, such as moving averages, often suffer from spectral aliasing between high- and
low-frequency features, thereby impairing forecasting precision. Furthermore, current multi-
scale frameworks frequently neglect high-frequency patterns, restricting their capacity to
model complex nonlinear dependencies. To overcome these limitations, we propose AFMT, a
novel time series forecasting framework that integrates Adaptive Frequency-Domain
Decomposition with a multi-scale patch-wise Transformer architecture. Specifically, we
design a dynamic filter capable of adaptively isolating high- and low-frequency components
based on their spectral distributions, effectively mitigating feature entanglement.
Subsequently, a multi-scale patching strategy enables independent modeling of these
components through Transformer blocks, followed by a learnable frequency-aware fusion
mechanism, thereby enhancing feature independence and boosting predictive accuracy.
Extensive empirical studies on eight public datasets demonstrate that AFMT consistently
surpasses state-of-the-art approaches in both forecasting accuracy and generalization ability,
validating its effectiveness in frequency-domain decomposition and multi-scale temporal
modeling.
Keywords: Time series forecasting; Adaptive frequency domain decomposition; Multi-scale
patches; Transformer; Deeply separable convolution

Hebin Liu, Qizhi Xu, Xiaolin Han, Biao Wang, Xiaojian Yi,
Attention on the key modes: Machinery fault diagnosis transformers through variational
mode decomposition,
Knowledge-Based Systems,
Volume 289,
2024,
111479,
ISSN 0950-7051,
[Link]
([Link]
Abstract: Machinery signals typically consist of multiple sub-signals in different frequency
bands, while existing Transformer-based fault diagnosis methods often lack attention to key
fault frequencies, causing interference in fault diagnosis. Therefore, an innovative
Transformer structure for fault diagnosis based on variational mode decomposition (VMD) is
proposed. First, to address the difficulty in identifying signal features arising from coupling of
multiple frequency bands, a mode encoder based on VMD is proposed to decompose the
coupled modes and calculate the key modes. Second, a position encoding method based on
central frequency is proposed to address the lack of attention to signal’s frequency in
existing position encoding methods. Third, fault characteristic frequency is used to verify the
frequency band attention scores, improving the interpretability and reliability of the network
in response to the lack of internal interpretability in fault diagnosis methods based on deep
learning. Finally, the proposed method was validated on bearing vibration dataset and motor
sound dataset. The results showed that the method has high diagnostic accuracy, and could
capture the intrinsic modes of different faults.
Keywords: Fault diagnosis; Transformer; Mode decomposition

Tong Hou, Hongqing Zhu, Ziying Wang, Zhong Zheng, Kai Chen, Ying Wang, Bingcang Huang,
Low-rank fused modality assisted magnetic resonance imaging reconstruction via an
anatomical variation adaptive transformer,
Pattern Recognition,
Volume 175,
2026,
113044,
ISSN 0031-3203,
[Link]
([Link]
Abstract: In clinical practice, precise magnetic resonance imaging (MRI) reconstruction from
undersampled data is crucial. While multi-modal approaches can enhance reconstruction
quality, acquiring fully sampled auxiliary information is often time-consuming. Given that
computed tomography (CT) images are routinely obtained during clinical examinations, this
paper utilizes an anatomical variation adaptive transformer (AVAT) assisted by low-rank
fused CT-MRI modality to propose an MRI reconstruction network (ATLF-Net). Specifically,
this method leverages a CT-MRI fused modality to assist in MRI reconstruction. The ATLF-Net
encompasses fusion and reconstruction processes. The fusion process aims to generate a CT-
MRI fused modality that minimizes the gap between it and the MRI modality, serving as
auxiliary information for reconstruction. The proposed ATLF-Net comprises the global
feature-aware block (GFAB), the local feature-aware block (LFAB), and the low-rank fusion
module (LRFM). GFAB and LFAB extract global and local information from shallow features,
respectively. LRFM fuses CT and undersampled MRI through modality-specific low-rank
factors. During the reconstruction process, an AVAT is developed to extract complex and
elongated pathological features. Extensive experiments show that the proposed ATLF-Net
achieves robust performance and high-quality reconstructed images with few parameters
compared to benchmarks, across various acceleration rates on public datasets. The code for
this project is available at [Link]
Keywords: MRI Reconstruction; Multi-modality; Low-rank; Anatomical variation adaptive
transformer

Jiming Lv, Daiyin Zhu, Zhe Geng, Shengliang Han, Yu Wang, Zheng Ye, Tao Zhou, Hongren
Chen, Jiawei Huang,
Recognition for SAR deformation military target from a new MiniSAR dataset using multi-
view joint transformer approach,
ISPRS Journal of Photogrammetry and Remote Sensing,
Volume 210,
2024,
Pages 180-197,
ISSN 0924-2716,
[Link]
([Link]
Abstract: Accurately detecting ground armored weapons is crucial for achieving initiative
advantages in military operations. Generally, satellite or airborne synthetic aperture radar
(SAR) systems face limitations due to their revisit cycles and fixed flight trajectories, resulting
in single-view imaging of targets, thereby hampering the recognition of small SAR ground
targets. In contrast, MiniSAR possesses the capability to capture the multi-view of a target by
acquiring images from different azimuth angles. In this research, our team utilizes a self-
developed MiniSAR system to generate multi-view SAR images of real ground armored
targets and recognize targets. However, the recognition of small targets in SAR images
encounters two significant difficulties. First, small targets in SAR images are prone to
interference from background noise. Second, SAR target deformation arises from variations
in depression angles and imaging processes. To tackle these difficulties, this paper proposes
a novel SAR ground deformation target recognition approach based on a joint multi-view
transformer model. The method first preprocesses SAR images based on a low-frequency
priori SAR image denoising method. Next, it obtains multi-view joint information through a
self-attentive mechanism, inputs joint features to the transformer structure. The outputs are
jointly updated by a multi-way averaging adaptive loss function to improve the recognition
accuracy of deformed targets. The experimental results demonstrate the superiority of the
proposed method in SAR ground deformation target recognition, outperforming other
representative approaches such as information fusion of target and shadow (IFTS) and Vision
Transformer (ViT). It is concluded that the proposed method has high recognition accuracies
of 98.37% and 93.86 % on the moving and stationary target acquisition and recognition
(Mstar) and our SAR images dataset, respectively, in the field of SAR ground deformation
target recognition. We have included links to the code and data in the abstract of this paper
for ease of access. The source code and sample dataset are available at
[Link]
Keywords: MiniSAR; Transformer network; Denoising with low-frequency prior information
(LFPD); Recognition of deformation targets; Cross attention mechanism

Jinyuan Zhang, Zhineng Zheng, Tianqing Ling,


Transformer in civil engineering defect detection: A survey,
Journal of Traffic and Transportation Engineering (English Edition),
Volume 12, Issue 5,
2025,
Pages 1330-1359,
ISSN 2095-7564,
[Link]
([Link]
Abstract: Detecting structural and functional defects in large-scale civil infrastructure during
operation is of paramount importance, and employing intelligent algorithms for detection
holds significant value. Deep learning technology emerges as the primary avenue to
accomplish this task. Recently, Transformer self-attention models have garnered attention as
alternatives to deep convolutional neural networks due to their robust parallel computing
capabilities and adeptness in modeling long-range dependencies. Harnessing this paradigm
shift, the field of civil engineering has delved into exploring Transformer's applicability in
intelligent defect detection tasks. Motivated by this, we conducted a systematic investigation
into the latest engineering defect detection applications of Transformers. Our survey
encompasses over 40 engineering detection algorithms based on Transformers, with
roadways, tunnels structures, and bridges serving as primary application scenarios. Lastly,
we discuss key challenges encountered in these applications and provide insights into future
development directions. This survey marks a systematic overview of Transformer
applications in the domain of civil engineering. Its objective is to aid researchers in grasping
the concept and architecture of Transformer models, and to furnish suitable reference
algorithms and application strategies for intelligent detection research in civil infrastructure.
Keywords: Civil engineering; Transformer; Defect detection; Deep learning; Image processing

Ayesha Ibrahim, Muhammad Zakir Khan, Muhammad Imran, Hadi Larijani, Qammer H.
Abbasi, Muhammad Usman,
RadSpecFusion: Dynamic attention weighting for multi-radar human activity recognition,
Internet of Things,
Volume 33,
2025,
101682,
ISSN 2542-6605,
[Link]
([Link]
Abstract: This paper presents RadSpecFusion, a novel dynamic attention-based fusion
architecture for multi-radar human activity recognition (HAR). Our method learns activity-
specific importance weights for each radar modality (24 GHz, 77 GHz, and Xethru sensors).
Unlike existing concatenation or averaging approaches, our method dynamically adapts
radar contributions based on motion characteristics. This addresses cross-frequency
generalization challenges, where transfer learning methods achieve only 11%–34% accuracy.
Using the CI4R dataset with spectrograms from 11 activities, our approach achieves 99.21%
accuracy, representing a 15.8% improvement over existing fusion methods (83.4%). This
demonstrates that different radar frequencies capture complementary information about
human motion. Ablation studies show that while the three-radar system optimizes
performance, dual-radar combinations achieve comparable accuracy (24GHz+77GHz: 96.1%,
24GHz+Xethru: 95.8%, 77GHz+Xethru: 97.2%), enabling flexible deployment for resource-
constrained applications. The attention mechanism reveals interpretable patterns: 77 GHz
radar receives higher weights for fine movements (superior Doppler resolution), while 24
GHz dominates gross body movements (better range resolution). The system maintains
71.4% accuracy at 10 dB SNR, demonstrating environmental robustness. This research
establishes a new paradigm for multimodal radar fusion, moving from cross-frequency
transfer learning to adaptive fusion with implications for healthcare monitoring, smart
environments, and security applications.
Keywords: Human activity recognition; Multi-modal fusion; Attention mechanisms; Cross-
frequency transfer learning

Qing Snyder, Qingtang Jiang, Erin Tripp,


Integrating self-attention mechanisms in deep learning: A novel dual-head ensemble
transformer with its application to bearing fault diagnosis,
Signal Processing,
Volume 227,
2025,
109683,
ISSN 0165-1684,
[Link]
([Link]
Abstract: In this paper, we propose a novel dual-head ensemble Transformer (DHET)
algorithm for the classification of signals with time–frequency features such as bearing
vibration signals. The DHET model employs a dual-input time–frequency architecture,
integrating a 1D Transformer model and a 2D Vision Transformer model to capture the
spatial and time–frequency features. By utilizing data from both the time and time–
frequency domains, the proposed algorithm broadens its feature extraction capabilities and
enhances the model’s capacity for generalization. In our DHET structure, the original
Transformer model leverages self-attention mechanisms to consider relationships among
signal input segmentations, which makes it effective at capturing long-range dependencies
in signal data, while the Vision Transformer model takes 2D images as input and creates the
image patches for embedding and each patch is linearly embedded into a flat vector and
treated as a ‘token,’ then the ‘tokens’ are processed by the Transformer layers to learn global
contextual representations, enabling the model to perform signal classification task. This
integration notably enhances the performance and capability of the model. Our DHET is
especially effective for rolling bearing fault diagnosis. The simulation results show that the
proposed DHET has higher classification accuracy for bearing fault diagnosis and
outperforms CNN-based methods.
Keywords: Short-time Fourier transform; Transformer; Dual-head ensemble Transformer;
Deep learning; Bearing fault diagnosis

Mohammed A.A. Al-qaness, Abdelghani Dahou, Mohamed Abd Elaziz, Ahmed M. Helmi,
Human activity recognition and fall detection using convolutional neural network and
transformer-based architecture,
Biomedical Signal Processing and Control,
Volume 95, Part B,
2024,
106412,
ISSN 1746-8094,
[Link]
([Link]
Abstract: Human Activity Recognition (HAR) and fall detection, as applications within the
field of biomedical signal processing, are increasingly pivotal in enhancing patient care,
preventive healthcare, and rehabilitation. Fall therapy is one of the most expensive
treatments, usually taking a long time to complete. A single fall accident might result in
serious injuries, long-term incapacity, or even death. As a result, a reliable and cost-effective
fall detection system is essential. Wearable sensors have received wide attention due to
their availability and capability to capture different human motions. Thus, in the current
study, we develop a comprehensive HAR system for multi-classification tasks to recognize
several human actions, such as walking, sitting, standing, falling, and others. At the same
time, a binary classification of this model is developed to recognize fall and non-fall actions,
which can be used to track elderly actions and send an alert in case of falling to do necessary
rescue actions. The developed system is built using a Parallel Convolutional Neural Network
and Transformer-based architecture (PCNN-Transformer). PCNN-Transformer benefits from
the parallel architecture and the residual mapping mechanism to learn temporal feature
representations from the sensors’ data. The CNN blocks are aligned in parallel alongside
several Transformer-based encoders, followed by a concatenation operation to sum up the
extracted features from the input data ( sensors data). Moreover, the CNN blocks implement
a residual mapping mechanism to reduce the model complexity and training time. The
proposed model is tested using several open-source datasets: SisFall, UniMib-SHAR, and
MobiAct. It recorded high accuracy rates compared to several deep learning models. For
instance, in the binary classification (fall detection), the proposed model achieved an
average accuracy of 99.95%, 98.68%, and 99.71% for SisFall, UniMib-SHAR, and MobiAct,
respectively.
Keywords: Fall detection; Human activity recognition (HAR); Pattern recognition; Wearable
sensors; Deep learning

Yang Qin, Huiming Xie, Shuxue Ding, Yujie Li, Benying Tan,
Enhancing vision-and-language transformers through two-stage generative alignment pre-
training,
Engineering Applications of Artificial Intelligence,
Volume 163, Part 3,
2026,
113076,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Vision-language pre-training (VLP) models based on transformer architectures have
achieved significant success in bridging the gap between natural language processing and
computer vision. However, the alignment between visual and textual semantic objects
remains a major challenge, particularly when noisy image–text pairs are used for training,
which can lead to incorrect semantic associations and degrade model performance. In this
paper, we propose a novel two-stage generative alignment pre-training framework, called
VL-GAP (Vision-Language Generative-Alignment Pre-training), designed to improve object
alignment quality in vision-language models. In the first stage, we utilize high-quality
annotated datasets to perform supervised learning, introducing a center-point strategy for
automated object alignment and optimizing multiple loss functions to achieve precise visual-
text alignment. The second stage leverages large-scale noisy datasets for self-supervised
learning, where momentum models and confidence-based pseudo-label filtering are
employed to enhance the model’s robustness to noise. Experimental results demonstrate
that VL-GAP outperforms state-of-the-art models in various downstream tasks, highlighting
the importance of object alignment quality over data scale in improving VLP model
performance. Our approach provides new insights into effective handling of noisy data and
advances the capability of vision-language models to understand and generate coherent
multimodal descriptions.
Keywords: Vision-and-language; Transformer; Generative-alignment pre-training; Object
alignment

Yu Si, Zhaofeng He, Fan Zhang, Xiaoyun Sun, Yong Chen, Haiqing Zheng,
Cost-effective and real-time landslide monitoring method based on ultra-wideband using
ultra-wideband transformer neural network,
Engineering Applications of Artificial Intelligence,
Volume 160, Part A,
2025,
111851,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Landslides rank among the most destructive natural phenomena, posing
substantial risks to human safety, infrastructure, and ecological systems. Their frequent
occurrence in topographically complex regions demands urgent development in real-time
monitoring solutions. Current monitoring methodologies, however, are constrained by
prohibitive costs, limited temporal resolution, and high-power consumption. These factors
create substantial implementation barriers to implementing landslide monitoring systems.
To address these limitations, this study proposes an economical real-time monitoring
method leveraging ultra-wideband (UWB) technology for landslide detection. The
implementation of a dual-Microcontroller Unit (MCU) distributed hardware architecture
enables high-accuracy ranging capabilities and high real-time performance. To enhance the
spatial resolution of UWB systems in landslide monitoring, we propose an optimized sensor
deployment structure and a novel deep learning architecture called Ultra-wideband
Transformer (UWBformer). This network utilizes differential UWB-ranging data to predict
spatial displacement at monitoring locations, specifically the displacement distance,
horizontal angle, and pitch angle. UWBformer incorporates a spatial multi-head attention
mechanism and a dual-channel architecture processing both time-domain and frequency-
domain features. It is specifically designed to mitigate ranging error propagation and
enhance prediction stability by focusing on relative distance changes rather than absolute
ranging accuracy. Empirical results demonstrate UWBformer's superior performance in
predicting displacement distance, horizontal angle, and pitch angle, outperforming the
conventional Caffery-Taylor (C-T) localization approach and established deep learning
benchmarks. Field tests incorporated 3σ criterion and Kalman filtering alongside to pre-
process raw measurements, thereby enhancing data stability. Comprehensive validation
across field tests demonstrates UWBformer's capability to maintain accurate spatial
displacement estimation under harsh environments.
Keywords: Landslide monitoring; Ultra-wideband sensors; Ultra-wideband Transformer;
Deep learning; Cost-effective monitoring; Real-time monitoring

Yajing Wang, Xiuchen Wang, Miaomiao Kang, Zhihui Zhang, Yichen Yang, Wei Zeng, Zhe Liu,
Recent advances in graphene-based materials for radar and infrared stealth application,
Composites Part A: Applied Science and Manufacturing,
Volume 192,
2025,
108807,
ISSN 1359-835X,
[Link]
([Link]
Abstract: Graphene has attracted attention in the field of electromagnetic stealth due to its
excellent electrical properties. In this work, we first describe how microwave absorption
properties can be enhanced by using graphene to construct heterogeneous interfaces and
structural defects through conventional composite preparation methods. Meanwhile, the
infrared stealth performance is discussed by tuning the charge density and Fermi energy
levels of graphene as well as by constructing three-dimensional porous structures. Then, the
current state of research on microwave and infrared stealth based on metamaterials and
metasurface structure design strategies is reviewed. In addition, the research progress of
radar-infrared compatible stealth technology using three strategies of material composite,
metamaterial and metasurface structure design is summarized. Finally, the advantages and
limitations of graphene-based stealth materials prepared using different strategies are
analyzed, as well as the current challenges.
Keywords: Graphene-based composites; Metamaterials; Metasurfaces; Stealth

Bo Xu, Jialin Yan, Hao Tang, Yonghui Zhang, Maozheng Wang, Jinxiong Gao,
An efficient compressive CNN and transformer hybrid framework for long-term dissolved
oxygen prediction in aquaculture,
Information Processing in Agriculture,
2026,
,
ISSN 2214-3173,
[Link]
([Link]
Abstract: Current methods for predicting water quality do not fully consider the nonlinear
coupling relationship between multiple parameters, given the multifaceted nature of water
quality factor evolution in aquaculture systems. This results in limited prediction accuracy
and generalization ability, making it impossible to predict water quality parameters over an
extended period. This paper proposes a multi-factor related long-term dissolved oxygen
concentration forecasting framework, leveraging the synergistic integration of efficient
compression CNN and Transformer. First, to further improve the short-term feature fusion of
multi-sensor data, a CNN is used to learn localized spatial representations from multi-sensor
measurement inputs; then, a compressed pooling strategy with binary coding is introduced
to mitigate network accuracy degradation during pooling and achieve efficient and accurate
feature coding; finally, the feature fusion results are inputted into Transformer to improve
the global temporal capture ability of the model, achieving long-term water quality
parameter prediction. Empirical evidence demonstrates the superior performance of this
approach in long-term dissolved oxygen concentration forecasting. Compared with
traditional max and average pooling methods, the proposed compressed pooling strategy
improves model convergence speed by 13.59%. Compared with LSTM, GRU, Transformer,
CNN-Transformer (Max-Pool), Strided Convolution-Transformer, iTransformer, and Informer
models, the Coefficient of Determination (R2) can reach 0.969 ± 0.004 and 0.963 ± 0.003 at
prediction step lengths of 96 (4 days) and 168 (7 days). The Root Mean Squared Error (RMSE)
and Mean Absolute Error (MAE) were reduced to 0.022 ± 0.004 mg/L and 0.017 ± 0.002
mg/L for the 96-step prediction horizon, and to 0.025 ± 0.001 mg/L and 0.020 ± 0.001 mg/L
for the 168-step horizon, respectively. The novel predictive framework engineered in this
paper can optimize faster during the parameter learning, enhancing the model training
efficiency and improving the accuracy of long-term predictions.
Keywords: Long-term water quality prediction; Compressed pooling strategy; Binary coding;
Multi-factor correlation

Hui Bi, Chengjie Cai, Jiawei Sun, Shihao Ge, Huazhong Shu, Xinye Ni,
DRTNet: Dual-route transformer network for thyroid ultrasound segmentation based on
Bbox-supervised learning,
Knowledge-Based Systems,
Volume 324,
2025,
113781,
ISSN 0950-7051,
[Link]
([Link]
Abstract: Background and Objective:
Accurate nodule delineation plays a significant role in the intelligent diagnosis of thyroid
disease. However, the labels accessing is difficult since it is time-consuming and laborious. To
mitigate the over-dependence of the segmentation accuracy on the labels, we proposed a
dual-route transformer network (DRTNet) based on Bbox-supervised learning that only
requires the rough bounding rectangle as labels instead of a precise boundary for thyroid
ultrasound segmentation.
Methods:
DRTNet incorporates double-branch foreground class activation mappings (CAMs) into
Transformers to combine key areas. Meanwhile, double-branch architecture dynamically
adjusts the feature distribution of nodules of different sizes in frequency channels and
spatial dimensions, effectively addressing the localization of nodules of different sizes.
Moreover, ultrasound prior background-aware pooling (UPBAP) is proposed in both
branches to deal with the ambiguous boundary of thyroid nodules. Finally, adaptive
uncertainty estimate multi-scale consistency (AUEMC) is proposed to help mitigate the risk
of excessive over-fitting because of pseudo annotations, which further guarantees
consistency among nodules with diverse resolutions.
Results:
Substantial improvement of segmentation accuracy is shown on the public thyroid dataset of
TN3k and DDTI dataset with Dice similarity coefficient (DSC) of 84.94% and 83.98%, with
95% of the asymmetric Hausdorff distance (HD95) of 27.69 and 29.18, respectively. And our
private dataset has a DSC of 84.39% and HD95 of 14.53.
Conclusions:
The proposed DRTNet used rectangular box labeled for thyroid ultrasound images based on
Bbox-supervised learning. The experimental results show that the DRTNet is comparable to
these fully supervised methods. Code is available at [Link]
Keywords: Ultrasound image segmentation; Thyroid nodule segmentation; Transformer-
based network; Bbox-supervised segmentation; DRTNet

Nirupam Das, Refat Noor Swarna, Md. Selim Hossain,


Deep learning-based circular disk type radar target detection in complex environment,
Physical Communication,
Volume 58,
2023,
102014,
ISSN 1874-4907,
[Link]
([Link]
Abstract: Target identification is one of the most popular radar uses in real life. Target
identification is a classifier that analyzes whether a signal contains an echo from a target
(target-present) or is merely noise (target-absent). Deep learning techniques are a popular
topic in classification, and they have evinced to be effective in a range of applications. In this
paper, a 64 layers Circular Disk type RADAR Target Detection (CDRTD) model is proposed
based on Transfer Learning using the SqueezeNet architecture of Convolutional Neural
Network (CNN) that functions directly with processed radar target return eco signal and
minimize the requirement of conventional laborious radar signal processing. Further, the
proposed 64 layers SqueezeNet-based CNN CDRTD model was then implemented to identify
circular disk type targets in complex environment. Finally, the target return eco data was
tested to identify the circular disk type radar target in complex environments. We further
analyzed target detection probability, false alarm rate, precision, recall, F1 in a complex
environment and compared it with the ideal case. We found that our proposed CDRTD
model can classify 83.3% of the test samples correctly with an overall accuracy of 94.59% in
a noisy and cluttered environment whereas 100% of the test samples are classified correctly
with an overall accuracy of 100% in an ideal environment.
Keywords: Deep learning; SqueezeNet-CNN; Circular disk target; Target detection; Complex
environment

Viet Dinh Le, Gyu-Hyun Go, Sayali Pangavhane, Chang Kyoon Yoo,
Developing fully convolutional networks with permittivity-based class mapping for tunnel
lining defects detection in ground penetration radar scan data,
Engineering Applications of Artificial Intelligence,
Volume 166, Part B,
2026,
113670,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Non-destructive testing using ground penetrating radar (GPR) plays a vital role in
identifying defects in tunnel linings, such as cavities, delamination, and interlayers, to ensure
the safety and longevity of underground structures in mega-cities. However, the absence of
automated tools for generating realistic simulation datasets and the limitations of existing
methods in handling complex defect scenarios, such as cavity and delamination, pose
significant challenges for accurate detection. This study addresses these gaps by introducing
KIT-GPR, a numerical simulation tool, and a fully convolutional network (FCN) for automated
defect detection. KIT-GPR has employed the Finite-Difference Time-Domain method to
simulate electromagnetic wave propagation through multi-layered tunnel structures,
resulting in high-fidelity B-scan data. Its innovative application of the Perlin noise algorithm,
commonly used in game design, generates realistic interlayers between grout and rock
layers, thereby improving the accuracy of the simulation model. A modified FCN, designed
with a 320x320 square input size to accommodate versatile data and a 256-class output
corresponding to dielectric constant ranges, was trained using 1000 paired B-scan and
permittivity datasets. This approach achieves a Root Mean Square Error of approximately 0.8
in most scenarios, including interlayers with delamination. Furthermore, the predicted
delamination locations using FCN from B-scan data of a real tunnel closely align with the
results from endoscopic imaging. This demonstrates that the FCN prediction model has
significant promise as a scalable and efficient solution for early defect detection, thereby
greatly enhancing the safety of underground infrastructure.
Keywords: Ground penetrating radar; Fully convolutional network; Local defects; Tunnel
lining

Jing Liu, Jing Ju, Kai Huang, Xiao Guan,


Function prediction of grain proteins based on graph transformer and protein interaction
network,
Food Bioscience,
Volume 73,
2025,
107592,
ISSN 2212-4292,
[Link]
([Link]
Abstract: Grain proteins are essential for human health and food security. However,
accurately predicting their biological functions remains challenging due to the difficulty of
simultaneously capturing both local structural details and global positional information from
protein data. To address this issue, a novel graph-based deep learning model, SPE-GTN
(Structural and Positional Encoding Graph Transformer Network), is proposed for grain
protein function prediction. In this model, structural and positional encodings are embedded
into a protein complex graph to enable the joint extraction of local and global features.
Graph Convolutional Networks (GCNs) are employed to aggregate neighborhood
information, whereas Transformer mechanisms are used to model long-range dependencies
among protein nodes. The proposed SPE-GTN model is evaluated on four datasets
comprising proteins from wheat, soybean, maize, and indica rice, representing a diverse
range of grain types. Experimental results demonstrate that SPE-GTN achieves a 13.6 %
improvement in prediction accuracy and a 9.4 % enhancement in F1-score. Compared to
state-of-the-art methods. Theoretical analysis further validates its capacity to effectively
capture complex relationships within protein interaction networks. These findings highlight
the effectiveness and generalizability of SPE-GTN in real-world grain protein function
prediction tasks and provide novel insights into protein bioinformatics and agricultural
genomics.
Keywords: Grain protein function prediction; GNN; Transformer; Positional encoding;
Structural encoding

Jongyun Byun, Jaehoon Cha, Jeyan Thiyagalingam, Hyeon-Joon Kim, Changhyun Jun,
Enhancing rainfall prediction accuracy through image fusion of radar and numerical weather
prediction models,
Expert Systems with Applications,
Volume 303,
2026,
130516,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Precipitation is one of the most challenging atmospheric phenomena to predict
due to the complexity involved in solving dynamic and thermodynamic atmospheric
equations. To address this challenge, extensive research has been conducted to enhance the
precision of numerical weather prediction models and radar-based extrapolation data, in
conjunction with the development of various blending techniques. However, traditional
methods have proven insufficient in capturing the diversity and nonlinearity of weather
phenomena. In response to these limitations, this study introduces a novel methodology
that leverages machine-learning based image fusion models to merge radar-based
extrapolation and numerical weather prediction rainfall datasets, thereby enhancing
prediction accuracy. An image fusion model was developed using radar-based extrapolation
data and numerical weather prediction data as input datasets, with radar observation data
utilized as target dataset. To identify the most suitable image fusion model for capturing the
complex patterns of rainfall data, two experiments were conducted: 1) Impact of model
topology, and 2) Effect of model size. A systematic analysis of the model outputs was
performed using eight evaluation metrics categorized under pixel-based metrics, feature-
based metrics, structural similarity metrics, and categorical verification metrics.
Experimental results indicated that image fusion model based on a Residual Network
(ResNet) outperformed other models in terms of model topology. Regarding model size, it
was observed that the performance did not increase proportionally with the number of
residual blocks; the most suitable performance was achieved with a specific number of
residual blocks (Case 5: 8 blocks). Additionally, the metrics compared with radar observation
data indicated that the proposed model delivered superior performance, thus offering a
high-accuracy rainfall prediction methodology.
Keywords: Image fusion; Radar; Numerical weather prediction; Deep learning; Precipitation

Guomin Xie, Zijian Zhang, Sen Xie, Zhaowei Yuan, Hao Liu,
CPWformer-DEC: improved Transformer with class-priority weather attention and dynamic
error compensation for photovoltaic power forecasting,
Expert Systems with Applications,
Volume 301,
2026,
130580,
ISSN 0957-4174,
[Link]
([Link]
Abstract: With the rapid expansion of photovoltaic (PV) capacity, accurate short-term and
medium-term PV power forecasting has become critical for grid dispatch and energy storage
coordination. However, the sudden and drastic fluctuations of meteorological variables over
a short period of time, as well as the abrupt transitions between contrasting weather states,
severely hinder the forecasting accuracy. Therefore, improved Transformer with class-
priority weather attention and dynamic error compensation (CPWformer-DEC) for
photovoltaic power forecasting is developed. In detail, based on the Transformer network,
sample-level Gaussian mixture model (GMM) categorization is firstly used to finely encode
the weather patterns. Then, class-priority weather attention and multi-scale convolution
module are added to balance intra-class and inter-class information. Moreover, a lightweight
long short-term memory (LSTM) is utilized for online residual correction. The experimental
results indicate that CPWformer-DEC achieves substantially lower errors under diverse
conditions. Compared with the suboptimal model, MAE is reduced by 35.4%, 64.8%, and
70.1% respectively in three typical cases, thereby supporting day-ahead and intra-day
dispatch and improving reserve allocation and storage scheduling in PV integrated power
systems.
Keywords: Photovoltaic power forecasting; Transformer; Class-priority weather attention;
Multi-scale fusion; Dynamic error compensation

Haoyu Zhang, Stephen Wu, Xiangyun Luo, Yong Huang, Hui Li,
Efficient matching of Transformer-enhanced features for accurate vision-based displacement
measurement,
Automation in Construction,
Volume 171,
2025,
105962,
ISSN 0926-5805,
[Link]
([Link]
Abstract: Computer vision technology and monitoring videos have been employed to obtain
structural displacement measurements. Noniterative algorithms are mainly designed for
rapid tracking of the motions of individual image points, rather than dense motion fields.
Iterative algorithms are limited to estimating motion fields with small amplitudes and
require high computation cost to achieve high accuracy. This paper introduces a noniterative
method for vision-based measurements that balances speed and density. The method
employs an attention-based matching strategy applied to Transformer-enhanced image
features. Motion priors and a physics-informed denoising approach are integrated to
improve measurement accuracy. Tested on challenging truss and cable-stayed bridge
vibration videos, the method demonstrated superior displacement measurement
performance compared to conventional approaches. It also achieved greater robustness to
brightness changes and partial occlusions while requiring minimal human intervention. This
method supports the development of automated and affordable vibration monitoring
systems.
Keywords: Attention mechanism; Subpixel feature matching; Transformer; Computer vision;
Displacement measurement; Vibration monitoring

Jackson S. Zaunegger, Paul G. Singerman, Ram M. Narayanan, Muralidhar Rangaswamy,


RadarTD: A Radar Text Dataset for multi-parameter optimization,
Natural Language Processing Journal,
Volume 12,
2025,
100178,
ISSN 2949-7191,
[Link]
([Link]
Abstract: This paper introduces the radar text dataset (RadarTD) for technical language
modeling. This dataset is comprised of sentences containing radar parameters, values, and
units determined from published radar literature. Additionally, each statement is assigned a
sentiment, goal priority, and goal direction label. In this work, we show how RadarTD may be
used to train simple Natural Language Processing (NLP) models to identify the attributes of
each sentence listed in RadarTD. Once the NLP models have identified these attributes from
text, we can use this information to develop Language Based Cost Functions (LBCF). Our
study shows that the proposed text classification model achieves a classification accuracy
between 96.7% and 97.8%, while the proposed named entity recognition model achieves an
F1 score of 99.7. These findings suggest that the developed models are capable of achieving
good performance for both text classification and named entity recognition for autonomous
radar applications. We then illustrate an example of how these models could be used with
Language Based Cost Functions to develop multi-parameter radar optimization schemes. We
also provide a method of providing scalarization weights for each parameter, to improve the
results of the optimization process.
Keywords: Text classification; Named entity recognition; Language modeling; Language-
based cost functions; Multi-parameter optimization; Cognitive radar

Guowen Li, Zihang Huang, Teng Fei, Dunxin Jia, Meng Bian,
Listen to the road: acoustic traffic monitoring on edge platforms via Lightweight Noise
Spectrogram Transformer (LNST),
Pervasive and Mobile Computing,
Volume 115,
2026,
102132,
ISSN 1574-1192,
[Link]
([Link]
Abstract: Accurate real-time traffic flow monitoring is crucial for intelligent transportation
systems (ITS), enabling optimized traffic management, urban planning, and policy-making.
However, conventional methods face cost, deployment, weather, and privacy challenges.
Addressing these shortcomings, this study investigates the potential of utilizing ubiquitous
traffic noise, an inherently accessible, cost-efficient, non-intrusive, and privacy-preserving
signal, as a viable data source. We propose the Lightweight Noise Spectrogram Transformer
(LNST), a novel deep learning model for analyzing traffic noise spectrograms as a Proof of
Concept. LNST leverages the Transformer architecture's self-attention mechanism to
effectively capture long-range temporal and spectral dependencies crucial for interpreting
complex traffic acoustics. Trained and evaluated on diverse urban traffic scenarios, LNST
demonstrates significant advantages. Experimental results show it consistently outperforms
baseline models, achieving superior prediction accuracy (MSE, MAE, R²). Furthermore,
through transfer learning and model pruning, LNST achieves high computational efficiency
with substantially fewer parameters and faster inference speeds. Its lighter design also
ensures its feasibility for deployment on resource-constrained edge computing platforms.
This work validates the practicality of acoustic sensing for traffic monitoring and presents an
accurate, computationally efficient, and LNST as a cost-effective, easily deployable, and
privacy-respecting solution, offering a valuable supplementary tool for advancing ITS.
Keywords: Traffic flow monitoring; Traffic noise; Transformer; Intelligent transportation
system; Edge computing

Abrar M. Alajlan, Marwah M. Almasri,


Autonomous Cyber-Physical System for Anomaly Detection and Attack Prevention Using
Transformer-Based Attention Generative Adversarial Residual Network,
Computers, Materials and Continua,
Volume 85, Issue 3,
2025,
Pages 5237-5262,
ISSN 1546-2218,
[Link]
([Link]
Abstract: Cyber-Physical Systems integrated with information technologies introduce
vulnerabilities that extend beyond traditional cyber threats. Attackers can non-invasively
manipulate sensors and spoof controllers, which in turn increases the autonomy of the
system. Even though the focus on protecting against sensor attacks increases, there is still
uncertainty about the optimal timing for attack detection. Existing systems often struggle to
manage the trade-off between latency and false alarm rate, leading to inefficiencies in real-
time anomaly detection. This paper presents a framework designed to monitor, predict, and
control dynamic systems with a particular emphasis on detecting and adapting to changes,
including anomalies such as “drift” and “attack”. The proposed algorithm integrates a
Transformer-based Attention Generative Adversarial Residual model, which combines the
strengths of generative adversarial networks, residual networks, and attention algorithms.
The system operates in two phases: offline and online. During the offline phase, the
proposed model is trained to learn complex patterns, enabling robust anomaly detection.
The online phase applies a trained model, where the drift adapter adjusts the model to
handle data changes, and the attack detector identifies deviations by comparing predicted
and actual values. Based on the output of the attack detector, the controller makes decisions
then the actuator executes suitable actions. Finally, the experimental findings show that the
proposed model balances detection accuracy of 99.25%, precision of 98.84%, sensitivity of
99.10%, specificity of 98.81%, and an F1-score of 98.96%, thus provides an effective solution
for dynamic and safety-critical environments.
Keywords: Cyber-physical systems; cyber threats; generative adversarial networks; residual
networks; and attention algorithms

Canfeng Liu, Binhui Wang, Hui Dong, Yihan Pan, Jiawen Lin, Jintian Yang, Yihui Tao, Hao Sun,
Time series analysis of nucleic acid reactions via a generalized transformer model,
Chemometrics and Intelligent Laboratory Systems,
Volume 267,
2025,
105522,
ISSN 0169-7439,
[Link]
([Link]
Abstract: The contemporary landscape of medical diagnostics and therapeutic interventions
has witnessed a remarkable surge in the production of time series data. Artificial intelligence
(AI), particularly the deep learning, has presented promising values in investigating the high-
dimension and meaningful significance hidden behind these diagnostic data. In this work,
we propose a novel analytics for intelligent nucleic acid amplification tests (NAAT) based on
deep learning and paper microfluidics. On-chip amplification data were straightforwardly fed
to a deep learning model derived from Transformer neural network. To facilitate the
development and deployment of the approach, we conducted a lightweight processing of
the Transformer model. Then, the capacity of the model for accurately predicting the
reaction trend and end-point value was validated. We also employed ablation experiments
to evaluate the effects of various parameters on prediction performance followed by
optimizing the model. Then, three clinical datasets including 706 positive and 205 negative
samples obtained from Fujian Provincial Hospital were used to verify the generalization of
the approach. Without any modification of the model structure and hyperparameters,
accuracy, sensitivity, and specificity by the presented approach were 98.28 %, 97.52 % and
99.02 %. Further comparison studies based on the nine different AI algorithms including
recurrent neural network and long-short term memory were performed. The presented
study holds potential to facilitating routine diagnostic tasks for preventing pandemic and
propelling the development of smart portable instruments.
Keywords: NAAT; Deep learning; Generalized transformer model; Predictive analytics

Ti Xiang, Pin Lv, Liguo Sun, Yipu Yang, Jiuwu Hao,


TCM Model for improving track sequence classification in real scenarios with Multi-Feature
Fusion and Transformer Block,
Knowledge-Based Systems,
Volume 283,
2024,
111202,
ISSN 0950-7051,
[Link]
([Link]
Abstract: The shipping industry has experienced rapid growth in recent years, prompting a
need for advanced target recognition technology based on marine radar. This paper
introduces the Track Classification Model (TCM), a novel approach for classifying track
sequences in real scenarios. The TCM utilizes a feature extraction network based on multi-
feature fusion, taking radar echo images and motion information of the target as input, to
improve classification accuracy. Additionally, the paper also presents a dataset production
method that addresses the issue of missing labels, a critical problem in track sequence
classification. Through ablation experiments, the paper demonstrates the effectiveness of
the design strategy, with the multi-feature fusion network successfully extracting features
and achieving superior performance over single feature extraction networks. The results
show that increasing the number of input track points and raising the upper limit of the
input sequence leads to improved classification accuracy. Finally, in real scenarios, the
proposed model outperforms other algorithms, showcasing its high engineering application
value.
Keywords: Track classification; Multi-feature fusion; Marine radar; Transformer

Shuai Liu, Tong Yu, Jun Zhou, Guanglong Xing, Huanfa Chen,
Cross-source transformer-based neighborhood contrastive learning for joint classification of
hyperspectral and LiDAR Data,
Information Fusion,
Volume 124,
2025,
103225,
ISSN 1566-2535,
[Link]
([Link]
Abstract: The fusion of the hyperspectral image (HSI) and light detection and ranging (LiDAR)
data has demonstrated significant potential in the land cover classification task. Although
deep learning has shown remarkable success in the joint classification of HSI and LiDAR data,
the large amount of unlabeled multisource remote sensing data is not fully utilized.
Additionally, effectively integrating HSI and LiDAR data remains a challenging task, and the
semantic relationship of neighborhood regions needs to be further exploited. In this paper,
we propose a cross-source transformer-based neighborhood contrastive learning model
(CTNCLM), which acquires a more discriminative feature representation from unlabeled data
through the pre-training stage. Considering the semantic correlation between neighboring
image patches, CTMCLM achieves the joint classification of HSI and LiDAR data at a more
precise level. A cross-patch contrastive learning (CPCL) module is proposed to calculate the
similarity between original patches and neighborhood patches. Furthermore, a cross-source
transformer (CST) with cross-source attention is proposed to fuse the multi-source data,
which exploits the intermodal information interaction between the HSI and LiDAR data.
Extensive experiments on three public datasets demonstrate the superior classification
performance of the proposed method compared with several state-of-the-art methods.
Keywords: HSI-LiDAR classification; Self-supervised learning; Momentum contrast;
Neighborhood contrastive learning; Cross-source transformer

Meng Zhang, Jilong Liu, Bing Han, Shengli Dong, Tong Cui, Yan Ren,
Adaptive processing of non-stationary ship lubrication signals under sporadic impulsive
noise,
Ocean Engineering,
Volume 342, Part 2,
2025,
122693,
ISSN 0029-8018,
[Link]
([Link]
Abstract: During the 2022 transpacific voyage of the tanker “Yuan Fu Yang”, we encountered
a challenging problem in machinery monitoring: sensor signals were severely corrupted by
sporadic impulsive noise superimposed on already non-stationary patterns. The harsh
marine environment–with intense mechanical vibrations and salt-induced sensor
degradation–created signal characteristics that deviated significantly from the Gaussian
assumptions of classical methods. To tackle this real-world challenge, we developed an
adaptive three-stage processing framework. Our first innovation, Adaptive Spectral
Enhancement Signal Extension (ASE-SE), addresses the dyadic length requirements of
wavelet transforms through context-aware autoregressive modeling that maintains signal
dynamics while satisfying computational constraints. Recognizing that conventional
denoising approaches often learn padding artifacts, we then designed a Learnable Wavelet
Packet Transform with Focus mechanism (LWPT-Focus), which cleverly masks synthetic
regions during training to avoid this pitfall. Finally, given the time-varying nature of ship
signals, we created the Frequency-Adaptive Decomposition Linear Model (FADLinear) that
dynamically adjusts trend-seasonal separation according to local spectral features. Testing
on 84,951 samples spanning 12 lubrication system parameters revealed that our approach
achieved an MSE of 0.0386 for 96-step predictions—outperforming various state-of-the-art
baselines including Transformer variants, yet maintaining the computational efficiency
essential for shipboard deployment. This work offers a practical solution for maritime
predictive maintenance under conditions where conventional methods struggle. The code
and dataset for this study are publicly available at: [Link]
padd-foucs-lwpt-fadlinear.
Keywords: Frequency-adaptive modeling; Non-stationary signals; Ship lubrication signals;
Sporadic noise; Time series forecasting; Unsupervised denoising

Najamuddin, Usman Ullah Sheikh, Ahmad Zuri Sha’ameri,


Ensemble deep learning approach for marine vessel classification: Integrating CNN and
vision transformers with machinery feature enhancement,
Information Fusion,
Volume 126, Part A,
2026,
103570,
ISSN 1566-2535,
[Link]
([Link]
Abstract: Accurate classification of marine vessels is a challenging task due to the dynamic
nature of the underwater environment and the complexity of vessel acoustic signatures.
Though recent deep learning studies have significantly improved classification accuracies
under ideal conditions, their performance degrades in noisy conditions due to
environmental complexities and weak acoustic signals. This study proposes an ensemble
deep learning framework combining Convolutional Neural Networks (CNN) and Vision
Transformers (ViT) to enhance classification performance. First, weak machinery
narrowband frequency components are extracted using Coherently Averaged Power
Spectrum Estimation (CAPSE), improving feature representations. Then, a classifier using a
CNN module captures local spatial features, while the ViT models long-range dependencies
from Low Frequency Analysis and Recording (LOFAR), ensuring a more comprehensive
understanding of the acoustic patterns of onboard machinery. By leveraging the
complementary strengths of both architectures, the proposed model effectively improves
classification. The ensemble framework achieved an average accuracy of 98.32%, and F1-
score of 98.31% on DeepShip and 99.50% accuracy and F1-score on ShipsEar dataset, an
improvement over several state-of-the-art approaches while maintaining around 3.5 million
learnable parameters. Moreover, the ensemble approach demonstrates robustness across
varying noise conditions, maintaining an accuracy of over 60% at 0 dB where individual CNN
or ViT models show performance degradation. These results highlight the effectiveness of
the CNN-ViT fusion strategy in passive acoustic classification, offering a strong balance
between accuracy and efficiency. The model’s lightweight architecture makes it particularly
suitable for real-time applications in vessel monitoring, maritime surveillance, and
underwater target classification.
Keywords: Time–frequency representations (TFR); Signal-to-noise ratio; LOFAR gram;
Convolutional Neural Networks; Vision transformer

Pengkai Wang, Jonghoek Kim, Mitra Ghergherehchi, Mingxuan Zhang, Estrella Montero,
Luwei Liao, Zhong Yang, Hongyu Xu,
Transformer-based aerial robot tracking system in environments with wind disturbances,
Robotics and Autonomous Systems,
Volume 193,
2025,
105104,
ISSN 0921-8890,
[Link]
([Link]
Abstract: Unmanned aerial vehicles (UAVs) are increasingly used in agriculture, surveillance,
and search and rescue. However, maintaining stable flight and accurate navigation in
dynamic environments, especially with wind disturbances, remains a challenge. Traditional
navigation systems often struggle with unreliable sensor data, complicating pose estimation
and tracking. This article proposes an advanced master–slave UAV system combining a
transformer-based model with YOLO for enhanced tracking in wind-affected environments.
YOLO performs real-time object detection, extracting feature points matched with known
landmarks to estimate the UAV’s position. To address the challenges of wind disturbances,
we simulate various wind conditions and train the model under different wind disturbance
environments. Using transformer-based trajectory and pose predictions, we provide control
compensation to counteract the effects of wind disturbances, ensuring stable flight in
dynamic conditions. The pose estimation is refined by integrating visual data with inertial
measurement unit (IMU) data using transformer architectures. A vision-based formation
control strategy is introduced for precise relative positioning in multi-UAV formations.
Initially designed for three UAVs, this strategy is extended to handle larger formations and
complex geometric shapes, focusing on maintaining a triangle formation. A graph-based
dynamic formation control framework enables real-time adaptation to formation changes
and environmental conditions. The approach improves MPC control with a transformer
model, enhancing adaptability to wind disturbances. The system’s effectiveness is validated
using webots simulations, demonstrating its ability to track UAVs and adapt to challenging
environmental conditions. A theorem proves the convergence of the control law using
Lyapunov’s direct method, ensuring that formation errors decay over time. Comparative
experiments and webots simulations confirm the approach’s feasibility, validating its
robustness in maintaining precise formation control under dynamic environmental factors.
Finally, we validate the reliability of our method in real-world environments, confirming its
practical applicability.
Keywords: Transformer-based model; YOLO algorithm; Image processing; Target tracking;
Wind disturbances; Unmanned aerial vehicles; Formation control

Jinsheng Fan, Guo-An Yu, Mingmeng Zhao, Hucheng Zong,


Addressing multi-scale temporal variability: deep integration and application of the CNN and
transformer model in monthly streamflow prediction,
Expert Systems with Applications,
Volume 292,
2025,
128658,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Accurate monthly streamflow prediction is essential for effective water resource
management, hydropower operation, and ecological sustainability. However, streamflow
processes are inherently nonlinear and exhibit considerable multiscale temporal variability,
driven by both natural conditions and potential anthropogenic influences. To address these
challenges, we propose a novel hybrid deep learning model, ISVM-CovTransformer, which
integrates the Improved Sparrow Search Algorithm (ISSA), Variational Mode Decomposition
(VMD), Mutual Information (MI), and a composite CovTransformer architecture. Within this
framework, ISSA is utilized to optimize the parameters of VMD for efficient signal
decomposition, while MI is employed to identify informative input features with strong
predictive relevance. The CovTransformer model, combining Convolutional Neural Networks
(CNN) and Transformer layers, enables the simultaneous extraction of localized temporal
patterns and long-range dependencies, thereby enhancing the model’s ability to capture
complex runoff dynamics and improve prediction accuracy. Using monthly precipitation and
streamflow data from the Tangnaihai, Toudaoguai, and Huayuankou hydrological stations,
experimental results demonstrate that the proposed model outperforms baseline
approaches. Specifically, during the testing phase, the model achieved an NSC of 0.9686,
RMSE of 91.99 m3/s, MAE of 70.90 m3/s, R2 of 0.9702, and a PBIAS of −1.198 % at
Tangnaihai; an NSC of 0.9498, RMSE of 90.35 m3/s, MAE of 724.63 m3/s, R2 of 0.9554, and a
PBIAS of 2.573 % at Toudaoguai; and an NSC of 0.9302, RMSE of 174.36 m3/s, MAE of
44.42 m3/s, R2 of 0.9393, and a PBIAS of 3.309 % at Huayuankou. These findings confirm the
proposed model’s effectiveness for monthly streamflow forecasting and suggest that it
provides a theoretically sound and generalizable framework, with potential extensions to
related hydrological applications such as sediment transport modeling.
Keywords: Monthly streamflow prediction; VMD; CNN; Transformer; CovTransformer

Naveed Imran, Jian Zhang, Zheng Yang, Jehad Ali,


mm-FERP: An effective method for human personality prediction via mm-wave radar using
facial sensing,
Information Processing & Management,
Volume 62, Issue 1,
2025,
103919,
ISSN 0306-4573,
[Link]
([Link]
Abstract: mm-FERP (millimeter wave Facial Expression Recognition for Personality) explores
the use of mm-Wave radar technology, specifically the TI IWR1443, to assess personality
traits based on the OCEAN model through facial expression analysis. This research uniquely
combines psychological profiling with state-of-the-art technology to predict the OCEAN
(Openness, Conscientiousness, Extraversion, Agreeableness, Neuroticism) personality traits
by carefully analyzing facial muscle movements collected through mm-wave radar alongside
detailed questionnaire analysis. Our advanced mm-FERP system employs mm-wave radar
technology for the detection and analysis of facial expressions in a manner that is both non-
intrusive and privacy-centric, handling the ethical and privacy concerns associated with
traditional camera-based methods. Using a convolutional neural network (CNN), mm-FERP
effectively analyzes the complex patterns in mm-wave signals. This approach enables the
smooth transfer of model knowledge from extensive image-based (Scalograms) datasets to
the detailed understanding of mm-wave radar signals, significantly enhancing the model’s
predictive accuracy and efficiency in identifying personality traits via emotional behavior.
Our in-depth evaluation reveals mm-FERP’s remarkable potential to predict personality traits
through emotion recognition (Neutral, Smile, Angry, Sad, Amazed) with an impressive
accuracy of 97% across distances up to 0.47 m. We experiment in a controlled environment
with more than 50 participants from different age groups (18–35) including males and
females of different continents to train our model on different facial symmetry. Each
participant gives 50 samples 10 for each expression making a total of 2500 samples. We also
collected a self-assessment report from the same participants of 64 questions related to
psychological behavior to validate personality by correlating it with radar signal features on
question value weight (0.5–1.5). mm-FERP achieve an average score of 97.8% in precision,
97.2% in Recall, and 97.2% of F1. These results show mm-FERP’s ability as an innovative
approach for psychological behavioral analysis through mm-wave emotion recognition,
improving user experience design, and paving the path for interactive technologies that are
both personalized and psychologically insightful.
Keywords: mm-FERP; mm-wave Radar; Deep learning; Emotion recognition; Human facial
expression; OCEAN personality prediction

Xinyu Zhang, Chu Zhang, Rui He, Changwen Ma, Junhao Yao, Muhammad Shahzad Nazir,
Tian Peng,
A pyramidal attention-based transformer model based on improved differential innovation
search algorithm and feature extraction for solar radiation prediction considering relevant
factors,
Renewable Energy,
Volume 253,
2025,
123666,
ISSN 0960-1481,
[Link]
([Link]
Abstract: Accurate prediction of solar radiation intensity (SI) is crucial for power system
scheduling and site selection of photovoltaic power stations. This paper proposes a
multivariable solar radiation prediction model based on Time-varying Filter-based Empirical
Mode Decomposition (TVFEMD), Fuzzy Entropy (FE), Random Forest (RF), Triangular
Wandering Strategy Improved Differential Creative Search Algorithm (TDCS), and Pyramidal
Attention-based Transformer (Pyraformer). First, the Random Forest is used for feature
extraction of solar radiation data; then, TVFEMD is employed to break down solar radiation
data into constituent sub-modes, thereby mitigating the non-stationarity present in the data
sequence. Fuzzy Entropy-based aggregation is utilized to diminish the quantity of data
sequences, which, together with various features, forms a multivariable input feature
matrix. The Differential Creative Search Algorithm (DCS) is enhanced with a Triangular
Wandering Strategy to obtain Improved TDCS, which optimizes the hyperparameters of
Pyraformer for solar radiation prediction, thereby enhancing the model's predictive
performance. This paper analyzes the prediction metrics of the TDCS-RF-TVFEMD-FE-
Pyraformer multivariate model compared with nine other multivariate benchmark models.
The results show TVFEMD, FE, and RF boost models accuracy. Post TDCS optimization,
Pyraformer's RMSE and MAE outperform the baseline models by 10 %–50 %, with R and
SMAPE also outperforming the baseline models.
Keywords: Solar radiation prediction; Multivariate prediction; Pyraformer; Differential
creative search algorithm optimization

Chuanmao Fu, Meng Li, Bo Zhang, Hongbo Wang,


TBiSeg: A transformer-based network with bi-level routing attention for inland waterway
segmentation,
Ocean Engineering,
Volume 311, Part 2,
2024,
119011,
ISSN 0029-8018,
[Link]
([Link]
Abstract: Unmanned surface vehicles (USVs) for inland waterways have recently attracted
increasing attention in various fields. Accurate detection in navigable regions is crucial for
ensuring USV safety in autonomous navigation. However, the complex and variable
environment of inland waterways, such as confusable textures and irregular edge details,
continues to pose some problems in existing methods. Therefore, to acquire navigable
regions, this study proposed TBiSeg, a Vision Transformer-based efficient inland waterway
segmentation network, for obtaining pixel-level results. Bi-level routing attention is used to
improve the Transformer block, which enhances the understanding of inland water textures.
Additionly, this study combined global and local attention through a hierarchical encoder–
decoder architecture. To simulate inland waterway scenes as accurately as possible, this
study used two representative public datasets for data integration and data augmentation,
and conducted testing and cross-validating using multiple inland waterway datasets. Results
demonstrated that the model performed better than current state-of-the-art models in
segmentation accuracy and robustness in complex inland waterway environments while
showing impressive generalization. The datasets and code used in this paper is available at
[Link]
Keywords: Inland waterway segmentation; Vision transformer; Deep learning; Attention
mechanism

Saeid Dehghani-Dehcheshmeh, Mehdi Akhoondzadeh, Saeid Homayouni,


Oil spills detection from SAR Earth observations based on a hybrid CNN transformer
networks,
Marine Pollution Bulletin,
Volume 190,
2023,
114834,
ISSN 0025-326X,
[Link]
([Link]
Abstract: Oil spills are the main threats to marine and coastal environments. Due to the
increase in the marine transportation and shipping industry, oil spills have increased in
recent years. Moreover, the rapid spread of oil spills in open waters seriously affects the
fragile marine ecosystem and creates environmental concerns. Effective monitoring, quick
identification, and estimation of the volume of oil spills are the first and most crucial steps
for a successful cleanup operation and crisis management. Remote Sensing observations,
especially from Synthetic Aperture Radar (SAR) sensors, are a very suitable choice for this
purpose due to their ability to collect data regardless of the weather and illumination
conditions and over far and large areas of the Earth. Owing to the relatively complex nature
of SAR observations, machine learning (ML) based algorithms play an important role in
accurately detecting and monitoring oil spills and can significantly help experts in faster and
more accurate detection. This paper uses SAR images from ESA's Copernicus Sentinel-1
satellite to detect and locate oil spills in open waters under different environmental
conditions. To this end, a deep learning framework has been presented to identify oil spills
automatically. The SAR images were segmented into two classes, the oil slick and the
background, using convolutional neural networks (CNN) and vision transformers (ViT).
Various scenarios for the proposed architecture were designed by placing ViT networks in
different parts of the CNN backbone. An extensive dataset of oil spill events in various
regions across the globe was used to train and assess the performance of the proposed
framework. After the detection performance assessments, the F1-score values for the
standard DeepLabV3+, FC-DenseNet, and U-Net networks were 75.08 %, 73.94 %, and 60.85,
respectively. In the combined networks models (combination of CNN and ViT), the best F1-
score results were obtained as 78.48 %. Our results showed that these hybrid models could
improve detection accuracy and have a high ability to distinguish oil spill borders even in
noisy images. Evaluation metrics are increased in all the combined networks compared to
the original CNN networks.
Keywords: Oil spill detection; Deep convolutional neural networks; Vision transformers;
Synthetic aperture radar

Shengchun Wang, Haowen Li, Lianye Liu, Ronghui Cai, Zhonghai Yin, Huijie Zhu,
TSFI-Fusion: A dual-branch decoupled infrared and visible image fusion network based on
transformer and spatial-frequency interaction,
Optics and Lasers in Engineering,
Volume 195,
2025,
109287,
ISSN 0143-8166,
[Link]
([Link]
Abstract: Infrared and visible image fusion (IVIF) aims to generate high-quality images by
combining detailed textures from visible images with the target-highlight capabilities of
infrared images. However, many existing methods struggle to capture both shared and
unique features of each modality. They often focus only on spatial domain fusion, such as
pixel averaging, while overlooking valuable frequency domain information. This makes it
hard to retain fine details. To overcome these limitations, we propose TSFI-Fusion, a dual-
branch network that combines Transformer-based global understanding with spatial-
frequency detail enhancement. The two branches include a Transformer-based semantic
construction branch for capturing global features and a detail enhancement branch utilizing
an invertible neural network (INN) and a frequency domain compensation module (FDCM)
to integrate spatial and frequency information. We also design a dual-domain interaction
module (DDIM) to improve feature correlation across domains and a collaborative
information integration module (CIIM) to effectively merge features from both branches.
Additionally, we introduce a focal frequency loss to guide the model in learning important
frequency information. Experimental results demonstrate that TSFI-Fusion outperforms
existing methods across multiple datasets and metrics on the IVIF task. In downstream
applications such as object detection, it effectively enhances performance. Furthermore,
extended experiments on the MIF task reveal the robust generalization ability of the
proposed mechanism across diverse fusion scenarios. Our code will be available at
[Link]
Keywords: Infrared and visible image fusion; Transformer; Spatial-frequency information;
Medical image fusion

Murat Hi̇çyılmaz,
Bayesian optimization-enhanced vision transformer for damage detection in steel truss
structures using continuous wavelet transform analysis,
Structures,
Volume 85,
2026,
111059,
ISSN 2352-0124,
[Link]
([Link]
Abstract: The main contribution of this paper is to optimize the hyperparameters of the
Vision transformer (ViT) architecture using an adaptive Bayesian optimization process that
provides superior performance compared to state-of-the-art models in terms of accuracy
and efficiency while remaining within acceptable limits in terms of computational efficiency.
For this purpose, a Bayesian-optimized Image Transformer model (CwBot) is proposed for
damage detection in steel truss structures using continuous wavelet transform (CWT)
analysis. CWT spectrograms generated from a 72-bar steel truss model and measurements
of a real steel truss bridge are considered to better reveal the complex dynamic behavior of
steel truss structures. The ViT hyperparameters including patch size, embedding size,
transformer depth, multi-head attention heads, MLP size, learning rate, and batch size are
optimized and compared with seven different state-of-the-art transfer learning models. The
proposed model achieves 93.29 % accuracy for the 72-bar steel truss and 98.43 % accuracy
for the real steel truss bridge, outperforming other models.
Keywords: ViT; Bayesian optimization; CWT, Steel trusses; Damage detection

Jing Jiao, Hu Wang, Yamin Dang, Yingying Ren, Caiya Yue, Xiujuan Wu, Haomeng Cui, Xinlin
Wang,
Noise-resilient GNSS coordinate time series prediction using AVMD-sLSTM-transformer
hybrid model,
Advances in Space Research,
Volume 76, Issue 11,
2025,
Pages 6863-6881,
ISSN 0273-1177,
[Link]
([Link]
Abstract: High-precision modeling and analysis of Global Navigation Satellite System (GNSS)
coordinate time series constitute a fundamental basis for geophysical and crustal
deformation studies. To overcome the accuracy limitations of traditional methods under
noisy GNSS observations, we introduce a hybrid AVMD-sLSTM-Transformer prediction
model. The approach is validated using over a decade of observations from 584 global GNSS
stations. Initially, an improved Adaptive Variational Mode Decomposition (AVMD) method is
applied to decompose the preprocessed time series, with power spectral density analysis
guiding signal reconstruction and effectively suppressing noise. The reconstructed signals are
then input into a hybrid architecture combining scalar Long Short-Term Memory (sLSTM)
networks and Transformer models. This architecture leverages the sLSTM’s ability to capture
local temporal dynamics and the Transformer’s strength in modeling global spatiotemporal
dependencies, enabling robust long-term forecasting. The results indicate that significant
advantages in the north, east, vertical components, particularly achieving sub-millimeter
accuracy in the east direction. The proposed model reduces prediction errors by 65 %
compared to sLSTM only implementations and demonstrates 32 % improvement over
standard alone Transformer models. When compared to conventional Support Vector
Regression (SVM) and Autoregressive Integrated Moving Average (ARIMA) approaches, it
achieves over 80 % enhancement in mean absolute error (MAE) metrics. For horizontal
components, the model exhibits exceptional stability, validating the efficacy of the hybrid
architecture in synergizing sLSTM’s temporal modeling with Transformer’s global attention
mechanism. Notably, the model maintains residual errors within ±1 mm for stations affected
by coseismic offset in plate boundary zones, demonstrating robust adaptability to non-
stationary crustal deformation patterns.
Keywords: GNSS time series prediction; AVMD; sLSTM; Transformer

Jiajing Xie, Ying Chen, Shijie Luo, Wenxian Yang, Yuxiang Lin, Liansheng Wang, Xin Ding,
Mengsha Tong, Rongshan Yu,
Tracing unknown tumor origins with a biological-pathway-based transformer model,
Cell Reports Methods,
Volume 4, Issue 6,
2024,
100797,
ISSN 2667-2375,
[Link]
([Link]
Abstract: Summary
Cancer of unknown primary (CUP) represents metastatic cancer where the primary site
remains unidentified despite standard diagnostic procedures. To determine the tumor origin
in such cases, we developed BPformer, a deep learning method integrating the transformer
model with prior knowledge of biological pathways. Trained on transcriptomes from 10,410
primary tumors across 32 cancer types, BPformer achieved remarkable accuracy rates of
94%, 92%, and 89% in primary tumors and primary and metastatic sites of metastatic
tumors, respectively, surpassing existing methods. Additionally, BPformer was validated in a
retrospective study, demonstrating consistency with tumor sites diagnosed through
immunohistochemistry and histopathology. Furthermore, BPformer was able to rank
pathways based on their contribution to tumor origin identification, which helped to classify
oncogenic signaling pathways into those that are highly conservative among different
cancers versus those that are highly variable depending on their origins.
Keywords: cancer of unknown primary; CUP; biological pathway; transformer; tracing the
origin of cancer

Guokang Xu, Jianchuan Yin, Nini Wang, Zeguo Zhang,


Maritime man-overboard search using a lightweight and efficient end-to-end detection
transformer,
Journal of Safety Science and Resilience,
Volume 7, Issue 2,
2026,
100267,
ISSN 2666-4496,
[Link]
([Link]
Abstract: Maritime transportation plays a crucial role in global economic trade. However,
maritime accidents occur frequently, posing significant threats to the safety of seafarers. In
search and rescue scenarios for man-overboard, unmanned aerial vehicles (UAVs) are
gradually replacing manned aircraft and helicopters. Besides, detecting man-overboard is
challenging because of their small pixel size, weak signals, and indistinct features on the
ocean surface. Furthermore, existing detectors struggle to strike a balance between
lightweight design for UAVs and detection accuracy. To address this issue, the novel Man-
overboard Detection Transformer (MOB-DETR) is proposed. On the one hand, the Token
Enhancement layer is introduced, which conducts fine-grained filtering of spatial and
channel dimensions, reducing redundant encoding caused by background queries. On the
other hand, the Effusion Fusion Module, based on the RepViT Block, is proposed, effectively
eliminating computational redundancy by decoupling the interaction mechanisms between
spatial and channel dimensions. Additionally, to fill the existing gap in benchmark datasets
for detecting man-overboard, the ManOverboard benchmark dataset has been established.
In the experimental validation phase, MOB-DETR is conducted on ManOverboard and
SeaDronesSeev2. Ablation experiments show that MOB-DETR achieves 11.7 % better
lightweight performance and 14.4 % higher APsmall than baselines. Comparison
experiments on ManOverboard and SeaDronesSeev2 validate its effectiveness, offering an
efficient solution for man-overboard detection. Overall, this research not only advances
man-overboard detection but also significantly enhances the resilience of maritime
transportation, ultimately protecting seafarers' lives and ensuring the reliability of the
world's essential trade routes.
Keywords: Maritime safety; Small object detection; Maritime search; Man-overboard
detection; ManOverboard

Jiayi Cai, Zhaocheng Yang, Ping Chu, Juntao Guo, Jianhua Zhou,
Robust hand gesture detection and recognition using 4D millimeter-wave radar in a
ubiquitous scene,
Measurement,
Volume 253, Part C,
2025,
117545,
ISSN 0263-2241,
[Link]
([Link]
Abstract: In current research on HGR using radar sensors, hand gestures are typically
confined to a smaller region. However, in ubiquitous scenarios, unrestricted human body
movements and unexpected hand gesture motions usually occur, which results in a large
false alarms and recognition performance degradation. To address this issue, we propose a
robust hand gesture detection and recognition method in ubiquitous scenarios using
Frequency-Modulated Continuous Wave (FMCW) Multiple-Input Multiple-Output (MIMO)
radar. The core idea is to progressively define and classify motions in a cascaded manner,
gradually filtering out non-specific movements, reducing false positives, and enhancing the
applicability of HGR. Specifically, we first propose a suspected hand gesture motion
detection method to help identify suspicious hand gestures. Then, the velocity and position
features of the mutated signal and the stable signal are extracted. A mutated signal motion
recognition method based on a single-layer long short-term memory (LSTM) network is used
to effectively distinguish non-hand gesture motions from hand gestures. Finally, the two-
dimensional trajectory features are extracted, and cascaded with a LSTM network combined
a Gaussian probability model is developed to enhance the ability of open-set recognition.
Experimental results show that the proposed method can achieve the recognition accuracy
of 99.53% for designed hand gestures, the false alarm rate of 1.5% for unexpected hand
gestures and 0.11% for non-hand gesture motions.
Keywords: Hand gesture recognition; Non-hand gesture motion; Feature extraction;
Probability models; Ubiquitous scene

Liang Li, Guochu Chen, Haiyan Wang, Baojiang Li, Bin Wang, Zizhen Yi, Chunbo Zhao,
VHTformer: A joint query perception method for visual-haptic-textual information based on
Transformer,
Applied Soft Computing,
Volume 181,
2025,
113529,
ISSN 1568-4946,
[Link]
([Link]
Abstract: Multimodal information fusion research struggles with aligning heterogeneous
modalities and addressing data imbalance, especially when integrating visual, haptic, and
text—three modalities offering complementary perceptual and semantic features. Current
research focuses on Transformers for unimodal and vision-haptics bimodal tasks, neglecting
tri-modal integration. Leveraging text's semantic bridging capacity could address this
limitation in cross-sensory learning. We propose VHTformer, a Transformer-based
framework designed to unify visual, haptic, and textual modalities via joint query learning.
The model leverages hierarchical attention mechanisms: self-attention refines intra-modal
features (e.g., extracting texture from haptic signals or contextual semantics from text).
Meanwhile, cross-attention aligns spatial-semantic patterns across modalities through
learnable joint queries. This enables synergistic fusion of geometric shapes (vision), material
properties (haptics), and descriptive attributes (text). Experiments were conducted on three
multimodal datasets—ObjectFolder 2.0, Touch and Go, and ObjectFolder Real—covering a
total of 100 + object categories with diverse material and shape properties. To mitigate class
imbalance and ensure statistical reliability, we adopted stratified 5-fold cross-validation. In
addition, we conducted robustness evaluations under Gaussian noise injection to verify the
model's robustness. VHTformer achieves up to 99.55 % recognition accuracy and
demonstrates strong robustness, highlighting the value of tri-modal integration for
comprehensive object understanding.
Keywords: Multimodal information recognition; Cross-modal fusion; Transformer; Attention
mechanism; Semantic alignment

Yongxin Li, Yukun Xue, Zhihui Xin, Guisheng Liao, Penghui Huang,
Multi-modal cross Swin transformer network for multi-label classification landslide detection
with optical and SAR images of Luding,
International Journal of Applied Earth Observation and Geoinformation,
Volume 145,
2025,
104954,
ISSN 1569-8432,
[Link]
([Link]
Abstract: Natural phenomena such as earthquakes and heavy rainfall can trigger landslide
events hundreds or even thousands of times in a given area. Therefore, rapid and efficient
intervention in affected regions is essential. Existing studies have demonstrated good
detection performance using optical remote sensing images; however, optical data has
significant limitations under cloud cover. Moreover, current multi-modal fusion methods
struggle to effectively capture the nonlinear interactions between different modes when
processing data with large informational disparities. To address these issues, we developed
the China Luding multi-modal landslide dataset, which includes optical and polarimetric
synthetic aperture radar (PolSAR) images. We also propose a multi-modal network for
landslide detection based on multi-label classification, called multi-modal cross Swin
Transformer network (MCSTNet). Additionally, a new weighted asymmetric loss function
(WASL) is proposed to improve multi-label classification tasks. Our proposed network
consists of two stages: feature extraction and feature fusion. In the first stage, two
independent branches extract high-level semantic features from optical and PolSAR images.
In the second stage, a multi-modal cross multi-head self-attention (MCMSA) mechanism
fuses the high-level semantic features from the multi-modal information. Therefore, the
model enhances the learning and feature representation capabilities through parallel
processing the input data information in different representation subspaces. Compared to
other recent landslide detection methods, our results indicate that the proposed method is
more effective, achieving an F1 score of 88.24% for landslide detection, with greater
accuracy in identifying landslide-prone areas. Therefore, our model can provide reliable and
effective decision support for disaster emergency responses.
Keywords: Landslide detection; Multi-label classification; Luding multi-modal landslide
dataset; MCSTNet; WASL

Purabi Sharma, Kandarpa Kumar Sarma,


Collaborative and distributive intelligence in radar systems: Enhancing electronic jamming
discrimination,
Computers and Electrical Engineering,
Volume 124, Part A,
2025,
110357,
ISSN 0045-7906,
[Link]
([Link]
Abstract: The precise discrimination of radar jamming signals is decisive in executing
effective electronic counter-countermeasures (ECCM). Data-driven deep learning (DL)
models have proven effective for this task, but challenges persist in addressing critical real-
world issues such as limited data sharing, time- and location-dependent variations in hostile
interferences, real-time adaptability to evolving jamming tactics, and imbalanced data
distribution. In this paper, a novel approach is proposed that employs federated learning (FL)
for training radar signal jamming classifiers individually on a set of devices, aggregating,
sharing, and continuously updating the knowledge. This approach enables privacy-
preserving training, eliminating access to client-local data or centralized data storage,
ensuring knowledge sharing, and resilient response to jamming. This work delineates a
collaborative and distributive learning framework for radar jamming signals, employing two
hybrid models within an FL platform. A deep spectra spatio-temporal discriminator (DSSTD)
and a shallow spectra-temporal feed-forward self-attention-driven discriminator (SSTFSAD)
network have been implemented as classifiers on a distributed arrangement. The time–
frequency attributes of radar jamming signals as 2D-scalograms individually train these
models across remote FL nodes. Extensive evaluations are conducted in FL environments
under independently and identically distributed (IID) and non-IID data configurations,
simulating real-world settings with diverse data distributions. Experimental results
demonstrate that both proposed approaches are effective, with FL-driven DSSTD
outperforming FL-driven SSTFSAD by 7.7% in IID and 6.8% in non-IID setups, even at -5 dB
JNR. These results highlight the robustness and adaptability of the FL-driven DSSTD model,
offering a significant advancement in radar jamming signal discrimination for electronic
warfare applications.
Keywords: Federated learning; Radar jamming signal; Multi-scale CWT; Deep learning; Long
short-term memory; Self-attention

Bochao Zou, Zizheng Guo, Jiansheng Chen, Junbao Zhuo, Weiran Huang, Huimin Ma,
RhythmFormer: Extracting patterned rPPG signals based on periodic sparse attention,
Pattern Recognition,
Volume 164,
2025,
111511,
ISSN 0031-3203,
[Link]
([Link]
Abstract: Remote photoplethysmography (rPPG) is a non-contact method for detecting
physiological signals based on facial videos, holding high potential in various applications.
Due to the periodicity nature of rPPG signals, the long-range dependency capturing capacity
of the transformer was assumed to be advantageous for such signals. However, existing
methods have not conclusively demonstrated the superior performance of transformers
over traditional convolutional neural networks. This may be attributed to the quadratic
scaling exhibited by transformer with sequence length, resulting in coarse-grained feature
extraction, which in turn affects robustness and generalization. To address that, this paper
proposes a periodic sparse attention mechanism based on temporal attention sparsity
induced by periodicity. A pre-attention stage is introduced before the conventional attention
mechanism. This stage learns periodic patterns to filter out a large number of irrelevant
attention computations, thus enabling fine-grained feature extraction. Moreover, to address
the issue of fine-grained features being more susceptible to noise interference, a fusion stem
is proposed to effectively guide self-attention towards rPPG features. It can be easily
integrated into existing methods to enhance their performance. Extensive experiments show
that the proposed method achieves state-of-the-art performance in both intra-dataset and
cross-dataset evaluations. The codes are available at
[Link]
Keywords: Remote physiological measurement; Periodic sparse attention

Yiheng Xie, Xiaoping Rui, Yarong Zou, Heng Tang, Ninglei Ouyang, Yingchao Ren,
STDPNet: supervised transformer-driven network for high-precision oil spill segmentation in
SAR imagery,
International Journal of Applied Earth Observation and Geoinformation,
Volume 143,
2025,
104812,
ISSN 1569-8432,
[Link]
([Link]
Abstract: Oil spill incidents are one of the major factors damaging marine ecosystems, and
there is an urgent need for effective detection and identification technologies to quickly
locate oil spill contamination areas. Synthetic Aperture Radar (SAR) is capable of monitoring
the ocean surface under various weather and lighting conditions, but the SAR images often
contain dense speckle noise, and popular SAR oil spill image datasets typically lack sufficient
polarization information. To overcome these issues, this study introduces a novel
polarimetric decomposition method to generate synthetic color image datasets that
integrate multiple polarization features, thereby enhancing image texture and contrast. An
image denoising module is designed, which reduces noise interference in the color images
through an adaptive sampling approach. Furthermore, a novel Transformer-CNN
architecture model is proposed, integrating two modules: the Super Visual Attention
Transformer and the Directional Multi-Branch Scale Self-Calibration Module. The
segmentation performance of the model is comprehensively evaluated on three datasets,
and compared with state-of-the-art segmentation methods, demonstrating superior
classification accuracy and stability. This research provides an effective technical support for
accurate oil spill detection and marine ecosystem protection.
Keywords: Oil spill segmentation; SAR imagery; Proximity pixel compensation method;
Transformer-CNN model; Super visual attention transformer; Directional multi-branch scale
self-calibration method

Yuhao Wu, Bin Li, Jun Li, Yonglou Liang, Naiqiang Zhang, Anlai Sun,
Enhancing nighttime cloud detection for moderate resolution imagers using a transformer
based deep learning network,
Remote Sensing of Environment,
Volume 332,
2026,
115067,
ISSN 0034-4257,
[Link]
([Link]
Abstract: Accurate cloud detection is essential for the quantitative applications of satellite
imager observations, but nighttime cloud detection has challenges due to limited spectral
bands, for example, physical methods using only infrared (IR) bands without using spatial
textures as input for cloud detection often result in high uncertainties, especially in some
situations such as cryosphere surface. Although numerous segmentation-style deep learning
cloud detection algorithms have proposed in previous studies, they are inadequate for
nighttime due to the difficulty in acquiring two-dimensional truth data for training and
validation. To overcome these challenges, the Transformer based Nighttime Cloud Detection
(TNCD) framework, which integrates spatial features and utilizes an advanced Transformer
architecture with relative position encoding, layer scaling, and channel attention
mechanisms, is proposed and investigated for nighttime cloud detection. The model was
trained on labels derived from CALIOP data, utilizing a dataset comprising nearly one
hundred million segments from MODIS. Independent validation indicates that TNCD
achieves robust and consistent performance across various scenarios, with an overall
accuracy (OA) of 93.26 % and over 90 % in cryosphere regions. The proposed algorithm
avoids the pattern noise appeared in the traditional physical methodology due to the
utilization of auxiliary data at coarser resolutions, it also mitigates the negative impact of
stripes in IR images for cloud detection. Moreover, TNCD shows high transferable
practicability across sensors, with over 90 % OA for MERSI. More importantly, our research
underscores the importance of water vapor absorption bands for nighttime cloud detection
over the cryosphere. TNCD's high accuracy and robustness provide unique methodology that
could be used operationally for nighttime cloud detection.
Keywords: Nighttime cloud detection; Transformer; Deep learning; MODIS; MERSI; CALIPSO

Yizhen Jia, Hui Chen, Bang Huang, WenKai Jia, Wen-Qin Wang,
Riemannian gradient deep network for joint waveform and filter optimization in MIMO radar
against chopping forwarding jamming and clutter,
Signal Processing,
Volume 239,
2026,
110257,
ISSN 0165-1684,
[Link]
([Link]
Abstract: With the rise of digital radio frequency memory technology, active deception
jamming poses a significant threat to radar systems, especially in detecting targets amid
mainlobe jamming and non-Gaussian clutter. Traditional methods like space–time matched
filtering struggle in such scenarios. This study introduces the Riemannian gradient deep
network (RGDN), a framework for joint optimization of transmit waveforms and receive
filters to improve target detection. Unlike conventional signal-to-clutter noise ratio (SCNR)
maximization, RGDN leverages information geometry to maximize the Kullback–Leibler
Divergence (KLD) between targets and clutter. By modeling non-Gaussian data with a
Gaussian mixture distribution and constructing a Riemannian manifold, the framework
achieves effective jamming suppression through receive filter term in the loss function,
minimizing jamming effects while enhancing target-clutter distinguishability. To address non-
convex optimization, Riemannian gradient descent is integrated into a deep network.
Numerical experiments show that RGDN achieves superior detection performance compared
to SCNR maximization method.
Keywords: Riemannian gradient; KL divergence; Waveform design; MIMO radar; Mainlobe
deception jamming; Deep learning

Jing Wang, Chao Li, Lu Li, Zhihua Huang, Chao Wang, Hong Zhang, Zhengjia Zhang,
InSAR time-series deformation forecasting surrounding Salt Lake using deep transformer
models,
Science of The Total Environment,
Volume 858, Part 2,
2023,
159744,
ISSN 0048-9697,
[Link]
([Link]
Abstract: The free and open data policy of Sentinel-1 SAR images enables Radar
interferometry (InSAR) to perform time series surface deformation monitoring over large
areas. InSAR deformation monitoring and prediction can investigate the freeze-thaw cycles
of permafrost on the Qinghai-Tibet Plateau. However, the convolutional and recurrent
neural networks cannot accurately model long-term and complex relations in multivariate
time series data, it is challenging to implement time series deformation prediction with high
spatial resolution. In this paper, an innovative InSAR deformation prediction integrated
algorithm based on the transformer models is proposed to predict time series deformation
more accurately surrounding Salt Lake. Compared with the other solutions, the unique
feature of the proposed method is that: 1) this method takes advantage of the self-attention
mechanism to study complicated dynamic deformation features of permafrost caused by
temperature and other variables from InSAR time series deformation. 2) The transformer-
based model can more accurately simulate seasonal and non-seasonal deformation signals,
and is effective for short-term prediction of surface deformation in permafrost areas. The
InSAR deformation prediction results demonstrate that the InSAR deformation prediction
method achieves better prediction performance in predicting the deformation trends of
permafrost with a point scale compared with the prediction results of other models. Based
on the predicted deformation and the water extraction results, the expansion trends
surrounding Salt Lake are discussed and evaluated. The total area of Salt Lake increased by
57.32 km2 during the period 2015–2019. And Salt Lake maintained slowing expansion trend
from 2019 to 2022. The time series deformation forecasting method can be used as a
generic framework for modeling nonlinear deformation processes in complex permafrost
areas, and it reveals the potential impact of the Salt Lake outburst event on the deformation
processes and the degradation of permafrost.
Keywords: InSAR; Qinghai-Tibet Plateau; Deformation prediction; Transformer; Salt Lake;
Permafrost

Jinsheng Fan, Guo-An Yu, Mingmeng Zhao, Hucheng Zong,


MCPT-CAF-BiGRU: A multi-scale CNN and ProbSparse-Masked Transformer model with cross-
attention fusion and BiGRU for hourly wind speed forecasting,
Expert Systems with Applications,
Volume 307,
2026,
131081,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Wind speed forecasting is essential for reliable wind power integration but remains
challenging due to the inherent non-stationarity, multi-scale variability, and spatial
heterogeneity of wind fields. To address these issues, we propose MCPT-CAF-BiGRU, a hybrid
deep learning framework that combines Multi-scale Convolutional Neural Networks
(MCNN), ProbSparse-Masked Transformer, Bidirectional Gated Recurrent Unit (BiGRU), and a
Cross-Attention Fusion (CAF) mechanism. The MCNN captures local disturbances across
multiple temporal scales, while the ProbSparse-Transformer efficiently models long-range
dependencies, and BiGRU further enhances bidirectional contextual representation. The CAF
mechanism adaptively integrates local and global feature streams, improving robustness
under turbulent dynamics. Extensive experiments on four real-world wind farm datasets at
three heights (10, 30, and 50 m) show that MCPT-CAF-BiGRU consistently outperforms
eleven benchmark models in terms of RMSE, MAE, MAPE, and NSC, demonstrating superior
accuracy and generalization across diverse wind regimes. Ablation studies confirm the
contribution of each architectural component. These findings establish MCPT-CAF-BiGRU as
an effective solution for multi-altitude wind speed forecasting, providing methodological
innovation and practical support for intelligent wind energy scheduling.
Keywords: Wind speed forecasting; MCNN; PSM-Transformer; BiGRU; Cross-attention fusion

B M Tazbiul Hassan Anik, Zubayer Islam, Mohamed Abdel-Aty,


A time-embedded attention-based transformer for crash likelihood prediction at
intersections using connected vehicle data,
Transportation Research Part C: Emerging Technologies,
Volume 169,
2024,
104831,
ISSN 0968-090X,
[Link]
([Link]
Abstract: The real-time crash likelihood prediction model is an essential component of the
proactive traffic safety management system. Over the years, numerous studies have
attempted to construct a crash likelihood prediction model in order to enhance traffic safety,
but mostly on freeways. In the majority of the existing studies, researchers have primarily
used a deep learning-based framework to identify crash potential. Lately, Transformers have
emerged as a potential deep neural network that fundamentally operates through attention-
based mechanisms. Transformers exhibit distinct functional benefits over established deep
learning models like Recurrent Neural Networks (RNNs), Long Short-Term Memory networks
(LSTMs), and Convolutional Neural Networks (CNNs). First, they employ attention
mechanisms to accurately weigh the significance of different parts of input data, a dynamic
functionality that is not available in RNNs, LSTMs, and CNNs. Second, they are well-equipped
to handle dependencies over long-range data sequences, a feat RNNs typically struggle with.
Lastly, unlike RNNs, LSTMs, and CNNs, which process data in sequence, Transformers can
parallelly process data elements during training and inference, thereby enhancing their
efficiency. Apprehending the immense possibility of Transformers, this paper proposes
inTersection-Transformer (inTformer), a time-embedded attention-based Transformer model
that can effectively predict intersection crash likelihood in real-time. The inTformer is
basically a binary prediction model that predicts the occurrence or non-occurrence of
crashes at intersections in the near future (i.e., next 15 min). The proposed model was
developed by employing traffic data extracted from connected vehicles. Acknowledging the
complex traffic operation mechanism at intersection, this study developed zone-specific
models by dividing the intersection region into two distinct zones: within-intersection and
approach zones, each representing the intricate flow of traffic unique to the type of
intersection (i.e., three-legged and four-legged intersections). In the ‘within-intersection’
zone, the inTformer models attained a sensitivity of up to 73%, while in the ‘approach’ zone,
the sensitivity peaked at 74%. Moreover, benchmarking the optimal zone-specific inTformer
models against earlier studies on crash likelihood prediction at intersections and several
established deep learning models trained on the same connected vehicle dataset confirmed
the superiority of the proposed inTformer. Further, to quantify the impact of features on
crash likelihood at intersections, the SHAP (SHapley Additive exPlanations) method was
applied on the best performing inTformer models. The most critical predictors were average
and maximum approach speeds, average and maximum control delays, average and
maximum travel times, split failure percentage and count, and percent arrival on green.
Keywords: Connected Vehicles; Transformer; Intersection Safety; Real-Time Crash Likelihood;
Traffic Safety

Menghao Du, Zhenfeng Shao, Xiongwu Xiao, Jindou Zhang, Duowang Zhu, Jinyang Wang,
Timo Balz, Deren Li,
High-precision flood change detection with lightweight SAR transformer network and
context-aware attention for enriched-diverse and complex flooding scenarios,
ISPRS Journal of Photogrammetry and Remote Sensing,
Volume 231,
2026,
Pages 507-531,
ISSN 0924-2716,
[Link]
([Link]
Abstract: Floods are highly destructive natural disasters that threaten both society and the
environment. Given the all-weather, all-time imaging capability of synthetic aperture radar
(SAR), analyzing flood events using SAR imagery across diverse scenarios is essential for
developing high-precision and robust detection models. However, existing transformer-
based change detection methods achieve high precision, but their high computational cost
and large parameter sizes necessitate lightweight design while maintaining detection
accuracy. Moreover, existing studies focus on a few specific scenarios without thorough
validation and in-depth analysis of model strengths across diverse flood conditions with
imbalanced inundation ratios. To address these challenges, this paper proposes an adaptive
window and context-aware attention network (AWCA-Net) for high-precision SAR-based
flood change detection under diverse flooding scenarios, achieving a lightweight model
while maintaining the highest detection accuracy. AWCA-Net has three key advantages:
Firstly, the neighborhood feature enhancement module with contextual information (NECM)
strengthens the discrimination of subtle and heterogeneous flood changes. Secondly, the
large kernel grouping attention gate module based on high-low layer feature difference
(LGDM) leverages difference-weighted attention to effectively guide the selection of flood-
relevant features. Thirdly, the multi-scale convolutional attention module based on adaptive
window selection (MSAWM) dynamically adjusts kernel sizes to capture diverse flood change
patterns. To better train and evaluate AWCA-Net, we constructed VarFloods, the first large-
scale and enriched-diverse benchmark dataset for flood change detection that spans five
continents and includes diverse regions, scenarios, land cover types, causes, years, and
inundation ratios, which includes both GRD and preprocessed versions. We evaluated
AWCA-Net’s performance on three datasets. We found that: (1) AWCA-Net achieves the
highest-precision while maintaining significantly lower computational cost, outperforming
other state-of-the-art (SOTA) methods. On the two representative public datasets and the
enriched-diverse benchmark dataset (VarFloods-G and VarFloods-P), AWCA-Net improves
the IoU by 11.59 % to 42.57 % over the basic model, 2.16 % to 11.48 % over an advanced
transformer-based model, and 1.31 % to 1.68 % over the best comparative model, while
maintaining a computational cost of only 17.53G, which is just 8.7 % to 60 % of existing SOTA
models. (2) Difference-guided attention enhances detection in complex background regions,
neighborhood-based fusion improves performance in irregular terrains, and adaptive
convolution contributes to stable results across diverse flood scenarios. And the proposed
AWCA-Net demonstrates strong generalization and stability under diverse flood scenarios
with imbalanced inundation ratios. The dataset and code of AWCA-Net will be released at:
[Link]
Keywords: High-precision flood change detection; Context-aware attention; Adaptive
window selection; Diverse flood scenarios; VarFloods dataset; Synthetic aperture radar (SAR)

Zai Zhang, Bin Shi, Kai Sun, Hao Wu, Bo Dong,


RADAR: Relation-assisted dual-graph aligning recognition for grounded multimodal named
entity recognition,
Information Processing & Management,
Volume 63, Issue 3,
2026,
104552,
ISSN 0306-4573,
[Link]
([Link]
Abstract: Grounded Multimodal Named Entity Recognition (GMNER) requires the
simultaneous identification of textual entities and their corresponding visual regions within
images. However, the inability to model visual contextual semantics and the disorganized
processing of cross-modal features often lead existing methods to struggle with both visual
entity differentiation and bridging the modality gap. We propose RADAR (Relation-Assisted
Dual-graph Aligning Recognition), a novel framework that leverages visual relations derived
from scene graphs to encode structured context and enhance visual understanding. To
achieve fine-grained cross-modal alignment, we design an object-level alignment self-
attention mechanism and introduce a dual-graph strategy. Evaluated on the Twitter-GMNER
dataset (13,076 image-text pairs), RADAR achieves 60.91 % F1 score, a +4.5 % improvement
over the H-Index baseline. The method also demonstrates consistent gains in subtasks, with
+3.54 % improvement in EEG metric, validating its effectiveness in multimodal entity
alignment.
Keywords: Multimodal named entity recognition; Visual grounding; Visual scene graph;
Cross-modal alignment

Yihao Wu, Zhenrong Li, Shize Duan, Xuanzhang He, Dawei Dong, Liyan Yu,
A 87.5 – 104.4 GHz broadband power amplifier with RF switch using transmission line
connected in parallel with MCR transformer in 130 nm SiGe BiCMOS,
Microelectronics Journal,
Volume 168,
2026,
106979,
ISSN 1879-2391,
[Link]
([Link]
Abstract: This paper presents a power amplifier (PA) and RF switch architecture designed for
W-band phased array transmit channels. To achieve wideband performance at high
frequencies, an output matching network based on a transmission line connected in parallel
with transformer is proposed, along with a method to improve the flatness of in-band
impedance matching. This approach significantly enhances the current-handling capacity of
the output matching network. To ensure adequate gain, the performance of various active
circuit topologies is analyzed. According to post-layout simulation results, the proposed
circuit achieves a peak gain of 19.8 dB and a 3 dB bandwidth ranging from 87.5 GHz to 104.4
GHz. At 93 GHz, despite the insertion loss introduced by the RF switch, the power amplifier
maintains a saturated output power(Psat) of 16 dBm and achieves a power-added efficiency
(PAE) of 8.4%. S11 remains below −13.2 dB throughout the entire 3 dB bandwidth, reaching
a minimum of −20.9 dB.
Keywords: Power amplifier; RF switch; Broadband; Transformer; Transmission line

Son Minh Nguyen, Duc Viet Le, Paul J.M. Havinga,


Seeing the world from its words: All-embracing Transformers for fingerprint-based indoor
localization,
Pervasive and Mobile Computing,
Volume 100,
2024,
101912,
ISSN 1574-1192,
[Link]
([Link]
Abstract: In this paper, we present all-embracing Transformers (AaTs) that are capable of
deftly manipulating attention mechanism for Received Signal Strength (RSS) fingerprints in
order to invigorate localizing performance. Since most machine learning models applied to
the RSS modality do not possess any attention mechanism, they can merely capture
superficial representations. Moreover, compared to textual and visual modalities, the RSS
modality is inherently notorious for its sensitivity to environmental dynamics. Such
adversities inhibit their access to subtle but distinct representations that characterize the
corresponding location, ultimately resulting in significant degradation in the testing phase. In
contrast, a major appeal of AaTs is the ability to focus exclusively on relevant anchors in RSS
sequences, allowing full rein to the exploitation of subtle and distinct representations for
specific locations. This also facilitates disregarding redundant clues formed by noisy ambient
conditions, thus enhancing accuracy in localization. Apart from that, explicitly resolving the
representation collapse (i.e., none-informative or homogeneous features, and gradient
vanishing) can further invigorate the self-attention process in transformer blocks, by which
subtle but distinct representations to specific locations are radically captured with ease. For
that purpose, we first enhance our proposed model with two sub-constraints, namely
covariance and variance losses at the Anchor2Vec. The proposed constraints are
automatically mediated with the primary task towards a novel multi-task learning manner. In
an advanced manner, we present further the ultimate in design with a few simple tweaks
carefully crafted for transformer encoder blocks. This effort aims to promote representation
augmentation via stabilizing the inflow of gradients to these blocks. Thus, the problems of
representation collapse in regular Transformers can be tackled. To evaluate our AaTs, we
compare the models with the state-of-the-art (SoTA) methods on three benchmark indoor
localization datasets. The experimental results confirm our hypothesis and show that our
proposed models could deliver much higher and more stable accuracy.
Keywords: RSS fingerprints; Indoor localization; Deep learning; Transformers

Haoming Feng, Huaqing Li, Wenwen Zhu, Denghao Li, Yukun Huang,
Micro-motion enhanced multi-person activity recognition with millimeter-wave radar,
Measurement,
Volume 258, Part B,
2026,
119090,
ISSN 0263-2241,
[Link]
([Link]
Abstract: As a non-contact sensing device, millimeter-wave radar exhibits unique strengths
in human activity recognition (HAR). Existing methods rely on micro-Doppler signatures for
activity classification, but they often encounter feature aliasing in multi-person activity
recognition (MPAR) scenarios. Although point cloud-based approaches can distinguish
individual targets, they primarily extract static morphological features, neglecting the micro-
motion information of human joints, which is crucial for accurate activity recognition. To
address these limitations, we proposes an innovative MPAR framework that integrates
spatial point clouds and micro-motion features. First, an improved point cloud data
association algorithm is applied to achieve multi-target point cloud feature separation,
followed by a dynamic projection mechanism to construct time–Doppler feature maps.
Then, a torso micro-motion enhancement algorithm is designed to enhance the details of
human body movements. Finally, a CNN-LSTM hybrid network architecture with a temporal-
attention is constructed for action classification. Experimental results show that the
proposed micro-motion enhancement algorithm improves recognition accuracy by 27.1%
and 2.3%, compared to two traditional time–frequency analysis methods. Furthermore,
MPAR task in occlusion scenarios achieves recognition accuracy of 93.5%. In summary,
proposed framework not only retains the inherent advantages of millimeter-wave radar but
also significantly enhances multi-person activity recognition in complex scenarios.
Keywords: Human activity recognition (HAR); Multi-person activity recognition (MPAR);
Feature separability; Temporal attention; Occlusion scenarios

Rui He, Tian Peng, Xinyu Zhang, Zhigang Chen, Junhao Yao, Muhammad Shahzad Nazir, Chu
Zhang,
A novel hybrid model for state of health prediction in lithium batteries based on non-
stationary transformers optimized by tree-structured Parzen estimator considering health
factors,
Applied Energy,
Volume 402, Part C,
2026,
127030,
ISSN 0306-2619,
[Link]
([Link]
Abstract: Accurate prediction of State of Health (SOH) in lithium batteries is crucial for
improving the performance, prolonging the service life, preventing failures, and ensuring the
safe use of lithium batteries. This paper proposes a multivariate predictive correction model
for lithium battery SOH based on Time-Varying Filter Empirical Mode Decomposition
(TVFEMD), Pearson Correlation Coefficient (PCC), Kernel Principal Component Analysis
(KPCA), improved Bayesian algorithm, Non-stationary Transformers (NSTransformers), and
Regularized Online Sequential Extreme Learning Machine (ReOSELM). In order to reduce the
complexity of lithium battery data and health factors and to fully extract the features,
multiple methods are used for processing. Firstly, TVFEMD is used for the initial
decomposition of lithium battery health state data, then KPCA is applied to downsize the
decomposed data, and then PCC is selected for correlation analysis of the health factors to
select features with high correlation. Next, the NSTransformers model is employed for
predicting the lithium battery SOH, and a tree-structured Bayesian optimization algorithm,
namely, Tree-structured Parzen Estimator (TPE) is used to optimize the important
parameters of the NSTransformers model, enhancing the model's predictive performance.
Finally, the ReOSELM model is used to correct the initial prediction errors, and the initial
predicted values and error-corrected predicted values are summed to obtain the final
lithium battery SOH prediction values. This paper compares the prediction results of the
multivariate and univariate models. Compared with the other eight multivariate benchmark
models, the MAE and RMSE of the TVFEMD-PCC-KPCA-TPE-NSTransformers-ReOSELM
multivariate model proposed in this paper are reduced by about 0.1 %, and the R are
increased by more than 1 %, which verifies the superiority of the multivariate model
proposed in this paper in the prediction of lithium battery SOH.
Keywords: State of health of Lithium battery; TVFEMD; Tree-structured Parzen estimator;
Non-stationary transformers; ReOSELM

Runwei Guan, Shanliang Yao, Lulu Liu, Xiaohui Zhu, Ka Lok Man, Yong Yue, Jeremy Smith, Eng
Gee Lim, Yutao Yue,
Mask-VRDet: A robust riverway panoptic perception model based on dual graph fusion of
vision and 4D mmWave radar,
Robotics and Autonomous Systems,
Volume 171,
2024,
104572,
ISSN 0921-8890,
[Link]
([Link]
Abstract: With the development of Unmanned Surface Vehicles (USVs), the perception of
inland waterways has become significant to autonomous navigation. RGB cameras can
capture images with rich semantic features, but they would fail in adverse weather and at
night. As a perception sensor that has initially emerged in recent years, 4D millimeter-wave
radar (4D mmWave radar) can work in all weather and has more abundant point-cloud
features than ordinary radar, but it also suffers from water-surface clutter seriously.
Furthermore, the shape and outline of dense point cloud captured by 4D mmWave radar are
irregular. CNN-based neural networks treat features as 2D rectangle grids, which excessively
favor image modality and are unfriendly to radar modality. Therefore, we transform both
features of image and radar into non-Euclidean space as graph structures. In this paper, we
focus on robust panoptic perception in inland waterways. Firstly, we propose the first
Clutter-Point-Removal (CPR) algorithm for 4D mmWave radar, removing water-surface clutter
and improving the recall of radar targets. Secondly, we propose a high-performance
panoptic perception model based on the graph neural network called Mask-VRDet, fusing
features of vision and radar to simultaneously perform object detection and semantic
segmentation. To the best of our knowledge, Mask-VRDet is the first riverway panoptic
perception model based on vision-radar graphical fusion. It outperforms other single-modal
and fusion models, and achieves state-of-the-art performance on our collected dataset. We
release our code at [Link]
Keywords: Riverway panoptic perception; Fusion of vision and radar; Graph convolution
network; Radar clutter removal

M. Vidhyalakshmi, M.S. Siva Priya, K. Kumaran, Manivannan Kirankumar,


Efficient disaster detection using drone imagery: A depthwise separable convolutional
transformer with majority bounding box voting,
Engineering Applications of Artificial Intelligence,
Volume 165, Part B,
2026,
113490,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Effective disaster management heavily relies on the rapid and accurate
identification of affected individuals in disaster zones. Traditional disaster zone detection
models often pose significant drawbacks, including slow processing speed, inadequate
dataset size, and high computational costs. To address this limitation, this paper introduces a
novel Optimized Inception based Depthwise Separable cross-transformer (OptimizedInc-
DSCT) model for disaster management by using drone-assisted search and rescue
operations. A data aggregation mechanism is applied for collecting all information from
different drones covered in various disaster areas, and the base station is used as a data
aggregator that transfers information from drones to detection models. The collected data
are preprocessed to enhance image quality and generalization. The Depthwise Separable
Cross Transformer (DSCT) module is designed as a residual connection to improve high-level
feature extraction in the disaster detection framework. Additionally, the incorporation of
Depthwise Separable Convolutions (DSC) and cross-attention operations significantly
reduces computational load while maintaining high accuracy. In order to extract features
into disaster categories and avoid overfitting, a fully connected layer with are used for final
classification. A Majority Voting of Bounding Boxes (MVB) algorithm is incorporated to
accurately identify disaster regions by evaluating detected bounding boxes based on class
labels and confidence scores. The evaluation of the proposed model attained a higher
accuracy of 98.64 %, a precision of 98.57 %, and an F1-score of 98.47 % across various
drone-assisted datasets. Overall, the proposed framework demonstrates remarkable
potential for advancing disaster detection accuracy and ensuring timely response in critical
applications.
Keywords: Disaster zone detection; Inception and reception block; Cross transformer;
Depthwise separable convolutions; Majority voting of bounding boxes; Disaster
management application

Pengcheng Hu, Kai Yang, Heng Wang, Jiadui Chen, Haisong Huang, Jingwei Yang,
Laser welding penetration states recognition based on Gramian Angular Difference Field
images of spectral signals and ConvNeXt model,
Measurement,
Volume 260,
2026,
119786,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Methods using spectral signals for penetration recognition offer significant
research potential. Traditional approaches rely on manual feature extraction, which requires
expertise and is inherently subjective. This research proposes a novel method for visualizing
spectral information to recognize penetration states. In this study, spectral signals were
collected during the laser welding process, and the correlation between the spectral data
and the penetration states was established based on elemental content, boiling point and
atomic transition coefficient. A method based on the Gramian Angular Field was proposed,
converting one-dimensional data into two-dimensional images while largely preserving the
original features. A ConvNeXt-based model was developed and optimized, achieving an
accuracy of 95.34% after 100 rounds of training, which is an increase of 0.46% in accuracy
compared with the method of directly using spectral data for penetration recognition. In
addition, high-precision prediction of weld depth was also achieved based on the spectral
data.
Keywords: Laser welding; Penetration state recognition; Spectral signal; Gramian Angular
Field; ConvNeXt

Junzhi Liu, Lifeng Zhang,


A two-stage electrical capacitance tomography reconstruction method for greedy block
sparsity and multi-domain feature fusion transformer,
Expert Systems with Applications,
Volume 280,
2025,
127590,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Electrical capacitance tomography (ECT) is a visual real-time monitoring
technology. However, the inherent nonlinearity and ill-posedness of the ECT inverse problem
make achieving accurate image reconstruction a significant challenge. To improve the
accuracy of image reconstruction, a two-stage image reconstruction approach based on
greedy block sparse and multi-domain feature fusion transformer is proposed. Firstly, the
sparse prior characteristic of the reconstructed signal is exploited and the Bregman distance
is introduced to construct the objective function, which is solved by employing the Greedy
Kaczmarz method (GKM) to obtain the initial permittivity distribution. Subsequently, a U-
shaped multi-domain feature fusion transformer is used to extract multi-domain features
from the initial permittivity distribution. Finally, simulation and static experiments are
conducted, followed by a comparative analysis with Linear Back Projection (LBP), Landweber
iteration methods, ResNet18 and TransV-Net. The results indicate that, compared to
established reconstruction algorithms, the two-stage algorithm not only exhibits superior
robustness and stability but also significantly improves image reconstruction accuracy,
making it applicable to practical ECT imaging systems.
Keywords: Electrical capacitance tomography; Sparse reconstruction; Multi-domain
attention mechanism; Deep learning; Feature extraction

Kun Qian, Dingwei Zhu, Yutong Wu, Jian Shen, Shoujin Zhang,
TransIST: Transformer based infrared small target tracking using multi-scale feature and
exponential moving average learning,
Infrared Physics & Technology,
Volume 145,
2025,
105674,
ISSN 1350-4495,
[Link]
([Link]
Abstract: Small Unmanned Aerial Vehicle (UAV) tracking against complex sky backgrounds is
of significant importance in both the military and civilian domains, with traditional
correlation filters imposing high requirements on feature models. Consequently, a deep
learning model is designed to achieve effective infrared tracking of a small UAV target.
Specifically, a transformer model is used as the backbone, and a multi-scale attention is
introduced to obtain the perceptual features referring to small targets. Then, an edge
suppression model named side window filter is embedded to suppress the negative effect of
edge on small target tracking. Furthermore, the proposed model is trained using an
exponential moving average learning strategy, which results in more precise network
parameters. Experimental results demonstrate that the proposed Transformer-based
Infrared Small Target Tracking (TransIST) algorithm exhibits superior performance in public
near-infrared videos compared to current related algorithms, with enhanced stability over
correlation filtering-based trackers. The code will be available at [Link]
ayan/TransIST, contributing to the remote sensing community.
Keywords: Infrared tracking; Small objects; Transformer; Multi-scale dilated attention;
Exponential moving average; Edge suppression

Zhuangji Wang, Dennis Timlin, Xiaofei Gong, Yuki Kojima, Shan Hua, David Fleisher,
Wenguang Sun, Sahila Beegum, Vangimalla R. Reddy, Katherine Tully, Robert Horton,
TDR-Transformer: A transformer neural network model to determine soil relative
permittivity variations along a time domain reflectometry sensor waveguide,
Computers and Electronics in Agriculture,
Volume 237, Part C,
2025,
110730,
ISSN 0168-1699,
[Link]
([Link]
Abstract: Interpreting soil relative permittivity (εr) variations along a time domain
reflectometry (TDR) waveguide provides an opportunity to determine soil water content at
multiple depths using a vertically installed TDR sensor. Compared to placing sensors at
different depths, vertical sensor installation reduces measurement efforts and enhances
data-use-efficiency. Revealing εr variations includes two aspects: identifying εr change
positions and determining εr values. Traditional inverse analyses are not widely applied due
to their high computational demands. Machine learning-based methods, e.g., TDR-CNN,
provide a forward computational workflow to track εr change positions and reduce
computational load, but errors in εr values are relatively large. In this study, TDR-
Transformer is developed as a new waveform interpretation model to improve εr estimation
accuracy. Modified from the standard transformer architecture, an encoder with
convolutional neural layers is used to extract waveform geometric features, and a decoder
generates a sequence of εr values to represent εr variations. Attention is a mechanism that
can dynamically extract and process the relevant information within the data, which
processes and integrates the waveform geometric information in the encoder, ensures the
causality (time-order) of the waveform data in the decoder, and transfers information from
the encoder to the decoder. TDR-Transformer was trained and tested using simulated
waveforms where εr changes along the waveguides, but soil electrical conductivity (EC) was
assumed to be small and stable. The RMSE for εr values was within 0.5–1.6 % and the RMSE
of εr change positions was within 5–8 %. A soil infiltration experiment and a precipitation-
evaporation experiment illustrated applications of TDR-Transformer to observed waveforms.
Consequently, TDR-Transformer is a promising artificial intelligence model to interpret TDR
waveforms in soils with nonuniform εr, and fine-tuning TDR-Transformer is recommended
for specific commercial TDR sensor designs.
Keywords: Time Domain Reflectometry (TDR); TDR waveform interpretation; Nonuniform
relative permittivity (εr); Transformer neural network; Machine learning

Caiyi Sun, Dawei Wang, Mingming Xu, Shiqing Wei, Shanwei Liu, Zhongwei Li,
MAF-UFormer: Oil spill detection in SAR images using multi-scale alignment and fusion U-
shaped transformer network,
Regional Studies in Marine Science,
Volume 94,
2026,
104773,
ISSN 2352-4855,
[Link]
([Link]
Abstract: Synthetic aperture radar (SAR) has emerged as a vital technology for detecting oil
spills, even in challenging weather conditions. Deep learning models have demonstrated
significant potential in leveraging SAR images for oil spill detection, owing to their robust
feature extraction capabilities. However, considering the diversity of oil spill target scales
and the extraction of global and local information, there are still particular challenges in
accurately extracting oil spill areas from SAR images. Additionally, polarimetric information
can significantly enhance the separability of oil films and seawater. To overcome these
challenges, a Multi-scale Alignment and Fusion U-Shape Transformer Network (MAF-
UFormer) is proposed, which enhances feature representation by integrating multi-scale
fusion and agent attention mechanisms. To evaluate the effectiveness of MAF-UFormer, we
perform experiments on the publicly available Deep-SAR Oil Spill Detection (SOS) dataset.
The results demonstrate that MAF-UFormer achieves F1-Scores of 87.51 % and 83.03 % on
the Sentinel-1 and PALSAR subsets of SOS, respectively. To further validate the robustness of
MAF-UFormer, we create a new dataset, the Sentinel-1 Oil Spill Detection Dataset Part 1
(S1OSD-1). Experiments on S1OSD-1 demonstrate MAF-UFormer’s superior accuracy in oil
spill detection, outperforming existing methods. Given SAR’s capability to extract
polarimetric features that aid in distinguishing oil spills from seawater, we enhance S1OSD-1
by incorporating polarimetric data to construct Part 2 (S1OSD-2). On S1OSD-2, MAF-
UFormer achieves an additional 1.68 % improvement in F1-Score over S1OSD-1. These
results highlight the potential of MAF-UFormer for oil spill detection, offering vital technical
support for oil spill emergency response and marine environmental protection.
Keywords: Oil spill detection; SAR; Sentinel-1; Polarization feature

Yi Peng, Kui Wang, Chaoyang Wu, Lingyun Kong, Jianming Wu, Zhengyu Ren, Fei Yu, Yiyuan
Duan, Bo Wang, Jiaojiao Wei,
Framework for viscoelastic pavement layer moduli back-calculation from FWD data:
Integrating the spectral element method and Transformer-MLP network,
Construction and Building Materials,
Volume 503,
2025,
144454,
ISSN 0950-0618,
[Link]
([Link]
Abstract: Asphalt pavements experience progressive layer modulus degradation and
frequent distresses as the in-service time increases, leading to higher maintenance costs and
compromising operational safety and service life. Therefore, accurate assessment of layer
moduli and monitoring of structural bearing capacity are essential for achieving long-life
pavements. However, conventional back-calculation methods often assume linear elasticity,
which fails to capture the true viscoelastic pavement structure and often results in limited
accuracy. To address these limitations, this study develops a viscoelastic modulus back-
calculation framework. First, a forward model based on the Spectral Element Method (SEM)
was established to simulate the mechanical response of viscoelastic pavement structures
under Falling Weight Deflectometer (FWD) loading. After model validation, a parameter
sensitivity analysis was conducted to determine the reasonable value ranges for key
parameters, and a comprehensive database of 10,000 simulation-derived pavement cases
was generated. Subsequently, a Transformer-MLP hybrid network model was developed for
back-calculation. The results demonstrate that the model possesses excellent generalization
ability, with the difference in the average R2 between the cross-validation and independent
test sets being less than 1 %. The model achieved high prediction accuracy on the
independent test set, with R2 values of 0.998 for the subgrade modulus, 0.966 for the base
modulus, and 0.887 for the upper surface layer modulus. Furthermore, the model shows
strong robustness, maintaining reliable predictive performance when subjected to simulated
field FWD sensor noise and load fluctuations. By integrating the SEM-based forward model
with the Transformer-MLP network, this framework enables the accurate and efficient
evaluation of pavement structural bearing capacity, providing significant practical value for
construction quality control, structural monitoring, and maintenance decision-making.
Keywords: Spectral element method; Falling weight deflectometer; Layer modulus back-
calculation; Transformer-MLP; Hybrid network

Ming Ma, Rui Zhang, Feng Ma, Qianzong Bao,


Diffraction wave separation and imaging by encoder–decoder network embedded
Transformer,
Journal of Applied Geophysics,
Volume 244,
2026,
106023,
ISSN 0926-9851,
[Link]
([Link]
Abstract: High-precision detection of geological anomalies can be achieved through
diffractions generated when seismic or electromagnetic waves propagate in subsurface
discontinuities, e.g., caves and fractures. However, capturing the diffracted portions of the
full wavefield from acquired seismic or ground penetrating radar (GPR) data is challenging
due to the strong interference and waveform blending. Furthermore, compared with
reflections, diffractions possess low magnitudes and complex shapes. The aforementioned
factors hinder difficulty in deploying robust diffraction extraction and imaging across various
data domains. To enhance the accuracy of diffraction imaging and simplify the processing
steps, we have built a new intricate mapping from full wavefield in dip-angle domain
common image gather (Dip-ADCIG) to unique migrated diffractions with deep learning (DL)
technique. By virtue of the encoder–decoder framework, characteristics of diffracted waves
can be depicted, which are applied to classify disordered waveforms with improved
efficiency. Self-attention computation in the improved backbone Swin Transformer V2
ensures the coincident fidelity between input and prediction result. Apart from the
utilization of optimally configured encoder panel, mode of feature maps concatenating is
modified in decoder module so as to obtain the diffraction imaging of small-scale
heterogeneities. Through a stable training with a flood of data for the diverse designed
geological models, the new workflow can provide a high-resolution depth-domain imaging
of diffractions even with poor quality input gathers. Numerical and field data tests verify the
high performance and validity of our proposed method.
Keywords: Diffraction imaging; Deep learning; Dip-ADCIG; Self-attention computation

Yun Zhou, Haoyu Cui, Dong Liu, Wei Wang,


MSTCRB: Predicting circRNA-RBP interaction by extracting multi-scale features based on
transformer and attention mechanism,
International Journal of Biological Macromolecules,
Volume 278, Part 2,
2024,
134805,
ISSN 0141-8130,
[Link]
([Link]
Abstract: CircRNAs play vital roles in biological system mainly through binding RNA-binding
protein (RBP), which is essential for regulating physiological processes in vivo and for
identifying causal disease variants. Therefore, predicting interactions between circRNA and
RBP is a critical step for the discovery of new therapeutic agents. Application of various
deep-learning models in bioinformatics has significantly improved prediction and
classification performance. However, most of existing prediction models are only applicable
to specific type of RNA or RNA with simple characteristics. In this study, we proposed an
attractive deep learning model, MSTCRB, based on transformer and attention mechanism for
extracting multi-scale features to predict circRNA-RBP interactions. Therein, K-mer and KNF
encoding are employed to capture the global sequence features of circRNA, NCP and DPCP
encoding are utilized to extract local sequence features, and the CDPfold method is applied
to extract structural features. In order to improve prediction performance, optimized
transformer framework and attention mechanism were used to integrate these multi-scale
features. We compared our model's performance with other five state-of-the-art methods
on 37 circRNA datasets and 31 linear RNA datasets. The results show that the average AUC
value of MSTCRB reaches 98.45 %, which is better than other comparative methods. All of
above datasets are deposited in [Link] and
source code are available from [Link]
Keywords: Mutli-scale feature; CircRNA-RBP interaction; Transformer; Attention mechanism

Abel Corrêa Dias, Viviane Pereira Moreira, João Luiz Dihl Comba,
RoBIn: A Transformer-based model for risk of bias inference with machine reading
comprehension,
Journal of Biomedical Informatics,
Volume 166,
2025,
104819,
ISSN 1532-0464,
[Link]
([Link]
Abstract: Objective:
Scientific publications are essential for uncovering insights, testing new drugs, and informing
healthcare policies. Evaluating the quality of these publications often involves assessing their
Risk of Bias (RoB), a task traditionally performed by human reviewers. The goal of this work
is to create a dataset and develop models that allow automated RoB assessment in clinical
trials.
Methods:
We use data from the Cochrane Database of Systematic Reviews (CDSR) as ground truth to
label open-access clinical trial publications from PubMed. This process enabled us to
develop training and test datasets specifically for machine reading comprehension and RoB
inference. Additionally, we created extractive (RoBInExt) and generative (RoBInGen)
Transformer-based approaches to extract relevant evidence and classify the RoB effectively.
Results:
RoBIn was evaluated across various settings and benchmarked against state-of-the-art
methods, including large language models (LLMs). In most cases, the best-performing RoBIn
variant surpasses traditional machine learning and LLM-based approaches, achieving a
AUROC of 0.83.
Conclusion:
This work addresses RoB assessment in clinical trials by introducing RoBIn, two Transformer-
based models for RoB inference and evidence retrieval, which outperform traditional models
and LLMs, demonstrating its potential to improve efficiency and scalability in clinical
research evaluation. We also introduce a public dataset that is automatically annotated and
can be used to enable future research to enhance automated RoB assessment.
Keywords: Evidence-based medicine; Systematic reviews; Risk of bias; Deep learning; Natural
language processing; Machine reading comprehension; Classification models

Rui Wan, Weigang Meng, Tianyun Zhao, Wei Lu,


Overcoming radar sparsity and cross-view misalignment: A sparse-to-sparse fusion paradigm
for robust 3D object detection,
Digital Signal Processing,
Volume 171,
2026,
105831,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Millimeter-wave radar-camera fusion provides a cost-effective alternative to LiDAR
for 3D perception in autonomous driving. However, its potential is constrained by two
limitations: (1) Previous bird’s-eye view fusion methods struggle to accurately align and fuse
cross-modal features, while the fusion strategy incurs substantial computational redundancy
from background processing. (2) The extreme sparsity of radar points (typically < 5 % LiDAR
density) hinders robust geometric measurement. To address these challenges, we propose a
radar-camera fusion 3D detection framework that redefines cross-modal interaction by
transitioning from dense fusion to sparse-to-sparse paradigm. This transformation is
initiated by generating spatially-aware 3D object queries from images and radar sweeps-
leveraging image-derived seed points with radar depth to anchor queries to objects via a
perspective-guided object query generator. Moreover, we introduce adaptive radar pillar
diffusion within foreground regions to mitigate radar sparsity, allowing object queries to
capture geometric information from diffused pillars. Additionally, to maximize image
semantic clues, we further refine object boxes through image keypoint feature aggregation
using a keypoint-aware object refinement module. Our framework not only circumvents
traditional fusion bottlenecks but also achieves real-time inference at 23.4 FPS. Evaluated on
nuScenes dataset, it demonstrates competitive detection performance (65.0 % NDS, 57.8 %
mAP) and tracking precision (58.3 % AMOTA and 0.687m AMOTP). By overcoming sparsity
constraints and improving cross-modal fusion, this work establishes a new paradigm for
robust perception systems.
Keywords: Autonomous driving; 3D object detection; Deep learning; Radar-camera fusion;
Cross-attention

Fan Meng, Tao Song, Xianxuan Lin, Kunlin Yang,


A Data Fusion Approach to Synthesize Microwave Imagery of Tropical Cyclones from Infrared
Data using Vision Transformers,
Information Fusion,
2026,
104167,
ISSN 1566-2535,
[Link]
([Link]
Abstract: Microwave images with high spatiotemporal resolution are essential for observing
and predicting tropical cyclones (TCs), including TC positioning, intensity estimation, and
detection of concentric eyewall. Nevertheless, the temporal resolution of tropical cyclone
microwave (TCMW) images is limited due to satellite quantity and orbit constraints,
presenting a challenging problem for TC disaster forecasting. This research suggests a multi-
sensor data fusion approach, using high-temporal-resolution tropical cyclone infrared (TCIR)
images to generate synthetic TCMW images, offering a solution to this data scarcity problem.
In particular, we introduce a deep learning network based on the Vision Transformer (TCA-
ViT) to translate TCIR images into TCMW images. This can be viewed as a form of synthetic
data generation, enhancing the available information for decision-making. We integrate a
phase-based physical guidance mechanism into the training process. Furthermore, we have
developed a dataset of TC infrared-to-microwave image conversions (TCIR2MW) for training
and testing the model. Experimental results demonstrate the method’s capability in rapidly
and accurately extracting key features of TCs. Leveraging techniques like Mask and Transfer
Learning, it addresses the absence of TCMW images by generating MW images from IR
images, thereby aiding downstream tasks like TC intensity and precipitation forecasting. This
study introduces a novel approach to the field of TC image research, with the potential to
advance deep learning in this direction and provide vital insights for real-time observation
and prediction of global TCs. Our source code and data are publicly available online at
[Link]
Keywords: Tropical Cyclones; Vision Transformer; Data Fusion; Physical Guidance; Microwave
Image; Multimodal Remote Sensing; Synthetic Data Generation

Baiju Yan, Hao Zhang, Yicheng Yao, Changyu Liu, Pu Jian, Peng Wang, Lidong Du, Xianxiang
Chen, Zhen Fang, Yirong Wu,
Heart signatures: Open-set person identification based on cardiac radar signals,
Biomedical Signal Processing and Control,
Volume 72, Part A,
2022,
103306,
ISSN 1746-8094,
[Link]
([Link]
Abstract: Objective
Non-contact continuous biometric identification system based on the heartbeat signals has
attracted more attention due to its privacy-friendly properties. Current methods, however,
mostly focused on the traditional machine learning methods under the close-set condition.
This paper aims to investigate the feasibility of using the cardiac radar heartbeat signals and
deep learning techniques to identify person in the open-set environment.
Methods
A novel dipole deep learning model (DDLM) was proposed for the open-set person
identification problem without heartbeat segmentation and feature engineering. The
normalized heartbeat samples with time duration of 5 s were used as input and encoded
into the feature space, where the encoded features of the same person cluster closely
around the corresponding negative pole and repel far from positive pole, and those of
different persons separate loosely from each other. Finally, threshold on the distance from
the features to the dipoles in the feature space was set for each known identity.
Results
Extensive experiments conducted on a public dataset of clinically recorded vital signs
indicate that:(1) The proposed model shows high stability under close-set condition with an
accuracy higher than 99 % with 30 subjects.(2) The accuracy and the F1-score attain 93.42 %
and 93.57 % under the open-set condition with the maximum openness of 29.3 %,
respectively.
Conclusion
The proposed model shows high effectiveness in person identification using heartbeat
signals. The DDLM outperforms most of the existing methods under the close-set condition.
And the DDLM shows a promising future for person identification in open-set environment.
Keywords: Cardiac radar heartbeat signatures; Dipole deep learning model (DDLM); Machine
learning; Non-contact biometric identification; Open set identification

Yunfeng Fang, Zheng Tong, Tianqing Hei, Siqi Wang, Tao Ma,
Deep learning applications in ground-penetrating radar inversion: A review,
Measurement,
Volume 258, Part D,
2026,
119399,
ISSN 0263-2241,
[Link]
([Link]
Abstract: The complex nonlinear relationship between the subsurface medium and ground-
penetrating radar signals results in the pervasive ill-posedness and non-uniqueness of
conventional inversion methods. Deep learning, with its powerful feature extraction
capabilities and advantages in modeling complex nonlinear relationships, has unique
strengths in handling complex signals and nonlinear problems, making it especially suitable
for GPR inversion tasks. This paper reviews the latest applications of deep learning in GPR
inversion, summarizing the application strategies of deep learning from two perspectives:
data-driven and data-physics hybrid-driven. Commonly used model architectures and their
performance in signal feature extraction, multi-scale information fusion, and data
preprocessing are discussed, along with the application of various loss functions in inversion
tasks. Finally, current challenges, such as limited model generalization, model dependence
on the dataset and computational efficiency constraints, are discussed, and potential future
research directions are proposed to further advance deep learning in GPR inversion.
Keywords: Ground-penetrating radar; Deep learning; Inversion

Tianyang Li, Chao Wang, Sirui Tian, Bo Zhang, Fan Wu, Yixian Tang, Hong Zhang,
TACMT: Text-aware cross-modal transformer for visual grounding on high-resolution SAR
images,
ISPRS Journal of Photogrammetry and Remote Sensing,
Volume 222,
2025,
Pages 152-166,
ISSN 0924-2716,
[Link]
([Link]
Abstract: This paper introduces a novel task of visual grounding for high-resolution synthetic
aperture radar images (SARVG). SARVG aims to identify the referred object in images
through natural language instructions. While object detection on SAR images has been
extensively investigated, identifying objects based on natural language remains under-
explored. Due to the unique satellite view and side-look geometry, substantial expertise is
often required to interpret objects, making it challenging to generalize across different
sensors. Therefore, we propose to construct a dataset and develop multimodal deep
learning models for the SARVG task. Our contributions can be summarized as follows. Using
power transmission tower detection as an example, we have built a new benchmark of
SARVG based on images from different SAR sensors to fully promote SARVG research.
Subsequently, a novel text-aware cross-modal Transformer (TACMT) is proposed which
follows DETR’s architecture. We develop a cross-modal encoder to enhance the visual
features associated with the textual descriptions. Next, a text-aware query selection module
is devised to select relevant context features as the decoder query. To retrieve the object
from various scenes, we further design a cross-scale fusion module to fuse features from
different levels for accurate target localization. Finally, extensive experiments on our dataset
and widely used public datasets have demonstrated the effectiveness of our proposed
model. This work provides valuable insights for SAR image interpretation. The code and
dataset are available at [Link]
Keywords: Synthetic aperture radar (SAR); Power transmission tower; Visual grounding;
Multimodal; Deep learning

Raihan Ahamed Rifat, Fuyad Hasan Bhoyan, Md Humaion Kabir Mehedi, Md Kaviul Hossain,
Md. Jakir Hossen, M.F. Mridha,
ConMatFormer: A multi-attention and transformer integrated ConvNext based deep learning
model for enhanced diabetic foot ulcer classification,
Results in Engineering,
Volume 28,
2025,
108248,
ISSN 2590-1230,
[Link]
([Link]
Abstract: Diabetic foot ulcer (DFU) detection is a clinically significant yet challenging task due
to the scarcity and variability of publicly available datasets. Limited annotated samples
restrict the ability of conventional deep learning models to achieve robust generalization in
real-world clinical scenarios. To solve these problems, we propose ConMatFormer, a new
hybrid deep learning architecture that combines ConvNeXt blocks, multiple attention
mechanisms convolutional block attention module (CBAM) and dual attention network
(DANet), and transformer modules in a way that works together. This design facilitates the
extraction of better local features and understanding of the global context, which allows us
to model small skin patterns across different types of DFU very accurately. To address the
class imbalance, we used data augmentation methods. A ConvNeXt block was used to obtain
detailed local features in the initial stages. Subsequently, we compiled the model by adding a
transformer module to enhance long-range dependency. This enabled us to pinpoint the
DFU classes that were underrepresented or constituted minorities. Tests on the DS1
(DFUC2021) and DS2 (diabetic foot ulcer (DFU)) datasets showed that ConMatFormer
outperformed state-of-the-art (SOTA) convolutional neural network (CNN) and Vision
Transformer (ViT) models in terms of accuracy, reliability, and flexibility. The proposed
method achieved an accuracy of 0.8961 and a precision of 0.9160 in a single experiment,
which is a significant improvement over the current standards for classifying DFUs. In
addition, by 4-fold cross-validation, the proposed model achieved an accuracy of 0.9755
with a standard deviation of only 0.0031. We further applied explainable artificial
intelligence (XAI) methods, such as Grad-CAM, Grad-CAM++, and LIME, to consistently
monitor the transparency and trustworthiness of the decision-making process. These
human-readable tools enhance the comprehension of the explanations and can substantially
increase the practical use of our methodology. Our findings set a new benchmark for DFU
classification and provide a hybrid attention transformer framework for medical image
analysis.
Keywords: Diabetic foot ulcers classification; Multi-attention; Transformer; GradCam;
Explainable AI; LIME

Irfanullah Khan, Antonio Guerrieri, Edoardo Serra, Giandomenico Spezzano,


A hybrid deep learning model for UWB radar-based human activity recognition,
Internet of Things,
Volume 29,
2025,
101458,
ISSN 2542-6605,
[Link]
([Link]
Abstract: In today’s world, energy efficiency in buildings has become a top priority due to the
significant energy waste caused by the operation of inefficient electrical appliances.
Conventional methods of reducing energy waste cause discomfort for occupants inside
buildings. One promising way to optimize energy consumption is to synchronize appliance
operation with building occupants’ dynamic behavior. Internet of Things (IoT) technologies,
which allow for widespread data collection and execution of Machine Learning (ML)
algorithms, enabled the creation of Smart Buildings (SBs). SBs can learn patterns from the
inhabitant’s behavior residing in, and adjust their operations in accordance with these
behaviors. By doing so, these SBs could reduce energy waste, enhancing resource efficiency
and consequently reduce CO2 gas emissions. Furthermore, they could improve the overall
comfort of the living environment and help with sustainability initiatives. In this context, this
paper proposes a novel approach that uses a hybrid deep-learning model to recognize
complex human activities based on data collected from ultra-wideband (UWB) radar
technology. Our approach, called Hybrid Deep Learning Model for Activity Recognition
(HDL4AR), includes long-short-term memory (LSTM) and a one-dimensional convolutional
neural network (1D-CNN). We deploy a real-time case study by collecting data from 22
participants involved in 10 diverse activities at the headquarters of the ICAR-CNR in the IoT
Laboratory, Italy. Moreover, we conducted a comprehensive benchmark of the HDL4AR
approach against various statistical techniques and other deep learning models recently
introduced in the literature. Results show that our proposed approach outperformed
conventional methods and achieved an impressive accuracy of 98.42%.
Keywords: Internet of Things; Smart buildings; Human activity recognition; UWB radar;
Artificial intelligence; Neural networks; LSTM

Yangxiaoyue Liu, Ying Xin, Cong Yin,


A Transformer-based method to simulate multi-scale soil moisture,
Journal of Hydrology,
Volume 655,
2025,
132900,
ISSN 0022-1694,
[Link]
([Link]
Abstract: The Transformer model, as an emerging deep learning method, shows great
potential in spatiotemporal sequence-related simulation tasks. However, there is very
limited understanding of its performance in soil moisture (SM) simulation. Here, we present
a Transformer-based SM simulation network (SMSNet) to improve the quality of the widely
used Soil Moisture Active Passive (SMAP) SM product by reconstructing and downscaling the
pixels, and obtain daily SM products with resolutions of 9 km and 1 km. The model employs
spatiotemporal attention in separate Transformer structures to extract patterns of 10
dynamic (MODIS Bands 1–7, land cover, land surface temperature, and precipitation) and 8
static variables (soil bulk density, clay content, gravel content, sand content, silt content,
digital elevation model, latitude, and longitude) to establish the relationship between
variable pattern and SMAP SM distribution. The SM product reconstructed by SMSNet shows
good agreement with in-situ measurements (SCAN network: R = 0.639, RMSE =
0.086 m3/m3. UCSRN network: R = 0.665, RMSE = 0.097 m3/m3) at the Continental United
States. Moreover, it can effectively mitigate the overestimation degree of SMAP SM and
improve the accuracy in forest. The seamless mapping of SMAP SM is achieved using
reconstructed SM to fill the gap, and the gap-filling SM exhibits reasonable spatiotemporal
pattern. The downscaled 1 km SM performs similar accuracy degrees to those of the
reconstructed ones, proving the applicability of transferring 9 km scale established SMSNet
to 1 km scale SM simulation. The downscaled dataset can provide detailed SM
characteristics, which further enhances the merit of SMAP SM at regional analysis.
Moreover, both reconstructed and downscaled SMs exhibit accuracy superiority compared
to the corresponding results from Cubic Spline Interpolation, Random Forest, and
Convolutional Neural Network. Overall, our study highlights the benefits and potential of
SMSNet in generating SM products with favorable accuracy over diverse and vast regions.
Keywords: Soil moisture; Transformer; Simulation; Reconstruction; Downscaling

Van Ngoc Dang, Ngoc Chau Hoang, Quoc Cuong Nguyen, Minh Thuy Le,
Advancing robust human activity recognition via informative mmWave radar characteristics
and a lightweight spatio-spectro-temporal network,
Measurement,
Volume 256, Part A,
2025,
118056,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Human activity recognition (HAR) is increasingly important in aiding our daily life,
with millimeter-wave (mmWave) radar sensors emerging as a promising noninvasive solution
thanks to their excellent spatial and velocity resolution. Although existing radar-based
systems have shown strong performance, they primarily focus on micro-Doppler signatures
while neglecting angle information, which can hinder practical deployment in real-world
scenarios. Moreover, current state-of-the-art recognition models using mmWave radar often
require substantial computational resources, making integration into resource-constrained
devices challenging. This work proposes an efficient radar-based HAR system that leverages
angle and spectro-temporal information from micro-Doppler signatures. Our system utilizes
a multi-channel micro-Doppler representation corresponding to the number of virtual
antenna receivers as input. Then, a lightweight dilated convolutional network, namely SST-
DCN, extracts spatial-aware multi-scale spectro-temporal information through time-
frequency dilated convolutions. Experimental results on our real-world dataset demonstrate
the superiority of our approach compared to conventional features and other state-of-the-
art radar-based HAR systems.
Keywords: Human activity recognition; Millimeter-wave radar; Deep learning; Lightweight
network; Dilated convolution

Xinran Zhang, Fenghua He, Yu Yao, Zhaochen Lin,


A robust multiple target tracking framework with Transformer-based data association and
dynamic trajectory management in challenging environments,
Aerospace Science and Technology,
Volume 168, Part E,
2026,
111007,
ISSN 1270-9638,
[Link]
([Link]
Abstract: In this paper, a robust and reliable solution is proposed for the multiple target
tracking (MTT) problem with position-only measurements in challenging environments
involving unknown target numbers, high clutter, and frequent missed detections. Limited
positional measurements complicate the disambiguation of closely spaced trajectories, and
are further exacerbated by high false alarm and missed detection rates. To address these
issues, a learning-based data association method embedded in a tracking framework with a
dynamic score-based trajectory management strategy is presented. Specifically, a Cross
Transformer-based Data Association (CTDA) method is proposed, which leverages cross
architecture and Transformer mechanisms to effectively extract discriminative features from
position measurements. Furthermore, a multiple target tracking framework is presented
which integrates three key components: a Best Linear Unbiased Estimation (BLUE) filter that
maintains individual trajectories while providing optimal state estimates, a dynamic
trajectory scoring mechanism that continuously updates trajectory confidence based on
historical association performance provided by the association network, and a trajectory
management strategy that handles the full life-cycle management of trajectories, including
the initialization, maintenance, and termination of the complete trajectory. Comprehensive
simulation results demonstrate that the proposed method outperforms conventional
algorithms in terms of tracking accuracy, reliability, and computational efficiency under
complex conditions, including an over 50 % increase in data association performance and a
substantial 58 % reduction in computational time.
Keywords: Multiple target tracking (MTT); Data association; Cross Transformer; Score-based
framework; Trajectory scoring; Trajectory management

Haoyu Jiang, Xiaoliang Chen, Duoqian Miao, Hongyun Zhang, Xiaolin Qin, Shangyi Du, Peng
Lu,
3WD-DRT: A three-way decision enhanced dynamic routing transformer for cost-sensitive
multimodal sentiment analysis,
Information Sciences,
Volume 725,
2026,
122704,
ISSN 0020-0255,
[Link]
([Link]
Abstract: Accurately interpreting human emotion from language, facial expressions, and
vocal tones remains a fundamental challenge in artificial intelligence. Current Multimodal
Sentiment Analysis (MSA) models often struggle with two key issues. First, their static fusion
strategies fail to handle conflicting modalities, such as sarcasm. Second, their standard loss
functions ignore the asymmetric risks of severe misjudgments. To address these limitations,
we propose the Three-Way Decision Enhanced Dynamic Routing Transformer (3WD-DRT), a
framework operating on a "quality-aware, decision-driven" principle. It dynamically assesses
each modality’s quality using a three-way decision gate, implemented via a dedicated MLP,
to partition information into acceptance, deferment, or rejection pathways. This enables the
model to amplify informative signals, moderately scale uncertain ones (deferment), and
attenuate noisy or misleading ones. We also introduce a novel cost-sensitive loss function
that imposes greater penalties on major semantic errors, such as polarity misclassifications.
This approach better aligns the model’s training objective with human perception. Extensive
experiments on CH-SIMS, CH-SIMSv2, MOSI, and MOSEI datasets show that 3WD-DRT
consistently outperforms state-of-the-art methods, setting new benchmarks with F1-scores
of 87.08 % on MOSI and 88.26 % on MOSEI. This work provides a robust solution for MSA,
fostering more nuanced and reliable emotionally-aware AI systems.
Keywords: Multimodal sentiment analysis (MSA); Three-way decision theory; Dynamic
routing transformer; Emotion-aware fusion

Guolian Hou, Fan Zhang, Congzhi Huang, Ting Huang,


Joint prediction of SOH and RUL for Lithium-ion batteries by an enhanced Transformer
model with physical information constraints,
Energy,
Volume 336,
2025,
138435,
ISSN 0360-5442,
[Link]
([Link]
Abstract: It is essential to accurately predict the remaining useful life (RUL) and state of
health (SOH) for effective health management of lithium-ion batteries (LIBs). Current
methods face significant hurdles in practical application because they are heavily reliant on
data quantity and quality, and their generalizability and interpretability are often limited. To
address these challenges, an enhanced Transformer model with physical information
constraints is proposed for the joint prediction of SOH and RUL. Firstly, to effectively capture
the complex local and global patterns of LIBs degradation while maintaining computational
efficiency, the matrix long short-term memory block is integrated into the encoder in the
enhanced Transformer model. Secondly, to extract health features highly correlated with
battery degradation, an improved Z-Score wavelet filtering algorithm is designed to process
the raw signals. Six key features are selected from the filtered current, voltage, and capacity
profiles as model input. Then, to address the insufficient generalization, poor interpretability
of pure data-driven models, an implicit partial differential equation describing the LIBs
degradation is established. It incorporated into the loss function, effectively constraining
model training. Additionally, a model transfer strategy is devised to tackle the insufficient
data of specific batteries, the physics-informed component, encoding universal degradation
laws, is kept fixed, while the data-driven adaptation component is fine-tuned on limited
target data. Finally, the proposed method is validated on four different types of LIBs data.
When only one sample is used for training, the proposed model achieves average fitting
coefficient exceeding 0.99 for SOH prediction and maintains RUL prediction errors within 5
cycles. The fine-tuning strategy achieved a fitting accuracy of 0.99 using only 30 % of the
target domain data. Therefore, the designed model provides a new solution for the high-
precision battery health management that possesses both generalization capability and
physical interpretability.
Keywords: Lithium-ion batteries degradation; Transformer; Extended long short-term
memory; Physics-informed neural networks; Transfer learning

Shweta B. Thomas, Sangeetha Subbaraj, Deepika Rani Sona, Benedict Thomas,


Non-destructive GPR signal processing technique for thickness estimation of pavement, coal
and ice layers: A review,
Journal of Applied Geophysics,
Volume 233,
2025,
105601,
ISSN 0926-9851,
[Link]
([Link]
Abstract: In recent years, there has been a significant surge in the utilization of Ground
Penetrating Radar (GPR) for measuring the thickness of subsurface layers, and researchers in
this field have paid close attention to it. GPR enables users to achieve greater precision in
evaluating the quality and condition of underground materials. The traditional methods used
to measure the thickness of underground layers are time-consuming, hard to conduct and
not economical. GPR is one of the most recommended non-destructive geophysical methods
for routine subsurface inspections. This article is intended to highlight the application of GPR
for thickness estimation of distinct materials such as, pavement, ice and coal layers and
novel non-destructive testing (NDT) techniques adopted recently for thickness estimation.
This article presents an overview of Ground Penetrating Radar (GPR) methodologies for layer
thickness estimation, encompassing their advantages, disadvantages, and recent research
findings. By synthesizing existing literature, the potential applications of GPR while
addressing its inherent limitations are illustrated here. Furthermore, practical
recommendations are provided to enhance the effectiveness of GPR-based layer thickness
estimation techniques.
Keywords: Ground penetrating radar; Non-destructive techniques; Thickness estimation;
Underground layers

Wei Wang, Ruobing Song, Yunxiao Wu, Li Zheng, Wenyu Zhang, Zhaoxi Chen, Gang Li, Zhifei
Xu,
Deep learning-based automated diagnosis of obstructive sleep apnea and sleep stage
classification in children using millimeter-wave radar and pulse oximeter,
Sleep Health,
Volume 11, Issue 6,
2025,
Pages 859-867,
ISSN 2352-7218,
[Link]
([Link]
Abstract: Study objectives
Due to the high cost, complexity, and workload of polysomnography, a radar-based sleep
monitoring device, QSA600, has been developed as a more simplified alternative for
children. This study evaluates its agreement with polysomnography for obstructive sleep
apnea diagnosis and sleep staging.
Methods
This diagnostic accuracy study included 281 children (1-18 years) who underwent
simultaneous polysomnography and QSA600 monitoring at Beijing Children's Hospital from
September-November 2023. QSA600 recordings were automatically analyzed using a deep
learning model, while polysomnography data were manually scored.
Results
The obstructive apnea-hypopnea index (OAHI) obtained from QSA600 and polysomnography
demonstrates a high level of agreement with an intraclass correlation coefficient of 0.945
(95% CI: 0.93-0.96). Bland-Altman analysis indicated that the mean difference of obstructive
apnea-hypopnea index between QSA600 and polysomnography was −0.10 events/h (95% CI:
−11.15 to 10.96). The deep learning model evaluated through cross-validation showed good
sensitivity (81.8%, 84.3%, and 89.7%) and specificity (90.5%, 95.3%, and 97.1%) values for
diagnosing children with OAHI >1, OAHI >5, and OAHI >10. The area under the receiver
operating characteristic curve was 0.923, 0.955, and 0.988, respectively. For sleep stage
classification, the model achieved Kappa coefficients of 0.854, 0.781, and 0.734, with
corresponding overall accuracies of 95.0%, 84.8%, and 79.7% for Wake-Sleep classification,
Wake-REM-Light-Deep classification, and Wake-REM-N1-N2-N3 classification, respectively.
Conclusions
QSA600 has demonstrated high agreement with polysomnography in diagnosing obstructive
sleep apnea and performing sleep staging in children. The device is portable, low-burden,
and suitable for follow-up and long-term pediatric sleep assessment.
Keywords: Obstructive sleep apnea; Children; Deep learning; Millimeter-wave radar;
Portable sleep monitoring device; Polysomnography

Chen Jiang, Shuxia Lu, Xianghu Zhou, Tingting Ma, Junhai Zhai,
KNN improved Transformer for 3D object detection,
Signal Processing: Image Communication,
Volume 142,
2026,
117488,
ISSN 0923-5965,
[Link]
([Link]
Abstract: In recent years, 3D object detection in autonomous driving perception has gained
significant attention in the industry. Due to its characteristics, LiDAR has become the most
commonly used and essential sensor. However, voxel-based networks often lose context
information during the voxelization process, which negatively impacts the detection of small
objects. In this paper, we address the challenge of low accuracy in LiDAR-based detection,
especially for small object categories, by proposing an improved Transformer structure.
Transformers are a type of deep learning model known for their ability to capture long-range
dependencies and contextual relationships in data. In our approach, we incorporate a k-
Nearest Neighbors (KNN) algorithm, which is a method for identifying the closest points in
space, to enhance the spatial relationships between point clouds. This combination allows
the model to better capture context information, strengthen feature extraction, and
significantly reduce both missed and false detections. Our method is designed to be plug-
and-play, allowing it to be directly applied to existing point cloud detectors. We evaluate our
approach on the public KITTI and Astyx datasets. Experimental results show significant
improvements, especially in detecting small object categories, even in challenging
conditions.
Keywords: Autonomous vehicle; Transformer; Lidar; 3D object detection; Point cloud

Bingcai Wei, Di Wang, Zhuang Wang, Liye Zhang,


Single Image Desnow Based on Vision Transformer and Conditional Generative Adversarial
Network for Internet of Vehicles,
CMES - Computer Modeling in Engineering and Sciences,
Volume 137, Issue 2,
2023,
Pages 1975-1988,
ISSN 1526-1492,
[Link]
([Link]
Abstract: With the increasing popularity of artificial intelligence applications, machine
learning is also playing an increasingly important role in the Internet of Things (IoT) and the
Internet of Vehicles (IoV). As an essential part of the IoV, smart transportation relies heavily
on information obtained from images. However, inclement weather, such as snowy weather,
negatively impacts the process and can hinder the regular operation of imaging equipment
and the acquisition of conventional image information. Not only that, but the snow also
makes intelligent transportation systems make the wrong judgment of road conditions and
the entire system of the Internet of Vehicles adverse. This paper describes the single image
snow removal task and the use of a vision transformer to generate adversarial networks. The
residual structure is used in the algorithm, and the Transformer structure is used in the
network structure of the generator in the generative adversarial networks, which improves
the accuracy of the snow removal task. Moreover, the vision transformer has good scalability
and versatility for larger models and has a more vital fitting ability than the previously
popular convolutional neural networks. The Snow100K dataset is used for training, testing
and comparison, and the peak signal-to-noise ratio and structural similarity are used as
evaluation indicators. The experimental results show that the improved snow removal
algorithm performs well and can obtain high-quality snow removal images.
Keywords: Artificial intelligence; Internet of Things; vision transformer; deep learning; image
desnow

Mahdi Amini Sedeh, Saeed Sharifian,


A serverless satellite edge computing infrastructure to detect lake shoreline change based
on contrastive attention siamese transformer,
Engineering Applications of Artificial Intelligence,
Volume 154,
2025,
111015,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Supporting food security, industrial activities, environmental well-being, and
ecological balance, water resources are absolutely vital for life on Earth. Still, issues such as
water shortage, pollution, and melting glaciers point to the need for sustainable water
management techniques. Monitoring changes in the water surface depends on satellite
images and remote sensing technologies, but conventional satellite systems suffer with data
latency and bandwidth inefficiency. Through data processing closer to the source, latency
reduction, and performance enhancement, edge computing has transformed these
technologies. Still under study, though, more precise image analysis techniques—especially
those based on deep learning models—remain a subject of inquiry. This work contributes to
the changing world by proposing a novel serverless satellite edge computing infrastructure
with an emphasis on lake shoreline dilatation, a crucial factor in water management. We
developed transformer-based Siamese deep learning models as the foundation of the image
analysis method. We specifically adjusted the transformer-based Siamese network by
including contrastive attention in the fusion module of the model. The suggested model
beats state-of-the-art models in overall accuracy by at least 1.75 % by assessing it using
actual satellite photos of various lakes all around the planet. Better accuracy in identifying
lake shoreline dilatation aids in our understanding of water movement and long-term
responsible water resource management in addition to the suggested distributed serverless
satellite edge computing infrastructure. This study marks a significant first step towards
creatively tackling the problems with water availability and quality, therefore promoting a
whole strategy to protect this priceless resource for next generations as well as present
ones.
Keywords: Lake shoreline change detection; Serverless computing; Satellite edge; Siamese
transformer; Contrastive attention

AmirHossein Adibfar, Hassan Davani,


PrecipNet: A transformer-based downscaling framework for improved precipitation
prediction in San Diego County,
Journal of Hydrology: Regional Studies,
Volume 62,
2025,
102738,
ISSN 2214-5818,
[Link]
([Link]
Abstract: Study region
San Diego County (California, USA), with its complex topography and coastal climate
variability, requires high-resolution precipitation data to support hydrological modeling and
climate adaptation planning. However, the coarse spatial resolution of Global Climate
Models (GCMs) limits their applicability in such a diverse and hydrologically sensitive region.
Study focus
This study introduces a two-stage hybrid statistical downscaling framework that combines
Transformer-based deep learning with traditional machine learning for localized
precipitation prediction. The goal is to downscale coarse-resolution CMIP5 precipitation data
(2° × 2.5°, 3-h intervals) to a finer 10 km × 10 km grid appropriate for regional hydrological
applications. The first stage employs HydroFusionNet, a Transformer-based classifier, to
detect rainfall occurrence using spatial atmospheric predictors, thereby filtering out non-rain
periods and improving computational efficiency. The second stage applies two regression
models: a Random Forest with linear bias adjustment and PrecipNet, a Transformer-based
model.
New hydrological insights for the region
PrecipNet achieved a Mean Absolute Error (MAE) of 1.24 mm, Root Mean Square Error
(RMSE) of 1.62 mm, and R² of 0.94, outperforming the Random Forest baseline in accuracy
and spatial generalization. HydroFusionNet demonstrated 92.75 % classification accuracy,
enhancing rainfall detection. The framework reduces false positives, captures complex
rainfall dynamics, and provides context-aware uncertainty estimation—offering a scalable,
hydrologically meaningful tool for regional climate impact assessments and water resource
decision-making in topographically complex areas like San Diego County.
Keywords: GCM downscaling; Neural networks; Transformer; Random forest; PrecipNet;
HydroFusionNet
Pianzhang Duan, Li Wang, Cheng Fang, Ziying Song, Ming Gao, Mo Zhou, Ying Li, Yibo Zhang,
Wei Fan, Bin Xu,
Global relationship awareness 3-dimensional object detection using 4-dimensional radar,
Engineering Applications of Artificial Intelligence,
Volume 164, Part B,
2026,
113318,
ISSN 0952-1976,
[Link]
([Link]
Abstract: 4D (4-dimensional) radar sensing technology is essential for high-precision
autonomous driving perception systems, as its superior detection capabilities at increased
distances, compared to traditional LiDAR (Light Detection and Ranging). However, due to the
sparsity of point clouds and the low resolution of millimeter-wave radar, voxel-based
methods may fail to detect distant or closely adjacent objects, leading to inadequate
detection accuracy. To mitigate the accuracy issues arising from the sparse nature of point
clouds in such scenarios, we propose a novel object detection network: GRA-Net (Global
Relation-Aware object detection Network). By leveraging a self-attention mechanism, GRA-
Net effectively learns critical features from each radar pillar, enhancing the network’s
capacity to capture relevant information about nearby objects. Furthermore, we introduce a
global perception module that integrates key features within the pillars and global features,
mitigating the impact of point cloud sparsity, particularly in distant regions. We conducted a
series of experiments to evaluate the performance of GRA-Net. On the Astyx HiRes 2019
dataset, our method achieved 33.63 mAP (mean Average Precision) and 43.93 mAP at the
moderate level; On the View-of-Delft dataset, our method achieved 47.74 mAP in the entire
annotated area and 69.25 mAP in the driving corridor area.
Keywords: 4-dimensional radar; 3-dimensional object detection; Self-attention mechanism;
Autonomous driving

Tamer Saleh, Xingxing Weng, Shimaa Holail, Chen Hao, Gui-Song Xia,
DAM-Net: Flood detection from SAR imagery using differential attention metric-based vision
transformers,
ISPRS Journal of Photogrammetry and Remote Sensing,
Volume 212,
2024,
Pages 440-453,
ISSN 0924-2716,
[Link]
([Link]
Abstract: Flood detection from synthetic aperture radar (SAR) imagery plays an important
role in crisis and disaster management. Based on pre- and post-flood SAR images, flooded
areas can be extracted by detecting changes of water bodies. Existing state-of-the-art
change detection methods primarily target optical image pairs. The nature of SAR images,
such as scarce visual information, similar backscatter signals, and ubiquitous speckle noise,
pose great challenges to identifying water bodies and mining change features, thus resulting
in unsatisfactory performance. Besides, the lack of large-scale annotated datasets hinders
the development of accurate flood detection methods. In this paper, we focus on the
difference between SAR image pairs and present a differential attention metric-based
network (DAM-Net), to achieve flood detection. By introducing feature interaction during
temporal-wise feature representation, we guide the model to focus on changes of interest
rather than fully understanding the scene of the image. On the other hand, we devise a class
token to capture high-level semantic information about water body changes, increasing the
ability to distinguish water body changes and pseudo changes caused by similar signals or
speckle noise. To better train and evaluate DAM-Net, we create a large-scale flood detection
dataset using Sentinel-1 SAR imagery, namely S1GFloods. This dataset consists of 5,360
image pairs, covering 46 flood events during 2015–2022, and spanning 6 continents of the
world. The experimental results on this dataset demonstrate that our method outperforms
several advanced change detection methods. DAM-Net achieves 97.8% overall accuracy,
96.5% F1, and 93.2% IoU on the test set. Our dataset and code are available at
[Link]
Keywords: Flood detection; SAR imagery; S1GFloods dataset; Vision transformers

Jia Liu, Hang Gu, Fangmei Liu, Hao Chen, Zuhe Li, Gang Xu, Qidong Liu, Wei Wang,
CE-CDNet: A Transformer-Based Channel Optimization Approach for Change Detection in
Remote Sensing,
Computers, Materials and Continua,
Volume 83, Issue 1,
2025,
Pages 803-822,
ISSN 1546-2218,
[Link]
([Link]
Abstract: In recent years, convolutional neural networks (CNN) and Transformer
architectures have made significant progress in the field of remote sensing (RS) change
detection (CD). Most of the existing methods directly stack multiple layers of Transformer
blocks, which achieves considerable improvement in capturing variations, but at a rather
high computational cost. We propose a channel-Efficient Change Detection Network (CE-
CDNet) to address the problems of high computational cost and imbalanced detection
accuracy in remote sensing building change detection. The adaptive multi-scale feature
fusion module (CAMSF) and lightweight Transformer decoder (LTD) are introduced to
improve the change detection effect. The CAMSF module can adaptively fuse multi-scale
features to improve the model’s ability to detect building changes in complex scenes. In
addition, the LTD module reduces computational costs and maintains high detection
accuracy through an optimized self-attention mechanism and dimensionality reduction
operation. Experimental test results on three commonly used remote sensing building
change detection data sets show that CE-CDNet can reduce a certain amount of
computational overhead while maintaining detection accuracy comparable to existing
mainstream models, showing good performance advantages.
Keywords: Remote sensing; change detection; attention mechanism; channel optimization;
multi-scale feature fusion

Derek Ka-Hei Lai, Li-Wen Zha, Tommy Yau-Nam Leung, Andy Yiu-Chau Tam, Bryan Pak-Hei So,
Hyo-Jung Lim, Daphne Sze Ki Cheung, Duo Wai-Chi Wong, James Chung-Wai Cheung,
Dual ultra-wideband (UWB) radar-based sleep posture recognition system: Towards
ubiquitous sleep monitoring,
Engineered Regeneration,
Volume 4, Issue 1,
2023,
Pages 36-43,
ISSN 2666-1381,
[Link]
([Link]
Abstract: Sleep posture monitoring is an essential assessment for obstructive sleep apnea
(OSA) patients. The objective of this study is to develop a machine learning-based sleep
posture recognition system using a dual ultra-wideband radar system. We collected
radiofrequency data from two radars positioned over and at the side of the bed for 16
patients performing four sleep postures (supine, left and right lateral, and prone). We
proposed and evaluated deep learning approaches that streamlined feature extraction and
classification, and the traditional machine learning approaches that involved different
combinations of feature extractors and classifiers. Our results showed that the dual radar
system performed better than either single radar. Predetermined statistical features with
random forest classifier yielded the best accuracy (0.887), which could be further improved
via an ablation study (0.938). Deep learning approach using transformer yielded accuracy of
0.713.
Keywords: Obstructive sleep apnea; Deep learning; Sleep monitoring; Feature extraction;
Ablation study

Shuochen Han, Zhonghao Wang, Guochang Zhang, Chengyang Li, Haitao Zhu, Yanyan Wang,
ROV Trajectory Prediction Algorithm Based on Transformer-LSTM,
IFAC-PapersOnLine,
Volume 59, Issue 35,
2025,
Pages 454-459,
ISSN 2405-8963,
[Link]
([Link]
Abstract: This study proposes a Transformer-LSTM hybrid algorithm to address the
degradation in trajectory prediction accuracy for Remotely Operated Vehicles (ROVs) caused
by umbilical cable communication delays during collaborative operations with Unmanned
Surface Vehicles (USVs). The methodology integrates historical USV observation data with
ROV motion characteristics to construct multimodal feature vectors, employing the
Transformer’s multi-head self-attention mechanisms for global trajectory feature extraction
and LSTM networks for local temporal dependency optimization. A compensation
mechanism utilizing historical predictions ensures stability during USV observation failures.
Experimental results demonstrate that the proposed approach significantly outperforms
baseline methods in both normal and failure scenarios.
Keywords: ROV; trajectory prediction; Transformer; LSTM; cooperative operation

Wang Zhang, Tingting Li, Yuntian Zhang, Gensheng Pei, Xiruo Jiang, Yazhou Yao,
LTFormer: A light-weight transformer-based self-supervised matching network for
heterogeneous remote sensing images,
Information Fusion,
Volume 109,
2024,
102425,
ISSN 1566-2535,
[Link]
([Link]
Abstract: Matching visible and near-infrared (NIR) images is a major challenge in remote
sensing image fusion due to nonlinear radiometric differences. Deep learning has shown
promise in computer vision, but most methods rely on supervised learning with limited
annotated data in remote sensing. To address this, we propose a novel keypoint descriptor
approach that obtains robust feature descriptors via a self-supervised matching network.
Our light-weight transformer network, LTFormer, generates deep-level feature descriptors.
Furthermore, we implement an innovative triplet loss function, LT Loss, to enhance the
matching performance further. Our approach outperforms conventional hand-crafted local
feature descriptors and proves equally competitive compared to state-of-the-art deep
learning-based methods, even amidst the shortage of annotated data. Code and pre-trained
model are available at [Link]
Keywords: Image matching; Transformer; Light-weight; Heterogeneous remote sensing
images; Self-supervised learning

Wenhao Dong, Yueyang Li, Weiming Zeng, Lei Chen, Hongjie Yan, Wai Ting Siok, Nizhuan
Wang,
STARFormer: A novel spatio-temporal aggregation reorganization transformer of FMRI for
brain disorder diagnosis,
Neural Networks,
Volume 192,
2025,
107927,
ISSN 0893-6080,
[Link]
([Link]
Abstract: Many existing methods that use functional magnetic resonance imaging (fMRI) to
classify brain disorders, such as autism spectrum disorder (ASD) and attention deficit
hyperactivity disorder (ADHD), often overlook the integration of spatial and temporal
dependencies of the blood oxygen level-dependent (BOLD) signals, which may lead to
inaccurate or imprecise classification results. To solve this problem, we propose a spatio-
temporal aggregation reorganization transformer (STARFormer) that effectively captures
both spatial and temporal features of BOLD signals by incorporating three key modules. The
region of interest (ROI) spatial structure analysis module uses eigenvector centrality (EC) to
reorganize brain regions based on effective connectivity, highlighting critical spatial
relationships relevant to the brain disorder. The temporal feature reorganization module
systematically segments the time series into equal-dimensional window tokens and captures
multiscale features through variable window and cross-window attention. The spatio-
temporal feature fusion module employs a parallel transformer architecture with dedicated
temporal and spatial branches to extract integrated features. The proposed STARFormer has
been rigorously evaluated on two publicly available datasets for the classification of ASD and
ADHD. The experimental results confirm that STARFormer achieves state-of-the-art
performance across multiple evaluation metrics, providing a more accurate and reliable tool
for the diagnosis of brain disorders and biomedical research. The official implementation
codes are available at: [Link]
Keywords: Brain disorder diagnosis; fMRI; Eigenvector centrality; Spatio-temporal
information integration; Transformer

Ibrahim Fayad, Philippe Ciais, Martin Schwartz, Jean-Pierre Wigneron, Nicolas Baghdadi,
Aurélien de Truchis, Alexandre d'Aspremont, Frederic Frappart, Sassan Saatchi, Ewan Sean,
Agnes Pellissier-Tanon, Hassan Bazzi,
Hy-TeC: a hybrid vision transformer model for high-resolution and large-scale mapping of
canopy height,
Remote Sensing of Environment,
Volume 302,
2024,
113945,
ISSN 0034-4257,
[Link]
([Link]
Abstract: Accurate and timely monitoring of forest canopy height is critical for assessing
forest dynamics, biodiversity, carbon sequestration as well as forest degradation and
deforestation. Recent advances in deep learning techniques, coupled with the vast amount
of spaceborne remote sensing data offer an unprecedented opportunity to map canopy
height at high spatial and temporal resolutions. Current techniques for wall-to-wall canopy
height mapping correlate remotely sensed information from optical and radar sensors in the
2D space to the vertical structure of trees using lidar's 3D measurement abilities serving as
height proxies. While studies making use of deep learning algorithms have shown promising
performances for the accurate mapping of canopy height, they have limitations due to the
type of architectures and loss functions employed. Moreover, mapping canopy height over
tropical forests remains poorly studied, and the accurate height estimation of tall canopies is
a challenge due to signal saturation from optical and radar sensors, persistent cloud cover,
and sometimes limited penetration capabilities of lidar instruments. In this study, we map
heights at 10 m resolution across the diverse landscape of Ghana with a new vision
transformer (ViT) model, dubbed Hy-TeC, optimized concurrently with a classification
(discrete) and a regression (continuous) loss function. This model achieves significantly
higher accuracy than previously employed convolutional-based approaches (ConvNets)
optimized with only a continuous loss function. Hy-TeC results show that our proposed
discrete/continuous loss formulation significantly increases the sensitivity for very tall trees
(i.e., > 35 m). Overall, Hy-TeC has significantly reduced bias (0.8 m) and higher accuracy
(RMSE = 6.6 m) over tropical forests for which other approaches show poorer performance
and oftentimes a saturation effect. The height maps generated by Hy-TeC also have better
ground sampling distance and better sensitivity to sparse vegetation. Over these areas, Hy-
TeC showed an RMSE of 3.1 m in comparison to a reference dataset while the baseline
ConvNet model had an RMSE of 4.3 m. Hy-TeC, which was used to generate a height map of
Ghana using free and open access remotely sensed data with Sentinel-2 and Sentinel-1
images as predictors and GEDI height measurements as calibration data, has the potential to
be used globally.
Keywords: Canopy height; GEDI; Sentinel-1; Sentinel-2; Vision transformers, deep learning,
knowledge distillation

Tao Tan, Xiuping Li, Yubing Li, Shuai Wu, Bohai Fang, Changkai Zhang, Yujian Qin,
A transformer-based current reuse CMOS Armstrong VCO using gm-boosted technique,
Microelectronics Journal,
Volume 163,
2025,
106753,
ISSN 1879-2391,
[Link]
([Link]
Abstract: This paper presents a low power LC Armstrong VCO based on current reuse
structure and gm-boosted technique with a 5-port 3-coil transformer. The applied current
reuse structure converts the traditional parallel cross-coupled transistor pair into a series
PMOS-NMOS pair, reducing the overall power consumption. Moreover, the phase-shifted
between ids and Vds of the transistors reduces power dissipated in the devices by leveraging
the large language model optimization results. The proposed gm-boosted technique is
employed by two transistors in series and capacitors in parallel. With the proposed topology,
the negative conductance, oscillation amplitude, and the Perturbation Projection Vector
(PPV) are improved, thus better phase noise performance. The 5-port 3-coil transformer
integrates all inductors to minimizing the area. Fabricated in GlobalFoundries 0.11-μm CMOS
technology, the VCO achieves a power consumption of 2 mW at a 1.2 V supply voltage, with
a phase noise of −114.5 dBc/Hz @1MHz offset at 11.2 GHz.
Keywords: Armstrong VCO; gm-boosted; Current reuse; 5-port transformer; Voltage-
controlled oscillator (VCO)

Senguo Cao, Congde Lu, Xiao Wang, Peng Zhang, Guanglai Jin, Wenlong Cai,
ME-YOLO: A novel real-time detection network for pavement interlayer distress using
ground-penetrating radar,
Journal of Applied Geophysics,
Volume 245,
2026,
106057,
ISSN 0926-9851,
[Link]
([Link]
Abstract: Interlayer distress detection in asphalt pavement is critical for highway
maintenance, as timely identification of pavement distress can ensure operational safety,
reliability, and extended service life. However, the problems of feature information loss and
the substantial confusable backgrounds significantly hinder detection accuracy. To address
these limitations, we propose an enhanced network specifically designed for automated
interlayer distress detection named ME-YOLO. Firstly, we design a Multiscale Adaptive
Feature Fusion (MAFF) module, which aggregates more scale information by Adaptive
Spatial Feature Fusion (ASFF). This design links all feature scales to make discriminative
features in each scale propagate directly to subsequent modules, enriching semantic
representations and mitigating the risk of feature loss, while leveraging shallow-layer
features to strengthen spatial localization. Furthermore, the Efficient Partial Self-Attention
(EPSA) module is introduced to suppress background interference in complex environments.
Unlike conventional transformers, EPSA adopts partial self-attention operations with multi-
path fusion, which can enable the network to acquire global representation capability with
low computational overhead. Extensive experiments indicate that the ME-YOLO network
outperforms the given state-of-the-art models, including Faster-RCNN, RT-DETR, YOLOv8s,
and YOLOv11s, on the interlayer distress dataset. Compared to YOLOv5s, ME-YOLO achieves
improvements of 2.2% in mAP0.5 and 3.5% in mAP0.5:0.95, while maintaining an inference
speed of 6.7 ms per image. The source code will be available at
[Link]
Keywords: Asphalt pavement; Interlayer distress; Ground penetrating radar; Transformer;
Multiscale feature fusion

Zhen Wen, Zongxuan Li, Shuping Tao, Yu Zhao, Yifan Li, Xinlong Wang,
Hybrid Mamba-Transformer network for phase unwrapping in optical interferometry,
Optics Communications,
Volume 601,
2026,
132726,
ISSN 0030-4018,
[Link]
([Link]
Abstract: Phase unwrapping (PU) is crucial in optical interferometry, as accurate phase
information directly affects quantitative analysis and precise reconstruction quality.
Conventional PU methods suffer from performance degradation under severe noise or
undersampling. With the rise of deep learning, recent advancements in PU have been
improved upon CNN and Transformer-based frameworks. Nonetheless, CNNs lack sufficient
capability to model spatial dependencies of wrapped phase, while Transformer architectures
suffer from quadratic computational complexity. State space models like Mamba have
recently become attractive solutions due to efficient linear complexity in capturing long-
range dependencies. Inspired by this, we propose HMTPU, an innovative hybrid architecture
for PU that integrates the strengths of Mamba and Transformer to achieve high performance
with computational efficiency. Specifically, we integrate Transformer layers after Mamba
layers to strengthens the model’s capability in modeling long-range spatial relationships and
improves its effectiveness in processing local wrapped phase information. Within the
Mamba, a geometric transformable selective scan module is designed to enhance the
acquisition of global spatial context through efficient state space modeling. And a
deformable local enhanced window attention is introduced to refine local representations
and handle structural variations in the Transformer. Additionally, we employ an enhanced
feedforward network that leverages context broadcasting and the gating mechanism to
facilitate efficient cross-channel interaction. Extensive experiments demonstrate that
HMTPU outperforms current advanced PU techniques. Testing on real-world datasets of
dynamic candle flames and holographic tomography shows the generalization capability of
our PU method.
Keywords: Phase unwrapping; Mamba; Transformer; Optical interferometry
M. Mortazavi, Z. Moravej, G.B. Gharehpetian,
Detection and localization of LV winding radial deformation in transformers using
electromagnetic waves - a feasibility study,
International Journal of Electrical Power & Energy Systems,
Volume 155, Part B,
2024,
109602,
ISSN 0142-0615,
[Link]
([Link]
Abstract: Recently, online methods based on electromagnetic waves have been proposed to
detect the mechanical defects of high voltage (HV) windings in power transformers. In this
article, for the first time, the possibility of online detecting and locating the radial
deformation (RD) of low voltage (LV) windings using electromagnetic waves is presented. In
the proposed method, a high frequency and wideband signal is sent to a simplified model of
transformer winding via small antennas. The antennas are connected to the inner side of the
oil tank cover through Radio Frequency (RF) cables, so that online monitoring can be realized
with minimal changes in the transformer structure. The data of the reflected signals in
different states is recorded in a database and sorted as primary data bank. By comparing the
results of the sound state with measured signals, the possible faults can be detected. In this
paper, the detection and location of radial deformation is estimated by using regression tree
based on three features. Several cases are simulated by Computer Simulation Technology
(CST) software, and their verification are conducted by a setup in laboratory. Based on
comparison results, it can be claimed that the proposed method has LV windings radial
deformation online detecting and locating ability with an acceptable accuracy.
Keywords: Power transformers; LV winding; Electromagnetic Waves; Regression; Radial
deformation; Wideband antenna

Yukai Kong, Xianxiang Yu, Jiachen Li, Kui Xiong, Guolong Cui,
Non-uniform pulse intervals based intra-pulse forwarding jamming detection and
recognition in clutter circumstance,
Signal Processing,
Volume 238,
2026,
110193,
ISSN 0165-1684,
[Link]
([Link]
Abstract: The detection and identification of jamming is the prerequisite and key to the
implementation of anti-jamming measures in radar. In the target detection scenario of
airborne radar, strong clutter causes great difficulty in the detection and identification of
intra-pulse forwarding jamming. This paper proposes a jamming detection and recognition
method based on non-uniform pulse interval coupled with encoder–decoder network.
Specifically, the emission mechanism with non-uniform pulse interval is utilized to disrupt
the echo order of clutter and target, which ensure that only can the jamming gain full
coherent accumulation gain. Subsequently, the jamming signal is recovered using pulse
selection and inverse Fourier transform. Eventually, the combination of multiple loss
functions based-encoder–decoder network is utilized to learn both useful information from
the labels and valid semantic information from the time-frequency feature of the recovered
jamming signal. This can improve the accuracy of jamming recognition. The experimental
results shows that the proposed algorithm achieves more than 90% jamming detection
accuracy and over 94% jamming identification accuracy at JCNR>-15 dB even under the
limitation of insufficient training data.
Keywords: Intra-pulse forwarding jamming; Clutter; Non-uniform pulse interval; Encoder–
decoder network; Combination of multiple loss functions

Chaojie Fan, Shuxiang Lin, Baoquan Cheng, Diya Xu, Kui Wang, Yong Peng, Sam Kwong,
EEG-TransMTL: A transformer-based multi-task learning network for thermal comfort
evaluation of railway passenger from EEG,
Information Sciences,
Volume 657,
2024,
119908,
ISSN 0020-0255,
[Link]
([Link]
Abstract: The evaluation of thermal comfort for railway passengers holds considerable
importance, not only in reducing energy consumption but also in enhancing the passengers'
experience. This paper presents a Transformer-based multi-task learning network
(TransMTL) designed for railway passenger thermal comfort evaluation using EEG. We
utilized manual features to extract temporal and frequency information, while a Transformer
encoder distilled spatial information. The multi-task learning structure enhances model
robustness by leveraging thermal comfort task correlations. We conducted experiments
during winter and summer with high-speed railway passengers, establishing a
comprehensive EEG dataset. The results demonstrated that our proposed EEG-TransMTL
model outperformed classical machine learning and deep learning models in all four thermal
comfort evaluation tasks, achieving accuracy rates of 65.00%, 66.70%, 80.38%, and 71.01%,
respectively. We enhanced model interpretability by visualizing attention weights from the
Transformer encoder, identifying key EEG channels. A simplified model utilizing only eight
crucial channels also delivered notable performance. This research provides a practical and
neuro-mechanism interpretable solution for thermal comfort evaluation.
Keywords: Electroencephalogram; Railway passenger; Thermal comfort evaluation; Deep
learning; Interpretable neural network

Pengfei Zheng, Anxue Zhang, Zhensheng Shi, Sen Wang, Yi'an Ma, Zhaodan Liu,
TLAD-YOLO: Lightweight network for intelligent detection of railway tunnel lining anomalies
using ground penetrating radar,
Journal of Applied Geophysics,
Volume 241,
2025,
105869,
ISSN 0926-9851,
[Link]
([Link]
Abstract: Ground Penetrating Radar (GPR) B-scan images and the you only look once (YOLO)
series are widely used for tunnel lining intelligent inspections to ensure quality. However, in
practical applications, lightweight YOLO detection networks fail to meet the requirements of
accuracy and robustness. In view of this, a tunnel lining anomalies detection YOLO (TLAD-
YOLO) is proposed for the intelligent detection of railway tunnel lining anomalies based on
GPR B-scan images. TLAD-YOLO introduces lightweight spatial and channel synergistic multi-
shape attention (SCSMSA) to enhance the detection accuracy of complex scenes and multi-
size abnormal objects, while ghost convolution is used to reduce parameters and
computation. The experiments are conducted on a dataset consisting of 47 railway tunnels.
Furthermore, we propose a multi-scale data augmentation to further expand the dataset,
which improves the detection accuracy. The experimental results demonstrate that TLAD-
YOLO is an accurate and lightweight detection network, outperforming SOTA detection
networks in non-destructive testing of railway tunnels. On the tunnel engineering
verification platform and newly built railway tunnels, TLAD-YOLO demonstrates remarkable
robustness.
Keywords: Tunnel lining; Anomaly detection; Ground penetrating radar (GPR); You only look
once (YOLO); B-scan images

Chenglong Liu, Yuchuan Du, Guanghua Yue, Yishun Li, Difei Wu, Feng Li,
Advances in automatic identification of road subsurface distress using ground penetrating
radar: State of the art and future trends,
Automation in Construction,
Volume 158,
2024,
105185,
ISSN 0926-5805,
[Link]
([Link]
Abstract: Affected by soil erosion and material deterioration, road subsurface is prone to
distress such as cavities, water-rich, and cracks. Ground penetrating radar (GPR), as a real-
time geophysical survey method that uses electromagnetic radiation to image the
subsurface, offers promising non-destructive solutions to road subsurface health monitoring.
However, the interpretation of GPR signals is non-intuitive and obscure in terms of distress
identification, whose performance is also limited by the heterogeneous road condition. In
conjunction with knowledge diagram analysis, a state-of-the-art review is applied to
summarize the advances in the automatic identification of road subsurface distress (RSD).
The algorithms based on the single-channel waveform (A-scan), two-dimensional profile (B-
scan), and three-dimensional data (C-scan) are elaborated from the perspectives of rule-
based recognition algorithm, machine learning algorithm, and deep learning algorithm. In
comparison to analytical methods, the emerging deep learning models have a powerful
ability to extract complex features from multi-dimensional GPR radargrams, enhancing the
efficiency and accuracy of road subsurface distress detection. Recommendations for model
selection are compiled from existing literature together with empirical evidence. The most
significant variables that influence the model selections are thought to be the type of
identified RSD, training sample quality and quantity, prior knowledge, and computational
cost. Some challenges, such as insufficient training samples and diverse road structures, are
presented. Future trends are concluded to draw the implications for GPR research.
Keywords: Road subsurface distress detection; GPR; Automatic identification; Machine
learning; Deep learning

Haiyan Yao, Yuefei Xu, Qiang Guo, Shizhe Chen, Bin Lu, Yuanjun Huang,
Study on transformer fault diagnosisbased on improved deep residual shrinkage network
and optimized residual variational autoencoder,
Energy Reports,
Volume 13,
2025,
Pages 1608-1619,
ISSN 2352-4847,
[Link]
([Link]
Abstract: The transformer as the core equipment in the power system, its fault diagnosis has
a vital role in ensuring the safe and stable operation of the power grid. However, traditional
transformer fault diagnosis methods often rely on manual experience or simple models,
which are difficult to meet the demand for efficient and accurate diagnosis when faced with
complex and evolving fault patterns. In this study, a new method for transformer fault
diagnosis based on improved deep residual shrinkage network (DRSN) and optimized
residual variational autoencoders (ORVAE) is proposed. Firstly, this study improves the DRSN
to enhance its feature extraction capability. By designing a specific shrinkage mechanism,
the improved DRSN can reduce the information loss in the face of complex data, greatly
improve the extraction ability of the key features of the transformer operating state, and
thus improve the accuracy of fault recognition. Secondly, in view of the difficulty and high
cost of transformer fault sample data collection, this study introduces a residual connection
structure based on the traditional variational autoencoder (VAE), and constructs the ORVAE
method to effectively address the challenge of insufficient data. The results show that the
fault recognition rate of the proposed method on the real transformer fault dataset reaches
97.14 %, which is better than the traditional method, showing excellent diagnostic
performance and strong practical application potential. Compared with the existing
technologies, this method not only improves the accuracy of transformer fault diagnosis, but
also provides new ideas and technical support for the intelligent development of power
system. This study offers an innovative solution for the field of fault diagnosis of power
equipment, and providing a strong technical guarantee for fault prediction and maintenance
in future smart grids.
Keywords: Transformer; Fault diagnosis; Improved DRSN; Shrinkage mechanism; Feature
extraction; ORVAE; Recognition rate

Nadav Cohen, Itzik Klein,


Adaptive Kalman-Informed Transformer,
Engineering Applications of Artificial Intelligence,
Volume 146,
2025,
110221,
ISSN 0952-1976,
[Link]
([Link]
Abstract: The extended Kalman filter (EKF) is a widely adopted method for sensor fusion in
navigation applications. A crucial aspect of the EKF is the online determination of the
process noise covariance matrix reflecting the model uncertainty. While common EKF
implementation assumes a constant process noise, in real-world scenarios, the process noise
varies, leading to inaccuracies in the estimated state and potentially causing the filter to
diverge. Model-based adaptive EKF methods were proposed and demonstrated performance
improvements to cope with such situations, highlighting the need for a robust adaptive
approach. In this paper, we derive an adaptive Kalman-informed transformer (A-KIT)
designed to learn the varying process noise covariance online. Built upon the foundations of
the EKF, A-KIT utilizes the well-known capabilities of set transformers, including inherent
noise reduction and the ability to capture nonlinear behavior in the data. This approach is
suitable for any application involving the EKF. In a case study, we demonstrate the
effectiveness of A-KIT in nonlinear fusion between a Doppler velocity log and inertial
sensors. This is accomplished using real data recorded from sensors mounted on an
autonomous underwater vehicle operating in the Mediterranean Sea. We show that A-KIT
outperforms the conventional EKF by more than 49.5% and model-based adaptive EKF by an
average of 35.4% in terms of position accuracy.
Keywords: Inertial sensing; Navigation; Set transformer; Kalman filter; Sensor fusion;
Autonomous underwater vehicle

Purabi Sharma, Kandarpa Kumar Sarma,


Attention driven CWT-deep learning approach for discrimination of Radar PRI modulation,
Physical Communication,
Volume 62,
2024,
102237,
ISSN 1874-4907,
[Link]
([Link]
Abstract: With the proliferation of radio frequency (RF) systems and radar applications,
Electronic Warfare (EW) is receiving increasing importance. The analysis of the radar signals
is a critical EW task that decides the nature of counter employments. In an Electronic
Support (ES) system, the challenge is to detect hostile radiation sources efficiently and
trigger a counter response. Detection of types of Pulse Repetition Interval (PRI) modulation
of radar signal significantly facilitates the manifestation of RF emitters during recognition
which is difficult in a dense EW environment. Recent developments in artificial intelligence
(AI) methods suggest that this emerging technology can be effective for such purposes. In
this direction, an automatic approach for recognizing several kinds of complex PRI
modulation based on Continuous Wavelet Transform (CWT) and a combination of the vanilla
Convolutional Neural Network (CNN), a multi-head self-attention (MHSA) mechanism and
the popular Long Short-Term Memory (LSTM) is proposed. The CWT is used to decompose
the PRI modulation sequence and obtain different time–frequency components. Further,
aided by the proposed CNN-MHSA-LSTM combination, the features extracted from the CWT
2D-scalograms are used to execute PRI modulation discrimination. In this method, the
vanilla CNN is employed for the extraction of deep features to figure out the class details
while capturing the spatial attributes. Thereafter, to improve the discriminative power of the
entire framework a MHSA mechanism is used. The temporal attributes are acquired by the
LSTM which works in concert with the CNN for executing the detection of the PRI classes
based on the extracted features. Also to assess the effectiveness of the proposed method,
three models based on ResNet, popular CNN and SqueezeNet are implemented for
benchmark comparison in terms of overall performance and complexity. The simulation
results show that the proposed method enhances performance and achieves robustness in
the noise-filled and imperfect channel knowledge environment. The best recognition
accuracy is 98.3% with 50% spurious pulses in the environment which fluctuates with
imperfect channel knowledge cases.
Keywords: Electronic warfare; PRI modulation; CWT; Convolutional Neural Network; Self-
attention mechanism; Long Short Term Memory

Min Wang, Hua Wang, Fan Zhang,


Correctformer: A transformer architecture for correcting periodic drift in time-series
forecasting,
Neural Networks,
Volume 196,
2026,
108375,
ISSN 0893-6080,
[Link]
([Link]
Abstract: Time-series forecasting is crucial both in production and daily life. Although the
Transformer architecture has demonstrated strong potential in this field and significantly
improved model performance through continuous iterations of attention mechanism
innovations, these advancements have primarily focused on capturing long-range
dependencies within sequences. However, insufficient attention has been paid to the
inherent periodic patterns in time-series data. Moreover, we observe that applying attention
mechanisms can lead to periodic blurring and even periodic drifts, making it challenging for
the model to capture the true dynamic patterns of time series, ultimately resulting in
degraded forecasting performance. To address this issue, this study proposes Correctformer,
a Transformer architecture integrating periodic embedding and periodic correction to
enhance the capture of periodic features both before and after the attention mechanism.
Specifically, periodic embedding encodes the periodic structural information within time-
series data, enabling the model to better perceive and learn periodic characteristics. Periodic
correction dynamically adjusts the periodic attributes of the data to rectify periodic drift,
restoring the stability of time-series periodicity. Experimental results demonstrate that this
approach provides significant advantages in handling data with complex periodic
characteristics and offers a more suitable Transformer-based architecture for time-series
modeling.
Keywords: Time-series forecasting; Periodic drift; Global periodicity; Periodic embedding;
Periodic correction

Zixuan Wang, Gang Liu, Hanlin Xu, Yao Qian, Rui Chang, Durga Prasad Bavirisetti,
Transformer architecture with illumination aware mechanisms for low-light image
enhancement via Retinex decomposition,
Engineering Applications of Artificial Intelligence,
Volume 162, Part B,
2025,
112414,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Enhancing low-light images is a complex task that involves not only restoring
brightness but also preserving color fidelity and reducing noise interference. In this paper,
we propose a novel Retinex-based Transformer Model with Illumination Aware Mechanisms
(TIMRetinex-Net), which achieves physically interpretable modeling through a
decomposition network guided by Retinex theory. To adapt to light variations in different
regions, we randomly apply gamma transformations to several subregions of the
illumination component and use a Color Estimation Module to capture the color global
distribution of the natural scene in the reflection component. By modeling the color global
distribution and repairing the degraded regions collaboratively, we alleviate the issue of
being highly sensitive to data usage during training and improve the model’s ability to
handle unknown scenes. The Illumination and Reflection Adjustment Transformer Network
(IRAT-Net) produces enhanced images, achieving a balanced enhancement of detail and
color. In addition, IRAT-Net incorporates an attention mechanism into the feature extraction
layer and introduces the Illumination-Guided Information Aggregation Module to adaptively
estimate lighting conditions. In the field of image processing, our method based on artificial
intelligence was evaluated on five datasets and compared with twelve state-of-the-art
methods. The results demonstrated strong alignment with the ground truth, with our
method achieving superior performance in both subjective and objective assessments.
Keywords: Low-light image enhancement; Retinex decomposition; Transformer; Image
restoration; Deep learning

Shuai Wang, Dehao Zhang, Ammar Belatreche, Yichen Xiao, Hongyu Qing, Wenjie Wei, Malu
Zhang, Yang Yang,
Ternary spike-based neuromorphic signal processing system,
Neural Networks,
Volume 187,
2025,
107333,
ISSN 0893-6080,
[Link]
([Link]
Abstract: Deep Neural Networks (DNNs) have been successfully implemented across various
signal processing fields, resulting in significant enhancements in performance. However,
DNNs generally require substantial computational resources, leading to significant economic
costs and posing challenges for their deployment on resource-constrained edge devices. In
this study, we take advantage of spiking neural networks (SNNs) and quantization
technologies to develop an energy-efficient and lightweight neuromorphic signal processing
system. Our system is characterized by two principal innovations: a threshold-adaptive
encoding (TAE) method and a quantized ternary SNN (QT-SNN). The TAE method can
efficiently encode time-varying analog signals into sparse ternary spike trains, thereby
reducing energy and memory demands for signal processing. QT-SNN, compatible with
ternary spike trains from the TAE method, quantifies both membrane potentials and synaptic
weights to reduce memory requirements while maintaining performance. Extensive
experiments are conducted on two typical signal-processing tasks: speech and
electroencephalogram recognition. The results demonstrate that our neuromorphic signal
processing system achieves state-of-the-art (SOTA) performance with a 94% reduced
memory requirement. Furthermore, through theoretical energy consumption analysis, our
system shows 7.5× energy saving compared to other SNN works. The efficiency and efficacy
of the proposed system highlight its potential as a promising avenue for energy-efficient
signal processing.
Keywords: Quantization spiking neural networks; Neural encoding for signals; Neuritic signal
processing; Ternary spiking neural networks; Keyword spotting and EEG

V.I. Mel'nikov, V.V. Ivanov, I.A. Teplyashin,


The study of ultrasonic reflex-radar waveguide coolant level gage for a nuclear reactor,
Nuclear Energy and Technology,
Volume 2, Issue 1,
2016,
Pages 37-41,
ISSN 2452-3038,
[Link]
([Link]
Abstract: Results of experimental study of operation of ultrasonic reflex-radar waveguide
level gage in water coolant at elevated parameters with pressure up to 18MPa and
temperature up to 350°C are examined. In contrast to the known waveguide level gages,
traveltime of acoustic pulses along the waveguide from the radiator to the subsurface layer
and back is measured in the level gage under study. Waveguide consists of two acoustically

zero-order flexural waves and piezoelectric transformers operated at frequency of ∼800kHz


isolated waveguides – the radiating waveguide and the receiving waveguide. Waveguides of

are applied. Processing of received signals is performed by microprocessor-based electronic


circuit. Measurement uncertainty does not exceed ±10mm. Description of the experimental
setup and the experimental methodology is provided. The instrument works reliably and
does not require introducing corrections of readings when coolant thermal physical
properties change. The measurement instrument is intended for application in heat
exchanging equipment in thermal and nuclear power generation.
Keywords: Ultrasonic reflex-radar waveguide level gage; Acoustic waveguide; Piezoelectric
transformer; Water coolant operated at elevated parameters (350°C, 18MPa); Nuclear
power installation; Power generating equipment

Chunyu Zhu, Tinghao Zhang, Qiong Wu, Yachao Li, Qin Zhong,
An Implicit Transformer-based Fusion Method for Hyperspectral and Multispectral Remote
Sensing Image,
International Journal of Applied Earth Observation and Geoinformation,
Volume 131,
2024,
103955,
ISSN 1569-8432,
[Link]
([Link]
Abstract: There is an effective way to enhance the spatial resolution of hyperspectral remote
sensing images by fusing them with multispectral remote sensing images. However, most of
the existing deep fusion techniques adopt discretized explicit models to approximate the
complex continuous nonlinear mapping in the fusion process, leading to limitations in
enhancing the fidelity of spatial details. Additionally, existing algorithms commonly utilize
discrete methods such as bilinear or bicubic interpolation during the hyperspectral
upsampling process, leading to the loss of crucial spatial-spectral features. To this end, this
study proposes a novel Implicit Transformer Fusion Generative Adversarial Network (ITF-
GAN), which incorporates the continuity perception mechanism of implicit neural
representation with the powerful self-attention mechanism of the Transformer architecture,
which uses point-to-point implicit functions aiming to efficiently process information in both
spatial and spectral dimensions. Besides, a guided implicit neural sampling module is
introduced in the hyperspectral image up-sampling process to enhance the coordinated
expression of features in the spatial and spectral domains, which improves the spatial
resolution and spectral fidelity of the fused image during the upsampling process. A series of
fusion experiments including 4x, 8x, and 16x scale factors have shown that ITF-GAN has
significant advantages over current popular fusion algorithms in both objective evaluation
indicators and subjective visual evaluation.
Keywords: Implicit Neural Repersentation; Image fusion; Implicit Transformer; ITF-GAN

Sajid Ullah, Xi Chen, Han Han, Junhao Wu, Jinghan Dong, Ruiqing Liu, Weijie Ding, Min Liu,
Qingli Li, Honggang Qi, Yonggui Huang, Philip Lh Yu,
A novel hybrid ensemble approach for wind speed forecasting with dual-stage
decomposition strategy using optimized GRU and transformer models,
Energy,
Volume 329,
2025,
136739,
ISSN 0360-5442,
[Link]
([Link]
Abstract: Wind energy has attracted global interest owing to its sustainable and
environmentally friendly characteristics. Nevertheless, precisely forecasting wind speed can
be challenging due to its volatile and unpredictable nature. This paper presents a new hybrid
forecasting approach based on dual stage decomposition mechanism, namely TMQGDT for
wind speed prediction. At first, a decomposition technique called time-varying filtered based
empirical mode decomposition (TVFEMD) is utilized to decompose the original wind speed
data into several intrinsic mode functions (IMFs). Afterwards, multi-scale permutation
entropy (MPE) is used to assess the complexity of each IMF. Based on the entropy values,
the IMFs are further classified into high-frequency and low-frequency IMFs. To address the
significant volatility observed in the high-frequency IMFs, discrete wavelet transform (DWT)
method is employed to perform secondary decomposition. The low-frequency IMFs are
forecasted using gated recurrent unit (GRU) model optimized with quantum particle swarm
optimization (QPSO) algorithm, while the high-frequency IMFs are forecasted with the
Transformer model. The proposed model is trained and validated using four wind speed time
series datasets collected from Germany and China. Five individual models and six hybrid
models are compared against the proposed model to validate the forecasting performance
of the proposed TMQGDT model. The prediction outcomes reveals that the R2 of the model
is 0.973, 0.968, 0.956, and 0.996 on the four dataset test sets, which has improved by
3.39 %, 3.93 %, 5.53 %, and 0.50 %, respectively, compared to the TVFEMD-MPE-QPSO-GRU-
DWT-Autoformer model. The excellent accuracy performance of the TMQGDT model
indicates that developing a hybrid model based on deep learning techniques using
secondary decomposition mechanism and optimization algorithm can enhance the precision
of wind speed prediction.
Keywords: Wind speed prediction; Time-varying filtered based empirical mode
decomposition; Discrete wavelet transform; Quantum particle swarm optimization; Multi-
scale permutation entropy

Yu-Jin Jeon, Min Jeong Hong, Chan Seop Ko, So Jin Park, Hyein Lee, Won-Gyeong Lee, Dae-
Hyun Jung,
A hybrid CNN-Transformer model for identification of wheat varieties and growth stages
using high-throughput phenotyping,
Computers and Electronics in Agriculture,
Volume 230,
2025,
109882,
ISSN 0168-1699,
[Link]
([Link]
Abstract: Wheat (Triticum aestivum L.) is a major crop consumed and cultivated worldwide,
with various varieties bred to meet the growth characteristics and resistances required by
the climate conditions of each cultivation region. However, as global climate change
accelerates, rapid environmental shifts have led to crossover interaction, where previously
superior varieties undergo changes, making variety selection increasingly challenging. In
particular, there is a lack of research on methods for rapidly assessing growth rate, a key
characteristic of crossover interactions, during the variety selection process. This study
proposes a deep learning-based model and method for identifying wheat varieties and
growth stages using hyperspectral data obtained from six wheat varieties cultivated on a
high-throughput phenotyping platform. The proposed model, which combines a CNN and
Transformer, achieved 94.05% accuracy in variety detection, surpassing the performance of
related studies, and 99.24% accuracy in growth stage detection. This model enables high-
throughput monitoring of wheat variety and growth information effectively. Furthermore, if
applied to the wheat breeding process, the proposed model is expected to contribute to the
rapid selection of superior varieties suited to specific climate conditions.
Keywords: Phenotyping platform; Hyperspectral imaging; Self-attention mechanism;
Convolutional neural networks

Zhifei Liu, Kang Zheng, Yongze Song, Jianing Zhang,


Daily high-resolution PM2.5 mapping using spatiotemporal CNN-transformer-KAN model,
International Journal of Applied Earth Observation and Geoinformation,
Volume 144,
2025,
104900,
ISSN 1569-8432,
[Link]
([Link]
Abstract: Daily high-resolution mapping of fine particulate matter (PM2.5) is critical for air
quality monitoring and public health. However, current methods struggle to achieve high
accuracy over large spatial and temporal scales due to limitations in modeling complex
spatiotemporal dependencies. This study proposed a novel hybrid deep learning model—
CNN-Transformer-KAN Network (CTKNet)—which utilizes Convolutional Neural Networks
(CNN) to capture spatial features, Transformers for capturing long-range dependencies, and
the Kolmogorov–Arnold Network (KAN) for nonlinear representation learning. Utilizing
spatially continuous satellite aerosol optical depth (AOD) data and other multi-source
spatiotemporal inputs, CTKNet estimated daily PM2.5 at a spatial resolution of 1 km across
China for the period 2015–2020, marking the first application of KAN in PM2.5 estimation. It
outperformed existing models, achieving a high cross-validation coefficient of determination
(R2) of 0.95 (sample-based), 0.90 (station-based), and 0.78 (time-based), and corresponding
RMSEs of 8.13, 11.03, and 17.76 µg/m3. Yearly sample-based cross-validation R2 values
ranged from 0.91 to 0.96 with RMSEs below 11.74 µg/m3, while seasonal R2 values ranged
from 0.82 to 0.89 with RMSEs below 22.14 µg/m3. Analysis reveals a significant decline in
PM2.5 nationwide, especially in eastern and central China. Seasonal peaks occur in winter,
with minima in summer, influenced by meteorology. Spatially, PM2.5 is highest in eastern
and northern regions; urban agglomerations like Beijing–Tianjin–Hebei (BTH) show severe
pollution, while Pearl River Delta (PRD) exhibits the lowest levels due to favorable
conditions. CTKNet also holds promise for other fine-scale environmental mapping tasks
using multi-source spatiotemporal data.
Keywords: PM2.5 estimation; Kolmogorov–Arnold Network; Hybrid deep learning model;
Satellite AOD; Air quality assessment

Wenxu Zhang, Fosheng Zhang, Zhongkai Zhao, Feiran Liu,


Radar specific emitter identification via the Attention-GRU model,
Digital Signal Processing,
Volume 142,
2023,
104198,
ISSN 1051-2004,
[Link]
([Link]
Abstract: In radar specific emitter identification (SEI), various types of unintentional
modulation on pulse (UMOP) are selected as the features for discriminating between
different radars. Unintentional Phase Modulation on Pulse (UPMOP), a typical type of UMOP,
can provide crucial information for identifying radars. In most radar SEI algorithms,
sacrificing time efficiency for higher accuracy is a common trade-off. This paper proposes a
method to solve this problem by combining denoised UPMOP sequences with an Attention-
based Gated Recurrent Units (Attention-GRU) model, which showed an excellent
performance. Firstly, the cause of UPMOP is analyzed and the phase observation model of
radar emitter signals and mathematical model of UPMOP are given. Then, the least-squares
method is used to eliminate the linear trend of the phase observation model and obtain a
noised estimation of the UPMOP sequences. Thirdly, the uniform B-spline (UBS) curves are
then used to fit the noised estimation, resulting in a denoised and refined UPMOP sequence.
Finally, the Attention-GRU model is employed to extract features from the denoised UPMOP
sequences to identify radar emitters automatically. Results from simulation and measured
data experiments show that the overall recognition rate of the algorithm reaches over 93%
and the algorithm has excellent performance, with high identification accuracy and relatively
low time consumption, even in low signal-to-noise ratio (SNR) conditions.
Keywords: Radar specific emitter identification; Unintentional phase modulation on pulse;
Uniform B-spline curve; Time-series model; Attention mechanism

Fei Xiong, Weili Kou, Yuhan Xun, Yinuo He, Bo Hu, Xinchen Ye, Yongke Sun,
A unified Vision Transformer (ViT) backbone with Penalty Outside Point Loss for monocular
body measurement of Binglangjiang buffaloes,
Computers and Electronics in Agriculture,
Volume 240,
2026,
111166,
ISSN 0168-1699,
[Link]
([Link]
Abstract: Binglangjiang buffalo (Bubalus bubalis), a China’s national-level protected livestock
genetic resource, lacks comprehensive body measurement data due to its nervous
temperament, which impedes conservation and research. However, existing non-contact
measurement schemes often rely on expensive multi-camera setups (depth or point cloud),
leading to high equipment costs and complex deployment. To overcome these limitations,
we propose a cost-effective monocular camera-based body measurement method. It
employs a unified Vision Transformer-based backbone network for feature extraction,
seamlessly integrating 2D image keypoints detection along with depth estimation via a
dedicated head, and introduces a novel Penalty Outside Point Loss. This loss function
enhances keypoints localization accuracy by penalizing predictions outside the body region,
outperforming conventional loss functions in boundary-sensitive scenarios. Experimental
results show that the depth estimation achieves an absolute relative error of 0.159, an
average precision of 94.83% for keypoints detection and an interior point ratio of 95.14%.
The mean absolute percentage errors for body height, hip height, oblique body length, chest
circumference, and abdominal circumference are 7.37%, 6.79%, 13.67%, 8.39%, and 7.35%,
respectively. By applying this method, we have successfully completed body measurements
for 424 Binglangjiang buffaloes, effectively filling the long-standing gap in comprehensive
body measurement data for this breed. This study establishes a reliable, low-cost framework
for buffalo body measurement data, offering crucial technical support for efficient
conservation and advancing precision livestock management practices.
Keywords: Binglangjiang buffalo; Body measurement; Vision Transformer; Monocular depth
estimation; Keypoints detection

Xinyue Xin, Ming Li, Yan Wu, Peng Zhang, Dazhi Xu,
DCDLNet: A label-noise tolerant classification algorithm for polsar images based on dual-
band consistency and difference,
Knowledge-Based Systems,
Volume 334,
2026,
115120,
ISSN 0950-7051,
[Link]
([Link]
Abstract: With the advancement of technology, PolSAR systems can acquire multiple signals
by transmitting and receiving electromagnetic waves in different frequency bands, thereby
enabling the collection of richer ground observation information. However, due to the lack
of consideration for the concepts of dual-band consistency and dual-band difference,
existing fusion methods still encounter problems of incomplete semantic information and
low computational efficiency. Moreover, in practice, the process of sample labeling often
involves manual intervention, which inevitably introduces labeling errors. To tackle these
problems, we propose a novel label-noise tolerant classification framework called DCDLNet:
dual-band consistency and difference learning network. Specifically, to extract the rich
information contained in dual-band PolSAR data, the DCDLNet comprises two principal
parts. The first part is an inter-band difference acquisition module (IDAM), which learns
dual-band complementary information based on the concept of dual-band difference. The
second part is a spatial-domain and frequency-domain feature extraction (SFFE) module. It
acquires more discriminative information by capturing local spatial information in the
spatial-domain and global spatial information in the frequency-domain. Furthermore, by
integrating the concept of dual-band consistency and the fitting capabilities of neural
networks, DCDLNet adopts a cross-band and bidirectional supervised (CBS) strategy to
mitigate the impact of label noise during the training process. Experiments on measured
PolSAR datasets demonstrate that our method outperforms several existing approaches in
terms of dual-band fusion and noisy label processing.
Keywords: Polarimetric synthetic aperture radar; Image classification; Dual-band fusion;
Label noise

Stavros Kassinos, Alessio Alexiadis,


Beyond language: Applying MLX transformers to engineering physics,
Results in Engineering,
Volume 26,
2025,
104871,
ISSN 2590-1230,
[Link]
([Link]
Abstract: Transformer Neural Networks are driving an explosion of activity and discovery in
the field of Large Language Models (LLMs). In contrast, there have been only a few attempts
to apply Transformers in engineering physics. Aiming to offer an easy entry point to physics-
centric Transformers, we introduce a physics-informed Transformer model for solving the
heat conduction problem in a 2D plate with Dirichlet boundary conditions. The model is
implemented in the machine learning framework MLX and leverages the unified memory of
Apple M-series processors. The use of MLX means that the models can be trained and
perform predictions efficiently on personal machines with only modest memory
requirements. To train, validate and test the Transformer model we solve the 2D heat
conduction problem using central finite differences. Each finite difference solution in these
sets is initialized with four random Dirichlet boundary conditions, a uniform but random
internal temperature distribution and a randomly selected thermal diffusivity. Validation is
performed in-line during training to monitor against over-fitting. The excellent performance
of the trained model is demonstrated by predicting the evolution of the temperature field to
steady state for the unseen test set of conditions.
Keywords: Physics-informed transformers; MLX framework; Heat conduction

Haoyu Wang, Chuanjiang Li, Peng Ding, Shaobo Li, Tandong Li, Chenyu Liu, Xiangjie Zhang,
Zejian Hong,
A novel transformer-based few-shot learning method for intelligent fault diagnosis with
noisy labels under varying working conditions,
Reliability Engineering & System Safety,
Volume 251,
2024,
110400,
ISSN 0951-8320,
[Link]
([Link]
Abstract: Recent years have witnessed the success of Few-shot Learning (FSL) methods in
equipment reliability enhancement and fault diagnosis, by virtue of learning from limited
data and adapting to new operating conditions. However, due to sensor bias, manual
collection, and mislabeling, label noise is inevitably introduced into the dataset, which
further reduces the quality of supervised information contained in the few-shot dataset,
posing significant challenges for accurate fault diagnosis. In this paper, the problem of Few-
shot Fault Diagnosis with Noisy Labels (FFDNL) is studied for the first time, and a novel
method named Enhanced Transformer with Asymmetric Loss Function (ETALF) is proposed.
ETALF leverages the self-attention mechanism of the transformer to dynamically measure
the similarity between fault samples in the support set to enhance the model's robustness
against label noise, then naturally aggregates the similar samples into corresponding correct
prototypes. Furthermore, an asymmetric loss function is designed, which adaptively assigns
the model with larger penalties for incorrect category predictions and smaller penalties for
correct category predictions, thereby enhancing fault diagnostic performance through
inherent asymmetry. Comprehensive experiments are conducted on two benchmark
datasets, and the compared results with representative approaches validate the
effectiveness of our proposed ETALF in performing intelligent fault diagnosis using limited
and noise-labeled data under varying working conditions, which achieves accuracies of
97.77% and 95.78% with 0.2 noisy-level labels during meta-training and meta-testing on the
CWRU and KAIST datasets, respectively.
Keywords: Few-shot learning; Noisy label; Intelligent fault diagnosis; Transformer;
Asymmetric loss function

Tao Zhou, Dechen Yao, Jianwei Yang, Chang Meng, Ankang Li, Xi Li,
DRSwin-ST: An intelligent fault diagnosis framework based on dynamic threshold noise
reduction and sparse transformer with Shifted Windows,
Reliability Engineering & System Safety,
Volume 250,
2024,
110327,
ISSN 0951-8320,
[Link]
([Link]
Abstract: In real industrial environments, acquiring vibration data from bearings is often
challenging due to noise, resulting in network models that excel when trained on datasets
with sufficient samples but struggle with accurate fault identification in real-world scenarios,
inevitably threatening the reliability of fault diagnosis. To address this problem, this paper
proposes an end-to-end fault diagnosis framework (DRSwin-ST) based on sparse transformer
with a shift window and dynamic threshold noise reduction. The Swin-Transformer serves as
the backbone, leveraging a multi-head self-attention mechanism with a shift window to
capture global information. The 1.5-Entmax replaces Softmax in the self-attention
mechanism, sparsifying irrelevant information and allowing the model to focus on essential
details. The self-attention mechanism, combined with a multi-scale structure, forms a
forward feedback network to obtain rich fault feature information. In addition, the paper
integrates a large convolutional kernel and a dynamic soft-threshold noise reduction module
to construct a convolutional network in front of the transformer structure. This configuration
extracts fault feature information and removes the noise, enhancing the fault recognition
accuracy of the model. Experimental results on three diverse datasets demonstrate that
DRSwin-ST exhibits robustness and high accuracy even in scenarios with limited samples and
high noise, validating its exceptional performance.
Keywords: Few samples; High noise; DRSwin-ST; Fault diagnosis

Sheng Kuang, Jie Shi, Kiki van der Heijden, Siamak Mehrkanoon,
BAST-Mamba: Binaural Audio Spectrogram Mamba Transformer for binaural sound
localization,
Neurocomputing,
Volume 650,
2025,
130804,
ISSN 0925-2312,
[Link]
([Link]
Abstract: Accurate sound localization in reverberant environments is essential for human
auditory perception. Recently, Convolutional Neural Networks (CNNs) have been used to
model the binaural human auditory pathway. However, CNNs face limitations in capturing
global acoustic features. To address this issue, we propose a novel end-to-end Binaural
Audio Spectrogram Mamba Transformer (BAST-Mamba) model to predict sound azimuth in
both anechoic and reverberant conditions. We explore two implementation modes: BAST-
Mamba-SP and BAST-Mamba-NSP, which correspond to shared and non-shared parameter
configurations, respectively. Our best model BAST-Mamba-SP, equipped with subtraction-
based interaural integration and a hybrid loss function, achieves a state-of-the-art angular
distance (AD) error of 0.89°and mean squared error of 0.0004, significantly outperforming
baseline models. The model demonstrates generalization across acoustic environments,
robust hemifield symmetry and high accurate real-time localization performance (<4°AD at
300 ms). Moderate noise augmentation at 30 dB SNR yields the strongest noise resilience.
Explainability analyses highlight consistent frequency focus in the 2–3 kHz and 5.5–6.5 kHz
bands, aligning with known neurophysiological cues. These results validate the potential of
neurobiologically inspired Transformer for robust, high-precision sound localization and offer
new insights into human sound localization.
Keywords: Transformer; Sound localization; Binaural integration

Ge Junkai, Sun Huaifeng, Shao Wei, Liu Dong, Yao Yuhong, Zhang Yi, Liu Rui, Liu Shangbin,
GPR-TransUNet: An improved TransUNet based on self-attention mechanism for ground
penetrating radar inversion,
Journal of Applied Geophysics,
Volume 222,
2024,
105333,
ISSN 0926-9851,
[Link]
([Link]
Abstract: Convolutional Neural Networks (CNN) are widely applied to Ground Penetrating
Radar (GPR) inversion because they have strong data-driven capabilities and are suitable for
the data structure form of GPR. For CNN, the computation increases with the distance that
the convolutional block moves from one region to another when it calculates the
relationship between two regions. For GPR data, the target reflection exists in the
surrounding traces and full time-window of the target, which leads to high degree of remote
relationship. In this paper, we propose GPR-TransUNet, a deep-learning based inversion
network which use self-attention mechanism. According to the characteristics of GPR data,
regression network and GPR-Loss mechanism were used. Both numerical and model
experiments were arranged to test the performance of the network, and the result as well as
comparative analysis demonstrate the superiority of GPR-TransUNet. Finally, we applied this
method to the field GPR data of Guangxi as an attempt.
Keywords: GPR; Inversion; Deep learning

Sheng Li, J.C. Ji, Yadong Xu, Ke Feng, Ke Zhang, Jingchun Feng, Michael Beer, Qing Ni, Yuling
Wang,
Dconformer: A denoising convolutional transformer with joint learning strategy for
intelligent diagnosis of bearing faults,
Mechanical Systems and Signal Processing,
Volume 210,
2024,
111142,
ISSN 0888-3270,
[Link]
([Link]
Abstract: Rolling bearings are the core components of rotating machinery, and their normal
operation is crucial to entire industrial applications. Most existing condition monitoring
methods have been devoted to extracting discriminative features from vibration signals that
reflect bearing health status. However, the complex working conditions of rolling bearings
often make the fault-related information easily buried in noise and other interference.
Therefore, it is challenging for existing approaches to extract sufficient critical features in
these scenarios. To address this issue, this paper proposes a novel CNN-Transformer
network, referred to as Dconformer, capable of extracting both local and global
discriminative features from noisy vibration signals. The main contributions of this research
include: (1) Developing a novel joint-learning strategy that simultaneously enhances the
performance of signal denoising and fault diagnosis, leading to robust and accurate
diagnostic results; (2) Constructing a novel CNN-transformer network with a multi-branch
cross-cascaded architecture, which inherits the strengths of CNNs and transformers and
demonstrates superior anti-interference capability. Extensive experimental results reveal
that the proposed Dconformer outperforms five state-of-the-art approaches, particularly in
strong noisy scenarios.
Keywords: Rolling bearing; Fault diagnosis; Vibration signal; Dconformer; Complex working
conditions; Noisy scenarios

Yanming Gu, Zhuhua Hu, Yaochi Zhao, Jianglin Liao, Weidong Zhang,
MFGTN: A multi-modal fast gated transformer for identifying single trawl marine fishing
vessel,
Ocean Engineering,
Volume 303,
2024,
117711,
ISSN 0029-8018,
[Link]
([Link]
Abstract: In order to achieve sustainable development of marine fishery resources, effective
supervise of trawl fishing during forbidden fishing period is of great significance. This paper
addresses the challenges of poor generalization and the lack of unstructured information in
the precise identification of single trawler fishing behavior. We propose a Transformer
network with multi-source information fusion processing (MFGTN), which accurately
classifies fishing vessels as single trawl or non-single trawl vessels. Firstly, a private fishing
dataset of single trawl behavior is constructed by integrating AIS data with radar data,
named HaiNan_SingleTrawlVessel(HN_STV). Subsequently, as fused data lacks unstructured
information, it undergoes transformation into trajectory point images and recurrence plot
images to reveal the internal structure of the fused data. As such, a visual module is
introduced to handle the trajectory point images and recurrence plot images as a branch.
Simultaneously, the fused data are input into a Double-Tower Transformer with Dual-gate
structures to extract information in different dimensions of the time series and feature space
as two separate branches. The Fast Attention module replaces the traditional Attention
module to improve network speed and reduce memory consumption. Ultimately, the output
of the three branches are fused and controlled by a Dual-gate structure that can
autonomously learn to determine the network output. Experimental results show that
compared to the current best-performing methods, the method discussed herein on the
HN_STV dataset has improved the accuracy, recall, precision, and F1-score performance
indicators by 2.34%, 2.46%, 0.97%, and 1.39%, respectively. The AUC area on the ROC curve
increased by 4%. In a public dataset including three fishing activities, the proposed method
improved accuracy, recall, precision, and F1-score by 2.95%, 2.59%, 2.25%, and 2.70%,
respectively, and the AUC area on the ROC curve increased by 3%. And in all experiments,
our network incurs the lowest time cost. Therefore, the method proposed herein
demonstrates its advanced performance.
Keywords: Deep learning; Data fusion; Automatic identification system; Ship trajectory
classification; Recurrence plot image

Yanrong Wang, Zihan Wang, Wanqing Zeng, Jingbao Wang, Zhiqiang Wang, Yubin Lan,
Identification of the geographical origin of wolfberry by synergetic application of electronic
eye and near-infrared spectroscopy combined with a Swin Transformer multi-scale fusion
model,
Microchemical Journal,
Volume 213,
2025,
113800,
ISSN 0026-265X,
[Link]
([Link]
Abstract: The nutritional effects and commercial value of wolfberry largely depend on its
geographical origin. This study proposed a novel method to identify the origin of wolfberry
by applying an electronic eye (EE) and near-infrared (NIR) spectroscopy combined with a
Swin Transformer multi-scale fusion model (STMIFNet). First, the exterior image and internal
quality information of wolfberry samples are collected by EE and NIR spectroscopy,
respectively. Subsequently, the Continuous Wavelet Transform (CWT) is implemented to
convert the NIR spectra into a two-dimensional (2D) spectrogram, thereby enhancing the
analysis and interpretation of spectral information. A multi-scale fusion model is further
proposed to perform feature extraction and pattern recognition based on the obtained EE
images and NIR spectrograms. This model utilizes the Swin Transformer to extract local and
global multi-scale features and incorporates multiple Information Interactive Fusion (IIF)
modules to facilitate the interaction of information between the EE images and NIR
spectrograms. The experimental results indicate that the proposed method yields more
comprehensive and accurate identification compared to using EE or NIR spectroscopy
individually. Compared to traditional machine learning methods and deep learning models,
the proposed STMIFNet demonstrates superior recognition accuracy and stronger
generalization ability. On the test set, the model achieves an accuracy, precision, recall, and
F1-score of 99.00%, 99.02%, 99.00%, and 0.9899, respectively. This study provides a rapid,
efficient, and environmentally friendly method for identifying the geographical origin of
wolfberries, which has great potential for applications in traceability detection of other food
types.
Keywords: Origin of wolfberry; Electronic eye; Near-infrared Spectroscopy; Swin
transformer; Information interactive fusion

Wandi Wang, Mahdi Motagh, Zhuge Xia, Simon Plank, Zhe Li, Aiym Orynbaikyzy, Chao Zhou,
Sigrid Roessner,
A framework for automated landslide dating utilizing SAR-Derived Parameters Time-Series,
An Enhanced Transformer Model, and Dynamic Thresholding,
International Journal of Applied Earth Observation and Geoinformation,
Volume 129,
2024,
103795,
ISSN 1569-8432,
[Link]
([Link]
Abstract: Determining the timing of landslide occurrence is crucial for establishing an
accurate, comprehensive and systematic landslide inventory while assessing the potential for
reducing landslide risk. Unfortunately, many existing landslide inventories lack temporal
information such as the precise time of landslide events. Optical and Synthetic Aperture
Radar (SAR) sensors are the most commonly used remote sensing technologies for landslide
detection. Unlike optical sensors, SAR sensors are not affected by cloudy conditions and
provide valuable imagery regardless of sunlight availability. Therefore, SAR-derived
parameters, i.e., SAR amplitude, interferometric coherence, and polarimetric features (alpha
and entropy), offer a higher temporal resolution for detecting landslide occurrence times
compared to optical data. Despite the advantages, there is currently no universally accepted
automatic method for determining the time of landslide events using SAR data. This is due to
the lack of anomaly labels and the high time-series volatility in detecting landslide
occurrence times. Despite advances in deep-learning methods for anomaly detection in
time-series, only a few of them can address these challenges in our case. In this paper, we
propose an unsupervised multivariate transformed-based deep-learning model to
automatically and efficiently estimate landslide occurrence times using multivariate SAR-
derived parameters time-series analysis. The designed gated relative position can increase
robustness and temporal context information, by learning global temporal trends in the
time-series. Subsequently, the time-series of the anomaly score derived from the proposed
Transformer model is analyzed using an adaptive thresholding strategy to dynamically and
automatically mark anomalies related to the landslide occurrence. Our research focuses on
collapsed landslides characterized by dramatic changes in ground surface topography, with a
particular attention for the need of a prior knowledge about landslide boundaries. We assess
the performance of the proposed methodology for several collapsed landslides including the
July 21, 2020 Shaziba and 23 July, 2019 Shuicheng landslides in China, March 19, 2019 Takht
landslide in Iran, June 15, 2018 Jalgyz-Jangak and May 25, 2018 Kugart landslides in
Kyrgyzstan, July 7, 2018 Hitardalur landslide in Iceland, and January 25, 2019 Brumadinho
landslide in Brazil. In comparison to commonly used neural networks like the LSTM
algorithm, our proposed framework leads to a more accurate estimate for the time of
landslide failure using time-series of SAR-derived parameters. Furthermore, our results
suggest the great potential of SAR data to narrow the time period detected from optical data
when used in conjunction with them.
Keywords: Landslide; SAR; Anomaly detection; Deep-learning

Qiming Cheng, Yihong Su, Yang He, Yang Wu, Fei Liu, Ye Rao, Yunsong Chao, Kaifeng Wang,
Zhen Liu, Jun Liu, Yao Chen,
Enhanced radar echo extrapolation for precipitation nowcasting quality using the
convolutional Kolmogorov–Arnold networks,
Journal of Hydrology,
Volume 663, Part A,
2025,
134134,
ISSN 0022-1694,
[Link]
([Link]
Abstract: With the ongoing climate warming, recurrent extreme rainfall events have become
a pervasive global challenge. The integration of disaster warnings with precipitation
nowcasting can effectively mitigate both human casualties and economic losses. Currently,
deep learning techniques are widely employed for radar echo extrapolation as a primary
approach to precipitation nowcasting. However, the lack of physical constraints often leads
to blurriness in predicted images as the forecast time increases, ultimately resulting in a
decline in forecast quality. In this study, we proposed an evolution network that incorporates
Kolmogorov-Arnold networks (KANs) to extract physical motion information and enhance
the network’s ability to learn such information. The results demonstrated that employing
convolutional KANs (ConvKANs) as the fundamental module significantly reduced the
number of model parameters while achieving superior performance. ConvKAN proved to be
a highly effective foundational module for the evolution network, not only substantially
reducing the number of model parameters and effectively improving specific meteorological
metrics (CSI, POD) and the clarity metric of predicted images (Tenengrad), but also achieving
the best performance in precipitation forecasting. Notably, our CUX2&evOnet (K2) model
combining convolutional modules exhibited optimal performance. Compared to the CNN-
based model named CUX2&evOnet(C32), it improved CSI, POD, and Tenengrad by 2.02%,
5.17%, and 2.52%, respectively, while requiring only 2.37% of its evolution network
parameters. These findings confirm that deep learning models driven by the Kolmogorov-
Arnold theorem (KAT) exhibit enhanced proficiency in precipitation nowcasting.
Keywords: Precipitation nowcasting; Deep learning; Kolmogorov-Arnold networks; Physical
constraints; Radar echo extrapolation

Yongni Shao, Dan Chen, Binggan Wang, Chen Zhao, Yun Tang, Yan Peng, Huiping Zhang,
Yiming Zhu, Wenchao Tang,
Identification of Pueraria lobata origin using terahertz precision spectroscopy and CNN-
transformer hybrid network algorithm,
Spectrochimica Acta Part A: Molecular and Biomolecular Spectroscopy,
Volume 348, Part 1,
2026,
127212,
ISSN 1386-1425,
[Link]
([Link]
Abstract: This study addresses the underexplored potential of terahertz (THz) spectroscopy
for geographical origin authentication of Pueraria lobata. We developed a novel non-
destructive approach integrating THz spectroscopy with a CNN-Transformer hybrid network
to classify samples from eight Chinese regions. High-Performance Liquid Chromatography
(HPLC) validated the correlation between THz spectral features and bioactive components
(puerarin, daidzein, daidzin). Comparative analysis with Raman spectroscopy and five
machine learning algorithms demonstrated THz spectroscopy's superiority. Under optimal
conditions, the hybrid model reached an accuracy of 91.67%, which is significantly higher
than that of traditional methods (60.42% to 64.58%) and the standard CNN
architecture(85.42%). Additionally, it achieved perfect classification (F1-score = 1.000) for
the Jiangxi/Shaanxi [Link] results establish THz spectroscopy coupled with deep
learning as a robust, accurate tool for origin traceability in traditional medicine quality
control.
Keywords: Terahertz spectroscopy; Raman spectroscopy; Pueraria lobata; Origin
identification; CNN-transformer hybrid network

Shabing Ye, Qian He,


MIMO radar moving target detector with model-aided learning,
Digital Signal Processing,
Volume 161,
2025,
105130,
ISSN 1051-2004,
[Link]
([Link]
Abstract: The moving target detection (MTD) problem in multiple-input multiple-output
(MIMO) radar usually involves a large number of unknown parameters, such as clutter-plus-
noise parameters, target reflection coefficients, and target velocity. The presence of these
numerous unknowns poses significant challenges to traditional model-based detector such
as generalized likelihood ratio test (GLRT), whose performance quickly deteriorates when
the number of unknown parameters increases. Recently, learning-based method has been
recognized as an effective strategy to deal with MTD. However, it requires a large number of
training data to ensure accuracy, which is not easily acquired in practical scenarios. Statistical
model is derived from historical data, and therefore incorporating this prior knowledge into
learning network intuitively reduces the number of training data. In this work, we propose a
model-aided learning moving target detector that combines the prior knowledge provided
by the statistical model with the purely learning-based approach. Through simulations, we
compare the model-aided learning moving target detector with model-based GLRT and
purely learning-based detector to demonstrate its effectiveness.
Keywords: Moving target detection (MTD); Multiple-input multiple-output (MIMO) radar;
Model-based detector; Purely learning-based detector; Training data; Model-aided learning
moving target detector

Jiacheng Lai, Zhen Zhang, Bin Zeng, Lei Wang,


GTFKAN: A Novel Microbe-drug Association Prediction Model Based on Graph Transformer
and Fourier Kolmogorov-Arnold Networks,
Journal of Molecular Biology,
Volume 437, Issue 17,
2025,
169201,
ISSN 0022-2836,
[Link]
([Link]
Abstract: Microbes have been shown to be closely related to human health. In recent years,
lots of computational methods for predicting microbial-drug association have been
proposed. In this manuscript, we introduced a novel predictive model, called GTFKAN, to
identify potential microbe-drug associations by combining Graph Transformation Networks
(GTN) with Fourier Kolmogorov-Arnold Networks (FKAN). In GTFKAN, we would first
compute the Gaussian kernel and functional similarity of microbes and drugs respectively,
and then adopt random walk and restart (RWR) methods to enhance these similar features
to construct a new microbe-drug heterogeneous network HN. At the same time, we would
further calculate the cosine similarity of microbes and diseases to construct another
microbe-drug heterogeneous network LDIM. Next, we would input HN into GTN to derive
the location and structural features of microorganisms and drugs, and input LDIM into FKAN
to extract the hidden higher-order features of microorganisms and drugs, respectively.
Finally, we would integrate these two features extracted by GTN and FKAN and feed the
integrated features into the MLP classifier to infer potential microbial-drug associations.
Moreover, to evaluate the performance of GTFKAN, we compared it with state-of-the-art
methods based on well-known public datasets, and the experimental results show that
GTFKAN can achieve satisfactory predictive performance. In addition, the results of ablation
experiments and case studies also demonstrated the superiority of GTFKAN, which means
that GTFKAN may be a useful microbial-drug association prediction tool in the future.
Keywords: microbe-drug association; Graph Transformer Network; Fourier Kolmogorov-
Arnold Network; the Gaussian kernel and functional similarities; prediction model

Sumanta Das, Bhagyasree Chatterjee, Malini Roy Choudhury, Suman Dutta, Bhabani Prasad
Mondal, Amit Awasthi,
Synthetic aperture radar for a changing planet: A 25-year global synthesis in hazard
assessment, urban development, and ecological applications,
Ecological Informatics,
Volume 92,
2025,
103477,
ISSN 1574-9541,
[Link]
([Link]
Abstract: The increasing frequency and intensity of natural disasters, rapid urbanization, and
accelerating ecological degradation underscore the urgent need for robust monitoring tools.
Over the past two and a half decades, Synthetic Aperture Radar (SAR) technology has
transformed the landscape of remote sensing, offering unique capabilities for all-weather,
day-and-night imaging. SAR applications have expanded dramatically into critical domains
such as hazard assessment, urban development, and ecological management. Despite
significant progress, fragmented knowledge, uneven adoption across geographies, and
several technical, methodological, and application-specific challenges continue to constrain
its full potential. This systematic review synthesizes the global research trends, technological
evolution, application domains, and persisting bottlenecks in SAR utilization across three key
thematic areas: hazard assessment, urban development, and ecological management over
the past 25 years (2000–2024), supported by a comprehensive bibliometric analysis of
11,201 peer-reviewed publications indexed in the Scopus database. Unlike previous reviews
that often focus on narrow applications or specific sensor types, this review offers a holistic,
cross-sectoral, and longitudinal perspective, identifying emerging trends, underexplored
geographies, and growing intersections with novel computational frameworks. The study
employed a systematic review protocol based on PRISMA guidelines, combined with
bibliometric techniques using VOSviewer and Bibliometrix (R-package). The articles were
retrieved using a well-defined keyword query related to SAR, hazards, urban studies, and
ecosystems. Analyses included publication trends, co-authorship networks, keyword co-
occurrence, source impact, and thematic evolution. Cluster analysis identified four major
research themes and three temporal development phases. Results indicate that publications
increased nearly tenfold from 2000 to 2024, with a peak after 2015 due to the launch of
Sentinel-1 and the rise of open data. Major contributing countries include China, USA, Italy,
Germany, and India, with strong international collaborations. Interferometric SAR (InSAR)
and Polarimetric SAR (PolSAR) dominated hazard and urban studies, while multi-temporal-
based approaches emerged in ecological monitoring. Notably, integration with AI and cloud-
based geospatial platforms remains limited (<15 % of publications). Urban development
applications, especially subsidence and infrastructure monitoring, show the fastest growth,
while ecological applications lag, indicating a critical research gap. Furthermore, this review
underscores that while SAR has made substantial strides, methodological integration,
capacity building in the Global South, and translation into policy-oriented tools remain key
challenges. The findings advocate for multi-sensor synergy, open-data initiatives, and
interdisciplinary collaborations as pathways to expand SAR's impact. Overall, this review
serves as a reference point for researchers, practitioners, and policymakers aiming to
harness SAR for sustainable development and disaster resilience in the coming decades.
Keywords: Remote sensing; Disaster resilience; Interferometric SAR; Polarimetric SAR; Multi-
temporal analysis; Radar imaging

Bo Zhu, Li Jia, Quanke Pan, Hui Zhang,


Cross-domain battery SOH and RUL estimation via Domain-Adaptive Transformer,
Energy,
Volume 341,
2025,
139288,
ISSN 0360-5442,
[Link]
([Link]
Abstract: Accurate estimation of the state of health (SOH) and remaining useful life (RUL) of
lithium-ion batteries across heterogeneous chemistries and cycling protocols remains
challenging because of significant domain shifts and complex degradation dynamics. To
overcome these issues, this study proposes a Domain-Adaptive Transformer (DAT)
framework for cross-domain battery health prognostics. The framework integrates a
protocol-independent feature representation—constructed from voltage (V), capacity (Q),
and their differential features (ΔV, ΔQ)—with a hybrid CNN–Transformer backbone
enhanced by rotary positional embedding and a lightweight Extreme Learning Machine
(ELM) module. This design enables efficient modeling of long-range temporal dependencies,
rapid fine-tuning on new domains, and high computational efficiency. Comprehensive
experiments on three transfer tasks (cross-discharge-protocol, cross-charge-protocol, and
cross-chemistry) demonstrate that the proposed model consistently outperforms state-of-
the-art baselines including BiLSTM, TCN, PINN, and SDE-BiLSTM. After fine-tuning, our model
achieves RUL RMSE = 178 cycles (R2 = 0.853) and SOH MAE = 0.138% (R2 = 0.999),
confirming its robustness and adaptability across domains. These results highlight the
potential of the framework for generalizable, data-efficient, and practical battery health
prognostics in real-world applications.
Keywords: Domain adaption; Transformer; Battery health prognostics; Remaining useful life;
State of health

Naseeb Singh, V.K. Tewari, P.K. Biswas,


Vision transformers for cotton boll segmentation: Hyperparameters optimization and
comparison with convolutional neural networks,
Industrial Crops and Products,
Volume 223,
2025,
120241,
ISSN 0926-6690,
[Link]
([Link]
Abstract: For the automation of cotton harvesting operations, precise segmentation of
cotton bolls is important. In the past, various handcrafted image processing-based
algorithms and convolutional neural networks (CNNs) have been developed for this purpose.
Handcrafted algorithms often only extract low-dimensional features, while CNNs have
limitations to capture global features due to their small receptive fields. However, in recent
times, Vision Transformers (ViTs) have proven to have the ability to capture long-range
dependencies through the self-attention mechanism, thus resulting in superior
segmentation accuracy. In this study, ViTs were utilized to segment cotton bolls, and the
impact of various hyperparameters on their efficacy was investigated. Different ViT variants
were developed using varying combinations of hyperparameters. Among all developed ViT
variants, the model with a patch size of 16, hidden dimensions of 8, 6 no. of Multi-head Self
attention (MHSA) heads, 12 transformer layers, and multilayer perceptron (MLP) dimension
of 128 outperformed the others. This optimal configuration achieved precision, recall, mean
Intersection over Union (m-IoU), and cotton-IoU values of 0.94, 0.94, 0.93, and 0.89,
respectively. The findings show that increasing hidden dimensions and the number of
attention heads increased model complexity but did not necessarily improve performance.
The cotton-IoU score was found to be higher for the best-performing ViT model (cotton-IoU
= 0.89) compared to the CNN model (cotton-IoU = 0.84). These results indicate that the ViT
model outperforms the CNN model (having a comparable number of trainable parameters)
for the segmentation of cotton bolls. Hence, ViTs can be effectively utilized for semantic
segmentation tasks in agriculture with higher segmentation performance while requiring
lower computational power. This makes ViTs a suitable technique for the automation of the
cotton harvesting process on resource-constrained devices without compromising
performance. Future work should include the use of pure transformer architectures,
incorporating advanced techniques to further optimize performance and efficiency in
various agricultural tasks.
Keywords: Cotton; Semantic segmentation; Deep learning; Vision transformers; Automated
harvesting

Dongjie Liu, Dawei Li, Hongliang Ding, Yang Cao, Kun Gao,
Beyond vision: A unified transformer with bidirectional attention for predicting driver
perceived risk from multi-modal data,
Transportation Research Part C: Emerging Technologies,
Volume 179,
2025,
105270,
ISSN 0968-090X,
[Link]
([Link]
Abstract: Modeling driver perceived risk (or subjective risk) plays a critical role in improving
driving safety, as different drivers often perceive varying levels of risk under identical
conditions, prompting adjustments in their driving behavior. Driving is a complex activity
involving multiple cognitive and perceptual processes, such as visual information, driver
feedback, vehicle dynamics, and traffic and environmental conditions. However, existing
models for subjective risk perception have yet to fully address the need for integrating multi-
modal data. To address this gap, we present a Transformer-based model aimed at processing
multimodal inputs in a unified manner to enhance the prediction of subjective risk
perception. Unlike existing methodologies that extract features specific to each modality, it
employs embedding layers to transform images, unstructured, and structured fields into
visual and text tokens. Subsequently, bi-directional multimodal attention blocks with inter-
modal and intra-modal attention mechanisms capture comprehensive representations of
traffic scene images, unstructured traffic scene descriptions, structured traffic data,
environmental statistics, and demographics. Experimental results show that the proposed
unified model achieves superior predictive performance over existing benchmarks while
maintaining reasonable interpretability. Furthermore, the model is generalizable, making it
applicable to various multi-modal prediction tasks across different transportation contexts.
Keywords: Traffic safety; Driving behavior modeling; Subjective risk perception; Multi-modal
data fusion

Lu Wang, Bailiang Sun, Chunhui Zhao, Suleman Mazhar, Tomoaki Ohtsuki, P. Takis
Mathiopoulos, Fumiyuki Adachi,
SAR image change detection based on saliency region guidance and SIFT keypoint extraction,
Pattern Recognition,
Volume 172, Part B,
2026,
112471,
ISSN 0031-3203,
[Link]
([Link]
Abstract: Synthetic Aperture Radar (SAR) can operate under all-weather, all-day conditions,
playing a crucial role in regional change detection (CD). However, due to its unique imaging
principles, SAR images contain significant speckle noise and blurred boundary and detail
features, which reduces the detection accuracy and leads to missed detection and false
detection. To address these issues, this paper proposes a SAR image CD method based on
saliency region guidance and Scale-Invariant Feature Transform (SIFT) keypoint extraction to
reduce the interference of speckle noise. First, a saliency region guidance method is
introduced to analyze the saliency of local features in SAR images, extracting potentially
changed regions and reducing the interference of speckle noise. Second, the SIFT is
employed to extract keypoints in regions significantly different from the background in the
difference map, leveraging its robustness to speckle noise. By extracting keypoints, the
approximate location and extent of the changed regions are determined. These are, then,
fused with the saliency region information, enhancing the saliency weights of pixels around
keypoints for more extraction of change regions. Finally, a Vision Transformer (ViT) detection
network is used for SAR image CD, utilizing the combined saliency information from the
original saliency map and SIFT keypoints. This approach effectively integrates SIFT’s stable
description of local features with ViT’s modeling capability for global features, improving the
model’s accuracy and robustness.
Keywords: Saliency region guidance; Scale-invariant feature transform; Change detection;
Vision transformer; Synthetic aperture radar image

Sonia Relan, Radha Krishna Rambola,


Big Bird-disentangled representation and adaptive contrast transformer for abstractive text
summarization with contrast attention,
Engineering Applications of Artificial Intelligence,
Volume 162, Part D,
2025,
112536,
ISSN 0952-1976,
[Link]
([Link]
Abstract: ive text summarization, a core task in Natural Language Processing (NLP), seeks to
generate concise summaries by forming novel sentences from long documents. Despite its
promise, challenges persist, including semantic ambiguity, factual inconsistencies, repetition,
limited length control, and handling out-of-vocabulary words. To address these issues, we
propose Big Bird Disentangled Representation and Adaptive Contrast Transformer (BB-
DRACT), a framework designed to enhance semantic fidelity and content relevance. BB-
DRACT extends Big Bird with an Enhanced Representation Learning Layer (ERLL) positioned
after the encoder's sparse attention, where semantic, content, and length features are
disentangled using a length-conditioned contrastive objective and cross-factor decorrelation.
To refine decoder representations, we introduce Self-Adaptive Contrast Attention (SACA)
within multi-head attention, injecting a contrast-consistency signal and uncertainty-aware
gating to suppress redundancy and strengthen factual grounding. For final summary
generation, a Dynamic Summary Generation Module (DSGM) replaces conventional
greedy/beam search. DSGM integrates adaptive sampling within token selection to improve
diversity and factuality, while an explicit length control mechanism embedded in decoder
state updates ensures coherent summaries with precise length adherence. Extensive
evaluations on PubMed Abstract and CNN/Daily Mail datasets show BB-DRACT consistently
outperforms state-of-the-art models. It achieves ROUGE-1 scores of 46.81 and 47.98,
ROUGE-2 of 24.18 and 23.19, ROUGE-L of 43.2 and 42.92, BLEU of 66.68 and 66.21, Accuracy
of 87.99 % and 87.97 %, and F1-scores of 56.65 and 57.02, reflecting 3–7 % improvements
over strong baselines. Qualitative analysis confirms BB-DRACT minimizes hallucinations,
enhances factual alignment, and offers robust length controllability, making it a reliable
solution for real-world text summarization.
Keywords: Abstractive text summarization; Natural language processing; Contrastive
learning; Enhanced representation learning layer; Disentanglement learning; Big bird;
Dynamic summary generation module

Hongyang Liang, Jiajun Wang, Jun Zhang, Xiaoling Wang, Shiwei Guan, Hao Yu,
Improved BOTSORT multi-object tracking algorithm for robotic rollers using feature-level
fusion of millimeter-wave radar and camera sensors,
Information Fusion,
Volume 123,
2025,
103294,
ISSN 1566-2535,
[Link]
([Link]
Abstract: The operational safety of robotic rollers is of paramount importance, particularly in
the challenging construction environment of dam construction sites. However, factors like
low-illumination and intense vehicle vibrations can critically impair obstacle tracking and
decision-making processes. To address this issue, this study proposes an improved BOTSORT
multi-object tracking algorithm using feature-level fusion of millimeter-wave radar and
camera sensors. Initially, by utilizing convolutional and PS-ROI align networks, radar and
camera data are merged into feature maps, which are then processed by the improved
BOTSORT algorithm using YOLOv8 instead of YOLOX for precise obstacle detection in low-
illumination conditions. Additionally, an unscented Kalman filter module is employed to
predict nonlinear motion of objects within the image during vibrations, while radar data
refines the target association process, improving tracking accuracy under severe vibration
conditions. A case study of a large-scale hydropower project demonstrates that the
proposed method achieves 61.7 % mAP and 76.5 % MOTA, outperforming other obstacle
detection and multi-object tracking algorithms. The proposed method improves the safety
and reliability of robotic rollers under low-illumination and severe vibration working
conditions.
Keywords: BOTSORT; Feature-level fusion; Multi-object tracking; Millimeter-wave radar and
camera; Low-illumination; Robotic roller; Severe vibration

Yehao Wang, Zijian Liu, Yingying Jin, Xiaoliang Wang, Lingyu Xu, Lei Wang, Jie Yu, Wenjuan
Dai, Jingxia Gao, Feng Zhang,
Interpreting spatiotemporal dynamics of Ulva prolifera blooms in the southern yellow sea
using an attention-enhanced transformer framework,
Environmental Pollution,
Volume 384,
2025,
126999,
ISSN 0269-7491,
[Link]
([Link]
Abstract: Harmful algal blooms dominated by Ulva prolifera have posed recurring ecological
and economic challenges in the southern Yellow Sea. To better understand and predict the
complex spatiotemporal dynamics of these blooms, we developed an enhanced
Transformer-based deep learning framework, incorporating multi-head self-attention
mechanisms. This model dynamically captures spatial dependencies, providing a
comprehensive understanding of bloom dynamics. Utilizing twelve key marine
environmental factors, we systematically explored all possible feature combinations to
determine the optimal predictive subset. Experimental results demonstrated superior
predictive performance of the model (MAE: 0.0213, MSE: 0.0016, R2: 0.9923) compared to
conventional deep learning models and recent spatiotemporal deep learning models.
Training dynamics revealed efficient convergence, especially with comprehensive
environmental information. Spatial attention analysis revealed that offshore regions
consistently received higher attention, indicating their critical role as informative and
generalizable environmental references. Furthermore, exhaustive feature attribution
experiments identified an optimal combination of eight environmental factors—including
temperature, salinity, current velocity, precipitation, wind direction, dissolved iron,
phosphate, and silicate—were found to significantly enhance prediction accuracy. This study
highlights the capability of attention-enhanced Transformer models for interpretable and
precise ecological forecasting, providing valuable insights for targeted mitigation and
management of U. prolifera blooms.
Keywords: Harmful algal blooms; U. prolifera; Deep learning; Spatiotemporal dynamics;
Environmental factors; Southern yellow sea

Roohollah Enayati, Reza Ravanmehr, Vahe Aghazarian,


A transformer-based siamese network using a self-attention mechanism for change
detection in remote sensing data,
Signal Processing: Image Communication,
Volume 138,
2025,
117379,
ISSN 0923-5965,
[Link]
([Link]
Abstract: Efficient change detection in remote sensing data remains a critical challenge,
necessitating a comprehensive solution that addresses diverse data formats and
applications. Existing methods often specialize in specific complications, lacking a unified
approach for various data types and resolutions. In this study, we propose an innovative
methodology that integrates a Siamese network with a Transformer model using a self-
attention mechanism, providing a versatile solution for change detection. Our approach
introduces a Pairwise Learning Task and Feature Fusion techniques, leveraging Convolutional
and Transformer Layers to enhance precision. The Siamese network, developed for similarity
determination, is augmented with self-attention mechanisms, enabling the capture of
intricate feature relationships across sequences. Our method not only excels in binary
change mapping but also in feature type identification, adding valuable insights to the
remote sensing domain. Experimental results on Sentinel-2, QuickBird, and TerraSAR-X
datasets demonstrate that our approach achieves overall accuracies of 99.23 %, 98.96 %,
and 99.08 %, respectively, and F1-scores of 92.16 %, 95.74 %, and 96.21 %, significantly
outperforming existing state-of-the-art methods by margins of up to 4.3 % in accuracy and
7.7 % in F1-score. These results highlight the proposed method’s adaptability, precision, and
robustness. The methodology, focusing on efficiency and accuracy, presents a significant
advancement in remote sensing change detection, with promising applications in
environmental monitoring, urban planning, and disaster management.
Keywords: Remote sensing; Change detection; Siamese network; CNN; Transformer; Feature
fusion; Pairwise learning task
Ming Xie, Lei Xie, Ying Li, Bing Han,
Oil species identification based on fluorescence excitation-emission matrix and transformer-
based deep learning,
Spectrochimica Acta Part A: Molecular and Biomolecular Spectroscopy,
Volume 302,
2023,
123059,
ISSN 1386-1425,
[Link]
([Link]
Abstract: After oil spills are found at sea, the identification on oil species can help determine
the source of leakage and form the plan of post-accident treatment. Since the fluorometric
characteristics of petroleum hydrocarbon reflect its molecular structure, the composition of
oil spills could potentially be inferred using the fluorescence spectroscopy method. The
excitation-emission matrix (EEM) includes additional fluorescence information in the
dimension of excitation wavelength, which could be useful to identify oil species. This study
proposed an oil species identification model using transformer network. The EEMs of oil
pollutants are reconstructed into sequenced patch input that consists of the fluorometric
spectra obtained under the different excitation wavelengths. The comparative experiments
show that the proposed model can reduce the incorrect predictions and achieve higher
identification accuracies than the regular convolutional neural networks that have been used
in the previous studies. According to the structure of transformer network, an ablation
experiment is also designed to evaluate the contributions of different input patches and seek
for the optimal excitation wavelengths for oil species identification. The proposed model is
expected to identify oil species, and even other fluorescent materials, based on the
fluorometric spectra collected under multiple excitation wavelengths.
Keywords: Fluorescence spectroscopy; UV-induced fluorescence; Spectral analysis;
Transformer network; Oil species identification; Oil spill

Quan Zhou, Mingwei Wen, Bin Yu, Cuijuan Lou, Mingyue Ding, Xuming Zhang,
Self-supervised transformer based non-local means despeckling of optical coherence
tomography images,
Biomedical Signal Processing and Control,
Volume 80, Part 2,
2023,
104348,
ISSN 1746-8094,
[Link]
([Link]
Abstract: Optical coherence tomography (OCT) depends on the coherence characteristics of
scattered light to reveal tissue morphology. Therefore, OCT images are inevitably corrupted
by speckle noise. The non-local means (NLM) method is a popular method for OCT image
despeckling. However, the existing NLM algorithm cannot preserve image details well or
deliver sufficient noise reduction because they calculate the similarity weights based on the
grayscale information or human-designed features of an image. This letter presents a self-
supervised transformer based NLM method for despeckling OCT images. The presented
method computes the weight of the NLM using the deep features extracted by the self-
supervised transformer and adopts the boosting strategy to realize the effective OCT image
despeckling. The experiments on two OCT image datasets demonstrate that our algorithm
performs better than other compared denoising algorithms in terms of the quantitative
metrics and visual evaluation.
Keywords: Optical coherence tomography; Non-local means; Self-supervised transformer;
Boosting strategy; OCT image despeckling

Pahati Tuxunjiang, Chencui Huang, Zhen Zhou, Wenyi Zhao, Bingyan Han, Weixiong Tan,
Jingru Wang, HuangHanjiaerbieke Kukun, Wei Zhao, Rui Xu, Ainikaerjiang Aihemaiti, Yimuran
Subi, Jingyang Zou, Chao Xie, Yifan Chang, Yunling Wang,
Prediction of NIHSS Scores and Acute Ischemic Stroke Severity Using a Cross-attention Vision
Transformer Model with Multimodal MRI,
Academic Radiology,
Volume 32, Issue 9,
2025,
Pages 5453-5467,
ISSN 1076-6332,
[Link]
([Link]
Abstract: Rationale and Objectives
This study aimed to develop and evaluate models for classifying the severity of neurological
impairment in acute ischemic stroke (AIS) patients using multimodal MRI data.
Methods
A retrospective cohort of 1227 AIS patients was collected and categorized into mild
(NIHSS<5) and moderate-to-severe (NIHSS≥5) stroke groups based on NIHSS scores. Eight
baseline models were constructed for performance comparison, including a clinical model,
radiomics models using DWI or multiple MRI sequences, and deep learning (DL) models with
varying fusion strategies (early fusion, later fusion, full cross-fusion, and DWI-centered cross-
fusion). All DL models were based on the Vision Transformer (ViT) framework. Model
performance was evaluated using metrics such as AUC and ACC, and robustness was
assessed through subgroup analyses and visualization using Grad-CAM.
Results
Among the eight models, the DL model using DWI as the primary sequence with cross-fusion
of other MRI sequences (Model 8) achieved the best performance. In the test cohort, Model
8 demonstrated an AUC of 0.914, ACC of 0.830, and high specificity (0.818) and sensitivity
(0.853). Subgroup analysis shows that model 8 is robust in most subgroups with no
significant prediction difference (p > 0.05), and the AUC value consistently exceeds 0.900. A
significant predictive difference was observed in the BMI group (p < 0.001). The results of
external validation showed that the AUC values of the model 8 in center 2 and center 3
reached 0.910 and 0.912, respectively. Visualization using Grad-CAM emphasized the infarct
core as the most critical region contributing to predictions, with consistent feature attention
across DWI, T1WI, T2WI, and FLAIR sequences, further validating the interpretability of the
model.
Conclusion
A ViT-based DL model with cross-modal fusion strategies provides a non-invasive and
efficient tool for classifying AIS severity. Its robust performance across subgroups and
interpretability make it a promising tool for personalized management and decision-making
in clinical practice.
Keywords: Acute ischemic stroke; Vision transformer; Deep learning; Prediction; Multimodal
MRI Fusion; Cross-Attention

Hao Xu, Zhenhao Zhu, Hongbing Liu, Enrico Zio, Xiaolong Qiu, Yuchen Lu, Xianqiang Qu,
Spectral dynamic aggregation transformer and fitted swing-door algorithm for wind power
monitoring,
Energy,
Volume 341,
2025,
139396,
ISSN 0360-5442,
[Link]
([Link]
Abstract: Current research on wind energy monitoring predominantly focuses on power
prediction while often overlooking the advanced warning of sudden operational anomalies.
To this end, we propose a wind power monitoring model based on a spectral dynamic
aggregation transformer integrated with a fitted swing gate algorithm. First, the integration
of spectral and dynamic aggregation blocks within the Transformer framework yields an
accurate wind power prediction model that effectively alleviates the impact of data
fluctuations. On this basis, the MI method is utilized to quantify the nonlinear relationships
between multi-source meteorological variables and wind power output. By integrating STL
for residual analysis to extract salient features, the proposed approach not only enhances
the input quality of the prediction model but also provides a physically interpretable
foundation for the early warning module. This facilitates seamless integration of prediction
and early warning at the feature level. Furthermore, the prediction results and key features
jointly drive a two-tier early warning framework: a wind power early warning system is
constructed based on the random forest algorithm and the swing door algorithm, followed
by joint calibration with the predictive model. By leveraging multi-source data, the model is
capable of detecting anomalous power changes and ramp events, thereby ensuring efficient
anomaly identification and advanced warning. Through the case study, it is demonstrated
that the proposed model can achieve wind power prediction and advanced warning
functions, thereby providing robust support for flexible grid scheduling and efficient wind
power integration.
Keywords: Wind power; Power prediction; Spectral dynamic aggregation; Dual early warning
with fitted swing-door algorithm; Transformer

Guangwei Zhang, Yuyao Cheng, Qi Xia, Jian Zhang,


Curvature envelope area based rapid identification method of bridge distributional element
stiffness using microwave interference radar,
Mechanical Systems and Signal Processing,
Volume 197,
2023,
110390,
ISSN 0888-3270,
[Link]
([Link]
Abstract: The bearing capacity of a bridge mainly relies on its own stiffness, while, at
present, it remains a challenge to identify the bridge stiffness rapidly. In this paper, a rapid
and innovative identification method of bridge distributional element stiffness (BDES) is
proposed based on microwave interference radar technology and curvature envelope area
(CEA). The main contributions of the work are as follows: (1) Taking advantage of multi-point
displacement synchronous measurement, the self-developed microwave inference radar
makes possible the acquisition of rotation influence lines (RILs) through the measured multi-
point displacement influence lines (DILs). (2) Based on the one-to-one accurate
correspondence between CEA, moment envelope area (MEA) and stiffness (“point-to-point”
matching), and using the essential relationships among CEA and rotation, the bridge stiffness
can be derived by using the DILs and the MEA calculated from the calibrated moving load
(“point-to-line” matching). (3) Due to the radar feature of synchronous observation over a
large area and referring to the proposed stiffness identification method, the BDES of the
monitored area can be further obtained (“line-to-surface” matching). The proposed bridge
identification method was applied at a laboratory experiment to identify the BDES of a steel
beam for both undamaged and damaged conditions, and the results support the feasibility
of the proposed method in BDES identification.
Keywords: Microwave interference radar; Multi-point displacement influence line; Curvature
envelope area; Rapid identification; Bridge distributional element stiffness

Yonghui Liu, Qian Li, Inhi Kim,


Enhanced trajectory reconstruction from sparse and noisy GPS data: A progressive chunked
transformer approach,
Communications in Transportation Research,
Volume 5,
2025,
100200,
ISSN 2772-4247,
[Link]
([Link]
Abstract: Trajectory reconstruction from sparse and noisy GPS data is critical for applications
such as urban mobility analysis, transportation planning, and navigation systems. However,
large sampling intervals and the typically long output sequences required to reconstruct
coherent travel trajectories significantly increase computational complexity, particularly in
the presence of noise. To address these challenges, we propose a progressive chunked
transformer (ProChunkFormer), which is a deep learning method for trajectory
reconstruction that employs self-attention mechanisms and chunked processing to balance
efficiency with accuracy. ProChunkFormer first generates intermediate trajectories at a semi-
high frequency from low-frequency sampled data, and then the remaining trajectory is
divided into manageable blocks and reconstructed parallelly in the condition of the semi-
high-frequency trajectory. By combining progressive reconstruction with chunk processing,
ProChunkFormer not only mitigates the cumulative errors commonly observed in
autoregressive models but also alleviates the rapid increase in complexity associated with
reconstructing ultralong trajectories. Specifically, our approach achieves quadratic
optimization in time and space for attention modules, with cubic time savings compared
with autoregressive decoding. A case study using an open-source taxi trajectory dataset
confirms the effectiveness of our approach. The performance of ProChunkFormer is
comparable to that of autoregressive transformers while offering better running efficiency. It
improves the accuracy, F1 score (F1), mean absolute error (MAE), and road network mean
absolute error (MAE_RN) by 23.1%, 18.6%, 22.3%, and 25.1%, respectively, for trajectories
with a long interval time of up to 240 s. Furthermore, we investigate incorporating heuristic
information to guide trajectory reconstruction for each block. The experimental results
indicate an improvement in both the overall performance and convergence speed of the
model.
Keywords: Trajectory reconstruction; Transformer; Chunked processing; Heuristic-informed;
Parallel computing

Wei Zhang, Xinyu Zhang, Junyu Dong, Xiaojiang Song, Renbo Pang,
CIDM: A comprehensive inpainting diffusion model for missing weather radar data with
knowledge guidance,
ISPRS Journal of Photogrammetry and Remote Sensing,
Volume 221,
2025,
Pages 299-309,
ISSN 0924-2716,
[Link]
([Link]
Abstract: Addressing data gaps in meteorological radar scan regions remains a significant
challenge. Existing radar data recovery methods tend to perform poorly under different
types of missing data scenarios, often due to over-smoothing. The actual scenarios
represented by radar data are complex and diverse, making it difficult to simulate missing
data. Recent developments in generative models have yielded new solutions for the problem
of missing data in complex scenarios. Here, we propose a comprehensive inpainting
diffusion model (CIDM) for weather radar data, which improves the sampling approach of
the original diffusion model. This method utilises prior knowledge from known regions to
guide the generation of missing information. The CIDM formalises domain knowledge into
generative models, treating the problem of weather radar completion as a generative task,
eliminating the need for complex data preprocessing. During the inference phase, prior
knowledge of known regions guides the process and incorporates domain knowledge
learned by the model to generate information for missing regions, thus supporting radar
data recovery in scenarios with arbitrary missing data. Experiments were conducted on
various missing data scenarios using Multi-Radar/MultiSensor System data sourced from the
National Oceanic and Atmospheric Administration, and the results were compared with
those of traditional and deep learning radar restoration methods. Compared with these
methods, the CIDM demonstrated superior recovery performance for various missing data
scenarios, particularly those with extreme amounts of missing data, in which the restoration
accuracy was improved by 5%–35%. These results indicate the significant potential of the
CIDM for quantitative applications. The proposed method showcases the capability of
generative models in creating fine-grained data for remote sensing applications.
Keywords: Weather radar data; Comprehensive inpainting; Diffusion models; Extreme
missing cases; Knowledge guidance

Wenteng Lu, Yinan Chen, Xiaoxia Huang,


mmChainPose: Geometry-aware temporal chaining for robust human pose estimation from
mmWave point clouds,
Neurocomputing,
Volume 663,
2026,
132019,
ISSN 0925-2312,
[Link]
([Link]
Abstract: Human sensing has emerged as a crucial enabling technology for applications in
smart homes, elderly monitoring, virtual reality and security systems. Wireless signal-based
approaches have gained significant attention as device-free solutions. Among these,
millimeter wave (mmWave) technology stands out due to its proven capability for high-
precision sensing, as extensively validated in prior work. Recent advances in deep learning
have expanded what is possible with mmWave sensing, enabling accurate 3D human pose
estimation (HPE) from sparse radar point clouds. In this work, we introduce mmChainPose, a
novel transformer-based framework designed to reconstruct accurate 3D human pose
directly from sparse mmWave point clouds. mmChainPose focuses on decoding spatio-
temporal dependencies across mmWave point cloud sequences. It consists of two core
modules: a Spatio-Temporal Dynamic Graph Convolution Network (STDGN) that extracts
discriminative per-frame features while modeling structural relationships of consecutive
frames, and a Geometry-Aware Temporal Chain Transformer (GeoChainFormer) that
propagates enriched geometric context across frames via a chain-attention mechanism,
ensuring robust pose estimation even with incomplete data. Comprehensive experiments on
two benchmarks demonstrate state-of-the-art performance of mmChainPose, achieving 2.70
cm MPJPE / 0.8706 OKS on MARS and 6.58 cm MPJPE / 0.6320 OKS on mmBody.
Keywords: Millimeter wave sensing; Human pose estimation; Point cloud reconstruction;
Spatio-temporal model; Transformer

Qinglu-Ma, Jirui-Li, Zheng-Zou, Saleem Ullah,


Vehicle type classification based on acoustic signals using LT-MFCC,
Applied Acoustics,
Volume 239,
2025,
110836,
ISSN 0003-682X,
[Link]
([Link]
Abstract: Automatic Vehicle Classification (AVC) is an essential and meaningful detecting
technology in the Intelligent Transportation System (ITS). It can judge the traffic condition
according to the different vehicle noise the receiver acquires. We proposed an enhanced
feature extraction method based on Mel Frequency Cepstral Coefficients (MFCC) for vehicle
type classification using acoustic signals. To improve classification accuracy under strong
interference noise, an improved fusion algorithm integrating Teager Energy Operator (TEO)
and Linear Discriminant Analysis (LDA) was proposed to enhance the classification accuracy
of moving vehicles under strong interference noise, utilizing Logarithmic Teager-Mel
Frequency Cepstral Coefficients (LT-MFCC). The classification results from the MFCC and LT-
MFCC methods under three different intensities of interference noise were analysed in
experiments. The results showed that LT-MFCC consistently outperformed MFCC, with
classification accuracies of 95.68 %, 96.51 %, and 97.80 % compared to MFCC’s 76.78 %,
82.40 %, and 87.07 %. At the same time, further analysis using NOISEX-92 under various
Signal-to-Noise Ratio (SNR) confirmed LT-MFCC’s robustness, maintaining a 97 % accuracy
even at 1 dB SNR, approximately 32 % higher than MFCC. The experimental results
demonstrate that the proposed method can obtain more robust acoustic features and adapt
to more complicated traffic scenarios, achieving a superior classification performance.
Keywords: Automatic vehicle classification; MFCC; Ramp confluence area; Teager energy
operator; Linear discriminant analysis

Pengfei Jia, Helmi Zulhaidi Mohd Shafri, Shengrui Yu, Zhi Zheng, Shiqing You, Abdul Rashid
Mohamed Shariff,
Radar-optical fusion of Sentinel-1/2 for high-resolution NDVI reconstruction and landscape-
driven carbon flux assessment in Kuala Selangor, Malaysia (2020–2024),
International Journal of Applied Earth Observation and Geoinformation,
Volume 145,
2025,
104966,
ISSN 1569-8432,
[Link]
([Link]
Abstract: Reliable quantification of carbon fluxes in humid tropical regions is constrained by
persistent cloud cover, heterogeneous land mosaics, and the limited resolution of existing
products. To address these challenges, this study developed a Cloud-Resilient Fusion
Network (CRFNet) that integrates Sentinel-1 SAR backscatter with cloud-screened Sentinel-2
NDVI using a CNN–BiLSTM–multi-head attention architecture. The framework reconstructed
10 m NDVI time series in Kuala Selangor, Malaysia (2020–2024), achieving annual R2 above
0.82 and RMSE below 0.12, thereby improving temporal continuity under heavy cloud–
rainfall interference. The reconstructed NDVI was used to drive a light-use-efficiency model
for net ecosystem productivity (NEP) estimation, supported by temperature-based
heterotrophic respiration. Results showed a 7.4 % decline in mean annual NEP across five
years, with degraded mangroves and sloping croplands emerging as hotspots of sink-to-
source transitions. Landscape analysis revealed strong structure–function coupling: stable
forests and mangroves were characterized by large cohesive patches with largest patch index
values above 40 % and edge density below 20 m ha-1, while croplands and degraded slopes
exhibited higher patch numbers, reduced patch dominance, and greater edge complexity,
which increased carbon source risk. By linking fine-scale NDVI reconstruction with process-
based carbon modeling and landscape metrics, this study provides a transferable workflow
for high-resolution carbon flux monitoring and a robust scientific basis for carbon budget
assessment, ecosystem management, and carbon-neutrality planning in tropical monsoon
regions.
Keywords: Cloud-resilient fusion network (CRFNet); NDVI reconstruction; Radar–optical
fusion; Net ecosystem productivity; Landscape metrics

Dawei Gao, Yongsheng Zhu, Ke Yan, C. Guedes Soares,


Deep learning–based framework for regional risk assessment in a multi–ship encounter
situation based on the transformer network,
Reliability Engineering & System Safety,
Volume 241,
2024,
109636,
ISSN 0951-8320,
[Link]
([Link]
Abstract: A method based on the predictable Transformer network associated with a
clustering method is introduced to build a framework for the regional collision risk
assessment, which is an alternative to the traditional methods that have two problems: 1)
The indicators are calculated based on the current navigation status of ships, not considering
the dynamic characteristics and the variability of the ship's trajectory, which makes the
calculated indicators inaccurate not allowing an accurate risk assessment. 2) Many deep
learning–based algorithms used in ship trajectory prediction are not easy to be trained, as
the model structure makes the features unable to be processed in parallel. First, the ships
with potential collision risk are clustered by the Density–Based Spatial Clustering of
Applications with Noise (DBSCAN) algorithm to divide the hotspots. Then, the possible
locations of ships in the near future are calculated by a multi–step prediction model, i.e., the
designed Transformer network. Finally, the ship pairs' collision risk and the regional collision
risk are evaluated based on the predicted results. Based on the AIS data from the Yangtze
River, the effectiveness of the proposed framework is verified through regional risk
assessment for 41 moments and specific interpretation for three moments.
Keywords: Collision risk assessment; Regional risk; Trajectory prediction; Multi–ship
encounter; Transformer network

Ke Wang, Bingyang Zhu, Banteng Liu, Jingyao Liang, Tan Lv, Jianfeng Wu,
Automatic sleep staging based on single-channel ballistocardiogram signals and multiple
scales temporal feature analysis,
Engineering Applications of Artificial Intelligence,
Volume 156, Part B,
2025,
111249,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Traditional sleep staging studies often rely on contact sensors for signal data
acquisition, which may compromise the integrity of sleep data. Our research presents an
automatic sleep staging method utilizing non-contact single-channel Ballistocardiogram
(BCG) signals to address this limitation. This study proposes a multiple scale(multi-scale)
time window feature extraction method based on heart rate variability (HRV) and respiratory
rate variability (RRV) to establish a more precise correlation between BCG signals and sleep
stages. Additionally, we present an advanced two-layer stacked ensemble model designed to
enhance the accuracy and robustness of sleep staging. The innovative sleep staging model is
subjected to a rigorous 5-fold cross-validation on 10 diverse recordings, encompassing
10,614 sleep segments. Experimental results indicate that the proposed feature extraction
method constructs a Top-50 feature set with an average weight of 0.2082, representing a
195.8 % improvement compared to the 0.6158 of the traditional 30s feature set.
Additionally, the proposed classification model achieves an accuracy of 89.15 %,
outperforming traditional sleep staging models by 2 percentage points. It provides valuable
references for research based on HRV, RRV, and other biological information features. This
advancement enhances sleep monitoring, especially for home and mobile healthcare,
offering a more user-friendly experience and practical medical tools. "The BCG signal sleep
data source code is available at [[Link]
Sleepstaging.]."
Keywords: Sleep staging; Ballistocardiogram signal; Heart rate variability; Respiratory rate
variability; Multiple scale time window; Two-layer stacked model

Xiuxin Xia, Yatao Cheng, Zhuo Zhang, Zhijie Hua, Qun Wang, Yan Shi, Hong Men,
Advancing research on odor-induced sweetness enhancement: A EEG local-global fusion
transformer network for sweetness quantification combined with EEG technology,
Food Chemistry,
Volume 463, Part 4,
2025,
141533,
ISSN 0308-8146,
[Link]
([Link]
Abstract: Reducing sugar intake is crucial for health, and odor sweetening enhances food
enjoyment and quality perception. Current research relies on subjective manual sensory
evaluations, which are poorly reproducible. Traditional methods also fail to capture dynamic
neural responses to odor-induced sweetness. We propose an electroencephalogram local-
global fusion transformer network (EEG-LGFNet) model to decode this impact objectively.
Electroencephalogram data were collected from 16 subjects under different odor and
sucrose stimuli. The model captures complex neural signals by integrating local and global
feature extraction mechanisms. Its performance was validated across three-time windows,
demonstrating efficacy over various temporal ranges. Analysis of the coefficient of
determination across brain regions confirmed the importance of the frontal, central, and
parietal areas of sweetness perception. The EEG-LGFNet model excelled in quantifying odor-
enhanced sweetness, significantly outperforming state-of-the-art models. This research
offers new insights into odor sweetening, with applications in food development,
personalized nutrition, and neuroscience.
Keywords: Reducing sugar; Odor sweetening; Sensory prediction; Neural responses; Artificial
intelligence

Niantang Liu, Qunshan Zhao, Richard Williams, Si-Bo Duan, Yingwei Sun, Brian Barrett,
Ensemble modelling based on transfer learning for enhancing crop mapping through
synergistic integration of InSAR coherence and multispectral satellite data,
Computers and Electronics in Agriculture,
Volume 242,
2026,
111332,
ISSN 0168-1699,
[Link]
([Link]
Abstract: Recent advancements in remote sensing have enabled the integration of multi-
temporal and multi-modal data for agricultural applications, such as crop mapping. This
study proposes an innovative framework that explores the synergistic use of multi-temporal
Sentinel-1 Interferometric Synthetic Aperture Radar (InSAR) coherence alongside Sentinel-2
and RapidEye multispectral data to enhance crop mapping in smallholder croplands in Bei’an
county, China. Various deep learning models were evaluated, including the 3-Dimensional U-
Net (3D U-Net), Transformer, Attention-based Long Short-Term Memory (AtLSTM), and a
baseline machine learning Random Forest (RF) model, focusing on their transfer learning
capabilities in complex intercropping patterns. Our new architecture, Transformer-AtLSTM-
RF, uses ensemble learning to fuse features from different classifiers with a rule-based
strategy, facilitating multi-source feature fusion for enhanced crop classification
performance. Fine-tuning with region-specific data yielded high overall accuracy (OA), mean
F1 score, and mean intersection over union (mIoU) for two test sites: site A (OA: 96.2%,
mean F1: 92.7%, mIoU: 86.9%) and site B (OA: 90.7%, mean F1: 88.6%, mIoU: 79.7%).
Additionally, we assessed feature importance by visualizing critical temporal features during
the model inference process to improve an in-depth understanding of underlying patterns in
the feature learning process. Our findings demonstrate the effectiveness of integrating time
series SAR-derived and optical data with advanced models for mapping intercropping
systems.
Keywords: Crop mapping; InSAR; Coherence; Deep learning; Transfer learning; Feature
importance

Yuanbo Li, Wenwu Zhang, Songtao Lv, Jing Yu, Dongdong Ge, Jiawei Guo, Lin Li,
YOLOv11-CAFM model in ground penetrating radar image for pavement distress detection
and optimization study,
Construction and Building Materials,
Volume 485,
2025,
141907,
ISSN 0950-0618,
[Link]
([Link]
Abstract: Ground Penetrating Radar (GPR) is an effective technology for detecting
underground structures and has been widely utilized for monitoring road damage.
Traditional B-scan-based one-dimensional images often fail to preserve continuous spatial
information, thus inadequately reflecting the nuances of damage patterns. This paper
investigates the accurate recognition of hidden internal road damage using 3D-sliced C-scan
images. While YOLO is one of the most effective and rapid neural network models for object
detection, it still suffers from low recognition accuracy and a high rate of missed detections.
To address these issues, this study proposes an improved Convolution and Attention Fusion
Module (CAFM) fusion network model for YOLOv11, which combines the CAFM with the
C2PSA global-local feature extraction mechanism to significantly enhance the recognition
performance for complex road damage. Experimental comparisons between the YOLOv11m-
CAFM and the YOLOv11 model reveal that the combined metrics for the small (n/s) and large
(l/x) models are lower than those for the medium model (m). The YOLOv11m-CAFM
demonstrates strong performance in key metrics such as precision, recall, mAP50, and
mAP50:95, achieving values of 0.840, 0.850, 0.881, and 0.584, respectively, representing
improvements of 0 %, 4.6 %, 1.8 %, and 2.0 % over the baseline model. The confidence level
in detecting standardized targets (e.g., pipelines and well covers) exceeds 0.89. Borehole
validation confirms that the model's localization error is less than 0.15 m, and the detection
frame rate reaches 71 FPS, satisfying the requirements for rapid road assessment. This study
offers a novel method for the intelligent interpretation of GPR images, considering both
detection accuracy and real-time performance, which holds significant engineering
applications in identifying hidden road damages.
Keywords: Pavement disease detection; Ground penetrating radar; Neural network; Object
detection; Deep learning algorithm

Feng Liu, Kunde Yang, Guohui Li, Zipeng Li, Guangyu Gong,
Feature extraction of underwater acoustic signal based on variational mode decomposition
and fractional-order chaotic oscillator,
Measurement,
Volume 253, Part B,
2025,
117542,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Underwater acoustic signal (UAS) feature extraction plays a crucial role in marine
target recognition. As human activities and research in marine environments continue to
increase, the task of extracting meaningful features from UAS has become more challenging.
To address this issue, this paper proposed an enhanced feature extraction method that
improves the accuracy of marine target recognition. The method combines an improved
variational mode decomposition, fractional order Duffing (FOD) oscillator, improved
multiscale amplitude-aware permutation entropy (IMAAPE), and an improved least squares
support vector machine by turbulent flow of water-based optimization (TFWO-LSSVM). By
decomposing UAS into intrinsic mode functions (IMFs), the method selects the IMF with the
smallest IMAAPE value for feature extraction. The FOD oscillator is then used to detect the
line spectrum of the selected IMF and determine the frequency range. The frequency
corresponding to the smallest IMAAPE value is identified as the line spectrum frequency.
Finally, the extracted features are input into the TFWO-LSSVM for recognition, achieving a
recognition rate of 97.4%. This method demonstrates high accuracy, and the success rate of
actual marine biological signal extraction reaches 95.1%. The proposed method offers
significant advancements in marine target recognition, with potential applications in
environmental monitoring and marine biology research.
Keywords: Underwater acoustic signal; Mode decomposition; fractional order Duffing;
Feature extraction; Entropy

Kurnia Paranita Kartika Riyanti, Eko Setijadi, Gamantyo Hendrantoro, Nurhayati,


CPW feed and stub optimization for Vivaldi antennas in broadband ground-penetrating radar
applications,
AEU - International Journal of Electronics and Communications,
2026,
156219,
ISSN 1434-8411,
[Link]
([Link]
Abstract: Vivaldi antennas used in Ground Penetrating Radar (GPR) systems often experience
performance degradation at low frequencies due to inefficient radiation and impedance
mismatch between the feedline and the tapered slot structure. To address these limitations,
this paper presents a compact Vivaldi antenna employing a coplanar waveguide (CPW)
feedline integrated with a single-stub matching technique. The stub position is analytically
optimized to improve impedance matching over a broad frequency range. The CPW
configuration simplifies the antenna structure by placing both the feedline and ground plane
on the same substrate layer, resulting in a compact layout and facilitating broadband
operation. The proposed antenna is fabricated on an FR-4 substrate with a thickness of
1.6mm and a relative permittivity of ϵr=4.3, with overall dimensions of 80×80×1.635 mm3.
The measured results demonstrate broadband operation from 1.13 to 5 GHz. The antenna
achieves a peak gain of 10.89 dBi in simulation and 10.36 dBi in measurement at 1.9 GHz,
with a measured radiation efficiency of approximately 90.16% at the resonant frequency.
The corresponding simulated and measured S11 values of −61.21 dB and −39.13 dB indicate
effective impedance matching at resonance, with minor discrepancies attributed to
fabrication tolerances and measurement conditions. To assess practical feasibility, a
preliminary sandbox-based GPR experiment was conducted using a pair of identical
antennas in a bistatic configuration. The resulting A-scan response shows a distinct reflection
corresponding to a buried metallic target at an estimated depth of 37.7 cm, which agrees
well with the actual burial depth. These results indicate that the proposed antenna can
support broadband GPR sensing within the investigated frequency range, while further
system-level and field validations are recommended.
Keywords: Broadband ground penetrating radar; Coplanar waveguide; Single stub; Vivaldi
antenna

Yiwen Chen, Yuan Zhuang, Binliang Wang, Jianzhu Huai,


4D RadarPR: Context-Aware 4D Radar Place Recognition in harsh scenarios,
ISPRS Journal of Photogrammetry and Remote Sensing,
Volume 221,
2025,
Pages 210-223,
ISSN 0924-2716,
[Link]
([Link]
Abstract: Place recognition is a fundamental technology for uncrewed systems such as
robots and autonomous vehicles, enabling tasks like global localization and simultaneous
localization and mapping (SLAM). Existing Place recognition technologies based on vision or
LiDAR have made significant progress, but these sensors may degrade or fail in adverse
conditions. 4D millimeter-wave radar offers strong resistance to particles like smoke, fog,
rain, and snow, making it a promising option for robust scene perception and localization.
Therefore, we explore the characteristics of 4D radar point clouds and propose a novel
Context-Aware 4D Radar Place Recognition (4D RadarPR) method for adverse scenarios.
Specifically, we first adopt a point-based feature extraction (PFE) module to capture raw
point cloud information. On top of PFE, we propose a multi-scale context information fusion
(MCIF) module to achieve local feature extraction at different scales and adaptive fusion. To
capture global spatial relationships and integrate contextual information, the MCIF module
introduces a fusion block based on multi-head cross-attention to combine point-wise
features with local spatial features. Additionally, we explore the role of Radar Cross Section
(RCS) information in enhancing the discriminability of descriptors and propose a local RCS
relation-guided attention network to enhance local features before generating the global
descriptor. Extensive experiments are conducted on in-house datasets and public datasets,
covering various scenarios and including both long-range and short-range radar data. We
compared the proposed method with several state-of-the-art approaches, including
BevPlace++, LSP-Net, and Transloc4D, and achieved the best overall performance. Notably,
on long-range radar data, our method achieved an average Recall@1 of 89.9%,
outperforming the second-best method by 1.9%. Furthermore, our method demonstrated
acceptable generalization ability across diverse scenarios, showcasing its robustness.
Keywords: Place recognition; 4D millimeter-wave radar; Harsh scenarios; Multi-scale context
fusion; Long- and short-range radars; Relocalization

Jinuk Kwon, Jihun Hwang, Jee Eun Sung, Chang-Hwan Im,


Speech synthesis from three-axis accelerometer signals using conformer-based deep neural
network,
Computers in Biology and Medicine,
Volume 182,
2024,
109090,
ISSN 0010-4825,
[Link]
([Link]
Abstract: Silent speech interfaces (SSIs) have emerged as innovative non-acoustic
communication methods, and our previous study demonstrated the significant potential of
three-axis accelerometer-based SSIs to identify silently spoken words with high classification
accuracy. The developed accelerometer-based SSI with only four accelerometers and a small
training dataset outperformed a conventional surface electromyography (sEMG)-based SSI.
In this study, motivated by the promising initial results, we investigated the feasibility of
synthesizing spoken speech from three-axis accelerometer signals. This exploration aimed to
assess the potential of accelerometer-based SSIs for practical silent communication
applications. Nineteen healthy individuals participated in our experiments. Five
accelerometers were attached to the face to acquire speech-related facial movements while
the participants read 270 Korean sentences aloud. For the speech synthesis, we used a
convolution-augmented Transformer (Conformer)-based deep neural network model to
convert the accelerometer signals into a Mel spectrogram, from which an audio waveform
was synthesized using HiFi-GAN. To evaluate the quality of the generated Mel spectrograms,
ten-fold cross-validation was performed, and the Mel cepstral distortion (MCD) was chosen
as the evaluation metric. As a result, an average MCD of 5.03 ± 0.65 was achieved using four
optimized accelerometers based on our previous study. Furthermore, the quality of
generated Mel spectrograms was significantly enhanced by adding one more accelerometer
attached under the chin, achieving an average MCD of 4.86 ± 0.65 (p < 0.001, Wilcoxon
signed-rank test). Although an objective comparison is difficult, these results surpass those
obtained using conventional SSIs based on sEMG, electromagnetic articulography, and
electropalatography with the fewest sensors and a similar or smaller number of sentences to
train the model. Our proposed approach will contribute to the widespread adoption of
accelerometer-based SSIs, leveraging the advantages of accelerometers like low power
consumption, invulnerability to physiological artifacts, and high portability.
Keywords: Spoken speech synthesis; Three-axis accelerometer; Deep neural network;
Conformer; Silent speech interface

Zhijun Xiao, Maarten De Vos, Christos Chatzichristos, Kejun Dong, Yunyi Jiang, Zhongyu
Wang, Yuwei Zhang, Fei Ding, Chenxi Yang, Jianqing Li, Chengyu Liu,
Noncontact capacitive coupling ECG-Derived respiratory signals using the conformer based
time–frequency domain generative adversarial network,
Expert Systems with Applications,
Volume 289,
2025,
128360,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Respiratory monitoring and analysis is a key method for detecting sleep-related
diseases. This paper presents a novel approach for respiratory monitoring that utilizes
noncontact capacitive coupling electrocardiograms-derived respiration (cEDR) method. We
propose a Time-Frequency Domain Generative Adversarial Network (TF-GAN) method for
generating respiratory signals, and successfully apply it to capacitive coupling
electrocardiograms(cECG). First, we analyze the mechanism of respiratory coupling with
cECG and verify the feasibility of the theory. Then, using the developed device, we collect
cECG data from 16 subjects during the night and simultaneously collect respiratory signals as
a reference, to validate the feasibility of our approach. Next, we convert the collected cECG
data into time–frequency domain features using Short-Time Fourier Transform (STFT) and
input these features into a Convolution-augmented transformer (Conformer) based
Generative Adversarial Network(GAN) to generate the cEDR. The network architecture
integrates self-attention mechanisms and time–frequency domain enhancement
mechanisms to effectively extract the respiratory energy components. Finally, we compare
the generated respiratory signals with the reference signals. The experimental results show
that the generated respiratory signals exhibit a high correlation with the reference signals.
Specifically, 86.3 % of the signals have a absolute waveform correlation coefficient greater
than 0.5, indicating good reproduction of real breathing waveforms. Our proposed model
demonstrates superior performance in respiratory signal extraction, achieving a low Root
Mean Square Error (RMSE) of 0.96 ± 0.12 bpm and a high agreement rate of 94.83 % ±
0.30 % within the Bland–Altman limits. Additionally, the model maintains an effective
respiratory segment ratio of 67.56 % ± 8.89 %, even under poor cECG signal conditions,
showcasing its robustness and reliability.
Keywords: Bedside Respiratory Monitoring; Capacitive coupling electrocardiogram;
Generative Adversarial Network; Time-Frequency Domain Enhancement

Jay Shen Teoh, Chee Keong Tan, Vishnu Monn Baskaran, Wai Peng Wong,
Spiking the transformer: NeuViT with single-step attention for efficient high-resolution UAV
vision,
Results in Engineering,
Volume 29,
2026,
108955,
ISSN 2590-1230,
[Link]
([Link]
Abstract: High-resolution aerial imagery from unmanned aerial vehicles (UAVs) provides the
fine-grained spatial detail required for reliable perception during field missions, yet onboard
processing of such data is constrained by high computational, latency, and energy demands.
We present Neuromorphic Vision Transformer (NeuViT), a spiking vision transformer tailored
for real-time high-resolution object detection under the strict power and latency constraints
of UAV platforms. NeuViT integrates a single-step stateless spiking paradigm that removes
temporal dependencies and membrane state tracking, and a spike-driven multi-head self-
attention mechanism that replaces dense attention with binary spike coincidences,
eliminating all multiply-accumulate operations. Evaluated on the VisDrone2019 dataset at
1500 × 2000 resolution, NeuViT achieves 88.8 % lower energy consumption and 31 % lower
latency than Swin Transformer, while retaining 79.4 % of its detection accuracy. NeuViT
remains fully operational within a 10 W power budget and sustains more than 22 frames per
second, whereas baseline vision transformers, efficient convolutional neural networks, and
spiking neural networks fail to meet UAV deployment requirements. Through per-class
precision analysis, confusion matrix characterization, and spike distribution correlation, we
identify that accuracy loss concentrates on small, low-contrast objects where spike
thresholds suppress weak activations, informing an enhancement roadmap. Energy
projections across 45 nm, 16 nm, and 7 nm technology nodes confirm that NeuViT maintains
sub-10 W operation regardless of fabrication process. These results demonstrate that
neuromorphic principles, when pragmatically adapted to silicon constraints, can enable
high-resolution vision capabilities in resource-constrained aerial platforms. The code and
models are publicly available at [Link]
Keywords: Spiking neural networks; Neuromorphic computing; Unmanned aerial vehicles;
Energy-efficient neural network; Edge AI; Real-time object detection; High-resolution image
processing

Junyu Zhou, Yuting Fu, Sihan Dong, Yuemeng Liu, Han Sun, Yanmin Li, Xunbin Wei,
Multi-model deep learning on photoacoustic flow cytometry signals for real-time melanoma
circulating tumor cells detection and biological characterization,
Expert Systems with Applications,
Volume 308,
2026,
131123,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Background
Melanoma remains one of the most aggressive forms of skin cancer, with early detection
being critical for patient outcomes. This study introduces a novel photoacoustic
fingerprinting approach integrated with advanced machine learning for non-invasive
melanoma detection and circulating tumor cell (CTC) identification.
Methods
We developed a three-tiered photoacoustic fingerprinting system combining photoacoustic
flow cytometry (PAFC) with machine learning algorithms. A uniform PAFC configuration
employed a 532 nm laser for vascular localization followed by a 1064 nm laser targeting
melanin-rich melanoma cells. The study included 50 melanoma patients and healthy
controls, analyzing spectral features across multiple wavelengths. We compared self-
supervised learning architectures (PAFCMamba vs. Transformer) and developed a hybrid
CNN-Transformer model for simultaneous CTC identification, staging, and metastatic
dissemination prediction.
Results
The photoacoustic fingerprinting system achieved exceptional diagnostic discrimination
between melanoma patients and healthy controls. Random Forest achieved area under
curve (AUC) values up to 0.97. The PAFCMamba model outperformed the Transformer
architecture (accuracy 0.75 vs. 0.62, AUC 0.785 vs. 0.730). The hybrid CNN-Transformer
architecture achieved exceptional performance with AUCs up to 0.974 and precision > 94 %
in simultaneous CTC detection, staging, and metastasis prediction. High-immunogenicity
genes including MLANA, GPR89B/A, and PIGF were identified as potential immunotherapy
targets, with photoacoustic signatures serving as non-invasive surrogate biomarkers for
underlying molecular characteristics.
Conclusions
This study establishes photoacoustic fingerprinting as a clinically viable, non-invasive
approach for melanoma detection and CTC monitoring, achieving performance comparable
to conventional methods. The integration of machine learning with photoacoustic
biomarkers provides a scalable framework with interpretable features that facilitates clinical
translation.
Keywords: Photoacoustic fingerprinting; Circulating tumor cells; Deep learning; Non-invasive
diagnosis; Biomarkers; Precision medicine

Gevork B. Gharehpetian, Hossein Karami, Seyed-Alireza Ahmadi,


Chapter 8 - Syntactic aperture radar imaging,
Editor(s): Gevork B. Gharehpetian, Hossein Karami,
Power Transformer Online Monitoring Using Electromagnetic Waves,
Academic Press,
2023,
Pages 179-246,
ISBN 9780128228012,
[Link]
([Link]
Abstract: The main drawback of electromagnetic-based power transformer monitoring
methods described in previous chapters is the need for different defected situations data. In
other words, the output signals of different locations and extents should be provided to be
used in classification methods. This chapter describes the syntactic aperture radar (SAR)
imaging method to detect the defect's occurrence, location, and extent without any need for
defected situation data. In the SAR imaging method, a short ultra wideband pulse is
transmitted toward the transformer. The reflected signal is received, which represents one-
dimensional information from the high-voltage (HV) winding. Then, the position of antennas
is changed along the height of the transformer winding and 1D information is gathered for
all positions. Using Kirchhoff's migration algorithm, a two-dimensional image of a
transformer HV winding is generated. Comparing sound and defected situation images of
the winding can show axial displacement or radial deformation of the winding. Another
defect in the power transformer is partial discharge (PD), which emits a signal in the ultra
high-frequency (UHF) band. Suppose detection of axial displacement using the SAR imaging
method can be applied in the UHF band. In that case, both defects can be detected using
only one set of antennas, which can be more economical. In this chapter, the feasibility of
simultaneous detection of both defects is studied using simulation by CST software. It is
shown that using the same frequency band to detect both defects leads to wrong
conclusions about winding conditions if PD occurs during the SAR imaging method
procedure. The UHF stepped-frequency imaging method and generalized likelihood ratio test
method are proposed to detect axial displacement to eliminate the PD effect on the results.
Finally, a monitoring system is designed and implemented on a real transformer winding to
simultaneously detect axial displacement or radial deformation and PD occurrence.
Keywords: Dielectric window; Generalized likelihood ratio test; Online monitoring; Partial
discharge detection; Stepped-frequency radar imaging; Ultra-wideband system

Mohammed Amin Adoul, Mansour Abed, Adel Belouchrani,


Time-frequency readability enhancement of compact support kernel-based distributions
using image post-processing: Application to instantaneous frequency estimation of M-ary
frequency shift keying signals,
Digital Signal Processing,
Volume 127,
2022,
103535,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Time-frequency distributions (TFDs) based on time-lag kernels with compact
support (KCS) have proved their high performance in terms of resolution and crossterms
suppression. However, as for all kernel-based quadratic TFDs, these distributions suffer from
spreading out signal terms. This is due to the unavoidable smoothing effects of the kernel in
the ambiguity domain. The main objective of this paper is to improve concentration,
interference rejection and so time-frequency localization of this representation class. The
latter has the advantage of being tuned using a single parameter while external windows are
no longer needed. The KCS-TFDs, referred to as KCSDs, are first optimized using a selection
of the most used objective performance measures in the literature. Important signal
features are extracted as well through analysis of time slice plots. The obtained TF diagrams
are then enhanced using a specific method that includes two-dimensional Wiener filter,
automatic binarization and morphological image processing techniques. The enhanced plots
are compared to those obtained from the original TFDs using several tests on real-life and
multicomponent frequency modulated (FM) signals including the noise effects. Moreover, a
comparative study involving a selection of the best-performing reassignment time-frequency
distributions is provided. The obtained results show a significant improvement of
concentration, time-frequency localization of the autoterms as well as interference and
noise suppression. As viable applications, the proposed approach is used first to
instantaneous frequency (IF) estimation of several synthetic and real-life M-ary frequency
shift keying (MFSK) signals. It is shown that the IF estimator from the enhanced plots
performs better than smoothed pseudo Wigner-Ville distribution (SPWVD) and reassignment
post-processing-based TFDs in terms of mainlobe width (MLW) and variance, respectively,
even at low signal-to-noise ratio (SNR). On the other hand, time-frequency characterization
of continuous wave linear frequency modulation (CW-LFM) and pulse linear FM (PLFM)
radar signals is also investigated.
Keywords: Image processing; MFSK; Polynomial CB distribution (PCBD); Separable CB
distribution (SCBD); Time-frequency enhancement; Enhanced KCSD

Yiming Chen, Zhen Zhang, Hao Li, Shuyuan Yang, Xiangyi Wang, Lei Wang,
GTDEKAN: Graph-aware transformer and enhanced Kolmogorov-Arnold Network for
microbe-drug association prediction,
Expert Systems with Applications,
Volume 285,
2025,
127968,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Microbes play a crucial role in human health, influencing immune responses and
disease progression. Predicting microbe-drug associations is essential for advancing drug
discovery and enhancing clinical treatments. However, current predictive methods face
challenges in processing features, preserving information, and modeling complex
relationships within intricate network structures. To address these limitations, we propose
GTDEKAN, a novel prediction model that combines an Graph-aware Transformer with a Dual
Cross-Attention (DCA) module. This integration effectively resolves issues related to long-
term dependencies and local structural information loss by extracting both global and local
features from the heterogeneous microbe-drug network. The DCA module, featuring
Channel and Spatial Cross-Attention, enhances feature processing, minimizing information
loss across complex structures. Additionally, we incorporate the Enhanced Kolmogorov-
Arnold Network (EKAN) to generate predictions of microbe-drug associations. EKAN
improves the model’s ability to capture relationships between nodes, avoiding catastrophic
forgetting and addressing the limitations of traditional deep learning models. This
integration significantly boosts the model’s accuracy. Extensive experiments show that
GTDEKAN consistently outperforms existing methods, offering superior predictive
capabilities. Case studies of four drugs across multiple databases further validate the
model’s effectiveness, revealing previously unknown microbe-drug associations and
providing new avenues for therapeutic development.
Keywords: Channel Cross-Attention (CCA); Spatial Cross-Attention (SCA); Dual Cross-
Attention (DCA); Heterogeneous-network; Kolmogorov-Arnold networks; Microbe-drug
association prediction

Seunghyun Kim, Seunghwan Shin, Sangwon Lee, Kaewon Choi, Yusung Kim,
Learning Visual Clue for UWB-based multi-person pose estimation,
Knowledge-Based Systems,
Volume 284,
2024,
111289,
ISSN 0950-7051,
[Link]
([Link]
Abstract: Compared to camera image-based methods, radio frequency (RF) based pose
estimation has great potential for use in situations where the field of view is obstructed. In
this paper, we present a novel RF-based Pose Estimation framework with Transformer (RPET)
that operates in a fully end-to-end fashion and uses an easy-to-install portable radar. RPET
eliminates the need for complex preprocessing and hand-crafted post-processing modules,
such as region-of-interest (RoI) cropping, non-maximum suppression (NMS), and keypoint
grouping. We also introduce a novel concept called Visual Clue (VC), which mimics a pose
feature represented in image-based methods and improves the learning performance of
multi-person pose estimation from RF signals. Our experimental results demonstrate the
effectiveness of VC and the generalizability of our model to different environmental
conditions, including changes in location and obstructed views.
Keywords: RF-based Pose Estimation; Multi-person pose estimation; End-to-end learning

Shahid Akbar, Ali Raza, Wajdi Alghamdi, Hashim Ali, Quan Zou, Ximei Luo,
Identifying protein succinylation sites using generative transformer and a two-dimensional
representation with a deep capsule network,
iScience,
Volume 28, Issue 12,
2025,
114137,
ISSN 2589-0042,
[Link]
([Link]
Abstract: Summary
Protein succinylation is a vital post-translational modification that regulates diverse cellular
processes. Accurate identification of succinylation sites is crucial for understanding protein
function and development of targeted drugs. In this study, we propose an intelligent
computational model, iSucc-SnCNs, which encodes protein sequences using the ProtGPT2-
based protein language model. Structural representations are derived from SMR and PSSM
matrices to extract SMR-HOG, SMR-DCT, and PSSM-DWT features. The BTGA+KNN algorithm
selects top-ranked features from the hybrid feature vector. Finally, a self-normalized capsule
neural network (Sn-CapsNet) is trained using a BTGA-based optimal feature set. The
proposed iSucc-SnCNs achieved an accuracy of 92.92% and an AUC of 0.96, outperforming
traditional models by 17%. The generalization of the iSucc-SnCNs model on two independent
datasets (Ind-I and Ind-II) demonstrated improved performance by approximately 13% and
2%, respectively. These results highlight iSucc-SnCNs as a robust and efficient framework for
large-scale succinylation site prediction and protein function analyses in drug discovery.
Keywords: Structural biology; Bioinformatics
Kang Ni, Pengcheng Wang, Zhizhong Zheng, Yanfei Zhong,
Complex-valued mix transformer for SAR ship detection,
ISPRS Journal of Photogrammetry and Remote Sensing,
Volume 231,
2026,
Pages 1-16,
ISSN 0924-2716,
[Link]
([Link]
Abstract: Transformer-based object detection algorithms, known for their powerful global
modeling capabilities and unified modeling framework, have now been widely applied to
SAR ship target detection. However, most existing approaches still rely solely on SAR
amplitude data, neglecting the importance of phase data in capturing fine structural details
of SAR targets. This reliance not only makes existing methods vulnerable to speckle noise,
but also fails to capture the high-frequency structural information in SAR phase data and
inadequate capacity to integrate fine-grained structural details and contextual semantic
information of SAR data. To address this limitation, this paper proposes a complex-valued
SAR target detection framework based on a “dual-branch feature learning–feature fusion–
feature selection” strategy, named Complex-valued Mix Transformer (CVMT), which
incorporates both SAR amplitude and phase information. The network incorporates a super
feature learning module and a detail enhancement module, which respectively mitigate
speckle noise in amplitude features and enhance the extraction of high-frequency phase
details. Moreover, the integration of a high-level feature fusion mechanism and a query
selection strategy significantly enhances the interaction between amplitude and phase
information, leading to more accurate ship target localization. Experiments conducted on
the OpenSARShip and FAIR-CSAR datasets validate the necessity of incorporating phase data
for SAR ship target detection. Furthermore, CVMT not only effectively suppresses
background clutter and preserves target structural integrity, but also achieves superior
detection robustness in both offshore and nearshore scenarios, outperforming other related
algorithms. The source code is available at [Link]
Keywords: SAR ship detection; Transformer; Complex-valued data; Feature fusion

Nadiane Nguekeu Metepong Lagpong, Joseph Mvogo Ngono, Auguste Vigny Noumsi
Woguia, Pierre Ele, Adrien Arnaud Kemche Ghomsi,
Automatic Detection of Flooded Areas in Polarimetric Radar Images From the Sentinel-1
Satellite,
International Journal of Applied Geospatial Research,
Volume 16, Issue 1,
2025,
,
ISSN 1947-9654,
[Link]
([Link]
Abstract: ABSTRACT
The free availability of Synthetic Aperture Radar (SAR) data from the sentinel satellite offers
a unique opportunity for developing countries. The research work focuses on the floods in
the town of Yagoua. The choice of this area is based on the multitude of floods causing
enormous damage. Existing methods, primarily based on machine learning and deep
learning algorithms, present major limitations such as sensitivity to radar noise, algorithmic
complexity, and dependency on training data. The methodology proposed here uses the
Kolmogorov algebraic method algorithm, which will be applied to the pre-processed images.
The Fuzzy C-Means algorithm will then be used to generate a change map consisting of two
output classes (water and not water). coupling these two methods gives good results and
analysis of pre- and post-flood images resulted in an average improvement of 12% compared
to state-of-the-art methods. This approach enhances rapid and reliable flood monitoring.
Keywords: Radar; Polarimetric Radar; Algebraic Method; Climate Change; Flooded Areas;
Fuzzy C-Means

Junkai Liu, Xinwei Qian, Lu Peng, Dan Lou, Yiwen Li,


TEDR: A spatiotemporal attention radar extrapolation network constrained by optical flow
and distribution correction,
Atmospheric Research,
Volume 311,
2024,
107702,
ISSN 0169-8095,
[Link]
([Link]
Abstract: In recent years, deep learning has been widely applied to meteorological radar
extrapolation due to the shortcomings of traditional optical flow methods in predicting the
genesis and dissipation of radar echoes. However, it still faces challenges in addressing issues
of clarity and overall intensity attenuation caused by uncertainty. This study implemented a
dual-path spatiotemporal attention network that integrates optical flow techniques by
employing intra-frame static attention and inter-frame dynamic attention, which could
simulate motion fields and the overall intensity distribution of radar echoes separately. Our
approach effectively resolve the issues of systematic intensity attenuation and clarity
degradation introduced by deep learning methods. Through the comparisons of key metrics
such as MSE, SSIM, CSI20, CSI30, and CSI40, the results demonstrated significant
improvements over traditional approaches, particularly in CSI30 and CSI40, where the
metrics improved by more than 35 %.
Keywords: Radar extrapolation; Convective weather forecasting; Machine learning; Optical
flow

Yanjiao Song, Linyi Li, Yun Chen, Junjie Li, Zhe Wang, Zhen Zhang, Xi Wang, Wen Zhang,
Lingkui Meng,
GCT-GF: A generative CNN-transformer for multi-modal multi-temporal gap-filling of surface
water probability,
International Journal of Applied Earth Observation and Geoinformation,
Volume 141,
2025,
104596,
ISSN 1569-8432,
[Link]
([Link]
Abstract: Spatial and temporal data gaps present a significant challenge to high-frequency
surface water mapping using satellite imagery. Utilizing observations from temporally close
periods and multi-modal sensors for gap-filling is of critical importance. However,
discontinuous pixel values inherent to conventional water maps hinder the application of
deep learning methods, which are effective and popular for relevant studies. In this study, a
novel approach, termed “gap-filling of surface water probability”, is introduced to achieve
seamless surface water mapping. A new fused dataset tailored for this purpose was
constructed, consisting of paired synthetic aperture radar (SAR) and surface water
probability data with a 10-meter spatial resolution at a 10-day interval. A Generative CNN-
Transformer (GCT) for Gap-Filling (GF) of surface water probability, GCT-GF, was then
proposed to integrate the strengths of convolutional neural networks (CNNs) and
transformers to reconstruct gapless water probability images from multi-modal and multi-
temporal data. The GCT-GF employs a coarse-to-fine structure: information from different
time points is initially aggregated using a branched gated inpainting module, followed by
refinement and alignment of the coarse output under target SAR guidance. For adversarial
learning, a branched SN-PatchGAN discriminator is introduced to adapt to the multi-
temporal input. The results show that the GCT-GF surpasses the state-of-the-art relevant
methods in quantitative metrics and visual perception. The fusion of multi-modal, multi-
temporal inputs obvious enhance the gap-filling performance across varying gap ratios.
Applied to Baiyangdian, Poyang Lake Basin and Qinghai Lake, GCT-GF demonstrates its high
reliability on large scale scenes.
Keywords: Gap-filling; Surface water mapping; Data fusion; CNN-transformer; Generative
adversarial network (GAN)

Kaiyuan Li, Wei Chen, Yanyan Zou, Zhigang Wang, Xianzhong Zhou, Jihao Shi,
Optimized PSOMV-VMD combined with ConvFormer model: A novel gas pipeline leakage
detection method based on low sensitivity acoustic signals,
Measurement,
Volume 247,
2025,
116804,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Traditional acoustic leak detection methods rely on artificial sensor systems and
are expensive to implement. The signals collected by pipeline leak detection robots based on
low-cost microphone arrays have low signal-to-noise ratios and are difficult to capture signal
details, which affects the detection results. Therefore, this paper introduces a cost-effective
method for gas pipeline leakage detection using a combination of optimized Variational
Mode Decomposition (VMD) and the ConvFormer model. The optimized VMD reduces noise
in low-sensitivity acoustic signals, enhancing feature extraction. The ConvFormer model then
processes the spectrogram to detect leaks. Leakage experiments conducted on a 100 m gas
pipeline validated the method’s effectiveness. Results demonstrate a significant
improvement in noise reduction, with reductions in Mean Squared Error (MSE) and Mean
Absolute Error (MAE) by 20 %–30 % and 18 %–24 %, respectively. The method achieved a
high detection accuracy of 99.31 %, offering a reliable solution for intelligent pipeline
inspection robots.
Keywords: Leakage detection; Microphone arrays; VMD; ConvFormer

Shahzad Hussain, Iqra Mumtaz, Chong Wang, Pei Lv,


SF-YOLOv9: PGI based hybrid backbone with dual-path attention for small object detection in
aerial imagery,
Egyptian Informatics Journal,
Volume 33,
2026,
100888,
ISSN 1110-8665,
[Link]
([Link]
Abstract: Small object detection in aerial imagery is a challenging task due to the minimal
pixel information in dense clutter, scale variation, and complex backgrounds. YOLOv9 has
demonstrated the effectiveness of Programmable Gradient Information (PGI) in mitigating
feature degradation. However, its fully convolutional architecture lacks the capability for
global context modeling, which is critical for resolving ambiguities in small targets. To
address these limitations, we propose SF-YOLOv9, a hybrid architecture that enhances
YOLOv9c by improving the backbone through the integration of a novel PGI-Aware Swin
Fusion Block (Transformer-GELAN) at its final stage. This module effectively preserves high-
resolution local features while injecting long-range global context through Swin Transformer-
based fusion. It results in richer and more discriminative semantic representations. We
introduce a Dual-Path Spatial and Channel Attention Module (DSCAM) into the main
detection head and the reversible auxiliary branches of PGI. By refining attention across all
supervisory signals, DSCAM significantly improves gradient flow and feature fidelity during
PGI training, reducing missed detections and false positives. We evaluate SF-YOLOv9 on
VisDrone and NWPU-VHR-10 datasets to demonstrate the effectiveness of SF-YOLOv9. It
outperformed the baseline models, achieving 49.1% mAP@0.50 on VisDrone and 98.3%
mAP@0.50 on NWPU VHR-10 in small-object detection.
Keywords: Small object detection; Aerial images; YOLOv9; Swin transformer; DSCAM; UAV

Tianjiao Liu, Si-Bo Duan, Niantang Liu, Baoan Wei, Juntao Yang, Jiankui Chen, Li Zhang,
Estimation of crop leaf area index based on Sentinel-2 images and PROSAIL-Transformer
coupling model,
Computers and Electronics in Agriculture,
Volume 227, Part 2,
2024,
109663,
ISSN 0168-1699,
[Link]
([Link]
Abstract: Accurate estimation of leaf area index (LAI) is hindered by challenges in capturing
crop-specific spectral variability and integrating complex model-data relationships. To
address these issues, this study proposes a novel framework based on Sentinel-2 images,
coupling the PROSAIL physical model with a Transformer-based deep learning model. This
framework incorporates three key features contributing to its effectiveness. Firstly, Sentinel-
2 reflectance was generated using the PROSAIL model and refined through sample matching
to ensure optimal alignment with Sentinel-2 imagery specific to each crop type. Secondly,
the Maximum Information Coefficient (MIC) and Recursive Feature Elimination (RFE) were
employed to identify the most relevant spectral feature combinations for different crop
categories. Thirdly, a PROSAIL-Transformer coupling model was constructed based on
selected feature combinations to generate accurate Sentinel-2 LAI products. To validate the
proposed approach, field crop LAI measurements were collected at five plots within the
study area. Quantitative assessments demonstrate a coefficient of determination (R2) of
0.87, root mean square error (RMSE) of 0.48, and mean absolute error (MAE) of 0.36. The
proposed framework enables the production of time-series LAI maps at fine resolution,
facilitating dynamic crop monitoring and management in areas of high spatial heterogeneity.
Keywords: Leaf area index; Sentinel-2; PROSAIL-Transformer coupling model; Spectral
feature combinations

Sujin Jin, Homin Song, Jungoo Kang, Byoungjoon Yu, Seunghee Park,
Concrete crack reasoning: Explainable defect diagnosis incorporating generative Pretrained
transformer 4 and multimodal nondestructive testing data,
Automation in Construction,
Volume 181, Part C,
2026,
106661,
ISSN 0926-5805,
[Link]
([Link]
Abstract: Timely detection of aging concrete deterioration requires diagnostic methods
combining laboratory-level accuracy with field robustness. Existing models suffer from
opaque decision-making and limited integration of multimodal sensor data. This paper
presents Concrete Crack Reasoning (CCR), a field-validated diagnostic framework for
concrete bridge defects that overcomes these limitations via a two-module pipeline: (i)
sensor-specific convolutional neural networks distill features from ground penetrating radar,
impact echo, and ultrasonic testing into concise, human-readable sentences; (ii) a memory-
guided GPT-4 reasoning stage, operable in Baseline, Feedback, and Adaptive Refiner modes,
the last of which uses prompt retrieval for self-correction. On a 1088-sample, four-class
bridge-deck dataset, CCR raises top-1 accuracy from 49.4 % to 87.8 % and 97.9 %, reducing
the 95 % bootstrap confidence interval to ±1.8 %. Explanations in Adaptive Refiner mode
employ threshold-aware, cluster-referenced language consistent with expert practice,
enhancing auditability. With lightweight in-model retrieval and no external databases, CCR is
deployable on resource-constrained units, offering a practical path toward explainable AI-
assisted structural health monitoring.
Keywords: Structural health monitoring; Concrete crack detection; Multimodal non-
destructive testing; Explainable AI; GPT-4; Machine learning

Saidul Islam, Hanae Elmekki, Ahmed Elsebai, Jamal Bentahar, Nagat Drawel, Gaith Rjoub,
Witold Pedrycz,
A comprehensive survey on applications of transformers for deep learning tasks,
Expert Systems with Applications,
Volume 241,
2024,
122666,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Transformers are Deep Neural Networks (DNN) that utilize a self-attention
mechanism to capture contextual relationships within sequential data. Unlike traditional
neural networks and variants of Recurrent Neural Networks (RNNs), such as Long Short-Term
Memory (LSTM), Transformer models excel at managing long dependencies among input
sequence elements and facilitate parallel processing. Consequently, Transformer-based
models have garnered significant attention from researchers in the field of artificial
intelligence. This is due to their tremendous potential and impressive accomplishments,
which extend beyond Natural Language Processing (NLP) tasks to encompass various
domains, including Computer Vision (CV), audio and speech processing, healthcare, and the
Internet of Things (IoT). Although several survey papers have been published, spotlighting
the Transformer’s contributions in specific fields, architectural disparities, or performance
assessments, there remains a notable absence of a comprehensive survey paper that
encompasses its major applications across diverse domains. Therefore, this paper addresses
this gap by conducting an extensive survey of proposed Transformer models spanning from
2017 to 2022. Our survey encompasses the identification of the top five application domains
for Transformer-based models, namely: NLP, CV, multi-modality, audio and speech
processing, and signal processing. We analyze the influence of highly impactful Transformer-
based models within these domains and subsequently categorize them according to their
respective tasks, employing a novel taxonomy. Our primary objective is to illuminate the
existing potential and future prospects of Transformers for researchers who are passionate
about this area, thereby contributing to a more comprehensive understanding of this
groundbreaking technology.
Keywords: Transformer; Self-attention; Deep learning; Natural language processing (NLP);
Computer vision (CV); Multi-modality

Yakun Yang, Yu Chen, Yukang Xu, Yudong Wang, Guangwu Sun,


Inverse prediction of micro-nano fiber membrane preparation process based on deep
learning,
Polymer,
Volume 342,
2026,
129372,
ISSN 0032-3861,
[Link]
([Link]
Abstract: The performance of micro-nano-scale fibrous membrane materials is highly
dependent on their microstructure, which is governed by fabrication process parameters.
Traditional trial-and-error optimization methods are inefficient and costly because they
require repeated experiments. Consequently, this study proposes a deep learning-based
inverse design strategy to predict the process parameters from target fiber microstructures.
Scanning electron microscopy images of fibrous networks are collected and used as inputs to
a model incorporating an attention penalty mechanism to construct an intelligent
“microstructure–process parameter” mapping. The RT-Net model integrates an improved
ResNet-50 with a Transformer architecture and leverages transfer learning and microscopic
image enhancement to overcome the limitations under small-sample conditions. Seven
types of fibrous membranes prepared under different microfluidic spinning conditions are
used for training. The model achieves a validation accuracy of 98.81 % and an F1 score of
0.9857 for process parameter classification, outperforming the baseline models by 2.46 %
and 2.17 %, respectively. Further microfluidic spinning experiments based on the predicted
parameters, followed by statistical analysis, confirm the reliability of the model for inverse
prediction and structural reproduction. This method provides a novel and efficient approach
for the intelligent fabrication of micro-nano fibrous membranes and holds potential for
extension to the development of other functional textile materials.
Keywords: Micro-nano fiber materials; Inverse design; Deep learning; Transfer learning;
Attention mechanism; Process parameter classification; Functional textile materials

Farhatullah, Xin Chen, Deze Zeng, Rahmat Ullah, Rab Nawaz, Jiafeng Xu, Tughrul Arslan,
A deep learning approach for non-invasive Alzheimer’s monitoring using microwave radar
data,
Neural Networks,
Volume 181,
2025,
106778,
ISSN 0893-6080,
[Link]
([Link]
Abstract: Over 50 million people globally suffer from Alzheimer’s disease (AD), emphasizing
the need for efficient, early diagnostic tools. Traditional methods like Magnetic Resonance
Imaging (MRI) and Computed Tomography (CT) scans are expensive, bulky, and slow.
Microwave-based techniques offer a cost-effective, non-invasive, and portable solution,
diverging from conventional neuroimaging practices. This article introduces a deep learning
approach for monitoring AD , using realistic numerical brain phantoms to simulate scattered
signals via the CST Studio Suite. The obtained data is preprocessed using normalization,
standardization, and outlier removal to ensure data integrity. Furthermore, we propose a
novel data augmentation technique to enrich the dataset across various AD stages. Our deep
learning approach combines Recursive Feature Elimination (RFE) with Principal Component
Analysis (PCA) and Autoencoders (AE) for optimal feature selection. Convolution Neural
Network (CNN) is combined with Gated Recurrent Unit (GRU), Bidirectional Long Short Term
Memory (Bidirectional-LSTM), and Long Short-Term Memory (LSTM) to improve
classification performance. The integration of RFE-PCA-AE significantly elevates
performance, with the CNN+GRU model achieving an 87% accuracy rate, thus outperforming
existing studies.
Keywords: Alzheimer’s disease; Classification; Deep learning; Data augmentation; Microwave
scattering; Signal processing

Mingyang Du, Ping Zhong, Xiaohao Cai, Daping Bi, Aiqi Jing,
Robust Bayesian attention belief network for radar work mode recognition,
Digital Signal Processing,
Volume 133,
2023,
103874,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Understanding and analyzing radar work modes play a key role in electronic
support measure system. Many classifiers, for example those based on convolutional neural
network (CNN) and recurrent neural network (RNN), are available for recognizing radar work
modes as well as emitter types from their waveform parameters. However, the performance
of these methods may suffer significantly when confronting different types of signal
degradation, e.g., measurement error, lost pulse and spurious pulse. To tackle this issue, we
in this paper develop a Bayesian attention belief network (BABNet) based on Bayesian neural
networks in which the probability distribution over weights can help to enhance the model
robustness for corrupted data. In particular, we adopt pre-trained CNN as the Bayesian
inference prior. This not only accelerates the convergence speed, but also avoids the training
process getting stuck in bad local minima. Meanwhile, instead of using RNNs which are
difficult to be implemented in parallel, the combination of padding operation and attention
module in the proposed BABNet enables CNN, as the backbone, to process sequential data
with variable length. Extensive experiments are conducted to demonstrate the recognition
capability and robustness of the BABNet in different environments.
Keywords: Radar work mode; Pulse descriptor word; Attention mechanism; Bayesian neural
network; Robustness; Recognition

Yurui Zhao, Xiang Wang, Zhitao Huang,


BiLSTM-Filt: Neural network for radar word segmentation,
Neural Networks,
Volume 181,
2025,
106815,
ISSN 0893-6080,
[Link]
([Link]
Abstract: Radar word extraction is the analysis foundation for multi-function radars (MFRs)
in electronic intelligence (ELINT). Although neural networks enhance performance in radar
word extraction, current research still faces challenges from complex electromagnetic
environments and unknown radar words. Therefore, in this paper, we propose a promising
two-stage radar word extraction framework, consisting of segmentation and recognition. To
fill the vacancy of radar word segmentation, we establish the mathematical model from the
time series analysis viewpoint and design a novel segmentation neural network based on Bi-
direction Long Short-Term Memory with a filter module (BiLSTM-Filt). Specific radar word
structure characteristics are extracted by training the network and applied for detecting
radar words in the pulse train. To further improve segmentation performance, a bounding
box regression method is designed to merge information from sub-region structures.
Simulation experiments on a typical MFR, Mercury, reveal that the proposed method can
outperform the baseline methods within complex electromagnetic environments, containing
corrupted environments, various pulse backgrounds, and variable pulse train lengths. Due to
the artificial design structure, the proposed method can also make a trial on unknown radar
word segmentation.
Keywords: Electronic intelligence (ELINT); Multi-function radar (MFR); Radar word
extraction; Radar word segmentation; Neural network; Time series analysis

Qidi Shu, Xiaolin Zhu, Shuai Xu, Yan Wang, Denghong Liu,
RESTORE-DiT: Reliable satellite image time series reconstruction by multimodal sequential
diffusion transformer,
Remote Sensing of Environment,
Volume 328,
2025,
114872,
ISSN 0034-4257,
[Link]
([Link]
Abstract: Repetitive optical observations from satellites are crucial for monitoring earth
surface dynamics over time. However, optical satellite image time series is severely affected
by frequent data gaps due to clouds and shadows. While synthetic aperture radar (SAR)
provides cloud-penetrating capabilities to complement missing optical data, recent
advancements in time series reconstruction have shifted focus from incorporating single SAR
image to exploiting SAR time series. However, current methods still struggle for challenging
scenarios like highly dynamic surface, persistent data gaps, and exhibit poor resilience to
inaccurate cloud masks. In this research, we approach the time series reconstruction
problem from the perspective of conditional generation. We propose a multimodal diffusion
framework termed RESTORE-DiT, which firstly promotes the sequence-level optical-SAR
fusion through a diffusion framework. Specifically, date-matched SAR time series provide
under-cloud surface dynamics to guide the denoising process of cloudy areas, and date
information is embedded to account for irregular observation intervals and periodic
patterns. Extensive experiments on three regions have shown the proposed method
achieves state-of-the-art performance. RESTORE-DiT outperforms comparison methods by
2.87 dB in PSNR and a 27.2 % reduction in RMSE on France site. SAR and date information
together increase PSNR by 2.41 dB. The reconstructed optical image time series is verified to
accurately reflect the crop growth condition and support for long-term vegetation
observations. In addition, RESTORE-DiT can be easily extended to other conditional
reconstruction or prediction tasks for arbitrary time series image data, thus facilitating
spatiotemporal analysis research. The codes will be public available at:
[Link]
Keywords: Time series reconstruction; Diffusion model; Data fusion; Optical-SAR fusion;
Cloud removal

Marco Martino Rosso, Giulia Marasco, Salvatore Aiello, Angelo Aloisio, Bernardino Chiaia,
Giuseppe Carlo Marano,
Convolutional networks and transformers for intelligent road tunnel investigations,
Computers & Structures,
Volume 275,
2023,
106918,
ISSN 0045-7949,
[Link]
([Link]
Abstract: Visual inspections do not provide a reliable and objective assessment of the
conservation state of road tunnels. Although direct tests might represent a valid survey
approach, they would often lead to prohibitive costs if performed extensively. Therefore,
indirect techniques, such as ground-penetrating radar (GPR), have become fundamental to
supporting limited direct tests. The analysis of the GPR tunnel linings profiles is mainly hand-
operated. It permits the detection of various tunnel linings defects, characterizing a tunnel’s
global health state. In the present work, the authors developed an artificial intelligence (AI)
based automatic road tunnel defects hierarchical classification framework to improve the
efficiency of this powerful indirect surveying method. Adopting the most recent tools in
image processing provided by the deep learning (DL) community, the authors proposed a
convolutional neural network (CNN) with the acknowledged ResNet-50 architecture,
initialized through the transfer learning method. For the sake of comparisons, the authors
also adopted the state-of-art convolutional EfficientNet architecture. To further improve the
proposed framework, the authors investigated how the bidimensional Fourier transform
applied as a preprocessing procedure could affect the classification performances of the
ResNet-50 model. Finally, to further enhance the classification performance, the state-of-art
neural vision transformer (ViT) architecture has been adopted with the transfer learning
approach to the currently proposed defects classification framework.
Keywords: Deep learning; Vision Transformers; Road tunnels; Fourier transform;
Convolutional Neural Network; Structural Health Monitoring; Ground Penetrating Radar

Omar Elharrouss, Yassine Himeur, Yasir Mahmood, Saed Alrabaee, Abdelmalik Ouamane,
Faycal Bensaali, Yassine Bechqito, Ammar Chouchane,
ViTs as backbones: Leveraging vision transformers for feature extraction,
Information Fusion,
Volume 118,
2025,
102951,
ISSN 1566-2535,
[Link]
([Link]
Abstract: The emergence of Vision Transformers (ViTs) has marked a significant shift in the
field of computer vision, presenting new methodologies that challenge traditional
convolutional neural networks (CNNs). This review offers a thorough exploration of ViTs,
unpacking their foundational principles, including the self-attention mechanism and multi-
head attention, while examining their diverse applications. We delve into the core mechanics
of ViTs, such as image patching, positional encoding, and the datasets that underpin their
training. By categorizing and comparing ViTs, CNNs, and hybrid models, we shed light on
their respective strengths and limitations, offering a nuanced perspective on their roles in
advancing computer vision. A critical evaluation of notable ViT architectures—including DeiT,
DeepViT, and Swin-Transformer—highlights their efficacy in feature extraction and domain-
specific tasks. The review extends its scope to illustrate the versatility of ViTs in applications
like image classification, medical imaging, object detection, and visual question answering,
supported by case studies on benchmark datasets such as ImageNet and COCO. While ViTs
demonstrate remarkable potential, they are not without challenges, including high
computational demands, extensive data requirements, and generalization difficulties. To
address these limitations, we propose future research directions aimed at improving
scalability, efficiency, and adaptability, especially in resource-constrained settings. By
providing a comprehensive overview and actionable insights, this review serves as an
essential guide for researchers and practitioners navigating the evolving field of vision-based
deep learning.
Keywords: Vision transformers; Transformers; Deep learning; Computer vision; Attention

Zhenhua Li, Jiuxi Cui, Heping Lu, Feng Zhou, Yinglong Diao, Zhenxing Li,
Prediction method for instrument transformer measurement error: Adaptive decomposition
and hybrid deep learning models,
Measurement,
Volume 253, Part D,
2025,
117592,
ISSN 0263-2241,
[Link]
([Link]
Abstract: The measurement accuracy of current transformers is crucial for power system
protection and trade fairness. The high penetration of renewable energy into the power grid
has affected the transient performance of power systems, posing significant challenges for
accurate current transformer measurement. To address this issue, this paper proposes a
prediction model for transformer measurement accuracy based on an adaptive dual-modal
decomposition strategy and a hybrid deep learning architecture. The framework integrates
an enhanced Adaptive Time-Varying Filter (A-TVF), an enhanced Adaptive Variational Mode
Decomposition (A-VMD), the Residual Error Index (REI), and the Maximum Information
Coefficient (MIC). First, A-TVF preprocesses the collected data by setting REI as the
optimization objective to adaptively adjust filter construction parameters, including the B-
spline order, bandwidth threshold, and decomposition number, and decomposes the
collected ratio error sequence to reduce the non-stationarity of the original sequence.
Subsequently, indices such as PE and Kurt are used to screen the decomposed sub-
sequences and reconstruct the complex components. Then, A-VMD is applied to further
decompose the complex components, minimizing MIC by adaptively determining the
decomposition number, penalty factor, convergence accuracy, and fidelity parameters.
Afterward, the complexity of the subcomponents obtained from the secondary
decomposition is calculated, and the entire sequence is reconstructed. Finally, a hierarchical
prediction model integrating Temporal Convolutional Networks (TCN), Bidirectional Gated
Recurrent Units (BiGRU), and a Multi-Head Attention mechanism (MHA) is employed to
predict the reconstructed components and generate the final results. Experimental results
demonstrate that the proposed adaptive dual-modal decomposition method significantly
improves prediction performance: compared with non-decomposition models, RMSE, MAE,
and SMAPE were reduced by an average of 50.12%, 46.09%, and 37.70% in global
decomposition scenarios, and by 25.92%, 23.69%, and 19.96% in rolling decomposition
scenarios, respectively. These results validate the effectiveness of the proposed method in
reducing data complexity and improving the accuracy and stability of Ratio Error predictions.
Keywords: ECT; Ratio error prediction; Adaptive dual-modal decomposition; Decomposition
and combination strategy; Hybrid deep model; Measurement accuracy
Seyedeh Leili Mirtaheri, Ali Kafi Tafti, Hamid Heidari Soureshjani, Andrea Pugliese,
GreenBERT: A lightweight green transformer for automated prediction of software
vulnerability scores,
Array,
Volume 28,
2025,
100536,
ISSN 2590-0056,
[Link]
([Link]
Abstract: Timely assessment of software vulnerabilities is critical for effective patch
prioritization, yet manual Common Vulnerability Scoring System (CVSS) scoring remains slow
and resource-intensive. While transformer-based models such as BERT have advanced
automated scoring, their substantial computational demands conflict with sustainable, green
computing objectives. This paper introduces GreenBERT, a tailored ensemble of lightweight
student Transformers, each specialized on individual CVSS metrics through a targeted multi-
head knowledge distillation framework. By jointly optimizing alignment with ground-truth
labels and softened outputs from a fine-tuned BERT teacher, GreenBERT efficiently captures
complex vulnerability patterns while significantly reducing computational overhead.
Extensive experiments on the National Vulnerability Database (NVD) and a more challenging
COMBINED dataset demonstrate that GreenBERT achieves an average F1-score
improvement exceeding 6% over the BERT baseline, while simultaneously reducing inference
time by approximately 80% and cutting energy usage and CO2 emissions by about 70%.
These results position GreenBERT as a robust, scalable, and environmentally conscious
solution for high-performance vulnerability scoring, effectively reconciling the traditionally
conflicting goals of predictive accuracy and sustainable AI.
Keywords: Green computing; Software vulnerability assessment; Deep learning; Knowledge
distillation; Transformers

John Atanbori, Christos A. Frantzidis, Mohammed Al-Khafajiy, Aliyu Aliyu, Behnaz Sohani,
Kofi Appiah, Harriet Moore, Catherine Sanders, Alastair I. Ward,
Learning with noisy labels for classifying biological echoes in polarimetric weather radar
observations using artificial neural networks,
Neurocomputing,
Volume 634,
2025,
129892,
ISSN 0925-2312,
[Link]
([Link]
Abstract: The identification of biological echoes in radar data has revolutionized research
into airborne migratory species. Deep learning applied to polarimetric weather radar
observations can reveal signature patterns of mass movement by bio-scatterers such as
birds, bats, and insects. However, due to the difficulties in labelling bio-scatterers in these
data, threshold approaches have been proposed in the literature. In this research, we used
the depolarization ratio (DR) based on differential reflectivity (zDR) and the cross-correlation
coefficient (pHV), along with citizen scientist-reported data, to label bio-scatterers for deep
learning. This method of labelling biological echoes in radar signatures is prone to noise,
which impacts the accuracy of any model that relies on it. We introduce a novel semi-
supervised co-training approach that uses a bootstrap ensemble with a confidence
threshold. Our ensemble consists of the newly proposed STNet and two modified FNet
models, which incorporate co-learning through bootstrap sampling for label correction. This
innovative method significantly improves classification accuracy across all three multivariate
numerical datasets compared to baseline models that lack co-learning with bootstrap-based
label correction.
Keywords: Artificial neural networks (ANN); Ensemble classifiers; Radar bio-scatterer
classification; Semi-supervised co-training

Cencen Liu, Dongyang Zhang, Guoming Lu, Wen Yin, Jielei Wang, Guangchun Luo,
SRMamba-T: Exploring the hybrid Mamba-Transformer network for Single Image Super-
Resolution,
Neurocomputing,
Volume 624,
2025,
129488,
ISSN 0925-2312,
[Link]
([Link]
Abstract: Single Image Super-Resolution (SISR) has made significant advancements with both
CNN-based and Transformer-based models. However, CNNs often struggle to capture long-
range dependencies effectively, and Transformers, though powerful, are hindered by
quadratic computational complexity. In recent years, state space models (SSMs), such as
Mamba, have emerged as promising alternatives due to their ability to model long-range
dependencies with linear time complexity. In light of this, we propose SRMamba-T, a hybrid
model that strategically combines Mamba and Transformer architectures to balance
computational efficiency with high performance. Specifically, we employ Transformer layers
following Mamba layers to further enhance the model’s capability to handle long-range
spatial information and expand its effective receptive fields. To design a lightweight network,
we propose a multi-directional selective scanning module to reduce parameter count and
improve computational efficiency. Additionally, a feature fusion module serves as a
bottleneck to effectively integrate hierarchical features and enhance the model’s
representational ability. Comprehensive experimental evaluations across five widely
recognized benchmarks underscore our model’s effectiveness, demonstrating its superiority
over existing state-of-the-art (SOTA) methods. For example, our model achieved a significant
PSNR improvement of 0.28 dB for ×2 lightweight super-resolution on Urban100, with a
reduction in MACs by 38.7% (from 198.1G to 121.5G) compared to the SOTA model,
MambaIR.
Keywords: Single image super-resolution; Mamba; Transformer; Lightweight model

Faruk Keskin, Fesih Keskin, Gültekin Işık,


GAST: A graph-augmented spectral–spatial transformer with adaptive gated fusion for small-
sample hyperspectral image classification,
ISPRS Open Journal of Photogrammetry and Remote Sensing,
Volume 19,
2026,
100116,
ISSN 2667-3932,
[Link]
([Link]
Abstract: Accurate hyperspectral image (HSI) classification under scarce labels and class
imbalance requires models that couple long-range spectral reasoning with irregular local
spatial context. We present GAST, a Graph-Augmented spectral–spatial Transformer with
Adaptive Gated Fusion for Small-Sample Hyperspectral Image Classification. GAST pairs a
lightweight spectral Transformer with a GATv2-based spatial branch on an 8-neighbor pixel
graph, and fuses them via a center-conditioned, channel-wise gating mechanism that uses
the center-pixel representation to modulate all tokens in the patch. Unlike conventional
static fusion strategies (e.g., concatenation or summation) that assign fixed importance to
modalities regardless of image content, this adaptive fusion dynamically modulates the
spectral and spatial streams at the pixel level, allowing the model to prioritize spatial texture
for complex urban structures while shifting focus to spectral signatures for subtle vegetation
classes. Training is further stabilized by an imbalance-aware objective that switches between
weighted cross-entropy and focal loss according to a measured class ratio, and by a two-
stage Bayesian hyperparameter search that aligns capacity with scene statistics. Across eight
public benchmarks under a 5%-label protocol, GAST consistently matches or surpasses
recent hybrid graph-Transformer architectures while remaining compact and fast at
inference. Ablation studies confirm the complementary roles of both branches and the
benefit of gated fusion. The resulting architecture offers a strong accuracy–efficiency trade-
off and reliable performance across seeds, making it a practical solution for low-data HSI
applications. The code is publicly available at [Link]
Keywords: Hyperspectral image classification; Transformer; Graph attention network;
Spectral–spatial fusion; Deep learning

Yifan Chen, Haibin Zhang, Xiang Shen, Xiangsheng Chen, Dong Su, Jiuqi Wu,
A novel deep learning-based identification technology of cutting pile states during super-
large diameter shield tunnelling,
Tunnelling and Underground Space Technology,
Volume 164,
2025,
106836,
ISSN 0886-7798,
[Link]
([Link]
Abstract: Accurately identifying the position and quantity of piles is critical for ensuring the
safe tunnelling process in shield cutting pile projects. The vibration signals generated during
the shield cutting pile process contain abundant information. To address the challenge of
determining pile positions and quantities, this study proposes a method for the
identification of strata based on vibration characteristics, integrating the dual advantages of
knowledge-driven and data-driven approaches. The method includes a data processing
module, a knowledge-driven module, a transformer-based model (MT), and a
comprehensive evaluation module, and it has been validated in the Guangzhou Haizhu Bay
shield tunnel project. The results show that the developed method achieves an accuracy of
99.56% in the identification of strata types, improving by 1.33%, 1.11%, and 14.16%
compared to the MLP, RF, and LSTM models, respectively. As the number of cutting piles
increases, the frequency of vibration signals gradually rises, while the amplitude shows no
significant change. Based on this finding, the top five frequencies were used as input.
Position encoding was employed to effectively learn the positional information of the
frequency, enabling the MT model to achieve an accuracy of 65.71% in identifying multiple
piles, improving by 10.51%, 19.14%, and 23.46% compared to the MLP, RF, and LSTM
models, respectively. Comprehensive evaluation analysis indicates that this method
demonstrates superior recall and weighted accuracy, highlighting its strong flexibility and
applicability in engineering contexts.
Keywords: Shield tunnel; Cutting pile; Deep learning; Ground identification; Transformer

Lifu He, Zhongchu Huang, Haidong Shao, Zhangbo Hu, Yuting Wang, Jie Mei, Xiaofei Zhang,
Fault Diagnosis of Wind Turbine Blades Based on Multi-Sensor Weighted Alignment Fusion in
Noisy Environments,
Computers, Materials and Continua,
Volume 86, Issue 3,
2026,
,
ISSN 1546-2218,
[Link]
([Link]
Abstract: Deep learning-based wind turbine blade fault diagnosis has been widely applied
due to its advantages in end-to-end feature extraction. However, several challenges remain.
First, signal noise collected during blade operation masks fault features, severely impairing
the fault diagnosis performance of deep learning models. Second, current blade fault
diagnosis often relies on single-sensor data, resulting in limited monitoring dimensions and
ability to comprehensively capture complex fault states. To address these issues, a multi-
sensor fusion-based wind turbine blade fault diagnosis method is proposed. Specifically, a
CNN-Transformer Coupled Feature Learning Architecture is constructed to enhance the
ability to learn complex features under noisy conditions, while a Weight-Aligned Data Fusion
Module is designed to comprehensively and effectively utilize multi-sensor fault information.
Experimental results of wind turbine blade fault diagnosis under different noise
interferences show that higher accuracy is achieved by the proposed method compared to
models with single-source data input, enabling comprehensive and effective fault diagnosis.
Keywords: Wind turbine blade; multi-sensor fusion; fault diagnosis; CNN-transformer
coupled architecture

Zongbin Zhang, Xiaoqiao Huang, Chengli Li, Feiyan Cheng, Yonghang Tai,
CRAformer: A cross-residual attention transformer for solar irradiation multistep forecasting,
Energy,
Volume 320,
2025,
135214,
ISSN 0360-5442,
[Link]
([Link]
Abstract: In recent years, solar energy has gained widespread adoption in smart grids due to
its safety, environmental friendliness, abundance, and other advantages, driving the
application of photovoltaic (PV) power generation technology. Accurately predicting solar
irradiance is essential for ensuring the operational stability of PV power systems, making it a
critical challenge for maintaining grid security and stability. Although Transformer models in
deep learning have achieved significant advancements in solar irradiance forecasting,
existing studies often treat cross-batch time-series data (TSD) as independent. By
overlooking the complex coupling relationships between different data batches, they fail to
fully capture the underlying patterns in TSD under varying conditions. Moreover, handling
the long-term dependencies and short-term weather-induced fluctuations inherent in TSD
remains difficult. To address these issues, this paper proposes an efficient Transformer
model (CRAformer) based on Cross-Residual Attention (CRA) for multi-step solar irradiance
forecasting. CRAformer effectively captures the deep coupling relationships within TSD
through a residual scoring mechanism, which can dynamically adjust feature weights and
balance long-term dependencies with short-term variations. Furthermore, by incorporating
a dual-output mode and dual-attention strategy, the model can deconstruct complex data
structures and guide the prediction process with greater accuracy. Additionally, the newly
designed Convolutional Weighted Fusion Module (CWFM) enhances the model's capability
to recognize diverse patterns and characteristics in TSD. By dynamically regulating the
information transfer process, the CWFM improves the model's generalization, fitting
accuracy, and robustness. To evaluate CRAformer's performance, four prediction tasks with
varying time steps (24 h, 48 h, 72 h, 96 h) were designed using irradiance datasets from
different locations: Denver, Clark, and Folsom. The experimental results demonstrate that,
compared to the second-best model, iTransformer, CRAformer reduces the RMSE by an
average of 5.6 %, 3.9 %, and 5.6 % across the four prediction steps for the datasets from
Denver, Clark, and Folsom, respectively. These results indicate that CRAformer offers
significant advantages in multi-step solar irradiance forecasting, providing a valuable
reference for future model optimization.
Keywords: Multi-step irradiance forecasting; Cross-residual; Convolutional weighting; Dual-
output mode; Photovoltaic power generation

Shaopeng He, Mingjun Wang, Nicola Forgione, Andrea Pucciarelli, W.X. Tian, S.Z. Qiu, G.H.
Su,
A multi-task Transformer-Mamba-Seq framework for real-time estimation of spatiotemporal
thermal stratification in passive residual heat exchanger,
International Communications in Heat and Mass Transfer,
Volume 169, Part D,
2025,
109868,
ISSN 0735-1933,
[Link]
([Link]
Abstract: Passive Residual Heat Removal Heat Exchanger (PRHR HX) is a critical component in
Generation-III nuclear power systems. Its spatiotemporal thermal stratification
characteristics directly influence residual heat removal capacity and serve as key inputs for
multiphysics coupling analyses. However, the complexity of input conditions challenges
traditional simulation and AI approaches, particularly under abnormal and accident
scenarios. To address this, we propose a multi-task Transformer-Mamba-Seq framework that
integrates multi-head attention with a selective scan mechanism. Compared to conventional
models, it demonstrates superior performance in both 5-fold cross-validation and

computational costs—cutting parameters by ∼90 % and training time by at least 59 %. Our


independent tests. Furthermore, a sequential training strategy significantly reduces

framework enables real-time prediction of the 4D temperature field and thermal


stratification characteristics in PRHR HX with high accuracy (RMSE/MAPE/R2:
1.81 K/0.41 %/0.887). It achieves a speedup of over 1500× compared to CFD simulations.
This work provides an efficient and accurate tool for real-time thermal analysis of PRHR HX,
supporting the thermal safety of Generation-III nuclear systems and could offering low-cost,
high-resolution inputs for thermal stress and flow-induced vibration analyses.
Keywords: Passive residual heat exchanger; Thermal stratification; Real-time estimation;
Transformer; Mamba

Dunlu Peng, Meiling Chen, Yiqin Zhang, Zekun Tian,


Enhanced optic-flow extrapolation for Doppler radar nowcasting with Dynamic Weight
Attention,
Expert Systems with Applications,
Volume 267,
2025,
126168,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Doppler radar echo extrapolation is an important method for extreme weather
forecasting. However, traditional optical flow methods lack learnable components and are
not suitable for complex atmospheric changes. Consequently, researchers have turned to
deep neural networks for prediction. Yet, the predictions from this approach often suffer
from issues such as mean reversion and a lack of small- and medium-scale structure. This
paper proposes a novel approach that combines optical flow methods with deep neural
networks. By introducing an artificially defined momentum weight matrix based on prior
assumptions, predictions for any future time distance are generated from full-scale optical
flow. Additionally, we propose a full-scale advection extractor, leveraging the continuity of
distribution in mesoscale and small-scale atmospheric systems and focusing on the long- and
short-distance relationships within the contour surface distribution sequence, which
improves the prediction accuracy of fine-scale advection. The experimental results show
that, compared with other advanced methods, the proposed method demonstrates
advantages in predicting extreme radar echoes and maintaining the echo structure.
Specifically, it achieved an improvement of 24.1% and 21.3% on key indicators such as CSI
and HSS, respectively, and reached 0.948 on the SSIM. Building on this, the inference speed
of our method is comparable to other deep learning approaches, being 3.25 times faster
than flow-based methods.
Keywords: Nowcasting; Optical flow; Neural network; Doppler radar

Inam Abousaber, Hany El-Ghaish, Haitham F. Abdallah,


A novel explainable ECG classification with spatio-temporal transformers and hybrid loss
optimization,
Biomedical Signal Processing and Control,
Volume 110, Part B,
2025,
108142,
ISSN 1746-8094,
[Link]
([Link]
Abstract: This paper presents a novel deep learning framework for ECG arrhythmias
detection, integrating Spatio-Temporal Adaptive Embedding (STAE) Transformers and
Variational Autoencoders (VAEs) to improve classification accuracy and address class
imbalance. Traditional ECG classification models struggle to capture long-range temporal
dependencies and handle imbalanced datasets, leading to poor sensitivity for rare
arrhythmias. The proposed system employs STAE Transformers to model intricate temporal
and spatial relationships within ECG signals to overcome these challenges. At the same time,
VAEs generate diverse and realistic ECG samples to enhance model generalization,
particularly for underrepresented arrhythmias. Additionally, combining Focal Loss and Dice
Loss, a Hybrid Loss Function further optimizes performance by focusing on hard-to-classify
arrhythmias. The model is evaluated on the MIT-BIH Arrhythmias Database and PTB
Diagnostic ECG Database using 5-fold cross-validation, achieving an accuracy of 99.56% and
a macro F1-score of 95.40%, outperforming existing state-of-the-art methods, with a 3.5%
improvement in sensitivity for rare arrhythmias. To ensure interpretability, SHapley Additive
exPlanations (SHAP) and Gradient-weighted Class Activation Mapping (Grad-CAM) are
utilized, highlighting the QRS complex and RR intervals as the most critical features and
confirming that the model focuses on clinically relevant waveform regions. These results
demonstrate the effectiveness of our approach in developing an accurate, interpretable, and
robust deep learning system for ECG arrhythmias detection, paving the way for more reliable
clinical decision support systems.
Keywords: Electrocardiography (ECG); Arrhythmias Detection; Spatio-Temporal Transformer;
Variational Autoencoders (VAEs); Hybrid Loss Function; SHAP; Grad-CAM

Zhi Tang, Zikang Feng, Zuqiang Su, Maolin Luo, Guo Wu, Lin Bo,
Fault diagnosis for the gas-path system of an engine test bed based on multi-modal signals,
Neurocomputing,
Volume 670,
2026,
132526,
ISSN 0925-2312,
[Link]
([Link]
Abstract: The aerospace engine test bed, as a pivotal equipment for evaluating engine
reliability, necessitates rigorous monitoring of its health state to ensure the safe operation of
the engine. The multi-point, multi-modal sensor information in the test bed's gas-path
system is characterized by complexity and variability. Traditional fault diagnosis methods
often struggle with issues such as labor-intensive threshold setting, frequent false alarms,
and missed detections. Inspired by the self-attention mechanism, a dynamic aware diagnosis
network (DADN) is proposed based on multi-modal sensor fusion. Firstly, DADN employs
shift-aware attention to focus on a small number of critical points (offset sampling) in the
sensor signal so that the discriminative local temporal features are extracted. Secondly,
dynamic pooling tricks are employed to score the channel-wise features over the duration
and generate fixed-length abstract representations. Shift-aware attention and dynamic
pooling constitute a temporal feature learning pipeline from “fine-grained localization” to
“global dynamic pooling”. Finally, multi-point, multi-modal sensor signals are fed into the
DADN to capture inter-sensor interactions and fault-related patterns, thereby achieving
accurate fault identification in the gas-path system. Experimental results demonstrate that
DADN exhibits superior diagnostic accuracy and robustness, highlighting its potential for
advanced fault diagnosis in an engine test bed's gas-path system.
Keywords: Fault diagnosis; Multi-modal; Engine test bed; Gas-path system

Zhongrui Bai, Fanglin Geng, Hao Zhang, Xianxiang Chen, Lidong Du, Peng Wang, Pang Wu,
Gang Cheng, Zhen Fang, Yirong Wu,
Non-contact blood pressure estimation using FMCW radar: A two-stream approach focused
on central arterial activity,
Biomedical Signal Processing and Control,
Volume 106,
2025,
107718,
ISSN 1746-8094,
[Link]
([Link]
Abstract: This paper proposes a radar-based two-stream blood pressure (BP) estimation
framework (R2S-BP), focusing on central arterial activity. It separately analyzes central-
arterial pulse transit time (caPTT) and pulse wave morphology using multi-location Doppler
Cardiogram (DCG) data from millimeter wave FMCW radar. Specifically, phase information at
harmonic heart rate frequencies is used to compute time delay arrays, representing caPTT-
related features. Additionally, k-Shape clustering is employed to select optimal DCGs from
the neck and chest regions that contain BP-related morphological features. These features
are processed through a two-stream neural network combining BiLSTM, ResNet, and multi-
head attention modules. Subject-independent 9-fold cross-validation results show that the
standard deviations of the errors for systolic and diastolic BP are 7.33 and 5.36 mmHg,
respectively. The intra-subject correlation coefficient for both systolic and diastolic BP
averages 0.82. Comparative and ablation studies demonstrate the superiority of the two-
stream approach and the critical importance of its components. This approach integrates
physiologically guided manual feature construction with a deep learning model, fully
leveraging the capabilities of FMCW radar data.
Keywords: Non-contact blood pressure estimation; Two-stream neural network; Doppler
Cardiogram; Central-artery pulse transit time

Yonggang Qian, Yinghua Wang, Hongwei Liu, Zelong Wang, Feipeng Yu, Chunhui Qu,
MPRANet: Multi-scale perception and reference attention network for lightweight SAR target
recognition,
Neurocomputing,
Volume 668,
2026,
132310,
ISSN 0925-2312,
[Link]
([Link]
Abstract: Deep learning methods have been widely used in Synthetic Aperture Radar
Automatic Target Recognition (SAR ATR). However, challenges remain due to limited SAR
data and computational constraints on mobile devices, which hinder model training and
deployment. In this paper, we propose a Multi-scale Perception and Reference Attention
Network (MPRANet) for lightweight SAR ATR, which is a hybrid structure combining
convolutional networks and transformers, built upon the ShuffleNetV2 network. Specifically,
MPRANet introduces two key improvements compared to the CNN-based ShuffleNetV2.
Firstly, we replace the depthwise convolutions (DWConv) in the downsampling and basic
units of ShuffleNetV2 with the Multi-scale Parameter-Shared Convolution (MPConv) module.
MPConv enables the extraction of multi-scale features of SAR targets with almost no
additional parameters, thereby enhancing the network’s feature extraction capabilities.
Secondly, we propose a lightweight Reference Attention Transformer (RAformer) to capture
global information, addressing the issue of insufficient channel feature interaction in
ShuffleNetV2. In RAformer, a Local Linear Mapping Unit (LMU) is designed to perform linear
mappings, reducing the introduction of redundant features while ensuring its lightweight
and efficient nature. RAformer contains two modules: the Reference Vector Attention (RVA)
module, which efficiently models attention relationships, and the Lightweight Feedforward
Neural Network (LW-FFN) module, which enhances the network’s ability to capture
nonlinear representations. We evaluated the performance of MPRANet using publicly
available SAR datasets, including the MSTAR dataset, OpenSARShip dataset, and SAR-
AIRcraft-1.0 dataset. The experimental results demonstrate that MPRANet consistently
achieves superior recognition performance compared to other lightweight networks of
similar complexity.
Keywords: Synthetic aperture radar (SAR); Automatic target recognition (ATR); Convolutional
neural networks (CNN); Transformer; Lightweight

Yuheng Chen, Decheng Feng, Zhongshi Pei, Xiaoxuan Mao, Lulu Fan, Meng Xu, Yang Li,
Dongsheng Wang, Junyan Yi,
Identification and information acquisition of high-value construction solid waste combined
millimeter-wave radar and convolutional neural networks,
Waste Management,
Volume 194,
2025,
Pages 390-400,
ISSN 0956-053X,
[Link]
([Link]
Abstract: The accumulation of construction solid waste (CSW) leads to the waste of land
resources and environmental pollution, becoming a significant social problem. Identifying
the amount of high-value CSW is essential for assessing the value of accumulated CSW and
formulating appropriate recycling strategies. With the development of machine learning
technology, CSW recognition techniques combining image acquisition devices and
convolutional neural networks have been widely applied. However, most technologies are
based on 2D images, making it difficult to recognize high-value CSW in accumulated CSW.
This study proposes a new method to identify high-value CSW using millimeter wave radar
based on penetration properties of electromagnetic waves. First, efficient imaging of CSW
was achieved by optimizing the imaging algorithm. Then, CSW were classified in
combination with the selected convolutional neural network (CNN) method based on the
dataset constructed in the lab. At last, high-value CSW was screened by normalizing the
imaging algorithm. The findings indicate that the balance between imaging effect and
efficiency is achieved when the step speed and height are 200 mm/s and 4 mm. The
optimized imaging approach effectively captures images of CSW. Compared with SegNet and
PSPNet, DeepLabv3+ can identify complete bricks and reinforcing bars precisely. The
accuracy can reach 85.18 %. Moreover, the millimeter-wave radar can determine the
location and size of waste and can potentially acquire three-dimensional and buried
information about waste.
Keywords: Construction solid waste; Millimeter wave radar; Electromagnetic scattering;
Convolutional neural networks

Manoj Samal, K. Lakshmi Prasanna, Mithun Mondal,


Radial deformation detection and localization in transformer windings: A terminal measured
impedance approach,
e-Prime - Advances in Electrical Engineering, Electronics and Energy,
Volume 11,
2025,
100945,
ISSN 2772-6711,
[Link]
([Link]
Abstract: Radial deformations (RD) in transformer windings pose a significant threat to their
reliability, potentially leading to increased losses, reduced lifespan, and catastrophic failures.
Existing methods for detecting and localizing RD often rely on complex techniques such as
model fitting, circuit synthesis, and individual disc access, limiting their practicality and
generalizability. This paper introduces a novel, non-invasive approach for diagnosing RD
faults in transformer windings using terminal impedance measurements. By analysing the
winding’s frequency response and directly calculating ladder network model parameters
from these measurements, the proposed method effectively detects and assesses the
severity of RD faults based on capacitance changes. A robust eigenvalue-based technique is
developed to accurately pinpoint the fault location within the winding. The proposed
method is rigorously validated through circuit and Finite Element Method simulations, as
well as experimental studies on a high-voltage transformer, demonstrating high accuracy in
diagnosing RD faults across various scenarios, including single- and multiple-disc
displacements, non-uniform winding configurations, and varying severities of buckling. This
work provides a valuable tool for condition monitoring and predictive maintenance of power
transformers, enabling early fault detection and mitigation to enhance operational reliability.
Keywords: Driving point impedance; Frequency response analysis; Ladder network;
Parameter estimation; Subspace identification; Terminal measurements; Transformer
winding; Radial deformation

Yash Soni, Malhaar Goswami, Nishit Prabhakar Shetty, Dhiraj,


Millimeter-wave radar for intelligent sensing: A comprehensive review of techniques,
applications, and challenges,
Computers and Electrical Engineering,
Volume 128, Part A,
2025,
110696,
ISSN 0045-7906,
[Link]
([Link]
Abstract: Millimeter-wave (mmWave) radar sensing has established itself as a robust
technology across diverse applications, such as automotive, healthcare, security, and smart
homes. Its exceptional capacity to function effectively in varying environmental conditions,
detect concealed objects, sense physiological signals, and facilitate precise target detection
positions it as a pivotal enabler for next-generation sensing solutions. The survey employs
bibliometric analysis to critically evaluate the existing literature surrounding mmWave radar,
highlighting key research trends, notable publications, and the challenges faced within the
field. This work presents a comprehensive examination of mmWave radar-based sensing,
detailing its fundamental operating principles, signal processing methodologies,
advancements in hardware, and the latest developments in machine learning applications. It
also addreses the key challenges in signal processing, including resolution enhancement,
environmental adaptability, and data fusion with complementary sensors such as LiDAR and
cameras. Furthermore, explored the potential of deep learning techniques to enhance target
classification, activity recognition, gesture identification, and healthcare applications while
addressing concerns related to accuracy and precision. This survey also sheds light on
emerging trends by assessing the strengths, limitations, and prospects of mmWave radar
technology. This review aims to provide insightful guidance for researchers and practitioners
committed to advancing radar-based sensing and its real-world implementations.
Keywords: mmWave radar; FMCW radar; Wireless sensing; Machine learning; Deep learning;
Application taxonomy

Yukai Kong, Xianxiang Yu, Jiachen Li, Kui Xiong, Guolong Cui,
Non-uniform pulse intervals based intra-pulse forwarding jamming detection and
recognition in clutter circumstance,
Signal Processing,
Volume 238,
2026,
110193,
ISSN 0165-1684,
[Link]
([Link]
Abstract: The detection and identification of jamming is the prerequisite and key to the
implementation of anti-jamming measures in radar. In the target detection scenario of
airborne radar, strong clutter causes great difficulty in the detection and identification of
intra-pulse forwarding jamming. This paper proposes a jamming detection and recognition
method based on non-uniform pulse interval coupled with encoder–decoder network.
Specifically, the emission mechanism with non-uniform pulse interval is utilized to disrupt
the echo order of clutter and target, which ensure that only can the jamming gain full
coherent accumulation gain. Subsequently, the jamming signal is recovered using pulse
selection and inverse Fourier transform. Eventually, the combination of multiple loss
functions based-encoder–decoder network is utilized to learn both useful information from
the labels and valid semantic information from the time-frequency feature of the recovered
jamming signal. This can improve the accuracy of jamming recognition. The experimental
results shows that the proposed algorithm achieves more than 90% jamming detection
accuracy and over 94% jamming identification accuracy at JCNR>-15 dB even under the
limitation of insufficient training data.
Keywords: Intra-pulse forwarding jamming; Clutter; Non-uniform pulse interval; Encoder–
decoder network; Combination of multiple loss functions

Pardhu Thottempudi, Vijay Kumar, Rajkishor Kumar,


Dynamic multi-modal attention network for robust and real-time through-wall human
activity recognition,
Results in Engineering,
Volume 28,
2025,
107632,
ISSN 2590-1230,
[Link]
([Link]
Abstract: Through-wall human activity recognition (TW-HAR) has emerged as a critical area
of research due to its applications in healthcare, surveillance, and emergency response.
Conventional approaches relying on single-modality data, such as radar or WiFi, often face
challenges in complex environments, including noise, variability in sensor placement, and
environmental obstructions. These limitations are further exacerbated by factors such as
signal attenuation and scattering caused by diverse wall materials (e.g., concrete, brick,
drywall), misalignment between sensors and human subjects, and dynamic noise conditions,
all of which significantly degrade recognition performance. This paper presents a novel
Dynamic Multi-Modal Attention Network (DMAN) that integrates data from Radar, WiFi, and
Acoustic sensors to achieve robust and accurate human activity recognition. The proposed
framework employs a hybrid Convolutional Neural Network (CNN) and Bidirectional Long
Short-Term Memory (BiLSTM) architecture, which effectively captures spatial and temporal
features from multi-modal data. A dynamic attention mechanism is incorporated to prioritize
critical modality-specific features, mitigating the effects of noise and redundancy.
Comprehensive evaluations based on standard metrics—including accuracy, precision, recall,
and F1-score—demonstrate that the proposed DMAN significantly outperforms state-of-the-
art methods. The system achieved an average accuracy of 96.9% across six distinct activity
classes: walking, running, sitting, standing, falling, and empty room scenarios. Furthermore,
the system maintains high robustness under challenging conditions such as varying wall
materials and sensor misalignments, with a low inference time of 2.8 seconds per sample,
making it suitable for real-time applications. This work establishes the DMAN as a scalable
and reliable solution for TW-HAR, addressing key limitations of existing methods. Future
research directions include exploring additional sensor modalities and enhancing
computational efficiency for broader deployment in smart environments and real-time
monitoring scenarios.
Keywords: Through-wall human activity recognition (TW-HAR); Dynamic multi-modal
attention network (DMAN); Hybrid deep learning architecture (CNN + BiLSTM); Impulse
radio ultra-wideband (IR-UWB) radar; Multi-modal data fusion; Real-time activity
classification

Yipu Yang, Fan Yang, Liguo Sun, Ti Xiang, Pin Lv,


Multi-target association algorithm of AIS-radar tracks using graph matching-based deep
neural network,
Ocean Engineering,
Volume 266, Part 1,
2022,
112208,
ISSN 0029-8018,
[Link]
([Link]
Abstract: Automatic Identification System(AIS) and radar track association is a challenging
subject in dense scenes in which there are some undesirable factors, such as multiple
targets, complicated target movement patterns, and asynchronous track information,
causing inaccurate and inefficient track correlation. Therefore, this research focuses on the
optimization problem of AIS and radar track association in dense scenes. Time-series data of
tracks are transformed into the distribution features in a graph, which is free from the close
dependence of the traditional algorithm on the pre-processing of the time alignment. To this
end, an end-to-end deep network pipeline based on graph matching is proposed to
overcome the influence of the above factors. It involves a multiscale point-level feature
extractor to embed local features. Meanwhile, we devise a cluster-level graph neural
network(GNN) with self-cross attention, which can look for global cues that help us
disambiguate the correct correlation from complex tracks. Graph matching is estimated by
tackling a differentiable optimal transport problem, which minimizes the transport cost and
then achieves global optimal track association. Experiments show that the proposed method
outperforms other approaches and achieves an ideal score(the precision rate and the recall
rate are 0.941 and 0.91, respectively) in our built dataset.
Keywords: Automatic identification system (AIS); Radar track association; Graph matching;
Graph neural network; Optimal transport

Ting Dai, Liye Mei, Yue Zhang, Biao Tian, Rui Guo, Teng Wang, Shan Du, Shiyou Xu,
UAVs and birds classification using robust coordinate attention synergy residual split-
attention network based on micro-Doppler signature measurement by using L-band staring
radar,
Measurement,
Volume 222,
2023,
113692,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Developing unmanned aerial vehicles (UAVs) and birds surveillance technologies to
produce accurate descriptions and achieve high classification accuracy is critical in the field
of radar automatic target recognition (RATR). This article proposes a grayscale spectrogram
image-based UAVs and birds classification method using a robust coordinate attention
synergy residual Split-Attention network (RCA-ResNeSt) under the holographic staring radar
system. Specifically, the ResNet structure with Split-Attention is used as an m-D feature
extractor. The CrossNorm and SelfNorm (CNSN) mechanism is then incorporated into the
network to advance generalization robustness. After that, to consider the spatial direction of
the m-D signature, a coordinated attention (CA) mechanism is introduced at the tail end of
the network to enable fine-grained mining of potential m-D features. Experiments are
carried out using a designed radar system. The results show the superiority of the proposed
method over existing approaches in classification accuracy and noise robustness.
Keywords: Micro-doppler (m-D) signature; Time–frequency representation (TFR); Robust
coordinate attention synergy residual split-attention network (RCA-resNeSt); Holographic
staring radar; Radar automatic target recognition (RATR)

Jiachen Li, Jiaxian Hao, Yukai Kong, Xianxiang Yu, Zhaoyin Xiang, Guolong Cui, Wenmin Wang,
Few-shot jamming recognition based on NMF combined with multi-dimensional fusion
network,
Signal Processing,
Volume 237,
2025,
110089,
ISSN 0165-1684,
[Link]
([Link]
Abstract: Accurately identifying specific types of active jamming is essential for optimizing
radar resources and enhancing anti-jamming efficiency, particularly in the context of few-
shot sample sizes, as discussed in this study. We first employ non-negative matrix
factorization (NMF) to pre-process the radar signal. NMF enhances the feature
representation of data while simultaneously augmenting the sample size. Subsequently, we
propose a multi-dimensional fusion network (MDFN) designed to integrate high-dimensional
features and classify jamming signals effectively. The proposed method demonstrates
superior performance compared to existing approaches across twelve categories of jamming
in few-shot scenario. Experimental results are presented to validate the reliability and
effectiveness of the proposed method.
Keywords: Active jamming recognition; Non-negative matrix factorization (NMF); Multi-
dimensional fusion network (MDFN); Efficient channel attention (ECA); Few-shot samples

Wei Quan, Wenjing Cheng, Yike Yang, Haiquan Zhao, Zhaoyu Chen, Yunfan Luo,
A signal fingerprint feature extraction method based on decomposition and fusion for radar
emitter individual identification,
Digital Signal Processing,
Volume 164,
2025,
105257,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Radar emitter individual identification is one of the key technologies of modern
electronic countermeasure reconnaissance and electronic intelligence. With the
advancement of radar technology and the increasingly complex electromagnetic
environment, existing methods for identifying emitter are gradually becoming unable to
meet the performance requirements of modern radar individual identification. Aiming at
improving the adaptability of feature extraction for non-cooperative radar emitter signals
and the robustness of individual identification in the complex modern electronic warfare
environment, a signal fingerprint feature extraction method based on decomposition and
fusion is proposed. It firstly integrates signal decomposition and scattering convolution
networks (SCN) to adaptively extract the multi-scale intra-pulse feature of the signal, while
removing the potential noise of the redundant component by energy proportion. And then a
deep feature fusion model based on multi-head self-attention and residual connection is
proposed to fuse the multi-scale features and the time domain features to further extract
signal fingerprint of radar emitter. Experimental results based on the real radar emitter
signals demonstrate that the identification method proposed in this paper can more
effectively extract signal fingerprint features and the identification accuracy reaches 96.45%,
which outperforms other existing identification methods.
Keywords: Radar emitter individual identification; Signal fingerprint feature; Signal
decomposition; Scattering convolution networks (SCN); Fusion

Xuezhong Wang,
Electronic radar signal recognition based on wavelet transform and convolution neural
network,
Alexandria Engineering Journal,
Volume 61, Issue 5,
2022,
Pages 3559-3569,
ISSN 1110-0168,
[Link]
([Link]
Abstract: With the continuous use of various new radar systems and complex radar systems,
the electromagnetic environment is extremely deteriorated. The traditional emitter
recognition methods have been difficult to meet the requirements of recognition
performance in the rapidly changing battlefield environment. To solve this problem, a radar
electronic signal recognition algorithm based on wavelet transform and deep learning is
proposed in this paper. Starting from the radar reconnaissance system, the causes of signal
preprocessing are analyzed, and the methods of signal denoising, signal normalization, signal
intra pulse modulation recognition, multipath signal detection and suppression are deeply
studied. In particular, the denoising algorithm based on threshold wavelet transform is
proposed, which significantly improves the reliability of the algorithm. Aiming at the
individual feature extraction of emitter signal, the extraction methods of emitter signal time
domain feature, frequency domain feature, fuzzy function slice feature and cyclic spectrum
feature based on wavelet transform are studied and analyzed, which provides stable and
reliable classification features for emitter signal recognition. According to the characteristics
of radar emitter signal, an optimized convolution neural network is designed, and the
feature fusion processing is carried out at the decision-making level, which greatly improves
the recognition effect and enhances the robustness of the recognition system. Experiments
on radar data show that the fusion recognition rate is higher than any single feature and has
strong robustness. In addition, compared with the traditional SVM and elm networks, the
CNN network proposed in this paper can extract detailed features more effectively and
improve the recognition rate of electronic radar.
Keywords: Radar signal recognition; Wavelet transform; Convolution neural network;
Feature fusion; Noise reduction

Qi Liu, Xinyu Zhang, Yongxiang Liu,


SCNet: Scattering center neural network for radar target recognition with incomplete target-
aspects,
Signal Processing,
Volume 219,
2024,
109409,
ISSN 0165-1684,
[Link]
([Link]
Abstract: Most of the previous radar automatic target recognition (RATR) methods based on
high resolution range profile (HRRP) are designed under the assumption of complete target-
aspects, which assumes the HRRP samples with different target-aspects are complete in
training dataset or template library. Few works were concentrated on HRRP RATR with
incomplete target-aspects. However, it is extremely difficult and sometimes even impossible
to obtain HRRPs with complete target-aspects in real world applications. Therefore, it is
required to recognize targets of unseen target-aspects with an incomplete target-aspect
template library. Aiming at this problem, a scattering center neural network (SCNet) is
proposed for radar HRRP target recognition with incomplete target-aspects. Based on the
assumption that an HRRP sample can be represented by a linear combination of a few atoms
from a scattering center dictionary, we proposed a scattering center layer (SC-layer), which
encapsulates the dictionary coding into an implicit layer and treats the scattering center
dictionary as its parameter that is optimized in an end-to-end manner. We further proposed
a discriminative target-aspect frame dictionary by dividing HRRPs of each class into multiple
target-aspect frames and associating frame label information with each atom of it to enforce
discriminability in features. A multiple target-aspect prototype classifier is proposed, which
is more robust to intra-class variations and thus more suitable to handle the incomplete
target-aspects problem compared with softmax classifier. The classification is simply
implemented by matching HRRP features with each prototype and each HRRP is assigned to
the class with the nearest prototype. In order to integrate both intra-class compactness and
inter-class separation, we proposed a novel discriminative prototype loss. Taking the
advantage of the discriminative prototype loss, HRRPs belonging to the same class and the
same target-aspect frame are pulled closer to the corresponding prototype in the feature
space, while simultaneously pushing apart from other prototypes of different classes.
Experiments on the aircraft electromagnetic simulation dataset and the measured dataset
demonstrated the superior performance of the proposed method compared with other
HRRP-based RTAR methods under the condition of incomplete target-aspects.
Keywords: Incomplete target-aspects; Scattering center; Radar automatic target recognition
(RATR); High resolution range profile (HRRP); Dictionary learning; Neural network
Ayesha Ibrahim, Muhammad Zakir Khan, Muhammad Imran, Hadi Larijani, Qammer H.
Abbasi, Muhammad Usman,
RadSpecFusion: Dynamic attention weighting for multi-radar human activity recognition,
Internet of Things,
Volume 33,
2025,
101682,
ISSN 2542-6605,
[Link]
([Link]
Abstract: This paper presents RadSpecFusion, a novel dynamic attention-based fusion
architecture for multi-radar human activity recognition (HAR). Our method learns activity-
specific importance weights for each radar modality (24 GHz, 77 GHz, and Xethru sensors).
Unlike existing concatenation or averaging approaches, our method dynamically adapts
radar contributions based on motion characteristics. This addresses cross-frequency
generalization challenges, where transfer learning methods achieve only 11%–34% accuracy.
Using the CI4R dataset with spectrograms from 11 activities, our approach achieves 99.21%
accuracy, representing a 15.8% improvement over existing fusion methods (83.4%). This
demonstrates that different radar frequencies capture complementary information about
human motion. Ablation studies show that while the three-radar system optimizes
performance, dual-radar combinations achieve comparable accuracy (24GHz+77GHz: 96.1%,
24GHz+Xethru: 95.8%, 77GHz+Xethru: 97.2%), enabling flexible deployment for resource-
constrained applications. The attention mechanism reveals interpretable patterns: 77 GHz
radar receives higher weights for fine movements (superior Doppler resolution), while 24
GHz dominates gross body movements (better range resolution). The system maintains
71.4% accuracy at 10 dB SNR, demonstrating environmental robustness. This research
establishes a new paradigm for multimodal radar fusion, moving from cross-frequency
transfer learning to adaptive fusion with implications for healthcare monitoring, smart
environments, and security applications.
Keywords: Human activity recognition; Multi-modal fusion; Attention mechanisms; Cross-
frequency transfer learning

Purabi Sharma, Kandarpa Kumar Sarma,


Attention driven CWT-deep learning approach for discrimination of Radar PRI modulation,
Physical Communication,
Volume 62,
2024,
102237,
ISSN 1874-4907,
[Link]
([Link]
Abstract: With the proliferation of radio frequency (RF) systems and radar applications,
Electronic Warfare (EW) is receiving increasing importance. The analysis of the radar signals
is a critical EW task that decides the nature of counter employments. In an Electronic
Support (ES) system, the challenge is to detect hostile radiation sources efficiently and
trigger a counter response. Detection of types of Pulse Repetition Interval (PRI) modulation
of radar signal significantly facilitates the manifestation of RF emitters during recognition
which is difficult in a dense EW environment. Recent developments in artificial intelligence
(AI) methods suggest that this emerging technology can be effective for such purposes. In
this direction, an automatic approach for recognizing several kinds of complex PRI
modulation based on Continuous Wavelet Transform (CWT) and a combination of the vanilla
Convolutional Neural Network (CNN), a multi-head self-attention (MHSA) mechanism and
the popular Long Short-Term Memory (LSTM) is proposed. The CWT is used to decompose
the PRI modulation sequence and obtain different time–frequency components. Further,
aided by the proposed CNN-MHSA-LSTM combination, the features extracted from the CWT
2D-scalograms are used to execute PRI modulation discrimination. In this method, the
vanilla CNN is employed for the extraction of deep features to figure out the class details
while capturing the spatial attributes. Thereafter, to improve the discriminative power of the
entire framework a MHSA mechanism is used. The temporal attributes are acquired by the
LSTM which works in concert with the CNN for executing the detection of the PRI classes
based on the extracted features. Also to assess the effectiveness of the proposed method,
three models based on ResNet, popular CNN and SqueezeNet are implemented for
benchmark comparison in terms of overall performance and complexity. The simulation
results show that the proposed method enhances performance and achieves robustness in
the noise-filled and imperfect channel knowledge environment. The best recognition
accuracy is 98.3% with 50% spurious pulses in the environment which fluctuates with
imperfect channel knowledge cases.
Keywords: Electronic warfare; PRI modulation; CWT; Convolutional Neural Network; Self-
attention mechanism; Long Short Term Memory

Asim Saleem, Guoyun Lv, Safa Hussein Mohammed,


Deep Learning-Based Radar Fingerprinting for Open-Set Generalization Using Dynamic
Thresholding and Embedding Rejection,
Knowledge-Based Systems,
Volume 333,
2026,
115047,
ISSN 0950-7051,
[Link]
([Link]
Abstract: This paper presents a comprehensive framework for radar-specific emitter
identification (SEI), starting with the simulation of a large-scale radar signal dataset designed
to mimic real hardware impairments. By incorporating diverse distortions–such as phase
noise, frequency jitter, amplitude nonlinearity, and multipath reflections–for multiple radar
types and signal-to-noise ratio (SNR) conditions, we generate a realistic and challenging
dataset, SimRF-14, suitable for learning-based signal analysis. We utilized this dataset to
develop RAFNet, a hybrid deep learning model specifically designed for closed-set and open-
set radar emitter classification. The proposed architecture combines convolutional,
recurrent, and attention-based components to capture spatial, temporal, and contextual
features from normalized I/Q waveforms. For open-set recognition, we integrate the
OpenMax algorithm enhanced with Extreme Value Theory (EVT), where class-wise Weibull
modeling of embedding distances enables outlier detection. In addition, an SNR-adaptive
thresholding mechanism improves open-set reliability under varying noise conditions. The
proposed method achieved 97.43% unknown rejection at -20 dB SNR and maintained low
false positives (<4%) at high SNRs, validating its effectiveness and reliability for practical SEI
scenarios.
Keywords: Radar Specific Emitter Identification; Open-set Recognition; RF Fingerprinting;
SNR-Adaptive Classification; Deep Learning for SEI

Yanping Liao, Xinyang Wang, Fan Jiang,


LPI radar waveform recognition based on semi-supervised model all mean teacher,
Digital Signal Processing,
Volume 151,
2024,
104568,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Low probability of intercept (LPI) radar signal identification plays an important role
in electronic warfare, but most existing algorithms are proposed under the condition of
sufficient samples, ignoring the problem of a small amount of labeled data in the actual
electromagnetic environment. To solve the problem, in this paper, a semi-supervised
learning model All Mean Teacher (AMT) based on Mean Teacher (MT) is proposed. First, the
LPI radar signal is transformed into Time-frequency images (TFIs) by using the Choi-Williams
distribution, and Random Erasing is used for TFIs which improves the generalization ability of
the model. Then the Multi-headed Self-Attention Network (MSA-Net) is aimed to extract
features, combined with AMT to realize the automatic waveform recognition of radar
signals. MSA-Net facilitates feature information propagation by computing contrast costs on
TFIs between the student and teacher networks. It solves the problem that TFIs are not easy
to train for small amounts of labeled data, improving the accuracy of signal recognition in
semi-supervised learning scenarios. Experimental results show that the average recognition
accuracy of the proposed method is up to 85.7% at a signal-to-noise ratio of -8 dB.
Keywords: LPI radar signals; Time-frequency analysis; Mean teacher; Self-attention
mechanism

Lingsheng Li, Weiqing Bai, Chong Han,


Multiscale Feature Fusion for Gesture Recognition Using Commodity Millimeter-Wave Radar,
Computers, Materials and Continua,
Volume 81, Issue 1,
2024,
Pages 1613-1640,
ISSN 1546-2218,
[Link]
([Link]
Abstract: Gestures are one of the most natural and intuitive approach for human-computer
interaction. Compared with traditional camera-based or wearable sensors-based solutions,
gesture recognition using the millimeter wave radar has attracted growing attention for its
characteristics of contact-free, privacy-preserving and less environment-dependence.
Although there have been many recent studies on hand gesture recognition, the existing
hand gesture recognition methods still have recognition accuracy and generalization ability
shortcomings in short-range applications. In this paper, we present a hand gesture
recognition method named multiscale feature fusion (MSFF) to accurately identify micro
hand gestures. In MSFF, not only the overall action recognition of the palm but also the
subtle movements of the fingers are taken into account. Specifically, we adopt hand gesture
multiangle Doppler-time and gesture trajectory range-angle map multi-feature fusion to
comprehensively extract hand gesture features and fuse high-level deep neural networks to
make it pay more attention to subtle finger movements. We evaluate the proposed method
using data collected from 10 users and our proposed solution achieves an average
recognition accuracy of 99.7%. Extensive experiments on a public mmWave gesture dataset
demonstrate the superior effectiveness of the proposed system.
Keywords: Gesture recognition; millimeter-wave (mmWave) radar; radio frequency (RF)
sensing; human-computer interaction; multiscale feature fusion

Teng Huang, Yongfeng Chen, Bingjian Yao, Bifen Yang, Xianmin Wang, Ya Li,
Adversarial attacks on deep-learning-based radar range profile target recognition,
Information Sciences,
Volume 531,
2020,
Pages 159-176,
ISSN 0020-0255,
[Link]
([Link]
Abstract: Target recognition based on a high-resolution range profile (HRRP) has always been
a research hotspot in the radar signal interpretation field. Deep learning has been an
important method for HRRP target recognition. However, recent research has shown that
optical image target recognition methods based on deep learning are vulnerable to
adversarial samples. Whether HRRP target recognition methods based on deep learning can
be attacked remains an open question. In this paper, four methods of generating adversarial
perturbations are proposed. Algorithm 1 generates the nontargeted fine-grained
perturbation based on the binary search method. Algorithm 2 generates the targeted fine-
grained perturbation based on the multiple-iteration method. Algorithm 3 generates the
nontargeted universal adversarial perturbation (UAP) based on aggregating some fine-
grained perturbations. Algorithm 4 generates the targeted universal perturbation based on
scaling one fine-grained perturbation. These perturbations are used to generate adversarial
samples to attack HRRP target recognition methods based on deep learning under white-box
and black-box attacks. The experiments are conducted with actual radar data and show that
the HRRP adversarial samples have certain aggressiveness. Therefore, HRRP target
recognition methods based on deep learning have potential security risks.
Keywords: Adversarial attacks; Deep neural networks; Radar images; Target recognition

Di Wu, Shuwen Xu, Hui Liu, Pengcheng Guo,


Sample unbalanced HRRP ground target recognition based on improved Lightgbm,
Digital Signal Processing,
Volume 168, Part D,
2026,
105624,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Radar-based ground target recognition faces significant challenges, including
complex terrain, diverse target types, high recognition difficulty, and low accuracy.
Moreover, the non-cooperative nature of military targets limits access to comprehensive
target data, leading to sample imbalances that further degrade recognition performance.
Addressing these issues, this paper proposes a target recognition method based on
LightGBM, which balances model complexity and recognition accuracy. This method
integrates a weighted focal loss function with dual-stage ground clutter suppression and
enhancement techniques. Initially, during the data preprocessing phase, spherical
hypothesis clustering, coupled with the local outlier factor algorithm, is utilized to mitigate
ground target clutter. Subsequently, in the training phase for target recognition, the weights
of imbalanced samples are dynamically adjusted to augment the model's learning capacity
and heighten its focus on challenging targets. This approach dynamically adjusts the weights
of imbalanced samples, thereby enhancing the model's learning ability and increasing its
attention to difficult-to-classify instances. Additionally, to better accommodate complex
backgrounds and bolster the model's robustness, an adaptive weighting coefficient
adjustment mechanism is incorporated. Ultimately, ground targets are identified using a
LightGBM multi-classifier. Simulations based on actual radar seeker data have validated the
effectiveness of this method, and the recognition performance for six distinct target types
has been evaluated. Comparative analyses with other classifiers demonstrate that this
method exhibits superior performance in ground target recognition under conditions of
imbalanced samples.
Keywords: High resolution range profile; Lightgbm; Sample imbalance; Target recognition

Shuai Guo, Ting Chen, Penghui Wang, Jun Ding, Junkun Yan, Hongwei Liu,
Knowledge embedding fusion based on language model for enhanced radar target
recognition,
Signal Processing,
Volume 238,
2026,
110199,
ISSN 0165-1684,
[Link]
([Link]
Abstract: Traditional radar target recognition methods typically model only single echoes,
neglecting the crucial information that domain knowledge can provide for understanding
data. In this paper, we propose a knowledge embedding fusion (KEF) method for enhanced
high-resolution range profile (HRRP) recognition, which utilizes the target state descriptions
available during radar detection. KEF leverages a language model (LM) to integrate textual
knowledge with echo features for fusion recognition. It consists of three components: HRRP
feature extraction, measurement-based knowledge construction, and knowledge embedding
fusion module. First, we perform feature extraction on the HRRP to obtain echo tokens.
Next, in the knowledge construction module, the measurement statuses are standardized to
a natural language format, and the LM is utilized to extract semantic information, resulting in
text tokens. Finally, in the knowledge embedding fusion module, a cross-attention HRRP-text
fusion strategy is employed to facilitate interaction between echo tokens and textual tokens.
We also design a combination of HRRP-text matching loss and fusion classification loss to
guide model training. Experiments are conducted on a real measured dataset, and the
results indicate that KEF effectively enhances recognition performance across multiple
scenarios compared with approaches that only utilize echoes.
Keywords: High-resolution range profile (HRRP); Knowledge embedding fusion; Language
model (LM); Radar target recognition

Yaqin Zhao, Yuchen Liu, Qi Wang, Rongqian Yang, Longwen Wu,


A radar signal sorting algorithm based on intra-pulse multidimensional feature fusion,
AEU - International Journal of Electronics and Communications,
Volume 201,
2025,
155994,
ISSN 1434-8411,
[Link]
([Link]
Abstract: In modern electronic warfare, the increasing density and complexity of radar
signals reveal critical limitations in traditional inter-pulse parameter-based sorting methods,
including batch overlap, pulse leakage, and heightened sensitivity to parameter tolerances.
This paper presents a radar signal sorting algorithm leveraging intra-pulse multidimensional
feature fusion. We utilize variational mode decomposition to extract signal energy entropy
and mode coefficients, apply phase space reconstruction for computing correlation
dimension and Lyapunov exponent, and employ intrinsic time-scale decomposition to derive
sample entropy with correlation coefficients. These six-dimensional features are fused into a
discriminative feature matrix to enhance inter-class separability. An improved density-peak
clustering fuzzy C-means algorithm is proposed, which adaptively determines the cluster
number and initial centers via density-peak clustering and optimizes membership iteration
through fuzzy C-means to address the limitations of traditional clustering algorithms, such as
dependency on prior parameters and error accumulation. Hardware-in-the-loop
experiments demonstrate that the proposed algorithm outperforms most baseline methods
across a wide range of evaluation metrics. It exhibits superior noise robustness under low
signal-to-noise ratio (SNR) conditions. It achieves near-optimal performance under high SNR
conditions at 5 dB and above, with all metrics exceeding 96 %.
Keywords: Radar signal sorting; Intra-pulse features; Software-defined radio platform;
Clustering algorithm; Feature fusion; Hardware-in-the-loop

Xianwen Zhang, Wenying Wang, Xuanxuan Zheng, Yao Wei,


A novel radar target recognition method for open and imbalanced high-resolution range
profile,
Digital Signal Processing,
Volume 118,
2021,
103212,
ISSN 1051-2004,
[Link]
([Link]
Abstract: High-resolution range profile (HRRP) data often has imbalanced and open-ended
distribution in realistic radar target recognition task, requiring the recognition system classify
targets among majority and minority classes and detect unseen targets accurately. In this
paper, we propose a novel algorithm for open and imbalanced HRPP recognition tasks, by
learning from realistic data distribution and optimizing the accuracy over the majority,
minority and open classes. This method maps the HRRP data to the latent feature space,
enhances the feature by sharing knowledge from majority to minority class, and provides a
scalar indicating the familiarity to known classes. Dual-attention is developed to provide
strong discriminative feature representation. An angular penalty is employed in the loss
function to optimize the intra-class similarity and inter-class variability. Experiments on
measured data prove that the proposed algorithm outperforms other existing methods
under both close and open settings with accuracy increased significantly, and more
discriminative representation. This study provides a promising and effective approach for
open and imbalanced HRRP target recognition.
Keywords: High-resolution range profile (HRRP); Radar automatic target recognition (RART);
Imbalanced learning; Open set recognition

Wenxu Zhang, Fosheng Zhang, Zhongkai Zhao, Feiran Liu,


Radar specific emitter identification via the Attention-GRU model,
Digital Signal Processing,
Volume 142,
2023,
104198,
ISSN 1051-2004,
[Link]
([Link]
Abstract: In radar specific emitter identification (SEI), various types of unintentional
modulation on pulse (UMOP) are selected as the features for discriminating between
different radars. Unintentional Phase Modulation on Pulse (UPMOP), a typical type of UMOP,
can provide crucial information for identifying radars. In most radar SEI algorithms,
sacrificing time efficiency for higher accuracy is a common trade-off. This paper proposes a
method to solve this problem by combining denoised UPMOP sequences with an Attention-
based Gated Recurrent Units (Attention-GRU) model, which showed an excellent
performance. Firstly, the cause of UPMOP is analyzed and the phase observation model of
radar emitter signals and mathematical model of UPMOP are given. Then, the least-squares
method is used to eliminate the linear trend of the phase observation model and obtain a
noised estimation of the UPMOP sequences. Thirdly, the uniform B-spline (UBS) curves are
then used to fit the noised estimation, resulting in a denoised and refined UPMOP sequence.
Finally, the Attention-GRU model is employed to extract features from the denoised UPMOP
sequences to identify radar emitters automatically. Results from simulation and measured
data experiments show that the overall recognition rate of the algorithm reaches over 93%
and the algorithm has excellent performance, with high identification accuracy and relatively
low time consumption, even in low signal-to-noise ratio (SNR) conditions.
Keywords: Radar specific emitter identification; Unintentional phase modulation on pulse;
Uniform B-spline curve; Time-series model; Attention mechanism

Zhiyan Lin, Minming Gu, Keyu Pan, Wei-Ping Zhu,


Adaptive temporal convolutional network with multi-head EMA-gated attention for
continuous radar-based human activity recognition,
Biomedical Signal Processing and Control,
Volume 117,
2026,
109667,
ISSN 1746-8094,
[Link]
([Link]
Abstract: Continuous human activity recognition (HAR) using radar signals offers strong
potential for privacy-preserving clinical health monitoring. However, its performance is
limited by challenges such as multi-scale temporal variation, signal noise, and unstable
activity transitions. To address these issues, this study introduces a radar-based HAR
framework with three tailored components. First, an adaptive temporal convolutional
network (ATCN) uses learnable dilation rates and sampling offsets to flexibly capture both
abrupt and periodic motion patterns over time. Second, an exponential moving average
(EMA)-gated attention (EDGA) module integrates linear attention with exponential moving
average smoothing through a dynamic gating mechanism, effectively suppressing noise
while preserving temporal continuity. Third, an attention-guided multi-stage refinement
(AMSR) module refines coarse predictions using global attention-driven residual corrections,
thereby reducing segmentation noise and improving boundary precision. Experiments on a
77 GHz frequency modulated continuous wave (FMCW) radar dataset show that the
proposed model achieves 96.09% accuracy, demonstrating its strong potential for
continuous and unobtrusive activity monitoring in healthcare applications.
Keywords: ATCN; EDGA; AMSR; Continuous HAR; FMCW radar

Xu Si, Hao Wan, Peikun Zhu, Jing Liang,


A micro-Doppler spectrogram denoising algorithm for radar human activity recognition,
Signal Processing,
Volume 221,
2024,
109505,
ISSN 0165-1684,
[Link]
([Link]
Abstract: Radar signal recognition based on micro-Doppler spectrogram has been widely
used in human action recognition tasks. However, in practical application scenarios, radar
signals inevitably have noise, which leads to different degrees of deformation of the
spectrogram graph structure and affects the accuracy of subsequent recognition algorithms.
In this paper, we present “ACFL”, a novel algorithm for micro-Doppler spectrogram denoising,
which aims to reduce the impact of noise on human action recognition. ACFL employs
amplitude–frequency two-dimensional clustering and fuzzy logic clustering selection
mechanism to remove noise elements from the spectrogram. Moreover, to address the issue
of noise leakage or target missing under time-varying noise and action conditions, ACFL
adopts spectrogram segmentation based on short-term Rényi entropy. By dividing the
spectrogram into intervals with different time–frequency distributions, the dynamic
spectrogram denoise over time is achieved. Simulation and measured data experiments
demonstrate that the proposed algorithm not only achieves a higher-quality denoised
spectrogram but also significantly improves the accuracy of human action recognition under
noisy conditions.
Keywords: Human activity recognition; Micro-doppler; Spectrogram denoise; Fuzzy logic

Hao Wan, Xu Si, Peikun Zhu, Jing Liang,


Target recognition via discriminant information and geometrical structure co-learning using
radar sensor network,
Pattern Recognition,
Volume 157,
2025,
110931,
ISSN 0031-3203,
[Link]
([Link]
Abstract: The target recognition system based on radar sensor network (RSN) has recently
been widely studied in radar automatic target recognition (RATR). The system can observe
the target from multiple perspectives to achieve more robust target recognition. However,
the target feature information observed by different radar sensors may be interrelated,
complementary, or even contradictory. Therefore, a feature extractor is needed to capture
information with discriminant consistency in RSN data. In this work, we propose a new RSN
cooperative target recognition method, namely global discriminant information and local
geometrical structure co-learning (GILSC). Specifically, the proposed GILSC performs feature
extraction by projecting global discriminant information and local neighborhood information
into the common feature space. In this feature subspace, the discriminant information
carried by the high-resolution range profile (HRRP) from different radars is uniformized, and
the geometrical structure of the original HRRP is maintained. The experimental results on
the measured HRRP dataset and public simulation dataset obtained by the Air Force
Research Laboratory (AFRL) prove the effectiveness of GILSC. Compared with other RSN
cooperative target recognition methods, GILSC has the highest recognition rate under the
condition of low SNR, and the recognition time is greatly shortened. Compared with a single
radar, the target recognition rate can be fully improved by two or three radars’ cooperative
target recognition. In addition, experiments on feature subspace show that GILSC has better
feature cohesion.
Keywords: High-resolution range profile; Cooperative target recognition; Feature extraction;
Radar sensor network

Tingpei Huang, Rongyu Gao, Haotian Wang, Jianhang Liu, Shibao Li,
mBox: 3D object detection based on millimeter-wave radar,
Measurement,
Volume 246,
2025,
116568,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Millimeter-wave radar is utilized for 3D object detection in autonomous driving
due to its advantage of not being affected by lighting conditions. Previous millimeter-wave
radar object detection algorithms have not sufficiently used the distributional and statistical
properties of sparse point clouds. This paper introduces mBox, a 3D object detection
framework using only millimeter-wave radar. To eliminate unnecessary information, we
propose a background filtering algorithm based on subtraction(BGFS) that matches and
differentiates the point cloud frame by frame. We propose a voting-based algorithm for
generating object centers(CPGV), which expands the valid data to generate more accurate
initial anchor boxes. To utilize feature information across different scales and capture the
structure of each granularity, we propose a multi-scale feature fusion network based on the
attention mechanism(GLFF-Net). We conduct experiments using the Pointillism and Astyx
datasets. The results show that the mBox outperforms the comparative methods in terms of
3D mean average precision(mAP).
Keywords: Object detection; Point cloud; Millimeter-wave radar

Zhuangzhuang Tian, Wei Wang, Fengchuan Wu, Kai Zhou, Shengqi Liu, Huiqiang Zhang,
SAR target recognition based on CNN with 2-D dual-tree complex wavelet transform
decomposition,
Pattern Recognition,
Volume 172, Part C,
2026,
112585,
ISSN 0031-3203,
[Link]
([Link]
Abstract: Synthetic aperture radar (SAR) is an effective imaging and observation sensor that
has been widely applied in both military and civilian fields. Deep learning approaches have
gained prominence in SAR target recognition and received extensive attention. However,
these methods often struggle when data is scarce, leading to insufficient training and
challenges in effective feature extraction. To address this limitation, we propose a less data-
dependent feature extraction framework. Specifically, we introduce the dual-tree complex
wavelet transform (DTCWT) to capture multi-frequency feature details of SAR images,
integrated with convolutional neural network. This approach enables effective extraction of
high- and low-frequency information. By leveraging the characteristics of these frequency
features, low-frequency subbands are used to emphasize the global structural features in
the images, and high-frequency subbands are employed to identify the significance of
different regions in the images. In response to the aforementioned characteristics, we
introduced an attention mechanism to effectively incorporate high-frequency local
information into low-frequency global information, thereby enhancing feature
representation and recognition efficiency. Moreover, we propose an adaptive rotational
convolution, and apply it to the high-frequency feature extraction. The adaptive rotational
convolution can adapt to the directionally selective subbands with a single convolution
kernel. Experiments conducted on the MSTAR and SAR car datasets demonstrate that the
proposed method can achieve better recognition performance with fewer parameters,
especially on small-scale datasets. The ablation study also confirms the effectiveness of the
introduced DTCWT and rotational convolution.
Keywords: Synthetic aperture radar; Target recognition; Dual-tree complex wavelet
transform; Convolutional neural network
Wenxu Zhang, Xian Lei, Zhongkai Zhao, Fuli Sun,
A dual-decision-maker frequency domain cooperative jamming method against multi-
function radar based on PPO,
Digital Signal Processing,
Volume 169,
2026,
105709,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Optimizing jamming strategies is crucial for coping with complex electromagnetic
countermeasures in dynamic spectrum environments, Among them, the frequency agility
characteristic of the multi-function radar enables them to exhibit strong anti-jamming
capabilities by quickly changing the carrier frequency. Aiming at the problem that traditional
electronic countermeasure strategies show insufficient adaptability to this, a dual-decision-
maker collaborative jamming method based on the proximal policy optimization (PPO)
framework is proposed in this paper. The method achieves dynamic adaptive jamming
against radars through a cascaded collaborative mechanism involving the frequency band
decision maker and the bandwidth decision maker. The confrontation scenario between the
radar network and multiple jammers is established, and the penetration confrontation
process is abstracted and modeled as a markov decision process. The simulation results
show that compared with classic reinforcement learning algorithms such as deep Q network,
the proposed dual-decision-maker collaborative jamming method based on PPO exhibits
superior performance in action estimation and jamming success rate. Specifically, the
convergence speed of the average jamming gain is improved by more than 65 % compared
with the sub-optimal algorithm, while its final stable reward is approximately 10 % higher.
Meanwhile, the key performance indicator values fluctuate minimally under different
experimental conditions, showcasing good robustness and scalability, and achieving
intelligent optimization of frequency domain decision-making.
Keywords: Frequency agility; Deep reinforcement learning; Intelligent jamming decision;
Proximal policy optimization

Yong Liu, Chenyang Lu, Liang Li, Xiangchao Meng, Qiuping Jiang, Feng Shao,
Interactive feature fusion for camera-radar-based vehicle segmentation in bird’s-eye view,
Pattern Recognition,
Volume 172, Part D,
2026,
112698,
ISSN 0031-3203,
[Link]
([Link]
Abstract: Vehicle segmentation in Bird’s-Eye View (BEV) is a fundamental task for
autonomous driving, and integrating multi-modal sensory inputs, e.g., cameras and radars,
could enhance perception capability by leveraging their complementary strengths. However,
cross-modal feature fusion raises additional challenges due to the sparse and noisy
characteristics of radar data and the inherent misalignment between radar and camera
features. While existing fusion methods frequently leverage powerful attention mechanisms,
they often overlook the aforementioned heterogeneities and their impact on achieving
consistent, fine-grained alignment across modalities. We introduce the Interactively
Enhanced Camera-Radar Fusion (IECRF) framework, a novel approach that effectively bridges
cross-modal discrepancies in two stages through three new modules: Camera-Radar Feature
Aggregation (CRFA), Multi-Scale Radar Enhancer (MSRE), and Camera-Radar Feature Fusion
(CRFF). Specifically, the CRFA module explicitly models the complementary features of visual
appearance and radar geometry through two attention mechanisms, enabling fine-grained
alignment and interactive enhancement between the two modalities. The MSRE module
further refines radar representations through a modality-specific down- and up-sampling
design, amplifying salient targets while suppressing background noise in sparse radar
features. The aggregated features are then fused using the CRFF module at each stage for
lateral decoding. Extensive evaluations on the nuScenes dataset demonstrate that our IECRF
framework can operate with multiple backbones and configurations, achieving higher
vehicle segmentation accuracy even when using a lightweight EfficientNet backbone, which
is six times faster than the existing state-of-the-art approach equipped with an advanced ViT
backbone. The source code and trained models are available at
[Link]
Keywords: Vehicle segmentation; Bird’s-Eye View (BEV); Radar perception; Visual-radar
fusion; Autonomous driving

Chaofeng Huang, Xiaowo Xu, Fan Fan, Shunjun Wei, Xiaoling Zhang, Dongmei Liu, Min Gu,
A low-SNR-adaptive temporal network with smart mask attention for radar signal
modulation recognition,
Digital Signal Processing,
Volume 168, Part D,
2026,
105640,
ISSN 1051-2004,
[Link]
([Link]
Abstract: The automatic modulation recognition of radar signals is a key technology in
electronic warfare and communication systems. However, traditional handcrafted features
often struggle to achieve high recognition accuracy under low signal-to-noise ratio (SNR)
conditions. With the rapid development of artificial intelligence technologies, deep learning-
based approaches have emerged as a promising alternative for modulation recognition. In
this article, a low-SNR-adaptive network architecture is proposed, which integrates a
bidirectional temporal convolutional network (Bi-TCN) and dual-channel smart mask
attention (DSMA) modules. The DSMA adaptively highlights informative features and
suppresses noise through complementary attention masks, enhancing robustness in low-SNR
conditions. Experimental results demonstrate that the autocorrelation domain outperforms
both time and frequency domains, with recognition accuracy improvements of 13.33 % and
14.71 %, respectively. Compared to state-of-the-art models, the proposed network achieves
63 % accuracy at -20 dB and more than 99 % accuracy at -6 dB, significantly enhancing radar
signal modulation recognition.
Keywords: Modulation recognition; Deep learning; Radar signal analysis,

Futai Liang, Xin Chen, Song He, Zihao Song, Hao Lu,
An Aerial Target Recognition Algorithm Based on Self-Attention and LSTM,
Computers, Materials and Continua,
Volume 81, Issue 1,
2024,
Pages 1101-1121,
ISSN 1546-2218,
[Link]
([Link]
Abstract: In the application of aerial target recognition, on the one hand, the recognition
error produced by the single measurement of the sensor is relatively large due to the impact
of noise. On the other hand, it is difficult to apply machine learning methods to improve the
intelligence and recognition effect due to few or no actual measurement samples. Aiming at
these problems, an aerial target recognition algorithm based on self-attention and Long
Short-Term Memory Network (LSTM) is proposed. LSTM can effectively extract temporal
dependencies. The attention mechanism calculates the weight of each input element and
applies the weight to the hidden state of the LSTM, thereby adjusting the LSTM’s attention
to the input. This combination retains the learning ability of LSTM and introduces the
advantages of the attention mechanism, making the model have stronger feature extraction
ability and adaptability when processing sequence data. In addition, based on the prior
information of the multi-dimensional characteristics of the target, the three-point estimation
method is adopted to simulate an aerial target recognition dataset to train the recognition
model. The experimental results show that the proposed algorithm achieves more than 91%
recognition accuracy, lower false alarm rate and higher robustness compared with the multi-
attribute decision-making (MADM) based on fuzzy numbers.
Keywords: Aerial target recognition; long short-term memory network; self-attention; three-
point estimation

Keyu Pan, Wei-Ping Zhu, Bo Shi,


A multi-stage few-shot framework for extensible radar-based human activity recognition,
Signal Processing,
Volume 239,
2026,
110244,
ISSN 0165-1684,
[Link]
([Link]
Abstract: This paper proposes a novel framework for radar-based indoor human activity
recognition (HAR) using a multi-stage few-shot learning (FSL) paradigm. The core of our
approach lies in the design of a dynamic feature extraction architecture that exploits wavelet
convolution along with depthwise separable convolutions to effectively capture multi-scale
and multi-frequency information from radar signals. We also propose a meta-learning-
inspired mechanism that dynamically adjusts class weights for unseen categories, thereby
enhancing adaptability and recognition accuracy in few-shot scenarios. Extensive
experiments on five benchmark datasets demonstrate consistent performance gains over
state-of-the-art methods, with substantial improvements observed for both seen and
unseen classes. These findings highlight the robustness, scalability, and generalization
capability of our framework, underscoring its potential to advance radar-based HAR in
complex and diverse environments.
Keywords: HAR; FSL; WTConv; Meta-learning; FMCW radar

Runze Hu, Tong Wang, Weijun Huang, Weichen Cui,


A clutter suppression algorithm via holistic attention deep neural network for airborne radar,
Digital Signal Processing,
Volume 168, Part B,
2026,
105506,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Space-time adaptive processing (STAP) methods have favorable clutter suppression
performance under the condition of sufficient independent and identically distributed (i.i.d.)
samples in airborne radar systems. However, in practical situations, airborne radar often
encounters heterogeneous and non-stationary clutter environment, resulting in a shortage
of i.i.d. samples and degradation in the performance of STAP methods. In this article, a novel
STAP method based on holistic attention network is proposed to address these problems. To
start with, based on airborne radar clutter model, the clutter snapshot data considering ideal
and non-ideal cases are simulated for network training. In addition, a network based on
holistic attention mechanism is developed to perform the super-resolution of clutter spatial-
Doppler spectra. Low-resolution clutter spatial-Doppler spectrum inputs (which are
calculated by only a few snapshot data) are converted to high-resolution clutter spatial-
Doppler spectrum outputs by trained holistic attention network (HAN). As a final point, with
high-resolution clutter spectra, clutter plus noise covariance matrices (CNCMs) are obtained
to calculate adaptive weight vectors. This method achieves recovering clutter spatial-
Doppler spectra precisely nearly in real time with only a few samples under ideal and non-
ideal conditions due to the proposed network's exceptional capability for feature extraction
and pattern recognition. In comparison with recent sparse recovery based and convolutional
neural network (CNN) based STAP, the proposed algorithm shows better performance in
both ideal and non-ideal situations in a short execution time. Numerous simulations are
conducted to demonstrate the effectiveness and superiority of the proposed method in
convergence rate, clutter suppression performance and computational efficiency.
Keywords: Clutter suppression; Space-time adaptive processing; Airborne radar; Radar signal
processing; Convolutional neural network

Hongwei Ma, Yi Liao, Chunhui Ren,


Low probability of interception radar overlapping signal modulation recognition based on an
improved you-only-look-once version 8 network,
Engineering Applications of Artificial Intelligence,
Volume 137, Part A,
2024,
109150,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Low probability of interception (LPI) radar is widely used in modern electronic
warfare. With the increased radiation sources, multiple signals will arrive simultaneously.
The traditional feature extraction method has too many features, which brings great trouble
to the subsequent data processing. Most modulation recognition methods based on deep
learning only consider the single signal after preprocessing, and the generalization ability is
weak. This paper proposes a deep learning solution based on an improved you only look
once version 8 (YOLOv8) network with a global attention mechanism (GAM), achieving
recognition accuracy over 98% in a -10 dB signal-to-noise ratio (SNR) scenario to address the
above problems. We improved the performance of the original network, which can be used
in electronic countermeasures.
Keywords: Global attention mechanism; Deep learning; You only look once version 8;
Modulation recognition; Low probability of interception radar aliasing signal

Jongyun Byun, Jaehoon Cha, Jeyan Thiyagalingam, Hyeon-Joon Kim, Changhyun Jun,
Enhancing rainfall prediction accuracy through image fusion of radar and numerical weather
prediction models,
Expert Systems with Applications,
Volume 303,
2026,
130516,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Precipitation is one of the most challenging atmospheric phenomena to predict
due to the complexity involved in solving dynamic and thermodynamic atmospheric
equations. To address this challenge, extensive research has been conducted to enhance the
precision of numerical weather prediction models and radar-based extrapolation data, in
conjunction with the development of various blending techniques. However, traditional
methods have proven insufficient in capturing the diversity and nonlinearity of weather
phenomena. In response to these limitations, this study introduces a novel methodology
that leverages machine-learning based image fusion models to merge radar-based
extrapolation and numerical weather prediction rainfall datasets, thereby enhancing
prediction accuracy. An image fusion model was developed using radar-based extrapolation
data and numerical weather prediction data as input datasets, with radar observation data
utilized as target dataset. To identify the most suitable image fusion model for capturing the
complex patterns of rainfall data, two experiments were conducted: 1) Impact of model
topology, and 2) Effect of model size. A systematic analysis of the model outputs was
performed using eight evaluation metrics categorized under pixel-based metrics, feature-
based metrics, structural similarity metrics, and categorical verification metrics.
Experimental results indicated that image fusion model based on a Residual Network
(ResNet) outperformed other models in terms of model topology. Regarding model size, it
was observed that the performance did not increase proportionally with the number of
residual blocks; the most suitable performance was achieved with a specific number of
residual blocks (Case 5: 8 blocks). Additionally, the metrics compared with radar observation
data indicated that the proposed model delivered superior performance, thus offering a
high-accuracy rainfall prediction methodology.
Keywords: Image fusion; Radar; Numerical weather prediction; Deep learning; Precipitation
Jackson S. Zaunegger, Paul G. Singerman, Ram M. Narayanan, Muralidhar Rangaswamy,
RadarTD: A Radar Text Dataset for multi-parameter optimization,
Natural Language Processing Journal,
Volume 12,
2025,
100178,
ISSN 2949-7191,
[Link]
([Link]
Abstract: This paper introduces the radar text dataset (RadarTD) for technical language
modeling. This dataset is comprised of sentences containing radar parameters, values, and
units determined from published radar literature. Additionally, each statement is assigned a
sentiment, goal priority, and goal direction label. In this work, we show how RadarTD may be
used to train simple Natural Language Processing (NLP) models to identify the attributes of
each sentence listed in RadarTD. Once the NLP models have identified these attributes from
text, we can use this information to develop Language Based Cost Functions (LBCF). Our
study shows that the proposed text classification model achieves a classification accuracy
between 96.7% and 97.8%, while the proposed named entity recognition model achieves an
F1 score of 99.7. These findings suggest that the developed models are capable of achieving
good performance for both text classification and named entity recognition for autonomous
radar applications. We then illustrate an example of how these models could be used with
Language Based Cost Functions to develop multi-parameter radar optimization schemes. We
also provide a method of providing scalarization weights for each parameter, to improve the
results of the optimization process.
Keywords: Text classification; Named entity recognition; Language modeling; Language-
based cost functions; Multi-parameter optimization; Cognitive radar

Changlong Wang, Jiawei Jiang, Chong Han, Hengyi Ren, Lijuan Sun, Jian Guo,
Through-Wall Multihuman Activity Recognition Based on MIMO Radar,
Computers, Materials and Continua,
Volume 83, Issue 3,
2025,
Pages 4537-4550,
ISSN 1546-2218,
[Link]
([Link]
Abstract: Existing through-wall human activity recognition methods often rely on Doppler
information or reflective signal characteristics of the human body. However, static
individuals, lacking prominent motion features, do not generate Doppler information.
Moreover, radar signals experience significant attenuation due to absorption and scattering
effects as they penetrate walls, limiting recognition performance. To address these
challenges, this study proposes a novel through-wall human activity recognition method
based on MIMO radar. Utilizing a MIMO radar operating at 1–2 GHz, we capture activity data
of individuals through walls and process it into range-angle maps to represent activity
features. To tackle the issue of minimal variation in reflection areas caused by static
individuals, a multi-scale activity feature extraction module is designed, capable of extracting
effective features from radar signals across multiple scales. Simultaneously, a temporal
attention mechanism is employed to extract keyframe information from sequential signals,
focusing on critical moments of activity. Furthermore, this study introduces an activity
recognition network based on a Deformable Transformer, which efficiently extracts both
global and local features from radar signals, delivering precise human posture and activity
sequences. In experimental scenarios involving 24 cm-thick brick walls, the proposed
method achieves an impressive 97.1% accuracy in activity recognition classification.
Keywords: MIMO radar; human activity; Transformer; through-wall

Jinyang Xie, Kanghui Zhou, Lei Han, Liang Guan, Maoyu Wang, Yongguang Zheng, Hongjin
Chen, Jiaqi Mao,
Enhancing multi-task learning-based Tornado identification using spatial and temporal
information from weather radar images,
Applied Soft Computing,
Volume 184, Part B,
2025,
113834,
ISSN 1568-4946,
[Link]
([Link]
Abstract: Tornadoes, as dynamic weather phenomena, exhibit unique spatial and temporal
evolution characteristics that reflect their formation and development. Existing tornado
detection algorithms often struggle with high false alarm rates, primarily due to insufficient
capture of temporal correlations in tornado development. As an improvement, we propose a
multi-task tornado identification network with three-dimensional temporal and spatial
information (TS-MTINet). Taking continuous three-frame radar data as input, the Multi-
frame Temporal Interaction Block (MTIB) utilizes multi-head attention to model the dynamic
interaction information between the radar data, thus exploring in-depth the temporal
features during tornado development. Further, we design a Spatial-Temporal Enhancement
Module (STEM), which analyzes the difference information between continuous data to
extract local and global spatial and temporal feature variations about tornadoes. Based on
this architecture, TS-MTINet incorporates a multi-task learning framework to perform
tornado detection and number estimation tasks simultaneously, thus extracting
comprehensive information related to tornadoes. To validate the performance of the
proposed model, we construct the first Chinese tornado identification dataset with fine
radar features. The experimental results show that the proposed method shows significant
advantages in several evaluation metrics, especially in reducing false alarms. In practical case
studies, compared to the traditional TVS method, TS-MTINet achieves an increase in POD of
approximately 30% and a decrease in FAR of about 20% in several typical tornado events.
Particularly in environments with strong interference, TS-MTINet demonstrates higher
detection accuracy, reflecting greater robustness and practical value.
Keywords: Deep learning; Multi-task learning; Tornado identification; Weather radar;
Attention mechanisms

Xuning Wang, Yuan Chen, Fuhao Wang, Kai Zheng, Yi Zong,


mPCT-LSTM: A lightweight human activity recognition model for 3D point clouds in
millimeter-wave radar,
Digital Signal Processing,
Volume 164,
2025,
105263,
ISSN 1051-2004,
[Link]
([Link]
Abstract: The millimeter-wave radar-based human activity recognition technology shows
significant potential in various domains, such as health monitoring, sports analysis, and
smart homes. Traditional methods rely on 2D feature spectrograms, which struggle to
balance low computational complexity with high recognition accuracy. Transforming radar
echoes into 3D point clouds and designing appropriate neural network models can mitigate
this conflict. However, most existing point cloud processing models are intended for dense
point clouds generated by LiDAR or depth sensors, leaving a gap in effective algorithms for
sparse mmWave radar point clouds. To address this need, we propose a lightweight model
for human activity recognition using sparse mmWave radar point clouds: mPCT-LSTM. This
model extracts spatial features from point clouds of human activity through a Point Cloud
Transformer (PCT) module, which consists of an embedding layer and four stacked offset-
attention layers. These spatial features are fed into a Long Short-Term Memory (LSTM)
module to capture temporal relationships between point cloud frames. Experimental results
demonstrate that the mPCT-LSTM model achieves an average recognition accuracy of
97.26% across three public datasets, outperforming the state-of-the-art by 1.41%.
Additionally, the model’s computational complexity is only 0.09 GFLOPS, a reduction of 70%
compared to current solutions.
Keywords: Millimeter-wave radar; Human activity recognition; 3D point cloud; Deep
learning; Long short-term memory (LSTM)

Jiaxiang Zhang, Bo Wang, Xinrui Han, Min Zhao, Zhennan Liang, Xinliang Chen, Quanhua Liu,
A multi-radar emitter sorting and recognition method based on hierarchical clustering and
TFCN,
Digital Signal Processing,
Volume 160,
2025,
105005,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Radar emitter recognition is an important part of the electronic attack and defense
system, and its key task is to sort and identify various information of emitters from mixed
pulse streams. Pulse descriptor word, as an easily obtainable feature, is often used for
sorting and recognition. Aiming at the problems of high complexity and difficulty in handling
intra-class clustering and inter-class aliasing of existing sorting methods, a hierarchical
clustering method based on Kernel Density Estimation-Kullback-Leibler Divergence-Template
Matching (KDE-KLD-TM) is proposed. Firstly, the down-sampled data is used to construct
central clusters, greatly improving the processing speed. Then, with probability theory as the
theoretical support, inter-class clustering on all samples is performed based on maximum
posterior probability. Finally, based on cluster distribution similarity and periodic template
matching, intra-class merging and inter-class deinterleaving are completed. After sorting,
considering the periodic differences in pulse repetition intervals among different types of
emitter and the insufficient attention paid by existing recognition methods to this feature, a
time-frequency convolution network (TFCN) based emitter recognition method is proposed
for the first time in terms of pulse description word. Using time-frequency analysis (TFA) to
extract periodic features and using convolutional neural networks (CNN) for classification,
the one-dimensional sequence classification problem is treated as a two-dimensional image
classification problem. The proposed method is simulated in a typical scenario with intra-
class clustering and inter-class aliasing. The results show that the proposed method can sort
a total of 2.06 million aliased samples composed of 10 classes within 58.61 s, and the
recognition accuracy reaches 96.33%. The comparison with the baseline method proves the
effectiveness and progressiveness of the proposed method. Finally, the sorting and
recognition performance of the proposed algorithm in complex scenarios in the presence of
pulse loss is discussed.
Keywords: Radar emitter sorting and recognition; Hierarchical clustering; Time-frequency
convolution network; Intra-class clustering; Inter-class aliasing

Jian Chen, Lan Du, Guanbo Guo, Linwei Yin, Di Wei,


Target-attentional CNN for Radar Automatic Target Recognition with HRRP,
Signal Processing,
Volume 196,
2022,
108497,
ISSN 0165-1684,
[Link]
([Link]
Abstract: In this paper, a target-attentional convolutional neural network (TACNN) combining
the convolutional neural network (CNN) and attention mechanism is proposed for radar
high-resolution range profile (HRRP) target recognition. The TACNN takes one-dimensional
CNN (1-D CNN) as the feature extractor and has the capability to excavate abundant local
structural features of data. However, the HRRP contains non-target areas, where the
information is useless or even unfavorable. Furthermore, different parts of HRRP target
regions should have differences in contribution to the recognition task. Therefore, it is an
inadvisable approach that treats all local features alike and directly uses them for the
subsequent target recognition, which is adopted by a lot of models, such as the conventional
CNN. To tackle this problem, the TACNN introduces the attention mechanism on the basis of
1-D CNN. In detail, the constructed attention module adaptively assigns a weight to each
local feature of HRRP so as to locate the target areas and meanwhile enhance the interest of
model in valuable target information. Specially, the attention mechanism in TACNN is
realized via a bidirectional gated recurrent unit (Bi-GRU) network, where the attention
coefficients used for weighting up local features are generated with full consideration of
sequential relationship among different regional features in HRRP. Therefore, the learned
attention coefficients in our TACNN can better represent the importance of each local
feature to the recognition task, ultimately beneficial for the discovery of target information
with more discriminability. Experimental results on measured HRRP data show that the
proposed model can get more effectiveness in target recognition than related methods.
Keywords: Radar target recognition; High-resolution range profile (HRRP); One-dimensional
convolutional neural network (1-D CNN); Gated recurrent unit (GRU); attention mechanism

Xiaodan Wang, Rui Li, Jian Wang, Lei Lei, Yafei Song,
One-dimension hierarchical local receptive fields based extreme learning machine for radar
target HRRP recognition,
Neurocomputing,
Volume 418,
2020,
Pages 314-325,
ISSN 0925-2312,
[Link]
([Link]
Abstract: Radar automatic target recognition (RATR) aims at extracting meaningful target
features from the electromagnetic echo signal and utilizing the features to automatically
recognize the target types. The high-resolution range profile (HRRP) plays an important role
in RATR field, HRRP is the amplitude of the echo summation for target scattering centers in
each range cell of wideband radar. Using deep neural networks for HRRP radar target
recognition encounters the problem of storage overhead and slow convergence rate, to
resolve those issues, we propose a one-dimension local receptive fields based extreme
learning auto-encoder (1D ELM-LRF-AE) network for HRRP local structures and meaningful
representations learning. ELM-LRF-AE consists of an input layer, a random convolution layer,
a pooling layer, several local connected layers and an output layer, it reconstructs the input
with a greedy strategy that the input feature vectors are divided into several subgroups and
the i-th pooling feature vector is used to reconstruct the i-th grouping input feature vector.
Then we use the learned pooling feature vectors to replace the random pooling feature
vectors as the learned representations. We also stack several 1D ELM-LRF-AEs to build 1D
hierarchical local receptive fields based extreme learning machine (1D H-ELM-LRF) for high
level HRRP abstract representations learning and recognition. Experimental results on
simulated HRRP data set demonstrate the superior high recognition performance and high
computational efficiency of our algorithm.
Keywords: Local receptive fields based extreme learning machine; Auto-encoder; HRRP
recognition; Deep learning

Guanhong Lu, Lei Kou, Pei Niu, Gaohang Lv, Xiao Zhang, Jian Liu, Quanyi Xie,
GPRTransNet: A deep learning–based ground-penetrating radar translation network,
Tunnelling and Underground Space Technology,
Volume 161,
2025,
106557,
ISSN 0886-7798,
[Link]
([Link]
Abstract: Ground-penetrating radar (GPR) is an essential nondestructive testing tool widely
used in tunnel and road defect detection, underground object detection, and unstructured
terrain perception. Currently, most GPR inversion methods based on deep learning use
convolutional neural networks, which have limitations such as incomplete feature extraction
and low accuracy. Inspired by advancements in the NLP field, this paper proposes a novel
deep learning framework for GPR called GPRTransNet. The algorithm introduces a “wave to
permittivity” translation architecture, leveraging the “memory” function of recurrent neural
networks to translate GPR data into permittivity model images, similar to language
translation. In addition, the inclusion of the attention mechanism significantly enhances the
network’s ability to represent defects in complex scenarios, resulting in outstanding
translation performance. GPRTransNet has been validated on a simulated dataset, with the
newly proposed G-SSIM evaluation metric showing that the permittivity similarity of
GPRTransNet-Lite reaches 98.47% and that of GPRTransNet-Pro is 99.26%. Furthermore,
GPRTransNet demonstrates good translation performance on real data. Experimental results
show that the results of GPRTransNet’s translation of GPR are satisfactorily matched.
Keywords: Ground-Penetrating Radar Data Translation; Recurrent Neural Networks; Deep
Learning; Gated Recurrent Unit

Adil Ali Saleem, Hafeez Ur Rehman Siddiqui, Muhammad Amjad Raza, Sandra Dudley, Julio
César Martínez Espinosa, Luis Alonso Dzul López, Isabel de la Torre Díez,
Ultra Wideband radar-based gait analysis for gender classification using artificial intelligence,
Array,
Volume 27,
2025,
100477,
ISSN 2590-0056,
[Link]
([Link]
Abstract: Gender classification plays a vital role in various applications, particularly in
security and healthcare. While several biometric methods such as facial recognition, voice
analysis, activity monitoring, and gait recognition are commonly used, their accuracy and
reliability often suffer due to challenges like body part occlusion, high computational costs,
and recognition errors. This study investigates gender classification using gait data captured
by Ultra-Wideband radar, offering a non-intrusive and occlusion-resilient alternative to
traditional biometric methods. A dataset comprising 163 participants was collected, and the
radar signals underwent preprocessing, including clutter suppression and peak detection, to
isolate meaningful gait cycles. Spectral features extracted from these cycles were
transformed using a novel integration of Feedforward Artificial Neural Networks and
Random Forests , enhancing discriminative power. Among the models evaluated, the
Random Forest classifier demonstrated superior performance, achieving 94.68% accuracy
and a cross-validation score of 0.93. The study highlights the effectiveness of Ultra-wideband
radar and the proposed transformation framework in advancing robust gender classification.
Keywords: Gait; Ultra-wide band radar; Gender classification; Spectral features; Feed
forward artificial neural network; Ridge classifier; Hist gradient boosting

Jiahao Liu, Yiming Zhang, Liang Song, Zheng Tong,


SCB-ADAE: An attention-based deep autoencoder for ground penetrating radar signal
denoising,
Engineering Applications of Artificial Intelligence,
Volume 160, Part A,
2025,
111902,
ISSN 0952-1976,
[Link]
([Link]
Abstract: In buried object detection, recorded signals of a ground penetrating radar (GPR)
inevitably include noise interference owing to complex underground environments. Existing
rule- and data-driven denoising methods struggle to handle non-Gaussian and real-world
noise because the rule-driven ones rely on the assumptions of simplified noise
characteristics and the data-driven ones cannot capture fine- and global-scale features of a
GPR signal well. To address the problem, this study proposes an attention-based denoising
model called the Swin-Conv Block with Attention Denoising Autoencoder (SCB-ADAE). The
model first feeds a GPR signal into a SCB module, which extracts a tensor with the fine-scale
features in the signal, such as sharp reflective interfaces and abrupt amplitude variations.
The feature tensor then passes through an ADAE module that uses encoder-decoder
structure with the self-attention to enhances the representation of the global-scale signal
features. Finally, the feature tensor from the ADAE module is decoded by another SCB
module to generate a denoised GPR signal, where the tensor includes the fine-scale and
global features of the raw signal. An experiment with three types of GPR signals
demonstrates the effectiveness of the proposed model: radar signals with Gaussian noise,
radar signals with inhomogeneous-material noise, and real-world signals. radar signals with
Gaussian noise, radar signals with inhomogeneous-material noise, and real-world signals.
Experimental results demonstrate that the proposed model outperforms other state-of-the-
art denoising methods on denosing the three types of GPR signals, where the signal-to-noise
ratio, peak signal-to-noise ratio, and structural similarity index are improved to 20.64, 14.59,
and 0.366, respectively.
Keywords: Ground penetrating radar; Denoising; Attention-based model; Autoencoder

Xueqing Zhao, Ren Xu, Yutao Zhang, Andrew Ty Lau, Ruitian Xu, Xingyu Wang, Andrzej
Cichocki, Jing Jin,
A novel paradigm based on radar-like scanning for directional recognition in event-related
potentials based brain-computer interfaces,
Journal of Neuroscience Methods,
Volume 423,
2025,
110546,
ISSN 0165-0270,
[Link]
([Link]
Abstract: Background
Event-related potentials (ERPs) based brain-computer interface (BCI) systems have shown
significant potential for directional control applications. Existing paradigms are constrained
by the limited scalability of directional commands that demand interface reconfiguration for
varying target numbers.
New method
We propose a novel radar-like scanning (RS) paradigm for 32-directional recognition tasks to
address these limitations. This paradigm continuously scans through directions using a
sector-shaped visual stimulus, naturally evoking ERP responses without discrete directional
indicators. During the online experiments, an early-stopping strategy is employed to
enhance efficiency. Additionally, this study analyzes subjects' directional recognition
performance using EEGNet under three sector rotation periods. Thirteen subjects
participated in the experiments.
Results
The grand-averaged ERP amplitudes exhibited a stronger negative deflection in the parietal,
occipital, and temporoparietal regions. The results demonstrated that, with a 2 s rotation
period and early-stopping strategy, the best subject achieved an accuracy of 87.50 % with a
mean absolute angle error of 1.64°. When the directional error tolerance was set to 11.25°,
the subject-averaged accuracy reached 91.83 % under the same conditions. Longer rotation
periods led to better subject-averaged recognition performance. When the rotation period
was short (1 s), targets close to the scanning center were challenging to recognize.
Comparison with existing methods
Compared with others, the RS paradigm enables more fine-grained directional target
recognition and is unaffected by the target numbers.
Conclusions
The proposed paradigm demonstrates significant potential for applications in ERP-BCI
systems.
Keywords: Brain-computer interface (BCI); Electroencephalography (EEG); Event-related
potential (ERP); Radar-like scanning; Directional recognition

Yuwen Wu,
Fusion-based modeling of an intelligent algorithm for enhanced object detection using a
Deep Learning Approach on radar and camera data,
Information Fusion,
Volume 113,
2025,
102647,
ISSN 1566-2535,
[Link]
([Link]
Abstract: Object detection, the process of detecting and classifying objects within a given
environment, forms the foundational element. Multisensory fusion incorporates data from
diverse sensors, like radar and cameras, to refine the reliability and accuracy of detection.
Further, Radar and camera data fusion refine this process by integrating the unique strength
of both technologies, which leverage the radar's proficiency in adverse weather conditions
and the camera's high-resolution imaging. This incorporation enhances the object detection
systems, which enables them to effectively operate across the spectrum of scenarios, from
autonomous vehicles navigating challenging weather to surveillance systems monitoring
critical infrastructure. Deep learning (DL), a branch of machine learning (ML), empowers this
system with the capability to learn complex representations and patterns directly from the
data, which enables them to generalize and adapt to new situations. By integrating the
advanced methodology, we can develop strong perception system capable of interpreting
and detecting objects accurately in dynamic and diverse environments, from autonomous
vehicles navigating urban landscapes to surveillance systems monitoring complex
environments. This study designs an Intelligent Algorithm for Enhanced Object Detection
Using Deep Learning Approach on the Radar and Camera Data Fusion (IAEOD-DLRCDF)
technique. The presented IAEOD-DLRCDF technique uses multi-angle joint calibration where
the spatial sparse alignment of the heterogeneous data of the camera and Radar is realized
with image falsification disregarded. Besides, the IAEOD-DLRCDF technique applies YOLOv8
object detector for radar and camera target detection individually which are then integrated
with the image plane. Moreover, the detected objects are then classified via the
bidirectional long short-term memory (BiLSTM) model. Furthermore, the Adam optimizer is
used for the optimum hyperparameter selection of the BiLSTM network which results in a
better recognition rate. The performance assessment of the IAEOD-DLRCDF method is tested
under benchmark dataset. The empirical analysis stated that the IAEOD-DLRCDF method
gains better performance over other models.
Keywords: Object detection; Deep learning; Data fusion; Radar; YOLOv8; Adam optimizer;
Machine learning

Chenxuan LI, Weigang ZHU, Bakun ZHU, Yonggang LI,


Few-shot incremental radar target recognition framework based on scattering-topology
properties,
Chinese Journal of Aeronautics,
Volume 37, Issue 8,
2024,
Pages 246-260,
ISSN 1000-9361,
[Link]
([Link]
Abstract: The continuous emergence of new targets in open scenarios leads to a substantial
decrease in the performance of Inverse Synthetic Aperture Radar (ISAR) recognition systems.
Also, data scarcity further exacerbates the challenge of identifying new classes of ISAR
targets. In this paper, a few-shot incremental target recognition framework based on
Scattering-Topology Properties (STPIL) is proposed. Specifically, STPIL extracts scattering-
topology properties of ISAR targets as recognition features. Meanwhile, the pseudo-
incremental training strategy effectively alleviates the algorithm’s forgetting of old
knowledge, and improves compatibility with new classes. Besides, a feature embedding
network, with few parameters, is designed based on the graph neural network. This
embedding network is highly adaptable to changes in data distribution. Additionally, STPIL
fully considers the joint distribution and marginal distribution in scattering features, and
uses the Brownian distance metric module to make the scattering-topology features more
discriminative. Experimental results on both the simulation dataset and the public measured
data indicate that STPIL can effectively balance new classes with old classes, and has
superior performance to other advanced methods in the incremental recognition of targets.
Keywords: Brownian distance metric; Graph neural networks; Incremental learning; Inverse
Synthetic Aperture Radar (ISAR); Scattering

Kexin Huang, Yizhen Jia, Wen-Qin Wang,


FDA-MIMO Radar Target Recognition Based on SVM Classification with Multi-channel
Feature Extraction,
Procedia Computer Science,
Volume 221,
2023,
Pages 1513-1518,
ISSN 1877-0509,
[Link]
([Link]
Abstract: FDA-MIMO radar has good anti-active interference performance and has received
extensive attention due to its distance dependence. However, there is still a lack of target
and interference recognition methods for FDA-MIMO radar in complex environments.
Therefore, in this paper, we propose a multi-channel feature extraction classification method
based on the support vector machine (SVM), which realizes the classification of four active
interferences and targets. In short, we classify the fine signal characteristics of different
interference signals by extracting effective features and selecting efficient classifiers.
Simulation results show that when the interference to noise ratio (INR) is greater than -15
dB, the recognition accuracy is greater than 0.95. It also shows that the proposed method
can distinguish target and interference well in a variety of complex environments.
Keywords: Frequency diversity array; FDA-MIMO; feature extraction; multi-channel signal
processing; SVM; target recognition

Sachin Kishanrao Bhingikar, Rishi Raj Sharma, Ram Bilas Pachori,


Radar-based non-contact heart rate monitoring: A comprehensive review,
Digital Signal Processing,
Volume 168, Part D,
2026,
105630,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Heart rate (HR) measurement is an essential physiological parameter used in
various fields, such as healthcare and human-computer interaction. This article presents a
comparative evaluation of contact and non-contact approaches for heart rate monitoring,
with a specific emphasis on radar-based non-contact methods. It begins by examining
conventional contact-based techniques like electrocardiography and photoplethysmography,
outlining their operational principles, applications, and limitations. While contact methods
ensure high accuracy and reliability, they can be inconvenient for long-term monitoring in
ambulatory environments owing to their inconvenience. The discussion then shifts to non-
contact approaches with a focus on radar-based techniques, which offer advantages such as
being non-intrusive, capable of remote operation, and having the potential for continuous
monitoring without physical contact. The paper provides insights into the underlying
principles of radar-based heart rate detection along with its varied applications across
domains like healthcare, automotive systems, and smart environments while also addressing
implementation challenges. Additionally, it discusses future directions pertaining to radar-
based non-contact heart rate detection and concludes by highlighting the strengths and
limitations.
Keywords: Radar sensing; Heart rate (HR) detection; Non-contact sensing; Vital sign
monitoring; Radar signal processing

Haojie Wei, Min Fang, Haixiang Li, Yinan Wang, Zhanpeng Zheng,
Prototype-based dual-alignment of multi-source domain adaptation for radar emitter
recognition,
Signal Processing,
Volume 230,
2025,
109853,
ISSN 0165-1684,
[Link]
([Link]
Abstract: Recent research on radar emitter recognition has primarily focused on single-
source domain adaptation, neglecting the potential benefits of leveraging multiple source
domains. While multi-source domain adaptation (MDA) models have been extensively
developed for image data, they often fail to address the unique challenges presented by
radar emitter data, such as domain class hierarchy discrepancies and intra-class
compactness. To address these challenges, this study proposes a prototype-based dual-
alignment (PBDA) method specifically designed for multi-source radar emitter recognition
tasks. The PBDA method incorporates both domain-level and class-level alignment. For
domain-level alignment, a reconstruction loss is introduced to preserve classification-
relevant features during adversarial learning. For class-level alignment, a prototype
alignment loss is proposed to minimize distributional discrepancies of the same class across
different domains, reducing fine-grained distribution gaps between each source domain and
the target domain. To ensure compact sample distribution within the target domain and
avoid negative transfer effects, information entropy loss and classification consistency loss
are applied, guiding target domain samples toward the correct class prototypes.
Experimental validation on both image benchmark datasets and radar emitter datasets
demonstrates the effectiveness of the proposed PBDA method.
Keywords: Domain adaptation; Radar emitter recognition; Multi-source; Adversarial learning

Wenxu Zhang, Yajie Wang, Xiuming Zhou, Zhongkai Zhao, Feiran Liu,
An interference power allocation method against multi-objective radars based on optimized
proximal policy optimization,
Signal Processing,
Volume 230,
2025,
109785,
ISSN 0165-1684,
[Link]
([Link]
Abstract: Aiming at the problem of interference resource scheduling in cognitive electronic
warfare, a multi-objective interference power allocation method based on the proximal
policy optimization (PPO) framework is proposed in this paper. Firstly, the confrontation
between jammers and multi-objective radar networks is mapped as the interaction between
the agent and the environment, and the radar target detection model under suppression
interference is established. On this basis, an interference power allocation model against
multi-objective radars based on PPO framework is constructed. Moreover, a reward
normalization mechanism is introduced to optimize the reward setting, and an interference
power allocation method based on optimized PPO is proposed. Meanwhile, this paper
constructs a confrontation scenario in which the jammer covers the target aircraft to break
through the multi-objective radar network. Simulation experiments are conducted based on
this scenario to verify the effectiveness of the method proposed in this paper. The
interference power allocation method proposed in this paper can intelligently adjust the
power allocation scheme of the jammer according to the electromagnetic situation on the
battlefield, optimize the resource utilization of the jammer, and occupy the initiative on the
battlefield.
Keywords: Interference power allocation; Proximal policy optimization; Limited interference
resources; Reward normalization

Wei Yin, Ling-Feng Shi, Yifan Shi,


Indoor human action recognition based on millimeter-wave radar micro-doppler signature,
Measurement,
Volume 235,
2024,
114939,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Considering that millimeter-wave radar lacks sufficient data to support Transfer
Learning (TL) and Human Action Recognition (HAR), we propose a Heterogeneous Multi-
source Transfer Learning (HMTL) method and a data selection algorithm based on
Categorical Probability to obtain Hash Coding (CPHC). CPHC is utilized to select data from
multi-source datasets in the same domain for downscaling and matching to ensure the
similarity between the selected data features and the target task features to obtain better
performance on the target task. The experimental results show that the CPHC can
downscale data more efficiently than the traditional algorithm. HMTL can effectively
improve the classification accuracy of the network to the state-of-the-art (SOAT) level. RPME
addresses that the radar Micro-Doppler Signature (MDS) is flooded by noise due to the
cluttered indoor environment, making the MDS more obvious and thus improving the
classification accuracy.
Keywords: Transfer learning; Hash coding; Micro-doppler signatures; Human action
recognition; Data downscaling

Yuankang Ye, Feng Gao, Shaoqing Zhang, Chang Liu,


Improving precipitation nowcasting via multiphysical parameter fusion in radar echo
extrapolation,
Journal of Hydrology,
Volume 668,
2026,
134947,
ISSN 0022-1694,
[Link]
([Link]
Abstract: Radar-based precipitation nowcasting plays a vital role in short-term
hydrometeorological forecasting and water resource management. Existing modeling
methodologies typically simplify precipitation nowcasting to a task of spatiotemporal
sequence prediction based on radar echo reflectivity data. However, the reliance on
unimodal reflectivity data including intensity-only information restricts the model’s ability to
characterize the phase evolution and dynamic processes of hydrometeor particles,
ultimately leading to insufficient extrapolation accuracy. This study breaks through the
conventional unimodal data paradigm, aiming to capture the complex dynamic evolutionary
features of hydrometeor particles. We integrate radar echo reflectivity and four additional
physical parameters of hydrometeor particles into a deep learning framework and propose a
novel Physics-Informed Multimodal Echo Extrapolation neural network (PIEE). Furthermore,
we systematically investigate the individual contributions of each physical parameter to the
accuracy of radar echo extrapolation. Specifically, PIEE adopts a three-stage structure. First, a
multimodal encoder with a dual-branch attention-based fusion strategy is used to capture
diverse physical signals. Second, a novel gated spatiotemporal self-attention module is
designed for deep feature extraction. Finally, the decoding stage generates the extrapolated
radar echoes. Experimental results on a real multimodal radar echo dataset show that the
proposed model demonstrates superior performance in two aspects. First, under a unimodal
baseline architecture, the PIEE model clearly outperforms the comparison model. Second,
after fusing multiple physical parameters, the PIEE achieves significant improvements in all
the evaluated metrics, especially in the CSI and HSS metrics for the high echo intensity
region (≥ 40 dBZ), with improvements of up to 24.2% and 20.3%, respectively. Furthermore,
systematic ablation experiments on physical parameters quantify the effects of different
combination methods on extrapolation accuracy, highlighting the potential of physics-
informed, multimodal deep learning approaches in improving short-term hydrological
prediction accuracy, with implications for flood forecasting, early warning systems, and
hydrometeorological risk management at catchment scales.
Keywords: Deep learning; Precipitation nowcasting; Radar echo extrapolation;
Hydrometeorological forecasting

Jiayi Cai, Zhaocheng Yang, Ping Chu, Juntao Guo, Jianhua Zhou,
Robust hand gesture detection and recognition using 4D millimeter-wave radar in a
ubiquitous scene,
Measurement,
Volume 253, Part C,
2025,
117545,
ISSN 0263-2241,
[Link]
([Link]
Abstract: In current research on HGR using radar sensors, hand gestures are typically
confined to a smaller region. However, in ubiquitous scenarios, unrestricted human body
movements and unexpected hand gesture motions usually occur, which results in a large
false alarms and recognition performance degradation. To address this issue, we propose a
robust hand gesture detection and recognition method in ubiquitous scenarios using
Frequency-Modulated Continuous Wave (FMCW) Multiple-Input Multiple-Output (MIMO)
radar. The core idea is to progressively define and classify motions in a cascaded manner,
gradually filtering out non-specific movements, reducing false positives, and enhancing the
applicability of HGR. Specifically, we first propose a suspected hand gesture motion
detection method to help identify suspicious hand gestures. Then, the velocity and position
features of the mutated signal and the stable signal are extracted. A mutated signal motion
recognition method based on a single-layer long short-term memory (LSTM) network is used
to effectively distinguish non-hand gesture motions from hand gestures. Finally, the two-
dimensional trajectory features are extracted, and cascaded with a LSTM network combined
a Gaussian probability model is developed to enhance the ability of open-set recognition.
Experimental results show that the proposed method can achieve the recognition accuracy
of 99.53% for designed hand gestures, the false alarm rate of 1.5% for unexpected hand
gestures and 0.11% for non-hand gesture motions.
Keywords: Hand gesture recognition; Non-hand gesture motion; Feature extraction;
Probability models; Ubiquitous scene

Leiyao Liao, Lan Du, Jian Chen,


Class factorized complex variational auto-encoder for HRR radar target recognition,
Signal Processing,
Volume 182,
2021,
107932,
ISSN 0165-1684,
[Link]
([Link]
Abstract: In the field of radar automatic target recognition (RATR), the high-resolution range
profile (HRRP) has received intensive attention. Bar a few exceptions, almost all HRRP-based
ATR classification systems ignore the phase of the HRRP when the data is input to the
classifier, relying instead only on the magnitude of the complex HRRP samples. This
approach ignores the phase of the complex HRRPs, which reduces the information in the
signal. In this paper, we develop a novel class factorized complex variational auto-encoder
(CFCVAE) to utilize the phase of the high range resolution (HRR) radar echo for recognition.
The CFCVAE is a complex-valued neural network (CV-NN) consisting of the encoding and
decoding modules. In CFCVAE, the encoding module projects the observed data into the
latent space, and then the latent features are fed to the decoding module, which further
maps the latent features to data. In particular, the decoding module introduces the class
labels to partition the whole observations into some parts, each of which is depicted by a
specific class-decoder. Compared with the traditional variational auto-encoder (VAE)
containing a single decoder, the CFCVAE can give a more accurate description to the whole
dataset via multiple class-decoders, thus improving the characterization ability of features.
In addition, based on the class labels, the CFCVAE employs the conditional prior on the
latent variable to enhance the discrimination of features. Moreover, a complex
backpropagation algorithm is derived for CFCVAE training, and a sample is classified to the
class corresponding to the class-decoder with the minimum reconstruction error in the test
stage. Experimental evaluations on the measured data indicate that the proposed method
indeed achieves very promising target recognition performance.
Keywords: Variational auto-encoder (VAE); Complex-valued neural network (CV-NN); High
range resolution (HRR) radar; Radar automatic target recognition (RATR); Minimum
reconstruction error

Yue XU, Quan PAN, Zengfu WANG, Hua LAN, Shuling JIN,
A self-learning refined model and tracking for near space hypersonic vehicle by space-based
radar,
Chinese Journal of Aeronautics,
2025,
103840,
ISSN 1000-9361,
[Link]
([Link]
Abstract: The Near Space Hypersonic Vehicle (NSHV) features a unique design and
propulsion system, achieving exceptional speed, range, and maneuverability, which
challenge ground-based radars. Space-Based Radar (SBR) offers a breakthrough for tracking
NSHV targets, with all-weather operation and freedom from Earth’s curvature, but faces
complex coordinate transformations. Traditional models often overlook the NSHV’s dynamic
gliding trajectory, especially the impact of hidden control variables on maneuvering, causing
mismatches during rapid motion changes. This paper proposes a refined tracking model
unified in the ECEF coordinate frame, incorporating model parameters that implicitly encode
control laws, and presents an Expectation-Maximization Multi-swarm Cooperative Particle
Swarm Optimization (EM-MCPSO) framework for both NSHV tracking and model parameter
estimation to address this problem. To minimize conversion errors, a transformation matrix
directly represented by the state in the Earth-Centered Earth-Fixed (ECEF) coordinate is
derived. Then the hybrid aerodynamic acceleration coefficients are introduced to precisely
describe the dynamic behaviors, formulating target tracking as a joint estimation problem of
state and parameters within EM framework. Finally, a self-learning algorithm based on a
master-slave structured PSO is proposed to solve the optimization of the conditional
expectations of EM under strong nonlinearity, with a Proportional-Derivative (PD) controller
accelerating convergence, and updating the population structure with historical data.
Simulations of vertical gliding and horizontal maneuvers validate the algorithm’s
effectiveness.
Keywords: Dynamics modeling; Expectation Maximization (EM); Maneuvering target
tracking; Near Space Hypersonic Vehicle (NSHV); Particle Swarm Optimization (PSO)

Dongming Wu, Junpeng Shi, Zhiyuan Zhang, Zhihui Li, Fangling Zeng,
Generative-contrastive learning for open set radar emitter identification,
Signal Processing,
Volume 239,
2026,
110295,
ISSN 0165-1684,
[Link]
([Link]
Abstract: In traditional radar emitter identification (REI) tasks, both the training and testing
samples share the same distribution, and the model is trained solely to recognize known
targets. However, in non-cooperative electromagnetic environments, unknown classes are
often absent from the training data, which may be incorrectly classified as known classes. To
address this issue, we propose an innovative Generative-contrastive Learning method for
Open Set REI (GLOSE) from the perspective of feature space optimization. We first introduce
a conditional generative model derived from diffusion to generate stable interpolated
samples within the feature space, which are defined as an additional class to compress the
coverage of known classes, thereby enhancing the capability to handle unknown space.
Subsequently, we employ contrastive learning with an adaptive contrastive loss to further
optimize the discriminative power of the feature space, which applies varying levels of intra-
class similarity for different types of samples. Extensive experiments are conducted on a
simulated radar emitter dataset based on intra-pulse unintentional modulation and a real-
world automatic dependent surveillance-broadcast (ADS-B) dataset. The results
demonstrate that the proposed method significantly improves the detection capability of
unknown class samples while maintaining high classification accuracy for known classes.
Keywords: Radar emitter identification; Open set recognition; Generative-contrastive
learning; Intra-pulse unintentional modulation

Qihang Zhai, Xiongkui Zhang, Zilin Zhang, Jiabin Liu, Shafei Wang,
Online few-shot learning for multi-function radars mode recognition based on backtracking
contextual prototypical memory,
Digital Signal Processing,
Volume 141,
2023,
104189,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Modern multi-function radars (MFRs) can flexibly generate multiple fine-grained
working modes for different missions through programmable parameters. Precise
recognition of these working modes by analyzing the parameter combinations lays
foundation for situational awareness. This recognition becomes more challenging to
electromagnetic reconnaissance system when prior information of radiation source is
unavailable and effective labeled signal data is not sufficient in adversarial scenarios. The
presence of incremental mode involved in radar data flow requires that the recognition
model enables to dynamic adjust and online perceive current data. This issue incorporates
online learning to few-shot learning to accomplish efficiently recognition to a novel mode
with the support of small amount of data, and still retain model's ability to recognize existing
modes. This paper designed a backtracking contextual prototypical memory (BCPM) network
for online MFR mode recognition. The proposed method learns temporal and spatial
information from data stream as a reference predicting the signal segments as existing
working modes or a novel mode. The BCPM network also designs a backtracking module and
an alignment regularization term to utilize the data without annotation adequately and
obtain a more reliable category representation. The experimental results verified the
proposed method's good performance and high robustness to non-ideal factors and
distractors.
Keywords: Signal modulation recognition; Radar signal hierarchical modeling; Few-shot
learning; Online learning; Incremental learning

Zhipeng Qing, Kecheng Ge, Shunsheng Zhang, Jing Yang, Zhijin Wen, Youlei Pu,
An inverse synthetic aperture radar imaging framework based on multi-layer networks and
heat conduction attention,
Engineering Applications of Artificial Intelligence,
Volume 167, Part 1,
2026,
113708,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Accurate compensation is essential for achieving high-resolution inverse synthetic
aperture radar (ISAR) imaging. Traditional parametric methods usually rely on iterative
optimization of the objective function to compensate for target motion in radar echoes.
However, the iteration process is often computationally intensive, difficult to integrate into
deep learning frameworks, and may discard sufficiently acceptable intermediate solutions.
To address these challenges, this study proposes a deep unfolding-based translational
compensation network that combines unsupervised learning with gradient back-
propagation. A prototype network is incorporated to monitor the imaging process, enabling
early termination of iterations. Moreover, a U-shaped network architecture based on a heat
conduction attention is employed to enhance ISAR image resolution and focusing
performance. To solve the problem of offset or splitting in the imaging results caused by
residual motion errors, a learnable affine transformation is employed for automatic
centering. These modules are integrated into an echo-to-image ISAR imaging framework.
Experimental results on both simulated and real radar data demonstrate the framework’s
effectiveness and robustness.
Keywords: Inverse synthetic aperture radar imaging; Translational compensation; Heat
conduction; Deep unfolding network; Affine transformation

Bin Xu, Bo Chen, Jinwei Wan, Hongwei Liu, Lin Jin,


Target-Aware Recurrent Attentional Network for Radar HRRP Target Recognition,
Signal Processing,
Volume 155,
2019,
Pages 268-280,
ISSN 0165-1684,
[Link]
([Link]
Abstract: In this paper, we develop a Target-Aware Recurrent Attentional Network (TARAN)
for Radar Automatic Target Recognition (RATR) based on High-Resolution Range Profile
(HRRP) to make use of the temporal dependence and find the informative areas in HRRP,
since it reflects the distribution of scatterers in target along the range dimension.
Specifically, we utilize RNN to explore the sequential relationship between the range cells
within a HRRP sample and employ the attention mechanism to weight up each timestep in
the hidden state so as to discover the target area, which is more discriminative and
informative. Effectiveness and efficiency are evaluated on the measured data. Compared
with traditional methods, besides the competitive recognition performance, TARAN is also
more robust to time-shift sensitivity thanks to the memory function of RNN and attention
mechanism. Furthermore, detailed analysis of TARAN model are provided based on time
domain and spectrogram features.
Keywords: Attention mechanism; Radar Automatic Target Recognition (RATR); Recurrent
Neural Network (RNN); Spectrogram analysis
Pengfei Yang, Feng Wu, Minyang Liu, Ting Zhong, Fan Zhou,
Beyond pillars: Advancing 3D object detection with salient voxel enhancement of liDAR-4D
radar fusion,
Pattern Recognition,
Volume 173,
2026,
112841,
ISSN 0031-3203,
[Link]
([Link]
Abstract: The fusion of LiDAR and 4D radar has emerged as a promising solution for robust
and accurate 3D object detection in complex and adverse conditions. Existing methods
typically rely on pillar-based representations, which, although computationally efficient, fail
to provide fine-grained structural details necessary for precise object localization and
recognition. In contrast, voxel-based representations offer richer spatial information but face
challenges such as background noise and data quality disparity. To address these limitations,
we propose SVEFusion, a voxel-based 3D object detection framework that integrates LiDAR
and 4D radar data using a salient voxel enhancement mechanism. Our method introduces an
adaptive feature alignment module and a novel spatial neighborhood attention module for
efficient early-stage multi-modal voxel feature integration. Furthermore, we design a salient
voxel enhancement mechanism that assigns higher weights to foreground voxels using a
multi-scale weight prediction strategy, progressively refining weight accuracy with
supervision loss. Experimental results demonstrate that SVEFusion significantly outperforms
state-of-the-art methods, establishing a new benchmark in multi-modal 3D object detection.
The source code and network weighting for reproducibility are available at
[Link]
Keywords: Object detection; Lidar; 4D Radar; Multi-modal fusion; Autonomous driving

Yuanbo Li, Wenwu Zhang, Songtao Lv, Jing Yu, Dongdong Ge, Jiawei Guo, Lin Li,
YOLOv11-CAFM model in ground penetrating radar image for pavement distress detection
and optimization study,
Construction and Building Materials,
Volume 485,
2025,
141907,
ISSN 0950-0618,
[Link]
([Link]
Abstract: Ground Penetrating Radar (GPR) is an effective technology for detecting
underground structures and has been widely utilized for monitoring road damage.
Traditional B-scan-based one-dimensional images often fail to preserve continuous spatial
information, thus inadequately reflecting the nuances of damage patterns. This paper
investigates the accurate recognition of hidden internal road damage using 3D-sliced C-scan
images. While YOLO is one of the most effective and rapid neural network models for object
detection, it still suffers from low recognition accuracy and a high rate of missed detections.
To address these issues, this study proposes an improved Convolution and Attention Fusion
Module (CAFM) fusion network model for YOLOv11, which combines the CAFM with the
C2PSA global-local feature extraction mechanism to significantly enhance the recognition
performance for complex road damage. Experimental comparisons between the YOLOv11m-
CAFM and the YOLOv11 model reveal that the combined metrics for the small (n/s) and large
(l/x) models are lower than those for the medium model (m). The YOLOv11m-CAFM
demonstrates strong performance in key metrics such as precision, recall, mAP50, and
mAP50:95, achieving values of 0.840, 0.850, 0.881, and 0.584, respectively, representing
improvements of 0 %, 4.6 %, 1.8 %, and 2.0 % over the baseline model. The confidence level
in detecting standardized targets (e.g., pipelines and well covers) exceeds 0.89. Borehole
validation confirms that the model's localization error is less than 0.15 m, and the detection
frame rate reaches 71 FPS, satisfying the requirements for rapid road assessment. This study
offers a novel method for the intelligent interpretation of GPR images, considering both
detection accuracy and real-time performance, which holds significant engineering
applications in identifying hidden road damages.
Keywords: Pavement disease detection; Ground penetrating radar; Neural network; Object
detection; Deep learning algorithm

Yiwen Chen, Yuan Zhuang, Binliang Wang, Jianzhu Huai,


4D RadarPR: Context-Aware 4D Radar Place Recognition in harsh scenarios,
ISPRS Journal of Photogrammetry and Remote Sensing,
Volume 221,
2025,
Pages 210-223,
ISSN 0924-2716,
[Link]
([Link]
Abstract: Place recognition is a fundamental technology for uncrewed systems such as
robots and autonomous vehicles, enabling tasks like global localization and simultaneous
localization and mapping (SLAM). Existing Place recognition technologies based on vision or
LiDAR have made significant progress, but these sensors may degrade or fail in adverse
conditions. 4D millimeter-wave radar offers strong resistance to particles like smoke, fog,
rain, and snow, making it a promising option for robust scene perception and localization.
Therefore, we explore the characteristics of 4D radar point clouds and propose a novel
Context-Aware 4D Radar Place Recognition (4D RadarPR) method for adverse scenarios.
Specifically, we first adopt a point-based feature extraction (PFE) module to capture raw
point cloud information. On top of PFE, we propose a multi-scale context information fusion
(MCIF) module to achieve local feature extraction at different scales and adaptive fusion. To
capture global spatial relationships and integrate contextual information, the MCIF module
introduces a fusion block based on multi-head cross-attention to combine point-wise
features with local spatial features. Additionally, we explore the role of Radar Cross Section
(RCS) information in enhancing the discriminability of descriptors and propose a local RCS
relation-guided attention network to enhance local features before generating the global
descriptor. Extensive experiments are conducted on in-house datasets and public datasets,
covering various scenarios and including both long-range and short-range radar data. We
compared the proposed method with several state-of-the-art approaches, including
BevPlace++, LSP-Net, and Transloc4D, and achieved the best overall performance. Notably,
on long-range radar data, our method achieved an average Recall@1 of 89.9%,
outperforming the second-best method by 1.9%. Furthermore, our method demonstrated
acceptable generalization ability across diverse scenarios, showcasing its robustness.
Keywords: Place recognition; 4D millimeter-wave radar; Harsh scenarios; Multi-scale context
fusion; Long- and short-range radars; Relocalization

Lele Qu, Jinpeng Tao, Tianhong Yang,


Human sleep posture recognition and vital sign monitoring method using FWCW radar,
Measurement,
Volume 258, Part D,
2026,
119349,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Radar technology has the great potential for recognizing human sleep postures and
tracking respiration rate (RR) and heart rate (HR) during sleep. In this paper, we propose a
novel integrated framework that simultaneously performs sleep posture recognition and
vital sign monitoring using dual frequency-modulated continuous-wave (FMCW) radar
modules. A distinctive aspect of the proposed framework is its dynamic radar switching
strategy, where sleep posture is first classified by applying histogram of oriented gradients
(HOG) and support vector machine (SVM) to the merged range-time map (RTM). Based on
the identified posture, either the top or side radar is selectively engaged to capture
physiological signals from the most effective orientation. To obtain the robust and accurate
estimation of RR and HR, we further propose the GA-MVMD algorithm that integrates
genetic algorithm (GA) and multivariate variational mode decomposition (MVMD) to jointly
decompose chest wall displacement signals across multiple range bins. Experimental results
demonstrate that the proposed method can effectively enhance the recognition accuracy
with an average accuracy rate of 96.7 % for four typical sleep postures and the proposed GA-
MVMD algorithm can provide more accurate RR and HR estimation results.
Keywords: Frequency modulated continuous wave (FMCW) radar; Sleep posture recognition;
Vital sign monitoring

Zai Zhang, Bin Shi, Kai Sun, Hao Wu, Bo Dong,


RADAR: Relation-assisted dual-graph aligning recognition for grounded multimodal named
entity recognition,
Information Processing & Management,
Volume 63, Issue 3,
2026,
104552,
ISSN 0306-4573,
[Link]
([Link]
Abstract: Grounded Multimodal Named Entity Recognition (GMNER) requires the
simultaneous identification of textual entities and their corresponding visual regions within
images. However, the inability to model visual contextual semantics and the disorganized
processing of cross-modal features often lead existing methods to struggle with both visual
entity differentiation and bridging the modality gap. We propose RADAR (Relation-Assisted
Dual-graph Aligning Recognition), a novel framework that leverages visual relations derived
from scene graphs to encode structured context and enhance visual understanding. To
achieve fine-grained cross-modal alignment, we design an object-level alignment self-
attention mechanism and introduce a dual-graph strategy. Evaluated on the Twitter-GMNER
dataset (13,076 image-text pairs), RADAR achieves 60.91 % F1 score, a +4.5 % improvement
over the H-Index baseline. The method also demonstrates consistent gains in subtasks, with
+3.54 % improvement in EEG metric, validating its effectiveness in multimodal entity
alignment.
Keywords: Multimodal named entity recognition; Visual grounding; Visual scene graph;
Cross-modal alignment

Haoming Feng, Huaqing Li, Wenwen Zhu, Denghao Li, Yukun Huang,
Micro-motion enhanced multi-person activity recognition with millimeter-wave radar,
Measurement,
Volume 258, Part B,
2026,
119090,
ISSN 0263-2241,
[Link]
([Link]
Abstract: As a non-contact sensing device, millimeter-wave radar exhibits unique strengths
in human activity recognition (HAR). Existing methods rely on micro-Doppler signatures for
activity classification, but they often encounter feature aliasing in multi-person activity
recognition (MPAR) scenarios. Although point cloud-based approaches can distinguish
individual targets, they primarily extract static morphological features, neglecting the micro-
motion information of human joints, which is crucial for accurate activity recognition. To
address these limitations, we proposes an innovative MPAR framework that integrates
spatial point clouds and micro-motion features. First, an improved point cloud data
association algorithm is applied to achieve multi-target point cloud feature separation,
followed by a dynamic projection mechanism to construct time–Doppler feature maps.
Then, a torso micro-motion enhancement algorithm is designed to enhance the details of
human body movements. Finally, a CNN-LSTM hybrid network architecture with a temporal-
attention is constructed for action classification. Experimental results show that the
proposed micro-motion enhancement algorithm improves recognition accuracy by 27.1%
and 2.3%, compared to two traditional time–frequency analysis methods. Furthermore,
MPAR task in occlusion scenarios achieves recognition accuracy of 93.5%. In summary,
proposed framework not only retains the inherent advantages of millimeter-wave radar but
also significantly enhances multi-person activity recognition in complex scenarios.
Keywords: Human activity recognition (HAR); Multi-person activity recognition (MPAR);
Feature separability; Temporal attention; Occlusion scenarios

Guanliang Liu, Wenchao Chen, Bo Chen, Bo Feng, Penghui Wang, Hongwei Liu,
Supervised contrastive deep Q-Network for imbalanced radar automatic target recognition,
Pattern Recognition,
Volume 161,
2025,
111264,
ISSN 0031-3203,
[Link]
([Link]
Abstract: In the presence of limited and extremely imbalanced data, deep learning methods
for radar automatic target recognition (RATR) often suffer from significant performance
degradation and overfitting. To tackle this issue, we propose Supervised Contrastive Deep Q-
network (SCDQ), a novel end-to-end reinforcement learning method, for multi-class
imbalanced RATR. SCDQ formulates the imbalanced recognition problem as a Markov
decision process (MDP) and optimizes the classifier through an enhanced Q-learning
paradigm. In order to augment the model’s feature extraction capabilities under the
constraint of limited samples, we tightly integrate reinforcement learning (RL) with
supervised contrastive learning, introducing an innovative feature enhancement module. To
further enhance the model’s adaptability to challenging samples, we integrate a
meticulously designed priority sampling into the proposed SCDQ framework, denoted as
SCDQ-P. Experimental results on both simulated and real datasets demonstrate the reliability
and effectiveness of the proposed method.
Keywords: Imbalanced RATR; Deep learning; Deep reinforcement learning (DRL); Supervised
contrastive learning; Priority sampling

Ligen Chen, Nannan Zhu, Hongbo Chen, Yonghao Dong, Yue Zhang, Nian Cai,
A causality-inspired single-source domain generalized method for low-slow-small threat
target recognition through holographic Doppler radar,
Expert Systems with Applications,
Volume 287,
2025,
128104,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Low-slow-small (LSS) target monitoring is critical for airport safety management,
particularly when LSS objects such as unmanned aerial vehicles (UAVs) and birds
unexpectedly enter airport airspace, posing significant risks to regular flights and airport
operations. Recent advances have taken advantage of deep learning for the recognition of
LSS radar targets, achieving promising classification accuracy. However, existing LSS radar
target recognition approaches often rely on statistical correlations, including unstable
spurious correlations, which can undermine the generalization performance of classification
networks, limiting their effectiveness in all-time radar recognition. To address this, we
propose a causality-inspired single-source domain generalization method for radar LSS target
recognition. Our method introduces a Causal-Symmetric Transformation (CST) module for
data augmentation, combining Non-Causal Augmentation for global perturbations and
Symmetric Transformation for local motion reversal, enhancing data diversity and reducing
bias. Additionally, we propose a Causal Mining (CM) module with a Causal Consistency loss
to extract causal features that boost generalization. A Fourier-Aware Attention (FAA) module
leverages frequency-domain information to strengthen feature representation and preserve
causal information. Extensive experiments on four real-world datasets validate the
effectiveness of our approach.
Keywords: Radar target recognition; Single-source domain generalization; Causility-inspired
model; Low-slow-small target; Holographic Doppler radar

Ziwei Zhang, Mengtao Zhu, Yunjie Li, Yan Li, Shafei Wang,
Joint recognition and parameter estimation of cognitive radar work modes with LSTM-
transformer,
Digital Signal Processing,
Volume 140,
2023,
104081,
ISSN 1051-2004,
[Link]
([Link]
Abstract: The recent developed cognitive radars can implement flexible work modes with
programmable modulation types and optimized modulating values for each mode definition
parameter. Automatic analysis of these work modes is a significant challenge for modern
electromagnetic reconnaissance receivers. In this paper, a Multi-Output Multi-Structure
(MOMS) learning-based processing framework is proposed for Joint inter-pulse automatic
Modulation Recognition and Parameter Estimation (JMRPE-MOMS). We propose a label
construction method as a feature interpretation method of the network to facilitate MOMS
learning and utilize the correlations between labels for performance gain. Moreover, an
LSTM-Transformer is designed to mine deep time-series characteristics, which can model
local and global relationships and reduce quantization loss. The proposed framework can
perform joint modulation recognition and parameter estimation (JMRPE) tasks
simultaneously with flexible output structures including scalar output and vector output
with fixed or variable sizes. Extensive simulations are performed based on the simulated
radar work modes defined with pulse repetition interval (PRI) sequences. The simulation
results validate the effectiveness and superiority of the proposed method especially under
non-ideal electromagnetic environments.
Keywords: Radar work mode; Automatic modulation recognition; Modulation parameter
estimation; Multi-output learning; Transformer

Wentao He, Jianfeng Ren, Ruibin Bai, Xudong Jiang,


Radar gait recognition using Dual-branch Swin Transformer with Asymmetric Attention
Fusion,
Pattern Recognition,
Volume 159,
2025,
111101,
ISSN 0031-3203,
[Link]
([Link]
Abstract: Video-based gait recognition suffers from potential privacy issues and performance
degradation due to dim environments, partial occlusions, or camera view changes. Radar has
recently become increasingly popular and overcome various challenges presented by vision
sensors. To capture tiny differences in radar gait signatures of different people, a dual-branch
Swin Transformer is proposed, where one branch captures the time variations of the radar
micro-Doppler signature and the other captures the repetitive frequency patterns in the
spectrogram. Unlike natural images where objects can be translated, rotated, or scaled, the
spatial coordinates of spectrograms and CVDs have unique physical meanings, and there is
no affine transformation for radar targets in these synthetic images. The patch splitting
mechanism in Vision Transformer makes it ideal to extract discriminant information from
patches, and learn the attentive information across patches, as each patch carries some
unique physical properties of radar targets. Swin Transformer consists of a set of cascaded
Swin blocks to extract semantic features from shallow to deep representations, further
improving the classification performance. Lastly, to highlight the branch with larger
discriminant power, an Asymmetric Attention Fusion is proposed to optimally fuse the
discriminant features from the two branches. To enrich the research on radar gait
recognition, a large-scale NTU-RGR dataset is constructed, containing 45,768 radar frames of
98 subjects. The proposed method is evaluated on the NTU-RGR dataset and the MMRGait-
1.0 database. It consistently and significantly outperforms all the compared methods on
both datasets. The codes are available at: [Link]
Keywords: Micro-Doppler signature; Radar gait recognition; Spectrogram; Cadence velocity
diagram; Asymmetric Attention Fusion

Kuiyu Chen, Jingyi Zhang, Si Chen, Shuning Zhang,


Deep metric learning for robust radar signal recognition,
Digital Signal Processing,
Volume 137,
2023,
104017,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Signal recognition technology is a currently active area in both civilian and military
applications. Recently, deep learning has aroused extensive attempts in radar signal
recognition due to its remarkable capability of automatic feature extraction. However,
existing radar signal recognition networks overly depend on the probability-based decision
model, resulting in poor robustness. This paper develops a novel deep metric learning frame
to enhance the robustness of the recognition system. First, a multiscale atrous pyramid
network (MAPNet) is proposed to efficiently learn high-resolution and distinct feature
representation. Second, a variance loss is designed to constrain the intra-class feature
distribution in metric space. Third, according to the distribution of training signals in metric
space, recognition results are recalibrated to provide explicit rejection probabilities for
unknowns. Extensive experiments and evaluations demonstrate that the proposed model
can accurately classify known signals while robustly identifying unknown signals. The signal
database and model can be freely accessed at [Link]
Keywords: Radar signal recognition; Metric learning; Multiscale atrous pyramid; Variance
loss; Unknown signals

Zhiyong Huang, Guoyuan Xu, Xiaoning Zhang, Bo Zang, Huayang Yu,


Three-dimensional ground-penetrating radar-based feature point tensor voting for semi-
rigid base asphalt pavement crack detection,
Developments in the Built Environment,
Volume 21,
2025,
100591,
ISSN 2666-1659,
[Link]
([Link]
Abstract: Three-dimensional Ground penetrating radar (3D-GPR) has been widely applied in
nondestructive testing of concealed cracks within asphalt pavement. However, due to the
weak GPR echo characteristics of concealed cracks and their susceptibility to environmental
noise, automatic recognition of crack echo features has always faced significant challenges.
To address this issue, numerous semi-rigid base crack images were collected and extracted
using feature point tensor voting with 3D-GPR's efficient, non-destructive road structure
detection. In this paper, the radar image is gridded by the ECA-ResNet network, and the
center point of the detected crack grid is used as the feature point, and the continuous path
of the crack is reconstructed by the tensor voting algorithm. The results show that this
method achieves 90% crack extraction, which is superior to traditional target detection
networks such as YOLOv5 and Fast R-CNN, providing an effective tool for rapid non-
destructive detection of pavement cracks.
Keywords: Three-dimensional ground-penetrating radar; Tensor voting; ECA-ResNet; Base
course cracks

Xiaoyuan Zhang, Shaohang Jing, Jingshu Li, Yechao Bai, Feng Yan,
Cognitive radar recognition with Kolmogorov-Smirnov test and momentum gradient descent,
Digital Signal Processing,
Volume 163,
2025,
105212,
ISSN 1051-2004,
[Link]
([Link]
Abstract: The emission parameters of cognitive radars can adaptively change according to
the environment, which poses a challenge to radar electronic countermeasures (ECM). To
counter cognitive radars, it is essential to identify the cognitive characteristics. In this paper,
a method is proposed to recognize cognitive radars with power allocation function. The
signal-to-interference-plus-noise ratio (SINR) distribution of cognitive radars is derived
through feature functions, and hypothesis test is used to identify whether the target radar
has cognitive function by designing a Kolmogorov-Smirnov (K-S) detector to recognize
adaptive optimization power allocation. Subsequently, a momentum gradient descent
algorithm is used to optimize the signal of the jamming machine to reduce type II error
probability of radar recognition. K-S detector is simulated and compared with Afriat
detector, SVM and MLP detector. Results demonstrate that the K-S detector outperforms
both the Afriat and MLP detectors in identifying cognitive radars with dynamic power
allocation functionality. At the same detection probability, the K-S detector achieves a 2 dB
improvement over the MLP detector and a 4 dB improvement over the Afriat detector.
Keywords: Cognitive radar; Electronic countermeasures (ECM); Kolmogorov–Smirnov test;
Momentum gradient descent algorithm; Afriat theorem
Baiju Yan, Peng Wang, Lidong Du, Xianxiang Chen, Zhen Fang, Yirong Wu,
mmGesture: Semi-supervised gesture recognition system using mmWave radar,
Expert Systems with Applications,
Volume 213, Part B,
2023,
119042,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Gesture recognition has found versatile applications in natural human–computer
interaction (HCI). Compared with traditional camera-based or wearable sensors-based
solutions, gesture recognition using the millimeter wave (mmWave) radar has attracted
growing attention for its characteristics of contact-free, privacy-preserving and less
environment-dependence. Recently, most of studies adopted one of the Range Doppler
Image (RDI), Range Angle Image (RAI), Doppler Angle Image (DAI) or Micro-Doppler
Spectrogram extracted from the raw radar signal as the input of a deep neural network to
realize gesture recognition. However, the effectiveness of these four inputs in gesture
recognition has attracted little attention so far. Moreover, the lack of large amounts of
labeled data restricts the performance of traditional supervised learning network. In this
paper, we first conducted extensive experiments to compare the effectiveness of these four
inputs in the gesture recognition, respectively. Then we proposed a semi-supervised leaning
framework by utilizing few labeled data in the source domain and large amounts of
unlabeled data in the target domain. Specially, we combine the ∏-model and some specific
data augmentation tricks on the mmWave signal to realize the domain-independent gesture
recognition. Extensive experiments on a public mmWave gesture dataset demonstrate the
superior effectiveness of the proposed system.
Keywords: Gesture recognition; mmWave radar; Semi-supervised learning; ∏-model

Nanyu Jiang, Yuyuan Fang, Lei Zhang, Chao He, Zhenhua Wu,
IFM-PointNet++: Achieving efficient radar signal waveform recognition with instantaneous
frequency measurement,
Digital Signal Processing,
Volume 168, Part B,
2026,
105507,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Efficient and robust identification of radar signal waveforms is an essential task in
electronic reconnaissance. Current deep learning-based methods can obtain satisfying
accuracy, but they are usually with high computational burden. To address the issue, this
article develops an efficient algorithm IFM-PointNet++. This algorithm transfers the radar
signal waveform recognition into the point cloud recognition task by adopting the
instantaneous frequency measurement (IFM). By integrating IFM with PointNet++ network,
our method achieves superior efficiency and accuracy. To demonstrate this effectiveness, we
conduct comprehensive comparison experiments with YOLOv8 waveform recognition on the
time-frequency images. The results demonstrate that our proposed method significantly
accelerates signal waveform recognition while maintaining high accuracy.
Keywords: Radar signal waveform recognition; Signal recognition; PointNet++; Instantaneous
frequency measurement; Deep learning

Liheng Dong, Chengyang Tao, Zhaoxiang Zhang, Guiqing He, Yuelei Xu, Dong Liu,
An attitude-centric and cross-band infrared framework for aerial target intention
recognition,
Aerospace Science and Technology,
Volume 170,
2026,
111510,
ISSN 1270-9638,
[Link]
([Link]
Abstract: In terminal attack-defence scenarios, aerial target intention can be accurately
recognized based on trajectory and attitude information. Currently, radar-based intention
recognition methods hold a dominant position. However, radar is not adept at capturing
target attitude information (e.g., pitch, roll, and yaw) and can only perform intention
recognition based on trajectory information. To bridge this gap, this paper proposes an
attitude-centric and cross-band infrared aerial target intention recognition framework that
simultaneously extracts trajectory and attitude representations from multi-band infrared
images. The framework comprises two steps: pose estimation for predicting keypoints from
infrared images, followed by intention recognition based on these keypoints. For pose
estimation, a cross-band invariant representation learning method is proposed to reduce the
impact of band-bias, thereby improving the model’s generalization on multi-band infrared
images. For intention recognition, regularized adaptive adjacency matrix and parameter
fusion mechanisms are designed to effectively capture aerial target attitude representations,
forming an attitude-centric approach. Experiments demonstrate that the proposed method
significantly enhances the generalization of various pose estimation models on multi-band
infrared images. Additionally, with the introduction of attitude representations, the intention
recognition accuracy increases from 90.12 % to 96.64 %.
Keywords: Aerial target; Intention recognition; Pose estimation; Multi-band infrared;
Attitude

Hu Liu, Zhenghua Zhang, Jing Yang, Jörg Benndorf, Xiaofei Wang, Jiaqi Dong, Zitao Lin,
Guoliang Chen,
GhostPointNet: A deep learning-based method for ghost point noise detection in four-
dimensional (4D) millimeter-wave radar point clouds of underground mine,
Engineering Applications of Artificial Intelligence,
Volume 161, Part C,
2025,
112380,
ISSN 0952-1976,
[Link]
([Link]
Abstract: The high dust concentration, multi-metal supports, and narrow winding tunnels in
underground mines collectively lead to frequent ghost point noise in four-dimensional (4D)
millimeter-wave radar point clouds, posing serious challenges for mining perception and
localization. To address this, we propose a deep learning algorithm, named GhostPointNet,
for 4D millimeter-wave radar ghost point detection in underground mining environments.
From an artificial intelligence perspective, this model thoroughly considers the multi-modal
features of 4D millimeter-wave radar and the environmental complexity of underground
mines. It incorporates multi-parameterized spatial information inputs in both Cartesian and
Spherical coordinates, coupled with “Double T-Net” adaptive alignment correction, while
integrating non-spatial information such as radar power and Doppler data to achieve multi-
modal representation and end-to-end discrimination between ghost points and real points.
Experimental validation shows that GhostPointNet achieves excellent performance in
underground mining scenarios with 92.45 % accuracy and 95.84 % F1-score, outperforming
traditional filtering, clustering, and machine learning algorithms. From an engineering
application perspective, GhostPointNet is specifically designed for ghost noise detection in
underground mines. Even in complex scenarios such as mine tunnel intersections and turns,
it preserves critical structural points. Its end-to-end neural network simplifies post-
processing procedures, enhances operational efficiency, and provides stable and reliable
perceptual support for subsequent tasks such as autonomous mine locomotive navigation
and three-dimensional (3D) structure reconstruction. Experimental results demonstrate that
this method surpasses baseline approaches in ghost point detection, real point preservation,
and generalization capability, providing significant support for improving underground
mining safety and efficiency.
Keywords: Deep learning; Four-dimensional (4D) millimeter-wave radar; Ghost noise;
Underground mining; Point cloud segmentation

Ye Qiu, Zhenmiao Deng, Xiaohong Huang,


Unified complex-valued high-resolution frequency representation with cross-domain
attention for radar-based physiological state recognition,
Pattern Recognition,
Volume 172, Part B,
2026,
112488,
ISSN 0031-3203,
[Link]
([Link]
Abstract: High-resolution frequency domain analysis is pivotal in a wide range of critical
applications, including physiological signal processing, radar target detection, and
communication systems. In this study, we present a complex-valued neural network
designed for accurate estimation of frequency components encompassing both magnitude
and phase, the Unified-Complex High-Resolution Frequency Representation Module
(UHFreq). This method generates comprehensive high-resolution frequency domain
representations, addressing key limitations in current approaches that typically capture only
amplitude information, omit crucial phase details, and suffer from low resolution in
frequency domain outputs. Furthermore, conventional methods for physiological signal
detection and recognition require meticulous preprocessing steps, including demodulation
and filtering. In response to these challenges, we propose UHFreq-based Vital Sign Status
Detection Network (UVSD-Net), an application example of UHFreq, which classifies different
human physiological states starting from raw radar echoes. This model utilizes the UHFreq
structure as the frontend for the frequency domain representation of physiological signals
from raw radar echoes. The UVSD-Net architecture incorporates a dual-pathway design: one
pathway processes frequency domain features via UHFreq, while the other applies time
domain amplitude and phase information from the raw radar signals. Furthermore, a weight
redistribution mechanism is introduced across the different feature domains to enhance
cross-domain feature integration and interaction. This comprehensive end-to-end
framework offers a robust approach for analyzing time domain original signals and enables
effective execution of downstream tasks.
Keywords: Biomedical ignal recognition; Physiological pattern recognition; Frequency
analysis; Signal processing; Deep learning (DL)

Li Qiusheng, Zhu Huajuan,


Target classification with low-resolution radars based on cyclic bispectrum and improved
ACGAN,
Measurement,
Volume 259, Part B,
2026,
119715,
ISSN 0263-2241,
[Link]
([Link]
Abstract: To address the challenges of insufficient generalization and high noise sensitivity in
low-resolution radar target recognition under limited-sample conditions, this paper
proposes a joint optimization framework integrating cyclic bispectral analysis and an
improved Auxiliary Classifier Generative Adversarial Network (ACGAN). First, a third-order
cyclic cumulant spectral model is designed to extract modulation-specific signatures of
aircraft targets in the cyclostationary domain, effectively suppressing both Gaussian and
non-Gaussian noise while preserving discriminative features that are robust to low SNR
conditions (maintaining 92.7 % accuracy at 0 dB). Second, an enhanced ACGAN architecture
is developed by incorporating self-attention mechanisms and Wasserstein distance
optimization with gradient penalty, with spectral normalization and dynamic gradient
penalties introduced to stabilize training dynamics and improve synthetic sample fidelity.
Extensive experiments on a real-world dataset collected by a certain Chinese-made VHF-
band radar demonstrate that the proposed method achieves state-of-the-art performance,
with average recognition accuracies of 98.46 % and 98.52 % for approaching and departing
targets in complex noise environments, respectively, alongside a Kappa coefficient exceeding
0.97. Comprehensive comparisons with traditional methods (wavelet, HOS, FrFT), GAN
variants (WGAN-GP, AFGAN + ResNet, Diffusion-GAN), VAE-based approaches, few-shot
learning models (ProtoNet, MatchingNet), and a modern Vision Transformer (ViT) baseline
show consistent improvements of 1.38–12.19 % in accuracy. Ablation studies validate the
contributions of key components, where the self-attention module and Wasserstein
optimization improve accuracy by 1.27 % and 0.97 %, respectively. Furthermore, embedded
platform tests confirm the framework’s feasibility for real-time deployment (inference
time < 15 ms/sample), offering a robust solution for resource-constrained radar systems.
This work highlights the efficacy of unifying physics-inspired feature extraction with
stabilized deep generative models for practical radar recognition.
Keywords: Radar target recognition; Cyclic bispectral analysis; Auxiliary classifier generative
adversarial network (ACGAN); Self-attention mechanism; Few-shot learning

Liying Wang, Zongyong Cui, Yiming Pi, Changjie Cao, Zongjie Cao,
Low personality-sensitive feature learning for radar-based gesture recognition,
Neurocomputing,
Volume 493,
2022,
Pages 373-384,
ISSN 0925-2312,
[Link]
([Link]
Abstract: Radar-based sensing of gestures has gained tremendous attention with the recent
advancements in radar technologies. However, evident discrepancies exist in the gesture
samples due to hand flexibility and individual habits. It is challenging for traditional methods
to identify the gestures from unknown data sources. Cross-person (Cross-scenario)
recognition refers to a recognition where the training and test samples are from different
people (scenarios), respectively. To explore how the recognition performance is affected by
the individual habits, the reasons are analyzed and visualized through the experiments. On
this basis, HandNet is targeted proposed for the low personality-sensitive feature learning
and it has two main contributions. First, a Stepped Data Augmentation (SDA) is proposed to
reduce the sample interferences by non-coherent accumulating, and capture the inter-frame
dependencies. Second, a Focus on Generalization loss (FoG loss) is proposed to highlight the
generalized feature learning by res tricting the distances of inter-source features. Extensive
experiments demonstrate that HandNet effectively reduces the classifier’s sensitivity to the
personalized habits, and outperforms the existing state-of-the-art methods on the cross-
person and cross-scenario gesture recognition. To the best of our knowledge, it is the first
time to dedicate to addressing the radar-based gesture recognition with low personal
sensitivity, which is more suitable for practical scenarios.
Keywords: Convolutional neural network; Gesture recognition; Feature learning

Philipp Reitz, Tobias Veihelmann, Norman Franchi, Maximilian Lübke,


Dual radar vision: A feature fusion approach for advanced object detection in IoT radar
networks,
Machine Learning with Applications,
Volume 21,
2025,
100703,
ISSN 2666-8270,
[Link]
([Link]
Abstract: 60GHz radar technology is one of the most promising movement detector
solutions for Internet of Things (IoT) applications. However, challenges remain in accurately
classifying different objects and detecting small objects in a multi-target scenario. This work
investigates whether sensor fusion between multiple radars can enhance object detection
and classification performance. A one-stage detection architecture, designed based on the
features of the latest YOLO generations, is used to perform fusion based on range-Doppler
(RD) maps of two non-coherent spatially separated radars. A complete physical 3D
propagation simulation using ray tracing evaluates the fusion methods. This approach
enables precise ground truth, as all unprocessed signal components are known, and
guarantees a consistent, error-free reference. Results demonstrate that dynamic, attention-
based fusion significantly improves detection and classification compared to static fusion in
homogeneous and heterogeneous radar setups.
Keywords: Data fusion; Deep learning; FMCW radar; IoT; Radar networks; Ray tracing; YOLO

Jingang Wang, Songbin Li, Ke Shi,


Radar target tracking based on motion characteristic and distribution pattern matching,
Signal Processing,
Volume 236,
2025,
110034,
ISSN 0165-1684,
[Link]
([Link]
Abstract: The mitigation of false alarm rates under real-world radar operating conditions
represents a critical challenge in advancing radar target detection algorithms. This study
proposes that utilizing multi-frame correlation information through radar target tracking
constitutes an effective solution for suppressing false alarms. We present a radar target
tracking methodology that integrates motion characteristic and distribution pattern
matching, effectively leveraging multi-frame radar measurements and echo amplitude
information. This approach enables false alarm reduction through trajectory consistency
validation. Specifically, the method operates in two stages: First, state filtering based on
motion characteristics is applied to predict potential candidate regions for previous
detections within the current frame. Subsequently, within these candidate regions, a deep
learning-based similarity evaluation framework employing a self-supervised Siamese
network performs distribution pattern matching to establish optimal data associations.
Experimental validation demonstrates that the proposed method achieves a 5.82%
improvement in F1-score over benchmark algorithms, confirming its enhanced detection
reliability and operational effectiveness.
Keywords: Radar target tracking; Distribution pattern matching; Deep learning

Irfanullah Khan, Antonio Guerrieri, Edoardo Serra, Giandomenico Spezzano,


A hybrid deep learning model for UWB radar-based human activity recognition,
Internet of Things,
Volume 29,
2025,
101458,
ISSN 2542-6605,
[Link]
([Link]
Abstract: In today’s world, energy efficiency in buildings has become a top priority due to the
significant energy waste caused by the operation of inefficient electrical appliances.
Conventional methods of reducing energy waste cause discomfort for occupants inside
buildings. One promising way to optimize energy consumption is to synchronize appliance
operation with building occupants’ dynamic behavior. Internet of Things (IoT) technologies,
which allow for widespread data collection and execution of Machine Learning (ML)
algorithms, enabled the creation of Smart Buildings (SBs). SBs can learn patterns from the
inhabitant’s behavior residing in, and adjust their operations in accordance with these
behaviors. By doing so, these SBs could reduce energy waste, enhancing resource efficiency
and consequently reduce CO2 gas emissions. Furthermore, they could improve the overall
comfort of the living environment and help with sustainability initiatives. In this context, this
paper proposes a novel approach that uses a hybrid deep-learning model to recognize
complex human activities based on data collected from ultra-wideband (UWB) radar
technology. Our approach, called Hybrid Deep Learning Model for Activity Recognition
(HDL4AR), includes long-short-term memory (LSTM) and a one-dimensional convolutional
neural network (1D-CNN). We deploy a real-time case study by collecting data from 22
participants involved in 10 diverse activities at the headquarters of the ICAR-CNR in the IoT
Laboratory, Italy. Moreover, we conducted a comprehensive benchmark of the HDL4AR
approach against various statistical techniques and other deep learning models recently
introduced in the literature. Results show that our proposed approach outperformed
conventional methods and achieved an impressive accuracy of 98.42%.
Keywords: Internet of Things; Smart buildings; Human activity recognition; UWB radar;
Artificial intelligence; Neural networks; LSTM

Mingyang Du, Ping Zhong, Xiaohao Cai, Daping Bi, Aiqi Jing,
Robust Bayesian attention belief network for radar work mode recognition,
Digital Signal Processing,
Volume 133,
2023,
103874,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Understanding and analyzing radar work modes play a key role in electronic
support measure system. Many classifiers, for example those based on convolutional neural
network (CNN) and recurrent neural network (RNN), are available for recognizing radar work
modes as well as emitter types from their waveform parameters. However, the performance
of these methods may suffer significantly when confronting different types of signal
degradation, e.g., measurement error, lost pulse and spurious pulse. To tackle this issue, we
in this paper develop a Bayesian attention belief network (BABNet) based on Bayesian neural
networks in which the probability distribution over weights can help to enhance the model
robustness for corrupted data. In particular, we adopt pre-trained CNN as the Bayesian
inference prior. This not only accelerates the convergence speed, but also avoids the training
process getting stuck in bad local minima. Meanwhile, instead of using RNNs which are
difficult to be implemented in parallel, the combination of padding operation and attention
module in the proposed BABNet enables CNN, as the backbone, to process sequential data
with variable length. Extensive experiments are conducted to demonstrate the recognition
capability and robustness of the BABNet in different environments.
Keywords: Radar work mode; Pulse descriptor word; Attention mechanism; Bayesian neural
network; Robustness; Recognition

Van Ngoc Dang, Ngoc Chau Hoang, Quoc Cuong Nguyen, Minh Thuy Le,
Advancing robust human activity recognition via informative mmWave radar characteristics
and a lightweight spatio-spectro-temporal network,
Measurement,
Volume 256, Part A,
2025,
118056,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Human activity recognition (HAR) is increasingly important in aiding our daily life,
with millimeter-wave (mmWave) radar sensors emerging as a promising noninvasive solution
thanks to their excellent spatial and velocity resolution. Although existing radar-based
systems have shown strong performance, they primarily focus on micro-Doppler signatures
while neglecting angle information, which can hinder practical deployment in real-world
scenarios. Moreover, current state-of-the-art recognition models using mmWave radar often
require substantial computational resources, making integration into resource-constrained
devices challenging. This work proposes an efficient radar-based HAR system that leverages
angle and spectro-temporal information from micro-Doppler signatures. Our system utilizes
a multi-channel micro-Doppler representation corresponding to the number of virtual
antenna receivers as input. Then, a lightweight dilated convolutional network, namely SST-
DCN, extracts spatial-aware multi-scale spectro-temporal information through time-
frequency dilated convolutions. Experimental results on our real-world dataset demonstrate
the superiority of our approach compared to conventional features and other state-of-the-
art radar-based HAR systems.
Keywords: Human activity recognition; Millimeter-wave radar; Deep learning; Lightweight
network; Dilated convolution

Shanliang Yuan, Hongyuan Gao, Xiaoyuan Gu, Jingya Ma,


Synthetic aperture radar target recognition with limited training data based on frequency-
domain-assisted dual-stream attention hierarchical deformable convolutional networks,
Engineering Applications of Artificial Intelligence,
Volume 161, Part C,
2025,
112309,
ISSN 0952-1976,
[Link]
([Link]
Abstract: In the field of synthetic aperture radar (SAR) automatic target recognition (ATR),
the recognition performance of deep learning methods is often constrained by sample
scarcity. To address this issue, this paper proposes a frequency-domain-assisted dual-stream
attention hierarchical deformable convolutional network (DAHDF-Net). A frequency-domain-
assisted dataset construction method is designed to extend the representational diversity of
the sample feature space. Then, a parallel input architecture of frequency-space dual-stream
features is constructed, and the adaptive dynamic fusion of frequency-domain stream
features with spatial-domain stream features is realized by using a channel attention module
in the frequency-domain stream. To enhance the expressiveness of the network on the basis
of modulated deformation convolution, a hierarchical residual (H-residual) module is
designed. Moreover, a triple hybrid loss is proposed to address the issues of high
deformation, which makes classification difficult, and the fuzzy decision boundaries between
categories. Experiments were conducted on the moving and stationary target acquisition
and recognition (MSTAR) dataset and the FUSAR-Ship dataset. Among them, DAHDF-Net
achieves recognition accuracies of 99.75 % and 97.48 % on all training sets and 10-way 25-
shot under standard operating condition (SOC) of MSTAR, which demonstrated superior
recognition accuracy compared to current state-of-the-art methods.
Keywords: Synthetic aperture radar target recognition; Frequency domain; Deformable
convolution; Hybrid loss; Limited data

Hongping Zhou, Lei Wang, Minghui Ma, Zhongyi Guo,


Compound radar jamming recognition based on signal source separation,
Signal Processing,
Volume 214,
2024,
109246,
ISSN 0165-1684,
[Link]
([Link]
Abstract: To deal with various jamming signals based on digital radio frequency memories, a
compound jamming signal recognition method based on source signal separation is
proposed. In order to overcome label limitation of the supervised learning method in the
recognition process, this paper puts forward "Separation + Recognition" strategy. Firstly, the
received signals of multiple channels are preprocessed, and the number of signal sources is
analyzed through the single-source detection algorithm. Then the received compound
jamming signals are processed by source separation, and the independent single jamming
signals can be obtained. On this basis, a fast signal compensation algorithm is added to
compensate different source signals at overlapping time-frequency points, which effectively
increases the integrity of the separated signals. The separated jamming signals are then put
into the convolutional neural network for recognition, and the specific jamming types in the
compound jamming signals can be obtained. It has been proved that the recognition
accuracy of five kinds of compound jamming exceeds 90% when the jamming-to-noise ratio
is 0 dB.
Keywords: Active-jamming recognition; Deep learning; Compound jamming recognition;
Neural networks

Anand Dubey, Avik Santra, Jonas Fuchs, Maximilian Lübke, Robert Weigel, Fabian Lurz,
HARadNet: Anchor-free target detection for radar point clouds using hierarchical attention
and multi-task learning,
Machine Learning with Applications,
Volume 8,
2022,
100275,
ISSN 2666-8270,
[Link]
([Link]
Abstract: Target localization and classification from radar point clouds is a challenging task
due to the inherently sparse nature of the data with highly non-uniform target distribution.
This work presents HARadNet, a novel attention based anchor free target detection and
classification network architecture in a multi-task learning framework for radar point clouds
data. A direction field vector is used as motion modality to achieve attention inside the
network. The attention operates at different hierarchy of the feature abstraction layer with
each point sampled according to a conditional direction field vector, allowing the network to
exploit and learn a joint feature representation and correlation to its neighborhood. This
leads to a significant improvement in the performance of the classification. Additionally, a
parameter-free target localization is proposed using Bayesian sampling conditioned on a pre-
trained direction field vector. The extensive evaluation on a public radar dataset shows an
substantial increase in localization and classification performance.
Keywords: Multi-task learning; Radar detection; Scene understanding

Hao Yang, Shirong Zhou, Liyan Liu, Zhong Zhou,


A fast and accurate detection model of internal defects in tunnel lining for ground
penetrating radar image data,
Advanced Engineering Informatics,
Volume 68, Part C,
2025,
103812,
ISSN 1474-0346,
[Link]
([Link]
Abstract: Defects in tunnel linings accelerate structural deterioration, reduce service life, and
pose serious safety risks. Existing algorithms for detecting defect signals in ground-
penetrating radar (GPR) images often struggle to balance accuracy and efficiency, with
limited capacity to extract meaningful features. To address these limitations, this paper
proposes a lightweight algorithm, MGD-DETR, for accurate recognition of internal tunnel
lining defects, using RT-DETR as the base model. First, a Multi-HGNet backbone feature
extraction network is introduced to reduce model size (MS) and enhance dynamic fusion and
interaction between feature layers, thereby improving feature extraction. Second, the
lightweight convolution module GSConv replaces standard convolution operations to reduce
the parameter count. Third, a dual attention module (DAM) is integrated to dynamically
adjust spatial and channel feature weights, improving the model’s generalization
performance. Five models—RT-DETR, YOLO-LD, YOLOv10, YOLOv11, and SSD—were used for
comparative evaluation. Experimental results show that MGD-DETR outperforms the other
models across all metrics, achieving a mean average precision (mAP) of 0.834, mean F1
score (mF1) of 0.818, MS of 26.9 M, and frames per second (FPS) of 91.2f/s, enabling fast
and accurate recognition of defect signals and facilitate subsequent deployment into tunnel
detection mobile devices.
Keywords: Tunnel engineering; Lining defects; Deep learning; Ground-penetrating radar
images
Wenxu Zhang, Kang Luo, Fuli Sun, Zhongkai Zhao, Yunxiao Fu, Feiran Liu,
Radar working mode recognition for small samples based on the DBA-CIB-IMP method,
Physical Communication,
Volume 72,
2025,
102718,
ISSN 1874-4907,
[Link]
([Link]
Abstract: To address the issue that the reliability of radar working mode recognition
decreases as detected radar pulses decrease, a Dynamic Time Warping (DTW) Barycenter
Averaging-Circularly Integrated Bispectrum-Informer Multilayer Perceptron (DBA-CIB-IMP)
recognition method is proposed. This method uses CIB feature extraction to reduce the
complexity of the input without losing the feature information, and computes the global
feature similarity by DTW Barycenter Averaging (DBA). Generating samples based on the
original data set and expanding the data by generating a supplementary database through a
weighted average algorithm. The supplementary database is then fused with the original
database to complete the radar working mode recognition work by intelligent network
model. Higher recognition accuracies are achieved in scenarios with a limited training
samples. Recognition accuracy exceeds 90% at 0 dB SNR with low time spent.
Keywords: Radar working mode recognition; Multi-functional radar; Informer; Circularly
integrated bispectrum

Long Jin, Jiamin Pu, Hongjian Li,


FMCW radar-based heartbeat recognition using SE-DenseNet for vital signs monitoring,
AEU - International Journal of Electronics and Communications,
Volume 200,
2025,
155906,
ISSN 1434-8411,
[Link]
([Link]
Abstract: With the aging of society and people’s increasing concern about their health, non-
contact vital signs measurement provides new possibilities for home health monitoring. In
recent years, studies have shown that millimeter wave (mmW) radar has high sensitivity in
heart rate (HR) monitoring. However, traditional signal processing methods are susceptible
to environmental interference and they are difficult to accurately reconstruct heartbeat
signals. This paper proposes a heartbeat signal reconstruction method based on the SE-
DenseNet deep learning (DL) model. The model combines dense connections (DenseNet)
with the Squeeze-and-Excitation (SE) blocks to automatically extract and enhance key
physiological features in the two-dimensional feature matrix generated by radar signals,
thereby improving the accuracy of signal reconstruction. The experimental results
demonstrate that the proposed method achieves a HR estimation accuracy exceeding
97.68% in monitoring experiments at distances within 1.3 m, and its strong robustness has
been confirmed through tests with different subjects and during mild movement.
Keywords: Vital signs monitoring; Millimeter-wave (mmW) radar; Heart rate (HR); Signal
reconstruction; Deep learning (DL)
Keyu Pan, Wei-Ping Zhu, Mojtaba Hasannezhad,
Self-attention CNN based indoor human events detection with UWB radar,
Journal of the Franklin Institute,
Volume 361, Issue 14,
2024,
107090,
ISSN 0016-0032,
[Link]
([Link]
Abstract: In the era of smart homes and healthcare automation, the ability to accurately
monitor and detect indoor human activities is paramount. Ultra-wideband (UWB) radar has
emerged as a promising means for event detection, given its non-invasive nature and easy
deployment in diverse environments. However, despite the advances in radar-based event
detection, challenges remain, such as distinguishing between similar events like falls and
rapid sitting. To address these challenges, for the first time, we propose an impulse radio-
ultrawideband (IR-UWB) radar system to collect over ten thousand radar echo signals of
eight similar actions from different angles and design a self-attention-based low-complexity
convolutional neural network (CNN) model for event classification. The model leverages
global correlations in radar signal spectrograms to efficiently extract features. A comparative
simulation study is conducted to evaluate the detection accuracy of the proposed model and
some of the existing methods based on different dataset sizes and CNN configurations.
Moreover, the influence of different self-attention structures on precision and model
parameter count is analyzed. Our findings reveal that the proposed self-attention-based CNN
model significantly outperforms other traditional machine learning techniques while
maintaining a low level computational complexity.
Keywords: UWB radar system; CNN; Self-attention

Zilu Ying, Wenyu Ke, Yikui Zhai, Xinglin Liu, Jianhong Zhou, Pasquale Coscia, Angelo
Genovese,
Diffusion-augmented direct classification: A few-shot learning framework for Synthetic
Aperture Radar image automatic target recognition,
Engineering Applications of Artificial Intelligence,
Volume 166, Part B,
2026,
113648,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Deep learning-based (DL-based) synthetic aperture radar automatic target
recognition technology (SAR-ATR) has undergone extensive development, demonstrating
superiority over other competitive methods. However, the intrinsic requirement of deep
learning for a large labeled dataset restricts its practical application. Moreover, some DL-
based few-shot SAR-ATR methods are overly complex, hindering their deployment in real-
world applications. In addressing these obstacles, our solution introduces a straightforward
yet efficient few-shot learning approach titled Diffusion-Augmented Direct Classification for
few-shot SAR-ATR applications. The proposed method adopts a two-stage paradigm, where a
diffusion model first learns from unlabeled data and then produces synthetic samples to
train a recognition model. In the upstream stage, a lightweight diffusion-based image
generator build upon the shuffle-residual network structure is trained on a limited number
of annotated SAR images to generate artificial training samples for the downstream
recognition model. In the downstream stage, a Siamese network-based recognition model
and a similarity training procedure are proposed to train the model on a combination of real-
world and artificial samples, thereby improving recognition accuracy. A projection expansion
layer is proposed to improve the efficiency of cosine similarity loss in the downstream.
Experiments conducted on the Moving and Stationary Target Acquisition and Recognition
dataset demonstrated that our method outperformed other few-shot learning methods
concerning recognition accuracy in SAR-ATR tasks. Specifically, our method achieves over
73% accuracy in a 5-sample-per-class scenario and over 85% accuracy in a 10-samples-per-
class scenario. Source code of our paper is available at [Link]
Keywords: Synthetic Aperture Radar; Automatic Target Recognition; Denoising Diffusion
Probability model; Few-shot learning

Hui Wang, Qinghua Liu, Lijun Zhou,


Underground target localization method for ground penetrating radar based on deep
learning,
Measurement,
Volume 253, Part B,
2025,
117647,
ISSN 0263-2241,
[Link]
([Link]
Abstract: To tackle the challenge of subsurface target localization under interference in field
scenarios, a novel two-level cascade network referred to as dual cascade is proposed. The
first level, Cascade-1, is a deep feature extraction network designed to extract and eliminate
direct wave interference signals. On this basis, Cascade-2 is developed using domain
knowledge from ground penetrating radar as prior information, and it incorporates an
attention mechanism along with a feature fusion strategy to enhance the accuracy of target
feature hyperbola detection. Subsequently, the least squares method is employed to fit the
feature hyperbola, and location estimation is performed based on geometric equations. The
proposed cascade network model has demonstrated superior performance compared to
other algorithms, such as column-connection clustering algorithm, YOLOv9, and Faster R-
CNN, in terms of the composite metric F1, which validates the model’s effectiveness in
extracting the feature hyperbola. Additionally, the proposed localization method has
exhibited greater accuracy than the conventional full waveform inversion algorithm.
Keywords: Ground Penetrating Radar; Buried Target Location; Domain Knowledge; Deep
Learning; Cascade Network

Chuan Du, Long Tian, Bo Chen, Lei Zhang, Wenchao Chen, Hongwei Liu,
Region-factorized recurrent attentional network with deep clustering for radar HRRP target
recognition,
Signal Processing,
Volume 183,
2021,
108010,
ISSN 0165-1684,
[Link]
([Link]
Abstract: Feature extraction plays an essential role in radar automatic target recognition
(RATR) with high-resolution range profiles (HRRPs). Traditional feature extraction algorithms
usually ignore that different regions in HRRP contain the information with different
importance, resulting in their inadequacy in characterizing HRRP data. In this work, we
propose a region factorized recurrent attentional network (RFRAN) for HRRP-RATR by making
use of the temporal dependence through recurrent neural network (RNN) and automatically
finding the informative regions by a deep clustering mechanism in HRRP samples, which
reflects the distribution of scatterers in target along range dimension. Specifically, we
represent the temporal RNN hidden state using a region factorized encoder whose
parameters are conditioned on the HRRP region cluster centers. Moreover an attention
mechanism is used to weight up the different recognition contribution of each time step’s
hidden state. The aim of all the above modules is to achieve a more informative and
discriminative feature. Crucially, the loss function of RFRAN is differentiable, so all
components can be jointly trained with a gradient-based optimization. Compared with
traditional methods, besides the competitive recognition performance, RFRAN has a
promising interpretability thanks to the sequential region-specific hidden states.
Keywords: Region factorization; Clustering strategy; Attention mechanism; HRRP-RATR; RNN

Shuyu ZHENG, Libing JIANG, Qingwei YANG, Yingjian ZHAO, Zhuang WANG,
GS-orthogonalization OMP method for space target detection via bistatic space-based radar,
Chinese Journal of Aeronautics,
Volume 37, Issue 7,
2024,
Pages 333-351,
ISSN 1000-9361,
[Link]
([Link]
Abstract: A space-based bistatic radar system composed of two space-based radars as the
transmitter and the receiver respectively has a wider surveillance region and a better early
warning capability for high-speed targets, and it can detect focused space targets more
flexibly than the monostatic radar system or the ground-based radar system. However, the
target echo signal is more difficult to process due to the high-speed motion of both space-
based radars and space targets. To be specific, it will encounter the problems of Range Cell
Migration (RCM) and Doppler Frequency Migration (DFM), which degrade the long-time
coherent integration performance for target detection and localization inevitably. To solve
this problem, a novel target detection method based on an improved Gram Schmidt (GS)-
orthogonalization Orthogonal Matching Pursuit (OMP) algorithm is proposed in this paper.
First, the echo model for bistatic space-based radar is constructed and the conditions for
RCM and DFM are analyzed. Then, the proposed GS-orthogonalization OMP method is
applied to estimate the equivalent motion parameters of space targets. Thereafter, the RCM
and DFM are corrected by the compensation function correlated with the estimated motion
parameters. Finally, coherent integration can be achieved by performing the Fast Fourier
Transform (FFT) operation along the slow time direction on compensated echo signal.
Numerical simulations and real raw data results validate that the proposed GS-
orthogonalization OMP algorithm achieves better motion parameter estimation
performance and higher detection probability for space targets detection.
Keywords: Bistatic space-based radar; High-speed maneuvering space targets detection;
Range Cell Migration (RCM); Doppler Frequency Migration (DFM); Gram Schmidt (GS)-
orthogonalization Orthogonal Matching Pursuit (OMP) algorithm

Yonggang Qian, Yinghua Wang, Hongwei Liu, Zelong Wang, Feipeng Yu, Chunhui Qu,
MPRANet: Multi-scale perception and reference attention network for lightweight SAR target
recognition,
Neurocomputing,
Volume 668,
2026,
132310,
ISSN 0925-2312,
[Link]
([Link]
Abstract: Deep learning methods have been widely used in Synthetic Aperture Radar
Automatic Target Recognition (SAR ATR). However, challenges remain due to limited SAR
data and computational constraints on mobile devices, which hinder model training and
deployment. In this paper, we propose a Multi-scale Perception and Reference Attention
Network (MPRANet) for lightweight SAR ATR, which is a hybrid structure combining
convolutional networks and transformers, built upon the ShuffleNetV2 network. Specifically,
MPRANet introduces two key improvements compared to the CNN-based ShuffleNetV2.
Firstly, we replace the depthwise convolutions (DWConv) in the downsampling and basic
units of ShuffleNetV2 with the Multi-scale Parameter-Shared Convolution (MPConv) module.
MPConv enables the extraction of multi-scale features of SAR targets with almost no
additional parameters, thereby enhancing the network’s feature extraction capabilities.
Secondly, we propose a lightweight Reference Attention Transformer (RAformer) to capture
global information, addressing the issue of insufficient channel feature interaction in
ShuffleNetV2. In RAformer, a Local Linear Mapping Unit (LMU) is designed to perform linear
mappings, reducing the introduction of redundant features while ensuring its lightweight
and efficient nature. RAformer contains two modules: the Reference Vector Attention (RVA)
module, which efficiently models attention relationships, and the Lightweight Feedforward
Neural Network (LW-FFN) module, which enhances the network’s ability to capture
nonlinear representations. We evaluated the performance of MPRANet using publicly
available SAR datasets, including the MSTAR dataset, OpenSARShip dataset, and SAR-
AIRcraft-1.0 dataset. The experimental results demonstrate that MPRANet consistently
achieves superior recognition performance compared to other lightweight networks of
similar complexity.
Keywords: Synthetic aperture radar (SAR); Automatic target recognition (ATR); Convolutional
neural networks (CNN); Transformer; Lightweight

Jingpeng Gao, Sisi Jiang, Xiangyu Ji, Chen Shen,


Cross-domain prototype similarity correction for few-shot radar modulation signal
recognition,
Signal Processing,
Volume 223,
2024,
109575,
ISSN 0165-1684,
[Link]
([Link]
Abstract: The new classes of radar signals are increasingly difficult to acquire under non-
cooperative environments, which makes it difficult to support convolutional neural network
training with limited labeled samples. The few-shot learning (FSL) methods have shown
great performance in classification with limited labeled samples, but the FSL methods ignore
that the class distributions between the new and original tasks are significantly different,
resulting in a massive challenge in identifying new radar signals. To solve this problem, a
few-shot radar modulation signal recognition method based on cross-domain prototype
similarity correction (CDPSC) is proposed. Specifically, a residual feature tokenizer
transformer (RFTT) model embedded with a pooling token generation block is designed to
focus on the important features and improve the ability to represent samples. Meanwhile,
the proposed domain prototype similarity mapping (DPSM) strategy adaptively learns the
class mapping, reduces the inter-domain difference through feature distribution alignment,
and effectively corrects the target domain prototypes. In addition, we introduce a sample
prototype embedding (SPE) strategy in the training phase, which can reduce the intra-class
distance and increase the inter-class distance. Experimental results demonstrate that the
CDPSC method is superior to typical FSL methods in recognition accuracy under different
sample numbers.
Keywords: Cross-domain; Few-shot learning; Prototype similarity correction; Radar
modulation signal recognition

Xiaofang Pei, Yan Hu, Jun Zhu, Yun Dong, Peng Wang, Yinghua Ye, Ruiqi Shen,
Sustainable carbon-based materials for radar-infrared compatible stealth: Progress and
prospects,
Composites Part B: Engineering,
Volume 311,
2026,
113252,
ISSN 1359-8368,
[Link]
([Link]
Abstract: Stealth technology, as a cornerstone of modern defense and aerospace systems, is
increasingly challenged by multi-modal detection spanning radar, infrared, and emerging
sensor platforms. Carbon-based materials, owing to their lightweight nature, structural
tunability, and superior electromagnetic and thermal management capabilities, have
emerged as ideal candidates for radar-infrared compatible stealth applications. However,
traditional carbon sources demand energy-intensive processing and impose environmental
burdens that highlight the need for sustainable alternatives. This review provides a
comprehensive overview of recent advances in sustainable and environmentally friendly
carbon-based stealth materials, focusing on biomass (plant-, animal-, and microorganism-
derived), industrial by-products, and municipal or consumer wastes. Particular emphasis is
placed on the underlying mechanisms, highlighting that radar stealth originates from
dielectric and magnetic losses, while infrared stealth depends on temperature regulation
and emissivity control within the 3–5 μm and 8–14 μm atmospheric windows.
Microstructural engineering, heteroatom doping, and dielectric-magnetic synergy enable
broadband absorption at low filler loadings while providing tunable emissivity for infrared
suppression. Despite remarkable progress, key challenges remain in achieving simultaneous
broadband microwave absorption and low emissivity in critical infrared windows, as well as
ensuring structural uniformity and stability from complex renewable precursors. Looking
ahead, we propose a roadmap toward performance-sustainability-engineering synergy,
involving green precursor selection, scalable processing, multifunctional integration and
adaptive intelligent design. This Review thus bridges stealth performance with sustainability
imperatives, which provides strategic insights into the next generation of radar-infrared
compatible stealth materials.
Keywords: Carbon-based materials; Radar-infrared compatible stealth; Biomass; Waste
resources; Emissivity regulation

Do-Hyun Park, Min-Wook Jeon, Hyoung-Nam Kim,


Activity-dependent resolution adjustment for radar-based human activity recognition,
Signal Processing,
Volume 243,
2026,
110456,
ISSN 0165-1684,
[Link]
([Link]
Abstract: The rising demand for detecting hazardous situations has led to increased interest
in radar-based human activity recognition (HAR). Conventional radar-based HAR methods
predominantly rely on micro-Doppler spectrograms for recognition tasks. However,
conventional spectrograms employ a fixed resolution regardless of the varying
characteristics of human activities, leading to limited representation of micro-Doppler
signatures. To address this limitation, we propose a time-frequency domain representation
method that adaptively adjusts the resolution based on activity characteristics. This
approach adaptively adjusts the spectrogram resolution in a nonlinear manner, emphasizing
frequency ranges that vary with activity intensity and are critical to capturing micro-Doppler
signatures. We validate the proposed method by training deep learning-based HAR models
on datasets generated using our adaptive representation. Experimental results demonstrate
that models trained with our method achieve superior recognition accuracy compared to
those trained with conventional methods.
Keywords: Human activity recognition; Micro-doppler; Pattern analysis

Nima Roshandel, Constantin Scholz, Hoang-Long Cao, Milan Amighi, Hamed Firouzipouyaei,
Aleksander Burkiewicz, Sebastien Menet, Felipe Ballen-Moreno, Dylan Warawout Sisavath,
Emil Imrith, Antonio Paolillo, Jan Genoe, Bram Vanderborght,
mmPrivPose3D: A dataset for pose estimation and gesture command recognition in human-
robot collaboration using frequency modulated continuous wave 60Hhz RaDAR,
Data in Brief,
Volume 59,
2025,
111316,
ISSN 2352-3409,
[Link]
([Link]
Abstract: 3D pose estimation and gesture command recognition are crucial for ensuring
safety and improving human-robot interaction. While RGB-D cameras are commonly used
for these tasks, they often raise privacy concerns due to their ability to capture detailed
visual data of human operators. In contrast, using RaDAR sensors offers a privacy-preserving
alternative, as they can output point-cloud data rather than images. We introduce
mmPrivPose3D, a dataset of 3D RaDAR point-cloud data that captures human movements
and gestures using a single IWR6843AOPEVM RaDAR sensor with a frequency of 10 Hz
synchronized with 19 corresponding 3D skeleton keypoints as the ground truth. These
keypoints were extracted from RGB-D images captured by an Intel RealSense camera
recorded at 30 frames per second using the Nuitrack SDK, and labeled with gestures. The
dataset was collected from n = 15 participants. Our dataset serves as a fundamental
resource for developing machine learning algorithms to improve the accuracy of pose
estimation and gesture recognition using RaDAR data.
Keywords: Human-robot collaboration; IWR6843AOPEVM; RaDAR; Pose estimation; Gesture
command recognition

Hua Wang, Qiangyu Zeng, Hao Wang, Jianxin He, Tiantian Yu, Guangpu Liu,
Temporal super-resolution reconstruction of weather radar echoes using a deep learning
approach,
Expert Systems with Applications,
Volume 300,
2026,
130189,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Severe convective weather events are characterised by rapid evolution and high
destructive potential, requiring weather radars to provide observations with high temporal
resolution. However, current S-band weather radar systems, constrained by their volumetric
scanning strategies, often fail to capture the rapidly changing features of these systems
promptly. To address this limitation, we propose EMAIRA-VFI, a deep learning–based
method for temporal super-resolution reconstruction of radar echoes, which enhances the
temporal resolution of radar data to meet the demands of severe convective weather
monitoring. By introducing an inter-frame attention mechanism, the proposed method
effectively fuses spatiotemporal features from sequential radar echoes, enabling accurate
modelling of dynamic weather evolution and the generation of continuous, high-temporal-
resolution radar echoes. Compared with conventional temporal interpolation methods,
EMAIRA-VFI demonstrates significant improvements in both interpolation accuracy and the
preservation of fine-scale meteorological structures. Experimental results show that the
model not only enhances the capability of S-band radars in monitoring rapidly evolving
weather events but also provides a new perspective for spatiotemporal fusion and the
intelligent application of radar data. We have open-sourced the code for this work at
[Link]
Keywords: Temporal super-resolution; Radar echo; Inter-frame attention mechanism

Xiaolin Zhu, Dongli Wang, Yan Zhou, Zixin Zhang, Jianxun Li, Rui Su, Yongcan Weng, Tao Zhu,
Deep learning-based group activity recognition in videos: A survey,
Neurocomputing,
Volume 661,
2026,
131150,
ISSN 0925-2312,
[Link]
([Link]
Abstract: Group activity recognition (GAR), which aims to identify the activity performed by
a group of people in a given video, is one of the representative tasks for video
understanding. With the advancement of deep neural networks, deep learning-based
methods have emerged as a dominant alternative to handcrafted feature engineering, which
can automatically mine the feature representations from the input data. In this survey, we
present a comprehensive review with in-depth analysis of deep learning-based group
activity recognition from 2016 to 2024. Specifically, we first briefly outline the definition and
several major challenges of group activity recognition. Then, a detailed taxonomy is
introduced in terms of different supervision types, network types, modeling mechanism
types, and input types, which can better classify the existing state-of-the-art methods from
different perspectives. To the best of our knowledge, we are the first to develop such a
taxonomy for GAR. For a better understanding of the pros and cons of each type, the
detailed discussions under each type are also presented with further categorization.
Moreover, we also provide the datasets, evaluation metrics, and performance comparisons
for group activity recognition. In the end, we conclude the survey by suggesting future
potential research directions in this rapidly growing GAR field to facilitate new research
ideas.
Keywords: Group activity recognition (GAR); Deep learning; Modeling mechanism; Self-
supervised learning; Multimodal fusion; Graph convolutional network (GCN); Transformer

Yifan Chen, Meng Jiang, Chao Xia, Hang Zhao, Panpan Ke, Sheng Chen, Heng Ge, Keran Li, Xu
Wang, Yufei Wang, Yezi Chai, Qiming Liu, Zhengyu Tao, Yuyan Lyu, Yani Wu, Ao Shi, Yang Liu,
Hongyi Xin, Yu Zhong, Wei Zhang, Fuhua Yan, Weiwei Quan, Yingjia Xu, Dan Liu, Yumin Sun,
Xinli Li, Yuanyuan Tian, Lianming Wu, Shengxian Tu, Hongwei Ji, Bin Sheng, Jun Pu,
A novel deep learning system for STEMI prognostic prediction from multi-sequence cardiac
magnetic resonance,
Science Bulletin,
Volume 70, Issue 24,
2025,
Pages 4241-4252,
ISSN 2095-9273,
[Link]
([Link]
Abstract: ST-elevation myocardial infarction (STEMI) remains a leading cause of
cardiovascular morbidity and mortality worldwide, and accurate early risk stratification is
critical for implementing precision therapies in clinical practice. However, existing clinical risk
scores and manually derived imaging biomarkers have limited accuracy in predicting post-
STEMI outcomes. To address this gap, we developed DeepSTEMI, an end-to-end deep
learning system that integrates multi-sequence cardiac magnetic resonance (CMR) images
with clinical parameters for predicting 2-year major adverse cardiovascular events (MACE).
The system comprised two key algorithmic modules: a U-Net module that automatically
segments heart regions from raw CMR images and a Transformer-based module that
predicted future cardiovascular events. DeepSTEMI was developed using a multicenter
dataset (n = 610; 20,618 images) from STEMI patients enrolled in the EARLY-MYO-CMR
registry (NCT03768453), with external validation performed in 334 patients (9944 images)
from three independent cardiac centers. In external validation, DeepSTEMI demonstrated
superior predictive performance compared to conventional clinical risk scores and manual
CMR parameters (AUC 0.894, 95% CI: 0.823–0.965; overall accuracy 94.3%). The model
identified high-risk patients who exhibited a 20-fold MACE risk compared to low-risk
counterparts (HR 20.43, log-rank P < 0.001). SHapley Additive exPlanations (SHAP) analysis
revealed that DeepSTEMI’s predictive power stems from clinical-imaging synergy, enabling it
to capture complex pathological patterns. DeepSTEMI achieved consistently superior
performance over the Eitel score across all subgroups, with the greatest benefit observed in
women (NRI 1.597) and in patients imaged 4–7 d post-STEMI (NRI 1.442). Overall,
DeepSTEMI serves as an automated, scalable, and interpretable clinical copilot, which
advances post-STEMI risk stratification beyond the limitations of current paradigms.
Keywords: Myocardial infarction; Deep learning; Transformer; Prognostic prediction; Cardiac
magnetic resonance

Mohamed Aymen Ben Khalifa, Mourad El Koundi, Imed Riadh Farah,


Pushing boundaries in remote sensing: A comprehensive review of deep learning for spatial
super-resolution,
Remote Sensing Applications: Society and Environment,
Volume 40,
2025,
101809,
ISSN 2352-9385,
[Link]
([Link]
Abstract: Remote sensing image spatial super-resolution (RSISR) leverages deep learning to
overcome the limitations in spatial detail that are critical for applications such as precision
agriculture, environmental monitoring, and urban planning. Deep learning (DL) has
transformed RSISR, with advances in spatial detail reconstruction being driven by
convolutional neural networks (CNNs), generative adversarial networks (GANs),
transformers, and diffusion models (DDPMs). This comprehensive review synthesizes 749
papers from 2018 to 2025, assessing deep learning models, datasets, and techniques for
enhancing spatial resolution in remote sensing imagery. It presents a refined taxonomy of
DL-based RSISR models, critically comparing their accuracy, efficiency, and applicability
across the AID, DOTA, UC Merced, and NWPU-RESISC45 benchmark datasets using PSNR,
SSIM, and LPIPS. Convolutional neural networks (CNNs), notably residual and attention-
based architectures, have historically led to RSISR, yet generative models like DDPM are
increasingly prominent. Key challenges, including large-scale scenes, multi-band data,
atmospheric noise, and computational complexity, are assessed in relation to model
performance. By synthesizing trends and proposing future directions, including physics-
informed approaches, this review offers a rigorous foundation for advancing deep learning in
RSISR, supporting precision-driven Earth observation.
Keywords: Remote sensing; Super-resolution; Deep learning; Image enhancement; Spatial
super-resolution

Mingjun Cheng, Hong Jin, Qinfeng Zhao, Yurun Wang, Yanxi Wu, Shan Huang, Wenze Yue,
Deep learning for optimizing urban governance by "sensing-processing-responding" cycle:
Recent advances, future prospects and challenges,
Sustainable Cities and Society,
Volume 135,
2025,
106994,
ISSN 2210-6707,
[Link]
([Link]
Abstract: With accelerating urbanization, traditional governance models are increasingly
strained. Deep learning (DL) offers powerful solutions, but its application in urban
governance lacks a systematic framework and faces significant hurdles. This paper addresses
these gaps through a systematic review of 329 articles published from 2016 to 2025. We
introduce a novel Sensing-Processing-Responding framework to classify the technological
pathways of DL in urban governance. This framework organizes applications into three core
stages: (1) Sensing technologies (e.g., CNNs) for dynamic data acquisition; (2) Processing
technologies (e.g., RNNs, Transformers) for predictive modeling and analysis; and (3)
Responding technologies (e.g., LLMs) for automated decision support. Our analysis reveals
that while DL is widely applied in traffic forecasting, environmental monitoring, and disaster
response, its deployment is constrained by key challenges. We found a significant gap
between research and practice, with only 7.6% of studies demonstrating real-world
application. Furthermore, it concerns data privacy and model interpretability limit public
acceptance, although our review indicates that fewer than 10% of studies involve high-risk
personal data. Future progress depends on integrating emerging technologies like
multimodal large models and multi-agent systems. We conclude by advocating for a
paradigm shift from focusing purely on accuracy to prioritizing public value, fairness, and
transparency. This study provides a comprehensive roadmap for developing more intelligent,
resilient, and sustainable urban governance systems.
Keywords: Deep learning; Urban governance; Large models; Sensing-Processing-Responding
framework; Sustainable development; Literature review

Wenjun Hou, Hu Jin, Chuang Peng, Li Jiang,


A cognitive communication jamming strategy based on Transformer and Deep
Reinforcement Learning,
Computers and Electrical Engineering,
Volume 120, Part A,
2024,
109610,
ISSN 0045-7906,
[Link]
([Link]
Abstract: The advent of sophisticated communication technologies, such as cognitive radio
and anti-jamming techniques, has significantly elevated the challenge of disrupting enemy
communications. Nevertheless, the inherent openness of wireless communications remains
a vulnerability that can be exploited to interfere with them. Some contemporary
Reinforcement Learning (RL)-based jamming strategies examine methods for rapidly
identifying the optimal jamming strategy for a specific modulated signal. However, such
algorithms lack the flexibility and responsiveness required to effectively counter the enemy’s
evolving communication strategies. To address this issue, we propose a Transformer and
Deep Reinforcement Learning (DRL)-based jamming strategy that can be trained to identify
jamming methods for multiple digital and analog signals. In particular, the Transformer
Encoder is employed as a network for DRL to process the state information pertaining to the
enemy communication. Subsequently, the decision module of the Double Deep Q Network
(DDQN) is utilized to select the jamming action based on the processed information.
Furthermore, we have devised a reward function and constructed an invalid jamming list,
with the objective of selecting an action that requires low power consumption and enhances
the convergence speed of the algorithm. The experimental results demonstrate that the
algorithm proposed in this paper exhibits notable performance advantages in comparison to
other networks and DRL algorithms.
Keywords: Cognitive interference; Deep Reinforcement Learning; Transformer model;
Double Deep Q Network

Zheng Wang, Lei Wang, Yue-Chao Li, Lei Wang,


HMFCDA: Hierarchical deep learning with semantic embeddings for CircRNA-disease
association prediction,
Neurocomputing,
Volume 671,
2026,
132661,
ISSN 0925-2312,
[Link]
([Link]
Abstract: Recent studies have revealed the crucial roles of circRNAs in human disease
development, highlighting their significant potential as diagnostic biomarkers. However,
conventional biological experiments for identifying circRNA-disease associations remain
time-consuming and labor-intensive. To address this limitation, this study proposes a novel
hierarchical multimodal fusion framework, termed HMFCDA, which integrates hierarchical
deep learning with semantic embeddings to predict potential circRNA-disease associations.
The framework first employs a transformer model to encode primary vector representations
from circRNA sequences, which are subsequently refined through a graph attention network
to extract deeper features. Simultaneously, semantic representations for diseases are
generated using word embedding techniques. The circRNA deep features and disease
semantic embeddings are fused to form a joint representation, which is finally fed into a
Random Forest classifier for association prediction. Evaluated on the benchmark
CircR2Disease dataset, HMFCDA achieved an average prediction accuracy of 93.88% and an
AUC of 0.9767 under 5-fold cross-validation, significantly outperforming existing methods.
Furthermore, case studies demonstrated that 18 of the top 20 predicted circRNA-disease
associations were corroborated by recent literature evidence. These results underscore the
efficacy of HMFCDA as a powerful computational tool for circRNA-disease association
identification, providing reliable candidates and valuable insights for subsequent
experimental validation.
Keywords: CircRNA-disease associations; Hierarchical multimodal fusion; Transformer model;
Graph attention network; Random forest

Haiyan Yao, Yuefei Xu, Qiang Guo, Shizhe Chen, Bin Lu, Yuanjun Huang,
Study on transformer fault diagnosisbased on improved deep residual shrinkage network
and optimized residual variational autoencoder,
Energy Reports,
Volume 13,
2025,
Pages 1608-1619,
ISSN 2352-4847,
[Link]
([Link]
Abstract: The transformer as the core equipment in the power system, its fault diagnosis has
a vital role in ensuring the safe and stable operation of the power grid. However, traditional
transformer fault diagnosis methods often rely on manual experience or simple models,
which are difficult to meet the demand for efficient and accurate diagnosis when faced with
complex and evolving fault patterns. In this study, a new method for transformer fault
diagnosis based on improved deep residual shrinkage network (DRSN) and optimized
residual variational autoencoders (ORVAE) is proposed. Firstly, this study improves the DRSN
to enhance its feature extraction capability. By designing a specific shrinkage mechanism,
the improved DRSN can reduce the information loss in the face of complex data, greatly
improve the extraction ability of the key features of the transformer operating state, and
thus improve the accuracy of fault recognition. Secondly, in view of the difficulty and high
cost of transformer fault sample data collection, this study introduces a residual connection
structure based on the traditional variational autoencoder (VAE), and constructs the ORVAE
method to effectively address the challenge of insufficient data. The results show that the
fault recognition rate of the proposed method on the real transformer fault dataset reaches
97.14 %, which is better than the traditional method, showing excellent diagnostic
performance and strong practical application potential. Compared with the existing
technologies, this method not only improves the accuracy of transformer fault diagnosis, but
also provides new ideas and technical support for the intelligent development of power
system. This study offers an innovative solution for the field of fault diagnosis of power
equipment, and providing a strong technical guarantee for fault prediction and maintenance
in future smart grids.
Keywords: Transformer; Fault diagnosis; Improved DRSN; Shrinkage mechanism; Feature
extraction; ORVAE; Recognition rate

Purabi Sharma, Kandarpa Kumar Sarma,


Attention driven CWT-deep learning approach for discrimination of Radar PRI modulation,
Physical Communication,
Volume 62,
2024,
102237,
ISSN 1874-4907,
[Link]
([Link]
Abstract: With the proliferation of radio frequency (RF) systems and radar applications,
Electronic Warfare (EW) is receiving increasing importance. The analysis of the radar signals
is a critical EW task that decides the nature of counter employments. In an Electronic
Support (ES) system, the challenge is to detect hostile radiation sources efficiently and
trigger a counter response. Detection of types of Pulse Repetition Interval (PRI) modulation
of radar signal significantly facilitates the manifestation of RF emitters during recognition
which is difficult in a dense EW environment. Recent developments in artificial intelligence
(AI) methods suggest that this emerging technology can be effective for such purposes. In
this direction, an automatic approach for recognizing several kinds of complex PRI
modulation based on Continuous Wavelet Transform (CWT) and a combination of the vanilla
Convolutional Neural Network (CNN), a multi-head self-attention (MHSA) mechanism and
the popular Long Short-Term Memory (LSTM) is proposed. The CWT is used to decompose
the PRI modulation sequence and obtain different time–frequency components. Further,
aided by the proposed CNN-MHSA-LSTM combination, the features extracted from the CWT
2D-scalograms are used to execute PRI modulation discrimination. In this method, the
vanilla CNN is employed for the extraction of deep features to figure out the class details
while capturing the spatial attributes. Thereafter, to improve the discriminative power of the
entire framework a MHSA mechanism is used. The temporal attributes are acquired by the
LSTM which works in concert with the CNN for executing the detection of the PRI classes
based on the extracted features. Also to assess the effectiveness of the proposed method,
three models based on ResNet, popular CNN and SqueezeNet are implemented for
benchmark comparison in terms of overall performance and complexity. The simulation
results show that the proposed method enhances performance and achieves robustness in
the noise-filled and imperfect channel knowledge environment. The best recognition
accuracy is 98.3% with 50% spurious pulses in the environment which fluctuates with
imperfect channel knowledge cases.
Keywords: Electronic warfare; PRI modulation; CWT; Convolutional Neural Network; Self-
attention mechanism; Long Short Term Memory

Asim Saleem, Guoyun Lv, Safa Hussein Mohammed,


Deep Learning-Based Radar Fingerprinting for Open-Set Generalization Using Dynamic
Thresholding and Embedding Rejection,
Knowledge-Based Systems,
Volume 333,
2026,
115047,
ISSN 0950-7051,
[Link]
([Link]
Abstract: This paper presents a comprehensive framework for radar-specific emitter
identification (SEI), starting with the simulation of a large-scale radar signal dataset designed
to mimic real hardware impairments. By incorporating diverse distortions–such as phase
noise, frequency jitter, amplitude nonlinearity, and multipath reflections–for multiple radar
types and signal-to-noise ratio (SNR) conditions, we generate a realistic and challenging
dataset, SimRF-14, suitable for learning-based signal analysis. We utilized this dataset to
develop RAFNet, a hybrid deep learning model specifically designed for closed-set and open-
set radar emitter classification. The proposed architecture combines convolutional,
recurrent, and attention-based components to capture spatial, temporal, and contextual
features from normalized I/Q waveforms. For open-set recognition, we integrate the
OpenMax algorithm enhanced with Extreme Value Theory (EVT), where class-wise Weibull
modeling of embedding distances enables outlier detection. In addition, an SNR-adaptive
thresholding mechanism improves open-set reliability under varying noise conditions. The
proposed method achieved 97.43% unknown rejection at -20 dB SNR and maintained low
false positives (<4%) at high SNRs, validating its effectiveness and reliability for practical SEI
scenarios.
Keywords: Radar Specific Emitter Identification; Open-set Recognition; RF Fingerprinting;
SNR-Adaptive Classification; Deep Learning for SEI

Qing Snyder, Qingtang Jiang, Erin Tripp,


Integrating self-attention mechanisms in deep learning: A novel dual-head ensemble
transformer with its application to bearing fault diagnosis,
Signal Processing,
Volume 227,
2025,
109683,
ISSN 0165-1684,
[Link]
([Link]
Abstract: In this paper, we propose a novel dual-head ensemble Transformer (DHET)
algorithm for the classification of signals with time–frequency features such as bearing
vibration signals. The DHET model employs a dual-input time–frequency architecture,
integrating a 1D Transformer model and a 2D Vision Transformer model to capture the
spatial and time–frequency features. By utilizing data from both the time and time–
frequency domains, the proposed algorithm broadens its feature extraction capabilities and
enhances the model’s capacity for generalization. In our DHET structure, the original
Transformer model leverages self-attention mechanisms to consider relationships among
signal input segmentations, which makes it effective at capturing long-range dependencies
in signal data, while the Vision Transformer model takes 2D images as input and creates the
image patches for embedding and each patch is linearly embedded into a flat vector and
treated as a ‘token,’ then the ‘tokens’ are processed by the Transformer layers to learn global
contextual representations, enabling the model to perform signal classification task. This
integration notably enhances the performance and capability of the model. Our DHET is
especially effective for rolling bearing fault diagnosis. The simulation results show that the
proposed DHET has higher classification accuracy for bearing fault diagnosis and
outperforms CNN-based methods.
Keywords: Short-time Fourier transform; Transformer; Dual-head ensemble Transformer;
Deep learning; Bearing fault diagnosis

Abdul Hanan, Mehak Khan, Nieves Fernandez-Anez, Reza Arghandeh,


DeepBioFusion: Multi-modal deep learning based above ground biomass estimation using
SAR and optical satellite images,
Ecological Informatics,
Volume 90,
2025,
103277,
ISSN 1574-9541,
[Link]
([Link]
Abstract: Accurate estimation of forest above-ground biomass (AGB) is essential for
ecosystem conservation, sustainable forest management, and mitigating climate change and
wildfire risks. Traditional methods, such as manual field surveys, are labor-intensive and
limited in scope. This study presents DeepBioFusion, a multi-modal deep learning
framework that first estimates AGB for validation as ground truth generation by using LiDAR-
derived tree heights and a Tree Species map, employing allometry equations to relate tree
height to Diameter at Breast Height (DBH). After this initial estimation, the framework is
trained to predict AGB using high-resolution optical imagery and multiple bands of Synthetic
Aperture Radar (SAR), including X, C, and L bands. The use of SAR bands enables improved
canopy penetration, particularly in dense and cloud-covered forests. DeepBioFusion
leverages the complementary strengths of SAR and optical data to enhance the accuracy of
biomass predictions. Benchmarking against models like ResNet50 and Transformer, the
proposed model demonstrates superior performance in AGB estimation across diverse forest
environments. This study offers a scalable, cutting-edge approach to biomass monitoring,
advancing efforts in climate change mitigation and sustainable forest management.
Keywords: Above Ground Biomass; Synthetic Aperture Radar (SAR); Machine learning; Multi-
modal; Diameter at Breast Height (DBH); Optical and SAR fusion; Tree species; Allometry
equation; Remote sensing

Jiaqi Cai, Yiquan Wu,


Three-dimensional object detection for autonomous driving via deep learning: A review,
Engineering Applications of Artificial Intelligence,
Volume 161, Part C,
2025,
112238,
ISSN 0952-1976,
[Link]
([Link]
Abstract: With the rapid advancement of autonomous driving technology, there is an
increasing demand for highly accurate and real-time vehicle perception systems. From an
artificial intelligence (AI) perspective, three-dimensional (3D) object detection benefits from
recent progress in deep neural networks, convolutional neural networks (CNNs), and
transformer architectures, which provide powerful tools for feature extraction, spatial
reasoning, and multi-modal data fusion. These AI techniques enable robust, uncertainty-
aware predictions by effectively modeling complex sensor data. From an engineering
standpoint, 3D object detectors serve as critical components in autonomous driving systems
by translating AI-derived insights into real-time, high-accuracy vehicle perception. This paper
reviews the research progress of deep learning-based 3D object detection algorithms in
autonomous driving. First, commonly used data acquisition sensors are systematically
categorized, and the most widely adopted 3D detection datasets and evaluation metrics are
introduced; the standard 3D bounding-box representation and core network architectures
are also explained. Second, algorithms are classified according to input data type: 1) single-
modal methods, including vision-based, 3D data dimensionality reduction, point cloud-
based, transformer-based, mamba-based, and hybrid point cloud approaches; 2) multi-
modal methods, subdivided into serial fusion and parallel fusion strategies within network
pipelines. The characteristics, contributions, and limitations of each category are
summarized, and representative algorithms are compared across datasets to identify
research trends. Finally, current challenges, such as sparse data handling, domain shifts, and
real-time constraints, are examined, and prospective directions for future development are
proposed.
Keywords: Three-dimensional object detection; Autonomous driving; Deep learning;
Artificial intelligence; Transformer; Multi-modal fusion

Divya Kumawat, Ardeshir Ebtehaj, Sujay Kumar, Andreas Colliander,


Deep learning of seasonal peak snow water content of global boreal forest and arctic using
spaceborne L-band radiometry,
Remote Sensing of Environment,
Volume 330,
2025,
114963,
ISSN 0034-4257,
[Link]
([Link]
Abstract: Estimating peak snow water equivalent (SWE) across the Northern Hemisphere is
critical for assessing seasonal water availability for both ecosystems and human needs. This
study is the first to demonstrate a direct link between peak SWE and the temporal variability
of L-band surface emission under a moderately dense vegetation canopy. We introduce
SWEFormer, a novel deep transformer neural network that retrieves peak SWE primarily
using time series of L-band brightness temperatures from NASA’s Soil Moisture Active and
Passive (SMAP) satellite. The model is trained using an incremental learning approach that
transfers low-level information from reanalysis data for spatially coherent high-level learning
from sparse in situ observations. SWEFormer outperforms leading global products, including
ERA5, GlobSnow, and AMSR-based estimates, particularly in complex boreal watersheds,
where previous global SWE estimates suffer from significant uncertainties, as vegetation
canopy often markedly attenuates high-frequency microwave signatures of snowpack.
Keywords: L-band radiometry; Snow water equivalent; Deep transformer neural networks;
Boreal snow; SMAP satellite

Qinzhong Hou, Yonghao Yang, Jiatong Liang, Xiaoyan Huo, Junqiang Leng,
A deep transfer learning approach for Real-Time traffic conflict prediction with trajectory
data,
Accident Analysis & Prevention,
Volume 214,
2025,
107966,
ISSN 0001-4575,
[Link]
([Link]
Abstract: Recently, real-time traffic conflict prediction has drawn increasing attention due to
its significant potential in proactive traffic safety systems. While various statistical and
machine learning models have been developed for conflict prediction, transferability
remains a fundamental issue across these models. Specifically, the predictive performance of
a real-time conflict prediction model developed for a specific location can significantly
decline when directly applied to a new location without any modifications, primarily due to
substantial differences in traffic environments between these areas. To address this gap, this
study proposed a novel deep transfer learning approach aimed at enhancing the
transferability of real-time conflict prediction models. Initially, a real-time conflict prediction
framework was designed utilizing trajectory data for merging areas with consideration of
temporal variations in traffic flow characteristics. Subsequently, the Gated-Transformer, Fully
Convolutional Networks (FCN), Long Short-Term Memory Fully Convolutional Networks
(LSTM-FCN), and Multivariate Long Short-Term Memory Fully Convolutional Networks
(MLSTM-FCN) were employed as backbone feature extraction networks to capture the
hidden correlations between time-varying traffic flow characteristics and traffic conflicts.
After that, an independent transfer learning architecture was established to assess the
similarity of the distribution of traffic flow characteristics at different locations, based on the
maximum mean discrepancy criteria. For empirical evaluation, merging areas from the exiD
dataset were differentiated into source and target domains. The results demonstrated that
the Gated-Transformer model outperforms other baseline models (FCN, LSTM–FCN and
MLSTM–FCN) in both feature extraction and predictive performance, achieving an F1 score
of 0.864 and an area under the curve (AUC) of 0.980. Furthermore, the transfer learning
architecture can substantially enhance the predictive performance of a model trained in the
source domain when applied to the target domain. In particular, the F1 score and AUC for
the Gated-Transformer model improved by 11.9% and 10.2%, respectively, after
incorporating the transfer learning architecture. Finally, the optimal values of key model
parameters, including the sliding time window (6 s) and the prewarning time (5 s), were
recommended for practical applications through sensitivity analysis. This study illustrates the
potential of the deep transfer learning approach as a reliable and effective alternative to
improve the transferability of real-time conflict prediction models. Additionally, results from
this study can offer valuable insights for practical applications in traffic safety warning
systems, particularly in vehicle-to-infrastructure traffic environments.
Keywords: Real-time conflict prediction; Deep transfer learning; Gated-Transformer; Merging
area; Trajectory data

Sidra Ghayour Bhatti, Imtiaz Ahmad Taj, Mohsin Ullah, Aamer Iqbal Bhatti,
Transformer-based models for intrapulse modulation recognition of radar waveforms,
Engineering Applications of Artificial Intelligence,
Volume 136, Part B,
2024,
108989,
ISSN 0952-1976,
[Link]
([Link]
Abstract: The increasing prevalence of low probability of intercept (LPI) radars in electronic
warfare (EW) systems highlights the need to effectively recognize phase-coded radar
waveforms intercepted at radar warning receivers (RWRs) from various threat emitters. The
complexities of the electromagnetic (EM) spectrum necessitate the implementation of an
automatic modulation recognition system (AMRS) within the RWR. However, a major
challenge is accurately identifying phase-coded waveforms with high accuracy at low signal-
to-noise ratios (SNRs). This research addresses the challenge by exploring three artificial
intelligence (AI)-driven AMRS architectures for identifying phase-coded waveforms using
short-time Fourier transform (STFT): vision transformer (ViT), vicinity vision transformer
(VViT), and deep convolutional neural network (DCNN). Unlike recent methods focusing on
amplitude spectra, our research delves into the phase spectra for the feature extraction of
phase-coded waveforms. We leverage phase-based features extracted from intercepted
phase-coded waveforms to classify six types of phase-coded signals using these AMRS
architectures across SNR levels ranging from −16 dB to 8 dB. The simulation experiments
show that these methods are effective at an SNR of −16 dB, with VViT and ViT achieving
recognition accuracies of 93% and 92.7%, respectively. Both outperform the DCNN, which
achieves an RA of 89% at the same SNR. This approach promises to enhance situational
awareness and decision-making in EW operations by improving phase-coded radar
waveform recognition and enabling appropriate countermeasure deployment.
Keywords: Automatic modulation recognition system; Feature extraction; Low probability of
intercept; Short time Fourier transform

Wenqiang Hua, Yi Wang, Zihan Yang,


Knowledge and data co-driven deep learning model for PolSAR image classification,
Results in Engineering,
Volume 29,
2026,
108947,
ISSN 2590-1230,
[Link]
([Link]
Abstract: Deep learning method has become the mainstream technology for polarimetric
synthetic aperture radar (PolSAR) image interpretation in recent years. However, traditional
deep learning-based approaches to PolSAR image classification are mostly data-driven, and
the prediction results are highly dependent on the quantity of labeled samples, while
ignoring the potential physical characteristics of PolSAR data. Moreover, owing to the unique
microwave imaging mechanism in PolSAR technology, some targets with complex scattering
are difficult to distinguish visually. Therefore, a deep learning model driven by knowledge
and data is proposed for PolSAR image terrain classification. This model embeds physical
knowledge into deep neural network and is guided to learn high-level semantic features with
physical perception capabilities. Meanwhile, in the output space, according to the general
physical scattering laws followed by PolSAR terrain targets, the corresponding regularized
loss function is designed to constrain the network to learn. This design helps to train models
with limited training samples, promotes the rapid convergence of the deep learning model,
and improves the accuracy of PolSAR image classification under limited labeled samples.
Comprehensive experiments are performed on three widely utilized PolSAR datasets to
validate the efficacy and superior performance of the proposed approach.
Keywords: Deep learning; Images classification; Knowledge and data co-driven; Polarimetric
synthetic aperture radar

Dawei Li, Jingnan Wang, Kefeng Deng, Di Zhang, Chengwu Zhao, Hongze Leng, Yingfang Wen,
Yudi Liu, Kaijun Ren, Junqiang Song,
Review on deep learning quantitative precipitation nowcasting: Advances and challenges,
Expert Systems with Applications,
Volume 305,
2026,
130775,
ISSN 0957-4174,
[Link]
([Link]
Abstract: In recent decades, extreme precipitation events have increased dramatically due to
global warming, resulting in significant casualties and economic losses. Quantitative
Precipitation Nowcasting (QPN), which predicts precipitation intensity within a six-hour
timeframe, plays a critical role in public safety, infrastructure protection, transportation
management, outdoor event planning, and flood prevention systems. Building on successful
innovations in computer vision, deep learning advancements have substantially improved
prediction accuracy and transformed QPN methodologies. However, despite the
proliferation of research in this rapidly advancing field, comprehensive surveys that
systematically examine mainstream techniques and identify key challenges remain limited.
This paper provides a thorough review of current deep learning approaches in precipitation
nowcasting, examining important challenges, analyzing methodologies across the QPN
development lifecycle, and exploring promising research directions. Through systematic
synthesis of emerging developments, we aim to foster interdisciplinary collaboration and
stimulate continued innovation in this essential field.
Keywords: Quantitative precipitation nowcasting; Computer vision; Multidisciplinary
cooperation

Shenghua Lv, Xiaowei Zhang, Xuan Zhao, Meng Li, Jianghao Zhang, Chen Lin, Jian Wen,
Rapid and accurate assessment of filed scale soil moisture using ground-penetrating radar
deep learning-based inversion,
Measurement,
2026,
120594,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Accurate quantification of soil moisture content (SMC) is essential for sustaining
plant growth and maintaining ecosystem stability. Current SMC monitoring approaches are
subject to several constraints: satellite-based remote sensing frequently suffers from
inadequate spatial resolution for large-scale precision, while point-scale techniques are
incapable of effectively capturing soil moisture spatiotemporal variations and are unsuitable
for large-area monitoring. Ground Penetrating Radar (GPR), as an efficient and non-
destructive subsurface detection technique, has been widely employed for estimating soil
water content. However, existing methods based on full-waveform inversion exhibit strong
dependency and involve computationally expensive processes, resulting in insufficient
efficiency when handling large-scale GPR data. To overcome these challenges, this study
proposes a GPR deep learning-based inversion framework for rapid and precise SMC
estimation. A 900 MHz GPR system was deployed to survey an experimental site equipped
with pre-installed moisture sensors. The results demonstrated strong agreement between
GPR-derived SMC values and measurements obtained via Time Domain Reflectometry (TDR).
Furthermore, continuous high-temporal-resolution data collection verified the capability of
the method to characterize the spatiotemporal dynamics of SMC. Notably, for regional soil
moisture content (RSMC) estimation, the proposed method achieved a substantially lower
error (0.0019 m3) compared to conventional point-based measurements (0.0146 m3) and
stratified estimation (0.0122 m3). This methodology provides a robust technical foundation
for accurate field-scale SMC monitoring and exhibits significant potential for use in ecological
surveillance, precision agriculture irrigation, and sustainable water resource management.
Keywords: Ground-penetrating radar; Non-destructive testing; Soil moisture content; Deep
learning-based inversion; Spatio-temporal evolution

Bo Da, Xianyou Chang, Lilin Zhu, Hao Huang, Jiayang Ma, Weili Song, Da Chen,
Flexural behavior prediction of reinforced seawater sea sand concrete beams based on deep
learning techniques,
Construction and Building Materials,
Volume 495,
2025,
143576,
ISSN 0950-0618,
[Link]
([Link]
Abstract: To accurately assess the flexural behavior of reinforced seawater sea sand concrete
beams (RSSSCB), the influence of input features was investigated using an experimental
dataset with 163 ultimate bending moment (Mu) sets. For optimal model selection, deep
learning algorithms (CNN, Transformer, LSTM) were compared with traditional machine
learning methods (XGBoost, Random Forest) via comprehensive metrics: mean absolute
error, root mean square error, and coefficient of determination. This integrated methodology
enabled precise characterization of flexural behavior and quantitative interpretation of
feature importance correlations. The results indicate that: The CNN exhibited superior
predictive performance on both training and testing datasets compared to Transformer,
LSTM, XGBoost, and Random Forest algorithms, increased by 9.2 % and 29.33 %, 13.1 % and
11.49 %, −2.1 % and 16.87 %, 0 % and 24.36 %, respectively. This finding indicates that the
CNN model demonstrates a notable advantage in terms of prediction accuracy and
generalization capability. Furthermore, relative to the empirical formulations in GB 50010–
2010, JGJ 12–2006, and Da et al., the CNN model exhibited accuracy enhancements of 21 %,
25 %, and 19 %, respectively. Through the interpretation of the CNN model, it is found that
the cross-sectional area of longitudinal reinforcement in the tension zone and height of
concrete in the compression zone have a significant influence on its predictive performance.
Additionally, a noteworthy positive correlation was observed between rectangular cross-
section height and concrete height in the compression zone.
Keywords: Reinforced seawater sea sand concrete beams; Flexural behavior; Deep learning;
CNN; Predictive modeling

Kai Zhao, Zhongqi Sun, Hao Jiang, Zhixuan Zou, Qiong Wu, Yanjie Xin, Xiangru Liu, Huijie
Jiang,
Transformer-based integration of radiomics and deep learning for differentiating lipid-poor
adrenal adenomas from malignant tumors,
Meta-Radiology,
Volume 3, Issue 4,
2025,
100183,
ISSN 2950-1628,
[Link]
([Link]
Abstract: Purpose
To evaluate the effectiveness of a Transformer model based on contrast-enhanced computed
tomography (CECT) that integrates radiomics and deep learning features in differentiating
adrenal lipid-poor adenomas (LPA) and malignant tumors (MT).
Methods
This retrospective study included 282 patients with adrenal tumors from two medical
centers between October 2018 and October 2024. The patients were classified into adrenal
(LPA) and adrenal (MT) groups. Radiomics and deep learning features were extracted from
CECT images. A total of 240 patients from the first center were randomly divided into
Training Set and Test Set at a 7:3 ratio, while 42 patients from the second center served as an
External Validation Set. A Transformer algorithm was employed to integrate radiomics and
deep learning features for building predictive models. Its self-attention mechanism was
utilized to capture intrinsic associations within each feature type and to uncover hidden
information related to clinical outcomes. Additionally, a Radiomics model, a Deep Learning
model (DL_model), and a Traditional Combined model integrating radiomics and deep
learning features were constructed. Model performance was assessed using the area under
the receiver operating characteristic (ROC) curve (AUC) and radar chart. Calibration curves
and decision curve analysis (DCA) were employed to assess the predictive accuracy and
clinical net benefit of the models. Furthermore, radiomics feature activation maps and
gradient-weighted class activation mapping (Grad-CAM) were utilized to visualize radiomics
and deep learning features, respectively.
Results
The Transformer model achieved the best predictive performance in the training, test, and
external validation sets, with AUCs of 0.949, 0.917, and 0.852, respectively. The DeLong test
indicated that the performance differences between this model and the other models were
statistically significant. Furthermore, the radar chart illustrated that the Transformer model
achieved superior overall performance, and DCA confirmed its higher clinical net benefit
compared with the other models.
Conclusion
The Transformer model that integrates radiomics and deep learning features can accurately
distinguish between LPA and MT. Furthermore, the visual analysis of radiomics feature
activation maps and Grad-CAM intuitively illustrates the distribution of radiomics and deep
learning features, enhancing their potential for clinical application in preoperative
assessment of adrenal tumors.
Keywords: Lipid-poor adrenal adenomas; Computed tomography; Radiomics; Deep learning;
Transformer

Gabriela Czibula, Andrei Mihai, Paul-Dumitru Orăşan, Istvan Gergely Czibula, Eugen Mihuleţ,
Sorin Burcea,
SepConv-ens: An ensemble of separable convolution-based deep learning models for
weather radar echo temporal extrapolation,
Procedia Computer Science,
Volume 246,
2024,
Pages 666-675,
ISSN 1877-0509,
[Link]
([Link]
Abstract: The paper addresses the topic of radar echo temporal extrapolation which is of
major interest in both operational and research meteorology. Weather radar measurements
are an important data source used by operational meteorologists for weather analysis, radar
refectivity having a significant influence on short-term heavy rainfall prediction. Thus,
extrapolating radar products’ values is important for early storm evolution assessment. The
paper proposes SepConv-ens approach for temporal extrapolation of radar observations
using an ensemble of three separable convolution-based deep learning models. Experiments
performed on real radar data from the Romanian National Meteorological Administration
(NMA) highlight a good performance of SepConv-ens in predicting radar data up to more
than 40 minutes ahead and a good correlation between the radar measurements and the
predictions in terms of spatial and intensity evolution of the radar echoes. SepConv-ens is
integrated in the operational visualization software utilized by the Romanian NMA and is the
first attempt, at the national level, to offer an artificial intelligence-based automated
assistance for operational meteorologists.
Keywords: deep learning; convolutional neural network; separable convolution; nowcasting;
weather radar 2000 MSC: 68T07; 68T10

Kundan Meshram, Aryan Saurabh, Vinay Kharole, Chatrabhuj, Umank Mishra, Kennedy C.
Onyelowe, Viroon Kamchoom, Krishna Prakash Arunachalam,
Design of an integrated model for pothole detection and repair optimization using
multimodal transformers and hybrid deep learning,
Case Studies in Construction Materials,
Volume 23,
2025,
e05431,
ISSN 2214-5095,
[Link]
([Link]
Abstract: The detection and timely repair of potholes are crucial for maintaining road safety
and minimizing vehicle damage. However, existing methods often suffer from limitations
such as reliance on single-modal data, poor generalization across diverse environments, and
suboptimal resource management. To address these challenges, we propose a
comprehensive framework for enhanced pothole detection and repair optimization using
advanced deep learning techniques. Our approach integrates four key methodologies:
Multimodal Enhanced Pothole Detection with Person-Level Data (M-E-Pot holeNet), Hybrid
Machine Learning-Deep Learning for Classification (Hybrid-Pot holeNet), Deep
Reinforcement Learning for Pot hole Detection and Repair Optimization (DRL-Pot holeOpt),
and Transfer Learning for Pothole Detection in Diverse Environments (TL-Pot holeAdaptNet).
M-E-Pot holeNet employs a Self-Supervised Multimodal Transformer (SSMT) to fuse camera,
accelerometer, and crowdsourced smartphone data, achieving robust detection with a 97 %
accuracy and under 2 % false positive rate. Hybrid-Pot holeNet combines Graph Attention
Networks (GAT) and XGBoost, modeling spatial road features to classify potholes with 95 %
accuracy and an F1-Score of 0.92. DRL-Pot holeOpt uses Soft Actor-Critic (SAC) with Bayesian
Optimization to efficiently schedule repair tasks, reducing repair costs by up to 20 % and
crew travel time by 15–25 %. Finally, TL-Pot holeAdaptNet leverages Domain-Adversarial
Neural Networks (DANN) to ensure cross-domain adaptability, with 90 % accuracy in new
environments and a 40–50 % reduction in domain discrepancy. This multi-faceted approach
addresses the limitations of previous work by providing scalable, real-time, and resource-
optimized solutions for pothole detection and maintenance, offering significant
improvements in accuracy, cost efficiency, and adaptability.
Keywords: Pothole detection; Multimodal data; Hybrid deep learning; Reinforcement
learning; Transfer learning; Scenarios

Wonsu Kim, Chang-Hoo Jeong, Seongchan Kim,


Improvements in deep learning-based precipitation nowcasting using major atmospheric
factors with radar rain rate,
Computers & Geosciences,
Volume 184,
2024,
105529,
ISSN 0098-3004,
[Link]
([Link]
Abstract: Recently, deep learning-based precipitation nowcasting has been investigated and
its usefulness has been recognized. However, existing approaches have treated precipitation
nowcasting as a spatiotemporal sequence prediction problem and have mainly used only
radar images. Radar images show the distribution of water or ice droplets, but are limited in
providing information about the dynamic or thermodynamic processes involved in the
development of the rainfall system. In this study, we introduced a deep neural network that
can utilize atmospheric factors that play major roles in the development of rainfall systems
over the Korean Peninsula, including divergence at 925 hPa and total column water vapor, in
addition to radar images. We also proposed a loss function based on mean categorical
scores (e.g., critical success index and false alarm ratio) to minimize the performance error.
In a quantitative evaluation of 1- to 6-hour forecast results, our deep learning models
outperformed the conventional approach based on radar echo extrapolation and produced
higher equitable threat scores (ETSs) for moderate to severe rain (i.e., 5, 10, and 20 mm h−1
thresholds). By applying the proposed loss function and using the divergence at 925 hPa as
an additional input, the deep-learning model not only obtained the highest ETS values, but
also revealed its potential to predict new developments of some heavy rainfall systems. The
relationship between the presented deep learning-based precipitation nowcasting and the
distribution features of the additional atmospheric factors is discussed further in the case
study.
Keywords: Deep neural network; Encoder–forecaster; Nowcasting; Radar; Reanalysis

Jiangfan Feng, Xi Fu, Shaokang Dong,


A novel deep learning approach for high-precision rainfall intensity inversion using urban
surveillance audio,
Advances in Space Research,
Volume 77, Issue 2,
2026,
Pages 1648-1663,
ISSN 0273-1177,
[Link]
([Link]
Abstract: Accurate rainfall monitoring in urban environments remains challenging due to
sparse rain gauge networks, radar signal attenuation effects, and high deployment costs of
integrated remote sensing systems. Current audio-based approaches exhibit limitations in
capturing multi-scale rainfall patterns, modeling time–frequency dependencies, and
enforcing effective constraints for regression tasks. To address these critical gaps, we
propose MS-TF RainNet, a novel deep learning framework that enables high-precision
rainfall intensity inversion from urban surveillance audio. The architecture comprises three
synergistic components: (1) A hierarchical multi-scale feature extraction module processing
Mel Frequency Cepstral Coefficients (MFCC) through parallel convolutional branches with
varying receptive fields, enabling the simultaneous capture of local rainfall features and
global rainfall patterns. (2) A dual-domain attention mechanism combining temporal
attention for transient noise suppression and frequency attention for spectral feature
amplification. (3) Deep supervision with auxiliary regression heads enforcing hierarchical
feature consistency, mitigating gradient vanishing in deep networks through intermediate-
layer constraints. Evaluated on the Surveillance Audio Rainfall Intensity Dataset (SARID), MS-
TF RainNet achieves an RMSE of 0.7708 mm/h and R2 of 0.8196 under denoised conditions,
outperforming a Transformer-based baseline model by 14.94% in RMSE and 9.18% in R2. In
noisy environments, it maintains robustness with RMSE of 0.8443 mm/h and R2 of 0.7983.
This work presents a transformative advancement for urban hydrometeorology by leveraging
existing surveillance infrastructure, offering a cost-effective solution that outperforms
conventional methodologies while achieving fine-grained rainfall monitoring.
Keywords: Rainfall intensity inversion; Urban surveillance audio; Multi-scale feature
extraction; Time–frequency; Attention mechanism

Nirupam Das, Refat Noor Swarna, Md. Selim Hossain,


Deep learning-based circular disk type radar target detection in complex environment,
Physical Communication,
Volume 58,
2023,
102014,
ISSN 1874-4907,
[Link]
([Link]
Abstract: Target identification is one of the most popular radar uses in real life. Target
identification is a classifier that analyzes whether a signal contains an echo from a target
(target-present) or is merely noise (target-absent). Deep learning techniques are a popular
topic in classification, and they have evinced to be effective in a range of applications. In this
paper, a 64 layers Circular Disk type RADAR Target Detection (CDRTD) model is proposed
based on Transfer Learning using the SqueezeNet architecture of Convolutional Neural
Network (CNN) that functions directly with processed radar target return eco signal and
minimize the requirement of conventional laborious radar signal processing. Further, the
proposed 64 layers SqueezeNet-based CNN CDRTD model was then implemented to identify
circular disk type targets in complex environment. Finally, the target return eco data was
tested to identify the circular disk type radar target in complex environments. We further
analyzed target detection probability, false alarm rate, precision, recall, F1 in a complex
environment and compared it with the ideal case. We found that our proposed CDRTD
model can classify 83.3% of the test samples correctly with an overall accuracy of 94.59% in
a noisy and cluttered environment whereas 100% of the test samples are classified correctly
with an overall accuracy of 100% in an ideal environment.
Keywords: Deep learning; SqueezeNet-CNN; Circular disk target; Target detection; Complex
environment

Yiming Xiao, Ali Mostafavi,


DamageCAT: A deep learning transformer framework for typology-based post-disaster
building damage categorization,
International Journal of Disaster Risk Reduction,
Volume 128,
2025,
105704,
ISSN 2212-4209,
[Link]
([Link]
Abstract: Rapid, accurate, and descriptive building damage assessment is critical for directing
post-disaster resources, yet current automated methods typically provide only binary
(damaged/ undamaged) or ordinal severity scales. This paper introduces DamageCAT, a
framework that advances damage assessment through typology-based categorical
classifications. We contribute: (1) the BD-TypoSAT dataset containing satellite image triplets
from Hurricane Ida with four damage categories – partial roof damage, total roof damage,
partial structural collapse, and total structural collapse – and (2) a hierarchical U-Net-based
transformer architecture for processing pre- and post-disaster image pairs. Our model
achieves 0.737 IoU and 0.846 F1-score overall, with cross-event evaluation demonstrating
transferability across Hurricane Harvey, Florence, and Michael data. While performance
varies across damage categories due to class imbalance, the framework shows that typology-
based classification can provide more actionable damage assessments than traditional
severity-based approaches, enabling targeted emergency response and resource allocation.
Keywords: Damage assessment; Satellite imagery; Transformers; Damage description
Ali Jamali, Masoud Mahdianpari, Fariba Mohammadimanesh, Saeid Homayouni,
A deep learning framework based on generative adversarial networks and vision transformer
for complex wetland classification using limited training samples,
International Journal of Applied Earth Observation and Geoinformation,
Volume 115,
2022,
103095,
ISSN 1569-8432,
[Link]
([Link]
Abstract: Wetlands have long been recognized among the most critical ecosystems globally,
yet their numbers quickly diminish due to human activities and climate change. Thus, large-
scale wetland monitoring is essential to provide efficient spatial and temporal insights for
resource management and conservation plans. However, the main challenge is the lack of
enough reference data for accurate large-scale wetland mapping. As such, the main
objective of this study was to investigate the efficient deep-learning models for generating
high-resolution and temporally rich training datasets for wetland mapping. The Sentinel-1
and Sentinel-2 satellites from the European Copernicus program deliver radar and optical
data at a high temporal and spatial resolution. These Earth observations provide a unique
source of information for more precise wetland mapping from space. The second objective
was to investigate the efficiency of vision transformers for complex landscape mapping. As
such, we proposed a 3D Generative Adversarial Network (3D GAN) to best achieve these two
objectives of synthesizing training data and a Vision Transformer model for large-scale
wetland classification. The proposed approach was tested in three different study areas of
Saint John, Sussex, and Fredericton, New Brunswick, Canada. The results showed the ability
of the 3D GAN to stimulate and increase the number of training data and, as a result,
increase the accuracy of wetland classification. The quantitative results also demonstrated
the capability of jointly using data augmentation, 3D GAN, and Vision Transformer models
with overall accuracy, average accuracy, and Kappa index of 75.61%, 73.4%, and 71.87%,
respectively, using a disjoint data sampling strategy. Therefore, the proposed deep learning
method opens a new window for large-scale remote sensing wetland classification.
Keywords: Generative adversarial network; Convolutional neural network; Wetland
classification; New Brunswick; Vision Transformer (ViT); Deep learning

Yushu Zhang, Gang Li, Xiao-Ping Zhang, You He,


A deep learning model based on transformer structure for radar tracking of maneuvering
targets,
Information Fusion,
Volume 103,
2024,
102120,
ISSN 1566-2535,
[Link]
([Link]
Abstract: The motion complexity of maneuvering target causes the estimation uncertainty of
target motion model, resulting in state estimation error. Especially for strong maneuvering
target, the drastic change of target motion models makes the tracking methods hard to
adapt and provide accurate state estimation promptly. To solve the state estimation problem
of strong maneuvering targets, we propose a new transformer maneuvering target tracking
model based on deep learning, named TrMTT model. The TrMTT model uses a new residual
mapping between the observation trajectory and the real trajectory to estimate the target
states, and is composed of the encoder and decoder branches while the two have the same
input of observation trajectory. The encoder extracts the self-attention information for the
input at each layer while the decoder implements cross-attention extraction and fusion
between features in different layers, thus providing more correlation information between
states for learning the transition law of rapidly changing states. Moreover, we propose an
input module before the encoder–decoder structure to code the state features of the
observation trajectory, and apply two kinds of normalization layers in the input module and
the encoder–decoder structure, to project the input into a feature space which facilitates
extracting the correlation information between states. Simulation results show that the
proposed TrMTT model is superior in performance for maneuvering target tracking
compared with other existing approaches.
Keywords: Maneuvering target; Transformer; Radar tracking; Deep learning

Jiaquan Wan, Junchao Wang, Wei Zhang, Hao Song, Congyi Nai, Fengchang Xue, Tao Yang,
Chunxiang Shi, Quan J. Wang, Baoxiang Pan,
RadarDiT: An advanced radar echo extrapolation model for three gorges reservoir area via
diffusion transformer,
Journal of Hydrology: Regional Studies,
Volume 61,
2025,
102703,
ISSN 2214-5818,
[Link]
([Link]
Abstract: Study region
The Three Gorges Reservoir Area (TGRA)
Study focus
TGRA faces increasing vulnerability to extreme precipitation events driven by complex
convective weather systems. Radar echo extrapolation—predicting future precipitation
patterns from current radar data—is essential for early warning systems but faces significant
challenges in this topographically complex region. While data-driven approaches have
advanced the field, current convolutional neural network-based diffusion models struggle
with the TGRA's dynamic meteorological conditions due to their reliance on translational
invariance, which often fails to capture rapid weather transitions in complex terrain.
New hydrogeological insights from the region
To address these limitations, we introduce RadarDiT, a Vision Transformer-based diffusion
model specifically engineered for radar extrapolation in the TGRA. First, we develop a five-
year radar dataset capturing diverse convective weather phenomena unique to this region.
Then, leveraging this dataset, RadarDiT employs multi-layer Vision Transformers that
effectively model global dependencies and complex spatial relationships, enabling accurate
prediction of convective cell evolution. Our model demonstrates superior performance in
maintaining strong echo and spatial coherence over longer forecast horizons. Quantitative
evaluations across multiple metrics and thresholds confirm RadarDiT's enhanced skill in
forecasting heavy precipitation events, with particular improvements in Critical Success
Index at higher radar echo values. This work establishes a foundation for more reliable
nowcasting systems in regions with complex terrain and dynamic weather patterns, directly
supporting enhanced disaster preparedness and response strategies.
Keywords: Radar Echo Extrapolation; Three Gorges Reservoir Area; Diffusion Model; Vision
Transformer; Nowcasting

Xiaole Han, Jintao Liu, Jian Ye, Zihe Wang, Pengfei Wu, Hai Yang,
Deep Learning-Based GPR interpretation of soil thickness in headwater hillslopes,
Geoderma,
Volume 462,
2025,
117530,
ISSN 0016-7061,
[Link]
([Link]
Abstract: Soil thickness strongly influences eco-hydrological and geomorphic processes, yet
conventional measurements such as auger drilling are invasive, labor-intensive, and
unsuitable for large-scale surveys. Ground-penetrating radar (GPR) provides a non-invasive
alternative, but its manual interpretation remains slow and prone to observer bias. To
address this challenge, we developed a fully automated framework that couples a hybrid
CNN-Transformer deep learning architecture with optimized signal filtering to predict soil
thickness directly from GPR profiles. The convolutional layers extract local waveform
features, while the attention mechanism captures long-range dependencies. Using field data
from a steep headwater hillslope (H1) in the Taihu Basin, China, we compared five filtering
strategies—median, Savitzky-Golay, Gaussian, moving average, and none—and found that
median filtering yielded the most accurate results (R2 up to 0.92, CCC of 0.96, RMSE near
10 cm). We further identified optimal filter window sizes (61–101 samples) and a training
duration threshold (≥500 epochs) that ensured stable and accurate predictions. Cross-site
validation on an independent hillslope (H2) without retraining showed that the pretrained
CNN-Transformer model achieved the highest R2 (0.80), CCC (0.89), and lowest RMSE
(11.3 cm), outperforming traditional machine learning models (CNN, MLP, RF, SVM) in
transferability. These findings demonstrate that integrating CNN-Transformer architectures
with appropriate signal filtering enables scalable, accurate, and objective soil thickness
mapping in complex terrain. The proposed approach also holds promise for broader GPR-
based subsurface applications, including soil horizon delineation and root system detection.
Keywords: Ground-penetrating radar; Soil thickness; Transformer; Headwater hillslopes;
Median filtering

Enzhe Sun, Yongchuan Cui, Peng Liu, Jining Yan,


A decade of deep learning for remote sensing spatiotemporal fusion: Advances, challenges,
and opportunities,
Information Fusion,
Volume 126, Part B,
2026,
103675,
ISSN 1566-2535,
[Link]
([Link]
Abstract: Remote sensing spatiotemporal fusion (STF) addresses the fundamental trade-off
between temporal and spatial resolution by combining high temporal-low spatial and high
spatial-low temporal imagery. This paper presents the first comprehensive survey of deep
learning advances in remote sensing STF over the past decade. We establish a systematic
taxonomy of deep learning architectures including Convolutional Neural Networks (CNNs),
Transformers, Generative Adversarial Networks (GANs), diffusion models, and sequence
models, revealing significant growth in deep learning adoption for STF tasks. Our analysis
reveals that CNN-based methods dominate spatial feature extraction, while Transformer
architectures show superior performance in capturing long-range temporal dependencies.
GAN and diffusion models demonstrate exceptional capability in detail reconstruction,
substantially outperforming traditional methods in structural similarity and spectral fidelity.
Through comprehensive experiments on seven benchmark datasets comparing ten
representative methods, we validate these findings and quantify the performance trade-offs
between different approaches. We identify five critical challenges: time-space conflicts,
limited generalization across datasets, computational efficiency for large-scale processing,
multi-source heterogeneous fusion, and insufficient benchmark diversity. The survey
highlights promising opportunities in foundation models, hybrid architectures, and self-
supervised learning approaches that could address current limitations and enable
multimodal applications. The specific models, datasets, and other information mentioned in
this article have been collected in: [Link]
Spatiotemporal-Fusion-Survey.
Keywords: Spatiotemporal fusion; Deep learning; Remote sensing; Literature review

Lixing Shi, Xueling Liang, Wenchao Chen, Yaoqiang Liu, Tong Ding, Kun Qin, Bo Chen,
Hongwei Liu,
Masked variational transformer for complex clutter modeling and target detection,
Signal Processing,
Volume 239,
2026,
110236,
ISSN 0165-1684,
[Link]
([Link]
Abstract: Weak target detection commonly encounters intense clutter interference, which
overshadows weak signals and complicates the task. Taking advantage of the powerful data
mining capability of neural networks, more and more deep learning-based methods are
applied to radar target detection. Among the approaches, those founded upon unsupervised
learning methodologies exhibit remarkable merit because they dispense with the
requirement for target samples within the training step, making them highly applicable in
practical target detecting scenarios. However, existing methods suffer from limitations in
leveraging the range-Doppler (R-D) two-dimensional correlation and finely modeling in
multiple clutter scenarios. In this paper, an unsupervised Transformer-based detector (TrDet)
is proposed to break through the boundary of modeling capability. First, with the designed
two-dimensional position embedding (2-DPE) and global query embedding (GQE)
techniques, an unsupervised training strategy for R-D spectrum based on Transformer
framework is utilized to achieve refined clutter modeling. Then, radar target detection is
formulated as an out-of-distribution (OOD) detection task to mitigate clutter interference.
Moreover, the masked variational Transformer-based detector (MVTrDet) is further
proposed to prevent target information leakage when the target is in close proximity to the
clutter in Doppler domain. Compared with several relative algorithms, our proposed
methods are better suited for radar target detection in complex clutter environments. The
experimental results derived from both measured data and simulated data verify the
effectiveness of our proposed methods.
Keywords: Radar target detection; Clutter modeling; Range-Doppler (R-D) spectrum;
Unsupervised learning; Out-of-distribution detection; Transformer

Binyu Xiong, Yuntian Chen, Dali Chen, Jun Fu, Dongxiao Zhang,
Deep probabilistic solar power forecasting with Transformer and Gaussian process
approximation,
Applied Energy,
Volume 382,
2025,
125294,
ISSN 0306-2619,
[Link]
([Link]
Abstract: Solar power generation encounters instability and unpredictability issues due to
the uncertainty of weather changes. Consequently, probabilistic forecasting of solar power is
essential for the effective management and integration of solar energy into the power grid,
substantially enhancing the reliability and efficiency of the electrical system. Among various
methods, time series analysis for probabilistic forecasting, which leverages historical data to
predict future solar power generation, has become a significant area of research due to
advancements in deep learning. However, existing methods often fall short in accuracy and
operational efficiency. This paper introduces an innovative deep learning framework tailored
for probabilistic forecasting of solar power generation. Considering the unique distribution
characteristics of solar power data, a novel data preprocessing method integrating Box–Cox
and Z-score transformations is applied to the input time series data. Subsequently, a novel
probabilistic time series forecasting method, leveraging a Transformer network enhanced
with Gaussian process approximation, predicts solar power generation for the forthcoming
24 h. The delta method is then employed to reverse transform the forecasts into actual
predicted values. Comparative analyses using a real-world solar power dataset demonstrate
that the proposed model outperforms existing probabilistic forecasting networks in
deterministic, probabilistic, and interval forecasting tasks. Compared to the commonly used
probabilistic forecasting method MC Dropout, our method decreases the CRPS index by
22.6% on the Shenzhen dataset and 39.7% on the Xingtai dataset. Furthermore, the
proposed model exhibits superior computational efficiency, reflecting an optimal balance
between accuracy and computational demands.
Keywords: Probabilistic forecasting; Solar power; Transformer network; Gaussian process
approximation

Yu Zhao, Yifang Ban,


RADARSAT constellation mission compact polarisation SAR data for burned area mapping
with deep learning,
International Journal of Applied Earth Observation and Geoinformation,
Volume 141,
2025,
104615,
ISSN 1569-8432,
[Link]
([Link]
Abstract: Monitoring wildfires has become increasingly critical due to the sharp rise in
wildfire incidents in recent years. Optical satellites like Sentinel-2 and Landsat are
extensively utilised for mapping burned areas. However, the effectiveness of optical sensors
is compromised by clouds and smoke, which obstruct the detection of burned areas. Thus,
satellites equipped with Synthetic Aperture Radar (SAR), such as dual-polarisation Sentinel-1
and quad-polarisation RADARSAT-1/-2 C-band SAR, which can penetrate clouds and smoke,
are investigated for mapping burned areas. However, there is limited research on using
compact polarisation (compact-pol) C-band RADARSAT Constellation Mission (RCM) SAR data
for this purpose. This study aims to investigate the capacity of compact polarisation RCM
data for burned area mapping through deep learning. Compact-pol m-χ decomposition and
Compact-pol Radar Vegetation Index (CpRVI) are derived from the RCM Multi-Look Complex
product. A deep-learning-based processing pipeline incorporating ConvNet-based and
Transformer-based models is applied for burned area mapping, with three different input
settings: using only log-ratio dual-polarisation intensity images, using only compact-pol
decomposition plus CpRVI, and using all three data sources. The training dataset comprises
46,295 patches, generated from 12 major wildfire events in Canada. The test dataset
includes seven wildfire events from the 2023 and 2024 Canadian wildfire seasons in Alberta,
British Columbia, Quebec and the Northwest Territories. The results demonstrate that
compact-pol m-χ decomposition and CpRVI images significantly complement log-ratio
images for burned area mapping. The best-performing Transformer-based model, UNETR,
trained with log-ratio, m-χ m-decomposition, and CpRVI data, achieved an F1 Score of 0.718
and an IoU Score of 0.565, showing a notable improvement compared to the same model
trained using only log-ratio images (F1 Score: 0.684, IoU Score: 0.557). This is the first study
to demonstrate that RCM C-band SAR data and its derived features are effective for burned
area mapping.
Keywords: RADARSAT constellation mission; Burned area mapping; SAR; Compact
polarisation; Decomposition; Radar vegetation index; Deep learning

Najamuddin, Usman Ullah Sheikh, Ahmad Zuri Sha’ameri,


Ensemble deep learning approach for marine vessel classification: Integrating CNN and
vision transformers with machinery feature enhancement,
Information Fusion,
Volume 126, Part A,
2026,
103570,
ISSN 1566-2535,
[Link]
([Link]
Abstract: Accurate classification of marine vessels is a challenging task due to the dynamic
nature of the underwater environment and the complexity of vessel acoustic signatures.
Though recent deep learning studies have significantly improved classification accuracies
under ideal conditions, their performance degrades in noisy conditions due to
environmental complexities and weak acoustic signals. This study proposes an ensemble
deep learning framework combining Convolutional Neural Networks (CNN) and Vision
Transformers (ViT) to enhance classification performance. First, weak machinery
narrowband frequency components are extracted using Coherently Averaged Power
Spectrum Estimation (CAPSE), improving feature representations. Then, a classifier using a
CNN module captures local spatial features, while the ViT models long-range dependencies
from Low Frequency Analysis and Recording (LOFAR), ensuring a more comprehensive
understanding of the acoustic patterns of onboard machinery. By leveraging the
complementary strengths of both architectures, the proposed model effectively improves
classification. The ensemble framework achieved an average accuracy of 98.32%, and F1-
score of 98.31% on DeepShip and 99.50% accuracy and F1-score on ShipsEar dataset, an
improvement over several state-of-the-art approaches while maintaining around 3.5 million
learnable parameters. Moreover, the ensemble approach demonstrates robustness across
varying noise conditions, maintaining an accuracy of over 60% at 0 dB where individual CNN
or ViT models show performance degradation. These results highlight the effectiveness of
the CNN-ViT fusion strategy in passive acoustic classification, offering a strong balance
between accuracy and efficiency. The model’s lightweight architecture makes it particularly
suitable for real-time applications in vessel monitoring, maritime surveillance, and
underwater target classification.
Keywords: Time–frequency representations (TFR); Signal-to-noise ratio; LOFAR gram;
Convolutional Neural Networks; Vision transformer

Xiaofang Sun, Meng Wang, Junbang Wang, Guicai Li, Xuehui Hou,
Deep learning classification of winter wheat from Sentinel optical-radar image time series in
smallholder farming areas,
Advances in Space Research,
Volume 75, Issue 3,
2025,
Pages 2683-2695,
ISSN 0273-1177,
[Link]
([Link]
Abstract: As crop yield stagnation, climate change, and the rising demand for agricultural
products pose increasing challenges, mapping crop systems is becoming more and more
important. Winter wheat is one of the major cereal crops cultivated in China, ranking as the
third largest crop in terms of production and harvested area. Accurately mapping winter
wheat is necessary for implementing effective farm management practices. While many
studies have successfully produced high spatiotemporal resolution land cover maps,
relatively few map products of crop types are available in China. The growing archive of
satellite image time series provides enormous opportunities to map crops more closely. This
research presents a two-step method to map winter wheat based on Sentinel-1 and
Sentinel-2 time-series data from Shandong Province using the deep learning approaches.
The winter crops were firstly mapped using time-series optical vegetation indices employing
the deep learning methods. Then winter wheat was extracted from the winter crops mask by
coupling optical and synthetic aperture radar time-series images. The results indicated that
the precision of mapping winter wheat using Temporal Convolution Neural Networks
(TempCNN) achieved the highest precision in mapping winter wheat, with an overall
accuracy of 93.7 %, a kappa coefficient of 0.907, and an F1-score of 0.989. This was followed
sequentially by the Residual 1D convolutional neural networks (ResNet), the Multi-Layer
Perceptron (MLP), and the Lightweight Temporal Self-Attention Encoder (L-TAE). The
Temporal Attention Encoder (TAE) model demonstrated the lowest precision among the
compared models. The results agree well with independent county-level official census
winter wheat area data (R2 = 0.936). The proposed framework can also be applied in other
regions to generate maps of different crops, so future work can extend the proposed model
to other agricultural regions, where an increased number of crop types and natural
vegetation types can be included and tested.
Keywords: Sentinel-1; Sentinel-2; Classification; Winter wheat mapping; Time series; Deep
learning

Wenyu Wang, Chenyang Wang, Libo Zhang, Yuchen Yan, Linxiu Wang, Jin Guo,
Monitoring mining-induced subsidence from satellite imagery using transformer-based deep
learning trained on gridded subsidence measurements,
Journal of Environmental Management,
Volume 394,
2025,
127536,
ISSN 0301-4797,
[Link]
([Link]
Abstract: The inherent concealment of underground coal mining makes it difficult for
environmental protection authorities to detect and regulate illicit activities. These mining
activities are only identified after severe environmental damage has occurred, such as
farmland flooding or structural cracks in residential buildings. By the time enforcement
actions are taken, the opportunity for early intervention is lost, and ecological restoration
becomes nearly impossible. Using artificial intelligence (AI) to analyse satellite imagery for
monitoring land subsidence in coal mining-affected areas is considered a promising solution.
However, two major research gaps remain unresolved. First, the lack of ground-truth
subsidence measurements limits the amount of training data available for AI models.
Second, traditional convolutional neural network (CNN) architectures, such as VGGNet and
ResNet, often fail to achieve satisfactory classification accuracy in this context. In this study,
a Vision Transformer (ViT-Base) model was trained using 191,630 land subsidence grid
measurements paired with high-resolution satellite images. The model achieved an overall
accuracy of 94 % in identifying land subsidence in the region corresponding to the training
data. To further evaluate its generalizability, ten representative mining-affected areas were
selected from China’s top ten coal-producing provinces, each providing 250 independent
subsidence grid measurements paired with high-resolution satellite imagery. The overall
accuracies obtained were ranging from 77.2 % to 84.8 %. These results demonstrate that ViT-
Base consistently outperforms conventional models in identifying mining-induced land
subsidence from satellite imagery, maintaining high accuracy across diverse geographic and
geological settings while requiring less training data. The proposed model thus addresses key
research gaps and provides a practical tool for the monitoring and management of mining-
induced land subsidence.
Keywords: Mining-induced land subsidence; Vision transformer (ViT); Satellite imagery
analysis

Xiaoqing Hu, Hongyan Zhu,


MFF-MTT: a multi-feature fusion-based deep learning algorithm for maneuvering target
tracking,
Information Fusion,
Volume 130,
2026,
104093,
ISSN 1566-2535,
[Link]
([Link]
Abstract: In target tracking applications, traditional model-driven algorithms suffer from the
model mismatch due to the lack of prior knowledge. Recently, some data-driven algorithms
have been showing increasing potential in dealing with uncertain target maneuvering
behaviors. To further enhance robustness to high maneuverability, we propose a multi-
feature fusion-based deep learning algorithm for maneuvering target tracking (MFF-MTT) by
combining the convolution and transformer network. Thereinto, the convolution network
extracts the local information to capture the transition law of rapidly changing states. The
Multi-Head Self-Attention (MHSA) in transformer network enables MFF-MTT to exploit the
global information by weighting different parts of input sequence and integrating diverse
subspace representations of queries, keys, and values. The local and global features are then
fused in two forms of merge and cross to capture the short-term maneuvers and long-term
trends of the trajectory jointly. Moreover, we also develop a novel encoder-decoder
framework that decodes the fused features by Bi-directional Long Short-Term Memory (Bi-
LSTM). In this way, a comprehensive understanding about the inherent structure of the data
can be obtained to facilitate the high-accuracy state estimation. Extensive simulation results
demonstrate that the proposed MFF-MTT outperforms other comparative methods on
estimation precision and robustness in maneuvering target tracking scenarios.
Keywords: Maneuvering target tracking; Multi-feature fusion; Transformer; Convolution; Bi-
LSTM; Trajectory dataset

Jinsheng Fan, Guo-An Yu, Mingmeng Zhao, Hucheng Zong,


Addressing multi-scale temporal variability: deep integration and application of the CNN and
transformer model in monthly streamflow prediction,
Expert Systems with Applications,
Volume 292,
2025,
128658,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Accurate monthly streamflow prediction is essential for effective water resource
management, hydropower operation, and ecological sustainability. However, streamflow
processes are inherently nonlinear and exhibit considerable multiscale temporal variability,
driven by both natural conditions and potential anthropogenic influences. To address these
challenges, we propose a novel hybrid deep learning model, ISVM-CovTransformer, which
integrates the Improved Sparrow Search Algorithm (ISSA), Variational Mode Decomposition
(VMD), Mutual Information (MI), and a composite CovTransformer architecture. Within this
framework, ISSA is utilized to optimize the parameters of VMD for efficient signal
decomposition, while MI is employed to identify informative input features with strong
predictive relevance. The CovTransformer model, combining Convolutional Neural Networks
(CNN) and Transformer layers, enables the simultaneous extraction of localized temporal
patterns and long-range dependencies, thereby enhancing the model’s ability to capture
complex runoff dynamics and improve prediction accuracy. Using monthly precipitation and
streamflow data from the Tangnaihai, Toudaoguai, and Huayuankou hydrological stations,
experimental results demonstrate that the proposed model outperforms baseline
approaches. Specifically, during the testing phase, the model achieved an NSC of 0.9686,
RMSE of 91.99 m3/s, MAE of 70.90 m3/s, R2 of 0.9702, and a PBIAS of −1.198 % at
Tangnaihai; an NSC of 0.9498, RMSE of 90.35 m3/s, MAE of 724.63 m3/s, R2 of 0.9554, and a
PBIAS of 2.573 % at Toudaoguai; and an NSC of 0.9302, RMSE of 174.36 m3/s, MAE of
44.42 m3/s, R2 of 0.9393, and a PBIAS of 3.309 % at Huayuankou. These findings confirm the
proposed model’s effectiveness for monthly streamflow forecasting and suggest that it
provides a theoretically sound and generalizable framework, with potential extensions to
related hydrological applications such as sediment transport modeling.
Keywords: Monthly streamflow prediction; VMD; CNN; Transformer; CovTransformer

Jiahao Liu, Yiming Zhang, Liang Song, Zheng Tong,


SCB-ADAE: An attention-based deep autoencoder for ground penetrating radar signal
denoising,
Engineering Applications of Artificial Intelligence,
Volume 160, Part A,
2025,
111902,
ISSN 0952-1976,
[Link]
([Link]
Abstract: In buried object detection, recorded signals of a ground penetrating radar (GPR)
inevitably include noise interference owing to complex underground environments. Existing
rule- and data-driven denoising methods struggle to handle non-Gaussian and real-world
noise because the rule-driven ones rely on the assumptions of simplified noise
characteristics and the data-driven ones cannot capture fine- and global-scale features of a
GPR signal well. To address the problem, this study proposes an attention-based denoising
model called the Swin-Conv Block with Attention Denoising Autoencoder (SCB-ADAE). The
model first feeds a GPR signal into a SCB module, which extracts a tensor with the fine-scale
features in the signal, such as sharp reflective interfaces and abrupt amplitude variations.
The feature tensor then passes through an ADAE module that uses encoder-decoder
structure with the self-attention to enhances the representation of the global-scale signal
features. Finally, the feature tensor from the ADAE module is decoded by another SCB
module to generate a denoised GPR signal, where the tensor includes the fine-scale and
global features of the raw signal. An experiment with three types of GPR signals
demonstrates the effectiveness of the proposed model: radar signals with Gaussian noise,
radar signals with inhomogeneous-material noise, and real-world signals. radar signals with
Gaussian noise, radar signals with inhomogeneous-material noise, and real-world signals.
Experimental results demonstrate that the proposed model outperforms other state-of-the-
art denoising methods on denosing the three types of GPR signals, where the signal-to-noise
ratio, peak signal-to-noise ratio, and structural similarity index are improved to 20.64, 14.59,
and 0.366, respectively.
Keywords: Ground penetrating radar; Denoising; Attention-based model; Autoencoder

Hongping Zhou, Chengwei Zhang, Peng Peng, Zhongyi Guo,


MambaMTT: A deep learning method based on mamba structure for maneuvering target
tracking,
Signal Processing,
Volume 239,
2026,
110285,
ISSN 0165-1684,
[Link]
([Link]
Abstract: The research area of maneuvering target-tracking in radar system has emerged as a
critical and valuable research frontier. The diverse and unpredictable movements of
maneuvering targets make it hard to estimate their state accurately. This challenge often
makes previous methods unreliable in dealing with maneuvering targets, especially the
highly maneuvering ones. Although Transformer-based models possess global modeling
capabilities, they encounter computational challenges when applied to long trajectory
sequences due to their inherent computational complexity. To address the problem of highly
maneuvering targets tracking, this paper proposes a Mamba-based maneuvering target
tracking algorithm, termed as MambaMTT. MambaMTT is specifically designed to model
trajectory change patterns by focusing on local and global feature correlation information of
the trajectory sequence, facilitating the effective processing of highly maneuvering targets.
Meanwhile, we introduce a multi-scale feature fusion module to capture spatial and
temporal correlations at different scales within the trajectory sequence. With this module,
the MambaMTT network is able to capture local and global trajectory features more
efficiently, improving the model’s adaptability and accuracy for complex maneuvering
targets. Experimental results show that the proposed MambaMTT algorithm exhibits higher
tracking efficiency and accuracy in various maneuvering target tracking, and it also has
better generalization ability on data beyond the training range.
Keywords: Maneuvering target-tracking; Radar system; Mamba; Spatial and temporal
correlations

Ghada Atteia, Mohammed Dabboor, Konstantinos Karantzalos, Maali Alabdulhafith,


Hybrid Deep Learning and Optimized Feature Selection for Oil Spill Detection in Satellite
Images,
Computers, Materials and Continua,
Volume 84, Issue 1,
2025,
Pages 1747-1767,
ISSN 1546-2218,
[Link]
([Link]
Abstract: This study explores the integration of Synthetic Aperture Radar (SAR) imagery with
deep learning and metaheuristic feature optimization techniques for enhanced oil spill
detection. This study proposes a novel hybrid approach for oil spill detection. The introduced
approach integrates deep transfer learning with the metaheuristic Binary Harris Hawk
optimization (BHHO) and Principal Component Analysis (PCA) for improved feature
extraction and selection from input SAR imagery. Feature transfer learning of the MobileNet
convolutional neural network was employed to extract deep features from the SAR images.
The BHHO and PCA algorithms were implemented to identify subsets of optimal features
from the entire feature dataset extracted by MobileNet. A supplemented hybrid feature set
was constructed from the PCA and BHHO-generated features. It was used as input for oil spill
detection using the logistic regression supervised machine learning classification algorithm.
Several feature set combinations were implemented to test the classification performance of
the logistic regression classifier in comparison to that of the proposed hybrid feature set.
Results indicate that the highest oil spill detection accuracy of 99.2% has been achieved
using the logistic regression classification algorithm, with integrated feature input from
subsets identified using the PCA and the BHHO feature selection techniques. The proposed
method yielded a statistically significant improvement in the classification performance of
the used machine learning model. The significance of our study lies in its unique integration
of deep learning with optimized feature selection, unlike other published studies, to
enhance oil spill detection accuracy.
Keywords: Oil spill; machine learning; deep learning; classification; metaheuristic
optimization

Yuwen Wu,
Fusion-based modeling of an intelligent algorithm for enhanced object detection using a
Deep Learning Approach on radar and camera data,
Information Fusion,
Volume 113,
2025,
102647,
ISSN 1566-2535,
[Link]
([Link]
Abstract: Object detection, the process of detecting and classifying objects within a given
environment, forms the foundational element. Multisensory fusion incorporates data from
diverse sensors, like radar and cameras, to refine the reliability and accuracy of detection.
Further, Radar and camera data fusion refine this process by integrating the unique strength
of both technologies, which leverage the radar's proficiency in adverse weather conditions
and the camera's high-resolution imaging. This incorporation enhances the object detection
systems, which enables them to effectively operate across the spectrum of scenarios, from
autonomous vehicles navigating challenging weather to surveillance systems monitoring
critical infrastructure. Deep learning (DL), a branch of machine learning (ML), empowers this
system with the capability to learn complex representations and patterns directly from the
data, which enables them to generalize and adapt to new situations. By integrating the
advanced methodology, we can develop strong perception system capable of interpreting
and detecting objects accurately in dynamic and diverse environments, from autonomous
vehicles navigating urban landscapes to surveillance systems monitoring complex
environments. This study designs an Intelligent Algorithm for Enhanced Object Detection
Using Deep Learning Approach on the Radar and Camera Data Fusion (IAEOD-DLRCDF)
technique. The presented IAEOD-DLRCDF technique uses multi-angle joint calibration where
the spatial sparse alignment of the heterogeneous data of the camera and Radar is realized
with image falsification disregarded. Besides, the IAEOD-DLRCDF technique applies YOLOv8
object detector for radar and camera target detection individually which are then integrated
with the image plane. Moreover, the detected objects are then classified via the
bidirectional long short-term memory (BiLSTM) model. Furthermore, the Adam optimizer is
used for the optimum hyperparameter selection of the BiLSTM network which results in a
better recognition rate. The performance assessment of the IAEOD-DLRCDF method is tested
under benchmark dataset. The empirical analysis stated that the IAEOD-DLRCDF method
gains better performance over other models.
Keywords: Object detection; Deep learning; Data fusion; Radar; YOLOv8; Adam optimizer;
Machine learning

Mahya G.Z. Hashemi, Ehsan Jalilvand, Hamed Alemohammad, Pang-Ning Tan, Narendra N.
Das,
Review of synthetic aperture radar with deep learning in agricultural applications,
ISPRS Journal of Photogrammetry and Remote Sensing,
Volume 218, Part A,
2024,
Pages 20-49,
ISSN 0924-2716,
[Link]
([Link]
Abstract: Synthetic Aperture Radar (SAR) observations, valued for their consistent
acquisition schedule and not being affected by cloud cover and variations between day and
night, have become extensively utilized in a range of agricultural applications. The advent of
deep learning allows for the capture of salient features from SAR observations. This is
accomplished through discerning both spatial and temporal relationships within SAR data.
This study reviews the current state of the art in the use of SAR with deep learning for crop
classification/mapping, monitoring and yield estimation applications and the potential of
leveraging both for the detection of agricultural management practices. This review
introduces the principles of SAR and its applications in agriculture, highlighting current
limitations and challenges. It explores deep learning techniques as a solution to mitigate
these issues and enhance the capability of SAR for agricultural applications. The review
covers various aspects of SAR observables, methodologies for the fusion of optical and SAR
data, common and emerging deep learning architectures, data augmentation techniques,
validation and testing methods, and open-source reference datasets, all aimed at enhancing
the precision and utility of SAR with deep learning for agricultural applications.
Keywords: SAR; Deep learning; Crop classification; Phenology; Yield prediction; Agricultural
management practice

Guozheng Wang, Qinzhe Lv, Liyi Liu, Rong Yang, Bowen Bie, Yaojun Wu, Yinghui Quan,
An end-to-end deep learning framework for separation and parameter measurement of
composite intermittent sampling repeater jamming,
Aerospace Science and Technology,
Volume 168, Part H,
2026,
111218,
ISSN 1270-9638,
[Link]
([Link]
Abstract: To address the severe challenge posed by composite intermittent sampling
repeater jamming (ISRJ) to aerospace radar, we propose a novel end-to-end deep learning
framework for simultaneous signal separation and parameter estimation. At the core of this
framework is a custom deep separation network (CISRJ-SN), which features a unique hybrid
attention architecture. This architecture synergistically fuses one-dimensional convolution
for local feature extraction with a gated attention unit for global dependency modeling,
thereby achieving high-fidelity jammer signal separation even at a low Jammer-to-Noise
Ratio (JNR). Results from both simulations and real-world hardware-in-the-loop experiments
collectively validate the superior performance of our framework. In two-component and
multi-component scenarios, it achieves accuracy rates of 99.5 % and 94.4 % respectively,
marking a significant improvement over the next-best methods. In addition to high accuracy,
the framework demonstrates exceptional estimation precision, reducing the Mean Absolute
Error (MAE) of key parameters by over 60 %. This proves the high stability and reliability of
its estimation results, offering a promising solution for future intelligent sense-and-
countermeasure closed-loop systems.
Keywords: Composite intermittent sampling repeater jamming; Parameter estimation; Deep
learning; Radar anti-jamming; Attention mechanism

Elise Colin, Nicolas Trouvé,


4 - Deep learning for artificial SAR image generation: Generative AI for SAR,
Editor(s): Michael Schmitt, Ronny Hänsch,
Deep Learning for Synthetic Aperture Radar Remote Sensing,
Elsevier,
2026,
Pages 75-97,
ISBN 9780443363443,
[Link]
([Link]
Abstract: This chapter explores methods for generating synthetic Aperture Radar (SAR)
images, tracing the evolution from traditional physics-based simulators to modern
generative AI techniques. After reviewing classical approaches — based on geometrical
optics, ray tracing, and Maxwell-based models — we highlight their limitations in terms of
scalability and realism. This has led to the rise of data-driven methods using deep generative
models. GANs, diffusion models, and text-to-image transformers now offer the ability to
synthesize realistic SAR scenes without explicitly modeling radar signal physics. The chapter
introduces a modular methodology including the construction of training multimodal
datasets, automated captioning, optical-to-SAR translation with GANs, fine-tuning of
diffusion models, and domain-specific evaluation of synthetic outputs. Practical examples
demonstrate how these components can be used for simulation, data augmentation, and
super-resolution. We conclude by discussing the future of SAR image generation, where
hybrid models combining physical modeling with learned representations may emerge to be
able to handle the unique characteristics of SAR data.
Keywords: Simulation; Generative networks; cGAN; Transformer; Latent diffusion

Shulin Pang, Zhanqing Li, Lin Sun, Biao Cao, Zhihui Wang, Xinyuan Xi, Xiaohang Shi, Jing Xu,
Jing Wei,
Enhancing cloud detection across multiple satellite sensors using a combined Swin
Transformer and UPerNet deep learning model,
Remote Sensing of Environment,
Volume 334,
2026,
115206,
ISSN 0034-4257,
[Link]
([Link]
Abstract: Cloud detection is crucial in many applications of satellite remote sensing data.
Traditional cloud detection methods typically operate at the pixel level, relying on
empirically tuned thresholds or, more recently, machine learning classification schemes
based on training datasets. Motivated by the success of the Transformer with its self-
attention mechanism and convolutional neural networks for enhanced feature extraction,
we propose a new encoder-decoder method that captures global and regional contexts with
multi-scale features. This new model takes advantage of two advanced deep-learning
techniques, the Swin Transformer and UPerNet (named STUPmask), demonstrating
improved cloud detection accuracy and strong adaptability to diverse imagery types,
spanning spectral bands from visible to thermal infrared and spatial resolutions from meters
to kilometers, across a wide range of surface types, including bright scenes such as ice and
desert, globally. Training and validation of the STUPmask model are conducted using data
obtained from the Landsat 8 and Sentinel-2 Manually Cloud Validation Mask datasets on a
global scale. STUPmask accurately estimates cloud amount with a marginal difference
against reference masks (0.27 % for Landsat 8 and −0.81 % for Sentinel-2). Additionally, the
model captures cloud distribution with a high overall classification accuracy (97.51 % for
Landsat 8 and 96.27 % for Sentinel-2). Notably, it excels in detecting broken, thin, and semi-
transparent clouds across diverse surfaces, including bright surfaces like urban and barren
lands, especially with acceptable accuracy over snow and ice. These encompass the majority
of challenging scenes encountered by cloud identification methods. It also adapts to cross-
sensor satellite data with varying spatial resolutions (4 m–2 km) from both Low-Earth-Orbit
(LEO) and Geostationary-Earth-Orbit (GEO) platforms (including GaoFen-2, MODIS, and
Himawari-8), with an overall accuracy of 94.21–97.11 %. The demonstrated successes in the
automatic identification of clouds with a variety of satellite imagery of different spectral
channels and spatial resolutions render the method versatile for a wide range of remote
sensing studies.
Keywords: Cloud detection; Cross-sensor; STUPmask; Swin Transformer; UPerNet

Farhana Ahmed Chowdhury, Md Kamal Hosain, Md Sakib Bin Islam, Md Shafayet Hossain,
Promit Basak, Sakib Mahmud, M. Murugappan, Muhammad E.H. Chowdhury,
ECG waveform generation from radar signals: A deep learning perspective,
Computers in Biology and Medicine,
Volume 176,
2024,
108555,
ISSN 0010-4825,
[Link]
([Link]
Abstract: Cardiovascular diagnostics relies heavily on the ECG (ECG), which reveals significant
information about heart rhythm and function. Despite their significance, traditional ECG
measures employing electrodes have limitations. As a result of extended electrode
attachments, patients may experience skin irritation or pain, and motion artifacts may
interfere with signal accuracy. Additionally, ECG monitoring usually requires highly trained
professionals and specialized equipment, which increases the treatment's complexity and
cost. In critical care scenarios, such as continuous monitoring of hospitalized patients,
wearable sensors for collecting ECG data may be difficult to use. Although there are issues
with ECG, it remains a valuable tool for diagnosing and monitoring cardiac disorders due to
its non-invasive nature and the detailed information it provides about the heart. The goal of
this study is to present an innovative method for generating continuous ECG waveforms
from non-contact radar data by using Deep Learning. The method can eliminate the need for
invasive or wearable biosensors and expensive equipment to collect ECGs. In this paper, we
propose the MultiResLinkNet, a one-dimensional convolutional neural network (1D CNN)
model for generating ECG signals from radar waveforms. With the help of a publicly
accessible radar benchmark dataset, an end-to-end DL architecture is trained and assessed.
There are six ports of raw radar data in this dataset, along with ground truth physiological
signals collected from 30 participants in five distinct scenarios: Resting, Valsalva, Apnea, Tilt-
up, and Tilt-down. By using strong temporal and spectral measurements, we assessed our
proposed framework's ability to convert ECG data from Radar signals in three distinct
scenarios, namely Resting, Valsalva, and Apnea (RVA). ECG segmentation performed better
by MultiResLinkNet than by state-of-the-art networks in both combined and individual
cases. As a result of the simulations, the resting, valsalva, and RVA scenarios showed the
highest average temporal values, respectively: 66.09523 ± 19.33, 60.13625 ± 21.92, and
61.86265 ± 21.37. In addition, it exhibited the highest spectral correlation values
(82.4388 ± 18.42 (Resting), 77.05186 ± 23.26 (Valsalva), 74.65785 ± 23.17 (Apnea), and
79.96201 ± 20.82 (RVA)), along with minimal temporal and spectral errors in almost every
case. The qualitative evaluation revealed strong similarities between generated and actual
ECG waveforms. As a result of our method of forecasting ECG patterns from remote radar
data, we can monitor high-risk patients, especially those undergoing surgery.
Keywords: ECG; Raw radar data; MultiResLinkNet; CNN; Deep learning

Farhad Mortezapour Shiri, Thinagaran Perumal, Norwati Mustapha, Raihani Mohamed,


Deep Learning and Federated Learning in Human Activity Recognition with Sensor Data: A
Comprehensive Review,
CMES - Computer Modeling in Engineering and Sciences,
Volume 145, Issue 2,
2025,
Pages 1389-1485,
ISSN 1526-1492,
[Link]
([Link]
Abstract: Human Activity Recognition (HAR) represents a rapidly advancing research domain,
propelled by continuous developments in sensor technologies and the Internet of Things
(IoT). Deep learning has become the dominant paradigm in sensor-based HAR systems,
offering significant advantages over traditional machine learning methods by eliminating
manual feature extraction, enhancing recognition accuracy for complex activities, and
enabling the exploitation of unlabeled data through generative models. This paper provides
a comprehensive review of recent advancements and emerging trends in deep learning
models developed for sensor-based human activity recognition (HAR) systems. We begin
with an overview of fundamental HAR concepts in sensor-driven contexts, followed by a
systematic categorization and summary of existing research. Our survey encompasses a wide
range of deep learning approaches, including Multi-Layer Perceptrons (MLP), Convolutional
Neural Networks (CNN), Recurrent Neural Networks (RNN), Long Short-Term Memory
networks (LSTM), Gated Recurrent Units (GRU), Transformers, Deep Belief Networks (DBN),
and hybrid architectures. A comparative evaluation of these models is provided, highlighting
their performance, architectural complexity, and contributions to the field. Beyond
Centralized deep learning models, we examine the role of Federated Learning (FL) in HAR,
highlighting current applications and research directions. Finally, we discuss the growing
importance of Explainable Artificial Intelligence (XAI) in sensor-based HAR, reviewing recent
studies that integrate interpretability methods to enhance transparency and trustworthiness
in deep learning-based HAR systems.
Keywords: Human activity recognition (HAR); machine learning; deep learning; sensors;
Internet of Things; federated learning (FL); explainable AI (XAI)

Yuhao Wu, Bin Li, Jun Li, Yonglou Liang, Naiqiang Zhang, Anlai Sun,
Enhancing nighttime cloud detection for moderate resolution imagers using a transformer
based deep learning network,
Remote Sensing of Environment,
Volume 332,
2026,
115067,
ISSN 0034-4257,
[Link]
([Link]
Abstract: Accurate cloud detection is essential for the quantitative applications of satellite
imager observations, but nighttime cloud detection has challenges due to limited spectral
bands, for example, physical methods using only infrared (IR) bands without using spatial
textures as input for cloud detection often result in high uncertainties, especially in some
situations such as cryosphere surface. Although numerous segmentation-style deep learning
cloud detection algorithms have proposed in previous studies, they are inadequate for
nighttime due to the difficulty in acquiring two-dimensional truth data for training and
validation. To overcome these challenges, the Transformer based Nighttime Cloud Detection
(TNCD) framework, which integrates spatial features and utilizes an advanced Transformer
architecture with relative position encoding, layer scaling, and channel attention
mechanisms, is proposed and investigated for nighttime cloud detection. The model was
trained on labels derived from CALIOP data, utilizing a dataset comprising nearly one
hundred million segments from MODIS. Independent validation indicates that TNCD
achieves robust and consistent performance across various scenarios, with an overall
accuracy (OA) of 93.26 % and over 90 % in cryosphere regions. The proposed algorithm
avoids the pattern noise appeared in the traditional physical methodology due to the
utilization of auxiliary data at coarser resolutions, it also mitigates the negative impact of
stripes in IR images for cloud detection. Moreover, TNCD shows high transferable
practicability across sensors, with over 90 % OA for MERSI. More importantly, our research
underscores the importance of water vapor absorption bands for nighttime cloud detection
over the cryosphere. TNCD's high accuracy and robustness provide unique methodology that
could be used operationally for nighttime cloud detection.
Keywords: Nighttime cloud detection; Transformer; Deep learning; MODIS; MERSI; CALIPSO

Linh Trinh, Siegfried Mercelis, Ali Anwar,


A comprehensive review of datasets and deep learning techniques for vision in unmanned
surface vehicles,
Ocean Engineering,
Volume 334,
2025,
121501,
ISSN 0029-8018,
[Link]
([Link]
Abstract: Unmanned Surface Vehicles (USVs) have emerged as a major platform in maritime
operations, capable of supporting a wide range of applications. USVs allow for difficult
unmanned tasks in harsh maritime environments. With the rapid development of USVs,
many vision tasks such as detection and segmentation become increasingly important.
Datasets play an important role in encouraging and improving the research and
development of reliable vision algorithms for USVs. In this regard, a large number of recent
studies have focused on the release of vision datasets for USVs. Along with the development
of datasets, a variety of deep learning techniques have also been studied, with a focus on
USVs. However, there is a lack of a systematic review of recent studies in both datasets and
vision techniques to provide a comprehensive picture of the current development of vision
on USVs, including limitations and trends. In this study, we provide a comprehensive review
of both USV datasets and deep learning techniques for vision tasks. Our review was
conducted using a large number of vision datasets from USVs. We elaborate several
challenges and potential opportunities for research and development in USV vision based on
a thorough analysis of current datasets and deep learning techniques.
Keywords: Unmanned surface vessels; Datasets; Deep learning; Computer vision

Yizhen Jia, Hui Chen, Bang Huang, WenKai Jia, Wen-Qin Wang,
Riemannian gradient deep network for joint waveform and filter optimization in MIMO radar
against chopping forwarding jamming and clutter,
Signal Processing,
Volume 239,
2026,
110257,
ISSN 0165-1684,
[Link]
([Link]
Abstract: With the rise of digital radio frequency memory technology, active deception
jamming poses a significant threat to radar systems, especially in detecting targets amid
mainlobe jamming and non-Gaussian clutter. Traditional methods like space–time matched
filtering struggle in such scenarios. This study introduces the Riemannian gradient deep
network (RGDN), a framework for joint optimization of transmit waveforms and receive
filters to improve target detection. Unlike conventional signal-to-clutter noise ratio (SCNR)
maximization, RGDN leverages information geometry to maximize the Kullback–Leibler
Divergence (KLD) between targets and clutter. By modeling non-Gaussian data with a
Gaussian mixture distribution and constructing a Riemannian manifold, the framework
achieves effective jamming suppression through receive filter term in the loss function,
minimizing jamming effects while enhancing target-clutter distinguishability. To address non-
convex optimization, Riemannian gradient descent is integrated into a deep network.
Numerical experiments show that RGDN achieves superior detection performance compared
to SCNR maximization method.
Keywords: Riemannian gradient; KL divergence; Waveform design; MIMO radar; Mainlobe
deception jamming; Deep learning

Wenkai Qiu, Haolong Chen, Huanlin Zhou,


Real-time identification of temperature-dependent thermo-mechanical properties of specific
pressure vessels via deep learning,
Engineering Analysis with Boundary Elements,
Volume 180,
2025,
106514,
ISSN 0955-7997,
[Link]
([Link]
Abstract: Pressure vessels are widely utilized in industrial production, and accurate
prediction of their thermoelastic properties is crucial for ensuring their safe and stable
operation. Consequently, this study proposes a novel fusion deep learning algorithm to
achieve rapid and accurate identification of the temperature-dependent thermo-mechanical
properties of specific pressure vessels for 3D transient thermoelastic problems. This model
integrates three network architectures: convolutional neural network, bidirectional long
short-term memory network and Transformer, significantly enhancing feature extraction
efficiency, parallel computing and nonlinear fitting abilities. Solving coupled thermoelastic
problem through finite element method yields temperature fields. Then, the temperature
fields are utilized as inputs for the deep learning model, while thermo-mechanical
properties, encompassing thermal conductivity, specific heat capacity, Young’s modulus, and
Poisson’s ratio, serving as outputs to monitor the training process. Numerical results indicate
that the proposed deep learning model accurately predicts temperature-dependent material
properties of specific pressure vessels in real-time, with strong generalization performance.
The proposed algorithm shows low sensitivity to the number of training samples and data
noise. Moreover, compared with other methods, it achieves higher prediction accuracy and
performance. The study provides an innovative and effective approach for identifying the
material properties of specific pressure vessels.
Keywords: Inverse problem; Convolutional neural network; Bidirectional long short-term
memory network; Transformer; Thermo-mechanical properties

Abdessamad Elmotawakkil, Adnane Al Karkouri, Nourddine Enneya,


Explainable machine and deep learning framework for newborn health monitoring: A
simulation-based approach,
Intelligent Hospital,
2025,
100044,
ISSN 3050-8371,
[Link]
([Link]
Abstract: Machine learning and deep learning techniques are increasingly being adopted in
neonatal health research to improve early risk detection and support clinical decision-
making. Progress, however, is limited by the scarcity and ethical constraints of real-world
neonatal datasets, which restrict model development and validation. To address this, we
employed a synthetic but medically realistic dataset simulating daily health records of
newborns to evaluate the performance of four classification models: k-Nearest Neighbor
(kNN), Support Vector Machine (SVM), Multi-Layer Perceptron (MLP), and TabTransformer.
Models were assessed using accuracy, precision, recall, F1-score, Cohen’s Kappa, confusion
matrices, and SHAP-based explainability analyses. Results indicated that the TabTransformer
consistently outperformed the other models, achieving the highest test accuracy (96.5 %)
and demonstrating superior recall for the minority at-risk class, highlighting its ability to
capture complex neonatal health patterns. MLP and SVM delivered competitive results,
whereas kNN, while interpretable, showed reduced generalization under class imbalance.
These findings demonstrate the promise of transformer based architectures for neonatal
health monitoring and emphasize the need to validate such models using real-world clinical
datasets in future studies.
Keywords: Machine Learning and Deep Learning; Newborn Health; Monitoring; Neonatal
Risk Prediction

Jiahao Deng, Yiqing Qian, Feifei Cui, Yanshuang Liu, Jialong Lai,
Research on lunar regolith of the Chang'E-4 landing site: An automated analysis method
based on deep learning framework,
Icarus,
Volume 425,
2025,
116338,
ISSN 0019-1035,
[Link]
([Link]
Abstract: On January 3, 2019, the Chang'E-4 lander successfully landed within the Von
Kármán crater, located in the South P ole-Aitken Basin (SPA) on the farside of the Moon
(45.5°S, 177.6°E), marking the first soft landing on the lunar farside. The lander, equipped
with the Lunar Penetrating Radar (LPR) system, aimed to provide insights into the structure
and evolution of the Moon. Previous research often relied on manually identifying
hyperbolic features to analyze the lunar shallow subsurface properties. This inefficient
approach may lead to subjective biases, resulting in unstable outcomes. This research
constructed an automatic analysis framework by integrating the Swin Transformer with a 3D
velocity spectrum, which is then applied to analyze the properties of the Chang'E-4 LPR data.
The experimental results indicate that the framework achieved a precision of 98.9 % and a
recall of 96.7 % in hyperbolic feature identification, with an F1 of 0.9782 and AP of 94.8 %.
Additionally, it has been experimentally validated that the framework can accurately invert
hyperbolic features' two-way travel time and velocity. Finally, the framework is applied to
analyze the lunar shallow subsurface structure and properties within the landing area of the
Chang'E-4 mission.

Keywords: Chang'E-4; Lunar penetrating radar; Swin transformer; 3D velocity spectrum

Shanika Edirisinghe, Bianca Schoen-Phelan, Svetlana Hensman,


Generative deep learning models for cloud removal in satellite imagery: A comparative
review of GANs and diffusion methods,
ISPRS Open Journal of Photogrammetry and Remote Sensing,
Volume 19,
2026,
100110,
ISSN 2667-3932,
[Link]
([Link]
Abstract: Satellite imagery provides essential geospatial data to support various remote
sensing applications, including environmental monitoring, disaster management, urban
planning, and land utilization studies. However, cloud cover often obstructs the clarity and
reliability of satellite images, reducing their usefulness. With advances in deep learning,
generative models — particularly Generative Adversarial Networks (GANs) and denoising
diffusion models — have emerged as promising solutions for cloud removal in satellite
imagery. This review systematically evaluates GAN-based and diffusion-based methods,
comparing their strengths, limitations, and performance across diverse geographic and cloud
conditions. The analysis shows that GANs generate visually realistic outputs through
adversarial training, while diffusion models offer superior spatial and structural fidelity due
to iterative noise reduction. Integrating auxiliary data such as Synthetic Aperture Radar (SAR)
imagery further enhances cloud removal accuracy. This review highlights current challenges
and identifies research gaps to support future innovation in satellite image restoration,
particularly in cloud removal and generative deep learning for remote sensing.
Keywords: GANs; Diffusion models; Cloud removal; Geospatial data

Guanliang Liu, Wenchao Chen, Bo Chen, Bo Feng, Penghui Wang, Hongwei Liu,
Supervised contrastive deep Q-Network for imbalanced radar automatic target recognition,
Pattern Recognition,
Volume 161,
2025,
111264,
ISSN 0031-3203,
[Link]
([Link]
Abstract: In the presence of limited and extremely imbalanced data, deep learning methods
for radar automatic target recognition (RATR) often suffer from significant performance
degradation and overfitting. To tackle this issue, we propose Supervised Contrastive Deep Q-
network (SCDQ), a novel end-to-end reinforcement learning method, for multi-class
imbalanced RATR. SCDQ formulates the imbalanced recognition problem as a Markov
decision process (MDP) and optimizes the classifier through an enhanced Q-learning
paradigm. In order to augment the model’s feature extraction capabilities under the
constraint of limited samples, we tightly integrate reinforcement learning (RL) with
supervised contrastive learning, introducing an innovative feature enhancement module. To
further enhance the model’s adaptability to challenging samples, we integrate a
meticulously designed priority sampling into the proposed SCDQ framework, denoted as
SCDQ-P. Experimental results on both simulated and real datasets demonstrate the reliability
and effectiveness of the proposed method.
Keywords: Imbalanced RATR; Deep learning; Deep reinforcement learning (DRL); Supervised
contrastive learning; Priority sampling

Sridhara Murthy B, Gandla Madhu, Varaganti Manisha, Muttukuri Madhu Krishna, Padmam
Divyavalli, Reteneni Naveen,
ODC-net: Scalable and efficient object detection and classification in multi-object CCTV
environments using deep transformer YOLO,
Franklin Open,
Volume 13,
2025,
100404,
ISSN 2773-1863,
[Link]
([Link]
Abstract: The increasing demand for real-time, accurate, and scalable surveillance systems is
driven by the rapid rise in urban Closed-Circuit Television (CCTV) deployments, with global
video surveillance expected to generate over 3 billion video hours daily. However, existing
multi-object detection and classification approaches struggle with poor scalability, temporal
inconsistencies, and sub-optimal accuracy, particularly in complex CCTV environments with
dynamic object motion. To address these limitations, this work proposes Object Detection
Classification Network (ODCNet), a scalable and efficient framework for multi-object
detection, classification, and tracking in CCTV video streams. The model leverages the MS-
COCO-2017 dataset for robust training on diverse object categories, followed by a novel
Hierarchical Spatial Temporal Aggregation (HSTA) feature extraction technique that enhances
spatio-temporal consistency and contextual learning. The core detection and classification
module are powered by Deep Transformer You Only Look at Once V8 (DT- YOLOV8),
combining transformer attention mechanisms with YOLO's real-time detection capabilities.
During testing, input CCTV videos undergo preprocessing through video-to-frame
conversion, with each frame processed using HSTA and DT- YOLOV8 for precise object
detection and classification. Furthermore, to ensure reliable object tracking across video
frames, a Feature Adaptive Continual-Learning Tracker (FACLT) is integrated, enabling
consistent object association and high-quality output video generation with real-time
annotations. Extensive experiments demonstrate the superior performance of ODCNet,
achieving an impressive 99.477 % accuracy, 99.334 % precision, 99.409 % recall, and 99.727
% F1-score, establishing its effectiveness for real-world multi-object surveillance in dynamic
CCTV environments.
Keywords: Cctv videos; Object detection; Real-time video analytics; Data preprocessing;
Hierarchical video intercorrelated similarity; Deep transformer

Siddik Barbhuiya, Vivek Gupta,


From gauged to ungauged: Large-scale deep learning rainfall-runoff modelling for reliable
streamflow estimation in India's diverse basins,
Environmental Modelling & Software,
Volume 194,
2025,
106696,
ISSN 1364-8152,
[Link]
([Link]
Abstract: Runoff estimation in India faces challenges due to diverse climate zones, complex
physiographic conditions, and variable rainfall patterns, limiting traditional hydrological
models and prompting exploration of advanced deep learning methods for improved
streamflow prediction. Existing deep learning hydrological models struggle to estimate
discharge at ungauged sites. In this study, we tested eight different deep learning models,
four recurrent neural networks (GRU, CudaLSTM, EALSTM, ARLSTM) and four attention-
based architectures (Transformer, Informer, Reformer, Linformer), across 144 watersheds in
the Indian subcontinent (ISC). Our training and testing datasets combined meteorological
forcing, catchment attributes, and observed discharge records. According to the results,
ARLSTM improved prediction accuracy, achieving a median Nash–Sutcliffe Efficiency (NSE) of
0.71 on test basins. ARLSTM performs exceptionally well in specific regions: tropical
monsoon areas (median NSE = 0.849), semi-arid regions (median NSE = 0.586), monsoon-
influenced subtropical zones (NSE = 0.688), and tropical wet–dry climates (NSE = 0.539),
especially in arid zones where traditional hydrological models often struggle. The
assessments of high- and low-flow frequencies and durations, mean discharge, and runoff
ratios underscore ARLSTM's capability to capture both extreme and average flow conditions.
ARLSTM's reliance on lagged streamflow limits its use in ungauged basins. To address this
issue, we developed a novel deep learning architecture, Ungauged Basin LSTM (UBLSTM), to
predict the runoff values for any ungauged basin in India. UBLSTM matches the performance
of ARLSTM, making it a better choice for areas in India that lack sufficient data or have
ungauged basins across various climate zones.

Ziwei Zhang, Mengtao Zhu, Yunjie Li, Yan Li, Shafei Wang,
Joint recognition and parameter estimation of cognitive radar work modes with LSTM-
transformer,
Digital Signal Processing,
Volume 140,
2023,
104081,
ISSN 1051-2004,
[Link]
([Link]
Abstract: The recent developed cognitive radars can implement flexible work modes with
programmable modulation types and optimized modulating values for each mode definition
parameter. Automatic analysis of these work modes is a significant challenge for modern
electromagnetic reconnaissance receivers. In this paper, a Multi-Output Multi-Structure
(MOMS) learning-based processing framework is proposed for Joint inter-pulse automatic
Modulation Recognition and Parameter Estimation (JMRPE-MOMS). We propose a label
construction method as a feature interpretation method of the network to facilitate MOMS
learning and utilize the correlations between labels for performance gain. Moreover, an
LSTM-Transformer is designed to mine deep time-series characteristics, which can model
local and global relationships and reduce quantization loss. The proposed framework can
perform joint modulation recognition and parameter estimation (JMRPE) tasks
simultaneously with flexible output structures including scalar output and vector output
with fixed or variable sizes. Extensive simulations are performed based on the simulated
radar work modes defined with pulse repetition interval (PRI) sequences. The simulation
results validate the effectiveness and superiority of the proposed method especially under
non-ideal electromagnetic environments.
Keywords: Radar work mode; Automatic modulation recognition; Modulation parameter
estimation; Multi-output learning; Transformer

Ziqi Zhou, Baichun Wang, Zirui Huang, Xiaohui Wu, Weidong Yang, Gang Guo, Shuichangtian
Qiu, Jiakuan Yang, Aijiao Zhou,
Floc image-driven deep learning enhanced by temporal windows and transformers for
carbon emission reduction in drinking water treatment plants,
Water Research,
Volume 289, Part A,
2026,
124868,
ISSN 0043-1354,
[Link]
([Link]
Abstract: Using machine learning (ML) and deep learning (DL) algorithms for precise
coagulant dosing in drinking water treatment plants (DWTPs) helps ensure drinking water
safety and supports greenhouse gas (GHG) emission reduction. The effectiveness of these
algorithms depends heavily on the availability of long-term data. Short-term data are used in
this study to explore the potential of four traditional ML algorithms and four DL algorithms
for precise coagulant dosing. Three strategies were introduced: an innovative method for
floc morphological feature extraction, selection of temporal windows, and integration of
transformer architecture. Based on these strategies, 16 different scenarios were
constructed, resulting in 96 models for analysis. Results show that without any strategy
applied, ML models achieved 5.0% higher R and 10.5% higher R² than DL models. This is due
to their simplicity, faster convergence, and suitability for low-dimensional data. However,
with the proposed strategies, DL models significantly improved and outperformed ML
models. Given the time-lagged dependencies across DWTP treatment units, optimized DL
models N better captured complex nonlinear temporal relationships. The best-performing
model was the temporal convolutional network (TCN) with floc morphological features, 4-h
temporal window, and transformer architecture, achieving R and R2 values of 0.99. The
model was trained with only one month of data and rapidly deployed. A weekly self-
updating mechanism was integrated to ensure long-term adaptability. The model has been
operating stably in a DWTP for over six months. It has reduced coagulant dosage by 20% and
carbon dioxide equivalent (CO2-eq) emissions by an estimated 70 tons annually. This study
demonstrates the strong potential of optimized DL algorithms to improve water purification
and reduce carbon emissions.
Keywords: Floc morphological feature; Temporal window; Transformer; Deep learning;
Carbon emission reduction

Wentao He, Jianfeng Ren, Ruibin Bai, Xudong Jiang,


Radar gait recognition using Dual-branch Swin Transformer with Asymmetric Attention
Fusion,
Pattern Recognition,
Volume 159,
2025,
111101,
ISSN 0031-3203,
[Link]
([Link]
Abstract: Video-based gait recognition suffers from potential privacy issues and performance
degradation due to dim environments, partial occlusions, or camera view changes. Radar has
recently become increasingly popular and overcome various challenges presented by vision
sensors. To capture tiny differences in radar gait signatures of different people, a dual-branch
Swin Transformer is proposed, where one branch captures the time variations of the radar
micro-Doppler signature and the other captures the repetitive frequency patterns in the
spectrogram. Unlike natural images where objects can be translated, rotated, or scaled, the
spatial coordinates of spectrograms and CVDs have unique physical meanings, and there is
no affine transformation for radar targets in these synthetic images. The patch splitting
mechanism in Vision Transformer makes it ideal to extract discriminant information from
patches, and learn the attentive information across patches, as each patch carries some
unique physical properties of radar targets. Swin Transformer consists of a set of cascaded
Swin blocks to extract semantic features from shallow to deep representations, further
improving the classification performance. Lastly, to highlight the branch with larger
discriminant power, an Asymmetric Attention Fusion is proposed to optimally fuse the
discriminant features from the two branches. To enrich the research on radar gait
recognition, a large-scale NTU-RGR dataset is constructed, containing 45,768 radar frames of
98 subjects. The proposed method is evaluated on the NTU-RGR dataset and the MMRGait-
1.0 database. It consistently and significantly outperforms all the compared methods on
both datasets. The codes are available at: [Link]
Keywords: Micro-Doppler signature; Radar gait recognition; Spectrogram; Cadence velocity
diagram; Asymmetric Attention Fusion

Lei Xia, Shurui Zhang, Yuhang Hu, Renli Zhang, Song Li, Weixing Sheng,
A deep learning-based maneuvering target tracking with temporal convolutional networks,
Signal Processing,
Volume 239,
2026,
110322,
ISSN 0165-1684,
[Link]
([Link]
Abstract: Traditional algorithms of the target tracking rely on predefined target motion
states to modify sensor observations. However, these algorithms struggle to accurately and
promptly model the maneuvering state of a target, thereby failing to provide precise state
estimation when the target exhibits maneuvering behavior. To address this challenge, we
propose a maneuvering target tracking algorithm based on temporal convolutional networks
(TcnMTT). The TcnMTT model employs a constant velocity model-based unscented Kalman
filter to decompose the input trajectory into high maneuver state and low maneuver state.
Furthermore, the model directly maps the input observations to the true trajectory through
a set of symmetric TCN networks. Additionally, TcnMTT incorporates an instance
normalization module to project features into a specific feature space and combines a
channel attention mechanism to extract feature correlations. Simulation results demonstrate
that the proposed TcnMTT model outperforms existing methods in tracking maneuvering
targets.
Keywords: Maneuvering target tracking; Temporal convolutional network; Radar tracking;
Deep learning algorithm

Junyu Zhou, Yuting Fu, Sihan Dong, Yuemeng Liu, Han Sun, Yanmin Li, Xunbin Wei,
Multi-model deep learning on photoacoustic flow cytometry signals for real-time melanoma
circulating tumor cells detection and biological characterization,
Expert Systems with Applications,
Volume 308,
2026,
131123,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Background
Melanoma remains one of the most aggressive forms of skin cancer, with early detection
being critical for patient outcomes. This study introduces a novel photoacoustic
fingerprinting approach integrated with advanced machine learning for non-invasive
melanoma detection and circulating tumor cell (CTC) identification.
Methods
We developed a three-tiered photoacoustic fingerprinting system combining photoacoustic
flow cytometry (PAFC) with machine learning algorithms. A uniform PAFC configuration
employed a 532 nm laser for vascular localization followed by a 1064 nm laser targeting
melanin-rich melanoma cells. The study included 50 melanoma patients and healthy
controls, analyzing spectral features across multiple wavelengths. We compared self-
supervised learning architectures (PAFCMamba vs. Transformer) and developed a hybrid
CNN-Transformer model for simultaneous CTC identification, staging, and metastatic
dissemination prediction.
Results
The photoacoustic fingerprinting system achieved exceptional diagnostic discrimination
between melanoma patients and healthy controls. Random Forest achieved area under
curve (AUC) values up to 0.97. The PAFCMamba model outperformed the Transformer
architecture (accuracy 0.75 vs. 0.62, AUC 0.785 vs. 0.730). The hybrid CNN-Transformer
architecture achieved exceptional performance with AUCs up to 0.974 and precision > 94 %
in simultaneous CTC detection, staging, and metastasis prediction. High-immunogenicity
genes including MLANA, GPR89B/A, and PIGF were identified as potential immunotherapy
targets, with photoacoustic signatures serving as non-invasive surrogate biomarkers for
underlying molecular characteristics.
Conclusions
This study establishes photoacoustic fingerprinting as a clinically viable, non-invasive
approach for melanoma detection and CTC monitoring, achieving performance comparable
to conventional methods. The integration of machine learning with photoacoustic
biomarkers provides a scalable framework with interpretable features that facilitates clinical
translation.
Keywords: Photoacoustic fingerprinting; Circulating tumor cells; Deep learning; Non-invasive
diagnosis; Biomarkers; Precision medicine

Xuhui Hu, Huimin Li, Chen Si, Yue Tang,


A hybrid deep learning framework for day-ahead PV forecasting using coarse-grained
weather data,
Renewable Energy,
Volume 261,
2026,
125306,
ISSN 0960-1481,
[Link]
([Link]
Abstract: Accurate day-ahead photovoltaic forecasting is critical for grid reliability and
market operations, especially when high-resolution weather data are unavailable. Existing
models often suffer reduced accuracy or data leakage risks when handling coarse-grained
meteorological inputs. This work proposes a novel deep learning framework that enhances
photovoltaic forecasting performance using low-resolution weather data while explicitly
avoiding information leakage during decomposition. The proposed method combines three
key components: (1) a Similar-Day Reconstruction strategy that identifies and ranks past
weather-photovoltaic profiles using multi-segment weighted similarity matching; (2) a dual-
branch prediction module that separates trend and volatility modelling, where Multi-
variable Variational Mode Decomposition is used for trend extraction in a leakage-free
manner; and (3) a Multi-Stream Learning Convolutional Network that fuses Long Short-Term
Memory networks, Convolutional Neural Networks, residual, and attention pathways to
capture multi-scale temporal features. Final forecasts are generated through weighted
integration of both branches. Extensive experiments on three photovoltaic datasets
demonstrate that the proposed approach outperforms benchmark models (e.g., Long Short-
Term Memory networks, Transformer) across multiple weather conditions. It maintains
robust prediction accuracy despite low-granularity weather inputs. The framework
demonstrates strong generalization, methodological rigor in avoiding information leakage,
and practical value for real-world photovoltaic forecasting applications where high-
resolution weather data may be unavailable.
Keywords: Similar days dataset construction; Multi-stream parallel architecture; Multi-
variable variational mode decomposition algorithm; Day-ahead; Photovoltaic output
prediction

Ruixing Wang, Wanying Gao, Jianfa Wu, Chunling Wei, Renjian Hao, Huida Yan,
Transformer-enhanced reinforcement learning for spacecraft evasion of asymmetric swarm
threats under complex multi-constraints,
Aerospace Science and Technology,
Volume 168, Part G,
2026,
111200,
ISSN 1270-9638,
[Link]
([Link]
Abstract: Aiming at the challenge of asymmetric threat targets characterized by ”large
numbers, high maneuverability, large fuel capacity and excellent measurement capability”
continuously approaching, a spacecraft threat evasion method based on a Transformer-
Decoder deep reinforcement learning (DRL) architecture is proposed under complex multi-
constraints such as fuel, maneuverability, time, lighting conditions, measurement capability,
and mission continuity. By introducing a multi-head attention mechanism, the method
dynamically allocates attention weights to different threat targets, enabling efficient evasion
of variable threats. Extensive simulation results demonstrate that the proposed approach
achieves superior learning efficiency, convergence speed, and generalization capability
compared with conventional DRL methods. Moreover, it maintains high success rates and
stability across swarms of different sizes and strategies, while outperforming conventional
methods in terms of maneuverability, fuel consumption, and mission continuity. These
results highlight the effectiveness of the proposed approach in enhancing the autonomous
evasion capability of spacecraft, providing a novel solution for safety control in complex
asymmetric threat environments.
Keywords: Spacecraft guidance; Intelligent decision-making; Asymmetric threat evasion;
Deep reinforcement learning; Attention mechanism

Tianwei Mou, Yang Liu, Lintao Tan, Lianhui Wu, Yaya Zhang, Chunming Li, Huan Zhou, Irene
D. Alabia,
Mapping subtle-featured oyster rafts with high-resolution imagery and deep learning
techniques,
ISPRS Journal of Photogrammetry and Remote Sensing,
Volume 231,
2026,
Pages 216-229,
ISSN 0924-2716,
[Link]
([Link]
Abstract: Raft farming is the primary method of oyster aquaculture, with both production
output and environmental impact directly linked to the density of oyster rafts. However, the
absence of accurate monitoring technologies to regulate raft density presents a significant
challenge for the sustainable development of the raft-based aquaculture industry. In this
study, we aimed to integrate deep learning techniques with high-resolution satellite imagery
to improve and support the effective management of offshore aquaculture. We applied a
Transformer-based semantic segmentation model (Mask2Former) alongside several
benchmark models (UNet, PSPNet, Segformer, and KNet) to the Jilin-1 high-resolution
satellite imagery obtained in 2024. This analysis focused on offshore oyster aquaculture in
Rushan City, China. Our results showed that Mask2Former outperformed the other models
in terms of generalizability and validation accuracy. It achieved a mean Intersection over
Union (mIoU) of 88.15 %, a mean accuracy (mAccuracy) of 92.22 %, and a mean F-score (mF-
score) of 93.39 % on the test dataset. Using the extraction maps generated by Mask2Former,
we identified over 40,000 oyster rafts in the offshore region of Rushan, covering an
aquaculture area of 20.43 km2. Compared to traditional optical imagery and extraction
methods, our approach was able to identify individual offshore aquaculture structures,
specifically oyster culture rafts, that are often overlooked. This allows for a more detailed
spatial and structural analysis. This case study underscores the significant potential of
integrating deep learning models with high-resolution satellite imagery to enhance the
sustainable management of offshore aquaculture.
Keywords: Offshore aquaculture; Deep-learning-based model; Semantic segmentation;
Aquaculture management

Yangfan Zhao, Deyi Chen, Yuxuan Wang, Baojie Nie, Yungang Zhao, Qi Li, Shilian Wang,
Dezhong Wang,
Transformer-based deep learning architecture for multivariable radioactive source term
inversion,
Journal of Environmental Radioactivity,
Volume 291,
2026,
107835,
ISSN 0265-931X,
[Link]
([Link]
Abstract: Inversion for the radioactive source term has received growing attention in the
post-Fukushima era. Under some special scenarios, the source term, including release rate,
height and position, is necessary for nuclear emergency response and consequence
assessment. Here, a transformer-based deep learning architecture was developed for
multivariable source term estimation. The CALMET-LAPMOD coupling model validated by
the Kincaid tracer experiment was employed to produce the datasets. The datasets were
systematically constructed for five representative scenarios with the following time-varying
parameters: release rate, release height, release location, coupling of release rate and
height, and coupling of all three variables. Subsequently, a Transformer model with Bayesian
optimization for adaptive hyperparameter tuning was developed. The results demonstrated
excellent performance in source term inversion, with a determination coefficient (R2) of
above 0.96 for release rate and height, and an average distance error of 1.19 km at a 95 %
confidence level for location prediction. Regarding the coupling of all three variables
scenario, the R2 for release rate and location remained above 0.92, whereas the height
achieved R2 of 0.72. Additionally, feature ablation analysis revealed that monitoring points
with high concentration values contribute significantly to inversion, providing quantitative
insights to optimize the monitoring network layout.
Keywords: Nuclear emergency; Atmospheric dispersion; Source term inversion; Transformer
architecture

Gang Xiong, Tao Zhen, Wenyu Huang, Bingxu Min, Wenxian Yu,
Fractal-domain deep learning with Transformer architecture for SAR ship classification,
ISPRS Journal of Photogrammetry and Remote Sensing,
Volume 230,
2025,
Pages 208-226,
ISSN 0924-2716,
[Link]
([Link]
Abstract: This paper extends deep learning from spatiotemporal-frequency to the fractal
domain, and to the best of our knowledge, introduces for the first time the concept of fractal
domain deep learning. Firstly, a Fractal Domain Transformer (FracFormer) model
architecture is proposed to address the challenging problem of SAR image target
classification in complex scenarios. Based on the Singularity Exponent-Domain Image
Feature Transform (SIFT), FracFormer transforms original images into the fractal-domain
feature images, utilizes fractal feature filters and combiners for iterative learning, and
ultimately achieves image classification through fractal feature mixers and classifiers.
Particularly, we derived the fractal feature filtering theorem based on SIFT and the feature
combination theorem based on SIFT, providing theoretical support for the design of the core
modules of FracFormer. On the OpenSARShip2.0 dataset, our model outperforms baseline
models, with improvements ranging from 0.37 % to 11.83 % on average. Besides, extensive
visualization analysis of the model’s fractal domain feature learning results indicates that
FracFormer accords with the two theorems, representing good interpretability. Furthermore,
FracFormer demonstrates fast convergence and strong generalization in low signal-to-noise
ratio scenarios. Specifically, at 0 dB sea clutter, it achieves a 9.96 % improvement in
classification performance over frequency domain GFNet and accelerates convergence by
approximately 36 %. The findings of this study are expected to provide new learning
paradigms and model architectures for the fields of deep learning and computer vision.
Keywords: Fractal domain deep learning; Fractal signal processing; SAR image recognition;
Fractal transformer; Learnalble fractal filtering

Wentao Zhao, Guoxiang Tong,


A deep learning based heart rate estimation method for millimeter wave radar,
Measurement,
Volume 255,
2025,
117923,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Contact-based vital sign detection technology has been widely used in the medical
field. However, contact-based devices may cause discomfort to users and suffer from user
dependency issues. Frequency Modulated Continuous Wave (FMCW) millimeter-wave radar
provides an efficient and accurate solution for heart rate and respiratory rate monitoring.
Nevertheless, due to the small amplitude of heartbeat micro-motion signals, they are
susceptible to noise interference such as respiratory harmonics, making accurate
measurement challenging. Aiming at the problem that traditional methods are difficult to
adapt to different environmental noises, we propose a heart rate estimation method based
on a convolutional neural network. The heart rate estimation accuracy is significantly
improved by recognizing the phase change patterns in radar signals. We reduce the
interference of respiratory harmonics on the heart rate micromotion signal by decomposing
the extracted phase signal into multiple frequency components using the Empirical Wavelet
Transform (EWT) algorithm. The proposed deep learning model is used by means of depth
convolution in order to balance the model size and accuracy. Additionally, we introduce
time–frequency channels into the convolutional neural network to further enhance its
feature extraction capability. Comparisons with various related works on heart rate
distribution, sensing range, and subjects demonstrate the lightweight nature and higher
accuracy of the proposed model. Extensive experiments are conducted on a dataset based
on the TI AWR1642BOOST radar. The results show that the proposed method achieves
outstanding performance, with an accuracy of ±5 BPM in heart rate monitoring.
Keywords: Deep learning; Heart rate (HR); Millimeter-wave (mmW) radar; Vital signs
monitoring; Moving target indication (MTI)

Mingqi Li, Pengxin Wang, Kevin Tansey, Fengwei Guo, Ji Zhou,


Improved leaf area index reconstruction in heavily cloudy areas: A novel deep learning
approach for SAR-Optical fusion integrating spatiotemporal features,
International Journal of Applied Earth Observation and Geoinformation,
Volume 142,
2025,
104745,
ISSN 1569-8432,
[Link]
([Link]
Abstract: The Leaf Area Index (LAI) is an essential parameter for assessing vegetation growth.
LAI derived from optical data can suffer from gaps caused by cloud cover. Synthetic Aperture
Radar (SAR) presents a solution with its all-weather observation capability. To address these
issues, this study proposes a new deep learning approach for reconstructing time series LAI
using SAR and optical data in two steps. Firstly, the two-dimensional Convolutional Neural
Network-Transformer (2D CNN-Transformer) is applied to bridge SAR and optical data.
Secondly, the 2D CNN-Transformer predicted LAI and the Sentinel-2 LAI are input into the
Enhanced Deep Convolutional Model for Spatiotemporal Image Fusion (EDCSTFN) model to
further improve the accuracy. The novelty lies in a two-step framework combining a 2D CNN-
Transformer for spatiotemporal feature extraction and a deep learning fusion algorithm
refining accurate LAI reconstruction. Results showed that the 2D CNN-Transformer achieved
a higher accuracy (R2 = 0.64, RMSE = 0.38 m2/m2) in establishing a relationship between
SAR and optical data, compared to 1D CNN, 2D CNN-LSTM, and 1D CNN-Transformer. In the
second step, the EDCSTFN reconstructed LAI achieved the highest accuracy of an R2 of 0.81
and an RMSE of 0.22 m2/m2, with an average R2 of 0.61 and RMSE of 0.37 m2/m2 across
croplands and forests in millions of pixels, further improving the accuracy based on the first
step. The approach effectively fills gaps in spatial details and achieves a more continuous
spatial distribution. The proposed approach demonstrates good generalizability in millions of
pixels under frequent cloud cover and complex surface conditions and provides a new
strategy for the fusion of optical and SAR data.
Keywords: SAR data; Optical data; Spatiotemporal features; LAI; Deep learning;
Spatiotemporal fusion

Yunfeng Fang, Zheng Tong, Tianqing Hei, Siqi Wang, Tao Ma,
Deep learning applications in ground-penetrating radar inversion: A review,
Measurement,
Volume 258, Part D,
2026,
119399,
ISSN 0263-2241,
[Link]
([Link]
Abstract: The complex nonlinear relationship between the subsurface medium and ground-
penetrating radar signals results in the pervasive ill-posedness and non-uniqueness of
conventional inversion methods. Deep learning, with its powerful feature extraction
capabilities and advantages in modeling complex nonlinear relationships, has unique
strengths in handling complex signals and nonlinear problems, making it especially suitable
for GPR inversion tasks. This paper reviews the latest applications of deep learning in GPR
inversion, summarizing the application strategies of deep learning from two perspectives:
data-driven and data-physics hybrid-driven. Commonly used model architectures and their
performance in signal feature extraction, multi-scale information fusion, and data
preprocessing are discussed, along with the application of various loss functions in inversion
tasks. Finally, current challenges, such as limited model generalization, model dependence
on the dataset and computational efficiency constraints, are discussed, and potential future
research directions are proposed to further advance deep learning in GPR inversion.
Keywords: Ground-penetrating radar; Deep learning; Inversion

Huayan Chen, Yongbin Liu, Yi Liang, Yu Chen, Wenbin Lin, Chaojiang Fu, Caisong Luo,
Multimodal deep learning framework for predicting the evolution of impact damage,
Advanced Engineering Informatics,
Volume 69, Part D,
2026,
104029,
ISSN 1474-0346,
[Link]
([Link]
Abstract: Accurate prediction of damage evolution in structures under impact loading is
crucial for engineering structural safety assessment. Traditional methods primarily rely on
post-hoc qualitative grading systems or simplified models. These approaches struggle to
effectively handle the complex nonlinear characteristics and multi-source heterogeneous
data inherent in impact processes. This study proposes a deep learning framework based on
multimodal fusion for predicting impact damage evolution. The framework first employs the
YOLOv8 instance segmentation model to perform precise mask segmentation on impact
process video frames, extracting time-varying geometric features of structural deformation,
which are subsequently fused with video features extracted via the ResNet18-CBAM-TSM
(RCT) architecture. Subsequently, a Bidirectional Multimodal Autoregressive Transformer
(BiMAR-Transformer) is constructed to achieve deep fusion of video features with physical
features. In addition, energy absorption rate (EAR) and support rotation angle (SRA) are used
to establish a reliable quantitative benchmark for damage evolution prediction. Impact
experimental validation conducted on castellated beams demonstrates that the proposed
framework reduces the Mean Absolute Error (MAE) in EAR prediction by 36.8% compared to
traditional GARCH models, with a Peak Prediction Error (PPE) of merely 1.81% for SRA.
Ablation experiments demonstrated that removing the instance segmentation module
resulted in a 17% reduction in the coefficient of determination (R2) compared to the
complete model, while removing the multimodal fusion module led to increases in PPE of
107.32% and 141.44% for EAR and SRA, respectively. The proposed framework establishes a
novel paradigm for predicting structural impact damage evolution, which is applicable to
various complex structures.
Keywords: Impact damage evolution; Multimodal fusion; Castellated beam; Deep learning;
Instance segmentation; Transformer

Yongjun Zhang, Xiaoshuan Zhang, Baotian Li, Wang Han, Shuran Feng,
VfiA: A vitality fusion identification approach for live fish based on wearable electrical
impedance sensing and deep learning technology,
Sensors and Actuators A: Physical,
Volume 394,
2025,
116947,
ISSN 0924-4247,
[Link]
([Link]
Abstract: This study aims to develop a novel vitality fusion identification approach (VfiA) that
integrates wearable electrical impedance sensing (WEIS) and deep learning technology for
non-invasive live fish monitoring. It builds upon the VMD-SSA-BiLSTM model, which
demonstrates superior identification performance compared to unimodal impedance
monitoring by integrating multi-impedance features and critical survival indicators (e.g.,
blood stress indexes, BSIs; respiratory rate, RR). Specifically, this model utilizes the Sparrow
Search Algorithm (SSA) to optimize a VMD-denoising-enhanced Bidirectional Long Short-
Term Memory (BiLSTM) classifier, thereby achieving higher classification accuracy.
Experimental results on groupers demonstrate that the proposed model attains superior
performance compared to SSA-Transformer, GWO-BiLSTM, PSO-BiLSTM, Transformer, and
BiLSTM within the 11–13℃ range, with an accuracy of 87.5 %, precision of 88.8 %, recall of
88.9 %, and an F1-score of 87.9 %. In conclusion, these findings validate VfiA as a robust
vitality assessment solution, highlighting its potential for optimizing temperature-zone
management in the smart live fish logistics industry.
Keywords: Impedance sensing; Deep learning; Vitality classification; Swarm intelligence
algorithm; Temperature optimization

Shuaiying Zhang, Zhen Dong, Huadong Lin, Zhendong Zhang, Jinran Wu, Sinong Quan,
Wentao An, Tong Li, Rajiv Pandey,
Integrating linear and circular polarization features for PolSAR land cover classification with
deep learning,
International Journal of Applied Earth Observation and Geoinformation,
Volume 146,
2026,
105090,
ISSN 1569-8432,
[Link]
([Link]
Abstract: Existing deep learning methods for ecological monitoring using polarimetric
synthetic aperture radar (PolSAR) imagery primarily rely on coherency (T) or covariance (C)
matrices derived from the linear polarization basis, often overlooking scattering information
inherent in alternative polarization representations. To address this limitation, this study
proposes a novel classification framework that explicitly incorporates a circular polarization
basis into the PolSAR deep learning workflow. A circular coherency matrix (Cir), analogous to
the conventional T matrix, was first derived through polarization basis transformation.
Subsequently, a multi-basis input scheme was introduced to fuse linear and circular
polarization features to enhance feature representation and information utilization. The
proposed framework was validated on two benchmark datasets using multiple deep learning
models, achieving state-of-the-art classification accuracies of 97.70% and 98.58%.Compared
with standard linear-basis approaches, the proposed scheme yielded accuracy
improvements of 2.86% over the T-matrix-based method and 2.26% over the C-matrix-based
method. In addition, the incorporation of circular polarization features significantly
enhanced physical interpretability, particularly for structurally complex targets such as
forests and buildings. Overall, the findings provide an effective technical pathway for
intelligent land cover classification and broader ecological monitoring. The source code and
datasets are available at [Link]
Implementation.
Keywords: PolSAR image classification; Deep learning; Circular polarization basis; Ecological
monitoring; Sustainable land use; Classification performance analysis

Raihan Ahamed Rifat, Fuyad Hasan Bhoyan, Md Humaion Kabir Mehedi, Md Kaviul Hossain,
Md. Jakir Hossen, M.F. Mridha,
ConMatFormer: A multi-attention and transformer integrated ConvNext based deep learning
model for enhanced diabetic foot ulcer classification,
Results in Engineering,
Volume 28,
2025,
108248,
ISSN 2590-1230,
[Link]
([Link]
Abstract: Diabetic foot ulcer (DFU) detection is a clinically significant yet challenging task due
to the scarcity and variability of publicly available datasets. Limited annotated samples
restrict the ability of conventional deep learning models to achieve robust generalization in
real-world clinical scenarios. To solve these problems, we propose ConMatFormer, a new
hybrid deep learning architecture that combines ConvNeXt blocks, multiple attention
mechanisms convolutional block attention module (CBAM) and dual attention network
(DANet), and transformer modules in a way that works together. This design facilitates the
extraction of better local features and understanding of the global context, which allows us
to model small skin patterns across different types of DFU very accurately. To address the
class imbalance, we used data augmentation methods. A ConvNeXt block was used to obtain
detailed local features in the initial stages. Subsequently, we compiled the model by adding a
transformer module to enhance long-range dependency. This enabled us to pinpoint the
DFU classes that were underrepresented or constituted minorities. Tests on the DS1
(DFUC2021) and DS2 (diabetic foot ulcer (DFU)) datasets showed that ConMatFormer
outperformed state-of-the-art (SOTA) convolutional neural network (CNN) and Vision
Transformer (ViT) models in terms of accuracy, reliability, and flexibility. The proposed
method achieved an accuracy of 0.8961 and a precision of 0.9160 in a single experiment,
which is a significant improvement over the current standards for classifying DFUs. In
addition, by 4-fold cross-validation, the proposed model achieved an accuracy of 0.9755
with a standard deviation of only 0.0031. We further applied explainable artificial
intelligence (XAI) methods, such as Grad-CAM, Grad-CAM++, and LIME, to consistently
monitor the transparency and trustworthiness of the decision-making process. These
human-readable tools enhance the comprehension of the explanations and can substantially
increase the practical use of our methodology. Our findings set a new benchmark for DFU
classification and provide a hybrid attention transformer framework for medical image
analysis.
Keywords: Diabetic foot ulcers classification; Multi-attention; Transformer; GradCam;
Explainable AI; LIME

Daniel Bin Lu, Semiha Ergan,


Behavioral modelling of roadway construction workers: Improving deep learning-based
trajectory prediction with contextual information in traffic work zones,
Advanced Engineering Informatics,
Volume 71, Part A,
2026,
104277,
ISSN 1474-0346,
[Link]
([Link]
Abstract: Construction workers face rising risks of fatal injuries from vehicle crashes in
roadway work zones. While transportation safety research has focused on motorists’
behavior, the behavior of roadway workers remains underexplored. Existing trajectory
prediction models, developed for pedestrians or generic construction workers, typically do
not account for the unique roadway work zone activities and traffic interactions faced by
roadway workers. This study leverages a virtual reality (VR) and traffic simulation-based
platform to capture detailed context data, such as roadwork activities and nearby vehicles in
the worker’s field of view. The study’s main objective is to evaluate whether including this
context improves trajectory prediction accuracy of deep learning-based models, particularly
gated recurrent units (GRU) and transformer architectures. Results indicate that
transformers can improve their trajectory prediction accuracy (i.e., lower miss-rate) when
accounting for both the worker’s behavioral and traffic context data compared to a
transformer trained on trajectory position data alone. These improvements in accuracy are
observed across different roadwork construction tasks (e.g., installing sensor cable,
distributing grout) and different proximities to traffic vehicles. These findings contribute to
the development of more precise roadway worker trajectory models for use in autonomous
vehicles and safety systems.
Keywords: Trajectory prediction; Roadway work zones; Worker safety; Virtual reality; Deep
learning

Saidul Islam, Hanae Elmekki, Ahmed Elsebai, Jamal Bentahar, Nagat Drawel, Gaith Rjoub,
Witold Pedrycz,
A comprehensive survey on applications of transformers for deep learning tasks,
Expert Systems with Applications,
Volume 241,
2024,
122666,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Transformers are Deep Neural Networks (DNN) that utilize a self-attention
mechanism to capture contextual relationships within sequential data. Unlike traditional
neural networks and variants of Recurrent Neural Networks (RNNs), such as Long Short-Term
Memory (LSTM), Transformer models excel at managing long dependencies among input
sequence elements and facilitate parallel processing. Consequently, Transformer-based
models have garnered significant attention from researchers in the field of artificial
intelligence. This is due to their tremendous potential and impressive accomplishments,
which extend beyond Natural Language Processing (NLP) tasks to encompass various
domains, including Computer Vision (CV), audio and speech processing, healthcare, and the
Internet of Things (IoT). Although several survey papers have been published, spotlighting
the Transformer’s contributions in specific fields, architectural disparities, or performance
assessments, there remains a notable absence of a comprehensive survey paper that
encompasses its major applications across diverse domains. Therefore, this paper addresses
this gap by conducting an extensive survey of proposed Transformer models spanning from
2017 to 2022. Our survey encompasses the identification of the top five application domains
for Transformer-based models, namely: NLP, CV, multi-modality, audio and speech
processing, and signal processing. We analyze the influence of highly impactful Transformer-
based models within these domains and subsequently categorize them according to their
respective tasks, employing a novel taxonomy. Our primary objective is to illuminate the
existing potential and future prospects of Transformers for researchers who are passionate
about this area, thereby contributing to a more comprehensive understanding of this
groundbreaking technology.
Keywords: Transformer; Self-attention; Deep learning; Natural language processing (NLP);
Computer vision (CV); Multi-modality

Yakun Yang, Yu Chen, Yukang Xu, Yudong Wang, Guangwu Sun,


Inverse prediction of micro-nano fiber membrane preparation process based on deep
learning,
Polymer,
Volume 342,
2026,
129372,
ISSN 0032-3861,
[Link]
([Link]
Abstract: The performance of micro-nano-scale fibrous membrane materials is highly
dependent on their microstructure, which is governed by fabrication process parameters.
Traditional trial-and-error optimization methods are inefficient and costly because they
require repeated experiments. Consequently, this study proposes a deep learning-based
inverse design strategy to predict the process parameters from target fiber microstructures.
Scanning electron microscopy images of fibrous networks are collected and used as inputs to
a model incorporating an attention penalty mechanism to construct an intelligent
“microstructure–process parameter” mapping. The RT-Net model integrates an improved
ResNet-50 with a Transformer architecture and leverages transfer learning and microscopic
image enhancement to overcome the limitations under small-sample conditions. Seven
types of fibrous membranes prepared under different microfluidic spinning conditions are
used for training. The model achieves a validation accuracy of 98.81 % and an F1 score of
0.9857 for process parameter classification, outperforming the baseline models by 2.46 %
and 2.17 %, respectively. Further microfluidic spinning experiments based on the predicted
parameters, followed by statistical analysis, confirm the reliability of the model for inverse
prediction and structural reproduction. This method provides a novel and efficient approach
for the intelligent fabrication of micro-nano fibrous membranes and holds potential for
extension to the development of other functional textile materials.
Keywords: Micro-nano fiber materials; Inverse design; Deep learning; Transfer learning;
Attention mechanism; Process parameter classification; Functional textile materials

El-Sayed M. El-kenawy, Amel Ali Alhussan, Ebrahim A. Mattar, Marwa Radwan,


Feature selection and hyperparameter tuning in transformer-based deep learning models for
photovoltaic power forecasting using the Swordfish Movement Optimization Algorithm
(SMOA),
International Journal of Electrical Power & Energy Systems,
Volume 174,
2026,
111509,
ISSN 0142-0615,
[Link]
([Link]
Abstract: Accurate photovoltaic (PV) power forecasting is essential for ensuring grid stability
and efficient energy management in modern power systems. However, the nonlinear and
intermittent nature of solar radiation limits the performance of traditional models. This
study proposes a hybrid forecasting framework combining PatchTST with the Swordfish
Movement Optimization Algorithm (SMOA) and its binary variant (bSMOA) to enhance
prediction accuracy and model interpretability. The PatchTST captures complex temporal
dependencies through self-attention and patch embedding, while bSMOA performs optimal
feature selection and SMOA fine-tunes hyperparameters to improve convergence and
generalization. Experimental results demonstrate that the proposed model achieved a Mean
Squared Error (5.64×10−7), Root Mean Squared Error (9.46×10−7), and Mean Absolute Error
(7.32×10−6), with a correlation coefficient (r=0.9598) and a coefficient of determination
(R2=0.9660). The proposed framework achieved a Willmott Index (WI=0.9632) and a Nash–
Sutcliffe Efficiency (NSE=0.9755). These results confirm the superior accuracy, robustness,
and efficiency of the PatchTST–SMOA framework. The model provides a scalable and
interpretable solution for short-term PV power forecasting, supporting predictive
maintenance, energy dispatching, and the sustainable integration of renewable energy into
intelligent grid systems.
Keywords: Photovoltaic power forecasting; Binary swordfish movement optimization
algorithm (bSMOA); Feature selection and hyperparameter optimization; Metaheuristic
deep learning integration

Farasath Hasan, Xintao Liu,


Advancing urban expansion modeling with a hybrid TRANSGAN deep learning approach,
Environmental Modelling & Software,
Volume 194,
2025,
106693,
ISSN 1364-8152,
[Link]
([Link]
Abstract: Urban expansion modeling is pivotal for sustainable urban planning, yet
conventional approaches often fail to capture intricate spatial and temporal dynamics. In this
study, we present TRANSGAN, the first framework combining Transformer networks and
Generative Adversarial Networks (GANs) for urban expansion simulation. By harnessing the
spatial learning strengths of Transformers alongside the generative capabilities of GANs,
TRANSGAN significantly outperforms traditional models, as evidenced by enhanced
predictive accuracy and spatial consistency. Trained on historical land use data in Hong Kong
and incorporating key drivers, such as proximity to CBDs, road networks, and elevation, the
model delivers highly realistic urban expansion forecasts for 2035 and 2045. Comparative
analyses with Transformer, GAN, U-Net, and Random Forest models demonstrate that
TRANSGAN achieves the highest F1 Score (0.9496), Precision (0.9396), FOM (0.8889), and
Recall (0.9428). This robust, interpretable, and scalable approach not only advances urban
expansion modeling but also provides critical insights for urban planners and policymakers.
Keywords: Urban expansion modeling; Deep learning; Transformer networks; Generative
adversarial network; Sustainable urban planning

Irfanullah Khan, Antonio Guerrieri, Edoardo Serra, Giandomenico Spezzano,


A hybrid deep learning model for UWB radar-based human activity recognition,
Internet of Things,
Volume 29,
2025,
101458,
ISSN 2542-6605,
[Link]
([Link]
Abstract: In today’s world, energy efficiency in buildings has become a top priority due to the
significant energy waste caused by the operation of inefficient electrical appliances.
Conventional methods of reducing energy waste cause discomfort for occupants inside
buildings. One promising way to optimize energy consumption is to synchronize appliance
operation with building occupants’ dynamic behavior. Internet of Things (IoT) technologies,
which allow for widespread data collection and execution of Machine Learning (ML)
algorithms, enabled the creation of Smart Buildings (SBs). SBs can learn patterns from the
inhabitant’s behavior residing in, and adjust their operations in accordance with these
behaviors. By doing so, these SBs could reduce energy waste, enhancing resource efficiency
and consequently reduce CO2 gas emissions. Furthermore, they could improve the overall
comfort of the living environment and help with sustainability initiatives. In this context, this
paper proposes a novel approach that uses a hybrid deep-learning model to recognize
complex human activities based on data collected from ultra-wideband (UWB) radar
technology. Our approach, called Hybrid Deep Learning Model for Activity Recognition
(HDL4AR), includes long-short-term memory (LSTM) and a one-dimensional convolutional
neural network (1D-CNN). We deploy a real-time case study by collecting data from 22
participants involved in 10 diverse activities at the headquarters of the ICAR-CNR in the IoT
Laboratory, Italy. Moreover, we conducted a comprehensive benchmark of the HDL4AR
approach against various statistical techniques and other deep learning models recently
introduced in the literature. Results show that our proposed approach outperformed
conventional methods and achieved an impressive accuracy of 98.42%.
Keywords: Internet of Things; Smart buildings; Human activity recognition; UWB radar;
Artificial intelligence; Neural networks; LSTM

Farhatullah, Xin Chen, Deze Zeng, Rahmat Ullah, Rab Nawaz, Jiafeng Xu, Tughrul Arslan,
A deep learning approach for non-invasive Alzheimer’s monitoring using microwave radar
data,
Neural Networks,
Volume 181,
2025,
106778,
ISSN 0893-6080,
[Link]
([Link]
Abstract: Over 50 million people globally suffer from Alzheimer’s disease (AD), emphasizing
the need for efficient, early diagnostic tools. Traditional methods like Magnetic Resonance
Imaging (MRI) and Computed Tomography (CT) scans are expensive, bulky, and slow.
Microwave-based techniques offer a cost-effective, non-invasive, and portable solution,
diverging from conventional neuroimaging practices. This article introduces a deep learning
approach for monitoring AD , using realistic numerical brain phantoms to simulate scattered
signals via the CST Studio Suite. The obtained data is preprocessed using normalization,
standardization, and outlier removal to ensure data integrity. Furthermore, we propose a
novel data augmentation technique to enrich the dataset across various AD stages. Our deep
learning approach combines Recursive Feature Elimination (RFE) with Principal Component
Analysis (PCA) and Autoencoders (AE) for optimal feature selection. Convolution Neural
Network (CNN) is combined with Gated Recurrent Unit (GRU), Bidirectional Long Short Term
Memory (Bidirectional-LSTM), and Long Short-Term Memory (LSTM) to improve
classification performance. The integration of RFE-PCA-AE significantly elevates
performance, with the CNN+GRU model achieving an 87% accuracy rate, thus outperforming
existing studies.
Keywords: Alzheimer’s disease; Classification; Deep learning; Data augmentation; Microwave
scattering; Signal processing

Xing Wang, Haiqin Chen, Ang Zhou, Ye Chen,


Rainfall intensity estimation at night using deep learning and urban surveillance cameras in
Jiangsu Province, China,
Journal of Hydrology: Regional Studies,
Volume 64,
2026,
103112,
ISSN 2214-5818,
[Link]
([Link]
Abstract: Study region
This study was conducted in the Yangtze River Delta of eastern China, focusing on the highly
urbanized corridor of Nanjing, Yangzhou, and Wuxi (Jiangsu Province), where the nighttime
surveillance videos were collected during 2022–2025.
Study focus
Nighttime rainfall measurement from surveillance video remains challenging due to low
visibility, uneven illumination, and complex background noise. To address these issues, this
study proposes NightRAIN-Net (Nighttime Rainfall Adaptive and Integrated Network), a novel
deep learning (DL) framework tailored for nighttime rainfall estimation. The framework
integrates two key modules: Rain-Adaptive Channel Enhancement, which adapts to
nighttime lighting variations to enhance raindrop visibility, and Selective Raindrop
Localization, which captures raindrop’s shape and structure, mitigating interference from
complex backgrounds. Furthermore, NightRAIN-Net integrates LSTM for modeling short-
term fluctuations in rainfall intensity and Transformer for learning long-range dependencies,
enabling robust performance across diverse precipitation types, from light drizzle to extreme
rainfall.
New hydrological insights
Real-world experimental results demonstrate that NightRAIN-Net achieves a Mean Absolute
Error (MAE) of 3.22 mm/h and a Root Mean Squared Error (RMSE) of 3.88 mm/h, while
remaining stable across different scenarios and varying camera parameters. It exhibits stable
performance across different rainfall scenarios, outperforming state-of-the-art methods.
These findings indicate that camera networks can provide scalable, near-continuous (24-
hour) high-frequency rainfall information in Yangtze River Delta, supporting urban
hydrological monitoring, rapid flood/urban waterlogging early warning, and disaster risk
mitigation.
Keywords: Nighttime Rainfall; Surveillance Cameras; Deep Learning; Rainfall Intensity
Estimation

PV Vinod, MD Behera, A Jaya Prakash, R Hebbar, SK Srivastav,


A novel multitask transformer deep learning architecture for joint classification and
segmentation of horticulture plantations using very High-Resolution satellite imagery,
Computers and Electronics in Agriculture,
Volume 227, Part 1,
2024,
109540,
ISSN 0168-1699,
[Link]
([Link]
Abstract: This study introduces MultiTaskDeiTUNet, a novel multitask deep learning
architecture designed to tackle the dual challenges of classifying tree plantation densities
and segmenting tree crowns in high-resolution (0.7 m) satellite imagery. The core challenge
lies in the overlapping spatial patterns of various tree species and densities, complicating the
accurate extraction and classification of individual tree crowns. MultiTaskDeiTUNet
integrates the Data-efficient Image Transformer (DeiT) with the U-Net model, harnessing
DeiT’s strength in contextual detail recognition and spatial dependency capture for density
classification, alongside U-Net’s proficiency in capturing low-level features for precise crown
segmentation. By addressing the complexities of high-resolution data handling and the
simultaneous execution of classification and segmentation tasks, MultiTaskDeiTUNet
achieves an average F1 score of 0.91 (± 0.03) and a precise tree crown segmentation with
mIoU of 0.73 (± 0.01). The DeiT backbone adeptly learns shared features such as canopy
shapes and spatial arrangements, which are crucial for both tasks and enhance overall
model performance. Ablation studies underscore the specialized roles of each component:
freezing DeiT’s weights results in reduced classification accuracy with an average F1 score of
0.48 (± 0.08), while freezing U-Net’s weights yields a reduced mIoU of 0.29 (± 0.12) This
differentiation highlights DeiT’s excellence in classification tasks and U-Net’s superiority in
segmentation. Substituting the DeiT model with a standard ViT model further highlights the
effectiveness of DeiT, as the ViT model demonstrated lower accuracy, with an average F1
score of 0.87 (±0.05) compared to DeiT’s F1 score of 0.91 (±0.03). Statistical analysis revealed
right-skewed distributions in tree crown areas across density categories. The efficacy of
MultiTaskDeiTUNet in tree plantation analysis indicates its potential applicability to a wide
range of horticultural plants. Customizing the architecture to species-specific characteristics
and varying image resolutions could provide valuable insights for improving management
and conservation practices across diverse agricultural and forest ecosystems.
Keywords: DeiT, U-Net; Transformer; Joint loss; Intersection of Union (IoU); Classification and
segmentation

Caihua Hao, Xinyong Mao, Tao Ma, Songping He, Bin Li, Hongqi Liu, Fangyu Peng, Lei Zhang,
A novel deep learning method with partly explainable: Intelligent milling tool wear
prediction model based on transformer informed physics,
Advanced Engineering Informatics,
Volume 57,
2023,
102106,
ISSN 1474-0346,
[Link]
([Link]
Abstract: With the trend of lightweight in the field of intelligent electric vehicles and 3C, the
demand for high precision machining of aluminum alloy parts is growing. And tool condition
monitoring (TCM) is very important for quality control of parts, so intelligent high-accuracy
wear prediction of aluminum alloy high precision machining tools has great industrial
application value at present and in the future. This paper presents a novel TCM model (Conv-
PhyFormer) of Transformer with physics informed. The model has excellent ability to capture
short-term and long-term dependencies from nonlinear cutting time series data when there
are few training samples. The embedded hard physical constraint and soft physical
constraint in the model make the model partially interpretable. Soft physical constraint in
the form of one-dimensional causal convolution can help the proposed model better learn
the local context. Hard physical constraint in the form of the mathematical equation
representing cutting physical knowledge are embedded, thus the model does not need to
learn this knowledge from time series data from scratch. A large number of analysis results
of aluminum alloy machining experimental data show that the proposed Conv-PhyFormer
has significantly superior prediction accuracy and robustness compared with the current
three popular deep learning models for TCM. Embedded soft and hard physical constraints
can significantly reduce the training epochs of Transformer prediction model.
Keywords: Tool condition monitoring; Deep learning; Transformer; Cutting physics
knowledge; Explainability; High Precision Machining

Muhammed Celik, Ozkan Inik,


Review of deep learning-based segmentation methods: Popular approaches, literature gaps,
and opportunities,
Displays,
Volume 91,
2026,
103225,
ISSN 0141-9382,
[Link]
([Link]
Abstract: One of the core challenges in computer vision is devising segmentation techniques
that can reliably disentangle objects from one another and from their backgrounds. To
identify gaps and inspire new solutions, this paper offers a comprehensive literature survey
of over two hundred deep-learning-based segmentation methods, evaluating their
performance across eleven benchmark datasets and common metrics. We categorize these
approaches into three principal families Hierarchical Feature Reconstruction Models, Region-
Proposal Based Architectures, and Hybrid and Advanced Architectural Approaches and chart
their historical evolution graphically. Our analysis shows that multi-scale and pyramid
networks, together with transformer-based designs, consistently achieve the highest mIoU
scores, yet often at the cost of increased computational complexity and annotation
requirements. We observe clear opportunities for future work in developing lightweight,
real-time-capable architectures, reducing reliance on dense pixel labels through semi- and
self-supervised learning, enhancing cross-domain generalization, and integrating multi-
modal inputs all while maintaining explainability and uncertainty quantification for real-
world deployment.
Keywords: Deep learning; Image segmentation; Benchmark datasets; Evaluation metrics

Zhigang Cheng, Zhizhou He, Peng Pan,


3D reconstruction of subsurface pipes and cavities using ground penetrating radar based on
deep learning,
NDT & E International,
Volume 158,
2026,
103579,
ISSN 0963-8695,
[Link]
([Link]
Abstract: Detecting subsurface pipes and cavities is important in urban infrastructure
management, but existing methods struggle to accurately reconstruct the 3D shapes of deep
subsurface objects. This study pioneers a new paradigm for this task by reformulating the ill-
posed permittivity regression problem as a 3D semantic segmentation problem. A novel
neural network, 3DReconNet, to predict the material type of each subsurface voxel from
ground penetrating radar (GPR) data was proposed. This approach leverages the intrinsic
relationship between material composition and reflected signal intensity to simultaneously
recover both geometry and material properties. A dataset of 3150 synthetic cases was
generated using full-scale simulation models and a Markov model-based algorithm to
simulate irregular cavities. The 3DReconNet adopts a U-shaped architecture and
incorporates residual connections to reduce information loss. The network is trained using
the Dice Loss function regularized with total variation (TV) constraints, which enhances
geometric consistency and reconstruction accuracy. The proposed method was validated
using both simulated and experimental data, and the qualitative as well as quantitative
results confirmed its effectiveness, robustness, and generalizability.
Keywords: Ground penetrating radar (GPR); Subsurface pipes; Subsurface cavities; Deep
learning; 3D reconstruction

Zhenhua Li, Jiuxi Cui, Heping Lu, Feng Zhou, Yinglong Diao, Zhenxing Li,
Prediction method for instrument transformer measurement error: Adaptive decomposition
and hybrid deep learning models,
Measurement,
Volume 253, Part D,
2025,
117592,
ISSN 0263-2241,
[Link]
([Link]
Abstract: The measurement accuracy of current transformers is crucial for power system
protection and trade fairness. The high penetration of renewable energy into the power grid
has affected the transient performance of power systems, posing significant challenges for
accurate current transformer measurement. To address this issue, this paper proposes a
prediction model for transformer measurement accuracy based on an adaptive dual-modal
decomposition strategy and a hybrid deep learning architecture. The framework integrates
an enhanced Adaptive Time-Varying Filter (A-TVF), an enhanced Adaptive Variational Mode
Decomposition (A-VMD), the Residual Error Index (REI), and the Maximum Information
Coefficient (MIC). First, A-TVF preprocesses the collected data by setting REI as the
optimization objective to adaptively adjust filter construction parameters, including the B-
spline order, bandwidth threshold, and decomposition number, and decomposes the
collected ratio error sequence to reduce the non-stationarity of the original sequence.
Subsequently, indices such as PE and Kurt are used to screen the decomposed sub-
sequences and reconstruct the complex components. Then, A-VMD is applied to further
decompose the complex components, minimizing MIC by adaptively determining the
decomposition number, penalty factor, convergence accuracy, and fidelity parameters.
Afterward, the complexity of the subcomponents obtained from the secondary
decomposition is calculated, and the entire sequence is reconstructed. Finally, a hierarchical
prediction model integrating Temporal Convolutional Networks (TCN), Bidirectional Gated
Recurrent Units (BiGRU), and a Multi-Head Attention mechanism (MHA) is employed to
predict the reconstructed components and generate the final results. Experimental results
demonstrate that the proposed adaptive dual-modal decomposition method significantly
improves prediction performance: compared with non-decomposition models, RMSE, MAE,
and SMAPE were reduced by an average of 50.12%, 46.09%, and 37.70% in global
decomposition scenarios, and by 25.92%, 23.69%, and 19.96% in rolling decomposition
scenarios, respectively. These results validate the effectiveness of the proposed method in
reducing data complexity and improving the accuracy and stability of Ratio Error predictions.
Keywords: ECT; Ratio error prediction; Adaptive dual-modal decomposition; Decomposition
and combination strategy; Hybrid deep model; Measurement accuracy

Nirdesh Kumar Sharma, Manabendra Saharia,


DeepSARFlood: Rapid and automated SAR-based flood inundation mapping using vision
transformer-based deep ensembles with uncertainty estimates,
Science of Remote Sensing,
Volume 11,
2025,
100203,
ISSN 2666-0172,
[Link]
([Link]
Abstract: Rapid and automated flood inundation mapping is critical for disaster
management. While optical satellites provide valuable data on flood extent and impact, their
real-time usage is limited by challenges such as cloud cover, limited vegetation penetration,
and the inability to operate at night, making real-time flood assessments difficult. Synthetic
Aperture Radar (SAR) satellites can overcome these limitations, allowing for high-resolution
flood mapping. However, SAR data remains underutilized due to less availability of training
data, and reliance on labor-intensive manual or semi-automated change detection methods.
This study introduces a novel end-to-end methodology for generating SAR-based flood
inundation maps, by training deep learning models on weak flood labels generated from
concurrent optical imagery. These labels are used to train deep learning models based on
Convolutional Neural Networks (CNN) and Vision Transformer (ViT) architectures, optimized
through multitask learning and model soups. Additionally, we develop a novel gain algorithm
to identify diverse ensemble members and estimate uncertainty through deep ensembles.
Our results show that ViT-based and CNN-ViT hybrid architectures significantly outperform
traditional CNN models, achieving a state-of-the-art Intersection over Union (IoU) score of
0.72 on the Sen1Floods11 test dataset, while also providing uncertainty quantification.
These models have been integrated into an open-source and fully automated, Python-based
tool called DeepSARFlood, and demonstrated for the Pakistan floods of 2022 and Assam
(India) floods of 2020. With its high accuracy, processing speed, and ability to estimate
uncertainty, DeepSARFlood is optimized for real-time deployment, processing a 1° × 1°
(12,100 km2) area in under 40 s, and will complement upcoming SAR missions like NISAR
and Sentinel 1-C for flood mapping.
Keywords: Flood mapping; Synthetic aperture radar; Deep ensembles; Uncertainty
estimation; Vision-transformer

Wei Wang, Ruobing Song, Yunxiao Wu, Li Zheng, Wenyu Zhang, Zhaoxi Chen, Gang Li, Zhifei
Xu,
Deep learning-based automated diagnosis of obstructive sleep apnea and sleep stage
classification in children using millimeter-wave radar and pulse oximeter,
Sleep Health,
Volume 11, Issue 6,
2025,
Pages 859-867,
ISSN 2352-7218,
[Link]
([Link]
Abstract: Study objectives
Due to the high cost, complexity, and workload of polysomnography, a radar-based sleep
monitoring device, QSA600, has been developed as a more simplified alternative for
children. This study evaluates its agreement with polysomnography for obstructive sleep
apnea diagnosis and sleep staging.
Methods
This diagnostic accuracy study included 281 children (1-18 years) who underwent
simultaneous polysomnography and QSA600 monitoring at Beijing Children's Hospital from
September-November 2023. QSA600 recordings were automatically analyzed using a deep
learning model, while polysomnography data were manually scored.
Results
The obstructive apnea-hypopnea index (OAHI) obtained from QSA600 and polysomnography
demonstrates a high level of agreement with an intraclass correlation coefficient of 0.945
(95% CI: 0.93-0.96). Bland-Altman analysis indicated that the mean difference of obstructive
apnea-hypopnea index between QSA600 and polysomnography was −0.10 events/h (95% CI:
−11.15 to 10.96). The deep learning model evaluated through cross-validation showed good
sensitivity (81.8%, 84.3%, and 89.7%) and specificity (90.5%, 95.3%, and 97.1%) values for
diagnosing children with OAHI >1, OAHI >5, and OAHI >10. The area under the receiver
operating characteristic curve was 0.923, 0.955, and 0.988, respectively. For sleep stage
classification, the model achieved Kappa coefficients of 0.854, 0.781, and 0.734, with
corresponding overall accuracies of 95.0%, 84.8%, and 79.7% for Wake-Sleep classification,
Wake-REM-Light-Deep classification, and Wake-REM-N1-N2-N3 classification, respectively.
Conclusions
QSA600 has demonstrated high agreement with polysomnography in diagnosing obstructive
sleep apnea and performing sleep staging in children. The device is portable, low-burden,
and suitable for follow-up and long-term pediatric sleep assessment.
Keywords: Obstructive sleep apnea; Children; Deep learning; Millimeter-wave radar;
Portable sleep monitoring device; Polysomnography

Xinrui Zhao, Zeng Liu, Qi Hu, Jianglong Sun, Xiaoyan Yang,


Significant wave height estimation and prediction from synthetic X-band radar data by
spatio-temporal deep neural networks,
Ocean Engineering,
Volume 339, Part 1,
2025,
122061,
ISSN 0029-8018,
[Link]
([Link]
Abstract: Estimation and prediction of real-time significant wave height (SWH) is a
fundamental requirement for the safety of offshore activities. This study utilizes a spatio-
temporal deep neural network model to effectively estimate and predict the SWH by
extracting key spatial and temporal features from synthetic X-band radar [Link] study
considers three deep neural networks: InceptionV3, ResNet and Vision Transformer (ViT) to
extract multi-scale spatial features from radar images for SWH estimation. Subsequently, a
gated recurrent unit (GRU) is employed on these spatial features to perform the time-series
SWH prediction. Irregular waves with sea state ranging from 4 to 6 were considered based
on the synthetic radar data. Results indicate that the ResNet model with deep residual
structure performs the best in both the estimation and prediction tasks, demonstrating
excellent generalization ability and adaptability to complex sea conditions. The ViT model
shows outstanding performance in scenarios without out-of-distribution data. While the
InceptionV3 model is inferior, it exhibits significant improvement for the SWH prediction
when the GRU is incorporated.
Keywords: Significant wave height(SWH); X-band marine radar; InceptionV3; ResNet; Vision
Transformer(ViT); Gated recurrent unit(GRU)

Tai Dinh, Dat Tran, Zdena Dobešová, Huynh Van Hong, Daniil Lisik, Rameesh Khan,
An efficient fusion-based deep learning framework for land use and land cover image
clustering,
Engineering Applications of Artificial Intelligence,
Volume 161, Part B,
2025,
112061,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Land use and land cover (LULC) analysis is vital for understanding spatial dynamics
and informing environmental management, urban planning, and sustainable development.
Traditional approaches, such as manual surveys and conventional image clustering methods,
often face limitations in scalability and adaptability. This paper presents a novel deep
learning framework that combines the Vision Transformer (ViT) and Variational Autoencoder
(VAE) to extract complementary feature representations for LULC image clustering. The ViT
tokenizes image patches to capture high-level semantic features, while the VAE models
latent structures to integrate contextual and structural information. To further improve
clustering performance, the framework incorporates Uniform Manifold Approximation and
Projection (UMAP) for dimensionality reduction followed by k-means++ clustering, enabling
a scalable and robust solution for diverse datasets. Experiments on multiple datasets,
including the Urban Atlas LULC 2018 dataset and recent LULC maps of Japan and Vietnam,
demonstrate the framework’s superior ability to capture complex LULC patterns compared
to traditional methods. The datasets and source code will be made publicly available at
[Link] This framework has broad applications across
geospatial and remote sensing engineering, civil and environmental engineering, agricultural
planning, transportation, and urban development.
Keywords: Artificial intelligence; Land use and land cover; Urban land use; Deep image
clustering; Transformer; Vision transformer; Variational autoencoder; Uniform manifold
approximation and projection for dimension reduction (UMAP); K-means++

Chongxing Ji, Yuan Xu,


trajPredRNN+: A new approach for precipitation nowcasting with weather radar echo images
based on deep learning,
Heliyon,
Volume 10, Issue 18,
2024,
e36134,
ISSN 2405-8440,
[Link]
([Link]
Abstract: :Short-term rainfall prediction is a crucial and practical research area, with the
accuracy of rainfall prediction, particularly for heavy rainfall, significantly impacting people's
lives, property, and even their safety. Existing models, such as ConvLSTM, TrajGRU, and
PredRNN, exhibit limitations in capturing fine-grained appearances due to insufficient
memory units or addressing positional misalignment issues, thereby compromising the
accuracy of model predictions. In this study, we propose trajPredRNN+, an innovative
approach that integrates the trajectory segmentation model and the PredRNN deep learning
model to address both limitations in nowcasting precipitation using weather radar echo
images. By incorporating attention mechanisms, the model demonstrates an enhanced focus
on short-term and imminent heavy rainfall events. To ensure improved stability during
training, a residual network is introduced. Lastly, a more rational and effective training loss
function is proposed, encompassing weight mechanism, SSIM index, and GAN loss. To
validate the proposed model, we conducted a comparative experiment and an ablation
experiment using the radar echo map dataset obtained from the Shenzhen Meteorological
Bureau. The results of these experiments demonstrate that our model has achieved
significant improvements across multiple key performance indicators.
Keywords: Radar echo map; Deep learning; Precipitation nowcasting; PredRNN; GAN

Hui Wang, Qinghua Liu, Lijun Zhou,


Underground target localization method for ground penetrating radar based on deep
learning,
Measurement,
Volume 253, Part B,
2025,
117647,
ISSN 0263-2241,
[Link]
([Link]
Abstract: To tackle the challenge of subsurface target localization under interference in field
scenarios, a novel two-level cascade network referred to as dual cascade is proposed. The
first level, Cascade-1, is a deep feature extraction network designed to extract and eliminate
direct wave interference signals. On this basis, Cascade-2 is developed using domain
knowledge from ground penetrating radar as prior information, and it incorporates an
attention mechanism along with a feature fusion strategy to enhance the accuracy of target
feature hyperbola detection. Subsequently, the least squares method is employed to fit the
feature hyperbola, and location estimation is performed based on geometric equations. The
proposed cascade network model has demonstrated superior performance compared to
other algorithms, such as column-connection clustering algorithm, YOLOv9, and Faster R-
CNN, in terms of the composite metric F1, which validates the model’s effectiveness in
extracting the feature hyperbola. Additionally, the proposed localization method has
exhibited greater accuracy than the conventional full waveform inversion algorithm.
Keywords: Ground Penetrating Radar; Buried Target Location; Domain Knowledge; Deep
Learning; Cascade Network

Yilin Bao, Xiangtian Meng, Xingnan Liu, Xue Wang, Zhengchao Qiu, Huanjun Liu, Mingchang
Wang, Abdul Mounem Mouazen,
Integrating bi-dynamic strategy and multivariate deep learning algorithms to predict high-
accuracy, long-term cropland soil organic matter,
International Soil and Water Conservation Research,
2025,
100598,
ISSN 2095-6339,
[Link]
([Link]
Abstract: Soil organic matter (SOM) is a key indicator for assessing soil health and carbon
neutrality, while environmental heterogeneity and dynamic sensitivity changes between
SOM and environmental variables can reduce prediction accuracy. This study developed a
deep learning model that accounts for spatial variability and sensitivity changes in the soil
environment, using 2284 soil samples and 64,802 Landsat TM/OLI images as inputs. A bi-
dynamic strategy is proposed, utilizing a Gaussian mixture model to dynamically partition
the study area and account for changes in environmental heterogeneity across periods. A
multivariate deep learning algorithm, T-CNN-GNN, which combines Transformer, a
convolutional neural network (CNN), and a graph neural network (GNN), was developed to
extract advanced spatio-temporal features, thereby enhancing the accuracy of long-term
SOM spatial distribution predictions. The results showed that the integration of the bi-
dynamic strategy and multivariate deep learning model achieved the highest prediction
accuracy, with a root mean square error (RMSE) of 9.49 g/kg and a coefficient of
determination (R2) of 0.77. The dynamic partitioning strategy effectively captured spatial
variations in environmental heterogeneity across periods. Over the past 40 years, the SOM
content in Northeast China decreased from 41.52 ± 0.24 g/kg to 37.94 ± 0.21 g/kg. This
study showed that SOM maps generated by the bi-dynamic strategy closely matched
measured SOM values, offering a promising method for long-term, high-accuracy soil
mapping.
Keywords: Soil organic matter; Gaussian mixed partitioning; Structural equation modelling;
Deep learning

Yanfei Peng, Jiang He, Qiangqiang Yuan, Shouxing Wang, Xinde Chu, Liangpei Zhang,
Automated glacier extraction using a Transformer based deep learning approach from multi-
sensor remote sensing imagery,
ISPRS Journal of Photogrammetry and Remote Sensing,
Volume 202,
2023,
Pages 303-313,
ISSN 0924-2716,
[Link]
([Link]
Abstract: Glaciers serve as sensitive indicators of climate change, making accurate glacier
boundary delineation crucial for understanding their response to environmental and local
factors. However, traditional semi-automatic remote sensing methods for glacier extraction
lack precision and fail to fully leverage multi-source data. In this study, we propose a
Transformer-based deep learning approach to address these limitations. Our method
employs a U-Net architecture with a Local-Global Transformer (LGT) encoder and multiple
Local-Global CNN Blocks (LGCB) in the decoder. The model design aims to integrate both
global and local information. Training data for the model were generated using Sentinel-1
Synthetic Aperture Radar (SAR) data, Sentinel-2 multispectral data, High Mountain Asia
(HMA) Digital Elevation Model (DEM), and Shuttle Radar Topography Mission(SRTM) DEM.
The ground truth was obtained for a glaciated area of 1498.06 km2 in the Qilian mountains
using classic band ratio and manual delineation based on 2 m resolution GaoFen (GF)
imagery. A series of experiments including the comparison between different models, model
modules and data combinations were conducted to evaluate the model accuracy. The best
overall accuracy achieved was 0.972. Additionally, our findings highlight the significant
contribution of Sentinel-2 data to glacier extraction.
Keywords: Glacier; Deep learning; Transformer; Multi-source data; Qilian mountain

Shoujin Sun, Chi Zhou, Guixian Sun, Zhexun Lian,


Multimodal fusion of ECG and chest X-ray with deep learning and radiomics for cardiac
diagnosis: A machine learning-based nomogram approach,
Journal of Radiation Research and Applied Sciences,
Volume 18, Issue 4,
2025,
102020,
ISSN 1687-8507,
[Link]
([Link]
Abstract: Objective
To develop and validate a comprehensive, multimodal diagnostic framework combining ECG,
chest X-ray imaging, radiomic features, heart rate variability (HRV), and clinical data for
accurate classification of major cardiac conditions.
Materials and methods
We enrolled 2285 patients from seven clinical centers, ultimately including 1987 for training
and 498 for external testing. ECG signals were converted into standardized 2D grayscale
images and processed using Vision Transformer (ViT) and EfficientNet models to extract
deep features. Chest X-ray images underwent preprocessing, segmentation by expert
radiologists, and feature extraction using the same deep learning architectures,
supplemented by handcrafted radiomic analysis. All features underwent intraclass
correlation coefficient (ICC) filtering (threshold ≥0.75) to ensure reliability. Dimensionality
reduction and feature selection were performed using PCA and LASSO. The resultant
features were fused with clinical variables (e.g., age, BMI, blood pressure, diabetes) and HRV
metrics. Several machine learning classifiers were trained and validated. Interpretability was
assessed using SHAP analysis, and a nomogram was developed for clinical decision-making.
Decision curve analysis (DCA) was applied to evaluate clinical utility across risk thresholds.
Results
The integrated model achieved a training accuracy of 96.0 % and external test accuracy of
94.0 %, with an AUC of 0.95. Fused ECG features outperformed single-model inputs
(AUC = 0.90), and radiomic features from chest X-ray provided robust discriminatory power
(AUC = 0.86). SHAP analysis revealed distinct feature contributions across cardiac categories,
and DCA showed superior net clinical benefit of the XGBoost model. Nomograms enabled
individualized risk assessment.
Conclusions
This study demonstrates that integrating multimodal data enhances cardiac disease
classification accuracy and interpretability. The proposed framework is a scalable,
reproducible solution for real-world clinical deployment.
Keywords: Radiomic features; Deep learning; Electrocardiogram imaging; Chest X-ray
analysis; Multimodal fusion; Cardiac diagnostics

Lv Zhou, TianLiang Chen, Fei Yang, YuanJin Pan, Ling Huang, Xiang Huang, PengDe Lai,
MSFlood-Net: A physically informed deep learning model integrating multi-source data for
flood inundation mapping,
Environmental Modelling & Software,
Volume 196,
2026,
106779,
ISSN 1364-8152,
[Link]
([Link]
Abstract: This study proposes MSFlood-Net, an enhanced U-Net-based deep learning model
for flood extent mapping. The model integrates Synthetic Aperture Radar (SAR), optical
imagery, and topographic inputs including the Digital Elevation Model (DEM) and Height
Above Nearest Drainage (HAND). Multi-scale attention and dilated convolutions enhance
feature representation in complex terrain. A multi-source dataset built upon the publicly
available GF-FloodNet dataset was constructed for training and evaluation. MSFlood-Net
achieves 97.187 % accuracy and a 96.756 % F1 score, outperforming U-Net and DeepLabV3
baselines. It shows strong robustness in delineating flood extent across rivers, reservoirs,
and urban transition areas. Physically informed inputs help reduce false positives from
shadows, clouds, and wet surfaces. MSFlood-Net provides a practical and scalable solution
for flood monitoring and supports integration into broader environmental modelling
systems.
Keywords: Flood inundation mapping; Deep learning; Multi-source remote sensing data;
Physically informed

Yash Soni, Malhaar Goswami, Nishit Prabhakar Shetty, Dhiraj,


Millimeter-wave radar for intelligent sensing: A comprehensive review of techniques,
applications, and challenges,
Computers and Electrical Engineering,
Volume 128, Part A,
2025,
110696,
ISSN 0045-7906,
[Link]
([Link]
Abstract: Millimeter-wave (mmWave) radar sensing has established itself as a robust
technology across diverse applications, such as automotive, healthcare, security, and smart
homes. Its exceptional capacity to function effectively in varying environmental conditions,
detect concealed objects, sense physiological signals, and facilitate precise target detection
positions it as a pivotal enabler for next-generation sensing solutions. The survey employs
bibliometric analysis to critically evaluate the existing literature surrounding mmWave radar,
highlighting key research trends, notable publications, and the challenges faced within the
field. This work presents a comprehensive examination of mmWave radar-based sensing,
detailing its fundamental operating principles, signal processing methodologies,
advancements in hardware, and the latest developments in machine learning applications. It
also addreses the key challenges in signal processing, including resolution enhancement,
environmental adaptability, and data fusion with complementary sensors such as LiDAR and
cameras. Furthermore, explored the potential of deep learning techniques to enhance target
classification, activity recognition, gesture identification, and healthcare applications while
addressing concerns related to accuracy and precision. This survey also sheds light on
emerging trends by assessing the strengths, limitations, and prospects of mmWave radar
technology. This review aims to provide insightful guidance for researchers and practitioners
committed to advancing radar-based sensing and its real-world implementations.
Keywords: mmWave radar; FMCW radar; Wireless sensing; Machine learning; Deep learning;
Application taxonomy

Hua Wang, Qiangyu Zeng, Hao Wang, Jianxin He, Tiantian Yu, Guangpu Liu,
Temporal super-resolution reconstruction of weather radar echoes using a deep learning
approach,
Expert Systems with Applications,
Volume 300,
2026,
130189,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Severe convective weather events are characterised by rapid evolution and high
destructive potential, requiring weather radars to provide observations with high temporal
resolution. However, current S-band weather radar systems, constrained by their volumetric
scanning strategies, often fail to capture the rapidly changing features of these systems
promptly. To address this limitation, we propose EMAIRA-VFI, a deep learning–based
method for temporal super-resolution reconstruction of radar echoes, which enhances the
temporal resolution of radar data to meet the demands of severe convective weather
monitoring. By introducing an inter-frame attention mechanism, the proposed method
effectively fuses spatiotemporal features from sequential radar echoes, enabling accurate
modelling of dynamic weather evolution and the generation of continuous, high-temporal-
resolution radar echoes. Compared with conventional temporal interpolation methods,
EMAIRA-VFI demonstrates significant improvements in both interpolation accuracy and the
preservation of fine-scale meteorological structures. Experimental results show that the
model not only enhances the capability of S-band radars in monitoring rapidly evolving
weather events but also provides a new perspective for spatiotemporal fusion and the
intelligent application of radar data. We have open-sourced the code for this work at
[Link]
Keywords: Temporal super-resolution; Radar echo; Inter-frame attention mechanism

Kuiyu Chen, Jingyi Zhang, Si Chen, Shuning Zhang, Huichang Zhao,


Automatic modulation classification of radar signals utilizing X-net,
Digital Signal Processing,
Volume 123,
2022,
103396,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Automatic modulation classification (AMC) of radar signals has long been a
challenge, especially in the area of electronic reconnaissance, where collecting and labeling
numerous signal samples are usually harsh and impracticable. In this article, a novel
recognition network is proposed for detecting radiation signals under intense noise
background with only simulation samples for model training. Owing to the X-shaped
structure, the recognition network is named X-net. A residual learning convolutional
denoising autoencoder (RLCDAE) and a supplementary classification network based on noise
level estimation (NLE) constitute the X-net. Via pre-training of RLCDAE, the robustness
against noise is enhanced. Then, a supplementary classification network further improves
the recognition performance under low signal-to-noise ratios (SNRs). To evaluate the
method, comparative experiments with some excellent algorithms on noise immunity and
recognition accuracy are conducted. The cognitive system can still recognize ten kinds of
radar signals with an overall precision of 96% even when the SNR is −8 dB. Furthermore,
simulation samples trained model is verified by measured data. Outstanding performance
proves the effectiveness and superiority of the proposed method on cognitive radar signals
under low SNRs.
Keywords: Automatic modulation classification; Radar signals; Residual learning
convolutional denoising autoencoder; Supplementary classification networks; Noise level
estimation; Measured data
Yiming Zhang, Jiaqi Li, Zheng Tong, Weiguang Zhang, Xiyuan Shen,
A direction-aware and expert-inspired network for internal crack size detection using on-site
ground penetrating radar data,
Engineering Applications of Artificial Intelligence,
Volume 165, Part A,
2026,
113414,
ISSN 0952-1976,
[Link]
([Link]
Abstract: The size of internal cracks is a key basis for determining maintenance measures.
Existing methods primarily utilize Ground Penetrating Radar (GPR) signals or images to
detect crack size, but still face two challenges: limited robustness in interpreting on-site data
based on GPR signals and the inability to directly characterize the crack size based on raw
GPR image features. To address these issues, this study has proposed a novel internal crack
size detection network, which was trained by a dataset with on-site GPR B-Scans and
interpreted crack size labels. In the proposed model, a deformable Cross Stage Partial (CSP)
block is first used to extract the irregular hyperbolic features of crack reflected waves from
B-Scans. Then, a directional fusion attention module is designed to construct direction-
aware channel attention and generate spatial interaction weights. Finally, a bipartite graph
matching detection head is proposed to emulate the expert behavior to analyze crack
reflected waves from a global B-Scan perspective, outputting trapezoidal-sized boxes to
detect internal crack size. The experimental results demonstrate that the proposed model
exceeds other state-of-the-art models on the tasks thanks to the channel-spatial weight
aggregation and global output strategy via bipartite graph matching. Additionally, the model
exhibits good stability across various antenna frequencies and pavement structures. The on-
site testing indicates that the predicted crack sizes sufficiently meet engineering
requirements in most scenarios, though challenges remain in detecting the bottom width of
small and water-saturated cracks.
Keywords: Asphalt pavement; Non-destructive testing; Ground penetrating radar; Internal
crack; Size detection

Kaixiang Zhang, Jiaxiang Zhang, Xinrui Han, Yilin Wang, Bo Wang, Quanhua Liu,
OSCJC: An open-set compound jamming cognition method for radar systems in high-
intensity electromagnetic warfare,
Defence Technology,
Volume 55,
2026,
Pages 436-455,
ISSN 2214-9147,
[Link]
([Link]
Abstract: In high-intensity electromagnetic warfare, radar systems are persistently subjected
to multi-jammer attacks, including potentially novel unknown jamming types that may
emerge exclusively under wartime conditions. These jamming signals severely degrade radar
detection performance. Precise recognition of these unknown and compound jamming
signals is critical to enhancing the anti-jamming capabilities and overall reliability of radar
systems. To address this challenge, this article proposes a novel open-set compound
jamming cognition (OSCJC) method. The proposed method employs a detection-
classification dual-network architecture, which not only overcomes the false alarm and
misdetection issues of traditional closed-set recognition methods when dealing with
unknown jamming but also effectively addresses the performance bottleneck of existing
open-set recognition techniques focusing on single jamming scenarios in compound
jamming environments. To achieve unknown jamming detection, we first employ a
consistency labeling strategy to train the detection network using diverse known jamming
samples. This strategy enables the network to acquire highly generalizable jamming
features, thereby accurately localizing candidate regions for individual jamming components
within compound jamming. Subsequently, we introduce contrastive learning to optimize the
classification network, significantly enhancing both intra-class clustering and inter-class
separability in the jamming feature space. This method not only improves the recognition
accuracy of the classification network for known jamming types but also enhances its
sensitivity to unknown jamming types. Simulations and experimental data are used to verify
the effectiveness of the proposed OSCJC method. Compared with the state-of-the-art open-
set recognition methods, the proposed method demonstrates superior recognition accuracy
and enhanced environmental adaptability.
Keywords: Radar compound jamming cognition; Open-set recognition; Detection-
classification dual-network; Time-frequency analysis; Contrastive learning

Ting Dai, Liye Mei, Yue Zhang, Biao Tian, Rui Guo, Teng Wang, Shan Du, Shiyou Xu,
UAVs and birds classification using robust coordinate attention synergy residual split-
attention network based on micro-Doppler signature measurement by using L-band staring
radar,
Measurement,
Volume 222,
2023,
113692,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Developing unmanned aerial vehicles (UAVs) and birds surveillance technologies to
produce accurate descriptions and achieve high classification accuracy is critical in the field
of radar automatic target recognition (RATR). This article proposes a grayscale spectrogram
image-based UAVs and birds classification method using a robust coordinate attention
synergy residual Split-Attention network (RCA-ResNeSt) under the holographic staring radar
system. Specifically, the ResNet structure with Split-Attention is used as an m-D feature
extractor. The CrossNorm and SelfNorm (CNSN) mechanism is then incorporated into the
network to advance generalization robustness. After that, to consider the spatial direction of
the m-D signature, a coordinated attention (CA) mechanism is introduced at the tail end of
the network to enable fine-grained mining of potential m-D features. Experiments are
carried out using a designed radar system. The results show the superiority of the proposed
method over existing approaches in classification accuracy and noise robustness.
Keywords: Micro-doppler (m-D) signature; Time–frequency representation (TFR); Robust
coordinate attention synergy residual split-attention network (RCA-resNeSt); Holographic
staring radar; Radar automatic target recognition (RATR)

Juan Chen, Song Yang, Shengli Zhang,


Advantage of quantum radar with intensity fluctuations,
Physics Letters A,
Volume 564,
2025,
131084,
ISSN 0375-9601,
[Link]
([Link]
Abstract: Quantum radar is a powerful tool for detecting weak targets. In theory, quantum
radar shows a 6 dB quantum advantage in error probability. It is important to determine
whether the advantages of a quantum illumination system exist in the presence of intensity
fluctuations. We studied the advantages of a typical quantum radar with a two-mode
Gaussian quantum entanglement state. By varying the standard deviation of the intensity
fluctuation, the advantages of the quantum radar were identified. It is shown that in the
weak signal intensity regime, the advantage of the quantum radar will decrease. However, in
the strong signal intensity regime, its advantages surpass that of the no-fluctuation case.
This contrasts with the intuitive notion that signal fluctuation would introduce distortion into
the radar system. Our results serve as proof of the advantage of quantum radar and offer
theoretical guidance for the development of more efficient quantum radar systems.
Keywords: Quantum radar; Quantum advantage; Intensity fluctuations

Wei Quan, Wenjing Cheng, Yike Yang, Haiquan Zhao, Zhaoyu Chen, Yunfan Luo,
A signal fingerprint feature extraction method based on decomposition and fusion for radar
emitter individual identification,
Digital Signal Processing,
Volume 164,
2025,
105257,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Radar emitter individual identification is one of the key technologies of modern
electronic countermeasure reconnaissance and electronic intelligence. With the
advancement of radar technology and the increasingly complex electromagnetic
environment, existing methods for identifying emitter are gradually becoming unable to
meet the performance requirements of modern radar individual identification. Aiming at
improving the adaptability of feature extraction for non-cooperative radar emitter signals
and the robustness of individual identification in the complex modern electronic warfare
environment, a signal fingerprint feature extraction method based on decomposition and
fusion is proposed. It firstly integrates signal decomposition and scattering convolution
networks (SCN) to adaptively extract the multi-scale intra-pulse feature of the signal, while
removing the potential noise of the redundant component by energy proportion. And then a
deep feature fusion model based on multi-head self-attention and residual connection is
proposed to fuse the multi-scale features and the time domain features to further extract
signal fingerprint of radar emitter. Experimental results based on the real radar emitter
signals demonstrate that the identification method proposed in this paper can more
effectively extract signal fingerprint features and the identification accuracy reaches 96.45%,
which outperforms other existing identification methods.
Keywords: Radar emitter individual identification; Signal fingerprint feature; Signal
decomposition; Scattering convolution networks (SCN); Fusion

Shuai Xu, Lutao Liu, Zhongkai Zhao,


Unsupervised recognition of radar signals combining multi-block TFR with subspace
clustering,
Digital Signal Processing,
Volume 151,
2024,
104552,
ISSN 1051-2004,
[Link]
([Link]
Abstract: In the realm of radar systems, the proliferation of new modulation techniques
introduces an increased level of complexity in the identification of both existing and
potentially novel modulation schemes. Conventional supervised recognition methodologies,
which rely heavily on labeled datasets, exhibit limitations in distinguishing amongst a
multitude of unlabeled signals. This paper introduces a novel framework that synergizes
Subspace Clustering (SC) with Multi-Block Time-Frequency Representations (TFRs, denoted
as mT), referred to as SC-mT. This framework leverages subspace clustering for the
unsupervised categorization of radar modulations and employs an innovative multi-block
strategy to augment classification precision. The process commences with the generation of
TFR datasets for radar signals, followed by the construction of two distinct multi-block
models, Model-A and Model-B. These models segment the radar TFRs into overlapping
multi-block sets, utilizing the concept of random receptive fields. The subspace clustering
algorithm is then applied to each block within the sub-TFR sets to procure an affinity matrix.
The culmination of this process involves the aggregation of the affinity matrices from all
blocks, facilitating the derivation of classification outcomes via spectral clustering. Empirical
analyses affirm that the sequential integration of the SC-mT algorithm surpasses the
classification efficacy of various conventional and cutting-edge algorithms. Notably, this
algorithm attains a classification accuracy surpassing 90% for ten unlabeled signals, even in
scenarios where the signal-to-noise ratio (SNR) is as low as -4 dB.
Keywords: Radar modulation; Unsupervised classification; Multi-block; Subspace clustering

Ayesha Ibrahim, Muhammad Zakir Khan, Muhammad Imran, Hadi Larijani, Qammer H.
Abbasi, Muhammad Usman,
RadSpecFusion: Dynamic attention weighting for multi-radar human activity recognition,
Internet of Things,
Volume 33,
2025,
101682,
ISSN 2542-6605,
[Link]
([Link]
Abstract: This paper presents RadSpecFusion, a novel dynamic attention-based fusion
architecture for multi-radar human activity recognition (HAR). Our method learns activity-
specific importance weights for each radar modality (24 GHz, 77 GHz, and Xethru sensors).
Unlike existing concatenation or averaging approaches, our method dynamically adapts
radar contributions based on motion characteristics. This addresses cross-frequency
generalization challenges, where transfer learning methods achieve only 11%–34% accuracy.
Using the CI4R dataset with spectrograms from 11 activities, our approach achieves 99.21%
accuracy, representing a 15.8% improvement over existing fusion methods (83.4%). This
demonstrates that different radar frequencies capture complementary information about
human motion. Ablation studies show that while the three-radar system optimizes
performance, dual-radar combinations achieve comparable accuracy (24GHz+77GHz: 96.1%,
24GHz+Xethru: 95.8%, 77GHz+Xethru: 97.2%), enabling flexible deployment for resource-
constrained applications. The attention mechanism reveals interpretable patterns: 77 GHz
radar receives higher weights for fine movements (superior Doppler resolution), while 24
GHz dominates gross body movements (better range resolution). The system maintains
71.4% accuracy at 10 dB SNR, demonstrating environmental robustness. This research
establishes a new paradigm for multimodal radar fusion, moving from cross-frequency
transfer learning to adaptive fusion with implications for healthcare monitoring, smart
environments, and security applications.
Keywords: Human activity recognition; Multi-modal fusion; Attention mechanisms; Cross-
frequency transfer learning

Purabi Sharma, Kandarpa Kumar Sarma,


Attention driven CWT-deep learning approach for discrimination of Radar PRI modulation,
Physical Communication,
Volume 62,
2024,
102237,
ISSN 1874-4907,
[Link]
([Link]
Abstract: With the proliferation of radio frequency (RF) systems and radar applications,
Electronic Warfare (EW) is receiving increasing importance. The analysis of the radar signals
is a critical EW task that decides the nature of counter employments. In an Electronic
Support (ES) system, the challenge is to detect hostile radiation sources efficiently and
trigger a counter response. Detection of types of Pulse Repetition Interval (PRI) modulation
of radar signal significantly facilitates the manifestation of RF emitters during recognition
which is difficult in a dense EW environment. Recent developments in artificial intelligence
(AI) methods suggest that this emerging technology can be effective for such purposes. In
this direction, an automatic approach for recognizing several kinds of complex PRI
modulation based on Continuous Wavelet Transform (CWT) and a combination of the vanilla
Convolutional Neural Network (CNN), a multi-head self-attention (MHSA) mechanism and
the popular Long Short-Term Memory (LSTM) is proposed. The CWT is used to decompose
the PRI modulation sequence and obtain different time–frequency components. Further,
aided by the proposed CNN-MHSA-LSTM combination, the features extracted from the CWT
2D-scalograms are used to execute PRI modulation discrimination. In this method, the
vanilla CNN is employed for the extraction of deep features to figure out the class details
while capturing the spatial attributes. Thereafter, to improve the discriminative power of the
entire framework a MHSA mechanism is used. The temporal attributes are acquired by the
LSTM which works in concert with the CNN for executing the detection of the PRI classes
based on the extracted features. Also to assess the effectiveness of the proposed method,
three models based on ResNet, popular CNN and SqueezeNet are implemented for
benchmark comparison in terms of overall performance and complexity. The simulation
results show that the proposed method enhances performance and achieves robustness in
the noise-filled and imperfect channel knowledge environment. The best recognition
accuracy is 98.3% with 50% spurious pulses in the environment which fluctuates with
imperfect channel knowledge cases.
Keywords: Electronic warfare; PRI modulation; CWT; Convolutional Neural Network; Self-
attention mechanism; Long Short Term Memory

Asim Saleem, Guoyun Lv, Safa Hussein Mohammed,


Deep Learning-Based Radar Fingerprinting for Open-Set Generalization Using Dynamic
Thresholding and Embedding Rejection,
Knowledge-Based Systems,
Volume 333,
2026,
115047,
ISSN 0950-7051,
[Link]
([Link]
Abstract: This paper presents a comprehensive framework for radar-specific emitter
identification (SEI), starting with the simulation of a large-scale radar signal dataset designed
to mimic real hardware impairments. By incorporating diverse distortions–such as phase
noise, frequency jitter, amplitude nonlinearity, and multipath reflections–for multiple radar
types and signal-to-noise ratio (SNR) conditions, we generate a realistic and challenging
dataset, SimRF-14, suitable for learning-based signal analysis. We utilized this dataset to
develop RAFNet, a hybrid deep learning model specifically designed for closed-set and open-
set radar emitter classification. The proposed architecture combines convolutional,
recurrent, and attention-based components to capture spatial, temporal, and contextual
features from normalized I/Q waveforms. For open-set recognition, we integrate the
OpenMax algorithm enhanced with Extreme Value Theory (EVT), where class-wise Weibull
modeling of embedding distances enables outlier detection. In addition, an SNR-adaptive
thresholding mechanism improves open-set reliability under varying noise conditions. The
proposed method achieved 97.43% unknown rejection at -20 dB SNR and maintained low
false positives (<4%) at high SNRs, validating its effectiveness and reliability for practical SEI
scenarios.
Keywords: Radar Specific Emitter Identification; Open-set Recognition; RF Fingerprinting;
SNR-Adaptive Classification; Deep Learning for SEI

Yanping Liao, Xinyang Wang, Fan Jiang,


LPI radar waveform recognition based on semi-supervised model all mean teacher,
Digital Signal Processing,
Volume 151,
2024,
104568,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Low probability of intercept (LPI) radar signal identification plays an important role
in electronic warfare, but most existing algorithms are proposed under the condition of
sufficient samples, ignoring the problem of a small amount of labeled data in the actual
electromagnetic environment. To solve the problem, in this paper, a semi-supervised
learning model All Mean Teacher (AMT) based on Mean Teacher (MT) is proposed. First, the
LPI radar signal is transformed into Time-frequency images (TFIs) by using the Choi-Williams
distribution, and Random Erasing is used for TFIs which improves the generalization ability of
the model. Then the Multi-headed Self-Attention Network (MSA-Net) is aimed to extract
features, combined with AMT to realize the automatic waveform recognition of radar
signals. MSA-Net facilitates feature information propagation by computing contrast costs on
TFIs between the student and teacher networks. It solves the problem that TFIs are not easy
to train for small amounts of labeled data, improving the accuracy of signal recognition in
semi-supervised learning scenarios. Experimental results show that the average recognition
accuracy of the proposed method is up to 85.7% at a signal-to-noise ratio of -8 dB.
Keywords: LPI radar signals; Time-frequency analysis; Mean teacher; Self-attention
mechanism

Tiantian Wang, Nan Yan, Chaosan Yang, Zeliang An, Gongjing Zhang, Yuqing Xu,
Electromagnetic signal recognition using multimodal tri-branch semantic fusion network in
the UAV-assist integrated sensing and communication systems,
Digital Signal Processing,
Volume 171,
2026,
105820,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Driven by the proliferation of integrated sensing and communication (ISAC)
systems, the accurate recognition of unauthorized unmanned aerial vehicle (UAV) signals in
dynamic electromagnetic environments has emerged as a critical challenge for spectrum
security and cognitive radio applications. Conventional automatic modulation recognition
(AMR) frameworks suffer from significant performance degradation in low signal-to-noise
ratio (SNR) regimes and exhibit limited adaptability to resource-constrained edge computing
platforms. To address these limitations, we propose a novel Multimodal Tri-branch Fusion
Network (MTF-Net) architecture that synergistically integrates time-frequency analysis with
statistical feature learning. The framework systematically processes binarized time-
frequency images (B-TFIs) and higher-order cumulant vectors through three collaboratively
operating branches: (1) A primary temporal feature extractor employing dilated convolution-
residual blocks (DCRBlocks) with hierarchical dilatation factors, incorporating channel
attention mechanisms to dynamically emphasize discriminative temporal patterns; (2) Dual
auxiliary branches based on Edge-Transformer modules (ETFormers), which achieve efficient
spatial-structural learning through depthwise separable convolutions (DSC) while capturing
long-range spectral dependencies via additive attention mechanisms with linear complexity;
(3) A hierarchical fusion module implementing cross-branch feature recalibration through
learnable parameter matrices. Extensive Monte Carlo experiments demonstrate that our
MTF-Net significantly outperforms traditional methods in recognition accuracy for radar and
communication signals under low SNR conditions, establishing a new benchmark for
lightweight AMR solutions in ISAC systems.
Keywords: Multi-modal feature fusion; Unmanned aerial vehicle(UAV); Integrated sensing
and communication (ISAC); Lightweight neural network; Transformer

Hao Wan, Shuai Yang, Jixiong Xiao,


Clutter suppression for ground penetrating radar echo signal based on layer division
processing,
Journal of Applied Geophysics,
Volume 243,
2025,
105923,
ISSN 0926-9851,
[Link]
([Link]
Abstract: To improve the accuracy of underground target identification, clutter must be
efficiently suppressed in ground-penetrating radar(GPR) echo signal. Classic methods such as
Robust Principal Component Analysis (RPCA) and Factor Group-Sparse Regularization (FGSR)
have been widely applied in GPR signal processing. RPCA separates background and target
signals based on low-rank and sparse decomposition. FGSR removes large-scale surface
clutter using morphological operations. However, both methods face limitations under non-
uniform subsurface conditions, where the background and target signals are highly complex
and overlapping. Based on the characteristics of echo signal, a clutter suppression method is
proposed, namely adaptive layer division processing combined with two-dimensional
wavelet transform. A Peplinski's heterogeneous soil model containing underground targets is
constructed in gprMax to evaluate the effectiveness and applicability of the proposed
method. By analyzing the statistical properties of kurtosis and skewness, adaptive layer
division processing is applied to preliminarily separate the direct wave, background clutter,
and target echo reflection signals. The two-dimensional wavelet transform is then applied to
suppress clutter in the target signal layer, and the final image is reconstructed. Simulation
results show that adaptive layer division processing enhances the clutter suppression
performance of conventional denoising methods. The proposed method, integrating two-
dimensional wavelet transform, demonstrates superior clutter suppression performance,
where the signal-to-clutter ratio (SCR) is improved to 15.45 dB, the image entropy is reduced
to 2.51, the improvement factor (IF) achieves a positive value of 1.63 dB, and the peak
signal-to-noise ratio (PSNR) rises to 27.63 dB. The proposed method provides an effective
approach for processing GPR echo signals under non-uniform subsurface conditions.
Keywords: Ground penetrating radar; Clutter suppression; Layer division processing; Two-
dimensional wavelet transform; gprMax

Chengxin Yang, Benoit Champagne, Wei Yi,


Transmit beamforming design for area surveillance and multi-target tracking in colocated
MIMO radar,
Signal Processing,
Volume 243,
2026,
110491,
ISSN 0165-1684,
[Link]
([Link]
Abstract: This paper addresses the optimization problem of transmit beamforming design for
area surveillance and multi-target tracking (MTT) in a colocated multiple-input multiple-
output (C-MIMO) radar system. We first establish the relationship between the detection
probability and the predictive Cramér-Rao lower bound (PCRLB) as performance metrics,
and the transmit signal correlation matrix as the design variable. The surveillance area,
defined as a circular sector bounded by a polar angle and the intersecting arc, is divided into
independent smaller sectors, each corresponding to a different illumination direction of the
C-MIMO radar. To maximize the efficient utilization of power resources, we then aim to
maximize the number of simultaneously illuminated sectors while achieving desired
detection probability and target tracking accuracy. Given that the formulated optimization
problem is an intractable non-convex mixed-integer nonlinear problem, we propose a
beamforming algorithm based on Quality of Service (QoS) to solve it efficiently. Simulation
results indicate that the proposed algorithm is capable of effectively maximizing the
illuminated area while consistently meeting the specified detection probability and MTT
accuracy requirements.
Keywords: Transmit beamforming; Colocated MIMO radar; Area surveillance; Multi-target
tracking

Yanwen Han, Xiaopeng Yan, Jiawei Wang, Sheng Zheng, Hongrui Yu, Jian Dai,
A sparse moving array imaging approach for FMCW radar with dual-aperture adaptive
azimuth ambiguity suppression and adaptive QR decomposition,
Defence Technology,
Volume 50,
2025,
Pages 254-271,
ISSN 2214-9147,
[Link]
([Link]
Abstract: Range-azimuth imaging of ground targets via frequency-modulated continuous
wave (FMCW) radar is crucial for effective target detection. However, when the pitch of the
moving array constructed during motion exceeds the physical array aperture, azimuth
ambiguity occurs, making range-azimuth imaging on a moving platform challenging. To
address this issue, we theoretically analyze azimuth ambiguity generation in sparse motion
arrays and propose a dual-aperture adaptive processing (DAAP) method for suppressing
azimuth ambiguity. This method combines spatial multiple-input multiple-output (MIMO)
arrays with sparse motion arrays to achieve high-resolution range-azimuth imaging. In
addition, an adaptive QR decomposition denoising method for sparse array signals based on
iterative low-rank matrix approximation (LRMA) and regularized QR is proposed to
preprocess sparse motion array signals. Simulations and experiments show that on a two-
transmitter-four-receiver array, the signal-to-noise ratio (SNR) of the sparse motion array
signal after noise suppression via adaptive QR decomposition can exceed 0 dB, and the
azimuth ambiguity signal ratio (AASR) can be reduced to below −20 dB.
Keywords: Frequency modulated continuous wave (FMCW); Sparse motion array; Range-
azimuth imaging; Azimuth ambiguity suppression; DAAP; Adaptive QR decomposition

Shuai Guo, Ting Chen, Penghui Wang, Jun Ding, Junkun Yan, Hongwei Liu,
Knowledge embedding fusion based on language model for enhanced radar target
recognition,
Signal Processing,
Volume 238,
2026,
110199,
ISSN 0165-1684,
[Link]
([Link]
Abstract: Traditional radar target recognition methods typically model only single echoes,
neglecting the crucial information that domain knowledge can provide for understanding
data. In this paper, we propose a knowledge embedding fusion (KEF) method for enhanced
high-resolution range profile (HRRP) recognition, which utilizes the target state descriptions
available during radar detection. KEF leverages a language model (LM) to integrate textual
knowledge with echo features for fusion recognition. It consists of three components: HRRP
feature extraction, measurement-based knowledge construction, and knowledge embedding
fusion module. First, we perform feature extraction on the HRRP to obtain echo tokens.
Next, in the knowledge construction module, the measurement statuses are standardized to
a natural language format, and the LM is utilized to extract semantic information, resulting in
text tokens. Finally, in the knowledge embedding fusion module, a cross-attention HRRP-text
fusion strategy is employed to facilitate interaction between echo tokens and textual tokens.
We also design a combination of HRRP-text matching loss and fusion classification loss to
guide model training. Experiments are conducted on a real measured dataset, and the
results indicate that KEF effectively enhances recognition performance across multiple
scenarios compared with approaches that only utilize echoes.
Keywords: High-resolution range profile (HRRP); Knowledge embedding fusion; Language
model (LM); Radar target recognition

Wenxu Zhang, Fosheng Zhang, Zhongkai Zhao, Feiran Liu,


Radar specific emitter identification via the Attention-GRU model,
Digital Signal Processing,
Volume 142,
2023,
104198,
ISSN 1051-2004,
[Link]
([Link]
Abstract: In radar specific emitter identification (SEI), various types of unintentional
modulation on pulse (UMOP) are selected as the features for discriminating between
different radars. Unintentional Phase Modulation on Pulse (UPMOP), a typical type of UMOP,
can provide crucial information for identifying radars. In most radar SEI algorithms,
sacrificing time efficiency for higher accuracy is a common trade-off. This paper proposes a
method to solve this problem by combining denoised UPMOP sequences with an Attention-
based Gated Recurrent Units (Attention-GRU) model, which showed an excellent
performance. Firstly, the cause of UPMOP is analyzed and the phase observation model of
radar emitter signals and mathematical model of UPMOP are given. Then, the least-squares
method is used to eliminate the linear trend of the phase observation model and obtain a
noised estimation of the UPMOP sequences. Thirdly, the uniform B-spline (UBS) curves are
then used to fit the noised estimation, resulting in a denoised and refined UPMOP sequence.
Finally, the Attention-GRU model is employed to extract features from the denoised UPMOP
sequences to identify radar emitters automatically. Results from simulation and measured
data experiments show that the overall recognition rate of the algorithm reaches over 93%
and the algorithm has excellent performance, with high identification accuracy and relatively
low time consumption, even in low signal-to-noise ratio (SNR) conditions.
Keywords: Radar specific emitter identification; Unintentional phase modulation on pulse;
Uniform B-spline curve; Time-series model; Attention mechanism

Xiaosong Tang, Feng Yang, Xu Qiao, Jialin Liu, Haitao Zuo, Liang Gao, Jianshe Zhao, Suping
Peng,
GPR-HIDiff: A diffusion-based model for horizontal interference suppression in urban
underground detection radar profiles,
Underground Space,
Volume 26,
2026,
Pages 458-478,
ISSN 2467-9674,
[Link]
([Link]
Abstract: Automated subsurface utility detection systems in construction rely heavily on the
quality of ground-penetrating radar (GPR) profiles, which are often degraded by high-
amplitude horizontal interference. Existing low-rank decomposition methods lack the
intelligence and flexibility required for multi-site data processing and involve labor-intensive
parameter tuning, impeding their integration into intelligent construction workflows. To
address these challenges, this paper proposes a horizontal interference suppression
algorithm based on a diffusion model, termed GPR-HIDiff. The proposed model replaces
conventional sequential convolutional operators with ResBlocks throughout the encoder,
intermediate layer, and decoder of the UNet architecture, enhancing training stability.
Lightweight agent attention modules are embedded between ResBlocks at each level to
improve global information modeling capability. A spatial attention mechanism is deployed
between the encoder and decoder to achieve adaptive spatial feature optimization.
Furthermore, the forward diffusion phase adopts a cosθ schedule-based strategy to ensure a
smooth temporal variation of noise variance. A standardized dataset comprising real-world
measured samples and finite difference time domain simulation samples of urban road
models has also been constructed. The effectiveness of the hybrid dataset, the introduced
modules, the robustness analysis, and the cosθ schedule is validated through training with
single/mixed datasets, ablation studies, evaluation of metric variations before and after the
introduction of different noise levels, and comparative experiments with constant, linear,
and cosθ schedules. Experimental results demonstrate that GPR-HIDiff significantly
outperforms both traditional methods and state-of-the-art deep learning models on both
simulated and real-world test samples. It effectively suppresses horizontal artifacts,
preserves target hyperbolic contours, and avoids excessive reduction of target scattering,
showcasing its exceptional performance. This method provides a powerful algorithmic
foundation for high-resolution GPR imaging and target detection.
Keywords: Ground-penetrating radar; Horizontal interference; Diffusion model; Agent
attention module; Spatial attention; Hybrid dataset

Zhiyan Lin, Minming Gu, Keyu Pan, Wei-Ping Zhu,


Adaptive temporal convolutional network with multi-head EMA-gated attention for
continuous radar-based human activity recognition,
Biomedical Signal Processing and Control,
Volume 117,
2026,
109667,
ISSN 1746-8094,
[Link]
([Link]
Abstract: Continuous human activity recognition (HAR) using radar signals offers strong
potential for privacy-preserving clinical health monitoring. However, its performance is
limited by challenges such as multi-scale temporal variation, signal noise, and unstable
activity transitions. To address these issues, this study introduces a radar-based HAR
framework with three tailored components. First, an adaptive temporal convolutional
network (ATCN) uses learnable dilation rates and sampling offsets to flexibly capture both
abrupt and periodic motion patterns over time. Second, an exponential moving average
(EMA)-gated attention (EDGA) module integrates linear attention with exponential moving
average smoothing through a dynamic gating mechanism, effectively suppressing noise
while preserving temporal continuity. Third, an attention-guided multi-stage refinement
(AMSR) module refines coarse predictions using global attention-driven residual corrections,
thereby reducing segmentation noise and improving boundary precision. Experiments on a
77 GHz frequency modulated continuous wave (FMCW) radar dataset show that the
proposed model achieves 96.09% accuracy, demonstrating its strong potential for
continuous and unobtrusive activity monitoring in healthcare applications.
Keywords: ATCN; EDGA; AMSR; Continuous HAR; FMCW radar

Yalan Wang, Weiwei Liu, Yibei Wang, Xiaoniu Peng, Zefeng Liu, Changping Wang, Anle Wang,
Multi-format frequency-hopping RF signals generation based on the soliton optoelectronic
oscillator incorporating FDML mechanism,
Optics & Laser Technology,
Volume 194,
2026,
114448,
ISSN 0030-3992,
[Link]
([Link]
Abstract: Emerging optoelectronic oscillators (OEO) are exploring ways to expand their signal
generation formats while retaining excellent phase noise performance to meet the demand
for high-quality signals in information systems such as radar and electronic warfare. Here, a
soliton OEO with reconfigurable operating states is proposed and experimentally
demonstrated for generating multi-format frequency-hopping radio-frequency (RF) signals.
By introducing a composite Fourier-domain mode-locking (FDML) mechanism into the
soliton oscillation state, the proposed scheme can extend the generated signal format from
single-frequency hopping to other hopping patterns −without any structural modifications-
simply by adjusting relative loop parameters. Dual filtering mechanisms in the OEO loop are
realized based on a phase-shift fiber Bragg grating (PS-FBG) cavity and the stimulated
Brillouin scattering (SBS) effect, respectively. Experiments verify the excellent functional
reconfiguration capability of the proposed OEO architecture, generating the aforementioned
signals with adjustable time-domain or frequency-domain parameters. Furthermore, when
the constructed OEO degenerates into a conventional OEO, it can generate different types of
dual continuous-wave signals, further demonstrating its potential as a high-quality arbitrary
waveform generator. Thus, the proposed method not only expands the scope of soliton OEO
but also the application field of the OEO technology.
Keywords: Soliton optoelectronic oscillator; Frequency-hopping; Multi-format; Stimulated
Brillouin scattering; Fourier-domain mode locking

Wentai Lei, Shixuan Yu, Tao Zhang,


HFL-YOLOv8: a hyperbolic feature-enhanced lightweight network for object detection in
ground penetrating radar images,
Applied Soft Computing,
Volume 188,
2026,
114403,
ISSN 1568-4946,
[Link]
([Link]
Abstract: Deep learning-based soft computing object detection algorithms for Ground
Penetrating Radar (GPR) B-Scan images face two significant challenges: Existing methods
struggle to account for the unique geometric characteristics of hyperbolas. This limitation
hampers the extraction of correlation information, particularly when detecting smaller
objects; The high resource consumption of these algorithms increases hardware
requirements, making them less efficient for practical deployment. To address these issues,
this paper proposes HFL-YOLOv8, a hyperbolic feature-enhanced lightweight object
detection network based on YOLOv8. The key contributions of HFL-YOLOv8 include: Three
convolution layers with varying dilation rates are employed to capture feature information
across different scales, improving the network’s ability to handle diverse object sizes. A
dynamic upsampling operator and a channel-position attention module are incorporated to
refine the detection of hyperbolic features, addressing the limitations in geometric
representation. A detection head with shared parameter convolution is used to reduce
computational overhead. Additionally, enhanced Sobel convolution and group normalization
convolution compensate for accuracy loss, ensuring robust detection. HFL-YOLOv8 achieves
an F1-score of 79.7 % and mAP50 of 81.8 %, representing significant performance
improvements. Moreover, the proposed network reduces the number of parameters by 14.8
% and computational costs by 13.6 %, offering enhanced accuracy and resource efficiency for
detecting small object bisectors in GPR B-Scan images.
Keywords: Ground penetrating radar; Radar signal processing; Deep learning; Object
detection; Lightweight

Karthik S., Rani Thottungal, S.A. Pasupathy,


Spatial Aware Feature Extraction for RADAR SLAM,
Measurement,
Volume 264,
2026,
120192,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Strong reflections from RADAR are often part of clutters and need reliable filtering
for before feature extraction. Traditonal intensity based filter result in clustered features. To
address this a Spatial Aware Feature Extraction (SAFE) method is proposed for radio
detection and ranging (RADAR) for Simultaneous Localization and Mapping (SLAM). The
objective is to improve odometry estimation as part of SLAM pipeline by ensuring a
balanced spatial distribution of feature extraction across the RADAR range. The proposed
method processes one-dimensional RADAR signals, representing returns per azimuth angle.
It extracts the strongest features based on signal strength and filters them using Euclidean
distance, ensuring a distributed spread across the RADAR range. This selection reduces
feature clustering and improves accuracy by spreading features across the environment,
addressing a limitation encountered when the number of features per signal is restricted. A
qualitative evaluation on the Oxford RADAR robocar dataset demonstrates the effect of
asserted Euclidean distance between features on the accuracy of the SLAM system. The
RADAR SLAM with SAFE achieved a mean translational error of 3.31% and a rotational error
of 0.82 degrees per meter on Oxford RADAR Robocar dataset.
Keywords: RADAR odometry; Feature extraction; Euclidean distance-based filtering; Filtering
techniques

Tingpei Huang, Rongyu Gao, Haotian Wang, Jianhang Liu, Shibao Li,
mBox: 3D object detection based on millimeter-wave radar,
Measurement,
Volume 246,
2025,
116568,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Millimeter-wave radar is utilized for 3D object detection in autonomous driving
due to its advantage of not being affected by lighting conditions. Previous millimeter-wave
radar object detection algorithms have not sufficiently used the distributional and statistical
properties of sparse point clouds. This paper introduces mBox, a 3D object detection
framework using only millimeter-wave radar. To eliminate unnecessary information, we
propose a background filtering algorithm based on subtraction(BGFS) that matches and
differentiates the point cloud frame by frame. We propose a voting-based algorithm for
generating object centers(CPGV), which expands the valid data to generate more accurate
initial anchor boxes. To utilize feature information across different scales and capture the
structure of each granularity, we propose a multi-scale feature fusion network based on the
attention mechanism(GLFF-Net). We conduct experiments using the Pointillism and Astyx
datasets. The results show that the mBox outperforms the comparative methods in terms of
3D mean average precision(mAP).
Keywords: Object detection; Point cloud; Millimeter-wave radar

Youwei MENG, Yaoyao LI, Shaoxiong CAI, Donglin SU,


A dynamic spectrum and power allocation method for co-located pulse radar and
communication system coexistence,
Chinese Journal of Aeronautics,
Volume 38, Issue 4,
2025,
103417,
ISSN 1000-9361,
[Link]
([Link]
Abstract: Airborne pulse radar and communication systems are essential for precise
detection and collision avoidance, ensuring that aircraft operate safely and efficiently. A
major challenge in spectrum sharing is the allocation of resources in both the time and
frequency domains, aiming to minimize inter-system interference as the available spectrum
fluctuates over time. In this paper, regarding maximization of detection probability and
spectrum utilization efficiency as two fundamental objectives, a novel Dynamic Spectrum
and Power Allocation based on Genetic Algorithm (GA-DSPA) model is proposed, which
dynamically allocates communication channel frequency and power under the constraints of
pulse radar detection probability and signal-to-interference-plus-noise ratio of
communication. To solve this bi-objective model, a non-dominated sorting-based multi-
objective genetic algorithm is developed. A novel environment perception strategy and
offspring sorting technique based on radar echoes are integrated into the optimization
framework. Simulation results indicate that by integrating environmental monitoring
mechanisms and dynamic adaptation strategies, the proposed method effectively tracks the
evolving Pareto-optimal Fronts (PoFs), thereby ensuring optimal performance for both co-
located pulse radar and communication systems. Hardware test results confirm that within
the GA-DSPA framework, the pulse radar achieves higher detection probabilities under
identical conditions, while the communication system realizes increased average
throughput.
Keywords: Communication systems; Dynamic multi-objective optimization; Electromagnetic
compatibility; Radar-communication coexistence; Spectrum and power allocation

Ge Junkai, Sun Huaifeng, Shao Wei, Liu Dong, Yao Yuhong, Zhang Yi, Liu Rui, Liu Shangbin,
GPR-TransUNet: An improved TransUNet based on self-attention mechanism for ground
penetrating radar inversion,
Journal of Applied Geophysics,
Volume 222,
2024,
105333,
ISSN 0926-9851,
[Link]
([Link]
Abstract: Convolutional Neural Networks (CNN) are widely applied to Ground Penetrating
Radar (GPR) inversion because they have strong data-driven capabilities and are suitable for
the data structure form of GPR. For CNN, the computation increases with the distance that
the convolutional block moves from one region to another when it calculates the
relationship between two regions. For GPR data, the target reflection exists in the
surrounding traces and full time-window of the target, which leads to high degree of remote
relationship. In this paper, we propose GPR-TransUNet, a deep-learning based inversion
network which use self-attention mechanism. According to the characteristics of GPR data,
regression network and GPR-Loss mechanism were used. Both numerical and model
experiments were arranged to test the performance of the network, and the result as well as
comparative analysis demonstrate the superiority of GPR-TransUNet. Finally, we applied this
method to the field GPR data of Guangxi as an attempt.
Keywords: GPR; Inversion; Deep learning

Wenxu Zhang, Xian Lei, Zhongkai Zhao, Fuli Sun,


A dual-decision-maker frequency domain cooperative jamming method against multi-
function radar based on PPO,
Digital Signal Processing,
Volume 169,
2026,
105709,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Optimizing jamming strategies is crucial for coping with complex electromagnetic
countermeasures in dynamic spectrum environments, Among them, the frequency agility
characteristic of the multi-function radar enables them to exhibit strong anti-jamming
capabilities by quickly changing the carrier frequency. Aiming at the problem that traditional
electronic countermeasure strategies show insufficient adaptability to this, a dual-decision-
maker collaborative jamming method based on the proximal policy optimization (PPO)
framework is proposed in this paper. The method achieves dynamic adaptive jamming
against radars through a cascaded collaborative mechanism involving the frequency band
decision maker and the bandwidth decision maker. The confrontation scenario between the
radar network and multiple jammers is established, and the penetration confrontation
process is abstracted and modeled as a markov decision process. The simulation results
show that compared with classic reinforcement learning algorithms such as deep Q network,
the proposed dual-decision-maker collaborative jamming method based on PPO exhibits
superior performance in action estimation and jamming success rate. Specifically, the
convergence speed of the average jamming gain is improved by more than 65 % compared
with the sub-optimal algorithm, while its final stable reward is approximately 10 % higher.
Meanwhile, the key performance indicator values fluctuate minimally under different
experimental conditions, showcasing good robustness and scalability, and achieving
intelligent optimization of frequency domain decision-making.
Keywords: Frequency agility; Deep reinforcement learning; Intelligent jamming decision;
Proximal policy optimization

Chaofeng Huang, Xiaowo Xu, Fan Fan, Shunjun Wei, Xiaoling Zhang, Dongmei Liu, Min Gu,
A low-SNR-adaptive temporal network with smart mask attention for radar signal
modulation recognition,
Digital Signal Processing,
Volume 168, Part D,
2026,
105640,
ISSN 1051-2004,
[Link]
([Link]
Abstract: The automatic modulation recognition of radar signals is a key technology in
electronic warfare and communication systems. However, traditional handcrafted features
often struggle to achieve high recognition accuracy under low signal-to-noise ratio (SNR)
conditions. With the rapid development of artificial intelligence technologies, deep learning-
based approaches have emerged as a promising alternative for modulation recognition. In
this article, a low-SNR-adaptive network architecture is proposed, which integrates a
bidirectional temporal convolutional network (Bi-TCN) and dual-channel smart mask
attention (DSMA) modules. The DSMA adaptively highlights informative features and
suppresses noise through complementary attention masks, enhancing robustness in low-SNR
conditions. Experimental results demonstrate that the autocorrelation domain outperforms
both time and frequency domains, with recognition accuracy improvements of 13.33 % and
14.71 %, respectively. Compared to state-of-the-art models, the proposed network achieves
63 % accuracy at -20 dB and more than 99 % accuracy at -6 dB, significantly enhancing radar
signal modulation recognition.
Keywords: Modulation recognition; Deep learning; Radar signal analysis,

Keyu Pan, Wei-Ping Zhu, Bo Shi,


A multi-stage few-shot framework for extensible radar-based human activity recognition,
Signal Processing,
Volume 239,
2026,
110244,
ISSN 0165-1684,
[Link]
([Link]
Abstract: This paper proposes a novel framework for radar-based indoor human activity
recognition (HAR) using a multi-stage few-shot learning (FSL) paradigm. The core of our
approach lies in the design of a dynamic feature extraction architecture that exploits wavelet
convolution along with depthwise separable convolutions to effectively capture multi-scale
and multi-frequency information from radar signals. We also propose a meta-learning-
inspired mechanism that dynamically adjusts class weights for unseen categories, thereby
enhancing adaptability and recognition accuracy in few-shot scenarios. Extensive
experiments on five benchmark datasets demonstrate consistent performance gains over
state-of-the-art methods, with substantial improvements observed for both seen and
unseen classes. These findings highlight the robustness, scalability, and generalization
capability of our framework, underscoring its potential to advance radar-based HAR in
complex and diverse environments.
Keywords: HAR; FSL; WTConv; Meta-learning; FMCW radar

Teng Li, Liwen Zhang, Youcheng Zhang, Qingmin Liao,


AXFL: Axial prior-guided cross-view fusion learning for radar semantic segmentation,
Expert Systems with Applications,
Volume 303,
2026,
130552,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Multiple 2D View spectrogram-based Radar Semantic Segmentation (MVRSS)
simultaneously leverages different radar-view spectrograms to capture comprehensive
spatial and velocity information of targets. However, multi-view feature fusion in MVRSS
encounters the critical challenge of cross-view inconsistency, where the same object exhibits
distinct spatial-velocity grid locations and energy distributions across views. Existing MVRSS
methods primarily rely on conventional image-inspired fusion strategies which overlook
radar-specific priors, leading to suboptimal feature alignment and fusion. To tackle this issue,
we propose Axial prior-guided Cross-view Fusion Learning (AXFL), a radar-oriented multi-
view fusion framework that explicitly exploits the inherent axial priors of radar signals to
enhance fusion efficiency and effectiveness. Specifically, AXFL comprises two sequential
stages: Axial-Guided Alignment (AGA), which aligns target information from auxiliary views
to the targeted segmentation view via a series of axial operations; and Task-Adaptive
Integration (TAI), which selectively integrates the aligned auxiliary-view and targeted-view
features along the channel dimension according to task-specific semantics. Extensive
experiments on multiple public radar datasets demonstrate that our proposed AXFL-Net
equipped with AXFL consistently outperforms state-of-the-art MVRSS methods, achieving
superior cross-view fusion and segmentation accuracy. The source code will be available at
[Link]
Keywords: Radar semantic segmentation; Deep learning; Autonomous driving; Feature
fusion; Radar prior

Runze Hu, Tong Wang, Weijun Huang, Weichen Cui,


A clutter suppression algorithm via holistic attention deep neural network for airborne radar,
Digital Signal Processing,
Volume 168, Part B,
2026,
105506,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Space-time adaptive processing (STAP) methods have favorable clutter suppression
performance under the condition of sufficient independent and identically distributed (i.i.d.)
samples in airborne radar systems. However, in practical situations, airborne radar often
encounters heterogeneous and non-stationary clutter environment, resulting in a shortage
of i.i.d. samples and degradation in the performance of STAP methods. In this article, a novel
STAP method based on holistic attention network is proposed to address these problems. To
start with, based on airborne radar clutter model, the clutter snapshot data considering ideal
and non-ideal cases are simulated for network training. In addition, a network based on
holistic attention mechanism is developed to perform the super-resolution of clutter spatial-
Doppler spectra. Low-resolution clutter spatial-Doppler spectrum inputs (which are
calculated by only a few snapshot data) are converted to high-resolution clutter spatial-
Doppler spectrum outputs by trained holistic attention network (HAN). As a final point, with
high-resolution clutter spectra, clutter plus noise covariance matrices (CNCMs) are obtained
to calculate adaptive weight vectors. This method achieves recovering clutter spatial-
Doppler spectra precisely nearly in real time with only a few samples under ideal and non-
ideal conditions due to the proposed network's exceptional capability for feature extraction
and pattern recognition. In comparison with recent sparse recovery based and convolutional
neural network (CNN) based STAP, the proposed algorithm shows better performance in
both ideal and non-ideal situations in a short execution time. Numerous simulations are
conducted to demonstrate the effectiveness and superiority of the proposed method in
convergence rate, clutter suppression performance and computational efficiency.
Keywords: Clutter suppression; Space-time adaptive processing; Airborne radar; Radar signal
processing; Convolutional neural network

Hongwei Ma, Yi Liao, Chunhui Ren,


Low probability of interception radar overlapping signal modulation recognition based on an
improved you-only-look-once version 8 network,
Engineering Applications of Artificial Intelligence,
Volume 137, Part A,
2024,
109150,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Low probability of interception (LPI) radar is widely used in modern electronic
warfare. With the increased radiation sources, multiple signals will arrive simultaneously.
The traditional feature extraction method has too many features, which brings great trouble
to the subsequent data processing. Most modulation recognition methods based on deep
learning only consider the single signal after preprocessing, and the generalization ability is
weak. This paper proposes a deep learning solution based on an improved you only look
once version 8 (YOLOv8) network with a global attention mechanism (GAM), achieving
recognition accuracy over 98% in a -10 dB signal-to-noise ratio (SNR) scenario to address the
above problems. We improved the performance of the original network, which can be used
in electronic countermeasures.
Keywords: Global attention mechanism; Deep learning; You only look once version 8;
Modulation recognition; Low probability of interception radar aliasing signal

Changlong Wang, Jiawei Jiang, Chong Han, Hengyi Ren, Lijuan Sun, Jian Guo,
Through-Wall Multihuman Activity Recognition Based on MIMO Radar,
Computers, Materials and Continua,
Volume 83, Issue 3,
2025,
Pages 4537-4550,
ISSN 1546-2218,
[Link]
([Link]
Abstract: Existing through-wall human activity recognition methods often rely on Doppler
information or reflective signal characteristics of the human body. However, static
individuals, lacking prominent motion features, do not generate Doppler information.
Moreover, radar signals experience significant attenuation due to absorption and scattering
effects as they penetrate walls, limiting recognition performance. To address these
challenges, this study proposes a novel through-wall human activity recognition method
based on MIMO radar. Utilizing a MIMO radar operating at 1–2 GHz, we capture activity data
of individuals through walls and process it into range-angle maps to represent activity
features. To tackle the issue of minimal variation in reflection areas caused by static
individuals, a multi-scale activity feature extraction module is designed, capable of extracting
effective features from radar signals across multiple scales. Simultaneously, a temporal
attention mechanism is employed to extract keyframe information from sequential signals,
focusing on critical moments of activity. Furthermore, this study introduces an activity
recognition network based on a Deformable Transformer, which efficiently extracts both
global and local features from radar signals, delivering precise human posture and activity
sequences. In experimental scenarios involving 24 cm-thick brick walls, the proposed
method achieves an impressive 97.1% accuracy in activity recognition classification.
Keywords: MIMO radar; human activity; Transformer; through-wall

Rui Li, Liulin Li, Hongping Zhou, Zhongyi Guo,


OAM-radar imaging: A review,
Defence Technology,
2025,
,
ISSN 2214-9147,
[Link]
([Link]
Abstract: Due to the spiral phase wavefront of vortex electromagnetic waves (VEMW), there
will exist the orbital angular momentum (OAM) in the VEMW, which supplies additional
spatial dimensional for the information process. In radar imaging, when the VEMW is used
to illuminate targets, it can be regarded as the superposition of multiple plane waves
incident on targets from various angles simultaneously. The target within the beam is
radiated by electromagnetic (EM) fields with different phase information. These differences
make the scattered echoes contain more target information. Providing a feasible technical
path to improve the traditional radar imaging technology. This paper introduces the
challenges faced by traditional radar imaging and the advantages of EM vortex wave
imaging. It briefly derives the principles of vortex radar imaging and azimuthal resolution
based on circular array antennas. Then, it reviews the development history and research
status of vortex radar imaging technology, which divides into six research directions: vortex
radar imaging model, Bessel function modulation, vortex radar gazing algorithm, EM vortex
synthetic aperture radar (SAR)/inverse synthetic aperture radar (ISAR) imaging, vortex radar
3D imaging, and vortex radar forward imaging. Finally, the future development trends of
vortex radar imaging technology are outlined and discussed.
Keywords: Vortex electromagnetic waves (VEMW); Orbital angular momentum (OAM);
Vortex radar imaging; Bessel function modulation

Muhammad Fahad Munir, Abdul Basit, Wasim Khan, Ahmed Saleem, Aleem Khaliq, Nauman
Anwar Baig,
Next-Gen solutions: Deep learning-enhanced design of joint cognitive radar and
communication systems for noisy channel environments,
Computers and Electrical Engineering,
Volume 120, Part A,
2024,
109663,
ISSN 0045-7906,
[Link]
([Link]
Abstract: In recent years, the dual-function radar and communication (DFRC) paradigm has
emerged as a focal point in addressing spectrum congestion challenges. However, prevailing
research heavily relies on computationally complex likelihood-based approaches for
communication signals with an added Gaussian noise based single waveform. Note that, a
single waveform for diverse scenarios e.g., presence of a communication receiver in the
radar main lobe, side lobe, etc., may lead to a deteriorated detection performance in a DFRC
design. Therefore, in this paper, we present a cognitive DFRC architecture that utilizes a
diverse set of orthogonal waveforms at the transmitter. Specifically, based on a perception-
action cycle, a QAM-based waveform is employed for communication when both the radar
target and communication receiver are within the main lobe, while a PSK-based waveform is
used when the radar target is in the main lobe and the communication receiver is in the side
lobes. Furthermore, to enhance the feature-based estimation, the communication receiver
integrates a Convolutional Neural Network (CNN) architecture designed to autonomously
learn and extract features from received signals with different Signal-to-Noise ratio (SNR).
Next, the adaptive nature of the system enables proficient discernment of the received
signal type and its corresponding SNR value. Moreover, deep learning techniques are applied
in realistic scenarios with various channel impairments to extract features from received
signals, departing significantly from likelihood-based methods and reducing computational
complexity. The proposed methodology’s effectiveness is validated through Monte Carlo
simulations, underscoring its potential to address challenges associated with DFRC under
real-world conditions.
Keywords: CNN based DFRC; DFRC modulation classification; Channel estimation by deep
learning; DFRC spectrum efficiency optimization; DFRC cognitive architecture

Jinyang Xie, Kanghui Zhou, Lei Han, Liang Guan, Maoyu Wang, Yongguang Zheng, Hongjin
Chen, Jiaqi Mao,
Enhancing multi-task learning-based Tornado identification using spatial and temporal
information from weather radar images,
Applied Soft Computing,
Volume 184, Part B,
2025,
113834,
ISSN 1568-4946,
[Link]
([Link]
Abstract: Tornadoes, as dynamic weather phenomena, exhibit unique spatial and temporal
evolution characteristics that reflect their formation and development. Existing tornado
detection algorithms often struggle with high false alarm rates, primarily due to insufficient
capture of temporal correlations in tornado development. As an improvement, we propose a
multi-task tornado identification network with three-dimensional temporal and spatial
information (TS-MTINet). Taking continuous three-frame radar data as input, the Multi-
frame Temporal Interaction Block (MTIB) utilizes multi-head attention to model the dynamic
interaction information between the radar data, thus exploring in-depth the temporal
features during tornado development. Further, we design a Spatial-Temporal Enhancement
Module (STEM), which analyzes the difference information between continuous data to
extract local and global spatial and temporal feature variations about tornadoes. Based on
this architecture, TS-MTINet incorporates a multi-task learning framework to perform
tornado detection and number estimation tasks simultaneously, thus extracting
comprehensive information related to tornadoes. To validate the performance of the
proposed model, we construct the first Chinese tornado identification dataset with fine
radar features. The experimental results show that the proposed method shows significant
advantages in several evaluation metrics, especially in reducing false alarms. In practical case
studies, compared to the traditional TVS method, TS-MTINet achieves an increase in POD of
approximately 30% and a decrease in FAR of about 20% in several typical tornado events.
Particularly in environments with strong interference, TS-MTINet demonstrates higher
detection accuracy, reflecting greater robustness and practical value.
Keywords: Deep learning; Multi-task learning; Tornado identification; Weather radar;
Attention mechanisms

Xuning Wang, Yuan Chen, Fuhao Wang, Kai Zheng, Yi Zong,


mPCT-LSTM: A lightweight human activity recognition model for 3D point clouds in
millimeter-wave radar,
Digital Signal Processing,
Volume 164,
2025,
105263,
ISSN 1051-2004,
[Link]
([Link]
Abstract: The millimeter-wave radar-based human activity recognition technology shows
significant potential in various domains, such as health monitoring, sports analysis, and
smart homes. Traditional methods rely on 2D feature spectrograms, which struggle to
balance low computational complexity with high recognition accuracy. Transforming radar
echoes into 3D point clouds and designing appropriate neural network models can mitigate
this conflict. However, most existing point cloud processing models are intended for dense
point clouds generated by LiDAR or depth sensors, leaving a gap in effective algorithms for
sparse mmWave radar point clouds. To address this need, we propose a lightweight model
for human activity recognition using sparse mmWave radar point clouds: mPCT-LSTM. This
model extracts spatial features from point clouds of human activity through a Point Cloud
Transformer (PCT) module, which consists of an embedding layer and four stacked offset-
attention layers. These spatial features are fed into a Long Short-Term Memory (LSTM)
module to capture temporal relationships between point cloud frames. Experimental results
demonstrate that the mPCT-LSTM model achieves an average recognition accuracy of
97.26% across three public datasets, outperforming the state-of-the-art by 1.41%.
Additionally, the model’s computational complexity is only 0.09 GFLOPS, a reduction of 70%
compared to current solutions.
Keywords: Millimeter-wave radar; Human activity recognition; 3D point cloud; Deep
learning; Long short-term memory (LSTM)

Huihui Ma, Haihong Tao, Yaxing Yue, Tiantian Zhong, Yunfei Fang, Le Wang,
Multiparameter Estimation for Bistatic EMVS-FDA-MIMO Radar with Arbitrarily Configured
Arrays,
Digital Signal Processing,
2026,
105928,
ISSN 1051-2004,
[Link]
([Link]
Abstract: This study explores the multiparameter estimation challenge within bistatic
frequency diverse array multiple-input-multiple-output (FDA-MIMO) radar system that
employs arbitrarily configured electromagnetic vector sensor (EMVS) arrays. The signal
reception model for the presented radar architecture is established. Building on this
foundation, a subspace-based algorithm is proposed to achieve accurate estimation of
spatial-polarization angles and ranges. First, rotation invariant structures in spatial domain
are formed by constructing several virtual steering matrices, from which the normalized
electromagnetic field vectors are derived. Then the two-dimensional direction-of-departure
(2D-DOD) and two-dimensional direction-of-arrival (2D-DOA) estimates are computed
through vector cross-product operation. Thereafter, polarization angles are determined
using least squares (LS) approach. Finally, by compensating the steering matrix with the
obtained 2D-DOD, the range estimation can be achieved. Furthermore, the developed
framework is evaluated for its identifiability, flexibility, computational demands, and Cramér-
Rao bound (CRB). It successfully estimates the targets’ spatial-polarization angles and
ranges, while also achieving automatic parameters pairing. Simulation results demonstrate
the validity of the developed approach.
Keywords: spatial-polarization angles estimation; range estimation; bistatic EMVS-FDA-
MIMO radar; arbitrary array

Xu Meng, Zhaogang Huang, Xin Deng, Hai Liu, Hongyuan Fang, Chao Liu, Xiaoyu Zhang, Jie
Cui,
Leakage detection and localization of buried water pipe using ground penetrating radar,
Measurement,
Volume 254,
2025,
117902,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Leakage in water distribution networks is a critical issue in urban cities worldwide.
Various non-destructive testing methods have been utilized to detect and localize water
leakages in buried pipes. Among these methods, ground penetrating radar (GPR) draws
attention to its advantages of fine resolution, high efficiency, and good portability. Previous
research has shown that water pipe leaks cause noticeable changes in GPR profiles.
However, most studies focus on comparing reflection patterns before and after water leaks,
typically through controlled experiments. The propagation mechanisms of electromagnetic
waves reflected from a leaky pipe and its surrounding media have not been investigated,
hindering direct extraction of leaked information in the GPR data. This paper proposes a co-
simulation algorithm to study electromagnetic wave propagation as water leaks appear and
evolve. Two laboratory experiments were conducted to validate the effectiveness of the co-
simulations. Results indicate that the wetting zone initially spread around the leak point and
finally presented a strawberry shape under the action of gravity. When the leaking water
reaches a certain volume, oscillating signals, composed of reflections from the pipe, the
heterogeneous wetting zone, and the creeping wave, occur in the GPR images. Based on the
findings, a wavelet-entropy-based method is proposed to localize the leak position from 3D
GPR data. Laboratory and field case results demonstrate that the proposed method is
effective for leak detection and possesses a fantastic application prospect.
Keywords: Non-destructive testing (NDT); Ground penetrating radar (GPR); Water pipe
leakage; Combined simulation; Leak localization

Yijiang Chen, Guohua Wei, Jiahao Bai, Xu Wang,


An efficient coherent integration method for RFPA radar signals in high-speed and multi-
target scenarios,
Digital Signal Processing,
Volume 171,
2026,
105838,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Although random frequency and pulse repetition interval (PRI) agile (RFPA) radar
signals provide strong electronic counter-countermeasures (ECCM) capabilities, coherent
integration under high-speed and multi-target scenarios remains a critical but insufficiently
explored challenge. We propose an efficient coherent integration method termed RFPA-KT-
NUFFT, which combines Keystone transform (KT) and nonuniform fast Fourier transform
(NUFFT). The truncated-sinc-interpolation-based KT exploits low computational complexity
to correct range migration and decouple the interdependencies between frequency and PRI.
Next, a tailored phase compensation function is developed to harness the frequency-agile
bandwidth, thereby enabling high range resolution and facilitating multi-target separation.
Finally, NUFFT efficiently handles the nonuniform slow-time sampling for coherent
integration. Monte Carlo experiments with a 500-sequence agile waveform library
demonstrate that the proposed method significantly reduces runtime while maintaining
integration performance comparable to existing optimal methods, offering a viable solution
that achieves a favorable trade-off between performance and efficiency in high-speed and
multi-target scenarios.
Keywords: Random frequency and pulse repetition interval agile; Coherent integration;
Keystone transform; Non-uniform fast fourier transform; Multi-target

Pengfei Wang, Peilin Shu, MingHao Yang, Hongqiu Zhang, Jianqi Wang, Cong Wang, Hongbo
Jia,
Dual-task physiological learning for radar-based continuous blood pressure monitoring:
classification-regularized regression,
Biomedical Signal Processing and Control,
Volume 112, Part C,
2026,
108790,
ISSN 1746-8094,
[Link]
([Link]
Abstract: Continuous blood pressure (BP) monitoring is critical for hypertension
management, yet conventional non-contact radar systems face challenges such as feature
space overlap under low signal-to-noise ratio conditions and inaccurate estimation during
physiological state transitions due to the susceptibility of millimeter-level cardiovascular
vibration signals to environmental interference. To address these challenges, we propose a
dual-task learning model integrating classification-constrained regression with multi-scale
spatiotemporal feature extraction. Our framework combines: (1) A hybrid ResNet-BiGRU
backbone capturing waveform morphology and hemodynamic continuity through multi-
scale convolutions (kernels = 15/7/3) and triple-layer bidirectional gating, enabling robust 2-
second-interval predictions; (2) A physiological regularization mechanism where
classification-derived BP-range probabilities (10 mmHg bins) constrain regression outputs,
suppressing implausible fluctuations during dynamic states. Validated on 30 subjects across
resting, Valsalva, and tilt-table tests, results indicate clinically relevant accuracy (SBP:
−0.21 ± 6.74 mmHg; DBP: 0.25 ± 4.81 mmHg) at 0.5 Hz sampling rate, while demonstrating
improved dynamic-state performance versus benchmarks with DBP RMSE reductions up to
8.4 % by dual-task strategy. This work suggests dual-task learning can mitigate radar-specific
SNR constraints and physiological nonstationarity while fulfilling clinical real-time monitoring
demands (beat-to-beat resolution), contributing to practical deployment of cuffless BP
devices.
Keywords: Blood pressure; Radar; Dual-task learning; Non-contact monitoring; ResNet;
Feature combination

Junkai Ge, Huaifeng Sun, Xiaodong Li, Xushan Lu, Xuening Wang, Li Li, Kejia Hu,
Decoding the stone Buddha: Three-dimensional ground penetrating radar attribute insights
into cracks and restoration history of Sumeru throne,
Journal of Cultural Heritage,
Volume 76,
2025,
Pages 39-51,
ISSN 1296-2074,
[Link]
([Link]
Abstract: The Northern Wei dynasty stone Buddha was built in 517 AD and is currently
housed in Qingdao Museum, located in Laoshan District, Qingdao, China (36°6 ′5.58″N,
120°28′23.42″E). Over the centuries, the natural weathering process and the damage caused
by various relocations has led to internal cracks on its Sumeru throne that threatens the
stability of the Buddha. Previous restoration attempts are visible on the surface of the
throne. To guarantee the quality and effectiveness of further restoration measures, it is
essential to thoroughly investigate the cracks developments and all invisible past restoration
efforts that might interfere future restoration. An ultra-wideband stepped-frequency
continuous wave (SFCW) ground penetrating radar (GPR) system was employed to perform a
non-invasive investigation of the Buddha Sumeru throne. We used a systematic imaging
method to tackle the challenges of detecting tiny internal features within the throne.
Leveraging scattering-based velocity estimation, advanced GPR signal enhancement, Stolt
migration, and envelope attribute extraction, this approach unveils a high-resolution three-
dimensional (3D) image, offering unprecedented insights into subsurface structures. The
obtained images revealed the internal cracks, details of past restoration effort, offering
valuable insights for guiding future restoration efforts. Finally, we discussed the advantages
of GPR for investigating stone statues.
Keywords: Stone Buddha; Crack detection; Restoration history; Ground-penetrating radar
(GPR); Attribute analysis

Jiahao Liu, Yiming Zhang, Liang Song, Zheng Tong,


SCB-ADAE: An attention-based deep autoencoder for ground penetrating radar signal
denoising,
Engineering Applications of Artificial Intelligence,
Volume 160, Part A,
2025,
111902,
ISSN 0952-1976,
[Link]
([Link]
Abstract: In buried object detection, recorded signals of a ground penetrating radar (GPR)
inevitably include noise interference owing to complex underground environments. Existing
rule- and data-driven denoising methods struggle to handle non-Gaussian and real-world
noise because the rule-driven ones rely on the assumptions of simplified noise
characteristics and the data-driven ones cannot capture fine- and global-scale features of a
GPR signal well. To address the problem, this study proposes an attention-based denoising
model called the Swin-Conv Block with Attention Denoising Autoencoder (SCB-ADAE). The
model first feeds a GPR signal into a SCB module, which extracts a tensor with the fine-scale
features in the signal, such as sharp reflective interfaces and abrupt amplitude variations.
The feature tensor then passes through an ADAE module that uses encoder-decoder
structure with the self-attention to enhances the representation of the global-scale signal
features. Finally, the feature tensor from the ADAE module is decoded by another SCB
module to generate a denoised GPR signal, where the tensor includes the fine-scale and
global features of the raw signal. An experiment with three types of GPR signals
demonstrates the effectiveness of the proposed model: radar signals with Gaussian noise,
radar signals with inhomogeneous-material noise, and real-world signals. radar signals with
Gaussian noise, radar signals with inhomogeneous-material noise, and real-world signals.
Experimental results demonstrate that the proposed model outperforms other state-of-the-
art denoising methods on denosing the three types of GPR signals, where the signal-to-noise
ratio, peak signal-to-noise ratio, and structural similarity index are improved to 20.64, 14.59,
and 0.366, respectively.
Keywords: Ground penetrating radar; Denoising; Attention-based model; Autoencoder

Shih-Lin Lin,
Advanced Multi-Channel Echo Separation Techniques for High-Interference Automotive
Radars,
Computers, Materials and Continua,
Volume 85, Issue 1,
2025,
Pages 1365-1382,
ISSN 1546-2218,
[Link]
([Link]
Abstract: This paper proposes an integrated multi-stage framework to enhance frequency
modulated continuous wave (FMCW) automotive radar performance under high noise and
interference. The four-stage pipeline is applied consecutively: (i) an improved independent
component analysis (ICA) blindly separates the two-channel echoes, isolating target and
interference components; (ii) a recursive least-squares (RLS) filter compensates amplitude-
and phase-mismatches, restoring signal fidelity; (iii) variational mode decomposition (VMD)
followed by the Hilbert-Huang Transform (HHT) extracts noise-free intrinsic mode functions
(IMFs) and sharpens their time-frequency signatures; and (iv) HHT-based beat-frequency
estimation reconstructs a clean echo and delivers accurate range information. Finally, key
IMFs are reconstructed into a clean signal, and a beat-frequency estimation via HHT confirms
accurate distance results, closely aligning with theoretical predictions. On synthetic data
with an input signal-to-noise ratio (SNR) of 12.7 dB, the pipeline delivers a 7.6 dB SNR gain,
yields a mean-squared error of 0.25 m2, and achieves a range root-mean-square error
(Range-RMSE) of 0.50 m. Empirical evaluations demonstrate that this enhanced ICA and
VMD/HHT scheme effectively restores the fundamental echo signature, providing a robust
approach for advanced driver assistance systems (ADAS).
Keywords: Automotive radar; FMCW; radar noise and interference; independent component
analysis (ICA); variational mode decomposition (VMD); hilbert-huang transform (HHT)

Maged Marghany,
Chapter 3 - Quantized synthetic aperture radar signal: a comprehensive exploration,
Editor(s): Maged Marghany,
Synthetic Aperture Radar Image Processing Algorithms for Nonlinear Oceanic Turbulence
and Front Modeling,
Elsevier,
2024,
Pages 51-88,
ISBN 9780443191558,
[Link]
([Link]
Abstract: This chapter presents a novel perspective on integrating quantum mechanics to
elucidate the complexities of synthetic aperture radar (SAR). Distinguishing itself from
previous works, this chapter avoids an exhaustive exploration of radar theory and signal
processing, as these topics have been thoroughly addressed elsewhere. Instead, the focus is
on introducing a pioneering concept of quantization to comprehend the mechanisms
involved in radar imaging of sea turbulence. Subsequent chapters will delve deeper into this
concept. This chapter seizes the opportunity to clarify the quantization of radar, providing a
lucid understanding of the term. The foundational quantization of electromagnetic waves,
crucial to radar operations, is initially examined. The discussion begins by elucidating why
photons, as fundamental units, exhibit quantization—an essential aspect for understanding
quantized electromagnetic waves. This inquiry reveals why photons possess discrete energy
amounts within a specific energy spectrum, departing from a continuous range. Expanding
beyond conventional discourse, the exploration extends to the detection of quantum
microwave propagation—a novel discussion in the realm of synthetic aperture radar
publications. Moreover, the quantization of the radar equation is implemented to introduce
a fresh perspective on the entanglement between radar cross section and sea surface
turbulence—an aspect previously unexplored in radar oceanography. Noteworthy is the
chapter’s foray into the novel domain of quantum pulse-compression ranging. In the
concluding sections, a comprehensive list of sensors associated with synthetic aperture
radar satellites is compiled, detailing their physical characteristics, including bands and
frequencies. In summary, the quantization of radar signals proposes an innovative
speculation regarding the entanglement of radar cross section with turbulence spectra, such
as the Kolmogorov energy spectra—a contribution unprecedented in the domain of radar
oceanography.
Keywords: Quantum physics; physics; optics; mesoscopic physics; mathematical physics;
quantum mechanics; instrumentation; physical chemistry; space physics; superconductivity;
physics education; computational physics; computer science; statistical applications;
radiation physics; emergent computing; plasma physics; materials characterization; atomic
physics; quantum cosmology

Mohammad Hossein Shirazi, Sira Yongchareon, Anuradha Singh, Jing Ma,


A survey on machine learning approaches for vital sign monitoring using radar,
Measurement,
Volume 253, Part D,
2025,
117707,
ISSN 0263-2241,
[Link]
([Link]
Abstract: The integration of machine learning methodologies with radar-based vital sign
monitoring represents a significant advancement in non-contact healthcare surveillance
systems. This systematic literature review synthesizes and critically analyzes research from
2020 to 2025, addressing substantive theoretical and methodological gaps in extant
literature. Our comprehensive taxonomic classification of machine learning paradigms
employed in this domain elucidates the progressive refinement from conventional
algorithmic approaches to sophisticated deep learning architectures, with particular
emphasis on hybrid neural network configurations optimized for physiological signal
extraction in non-stationary environments. Methodologically, this survey contributes a
rigorous evaluation framework comprising standardized assessment protocols, quantifiable
performance metrics, and cross-validation methodologies—elements conspicuously absent
in previous reviews. Empirical analysis demonstrates substantial correlations between
dataset demographic characteristics and algorithmic generalizability, with heterogeneous
participant cohorts yielding markedly enhanced performance across cardiac, respiratory, and
hemodynamic parameter estimation tasks. The review delineates four distinct
developmental phases in the field’s chronological evolution and provides analytical insight
into persistent technical challenges: motion artifact compensation, multi-subject
disambiguation, and the translation of laboratory efficacy to clinical utility. This
comprehensive examination of computational approaches for radar-based vital sign
monitoring establishes a theoretical foundation and methodological framework to guide
future research towards physiologically robust and clinically viable implementations.
Keywords: Non-intrusive vital sign monitoring; Machine learning; Radar

Aurora Polo-Rodríguez, Miguel Ángel Anguita-Molina, Ignacio Rojas-Ruiz, Javier Medina-


Quero,
Multi-occupant tracking with radar and wearable devices for enhanced accuracy in indoor
environments,
Engineering Applications of Artificial Intelligence,
Volume 154,
2025,
110872,
ISSN 0952-1976,
[Link]
([Link]
Abstract: This work explores the integration of millimetre-wave (mmWave) radar and a
minimal configuration of ultra-wideband (UWB) devices for enhanced multi-occupant
tracking in real domestic environments. Using a low-cost, non-intrusive, and rapidly
deployable device setup, our approach addresses key challenges in multi-occupant tracking,
including individual identification and ease of installation. While mmWave radar precisely
detects occupant presence, it lacks individual recognition and exhibits limited sensitivity.
This limitations are addressed by incorporating a minimal configuration of UWB (wearable
tags and ambient anchors), enabling individual identification through signal strength
measurements. Several data autoencoder models, such as long short-term memories
(LSTMs), Convolutional Neural Networks (CNN) or Transformers, were evaluated.
Experiments conducted in two real-world domestic settings, each with up to three
inhabitants, demonstrate the effectiveness of combining mmWave and UWB technologies
for indoor multi-occupant tracking. Our results show that ConvLSTM achieves the best
performance with a mean squared error (MSE) between 0,0142 and 0,0433 in single and
multi-occupation, respectively. These findings suggest promising applications for accurate
inhabitant tracking in ambient assisted living and other smart environment contexts.
Keywords: Multi-tracking; Ultra-wideband; Wave radar; Autoencoder models

Yuankang Ye, Feng Gao, Shaoqing Zhang, Chang Liu,


Improving precipitation nowcasting via multiphysical parameter fusion in radar echo
extrapolation,
Journal of Hydrology,
Volume 668,
2026,
134947,
ISSN 0022-1694,
[Link]
([Link]
Abstract: Radar-based precipitation nowcasting plays a vital role in short-term
hydrometeorological forecasting and water resource management. Existing modeling
methodologies typically simplify precipitation nowcasting to a task of spatiotemporal
sequence prediction based on radar echo reflectivity data. However, the reliance on
unimodal reflectivity data including intensity-only information restricts the model’s ability to
characterize the phase evolution and dynamic processes of hydrometeor particles,
ultimately leading to insufficient extrapolation accuracy. This study breaks through the
conventional unimodal data paradigm, aiming to capture the complex dynamic evolutionary
features of hydrometeor particles. We integrate radar echo reflectivity and four additional
physical parameters of hydrometeor particles into a deep learning framework and propose a
novel Physics-Informed Multimodal Echo Extrapolation neural network (PIEE). Furthermore,
we systematically investigate the individual contributions of each physical parameter to the
accuracy of radar echo extrapolation. Specifically, PIEE adopts a three-stage structure. First, a
multimodal encoder with a dual-branch attention-based fusion strategy is used to capture
diverse physical signals. Second, a novel gated spatiotemporal self-attention module is
designed for deep feature extraction. Finally, the decoding stage generates the extrapolated
radar echoes. Experimental results on a real multimodal radar echo dataset show that the
proposed model demonstrates superior performance in two aspects. First, under a unimodal
baseline architecture, the PIEE model clearly outperforms the comparison model. Second,
after fusing multiple physical parameters, the PIEE achieves significant improvements in all
the evaluated metrics, especially in the CSI and HSS metrics for the high echo intensity
region (≥ 40 dBZ), with improvements of up to 24.2% and 20.3%, respectively. Furthermore,
systematic ablation experiments on physical parameters quantify the effects of different
combination methods on extrapolation accuracy, highlighting the potential of physics-
informed, multimodal deep learning approaches in improving short-term hydrological
prediction accuracy, with implications for flood forecasting, early warning systems, and
hydrometeorological risk management at catchment scales.
Keywords: Deep learning; Precipitation nowcasting; Radar echo extrapolation;
Hydrometeorological forecasting

Jiayi Cai, Zhaocheng Yang, Ping Chu, Juntao Guo, Jianhua Zhou,
Robust hand gesture detection and recognition using 4D millimeter-wave radar in a
ubiquitous scene,
Measurement,
Volume 253, Part C,
2025,
117545,
ISSN 0263-2241,
[Link]
([Link]
Abstract: In current research on HGR using radar sensors, hand gestures are typically
confined to a smaller region. However, in ubiquitous scenarios, unrestricted human body
movements and unexpected hand gesture motions usually occur, which results in a large
false alarms and recognition performance degradation. To address this issue, we propose a
robust hand gesture detection and recognition method in ubiquitous scenarios using
Frequency-Modulated Continuous Wave (FMCW) Multiple-Input Multiple-Output (MIMO)
radar. The core idea is to progressively define and classify motions in a cascaded manner,
gradually filtering out non-specific movements, reducing false positives, and enhancing the
applicability of HGR. Specifically, we first propose a suspected hand gesture motion
detection method to help identify suspicious hand gestures. Then, the velocity and position
features of the mutated signal and the stable signal are extracted. A mutated signal motion
recognition method based on a single-layer long short-term memory (LSTM) network is used
to effectively distinguish non-hand gesture motions from hand gestures. Finally, the two-
dimensional trajectory features are extracted, and cascaded with a LSTM network combined
a Gaussian probability model is developed to enhance the ability of open-set recognition.
Experimental results show that the proposed method can achieve the recognition accuracy
of 99.53% for designed hand gestures, the false alarm rate of 1.5% for unexpected hand
gestures and 0.11% for non-hand gesture motions.
Keywords: Hand gesture recognition; Non-hand gesture motion; Feature extraction;
Probability models; Ubiquitous scene

Qiangyu Zeng, Ling Li, Hao Wang, Jianxin He, Hua Wang, Yao Gao,
MCDA-UNet: A satellite data-based model for radar composite reflectivity retrieval,
Atmospheric Research,
Volume 330,
2026,
108619,
ISSN 0169-8095,
[Link]
([Link]
Abstract: The weather radar network in China exhibits an uneven spatial distribution, with
dense coverage in the eastern regions and sparse deployment in the west, resulting in
substantial detection blind spots in areas with complex terrain. This severely limits the
continuity and precision of weather monitoring and early warning in these regions. To
address this challenge, a multi-channel deep learning model, MCDA-UNet, is proposed for
radar composite reflectivity retrieval, aiming to reconstruct and enhance radar echo patterns
in regions lacking radar coverage by leveraging the extensive spatial coverage and
continuous observation capabilities of geostationary meteorological satellites. The model
employs a multi-channel input architecture to extract features from different spectral bands,
while spatial and channel attention models are incorporated to improve the representation
of key meteorological information, thereby enhancing retrieval accuracy and regional
adaptability. Comparative experiments conducted under varying precipitation intensities
demonstrate that MCDA-UNet consistently outperforms existing models across multiple
evaluation metrics, particularly in reconstructing weather radar echo structures and edge
details. These results validate the model’s capability to adapt to the full dynamic range of
weather radar reflectivity and highlight its potential for accurate precipitation retrieval in
radar blind-spot regions.
Keywords: Weather radar composite reflectivity; Satellite data retrieval; Multi-channel
structure; Full dynamic range

Nan Xia, Siqi Wang, Weijia Lu,


Automotive radar co-channel interference mitigation and target detection based on third-
order cumulant and neural network,
Measurement,
Volume 266,
2026,
120494,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Automotive millimeter-wave (mmWave) radar systems face significant challenges
from noise and co-channel interference, which can obscure weak target reflections. This
paper proposes a novel target localization algorithm that combines a third-order cumulant
based interference suppression technique with a dual-stream deep neural network to
enhance detection performance. The third-order cumulant operation exploits higher-order
statistics to suppress co-channel interference among radars. A dual-stream convolutional
neural network then extracts image-based features from heatmaps and edge maps, enabling
robust detection and localization of weak and overlapping targets in complex multi-target
scenarios. The proposed method is evaluated on both simulated radar data and real
measurements from an automotive mmWave radar. Experimental results demonstrate that
the proposed method effectively mitigates interference and significantly improves target
detection accuracy and localization precision compared to baseline methods. These results
indicate that the integration of higher-order statistical filtering with advanced neural
networks can greatly enhance radar target localization in challenging interference
environments.
Keywords: Automotive radar; Co-channel interference mitigation; The third-order cumulant;
Deep learning; Target localization

Zhipeng Qing, Kecheng Ge, Shunsheng Zhang, Jing Yang, Zhijin Wen, Youlei Pu,
An inverse synthetic aperture radar imaging framework based on multi-layer networks and
heat conduction attention,
Engineering Applications of Artificial Intelligence,
Volume 167, Part 1,
2026,
113708,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Accurate compensation is essential for achieving high-resolution inverse synthetic
aperture radar (ISAR) imaging. Traditional parametric methods usually rely on iterative
optimization of the objective function to compensate for target motion in radar echoes.
However, the iteration process is often computationally intensive, difficult to integrate into
deep learning frameworks, and may discard sufficiently acceptable intermediate solutions.
To address these challenges, this study proposes a deep unfolding-based translational
compensation network that combines unsupervised learning with gradient back-
propagation. A prototype network is incorporated to monitor the imaging process, enabling
early termination of iterations. Moreover, a U-shaped network architecture based on a heat
conduction attention is employed to enhance ISAR image resolution and focusing
performance. To solve the problem of offset or splitting in the imaging results caused by
residual motion errors, a learnable affine transformation is employed for automatic
centering. These modules are integrated into an echo-to-image ISAR imaging framework.
Experimental results on both simulated and real radar data demonstrate the framework’s
effectiveness and robustness.
Keywords: Inverse synthetic aperture radar imaging; Translational compensation; Heat
conduction; Deep unfolding network; Affine transformation

Purabi Sharma, Kandarpa Kumar Sarma,


Collaborative and distributive intelligence in radar systems: Enhancing electronic jamming
discrimination,
Computers and Electrical Engineering,
Volume 124, Part A,
2025,
110357,
ISSN 0045-7906,
[Link]
([Link]
Abstract: The precise discrimination of radar jamming signals is decisive in executing
effective electronic counter-countermeasures (ECCM). Data-driven deep learning (DL)
models have proven effective for this task, but challenges persist in addressing critical real-
world issues such as limited data sharing, time- and location-dependent variations in hostile
interferences, real-time adaptability to evolving jamming tactics, and imbalanced data
distribution. In this paper, a novel approach is proposed that employs federated learning (FL)
for training radar signal jamming classifiers individually on a set of devices, aggregating,
sharing, and continuously updating the knowledge. This approach enables privacy-
preserving training, eliminating access to client-local data or centralized data storage,
ensuring knowledge sharing, and resilient response to jamming. This work delineates a
collaborative and distributive learning framework for radar jamming signals, employing two
hybrid models within an FL platform. A deep spectra spatio-temporal discriminator (DSSTD)
and a shallow spectra-temporal feed-forward self-attention-driven discriminator (SSTFSAD)
network have been implemented as classifiers on a distributed arrangement. The time–
frequency attributes of radar jamming signals as 2D-scalograms individually train these
models across remote FL nodes. Extensive evaluations are conducted in FL environments
under independently and identically distributed (IID) and non-IID data configurations,
simulating real-world settings with diverse data distributions. Experimental results
demonstrate that both proposed approaches are effective, with FL-driven DSSTD
outperforming FL-driven SSTFSAD by 7.7% in IID and 6.8% in non-IID setups, even at -5 dB
JNR. These results highlight the robustness and adaptability of the FL-driven DSSTD model,
offering a significant advancement in radar jamming signal discrimination for electronic
warfare applications.
Keywords: Federated learning; Radar jamming signal; Multi-scale CWT; Deep learning; Long
short-term memory; Self-attention
Farhana Ahmed Chowdhury, Md Kamal Hosain, Md Sakib Bin Islam, Md Shafayet Hossain,
Promit Basak, Sakib Mahmud, M. Murugappan, Muhammad E.H. Chowdhury,
ECG waveform generation from radar signals: A deep learning perspective,
Computers in Biology and Medicine,
Volume 176,
2024,
108555,
ISSN 0010-4825,
[Link]
([Link]
Abstract: Cardiovascular diagnostics relies heavily on the ECG (ECG), which reveals significant
information about heart rhythm and function. Despite their significance, traditional ECG
measures employing electrodes have limitations. As a result of extended electrode
attachments, patients may experience skin irritation or pain, and motion artifacts may
interfere with signal accuracy. Additionally, ECG monitoring usually requires highly trained
professionals and specialized equipment, which increases the treatment's complexity and
cost. In critical care scenarios, such as continuous monitoring of hospitalized patients,
wearable sensors for collecting ECG data may be difficult to use. Although there are issues
with ECG, it remains a valuable tool for diagnosing and monitoring cardiac disorders due to
its non-invasive nature and the detailed information it provides about the heart. The goal of
this study is to present an innovative method for generating continuous ECG waveforms
from non-contact radar data by using Deep Learning. The method can eliminate the need for
invasive or wearable biosensors and expensive equipment to collect ECGs. In this paper, we
propose the MultiResLinkNet, a one-dimensional convolutional neural network (1D CNN)
model for generating ECG signals from radar waveforms. With the help of a publicly
accessible radar benchmark dataset, an end-to-end DL architecture is trained and assessed.
There are six ports of raw radar data in this dataset, along with ground truth physiological
signals collected from 30 participants in five distinct scenarios: Resting, Valsalva, Apnea, Tilt-
up, and Tilt-down. By using strong temporal and spectral measurements, we assessed our
proposed framework's ability to convert ECG data from Radar signals in three distinct
scenarios, namely Resting, Valsalva, and Apnea (RVA). ECG segmentation performed better
by MultiResLinkNet than by state-of-the-art networks in both combined and individual
cases. As a result of the simulations, the resting, valsalva, and RVA scenarios showed the
highest average temporal values, respectively: 66.09523 ± 19.33, 60.13625 ± 21.92, and
61.86265 ± 21.37. In addition, it exhibited the highest spectral correlation values
(82.4388 ± 18.42 (Resting), 77.05186 ± 23.26 (Valsalva), 74.65785 ± 23.17 (Apnea), and
79.96201 ± 20.82 (RVA)), along with minimal temporal and spectral errors in almost every
case. The qualitative evaluation revealed strong similarities between generated and actual
ECG waveforms. As a result of our method of forecasting ECG patterns from remote radar
data, we can monitor high-risk patients, especially those undergoing surgery.
Keywords: ECG; Raw radar data; MultiResLinkNet; CNN; Deep learning

Cries Avian, Jenq-Shiou Leu, Hang Song, Jun-ichi Takada, Nur Achmad Sulistyo Putro,
Muhammad Izzuddin Mahali, Setya Widyawan Prakosa,
RCTrans-Net: A spatiotemporal model for fast-time human detection behind walls using
ultrawideband radar,
Computers and Electrical Engineering,
Volume 120, Part C,
2024,
109873,
ISSN 0045-7906,
[Link]
([Link]
Abstract: Ultrawideband (UWB) radar systems are becoming increasingly popular for
detecting human presence, even through walls. Recent advancements in signal processing
use deep learning techniques, which are known for their accuracy. While earlier methods
focused on spatial information using Convolutional Neural Networks (CNNs), newer research
highlights the importance of temporal information, such as how data peaks shift over time.
This study introduces RCTrans-Net, a deep-learning architecture that combines RCNet (a
Residual CNN) for spatial features with TransNet (a Transformer) for temporal features. This
fusion improves human presence classification in fast-time signal processing. Tested under
various conditions—different materials, body orientations, ranges, and radar heights—
RCTrans-Net achieved high performance with F1-scores of 0.997±0.000 for static,
0.967±0.004 for dynamic, and 0.978±0.001 for combined scenarios. The architecture
outperforms previous methods and offers real-time processing with an inference time of
about one millisecond.
Keywords: Human presence behind the wall; Residual network; Spatiotemporal'
Transformer; Ultrawideband radar system

Yuanbo Li, Wenwu Zhang, Songtao Lv, Jing Yu, Dongdong Ge, Jiawei Guo, Lin Li,
YOLOv11-CAFM model in ground penetrating radar image for pavement distress detection
and optimization study,
Construction and Building Materials,
Volume 485,
2025,
141907,
ISSN 0950-0618,
[Link]
([Link]
Abstract: Ground Penetrating Radar (GPR) is an effective technology for detecting
underground structures and has been widely utilized for monitoring road damage.
Traditional B-scan-based one-dimensional images often fail to preserve continuous spatial
information, thus inadequately reflecting the nuances of damage patterns. This paper
investigates the accurate recognition of hidden internal road damage using 3D-sliced C-scan
images. While YOLO is one of the most effective and rapid neural network models for object
detection, it still suffers from low recognition accuracy and a high rate of missed detections.
To address these issues, this study proposes an improved Convolution and Attention Fusion
Module (CAFM) fusion network model for YOLOv11, which combines the CAFM with the
C2PSA global-local feature extraction mechanism to significantly enhance the recognition
performance for complex road damage. Experimental comparisons between the YOLOv11m-
CAFM and the YOLOv11 model reveal that the combined metrics for the small (n/s) and large
(l/x) models are lower than those for the medium model (m). The YOLOv11m-CAFM
demonstrates strong performance in key metrics such as precision, recall, mAP50, and
mAP50:95, achieving values of 0.840, 0.850, 0.881, and 0.584, respectively, representing
improvements of 0 %, 4.6 %, 1.8 %, and 2.0 % over the baseline model. The confidence level
in detecting standardized targets (e.g., pipelines and well covers) exceeds 0.89. Borehole
validation confirms that the model's localization error is less than 0.15 m, and the detection
frame rate reaches 71 FPS, satisfying the requirements for rapid road assessment. This study
offers a novel method for the intelligent interpretation of GPR images, considering both
detection accuracy and real-time performance, which holds significant engineering
applications in identifying hidden road damages.
Keywords: Pavement disease detection; Ground penetrating radar; Neural network; Object
detection; Deep learning algorithm

Lele Qu, Jinpeng Tao, Tianhong Yang,


Human sleep posture recognition and vital sign monitoring method using FWCW radar,
Measurement,
Volume 258, Part D,
2026,
119349,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Radar technology has the great potential for recognizing human sleep postures and
tracking respiration rate (RR) and heart rate (HR) during sleep. In this paper, we propose a
novel integrated framework that simultaneously performs sleep posture recognition and
vital sign monitoring using dual frequency-modulated continuous-wave (FMCW) radar
modules. A distinctive aspect of the proposed framework is its dynamic radar switching
strategy, where sleep posture is first classified by applying histogram of oriented gradients
(HOG) and support vector machine (SVM) to the merged range-time map (RTM). Based on
the identified posture, either the top or side radar is selectively engaged to capture
physiological signals from the most effective orientation. To obtain the robust and accurate
estimation of RR and HR, we further propose the GA-MVMD algorithm that integrates
genetic algorithm (GA) and multivariate variational mode decomposition (MVMD) to jointly
decompose chest wall displacement signals across multiple range bins. Experimental results
demonstrate that the proposed method can effectively enhance the recognition accuracy
with an average accuracy rate of 96.7 % for four typical sleep postures and the proposed GA-
MVMD algorithm can provide more accurate RR and HR estimation results.
Keywords: Frequency modulated continuous wave (FMCW) radar; Sleep posture recognition;
Vital sign monitoring

Xuexiang Zhu, Linxin Zhao, Dian Zhang,


Novel Positioning via Ground Penetrating Radar Image Matching,
Procedia Computer Science,
Volume 259,
2025,
Pages 941-949,
ISSN 1877-0509,
[Link]
([Link]
Abstract: Positioning technology is crucial for numerous real-world applications, such as
navigation, robotics and autonomous driving. However, existing methods like GPS, Wi-Fi and
Bluetooth positioning technologies have limitations. such as weak signals, significant
interference, limited effective ranges and adverse weather conditions. Visual positioning, on
the other hand, is sensitive to lighting conditions and performs poorly when optical features
become obscured. Therefore, there is an urgent need for novel solutions to address
positioning challenges in complex environments. In this paper, we propose a new positioning
system based on enhanced Siamese deep network using ground penetrating radar (GPR)
data, aiming to achieve high-precision positioning. The proposed method employs a dual-
channel input mechanism to process query images and reference images respectively,
extract key features and calculate the similarities, thereby achieving precise positioning. The
model incorporates an attention mechanism and enhances feature extraction capabilities by
introducing the SE module. The training and verification results on the surface fingerprint
data set show that the image matching method based on the Siamese network significantly
improves the matching accuracy and positioning ability of GPR data. Experimental results
indicate that the proposed system achieved relatively high-precision positioning using single
ground penetrating radar modality alone, successfully overcome many limitations faced by
existing positioning technologies, thus, it provides a new direction for the development of
intelligent positioning and navigation systems. Ultimately, the findings in this work offers a
reliable positioning solution for applications such as autonomous driving, intelligent robots,
and drone navigation.
Keywords: ground penetrating radar; precise positioning; localization; Siamese network

Zai Zhang, Bin Shi, Kai Sun, Hao Wu, Bo Dong,


RADAR: Relation-assisted dual-graph aligning recognition for grounded multimodal named
entity recognition,
Information Processing & Management,
Volume 63, Issue 3,
2026,
104552,
ISSN 0306-4573,
[Link]
([Link]
Abstract: Grounded Multimodal Named Entity Recognition (GMNER) requires the
simultaneous identification of textual entities and their corresponding visual regions within
images. However, the inability to model visual contextual semantics and the disorganized
processing of cross-modal features often lead existing methods to struggle with both visual
entity differentiation and bridging the modality gap. We propose RADAR (Relation-Assisted
Dual-graph Aligning Recognition), a novel framework that leverages visual relations derived
from scene graphs to encode structured context and enhance visual understanding. To
achieve fine-grained cross-modal alignment, we design an object-level alignment self-
attention mechanism and introduce a dual-graph strategy. Evaluated on the Twitter-GMNER
dataset (13,076 image-text pairs), RADAR achieves 60.91 % F1 score, a +4.5 % improvement
over the H-Index baseline. The method also demonstrates consistent gains in subtasks, with
+3.54 % improvement in EEG metric, validating its effectiveness in multimodal entity
alignment.
Keywords: Multimodal named entity recognition; Visual grounding; Visual scene graph;
Cross-modal alignment

Qi Cheng, Shiwen Zhang, Xiaoyang Chen, Hongbiao Cui, Yunfei Xu, Shasha Xia, Ke Xia, Tao
Zhou, Xu Zhou,
Inversion of reclaimed soil water content based on a combination of multi-attributes of
ground penetrating radar signals,
Journal of Applied Geophysics,
Volume 213,
2023,
105019,
ISSN 0926-9851,
[Link]
([Link]
Abstract: Rapid, accurate, and non-destructive acquisition of the distribution of reclaimed
soil moisture information can provide data for the rapid monitoring of reclaimed soil in areas
experiencing coal mining subsidence. However, the water content inversion methods based
on ground penetrating radar (GPR) are mostly single-attribute analysis methods, which are
easily affected by the soil structure. This paper proposes a multi-attribute joint analysis
method, which can reduce the influence of complex soil structure on the prediction results.
Surveys and soil sampling using GPR were performed on a subsided reclamation area in
Huaibei City, Anhui Province, China. Correlation analysis of GPR attribute information and
volumetric water content (VWC) showed that frequency peak (FP), average envelope
amplitude (AEA), energy, instantaneous amplitude area, instantaneous frequency area, and
average instantaneous frequency values were significantly related to the VWC of reclaimed
soil. The applicability of different single-attribute analysis methods under the condition of
reclaimed soil was compared. Results showed that the predictive effects of FP and AEA
attributes were better than other attributes. Generally, the single-attribute analysis method
was greatly affected by the structure of reclaimed soil, so the accuracy and reliability of this
method need to be further optimized. The prediction results of the single- and multi-
attribute joint analysis methods in the structure of reclaimed soil were compared and
analyzed. This showed that the prediction accuracy and model reliability of the multi-
attribute method are both higher than those of the single-attribute method. The multi-
attribute method can overcome the problem of insufficient accuracy of water content
detection of ground-seeking radar in soil with a complex structure. Finally, the multi-
attribute method was used to obtain the water distribution information in the reclamation
area. Results of this study provide new methods and ideas for the prediction of water
content by GPR in soils with complex structure.
Keywords: Ground penetrating radar; Volumetric water content; Attribute analysis; Land
reclamation

Haoming Feng, Huaqing Li, Wenwen Zhu, Denghao Li, Yukun Huang,
Micro-motion enhanced multi-person activity recognition with millimeter-wave radar,
Measurement,
Volume 258, Part B,
2026,
119090,
ISSN 0263-2241,
[Link]
([Link]
Abstract: As a non-contact sensing device, millimeter-wave radar exhibits unique strengths
in human activity recognition (HAR). Existing methods rely on micro-Doppler signatures for
activity classification, but they often encounter feature aliasing in multi-person activity
recognition (MPAR) scenarios. Although point cloud-based approaches can distinguish
individual targets, they primarily extract static morphological features, neglecting the micro-
motion information of human joints, which is crucial for accurate activity recognition. To
address these limitations, we proposes an innovative MPAR framework that integrates
spatial point clouds and micro-motion features. First, an improved point cloud data
association algorithm is applied to achieve multi-target point cloud feature separation,
followed by a dynamic projection mechanism to construct time–Doppler feature maps.
Then, a torso micro-motion enhancement algorithm is designed to enhance the details of
human body movements. Finally, a CNN-LSTM hybrid network architecture with a temporal-
attention is constructed for action classification. Experimental results show that the
proposed micro-motion enhancement algorithm improves recognition accuracy by 27.1%
and 2.3%, compared to two traditional time–frequency analysis methods. Furthermore,
MPAR task in occlusion scenarios achieves recognition accuracy of 93.5%. In summary,
proposed framework not only retains the inherent advantages of millimeter-wave radar but
also significantly enhances multi-person activity recognition in complex scenarios.
Keywords: Human activity recognition (HAR); Multi-person activity recognition (MPAR);
Feature separability; Temporal attention; Occlusion scenarios

Ligen Chen, Nannan Zhu, Hongbo Chen, Yonghao Dong, Yue Zhang, Nian Cai,
A causality-inspired single-source domain generalized method for low-slow-small threat
target recognition through holographic Doppler radar,
Expert Systems with Applications,
Volume 287,
2025,
128104,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Low-slow-small (LSS) target monitoring is critical for airport safety management,
particularly when LSS objects such as unmanned aerial vehicles (UAVs) and birds
unexpectedly enter airport airspace, posing significant risks to regular flights and airport
operations. Recent advances have taken advantage of deep learning for the recognition of
LSS radar targets, achieving promising classification accuracy. However, existing LSS radar
target recognition approaches often rely on statistical correlations, including unstable
spurious correlations, which can undermine the generalization performance of classification
networks, limiting their effectiveness in all-time radar recognition. To address this, we
propose a causality-inspired single-source domain generalization method for radar LSS target
recognition. Our method introduces a Causal-Symmetric Transformation (CST) module for
data augmentation, combining Non-Causal Augmentation for global perturbations and
Symmetric Transformation for local motion reversal, enhancing data diversity and reducing
bias. Additionally, we propose a Causal Mining (CM) module with a Causal Consistency loss
to extract causal features that boost generalization. A Fourier-Aware Attention (FAA) module
leverages frequency-domain information to strengthen feature representation and preserve
causal information. Extensive experiments on four real-world datasets validate the
effectiveness of our approach.
Keywords: Radar target recognition; Single-source domain generalization; Causility-inspired
model; Low-slow-small target; Holographic Doppler radar

Yun Zhou, Yinglin Zhu, Haohao Ren, Jiahao Kang, Xuegang Wang,
Refined multi-modal feature learning framework for marine target detection using radar
sensor,
Digital Signal Processing,
Volume 170,
2026,
105816,
ISSN 1051-2004,
[Link]
([Link]
Abstract: The fusion of time and time-frequency characteristics in radar echoes offers a
novel approach for marine target detection. However, echo amplitude alone cannot fully
characterize the time-domain information, as it fails to capture the temporal correlation
between sampling points. Therefore, this article introduces the Gramian Angular Summation
Field (GASF) for processing raw radar echoes to obtain the temporal information. Concretely,
to enable the detector to utilize features from diverse signal representations of the same
target echoes, we first preprocess the echoes of radar with two signal processing methods,
GASF and STFT, which aim to reflect the temporal dependence and dynamic changes of
frequency components, respectively. Subsequently, we develop a dual-stream feature
extraction network, i.e., time-frequency self-attention learning and GASF-based spatial-
temporal correlation learning, to deeply extract the discriminative features from two
modalities of the same radar echo. Then, to overcome the heterogeneity of multimodal
features during feature fusion, we propose a cross-modal feature fusion strategy to map
multi-modal features to a unified space. Finally, the fused features are fed into the detection
module. Numerous evaluation experiments on the publicly available measured IPIX dataset
demonstrate that the proposed detector is competitive with some state-of-the-art detectors
for marine target detection.
Keywords: Radar target detection; Signal processing; Gramian angular summation field;
Deep learning; Short-time Fourier transform

Wentao He, Jianfeng Ren, Ruibin Bai, Xudong Jiang,


Radar gait recognition using Dual-branch Swin Transformer with Asymmetric Attention
Fusion,
Pattern Recognition,
Volume 159,
2025,
111101,
ISSN 0031-3203,
[Link]
([Link]
Abstract: Video-based gait recognition suffers from potential privacy issues and performance
degradation due to dim environments, partial occlusions, or camera view changes. Radar has
recently become increasingly popular and overcome various challenges presented by vision
sensors. To capture tiny differences in radar gait signatures of different people, a dual-branch
Swin Transformer is proposed, where one branch captures the time variations of the radar
micro-Doppler signature and the other captures the repetitive frequency patterns in the
spectrogram. Unlike natural images where objects can be translated, rotated, or scaled, the
spatial coordinates of spectrograms and CVDs have unique physical meanings, and there is
no affine transformation for radar targets in these synthetic images. The patch splitting
mechanism in Vision Transformer makes it ideal to extract discriminant information from
patches, and learn the attentive information across patches, as each patch carries some
unique physical properties of radar targets. Swin Transformer consists of a set of cascaded
Swin blocks to extract semantic features from shallow to deep representations, further
improving the classification performance. Lastly, to highlight the branch with larger
discriminant power, an Asymmetric Attention Fusion is proposed to optimally fuse the
discriminant features from the two branches. To enrich the research on radar gait
recognition, a large-scale NTU-RGR dataset is constructed, containing 45,768 radar frames of
98 subjects. The proposed method is evaluated on the NTU-RGR dataset and the MMRGait-
1.0 database. It consistently and significantly outperforms all the compared methods on
both datasets. The codes are available at: [Link]
Keywords: Micro-Doppler signature; Radar gait recognition; Spectrogram; Cadence velocity
diagram; Asymmetric Attention Fusion

Yuanzhi Su, Huiying Cynthia Hou, Chun Zhao, Zhuojun Nan,


PoseGraphNet: Pose prior and graph structure for 3D human pose estimation using
mmWave radar,
Measurement,
Volume 257, Part C,
2026,
118851,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Human pose estimation (HPE) is a crucial task in computer vision with extensive
applications in healthcare, surveillance, and human–computer interaction. Traditional HPE
research primarily utilizes RGB cameras, which may suffer from poor performance under
varying lighting conditions and raise privacy concerns. Recently, millimeter-wave (mmWave)
radar technology has emerged as a promising alternative, providing a non-invasive and
privacy-preserving solution for HPE. However, the progress in mmWave-based HPE is
hindered by the limited availability of high-quality datasets that encompass a diverse range
of poses and provide accurate data annotations. Current mmWave-based datasets for HPE
often feature only basic poses or rely on imprecise annotations, typically derived from pre-
trained image-based HPE models using synchronized RGB images, which can limit the
potential of derived models. This study introduces a pioneering approach to HPE by
synergizing wearable motion capture sensors with mmWave radar technology to create a
comprehensive and precise dataset tailored for enhancing HPE with mmWave radar.
Leveraging this dataset, we develop an innovative deep learning framework specifically
designed to explore the unique properties of radar signals for HPE. The performance of our
proposed model is evaluated and compared with several well-known deep learning models.
Extensive experimental results affirm the robustness of the dataset, establishing it as a
rigorous benchmark for mmWave radar-based HPE. The proposed methodology
demonstrates exceptional accuracy in estimating human poses from radar data, setting the
stage for its application in environments where privacy and complexity are critical concerns.
Keywords: Human pose estimation; 3D point cloud; Millimeter radar; Deep learning

Yuanjia Xia, Guobing Chen, Zhen Zhang, Shuang Zhao, Zhifang Fei, Kunfeng Li, Xiaoxiao Xia,
Zichun Yang,
Design, research progress and prospects of high temperature infrared/radar compatible
stealth materials,
Optical Materials,
Volume 165,
2025,
117156,
ISSN 0925-3467,
[Link]
([Link]
Abstract: Multispectrum-compatible stealth materials, and in particular, infrared/radar
compatible materials, constitute one of the most important research areas in the stealth
technology field. Although such materials have been extensively investigated at room
temperature, those intended for the high temperature power parts of weapons and
equipment have recently gained increasing attention. This study first analyses and
summarises several typical conventional infrared/radar compatible stealth materials
intended for high temperature conditions from a structural design and mechanistic
viewpoint, then briefly summarises the research status of infrared/radar compatible stealth
metamaterials applicable under high temperature conditions, and finally offers insights into
future development directions.
Keywords: High temperature; Radar wave; Infrared; Compatible stealth

Kuiyu Chen, Jingyi Zhang, Si Chen, Shuning Zhang,


Deep metric learning for robust radar signal recognition,
Digital Signal Processing,
Volume 137,
2023,
104017,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Signal recognition technology is a currently active area in both civilian and military
applications. Recently, deep learning has aroused extensive attempts in radar signal
recognition due to its remarkable capability of automatic feature extraction. However,
existing radar signal recognition networks overly depend on the probability-based decision
model, resulting in poor robustness. This paper develops a novel deep metric learning frame
to enhance the robustness of the recognition system. First, a multiscale atrous pyramid
network (MAPNet) is proposed to efficiently learn high-resolution and distinct feature
representation. Second, a variance loss is designed to constrain the intra-class feature
distribution in metric space. Third, according to the distribution of training signals in metric
space, recognition results are recalibrated to provide explicit rejection probabilities for
unknowns. Extensive experiments and evaluations demonstrate that the proposed model
can accurately classify known signals while robustly identifying unknown signals. The signal
database and model can be freely accessed at [Link]
Keywords: Radar signal recognition; Metric learning; Multiscale atrous pyramid; Variance
loss; Unknown signals

Oleg Berngardt, Ivan Lavygin,


Self-learning signal classifier for HF coherent scatter radars,
Advances in Space Research,
2025,
,
ISSN 0273-1177,
[Link]
([Link]
Abstract: The paper presents a method for automatic constructing a classifier for processed
data obtained by HF coherent scatter radars. The method is based only on the radar data
obtained, the results of automatic modeling of radio wave propagation in the ionosphere,
and mathematical criteria for estimating the quality of the models. The final classifier is the
model trained on data obtained by 12 radars of the SuperDARN and SECIRA networks over
two years for each radar. The model has 2,669 coefficients. For the classification, the model
uses both the calculated parameters of radio wave propagation in the model ionosphere and
the parameters directly measured by the radar. We calibrated elevation measurements using
meteor trail echoes. The analysis revealed 37 optimal classes, with 25 frequently observed.
The analysis made it possible to choose 14 classes from them, which are confidently
separated in other variants of model training. A preliminary interpretation of 10 of them was
carried out. The dynamics of observation of various classes and their dependence on the
geographical latitude of radars at different levels of solar and geomagnetic activity were
presented, the result aligns with known physical mechanisms. The analysis showed that the
most important parameters to identify the classes are the shape of the signal ray tracing
trajectory in its second half, the ray traced scattering height and the Doppler velocity
measured by the radar.
Keywords: Radar data processing; Neural networks; Automatic classification; Data-driven
analysis; SECIRA; SuperDARN

Xiaoyuan Zhang, Shaohang Jing, Jingshu Li, Yechao Bai, Feng Yan,
Cognitive radar recognition with Kolmogorov-Smirnov test and momentum gradient descent,
Digital Signal Processing,
Volume 163,
2025,
105212,
ISSN 1051-2004,
[Link]
([Link]
Abstract: The emission parameters of cognitive radars can adaptively change according to
the environment, which poses a challenge to radar electronic countermeasures (ECM). To
counter cognitive radars, it is essential to identify the cognitive characteristics. In this paper,
a method is proposed to recognize cognitive radars with power allocation function. The
signal-to-interference-plus-noise ratio (SINR) distribution of cognitive radars is derived
through feature functions, and hypothesis test is used to identify whether the target radar
has cognitive function by designing a Kolmogorov-Smirnov (K-S) detector to recognize
adaptive optimization power allocation. Subsequently, a momentum gradient descent
algorithm is used to optimize the signal of the jamming machine to reduce type II error
probability of radar recognition. K-S detector is simulated and compared with Afriat
detector, SVM and MLP detector. Results demonstrate that the K-S detector outperforms
both the Afriat and MLP detectors in identifying cognitive radars with dynamic power
allocation functionality. At the same detection probability, the K-S detector achieves a 2 dB
improvement over the MLP detector and a 4 dB improvement over the Afriat detector.
Keywords: Cognitive radar; Electronic countermeasures (ECM); Kolmogorov–Smirnov test;
Momentum gradient descent algorithm; Afriat theorem

Yongqiang Cui, Yiyang Zhang, Di Bai, Yi Diao, Yulei Wang,


3D map and mmWave radar-based self-localization for UAVs in GNSS-denied environments,
Vehicular Communications,
Volume 57,
2026,
100986,
ISSN 2214-2096,
[Link]
([Link]
Abstract: Reliable self-localization of unmanned aerial vehicles (UAVs) in dense urban
environments remains a major challenge due to the frequent unavailability or degradation of
Global Navigation Satellite Systems (GNSS) and other radio signals. This paper presents a
robust and cost-effective method for UAV self-localization by using vision and millimeter-
wave (mmWave) radar data in GNSS-denied environments. The approach generates an initial
dense point cloud through depth estimation and semantic segmentation, which is then
geometrically refined using sparse mmWave radar point cloud. A semantic-guided clustering
method is applied to the mmWave radar point cloud to remove noise and extract key
structural elements such as walls, which are later fused with vision-based depth information.
For positioning, image matching algorithm provides coarse localization, followed by fine
registration that leverages geometric features of windows to enhance precision.
Experimental results demonstrate that the proposed method can achieve self-localization
accuracy within 0.4 m, while maintaining low system complexity and deployment cost,
offering a practical solution for UAV self-localization in GNSS-denied urban scenarios.
Keywords: Millimeter-wave (mmWave) radar; Data fusion; 3D reconstruction; Point cloud
registration; Self-localization

Wenxu Zhang, Lin An, Wencheng Yang, Zhongkai Zhao, Feiran Liu,
Open set recognition of radar specific emitter based on adversarial reciprocal point learning,
Signal Processing,
Volume 238,
2026,
110137,
ISSN 0165-1684,
[Link]
([Link]
Abstract: Radar specific emitter identification (SEI) is a key technology in electromagnetic
spectrum control. Although the emergence of deep learning has promoted the development
of SEI, there are still many shortcomings in the current research results. Most of the
traditional deep learning algorithms are applicable to closed-set identification and can only
be used when the database is complete. In addition, individual differences in radar signals
are susceptible to noise interference, but traditional denoising methods are usually
independent of the feature extraction process, making it difficult to ensure that certain
individual information is not lost. Therefore, in this paper, we propose a new radar emitter
open set recognition method called adversarial reciprocal point learning with adaptive
denoising (ARPLAD). Firstly, we design a new feature extraction network for one-dimensional
signals, which combines deep residual shrinkage network with efficient attention mechanism
to autonomously denoise signals and focus on important parts of signal features. Secondly,
we train the network using adversarial reciprocal point learning combined with center loss
to extract discriminative features with compact intraclass distances and separable interclass
distances, which can efficiently discriminate unknown signals and reduce the risk of open set
identification. The experimental results show that ARPLAD exhibits excellent performance in
different conditions, providing an effective solution for SEI in open electromagnetic
environments.
Keywords: Deep learning; Radar specific emitter identification; Open set recognition;
Adversarial reciprocal point learning

Zeyu Chen, Jian Sun, Zhengda Huan, Ziyi Zhang,


Research on Vehicle Joint Radar Communication Resource Optimization Method Based on
GNN-DRL,
Computers, Materials and Continua,
Volume 86, Issue 2,
2025,
Pages 1-17,
ISSN 1546-2218,
[Link]
([Link]
Abstract: To address the issues of poor adaptability in resource allocation and low multi-
agent cooperation efficiency in Joint Radar and Communication (JRC) systems under dynamic
environments, an intelligent optimization framework integrating Deep Reinforcement
Learning (DRL) and Graph Neural Network (GNN) is proposed. This framework models
resource allocation as a Partially Observable Markov Game (POMG), designs a weighted
reward function to balance radar and communication efficiencies, adopts the Multi-Agent
Proximal Policy Optimization (MAPPO) framework, and integrates Graph Convolutional
Networks (GCN) and Graph Sample and Aggregate (GraphSAGE) to optimize information
interaction. Simulations show that, compared with traditional methods and pure DRL
methods, the proposed framework achieves improvements in performance metrics such as
communication success rate, Average Age of Information (AoI), and policy convergence
speed, effectively enabling resource management in complex environments. Moreover, the
proposed GNN-DRL-based intelligent optimization framework obtains significantly better
performance for resource management in multi-agent JRC systems than traditional methods
and pure DRL methods.
Keywords: Graph neural network; joint radar and communication; resource allocation; multi-
agent collaboration

Nanyu Jiang, Yuyuan Fang, Lei Zhang, Chao He, Zhenhua Wu,
IFM-PointNet++: Achieving efficient radar signal waveform recognition with instantaneous
frequency measurement,
Digital Signal Processing,
Volume 168, Part B,
2026,
105507,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Efficient and robust identification of radar signal waveforms is an essential task in
electronic reconnaissance. Current deep learning-based methods can obtain satisfying
accuracy, but they are usually with high computational burden. To address the issue, this
article develops an efficient algorithm IFM-PointNet++. This algorithm transfers the radar
signal waveform recognition into the point cloud recognition task by adopting the
instantaneous frequency measurement (IFM). By integrating IFM with PointNet++ network,
our method achieves superior efficiency and accuracy. To demonstrate this effectiveness, we
conduct comprehensive comparison experiments with YOLOv8 waveform recognition on the
time-frequency images. The results demonstrate that our proposed method significantly
accelerates signal waveform recognition while maintaining high accuracy.
Keywords: Radar signal waveform recognition; Signal recognition; PointNet++; Instantaneous
frequency measurement; Deep learning

Ye Qiu, Zhenmiao Deng, Xiaohong Huang,


Unified complex-valued high-resolution frequency representation with cross-domain
attention for radar-based physiological state recognition,
Pattern Recognition,
Volume 172, Part B,
2026,
112488,
ISSN 0031-3203,
[Link]
([Link]
Abstract: High-resolution frequency domain analysis is pivotal in a wide range of critical
applications, including physiological signal processing, radar target detection, and
communication systems. In this study, we present a complex-valued neural network
designed for accurate estimation of frequency components encompassing both magnitude
and phase, the Unified-Complex High-Resolution Frequency Representation Module
(UHFreq). This method generates comprehensive high-resolution frequency domain
representations, addressing key limitations in current approaches that typically capture only
amplitude information, omit crucial phase details, and suffer from low resolution in
frequency domain outputs. Furthermore, conventional methods for physiological signal
detection and recognition require meticulous preprocessing steps, including demodulation
and filtering. In response to these challenges, we propose UHFreq-based Vital Sign Status
Detection Network (UVSD-Net), an application example of UHFreq, which classifies different
human physiological states starting from raw radar echoes. This model utilizes the UHFreq
structure as the frontend for the frequency domain representation of physiological signals
from raw radar echoes. The UVSD-Net architecture incorporates a dual-pathway design: one
pathway processes frequency domain features via UHFreq, while the other applies time
domain amplitude and phase information from the raw radar signals. Furthermore, a weight
redistribution mechanism is introduced across the different feature domains to enhance
cross-domain feature integration and interaction. This comprehensive end-to-end
framework offers a robust approach for analyzing time domain original signals and enables
effective execution of downstream tasks.
Keywords: Biomedical ignal recognition; Physiological pattern recognition; Frequency
analysis; Signal processing; Deep learning (DL)

Yanwen Bai, Jibin Zheng, Hanxing Shao, Hongwei Liu,


A dual-driven hybrid tracking architecture for radar targets based on innovation,
Information Fusion,
Volume 129,
2026,
104056,
ISSN 1566-2535,
[Link]
([Link]
Abstract: Targets such as hypersonic missiles and stealth aircraft are characterized by
complex motion patterns, strong maneuverability, and anomalous radar measurement
statistics. Although model-driven radar target tracking methods offer physical
interpretability, they suffer from their dependence on explicit prior assumptions. Data-
driven methods can theoretically approximate arbitrarily complex motions through
nonlinear mappings, but suffer from poor interpretability, vulnerability to noise during
feature extraction, and loss of low-frequency maneuvering features due to sample
imbalance. Therefore, this paper proposes a Dual-Driven Hybrid Tracking Architecture Based
on Innovation (DDHTA), which fuses the advantages of both model-driven and data-driven
approaches. First, a model-driven approach is adopted for basic state estimation, and a Dual
Condition Judgment Adjustment (DCJA) method is proposed to adaptively adjust the
measurement error variance, thereby providing a high-quality baseline estimate for the
data-driven layer and reducing the interference of anomalous noise on feature extraction.
Further, in the data-driven layer, a Dual-Scale Temporal Network (DSTNet) is designed. By
learning the mapping from the innovation to the estimation errors, it combines the
strengths of causal dilated convolution and multi-head self-attention to provide dynamic
compensation, which corrects the estimation errors of the model-driven method. Numerical
simulation results demonstrate that the proposed method enhances the algorithm’s ability
to handle target maneuvers in complex environments, achieving higher tracking accuracy
and robustness.
Keywords: Radar tracking; Deep learning; Self-attention mechanism; Causal convolution
Junkai Liu, Xinwei Qian, Lu Peng, Dan Lou, Yiwen Li,
TEDR: A spatiotemporal attention radar extrapolation network constrained by optical flow
and distribution correction,
Atmospheric Research,
Volume 311,
2024,
107702,
ISSN 0169-8095,
[Link]
([Link]
Abstract: In recent years, deep learning has been widely applied to meteorological radar
extrapolation due to the shortcomings of traditional optical flow methods in predicting the
genesis and dissipation of radar echoes. However, it still faces challenges in addressing issues
of clarity and overall intensity attenuation caused by uncertainty. This study implemented a
dual-path spatiotemporal attention network that integrates optical flow techniques by
employing intra-frame static attention and inter-frame dynamic attention, which could
simulate motion fields and the overall intensity distribution of radar echoes separately. Our
approach effectively resolve the issues of systematic intensity attenuation and clarity
degradation introduced by deep learning methods. Through the comparisons of key metrics
such as MSE, SSIM, CSI20, CSI30, and CSI40, the results demonstrated significant
improvements over traditional approaches, particularly in CSI30 and CSI40, where the
metrics improved by more than 35 %.
Keywords: Radar extrapolation; Convective weather forecasting; Machine learning; Optical
flow

Li Qiusheng, Zhu Huajuan,


Target classification with low-resolution radars based on cyclic bispectrum and improved
ACGAN,
Measurement,
Volume 259, Part B,
2026,
119715,
ISSN 0263-2241,
[Link]
([Link]
Abstract: To address the challenges of insufficient generalization and high noise sensitivity in
low-resolution radar target recognition under limited-sample conditions, this paper
proposes a joint optimization framework integrating cyclic bispectral analysis and an
improved Auxiliary Classifier Generative Adversarial Network (ACGAN). First, a third-order
cyclic cumulant spectral model is designed to extract modulation-specific signatures of
aircraft targets in the cyclostationary domain, effectively suppressing both Gaussian and
non-Gaussian noise while preserving discriminative features that are robust to low SNR
conditions (maintaining 92.7 % accuracy at 0 dB). Second, an enhanced ACGAN architecture
is developed by incorporating self-attention mechanisms and Wasserstein distance
optimization with gradient penalty, with spectral normalization and dynamic gradient
penalties introduced to stabilize training dynamics and improve synthetic sample fidelity.
Extensive experiments on a real-world dataset collected by a certain Chinese-made VHF-
band radar demonstrate that the proposed method achieves state-of-the-art performance,
with average recognition accuracies of 98.46 % and 98.52 % for approaching and departing
targets in complex noise environments, respectively, alongside a Kappa coefficient exceeding
0.97. Comprehensive comparisons with traditional methods (wavelet, HOS, FrFT), GAN
variants (WGAN-GP, AFGAN + ResNet, Diffusion-GAN), VAE-based approaches, few-shot
learning models (ProtoNet, MatchingNet), and a modern Vision Transformer (ViT) baseline
show consistent improvements of 1.38–12.19 % in accuracy. Ablation studies validate the
contributions of key components, where the self-attention module and Wasserstein
optimization improve accuracy by 1.27 % and 0.97 %, respectively. Furthermore, embedded
platform tests confirm the framework’s feasibility for real-time deployment (inference
time < 15 ms/sample), offering a robust solution for resource-constrained radar systems.
This work highlights the efficacy of unifying physics-inspired feature extraction with
stabilized deep generative models for practical radar recognition.
Keywords: Radar target recognition; Cyclic bispectral analysis; Auxiliary classifier generative
adversarial network (ACGAN); Self-attention mechanism; Few-shot learning

Rui Wan, Weigang Meng, Tianyun Zhao, Wei Lu,


Overcoming radar sparsity and cross-view misalignment: A sparse-to-sparse fusion paradigm
for robust 3D object detection,
Digital Signal Processing,
Volume 171,
2026,
105831,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Millimeter-wave radar-camera fusion provides a cost-effective alternative to LiDAR
for 3D perception in autonomous driving. However, its potential is constrained by two
limitations: (1) Previous bird’s-eye view fusion methods struggle to accurately align and fuse
cross-modal features, while the fusion strategy incurs substantial computational redundancy
from background processing. (2) The extreme sparsity of radar points (typically < 5 % LiDAR
density) hinders robust geometric measurement. To address these challenges, we propose a
radar-camera fusion 3D detection framework that redefines cross-modal interaction by
transitioning from dense fusion to sparse-to-sparse paradigm. This transformation is
initiated by generating spatially-aware 3D object queries from images and radar sweeps-
leveraging image-derived seed points with radar depth to anchor queries to objects via a
perspective-guided object query generator. Moreover, we introduce adaptive radar pillar
diffusion within foreground regions to mitigate radar sparsity, allowing object queries to
capture geometric information from diffused pillars. Additionally, to maximize image
semantic clues, we further refine object boxes through image keypoint feature aggregation
using a keypoint-aware object refinement module. Our framework not only circumvents
traditional fusion bottlenecks but also achieves real-time inference at 23.4 FPS. Evaluated on
nuScenes dataset, it demonstrates competitive detection performance (65.0 % NDS, 57.8 %
mAP) and tracking precision (58.3 % AMOTA and 0.687m AMOTP). By overcoming sparsity
constraints and improving cross-modal fusion, this work establishes a new paradigm for
robust perception systems.
Keywords: Autonomous driving; 3D object detection; Deep learning; Radar-camera fusion;
Cross-attention

Yunfeng Fang, Zheng Tong, Tianqing Hei, Siqi Wang, Tao Ma,
Deep learning applications in ground-penetrating radar inversion: A review,
Measurement,
Volume 258, Part D,
2026,
119399,
ISSN 0263-2241,
[Link]
([Link]
Abstract: The complex nonlinear relationship between the subsurface medium and ground-
penetrating radar signals results in the pervasive ill-posedness and non-uniqueness of
conventional inversion methods. Deep learning, with its powerful feature extraction
capabilities and advantages in modeling complex nonlinear relationships, has unique
strengths in handling complex signals and nonlinear problems, making it especially suitable
for GPR inversion tasks. This paper reviews the latest applications of deep learning in GPR
inversion, summarizing the application strategies of deep learning from two perspectives:
data-driven and data-physics hybrid-driven. Commonly used model architectures and their
performance in signal feature extraction, multi-scale information fusion, and data
preprocessing are discussed, along with the application of various loss functions in inversion
tasks. Finally, current challenges, such as limited model generalization, model dependence
on the dataset and computational efficiency constraints, are discussed, and potential future
research directions are proposed to further advance deep learning in GPR inversion.
Keywords: Ground-penetrating radar; Deep learning; Inversion

Philipp Reitz, Tobias Veihelmann, Norman Franchi, Maximilian Lübke,


Dual radar vision: A feature fusion approach for advanced object detection in IoT radar
networks,
Machine Learning with Applications,
Volume 21,
2025,
100703,
ISSN 2666-8270,
[Link]
([Link]
Abstract: 60GHz radar technology is one of the most promising movement detector
solutions for Internet of Things (IoT) applications. However, challenges remain in accurately
classifying different objects and detecting small objects in a multi-target scenario. This work
investigates whether sensor fusion between multiple radars can enhance object detection
and classification performance. A one-stage detection architecture, designed based on the
features of the latest YOLO generations, is used to perform fusion based on range-Doppler
(RD) maps of two non-coherent spatially separated radars. A complete physical 3D
propagation simulation using ray tracing evaluates the fusion methods. This approach
enables precise ground truth, as all unprocessed signal components are known, and
guarantees a consistent, error-free reference. Results demonstrate that dynamic, attention-
based fusion significantly improves detection and classification compared to static fusion in
homogeneous and heterogeneous radar setups.
Keywords: Data fusion; Deep learning; FMCW radar; IoT; Radar networks; Ray tracing; YOLO

Mingyang Du, Ping Zhong, Xiaohao Cai, Daping Bi, Aiqi Jing,
Robust Bayesian attention belief network for radar work mode recognition,
Digital Signal Processing,
Volume 133,
2023,
103874,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Understanding and analyzing radar work modes play a key role in electronic
support measure system. Many classifiers, for example those based on convolutional neural
network (CNN) and recurrent neural network (RNN), are available for recognizing radar work
modes as well as emitter types from their waveform parameters. However, the performance
of these methods may suffer significantly when confronting different types of signal
degradation, e.g., measurement error, lost pulse and spurious pulse. To tackle this issue, we
in this paper develop a Bayesian attention belief network (BABNet) based on Bayesian neural
networks in which the probability distribution over weights can help to enhance the model
robustness for corrupted data. In particular, we adopt pre-trained CNN as the Bayesian
inference prior. This not only accelerates the convergence speed, but also avoids the training
process getting stuck in bad local minima. Meanwhile, instead of using RNNs which are
difficult to be implemented in parallel, the combination of padding operation and attention
module in the proposed BABNet enables CNN, as the backbone, to process sequential data
with variable length. Extensive experiments are conducted to demonstrate the recognition
capability and robustness of the BABNet in different environments.
Keywords: Radar work mode; Pulse descriptor word; Attention mechanism; Bayesian neural
network; Robustness; Recognition

Zhigang Cheng, Zhizhou He, Peng Pan,


3D reconstruction of subsurface pipes and cavities using ground penetrating radar based on
deep learning,
NDT & E International,
Volume 158,
2026,
103579,
ISSN 0963-8695,
[Link]
([Link]
Abstract: Detecting subsurface pipes and cavities is important in urban infrastructure
management, but existing methods struggle to accurately reconstruct the 3D shapes of deep
subsurface objects. This study pioneers a new paradigm for this task by reformulating the ill-
posed permittivity regression problem as a 3D semantic segmentation problem. A novel
neural network, 3DReconNet, to predict the material type of each subsurface voxel from
ground penetrating radar (GPR) data was proposed. This approach leverages the intrinsic
relationship between material composition and reflected signal intensity to simultaneously
recover both geometry and material properties. A dataset of 3150 synthetic cases was
generated using full-scale simulation models and a Markov model-based algorithm to
simulate irregular cavities. The 3DReconNet adopts a U-shaped architecture and
incorporates residual connections to reduce information loss. The network is trained using
the Dice Loss function regularized with total variation (TV) constraints, which enhances
geometric consistency and reconstruction accuracy. The proposed method was validated
using both simulated and experimental data, and the qualitative as well as quantitative
results confirmed its effectiveness, robustness, and generalizability.
Keywords: Ground penetrating radar (GPR); Subsurface pipes; Subsurface cavities; Deep
learning; 3D reconstruction

Hongping Zhou, Lei Wang, Minghui Ma, Zhongyi Guo,


Compound radar jamming recognition based on signal source separation,
Signal Processing,
Volume 214,
2024,
109246,
ISSN 0165-1684,
[Link]
([Link]
Abstract: To deal with various jamming signals based on digital radio frequency memories, a
compound jamming signal recognition method based on source signal separation is
proposed. In order to overcome label limitation of the supervised learning method in the
recognition process, this paper puts forward "Separation + Recognition" strategy. Firstly, the
received signals of multiple channels are preprocessed, and the number of signal sources is
analyzed through the single-source detection algorithm. Then the received compound
jamming signals are processed by source separation, and the independent single jamming
signals can be obtained. On this basis, a fast signal compensation algorithm is added to
compensate different source signals at overlapping time-frequency points, which effectively
increases the integrity of the separated signals. The separated jamming signals are then put
into the convolutional neural network for recognition, and the specific jamming types in the
compound jamming signals can be obtained. It has been proved that the recognition
accuracy of five kinds of compound jamming exceeds 90% when the jamming-to-noise ratio
is 0 dB.
Keywords: Active-jamming recognition; Deep learning; Compound jamming recognition;
Neural networks

Yan Cheng, Ke Mei, Hao Zeng,


Recognition of intrapulse modulation mode in radar signal with BRN-EST,
Journal of Electronic Science and Technology,
Volume 23, Issue 4,
2025,
100336,
ISSN 1674-862X,
[Link]
([Link]
Abstract: Neural network-based methods for intrapulse modulation recognition in radar
signals have demonstrated significant improvements in classification accuracy. However,
these approaches often rely on complex network structures, resulting in high computational
resource requirements that limit their practical deployment in real-world settings. To
address this issue, this paper proposes a Bottleneck Residual Network with Efficient Soft-
Thresholding (BRN-EST) network, which integrates multiple lightweight design strategies and
noise-reduction modules to maintain high recognition accuracy while significantly reducing
computational complexity. Experimental results on the classical low-probability-of-intercept
(LPI) radar signal dataset demonstrate that BRN-EST achieves comparable accuracy to state-
of-the-art methods while reducing computational complexity by approximately 50 %.
Keywords: Attention mechanism; Convolutional neural network; Low probability of intercept
radar; Recognition of intrapulse modulation

Hao Yang, Shirong Zhou, Liyan Liu, Zhong Zhou,


A fast and accurate detection model of internal defects in tunnel lining for ground
penetrating radar image data,
Advanced Engineering Informatics,
Volume 68, Part C,
2025,
103812,
ISSN 1474-0346,
[Link]
([Link]
Abstract: Defects in tunnel linings accelerate structural deterioration, reduce service life, and
pose serious safety risks. Existing algorithms for detecting defect signals in ground-
penetrating radar (GPR) images often struggle to balance accuracy and efficiency, with
limited capacity to extract meaningful features. To address these limitations, this paper
proposes a lightweight algorithm, MGD-DETR, for accurate recognition of internal tunnel
lining defects, using RT-DETR as the base model. First, a Multi-HGNet backbone feature
extraction network is introduced to reduce model size (MS) and enhance dynamic fusion and
interaction between feature layers, thereby improving feature extraction. Second, the
lightweight convolution module GSConv replaces standard convolution operations to reduce
the parameter count. Third, a dual attention module (DAM) is integrated to dynamically
adjust spatial and channel feature weights, improving the model’s generalization
performance. Five models—RT-DETR, YOLO-LD, YOLOv10, YOLOv11, and SSD—were used for
comparative evaluation. Experimental results show that MGD-DETR outperforms the other
models across all metrics, achieving a mean average precision (mAP) of 0.834, mean F1
score (mF1) of 0.818, MS of 26.9 M, and frames per second (FPS) of 91.2f/s, enabling fast
and accurate recognition of defect signals and facilitate subsequent deployment into tunnel
detection mobile devices.
Keywords: Tunnel engineering; Lining defects; Deep learning; Ground-penetrating radar
images

Tao Chen, Boyi Yang, Limin Guo, Lei Zhan,


Radar signal sorting via GraphSAGE convolution with instance-aware clustering,
Digital Signal Processing,
Volume 165,
2025,
105336,
ISSN 1051-2004,
[Link]
([Link]
Abstract: This study presents a graph-based radar signal sorting algorithm addressing scale
transformation challenges in image-based intelligence sorting. We develop a variable-
dimensional framework representing pulse description words as graph nodes, with carrier
frequency and other parameters as initial node features. Through iterative neighbor
sampling and feature aggregation, the network learns relational patterns between pulses via
trainable parameter matrices. Moreover, a novel magnetic loss function combining intra-
cluster attraction and inter-cluster repulsion provides weak supervision for feature space
optimization. Compared to the sorting algorithms based on image segmentation, the
proposed algorithm avoids issues related to scale transformation and pulse congestion.
Unlike algorithms based on recurrent neural networks or autoencoders, this algorithm
utilizes inductive aggregation to investigate the relationships between pulses, offering a
novel learning approach. Simulation results demonstrate that, in the scenarios involving
densely interleaved pulse sequences from multiple radar emitters, the proposed network
effectively deinterleaves the sequences using a magnetic cluster loss function. Even in
complex electromagnetic environments, with high missing pulse rates, high spurious pulse
rates, and significant measurement errors, the algorithm demonstrates strong robustness.
The algorithm demonstrates consistent alignment with expert analysis across field tests,
showing particular advantages in real-time processing and electromagnetic complexity
adaptation.
Keywords: Radar signal sorting; Weak supervision; Graph convolution network; Instance
segmentation; Magnetic loss; Deep cluster

Li Li, Jiaxin Shi, Tianshuang Qiu, Mingyan He,


AMTCC-PARAFAC: A convergent tensor framework for DOD–DOA–Doppler estimation in
bistatic MIMO radar under impulsive noise,
Physical Communication,
Volume 73,
2025,
102823,
ISSN 1874-4907,
[Link]
([Link]
Abstract: To address the severe performance degradation of parameter estimation in
impulsive noise environments, this paper proposes a novel tensor decomposition framework
based on adaptive maximum total complex correntropy (AMTCC) for robust joint parameter
estimation in bistatic MIMO radar systems. In the proposed method, for the first time, the
AMTCC criterion to reconstruct the parallel factor (PARAFAC) cost function, marking the
initial integration of complex correntropy theory with tensor decomposition. To optimize
performance, we incorporate an adaptive kernel bandwidth selection mechanism that
dynamically adjusts to impulsive noise environments, significantly enhancing parameter
estimation accuracy. Then, we develop a novel PARAFAC algorithm based on the AMTCC and
apply it to target parameter estimation in bistatic MIMO radar. The proposed algorithm
eliminates FLOS methods’ need for prior noise knowledge while concurrently suppressing
complex noise components and enabling automatic parameter pairing. Furthermore, we
provide theoretical analyses: (1) analyzed complex correntropy’s impulsive noise
suppression via nonlinear kernels, (2) proved the boundedness of the AMTCC cost function,
(3) analyzed robustness advantages of AMTCC-PARAFAC over existing decompositions and its
methodological positioning, (4) derived parameter Cramér–Rao bounds under α-stable
noise, and (5) determined target identifiability limits through factor matrices’ Kruskal rank
and dimension constraints. Simulation results demonstrate that the proposed algorithm
effectively suppresses complex-domain impulsive noise while eliminating FLOS methods’
dependence on prior noise statistics, achieving superior parameter estimation accuracy and
automatic pairing capability in α-stable noise environments.
Keywords: Impulsive noise; Bistatic MIMO radar; Maximum total complex correntropy;
Adaptive kernel bandwidth; PARAFAC

Pianzhang Duan, Li Wang, Cheng Fang, Ziying Song, Ming Gao, Mo Zhou, Ying Li, Yibo Zhang,
Wei Fan, Bin Xu,
Global relationship awareness 3-dimensional object detection using 4-dimensional radar,
Engineering Applications of Artificial Intelligence,
Volume 164, Part B,
2026,
113318,
ISSN 0952-1976,
[Link]
([Link]
Abstract: 4D (4-dimensional) radar sensing technology is essential for high-precision
autonomous driving perception systems, as its superior detection capabilities at increased
distances, compared to traditional LiDAR (Light Detection and Ranging). However, due to the
sparsity of point clouds and the low resolution of millimeter-wave radar, voxel-based
methods may fail to detect distant or closely adjacent objects, leading to inadequate
detection accuracy. To mitigate the accuracy issues arising from the sparse nature of point
clouds in such scenarios, we propose a novel object detection network: GRA-Net (Global
Relation-Aware object detection Network). By leveraging a self-attention mechanism, GRA-
Net effectively learns critical features from each radar pillar, enhancing the network’s
capacity to capture relevant information about nearby objects. Furthermore, we introduce a
global perception module that integrates key features within the pillars and global features,
mitigating the impact of point cloud sparsity, particularly in distant regions. We conducted a
series of experiments to evaluate the performance of GRA-Net. On the Astyx HiRes 2019
dataset, our method achieved 33.63 mAP (mean Average Precision) and 43.93 mAP at the
moderate level; On the View-of-Delft dataset, our method achieved 47.74 mAP in the entire
annotated area and 69.25 mAP in the driving corridor area.
Keywords: 4-dimensional radar; 3-dimensional object detection; Self-attention mechanism;
Autonomous driving

Liangang Qi, Hongzhuo Chen, Qiang Guo, Shuai Huang, Mykola Kaliuzhnyi,
GLS: A hybrid deep learning model for radar emitter signal sorting,
Digital Signal Processing,
Volume 161,
2025,
105117,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Radar emitter signal sorting is a pivotal aspect of radar reconnaissance signal
processing. The increasing density of the electromagnetic environment in modern radar
pulse streams, coupled with the growing complexity and variability of operational modes
and signal forms, results in extremely limited reference data. Consequently, most existing
sorting methods fall short of meeting the performance requirements of modern electronic
warfare. To enhance sorting performance under conditions of limited samples and labeled
data, this paper proposes a radar emitter signal sorting model based on ResGCN-BiLSTM-SE
(GLS). Firstly, we propose a novel adaptive weighted adjacency matrix construction method
that aggregates multi-scale information of local and global features. Based on this, for GLS
networks, the graph convolutional network (ResGCN) is combined with the bidirectional long
short-term memory (BiLSTM) network. The GCN is employed to extract attribute features
from interleaved radar pulse sequences, while the BiLSTM is utilized to deeply capture the
temporal dependence in interleaved pulse sequences after feature extraction. Finally, an
improved squeeze-and-excitation (SE) module is applied to perform weighted fusion of
critical channel information from both spatial and temporal features. Simulation results
demonstrate that the proposed method not only achieves higher accuracy under small
sample conditions compared to existing methods, but also exhibits strong robustness in
challenging scenarios involving measurement errors, missing pulses, and spurious pulses.
Keywords: Radar emitter signal sorting (RESS); Adaptive weighted adjacency matrix; GLS
model; Features fusion

Keyu Pan, Wei-Ping Zhu, Mojtaba Hasannezhad,


Self-attention CNN based indoor human events detection with UWB radar,
Journal of the Franklin Institute,
Volume 361, Issue 14,
2024,
107090,
ISSN 0016-0032,
[Link]
([Link]
Abstract: In the era of smart homes and healthcare automation, the ability to accurately
monitor and detect indoor human activities is paramount. Ultra-wideband (UWB) radar has
emerged as a promising means for event detection, given its non-invasive nature and easy
deployment in diverse environments. However, despite the advances in radar-based event
detection, challenges remain, such as distinguishing between similar events like falls and
rapid sitting. To address these challenges, for the first time, we propose an impulse radio-
ultrawideband (IR-UWB) radar system to collect over ten thousand radar echo signals of
eight similar actions from different angles and design a self-attention-based low-complexity
convolutional neural network (CNN) model for event classification. The model leverages
global correlations in radar signal spectrograms to efficiently extract features. A comparative
simulation study is conducted to evaluate the detection accuracy of the proposed model and
some of the existing methods based on different dataset sizes and CNN configurations.
Moreover, the influence of different self-attention structures on precision and model
parameter count is analyzed. Our findings reveal that the proposed self-attention-based CNN
model significantly outperforms other traditional machine learning techniques while
maintaining a low level computational complexity.
Keywords: UWB radar system; CNN; Self-attention

Long Jin, Jiamin Pu, Hongjian Li,


FMCW radar-based heartbeat recognition using SE-DenseNet for vital signs monitoring,
AEU - International Journal of Electronics and Communications,
Volume 200,
2025,
155906,
ISSN 1434-8411,
[Link]
([Link]
Abstract: With the aging of society and people’s increasing concern about their health, non-
contact vital signs measurement provides new possibilities for home health monitoring. In
recent years, studies have shown that millimeter wave (mmW) radar has high sensitivity in
heart rate (HR) monitoring. However, traditional signal processing methods are susceptible
to environmental interference and they are difficult to accurately reconstruct heartbeat
signals. This paper proposes a heartbeat signal reconstruction method based on the SE-
DenseNet deep learning (DL) model. The model combines dense connections (DenseNet)
with the Squeeze-and-Excitation (SE) blocks to automatically extract and enhance key
physiological features in the two-dimensional feature matrix generated by radar signals,
thereby improving the accuracy of signal reconstruction. The experimental results
demonstrate that the proposed method achieves a HR estimation accuracy exceeding
97.68% in monitoring experiments at distances within 1.3 m, and its strong robustness has
been confirmed through tests with different subjects and during mild movement.
Keywords: Vital signs monitoring; Millimeter-wave (mmW) radar; Heart rate (HR); Signal
reconstruction; Deep learning (DL)

Ayesha Jabbar, Muhammad Kashif Jabbar, Asif Jabbar, Ahmed S. Almasoud, Faijan Akhtar,
Maryam Zulfiqar, Tariq Mahmood, Amjad Rehman,
Enhancing radar tracking accuracy using combined Hilbert transform and proximal gradient
methods,
Results in Engineering,
Volume 24,
2024,
103479,
ISSN 2590-1230,
[Link]
([Link]
Abstract: Accurate radar tracking is crucial in defense, navigation, and surveillance
applications, where high precision and resilience to noise are essential. Traditional radar
tracking techniques, such as Kalman Filters and Particle Filters, often struggle with
performance limitations in noisy and non-linear environments, leading to inaccuracies in
target tracking. To address these challenges, we propose a hybrid radar tracking approach
combining the Hilbert Transform with the Proximal Gradient Method within a convex
optimization framework. This combination leverages the Hilbert Transform's signal
enhancement capabilities with the Proximal Gradient Method's optimization strength,
improving accuracy and robustness under challenging conditions. Experimental results
demonstrate that the proposed method achieves a 23% reduction in Mean Squared Error
(MSE) and a 20% increase in tracking accuracy compared to conventional methods,
alongside a Signal-to-Noise Ratio (SNR) of approximately 18.3 dB, indicating superior noise
resilience. While the hybrid method offers significant improvements, it does involve
increased computational complexity and may be sensitive to initial parameter settings,
requiring careful tuning for optimal performance. Nevertheless, this method represents a
promising advancement over traditional techniques, providing a more accurate and resilient
solution for modern radar tracking applications.
Keywords: Radar tracking; Proximal gradient method; Hilbert transform; Trajectory
estimation; Convex optimization

Hui Wang, Qinghua Liu, Lijun Zhou,


Underground target localization method for ground penetrating radar based on deep
learning,
Measurement,
Volume 253, Part B,
2025,
117647,
ISSN 0263-2241,
[Link]
([Link]
Abstract: To tackle the challenge of subsurface target localization under interference in field
scenarios, a novel two-level cascade network referred to as dual cascade is proposed. The
first level, Cascade-1, is a deep feature extraction network designed to extract and eliminate
direct wave interference signals. On this basis, Cascade-2 is developed using domain
knowledge from ground penetrating radar as prior information, and it incorporates an
attention mechanism along with a feature fusion strategy to enhance the accuracy of target
feature hyperbola detection. Subsequently, the least squares method is employed to fit the
feature hyperbola, and location estimation is performed based on geometric equations. The
proposed cascade network model has demonstrated superior performance compared to
other algorithms, such as column-connection clustering algorithm, YOLOv9, and Faster R-
CNN, in terms of the composite metric F1, which validates the model’s effectiveness in
extracting the feature hyperbola. Additionally, the proposed localization method has
exhibited greater accuracy than the conventional full waveform inversion algorithm.
Keywords: Ground Penetrating Radar; Buried Target Location; Domain Knowledge; Deep
Learning; Cascade Network

Dunlu Peng, Meiling Chen, Yiqin Zhang, Zekun Tian,


Enhanced optic-flow extrapolation for Doppler radar nowcasting with Dynamic Weight
Attention,
Expert Systems with Applications,
Volume 267,
2025,
126168,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Doppler radar echo extrapolation is an important method for extreme weather
forecasting. However, traditional optical flow methods lack learnable components and are
not suitable for complex atmospheric changes. Consequently, researchers have turned to
deep neural networks for prediction. Yet, the predictions from this approach often suffer
from issues such as mean reversion and a lack of small- and medium-scale structure. This
paper proposes a novel approach that combines optical flow methods with deep neural
networks. By introducing an artificially defined momentum weight matrix based on prior
assumptions, predictions for any future time distance are generated from full-scale optical
flow. Additionally, we propose a full-scale advection extractor, leveraging the continuity of
distribution in mesoscale and small-scale atmospheric systems and focusing on the long- and
short-distance relationships within the contour surface distribution sequence, which
improves the prediction accuracy of fine-scale advection. The experimental results show
that, compared with other advanced methods, the proposed method demonstrates
advantages in predicting extreme radar echoes and maintaining the echo structure.
Specifically, it achieved an improvement of 24.1% and 21.3% on key indicators such as CSI
and HSS, respectively, and reached 0.948 on the SSIM. Building on this, the inference speed
of our method is comparable to other deep learning approaches, being 3.25 times faster
than flow-based methods.
Keywords: Nowcasting; Optical flow; Neural network; Doppler radar

Ruma Adhikari, Alok Bhardwaj,


Development of flood detection framework integrating Synthetic Aperture Radar
polarimetry and machine learning for semi-urban vegetation systems,
Journal of Environmental Management,
Volume 397,
2026,
128208,
ISSN 0301-4797,
[Link]
([Link]
Abstract: Semi-urban vegetation system includes vegetation partly managed by humans and
partly growing naturally. It plays a vital role in ecosystem stability but are vulnerable to
floods across the world highlighting the need for strategies to enhance resilience to climate
change. However, detecting floods is difficult using Earth Observation due to complex
scattering between radar signals and varied surface conditions. To address this, Synthetic
Aperture Radar (SAR) polarimetry provides information to distinguish scattering patterns of
flooded from non-flooded areas. The study proposes a flood detection methodology using
Sentinel-1 SAR aimed at combining the Degree of Polarization (DOP) and Linear Polarization
Ratio (LPR) derived from Stokes parameters, and Eigenvalues of the SAR covariance matrix.
The proposed Flood Index (FI) integrates both amplitude and phase, unlike Normalized
Difference Flood Index (NDFI) and VH/VV ratio that use only intensity data; the phase data
helps separate smooth flooded surfaces from rough land or vegetation. A Random Forest
model trained on the FI with bootstrap sampling detects flood extents accurately in Japan
(2019 Typhoon Hagibis), India (2023 Delhi flood), and Greece (2023 Larissa flood). The
model achieves F1 scores between 0.81 and 0.86 and Intersection over Union scores
between 0.70 and 0.76. The proposed model is better than Otsu and NDFI across all study
sites by maintaining lower False Negative Rate (0.09–0.17) and moderate False Positive Rate
(0.19–0.39). Better transferability of the trained model is achieved across different flooded
areas for scalable flood management in semi-urban vegetation areas.
Keywords: Flood detection; Semi-urban vegetation system; SAR polarimetry; Machine
learning

Md. Alamgir Hossain,


FED-GEM-CN: A federated dual-CNN architecture with contrastive cross-attention for
maritime radar intrusion detection,
Array,
Volume 27,
2025,
100456,
ISSN 2590-0056,
[Link]
([Link]
Abstract: The escalating complexity of maritime operations and the integration of advanced
radar systems have heightened the susceptibility of maritime infrastructures to sophisticated
cyber intrusions. Ensuring resilient and privacy-preserving intrusion detection in such
environments necessitates innovative solutions capable of learning from distributed,
heterogeneous data sources without compromising sensitive information. This study
introduces FED-GEM-CN, a novel federated learning framework designed explicitly for
maritime radar intrusion detection. The proposed architecture integrates dual parallel
convolutional neural network (CNN) pipelines to independently process network and radar
modality features, which are subsequently fused via a multi-head cross-attention
mechanism to capture intricate inter-modal dependencies. To enhance feature
discriminability, a supervised contrastive learning paradigm is incorporated, while a gradient
episodic memory (GEM) buffer strategically retains challenging instances to bolster model
robustness against hard-to-detect intrusions. Operating under a federated learning scheme,
FED-GEM-CN facilitates collaborative model optimization across distributed radar nodes,
preserving data locality and mitigating privacy risks inherent in centralized approaches.
Experimental evaluations conducted on a comprehensive real-world maritime radar dataset
reveal that FED-GEM-CN achieves superior performance, attaining an overall accuracy
exceeding 99 % and macro F1-scores above 0.97 across federated rounds, with convergence
typically observed within 15 communication iterations. These findings substantiate the
efficacy of the proposed system in delivering robust, energy-efficient, and privacy-aware
intrusion detection tailored to the constraints of maritime radar networks. The approach
underscores a significant advancement toward deploying intelligent, distributed
cybersecurity solutions within critical maritime infrastructures.
Keywords: Federated learning for maritime cybersecurity; Dual-CNN architecture; Cross-
attention mechanism; Intrusion detection system; Privacy-preserving learning; Multi-head
attention networks; Marine radar security

Jingpeng Gao, Sisi Jiang, Xiangyu Ji, Chen Shen,


Cross-domain prototype similarity correction for few-shot radar modulation signal
recognition,
Signal Processing,
Volume 223,
2024,
109575,
ISSN 0165-1684,
[Link]
([Link]
Abstract: The new classes of radar signals are increasingly difficult to acquire under non-
cooperative environments, which makes it difficult to support convolutional neural network
training with limited labeled samples. The few-shot learning (FSL) methods have shown
great performance in classification with limited labeled samples, but the FSL methods ignore
that the class distributions between the new and original tasks are significantly different,
resulting in a massive challenge in identifying new radar signals. To solve this problem, a
few-shot radar modulation signal recognition method based on cross-domain prototype
similarity correction (CDPSC) is proposed. Specifically, a residual feature tokenizer
transformer (RFTT) model embedded with a pooling token generation block is designed to
focus on the important features and improve the ability to represent samples. Meanwhile,
the proposed domain prototype similarity mapping (DPSM) strategy adaptively learns the
class mapping, reduces the inter-domain difference through feature distribution alignment,
and effectively corrects the target domain prototypes. In addition, we introduce a sample
prototype embedding (SPE) strategy in the training phase, which can reduce the intra-class
distance and increase the inter-class distance. Experimental results demonstrate that the
CDPSC method is superior to typical FSL methods in recognition accuracy under different
sample numbers.
Keywords: Cross-domain; Few-shot learning; Prototype similarity correction; Radar
modulation signal recognition

Hairui Zhu, Shanhong Guo, Weixing Sheng,


RDJCNN: A micro-convolutional neural network for radar active jamming signal classification,
Engineering Applications of Artificial Intelligence,
Volume 123, Part C,
2023,
106417,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Timely awareness of jamming situations and classification of jamming categories
are vital for radars to suppress jamming, ensure viability, and maintain functions in complex
electromagnetic environments. To satisfy strict time requirements for radars on embedded
devices, a micro-dynamic convolutional neural network for jamming signal classification is
proposed in this paper. The proposed network takes the range-Doppler distribution obtained
from built-in radar signal processing as input. The proposed data augmentation algorithm,
together with the attention mechanism and the efficient convolutional architecture,
improves the generalization capability and reduces the computational complexity. In
addition, we propose a dynamic depth mechanism based on a task difficulty evaluator that
enables the network to be adjusted automatically and further reduces the average
computational complexity of classification. Simulation results verify the advantages of our
approach in size, accuracy, and efficiency. The proposed network achieved 98.82% and
85.00% top-1 accuracy in two datasets with only 1.73 M multiply–accumulate operations.
Keywords: Convolutional neural networks; Radar signal-processing; Jamming situation
awareness; Jamming signal classification

Chengming Zong, Zhizhong Lu, Yanbo Wei,


Wave height measurement based on feature fusion extracted from marine radar images,
Journal of Sea Research,
Volume 208,
2025,
102626,
ISSN 1385-1101,
[Link]
([Link]
Abstract: Obtaining wave information near the ship's location is not only crucial for ensuring
navigation safety, but also an important basis for meteorological forecasting and disaster
prevention, which is of great significance for marine engineering and scientific research. To
further improve the estimation accuracy of significant wave height (SWH) from non-
coherent X-band marine radar image, a wave height measurement method is proposed
based on the feature fusion and radial basis function (RBF) network. The wave slope and
signal-to-noise ratio (SNR) extracted from radar image and environmental factors such as
wave direction and wind information are introduced to establish the feature vector as the
input of RBF network. By training the RBF network model, accurate estimation of SWH is
achieved. The measured radar data is used for experimental verification, and the
experimental results show that the feature fusion method proposed has higher accuracy and
reliability in calculating SWH than the shadow statistical method and the traditional SNR-
based method, when the environmental factor of wind information and wave direction is
considered. The correlation coefficient between buoy record and estimated SWH
approaches 0.92, and the root mean square error deceases to 0.21 m.
Keywords: Marine radar image; Wave slope; Feature fusion; Radial basis function network;
Wave height

Hua Wang, Qiangyu Zeng, Hao Wang, Jianxin He, Tiantian Yu, Guangpu Liu,
Temporal super-resolution reconstruction of weather radar echoes using a deep learning
approach,
Expert Systems with Applications,
Volume 300,
2026,
130189,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Severe convective weather events are characterised by rapid evolution and high
destructive potential, requiring weather radars to provide observations with high temporal
resolution. However, current S-band weather radar systems, constrained by their volumetric
scanning strategies, often fail to capture the rapidly changing features of these systems
promptly. To address this limitation, we propose EMAIRA-VFI, a deep learning–based
method for temporal super-resolution reconstruction of radar echoes, which enhances the
temporal resolution of radar data to meet the demands of severe convective weather
monitoring. By introducing an inter-frame attention mechanism, the proposed method
effectively fuses spatiotemporal features from sequential radar echoes, enabling accurate
modelling of dynamic weather evolution and the generation of continuous, high-temporal-
resolution radar echoes. Compared with conventional temporal interpolation methods,
EMAIRA-VFI demonstrates significant improvements in both interpolation accuracy and the
preservation of fine-scale meteorological structures. Experimental results show that the
model not only enhances the capability of S-band radars in monitoring rapidly evolving
weather events but also provides a new perspective for spatiotemporal fusion and the
intelligent application of radar data. We have open-sourced the code for this work at
[Link]
Keywords: Temporal super-resolution; Radar echo; Inter-frame attention mechanism

Zhao Liu, Jingyi Wang, Huachun Tan, Fan Ding,


Long-term traffic flow prediction via spatiotemporal sequence reconstruction on highway
networks,
Expert Systems with Applications,
Volume 304,
2026,
130845,
ISSN 0957-4174,
[Link]
([Link]
Abstract: In recent years, graph neural networks with attention mechanisms have achieved
remarkable progress in traffic flow prediction. However, traffic data often exhibit a
composite structure in which stable, low-frequency travel patterns are superimposed on
occasional, high-frequency trips. This multi-scale frequency characteristic makes it
challenging to effectively identify and disentangle the underlying components, thereby
constraining the accuracy of long-term forecasting. To overcome this limitation, this paper
proposes a Long-term Traffic Flow Prediction framework based on Spatiotemporal Sequence
Reconstruction on Highway Networks (LTFP-STSR). LTFP-STSR first employs a multi-
dimensional feature embedding strategy that integrates dual-scale temporal features, traffic
flow features, and sinusoidal positional encodings to construct a unified high-dimensional
representation space. Leveraging frequency-domain information extracted via Fourier
transform, a frequency–amplitude–driven sequence decomposition module subsequently
decouples the representation into periodic and residual components. Finally, an attention-
guided frequency-domain spatiotemporal modeling mechanism is introduced to capture
intra-phase, inter-period, and cross-station dependencies jointly. The proposed LTFP-STSR
was validated on real-world traffic datasets and compared against several state-of-the-art
baselines. Experimental results demonstrate that LTFP-STSR consistently outperforms
existing methods, achieving the lowest MAE, MAPE, and RMSE across all prediction horizons,
with the most significant improvement observed at the 120-minute horizon (21.99, 0.094,
and 35.79, respectively). Overall, this work contributes a novel spatial–temporal
reconstruction strategy that aligns model representations with the patterned evolution of
long-term traffic flow, thereby enhancing both the accuracy and robustness of traffic
predictions.
Keywords: Traffic Flow Prediction; Spatiotemporal Sequence Reconstruction; Attention
Mechanism; Fast Fourier Transform; High-order Data Interaction

Yi Zhou, Yu Zhang, Changsheng Chen, Lele Li, Danya Xu, Robert C. Beardsley, Weizeng Shao,
Assessment of radar freeboard, radar penetration rate, and snow depth for potential
improvements in Arctic sea ice thickness retrieved from CryoSat-2,
Cold Regions Science and Technology,
Volume 231,
2025,
104408,
ISSN 0165-232X,
[Link]
([Link]
Abstract: The accuracy of Arctic sea ice thickness retrieved from the CryoSat-2 satellite is
significantly influenced by the sea ice surface roughness, snow backscatter, and snow depth.
In this study, four updated cases incorporating physical model-based radar freeboard, newly
estimated radar penetration rate, and well-validated satellite snow depth were constructed
to evaluate their potential improvements to the Alfred Wegener Institute's CryoSat-2 sea ice
thickness (AWI CS2). The updated cases were then compared with airborne remotely sensed
observations from the National Aeronautics and Space Administration's Operation IceBridge
(OIB) and CryoSat Validation Experiment (CryoVEx) in 2013 and 2014, as well as with ground-
based observations during the Multidisciplinary drifting Observatory for the Study of Arctic
Climate (MOSAiC) expedition from October 2019 to April 2020. The results showed that all
updated cases had the potential to improve the accuracy of sea ice thickness, maintaining
comparable correlation coefficients and significantly reducing statistical errors compared to
the AWI CS2. In the evaluation with OIB, CryoVEx, and MOSAiC, the four updated cases
reduced the root mean square error of AWI CS2 by up to 0.68 m (55 %) against OIB, 0.76 m
(53 %) against CryoVEx, and 0.47 m (76 %) against MOSAiC. The updated sea ice thicknesses
retained the main distribution patterns generated by AWI CS2, but generally showed thinner
sea ice thicknesses. From 2013 to 2018, the interannual variation trends between the
updated cases and AWI CS2 varied regionally, but both show significant decreasing trends
along the northern coasts of the Canadian Arctic Archipelago and Greenland. The updated
schemes provided new insights into the retrieval of sea ice thickness using CryoSat-2,
thereby further contributing to the quantification of the sea ice volume in the context of a
warming climate.
Keywords: Arctic; Sea ice thickness; Satellite retrieval; Assessment

Yadong Xie, Xu Yue, Junfei Zheng, Guangwei Chen, Lin Kong, Dongya Ren,
Mechanistic and spatiotemporal evolution of alkali-pumping in newly constructed bridge
pavements in subtropical environments,
Construction and Building Materials,
Volume 509,
2026,
145163,
ISSN 0950-0618,
[Link]
([Link]
Abstract: In hot and humid regions, early alkali-pumping frequently occurs in newly
constructed asphalt pavement layer on cement concrete bridge deck, with surface whitening
often observed even before the bridge is opened to traffic. This phenomenon severely
compromises the durability and service performance of the bridge deck system. To elucidate
its formation mechanism and dominant influencing factors, this study investigates a newly
built cement concrete bridge deck located in a typical subtropical climate zone, where
extensive surface whitening occurred even before the bridge was opened to traffic. A
combination of field investigation, permeability testing, ground penetrating radar (GPR),
computed tomography (CT) scanning, and X-ray analyses (XRD/XRF) was employed to
systematically explore the water migration pathways and the spatiotemporal characteristics
of alkali-pumping evolution. The results indicate that the early occurrence of alkali-pumping
is closely related to insufficient interlayer compaction, moisture accumulation in structural
depressions, and preferential infiltration through poorly drained zones such as shoulders and
joints. CT analysis demonstrated the presence of interconnected pores within the asphalt
layer, which serve as channels for upward moisture migration and calcium ion transport. XRD
and XRF tests confirmed that the alkali-pumping products are primarily composed of calcium
carbonate, originating from the free calcium components in the cement concrete decks. This
study advances the theoretical understanding of alkali-pumping in cement concrete bridge
decks under hot and humid environments and provides a scientific basis and technical
reference for improving structural design and early-stage damage prevention.
Keywords: Bridge deck pavement; Early alkali-pumping; Formation mechanism; Water
migration path; CT scanning technology

Bingjian Lu, Zhenyu Lu, Xiaowen Zhang, Quanbo Ge,


MFD-CANet: An intelligent identification model for thunderstorm wind gust via
multidimensional feature decoupling and spatiotemporal cross-attention,
Neurocomputing,
Volume 651,
2025,
130926,
ISSN 0925-2312,
[Link]
([Link]
Abstract: Thunderstorm wind gusts are extreme weather phenomena characterized by
strong surges and destructive power. Due to their complex evolution and the wide range of
influencing factors, traditional methods relying on a single observation, such as automatic
weather stations, radar, or satellite, often fail to provide accurate and comprehensive
monitoring. To address the issues of insufficient spatial resolution and low accuracy in
conventional models, this paper proposes an intelligent identification model for
thunderstorm wind gusts, called MFD-CANet. The model integrates multisource
observational data, including ground wind speed, radar reflectivity, and satellite brightness
temperature, to effectively exploit complementary information between different modalities
to accurately capture the spatial organizational structure and evolutionary characteristics of
thunderstorm wind gusts. Specifically, the model first introduces a multidimensional
nonlinear structural perception module to extract key storm evolution features in the meso-
and small-scale convective systems. It then employs a spatiotemporal cross-attention
mechanism to effectively integrate temporal dynamics with spatial distribution information,
enhancing the ability to identify the occurrence locations of thunderstorm wind gusts.
Finally, a loss function based on the Critical Success Index is designed to optimize the
balance between accuracy and false alarms. Experimental results demonstrate that MFD-
CANet achieves superior performance compared to existing mainstream methods, with
average improvements of 6.6 % in Probability of Detection (POD) and 7.8 % in Critical
Success Index (CSI), along with a 9.5 % reduction in False Alarm Rate (FAR). This model
provides more precise and reliable technical support for real-time monitoring and early
warning of severe convective weather.
Keywords: Thunderstorm wind gust; Intelligent identification; Multidimensional feature
decoupling; Spatiotemporal cross-attention

Huilin Liu, Xiaolong Hu, Yu Jiang, Tianyue Wan, Wanqi Ma,


SCTNet: Structured and causality-guided spatiotemporal diffusion network for unsupervised
traffic accident detection,
Information Processing & Management,
Volume 63, Issue 4,
2026,
104598,
ISSN 0306-4573,
[Link]
([Link]
Abstract: Accurate detections of traffic accidents are challenging, and existing detection
methods struggle to maintain robustness to diverse conditions. To address this, we propose
a structured and causality-guided spatiotemporal diffusion network (SCTNet) for
unsupervised traffic accident detection. The SCTNet framework integrates dual-phase patch
sampling (DPPS) to mitigate sampling bias between training and testing phases.
Spatiotemporal causal graph fusion (STCGF) captures the causal dependencies among
interacting agents, and a structured spatiotemporal noise (SSTN) mechanism enhances
temporal sensitivity and context consistency. The diffusion-based dual-stream design
enables the fusion of visual and motion information for robust spatiotemporal
representation learning. Experiments conducted on two traffic datasets show that SCTNet
achieves higher detection accuracy and stronger cross-domain generalization than existing
methods. More generally, our study contributes to data-driven decision making and research
on intelligent information systems in complex, dynamic transportation environments. The
source code is available at [Link]
Keywords: Traffic accident detection; Video anomaly detection; Spatiotemporal diffusion;
Dual-stream diffusion model; Deep learning

Hanmeng Xia, Kaicun Wang,


Hourly, kilometer-scale precipitation merged from rain gauge, ground-based radar and
satellites over east Asia: methods, evaluation and applications,
Journal of Hydrology,
Volume 662, Part C,
2025,
134148,
ISSN 0022-1694,
[Link]
([Link]
Abstract: Accurate precipitation estimation is essential for informed decision making and
planning in sectors such as civil infrastructure, public services, food production, disaster risk
management, and climate change adaptation. Each observational source (gauges, ground
radar, and satellites) offers distinct strengths and suffers from inherent limitations. Radar
supplies high temporal and spatial resolution. Satellite products provide broad areal
coverage. Gauges deliver point measurements with high accuracy. The individual
weaknesses of these sources can be mitigated by integrating multiple datasets. In this study,
gauge, ground radar, and satellite observations with differing native resolutions and
coverage were integrated to produce an hourly precipitation dataset at 0.01° spatial
resolution. The integration workflow comprised bias correction, machine learning based first
guess field generation, and optimal interpolation. The final dataset spans 2017 to 2022 and
the domain 72.6°E to 135.6°E and 15°N to 54.2°N. Each processing stage was systematically
evaluated using independent data subsets across subregions, seasons, and years. The results
show that the final merged dataset consistently outperforms the individual single-source
precipitation products across seasons, subregions, and intensity thresholds (including
extreme events). Using independent test gauges, the merged product achieved a
root-mean-square error (RMSE) of 0.65 mm/h, a critical success index (CSI) of 0.48 at the
0.1 mm/h threshold, and a correlation coefficient (CC) of 0.71. These metrics all exceed the
corresponding values of the standalone satellite and radar products. The merged field
exhibits higher accuracy and stability over diverse physiographic regions, seasonal regimes,
and rainfall intensities. It preserves the fine-scale structural detail and convective cell
capture capability contributed by radar while leveraging the broader spatial coverage of
satellite estimates, thereby further improving quantitative accuracy and spatial
completeness. The merged precipitation products from gauge, ground radar, and satellite
observations at hourly temporal resolution and 0.01° spatial resolution from 2017 to 2022
are available at the National Tibetan Plateau / Third Pole Environment Data Center
([Link] and Wang, 2024).
Keywords: Precipitation; Gauge; Ground radar; Satellite; Machine learning

Long He, Kun Zheng, Huihua Ruan, Shuo Yang, Jinbiao Zhang, Cong Luo, Siyu Tang, Yunlei Yi,
Yugang Tian, Jianmei Cheng,
A spatiotemporal mixed-enhanced generative adversarial network for radar-based
precipitation nowcasting,
Computers & Geosciences,
Volume 200,
2025,
105919,
ISSN 0098-3004,
[Link]
([Link]
Abstract: Skillful precipitation nowcasting with high resolution and detailed information
holds promise for providing reliable alerts about severe weather events to society. Radar
echo extrapolation is an essential method for precipitation nowcasting, but traditional
methods struggle to capture rapidly changing regions. Deep learning (DL)-based methods
exhibit superior performance. However, existing DL-based methods face challenges such as
low accuracy, particularly in producing clear forecasts over longer lead times and accurately
forecasting moderate to heavy rainfall events. To address these challenges, we developed a
novel radar-based precipitation nowcasting model, STMixGAN, which can be described as a
nonlinear proximity forecasting model. This model effectively aggregates global-to-local
information and imposes constraints to represent the complex evolution of rainfall
efficiently. Consequently, STMixGAN produces realistic and spatiotemporally consistent
predictions. Using radar observations from South China, STMixGAN successfully forecasted
radar maps for the next 1 h using 24 min of input data. Two traditional methods (Persistence
and Optical flow) and five DL-based methods (ConvLSTM, Rainformer, IAM4VP, REMNet, and
GAN-argcPredNet) were employed as benchmarks to validate STMixGAN’s forecasting
capabilities. The experimental results demonstrate STMixGAN’s superior performance and
provide valuable insights for enhancing heavy rainfall forecasting.
Keywords: Precipitation nowcasting; Spatiotemporal mixed enhancement; Generative
adversarial networks; Self-attention

Xuqian Bai, Zhitao Zhang, Haorui Chen, Long Qian, Tianjin Dai, Ruiqi Li, Shuailong Fan, Sisi
Jing, Junying Chen, Maosheng Ge,
A spatiotemporal fusion algorithm based on Fourier transform is developed to generate daily
surface soil moisture with 20 m spatial resolution,
Geoderma,
Volume 463,
2025,
117548,
ISSN 0016-7061,
[Link]
([Link]
Abstract: Accurate soil moisture data with detailed spatial and temporal resolutions are
essential for hydrological modeling, precision agriculture, and climate research. Nonetheless,
the intrinsic trade-off between spatial and temporal resolution in remote sensing limits the
accessibility of soil moisture products at granular scales. This study presents a
spatiotemporal fusion algorithm utilizing Fourier transform (STFFT), integrated with Random
Forest (RF), the Water Cloud Model (WCM), and the radiative transfer model (PROSAIL) to
create a comprehensive framework for downscaling surface soil moisture (SSM). Employing
Sentinel-1 and Sentinel-2 datasets, we downscaled Soil Moisture Active and Passive (SMAP)
soil moisture products to generate daily Soil Surface Moisture (SSM) maps at a 20-meter
spatial resolution for the study area. The findings indicate that STFFT is more adept at
accommodating SSM data marked by significant heterogeneity and scale discrepancies
compared to traditional spatiotemporal fusion algorithms. Furthermore, STFFT exhibits
computational efficiency and is independent of reference image selection. The
amalgamation of RF with WCM and PROSAIL adeptly elucidates the intricate correlations
between remote sensing variables and soil moisture; the suggested framework attains
precise soil moisture mapping, evidenced by an average correlation coefficient (R) of 0.892
and a root mean square error (RMSE) of 0.034 m3/m3 across diverse land cover types.
Compared to benchmark methods that produce an average R of 0.753 and an RMSE of
0.043 m3/m3, STFFT demonstrates markedly enhanced accuracy and robustness, particularly
in heterogeneous terrains. This study introduces an improved methodology for producing
fine-scale soil moisture products characterized by enhanced spatiotemporal continuity and
reliability.
Keywords: Surface soil moisture; Downscaling; Spatiotemporal fusion algorithm; SMAP;
Sentinel-1/2

Asim Saleem, Guoyun Lv, Safa Hussein Mohammed,


Deep Learning-Based Radar Fingerprinting for Open-Set Generalization Using Dynamic
Thresholding and Embedding Rejection,
Knowledge-Based Systems,
Volume 333,
2026,
115047,
ISSN 0950-7051,
[Link]
([Link]
Abstract: This paper presents a comprehensive framework for radar-specific emitter
identification (SEI), starting with the simulation of a large-scale radar signal dataset designed
to mimic real hardware impairments. By incorporating diverse distortions–such as phase
noise, frequency jitter, amplitude nonlinearity, and multipath reflections–for multiple radar
types and signal-to-noise ratio (SNR) conditions, we generate a realistic and challenging
dataset, SimRF-14, suitable for learning-based signal analysis. We utilized this dataset to
develop RAFNet, a hybrid deep learning model specifically designed for closed-set and open-
set radar emitter classification. The proposed architecture combines convolutional,
recurrent, and attention-based components to capture spatial, temporal, and contextual
features from normalized I/Q waveforms. For open-set recognition, we integrate the
OpenMax algorithm enhanced with Extreme Value Theory (EVT), where class-wise Weibull
modeling of embedding distances enables outlier detection. In addition, an SNR-adaptive
thresholding mechanism improves open-set reliability under varying noise conditions. The
proposed method achieved 97.43% unknown rejection at -20 dB SNR and maintained low
false positives (<4%) at high SNRs, validating its effectiveness and reliability for practical SEI
scenarios.
Keywords: Radar Specific Emitter Identification; Open-set Recognition; RF Fingerprinting;
SNR-Adaptive Classification; Deep Learning for SEI

Quan Shi, Xiaoliang Xu, Jianlin Li, Huafeng Deng, Qinghai Zhang, Delin Tan, Yu He,
Spatiotemporal effect driven landslide susceptibility mapping at fine scales: a deep learning
model based on multidimensional feature fusion and source data adaptation,
Engineering Applications of Artificial Intelligence,
Volume 156, Part B,
2025,
110924,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Landslides, as frequent natural hazards, pose severe threats to human life and
property safety. This study proposes a deep learning model named Local and Global Feature
Convolutional Network with Interferometric Synthetic Aperture Radar (LGCR-Net). The
model aims to address the issues of multi-source data, heterogeneity, and interference
present in Landslide Susceptibility Mapping (LSM), while also compensating for the
timeliness limitations of LSM. Firstly, the model introduces a parallel structure for extracting
multidimensional features, achieving effective feature fusion through a Multidimensional
Feature Fusion Module(MF). Subsequently, an image enhancement module is incorporated
to process the source data, significantly enhancing the model's stability and source data
adaptability. Furthermore, by integrating the Local and Global Features Convolutional
Network (LGC-Net) model with Interferometric Synthetic Aperture Radar (InSAR) technology,
the LGCR-Net model is formed, enabling the output LSM to possess timeliness. This model
has been successfully applied to the study of the Baihetan reservoir and its surrounding
areas in China. Experimental results indicate that after incorporating the MF and image
enhancement module, the model's Precision increased by 0.0151 and 0.0213, respectively,
while Recall improved by 0.0239 and 0.0334, respectively. Compared to traditional deep
learning models, the LGC-Net model demonstrates superior practicality and predictive
reliability. With the integration of InSAR technology, the improved LSM not only addresses
the model's sensitivity shortcomings in medium and low susceptibility areas but also
maintains predictive performance in high susceptibility areas, providing a more
comprehensive and accurate representation of the potential landslide risks within the study
area.
Keywords: Landslide susceptibility mapping; Multidimensional features; Interferometric
synthetic aperture radar; Multidimensional feature fusion module; Image enhancement
module; Integrating

Jianao Cai, Dongping Ming, Feng Liu, Wenyi Zhao, Mingzhi Zhang, Xiao Ling, Mengyuan Zhu,
Lu Xu, Tingting Lu, Ningjie Liu, Yanfei Wei, Ming Huang,
An enhanced spatiotemporal prediction method on landslide displacement with LDP-
ConvFormer and MT-InSAR observations,
ISPRS Journal of Photogrammetry and Remote Sensing,
Volume 232,
2026,
Pages 594-612,
ISSN 0924-2716,
[Link]
([Link]
Abstract: Landslide Displacement Prediction (LDP) implementation for Landslide Early
Warning Systems (LEWS) using the Multi-Temporal Interferometric Synthetic Aperture Radar
(MT-InSAR) technique poses significant challenges in the Three Gorges Reservoir Area
(TGRA). On the one hand, the limited revisit frequency of satellites fails to satisfy the high-
frequency monitoring requirements of LEWS. On the other hand, traditional LDP methods
concentrate on single-point modeling. It neglects the spatial correlation between
displacement points and the landslide surface. To enhance the low-frequency MT-InSAR
observations, this paper proposes a new hybrid algorithm that integrates the Kalman Filter
(KF) and LDP-ConvFormer to achieve enhanced spatiotemporal LDP. First, multi-orbit MT-
InSAR measurements are transformed to downslope displacement. Subsequently, the multi-
orbit downslope displacements are integrated by KF to generate time series data with
enhanced temporal resolution (5/7-day intervals). The KF estimations indicate that the
integrated higher-resolution time series achieves high accuracy, with an RMSE of 0.431 cm
and an R2 of 0.974 compared to GNSS. Finally, to overcome the limitation of single-point
modeling, a novel LDP-ConvFormer is constructed for enhanced spatiotemporal LDP. The
Spatiotemporal Displacement Prediction Transformer (STDP-Former) employs Local Spatial
Multi-Head Self-Attention (LSMHSA) and Temporal Multi-Head Self-Attention (TMHSA) to
capture the displacement dependencies between different spatial locations at the same time
steps and temporal relationships across different time steps, respectively. Additionally, the
spatiotemporal feature map is decomposed into trend and periodic components, which are
modeled separately and then summed for final predictions. Experimental results
demonstrate that the constructed model can accurately establish the nonlinear relationship
between the landslide displacement and its triggering factors. The LDP-ConvFormer
outperforms benchmark methods, achieving RMSE: 46.29 mm, MAE: 26.7 mm, SSIM:
0.8187, PSNR: 35.62, R2: 0.9574, and EVar: 0.9603. Moreover, LDP-ConvFormer shows
notable superiority in LDP over medium to long periods (60-90d) in the TGRA. The enhanced
spatiotemporal LDP method provides extremely valuable reference for LEWS of translational
landslides in the TGRA.
Keywords: MT-InSAR; Deep learning; Landslide displacement prediction; Kalman filter; TGRA

Yiming Liu, Huadong Guo, Lu Zhang, Dong Liang, Qi Zhu, Zhuoran Lv, Xinyu Dou, Xiaobing Du,
A study of PM2.5 transport pathways in China from 2000 to 2021 with a novel
spatiotemporal correlation method,
Geoscience Frontiers,
Volume 16, Issue 5,
2025,
102116,
ISSN 1674-9871,
[Link]
([Link]
Abstract: In the context of urbanization, air pollution has emerged as a significant
environmental challenge. A thorough understanding of their transport pathways, especially
at a national scale, is essential for environmental protection and policy-making. However, it
remains partially elusive due to the constraints of available data and analytical methods. This
study proposed a data-driven spatiotemporal correlation analysis method employing the
Dynamic Time Warping (DTW). We represented the first comprehensive attempt to chart the
long-term and nationwide transport pathways of PM2.5 utilizing an extensive dataset
spanning from 2000 to 2021 across China, which is crucial for understanding long-term air
pollution trends. Compared with traditional chemical transport models (CTMs), this data-
driven method can generate transport pathways of PM2.5 without requiring extensive
meteorological or emission data, and suggesting fundamentally consistent spatial
distribution and trends. Our analysis reveals that China’s transport pathways are notably
pronounced in the Northwest (34% of the total pathways in China), Southwest (22%), and
North (21%) regions, with less significant pathways in the Northeast (10%) region and
isolated occurrences elsewhere. Additionally, a notable decrease in the number of China’s
PM2.5 transport pathways, similar to annual average concentrations, was observed after
2013, aligning with stricter environmental regulations. Furthermore, we have demonstrated
the feasibility of applying our method to the transport pathways of other gaseous pollutants.
The approach is effective in detecting and quantifying air pollutants’ transport pathways,
even in regions like the Northwest with limited monitoring infrastructure, which may aid in
environmental decision-making. The study will notably improve the current understanding
of air pollutants’ transport process, providing a new perspective for studying the large-scale
spatiotemporal correlations.
Keywords: Air pollutants; Spatiotemporal correlation; Big Earth Data; Transport pathways;
PM2.5; Sustainable development goals

Weicheng Liu, Xia Shi, Wenjun Yan, Shang Gou, Douglas J. Parker, Zhuxia Xu,
Spatiotemporal patterns and propagation characteristics of convective activity on the
northeast slope of Tibetan Plateau: A high-resolution radar perspective,
Atmospheric Research,
Volume 331,
2026,
108609,
ISSN 0169-8095,
[Link]
([Link]
Abstract: The northeastern slope of the Tibetan Plateau, situated in a complex terrain and
the monsoon-westerly transition zone, experiences frequent convective storms and high
disaster risk. Based on CINRAD radar and ERA5 reanalysis data from 2015 to 2019, a high-
resolution convective climatology was established, and its environmental fields were
diagnosed. Results indicate a significant topographic anchoring effect on convection, with
persistent hotspots located in the Yellow River valley-Xinglong Mountain, the eastern Qilian
Mountains, and the sharp-bend reach of the Yellow River. Moderate convection dominates
(64.7 %), while deep convection has a low frequency but high local intensity. The most active
month seasonally is July, with June and August exhibiting similar levels of activity. The
convective activity in July is most active during the season, with levels in June and August
being similar. The configuration of synoptic patterns indicates a synergistic mode of “upper-
level trough, lower-level convergence, and strong moisture transport” for convective days.
Diurnal variation is characterized by a peak in the afternoon (13:00–18:00 BJT) and is
weakest in the early morning to morning, consistent with the synergistic trigger between
solar radiation and terrain convergence. Atmospheric environment diagnostics reveal that,
compared to non-convective events, convective events have higher CAPE, stronger updrafts,
greater vertical wind shear, and more abundant water vapor in the two hours prior to
triggering. The statistical distribution of storm scales exhibits an exponential decay with a
“long-tail” pattern, with approximately 80 %–85 % of convective events having a propagation
distance of less than 20 km and a max area of less than 200 km2. The propagation direction
of convective activity exhibits inter-monthly shifts, trending eastward/southward in June,
shifting northward to east-southeastward in July, and westward/northward in August. These
findings reveal the mechanisms by which large-scale circulation and local topography jointly
influence convection, providing critical scientific support for monitoring and early warning
systems for severe convection in the Tibetan Plateau and its surrounding areas, as well as for
disaster risk prevention and control.
Keywords: Tibetan Plateau; Convective storm; Radar climatology; Propagation
characteristics; Orographic forcing

Yufeng Chi, Kai Wang, Yin Ren, Hong Ye,


The spatiotemporal mechanism of surface compound ozone and heat (SCOH) potential risk
across urban China,
Ecological Indicators,
Volume 178,
2025,
114018,
ISSN 1470-160X,
[Link]
([Link]
Abstract: The synergistic effects of large-scale surface compound ozone and heat (SCOH)
present a more extensive and persistent risk to population exposure and environmental
safety compared to isolated extreme heat or ozone events. Quantifying the spatiotemporal
mechanisms and diffusivity of SCOH in urban areas is therefore critical for risk mitigation.
This study integrates the air pollutants spatiotemporal dataset named Multiple Air Pollutants
dataset (MuAP) and surface heat datasets to map the 1 km-scale time delay correlation
between surface ozone and heat. Combining BayesConvLightGBM and SHapley Additive
exPlanations (SHAP), the quantitative influence of urban factors such as building/canopy
height and road length on SCOH in predominant urban is examined through scene analysis
and diffusion potential analysis. The results show that SCOH has significant temporal and
spatial distribution characteristics. Based on more effective spatiotemporal response
BayesConvLightGBM modeling of SCOH (The BayesConvLightGBM’s R2 is 0.03–0.07 higher
than LightGBM), we found that buildings, roads, and trees have the ability to significantly
affect SCOH in urban, locally or globally. Meanwhile, more compact planning of urban areas
will help reduce the complex risk of SCOH. Even so, it is still important to be aware of the risk
of exposure of SCOH to the population at a range of 4 km or more during a 30-day time
delay period. This study deepens the quantification of nonlinear interactions between urban
infrastructure and SCOH propagation, the understanding of surface compound ozone and
heat, and strengthens key elements and quantification approaches using optimized machine
learning. This is of great significance to explain the spatiotemporal response of SCOH, and
provides an important reference for the study of compound exposure.
Keywords: SHAP; Urban; Surface compound ozone and heat; Machine learning

Chuangwei Xu, Jie Liu, Shiyuan Han, Xiaoqi Duan, Lei Xiang, Tong Zhang,
FourCastLSTM: A precipitation nowcasting model integrating global and local spatiotemporal
features,
Computers & Geosciences,
Volume 204,
2025,
105966,
ISSN 0098-3004,
[Link]
([Link]
Abstract: Accurate precipitation nowcasting is crucial for transportation, agriculture, urban
planning, and tourism, and it is highly beneficial in disaster prevention, resource allocation,
and service optimization. Existing precipitation nowcasting methods often integrate
convolution neural networks and recurrent neural networks or employ vision transformers
to capture spatiotemporal correlations. However, convolutional operators struggle to
capture global information, and vision transformers based global modeling may
overemphasize heavy rainfall while neglecting moderate and light precipitation. In this study,
Fourier nowCasting LSTM (FourCastLSTM) is introduced to effectively capture and fusion
spatiotemporal global and local features of precipitation, enhancing prediction accuracy for
different precipitation intensities. A Fourier nowCasting LSTM Cell (FourCastCell), which
combine the Adaptive Fourier Neural Operator (AFNO) with a simplified LSTM, is proposed
to reinforce the representation of global spatiotemporal precipitation patterns by replacing
traditional convolutional layers with AFNO. An Image Detail Enhancement module (IDE) is
adopted to strengthen local precipitation detail features by integrating difference
convolutional neural network. Finally, the adaptive feature fusion module embedded in the
IDE, can dynamically adjust the integration weights of global and local features based on the
specific spatiotemporal features of precipitation events, ensuring a balanced fusion of
features with different intensities. Experiments on synthetic datasets (MovingMNIST++) and
real-world datasets (RadarCIKM) demonstrate that the proposed FourCastLSTM outperforms
state-of-the-art approaches by 15.6 % and 9.6 % in B-MAE and B-MSE metrics, respectively.
Keywords: Precipitation nowcasting; Spatiotemporal prediction; Spatiotemporal
precipitation feature integration; Heavy rainfall

Yuchen Han, Yiyang Wang, Lei Wang, Changze Zhou, Shijin Yuan,
Extreme-oriented loss: Powering a novel framework for improving spatiotemporal met-
ocean forecasting,
Expert Systems with Applications,
Volume 309,
2026,
131115,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Forecasting spatiotemporal met-ocean series in imbalanced datasets poses a
significant research challenge that warrants substantial attention. Despite specialised
techniques for extreme event prediction, existing methods often prioritise extreme event
prediction accuracy at the expense of accuracy on normal samples. In this paper, we
innovatively proposed a pixel-wise Extreme-oriented Loss (EoL). Distinct from previous
studies that mainly concentrated on the overall prediction difficulty of frames, EoL uniquely
accentuates pixel-wise extreme characteristics and ingeniously integrates potential
imbalances in both temporal and spatial dimensions. The loss applies stronger penalties to
underestimated extreme values while attenuating penalties on their overestimation, thereby
improving the reliability of extreme-event modeling without compromising normal-event
accuracy. Furthermore, a multi-scale feature extraction module is introduced to effectively
capture features across various spatial scales. Additionally, a dependency enhancement
strategy is incorporated, making use of intermediate step prediction information as auxiliary
information to enhance the relatively long-term prediction accuracy. Extensive experimental
evaluations on real-world climate datasets from two distinct domains demonstrate that our
framework consistently outperforms representative state-of-the-art approaches, reducing
MAE and RMSE by up to 24% and yielding notable improvements in R2.
Keywords: Met-ocean forecasting; Spatiotemporal prediction; Data imbalanced; Extreme-
oriented loss
Zhifei Liu, Kang Zheng, Yongze Song, Jianing Zhang,
Daily high-resolution PM2.5 mapping using spatiotemporal CNN-transformer-KAN model,
International Journal of Applied Earth Observation and Geoinformation,
Volume 144,
2025,
104900,
ISSN 1569-8432,
[Link]
([Link]
Abstract: Daily high-resolution mapping of fine particulate matter (PM2.5) is critical for air
quality monitoring and public health. However, current methods struggle to achieve high
accuracy over large spatial and temporal scales due to limitations in modeling complex
spatiotemporal dependencies. This study proposed a novel hybrid deep learning model—
CNN-Transformer-KAN Network (CTKNet)—which utilizes Convolutional Neural Networks
(CNN) to capture spatial features, Transformers for capturing long-range dependencies, and
the Kolmogorov–Arnold Network (KAN) for nonlinear representation learning. Utilizing
spatially continuous satellite aerosol optical depth (AOD) data and other multi-source
spatiotemporal inputs, CTKNet estimated daily PM2.5 at a spatial resolution of 1 km across
China for the period 2015–2020, marking the first application of KAN in PM2.5 estimation. It
outperformed existing models, achieving a high cross-validation coefficient of determination
(R2) of 0.95 (sample-based), 0.90 (station-based), and 0.78 (time-based), and corresponding
RMSEs of 8.13, 11.03, and 17.76 µg/m3. Yearly sample-based cross-validation R2 values
ranged from 0.91 to 0.96 with RMSEs below 11.74 µg/m3, while seasonal R2 values ranged
from 0.82 to 0.89 with RMSEs below 22.14 µg/m3. Analysis reveals a significant decline in
PM2.5 nationwide, especially in eastern and central China. Seasonal peaks occur in winter,
with minima in summer, influenced by meteorology. Spatially, PM2.5 is highest in eastern
and northern regions; urban agglomerations like Beijing–Tianjin–Hebei (BTH) show severe
pollution, while Pearl River Delta (PRD) exhibits the lowest levels due to favorable
conditions. CTKNet also holds promise for other fine-scale environmental mapping tasks
using multi-source spatiotemporal data.
Keywords: PM2.5 estimation; Kolmogorov–Arnold Network; Hybrid deep learning model;
Satellite AOD; Air quality assessment

Zhiyun Yang, Hao Wu, Qi Liu, Xiaodong Liu, Yonghong Zhang, Xuefei Cao,
A self-attention integrated spatiotemporal LSTM approach to edge-radar echo extrapolation
in the Internet of Radars,
ISA Transactions,
Volume 132,
2023,
Pages 155-166,
ISSN 0019-0578,
[Link]
([Link]
Abstract: In recent years, the number of weather-related disasters significantly increases
across the world. As a typical example, short-range extreme precipitation can cause severe
flooding and other secondary disasters, which therefore requires accurate prediction of
extent and intensity of precipitation in a relatively short period of time. Based on the echo
extrapolation of networked weather radars (i.e., the Internet of Radars), different solutions
have been presented ranging from traditional optical-flow methods to recent deep neural
networks. However, these existing networks focus on local features of echo variations to
model the dynamics of holistic radar echo motion, so it often suffers from inaccurate
extrapolation of the radar echo motion trend, trajectory, and intensity. To address the
problem, this paper introduces the self-attention mechanism and an extra memory that
saves global spatiotemporal feature into the original Spatiotemporal LSTM (ST-LSTM) to form
a self-attention Integrated ST-LSTM recurrent unit (SAST-LSTM), capturing both spatial and
temporal global features of radar echo motion. And several these units are stacked to build
the radar echo extrapolation network SAST-Net. Comparative experiments show that the
proposed model has better performance on different real world radar echo datasets over
other recent methods.
Keywords: Radar echo extrapolation; Self-attention; Long short-term memory;
Spatiotemporal prediction

Chuyao Luo, Xinyue Zhao, Yuxi Sun, Xutao Li, Yunming Ye,
PredRANN: The spatiotemporal attention Convolution Recurrent Neural Network for
precipitation nowcasting,
Knowledge-Based Systems,
Volume 239,
2022,
107900,
ISSN 0950-7051,
[Link]
([Link]
Abstract: Precipitation nowcasting is an important task in the fields of transportation, traffic,
agriculture, and tourism. One of the main challenges is radar echo maps forecasting. It is
regarded as a spatiotemporal sequence prediction problem. The prevailing approaches
including the state-of-the-art methods are all based on the ConvRNN which combines the
Convolution Neural Network (CNN) and Recurrent Neural Network (RNN). However, the
feature flow delivered in multi-layer CNNs and RNN usually accompanies the information
loss. Therefore, these algorithms fail to model the long-term dependency and the heavy
rainfalls tend to be underestimated. In addition, they cannot predict the increasing intensity
trend of heavy rainfalls. In this paper, we propose a PredRANN model by embedding the
Temporal Attention Module (TAM) and Layer Attention Module (LAM) into the prediction
unit to preserve more representation from temporal and spatial dimensions respectively.
The extensive experimental results on both synthetic data sets and real world data sets
demonstrate the effectiveness and superiority of the proposed method over state-of-the-art
methods. Ablation studies also validate the developed TAM and LAM components. To
reproduce the results, we release the source code at:
[Link]
Keywords: Precipitation nowcasting; Spatiotemporal sequence prediction; Self attention

Hansheng Zeng, Yuqi Li, Ruize Niu, Chuanguang Yang, Shiping Wen,
Enhancing spatiotemporal prediction through the integration of Mamba state space models
and Diffusion Transformers,
Knowledge-Based Systems,
Volume 316,
2025,
113347,
ISSN 0950-7051,
[Link]
([Link]
Abstract: This paper presents an advanced architecture for spatiotemporal prediction MAD,
integrating Mamba modules with Diffusion Transformers for efficient spatiotemporal
modeling. The model consists of three phases: encoding, reconstruction, and prediction.
Initially, the encoder transforms raw spatiotemporal data into compact latent embeddings.
In the reconstruction phase, the Mamba module processes these embeddings through
normalization and bidirectional state space models, generating reconstructed
representations which are then decoded to restore the input data. The prediction phase
utilizes the Diffusion Transformer to model spatiotemporal features, incorporating time
embeddings and leveraging self-attention mechanisms to capture complex spatiotemporal
dependencies. Finally, the model jointly trains the reconstruction and prediction paths to
achieve high-precision spatiotemporal forecasts. Experimental results demonstrate the
model’s superior performance across various spatiotemporal prediction tasks, validating its
effectiveness and robustness. Our codes are available at
[Link]
Keywords: Deep learning; Spatio-temporal prediction; Mamba; Diffusion

Tuo Xie, Xinyao Yun, Gang Zhang, Hua Li, Kaoshe Zhang, Ruogu Wang,
Charging station cluster load prediction: Spatiotemporal multi-graph fusion technology,
Renewable and Sustainable Energy Reviews,
Volume 206,
2024,
114855,
ISSN 1364-0321,
[Link]
([Link]
Abstract: In recent years, single-station charging load prediction technology for electric
vehicles has gradually matured, but there are few prediction studies at the charging station
cluster level. Therefore, this research propose a load prediction framework for electric
vehicle charging station groups based on multi-graph fusion. First, a distance map, a traffic
network map, and a traffic density map are established to extract the topological
information of the charging station group, and the fusion operation is performed by
establishing the relationship between the influencing factors based on the mutual
correlation of the influencing factors and the multi-graph attention mechanism; Secondly, a
spatiotemporal prediction model was constructed, multi-level feature extraction was
performed, and multiple charging stations were predicted at the same time; Finally, taking
the cluster load of charging stations in an urban area as an example, a comparative
experiment was conducted to compare the model proposed in this study with the
mathematical model, the prediction performance of different variants of machine learning
models, common deep learning models and the model proposed in this study, and a
comparative test with multiple prediction horizons was conducted. The research results
show that the model proposed in this work improves the accuracy of multi-station
forecasting and provides new ideas for data-driven charging station cluster forecasting
research.
Keywords: Spatiotemporal forecasting of charging load; Attention mechanism; Graph neural
network; Maximum information coefficient; Spatiotemporal feature mining; Graph
convolutional neural network; Graph attention neural network; Multi-foresight prediction

Yuxiang Gao, Lu Liang, Tiecheng Su, Mingzhang Pan,


An embedded spatiotemporal hybrid model integrating multi-graphs and attention-driven
fusion for single- and multi-site photovoltaic power forecasting,
Energy Conversion and Management,
Volume 336,
2025,
119897,
ISSN 0196-8904,
[Link]
([Link]
Abstract: Single-site photovoltaic power prediction supports localized station management,
while multi-site prediction facilitates regional grid optimization. However, most existing
models focus solely on single-site prediction, with limited research addressing simultaneous
multi-site forecasting. Even rarer research can both leverage information from neighboring
stations for accurate single-site prediction and consider global spatiotemporal coupling for
multi-site prediction. To bridge this gap, this study introduces a novel spatiotemporal hybrid
model that integrates multi-graphs and attention-driven feature fusion to achieve both
tasks. The model embeds Chebyshev graph convolution into bidirectional Long Short-Term
Memory networks for spatiotemporal feature extraction. Using multi-graph structures, it
combines multi-source heterogeneous data to extract multi-view features with enhanced
granularity. Attention mechanisms dynamically refine and integrate these features,
emphasizing the most critical information. Additionally, an improved dung beetle optimizer
enhances hyper-parameter tuning through innovations such as chaotic initialization, search
range normalization, adaptive factors, and random perturbations, ensuring robust
performance. Experimental validation using datasets from eight DKASC stations in Australia
demonstrated the model’s effectiveness. It consistently outperformed benchmark models in
single-site prediction and achieved state-of-the-art results in multi-site forecasting, with an
R2 of 0.99457 and MSE, MAE, and RMSE values of 0.01340, 0.04626, and 0.11576,
respectively. These findings highlight the model’s potential for renewable energy forecasting
and efficient multi-scale grid management.
Keywords: Photovoltaic power prediction; Graph convolutional network; Attention
mechanism; Spatiotemporal hybrid models; Improved optimization algorithm

Renaldy Fredyan, Karli Eka Setiawan,


Spatiotemporal rainfall rorecasting through GAN-based deep learning with enhanced
temporal coherence,
Procedia Computer Science,
Volume 269,
2025,
Pages 372-379,
ISSN 1877-0509,
[Link]
([Link]
Abstract: Accurate short-term rainfall forecasting based on radar is becoming increasingly
important as extreme weather events occur more frequently on Earth. However, accurately
predicting rainfall remains a challenge for researchers due to the complex and dynamic
nature of the atmospheric system, making it difficult to accurately predict. To address this,
we introduce GAN Analysis, a deep learning model that combines spatial and temporal
prediction networks with a generative adversarial framework. This module uses multiple
decoding modules connected with neural network layers to better capture the temporal
structure in radar data so that the model can retain temporal information over a long period
of time. Furthermore, we also use an improved loss function to improve the clarity and
accuracy of the predicted rainfall map so that the model can learn well. Experimental results
show that the model achieves good performance especially under high-intensity rainfall
conditions.
Keywords: Generative Adversarial Network; Spatiotemporal; Radar Data; Rainfall Prediction

Yong Liu, Chenyang Lu, Liang Li, Xiangchao Meng, Qiuping Jiang, Feng Shao,
Interactive feature fusion for camera-radar-based vehicle segmentation in bird’s-eye view,
Pattern Recognition,
Volume 172, Part D,
2026,
112698,
ISSN 0031-3203,
[Link]
([Link]
Abstract: Vehicle segmentation in Bird’s-Eye View (BEV) is a fundamental task for
autonomous driving, and integrating multi-modal sensory inputs, e.g., cameras and radars,
could enhance perception capability by leveraging their complementary strengths. However,
cross-modal feature fusion raises additional challenges due to the sparse and noisy
characteristics of radar data and the inherent misalignment between radar and camera
features. While existing fusion methods frequently leverage powerful attention mechanisms,
they often overlook the aforementioned heterogeneities and their impact on achieving
consistent, fine-grained alignment across modalities. We introduce the Interactively
Enhanced Camera-Radar Fusion (IECRF) framework, a novel approach that effectively bridges
cross-modal discrepancies in two stages through three new modules: Camera-Radar Feature
Aggregation (CRFA), Multi-Scale Radar Enhancer (MSRE), and Camera-Radar Feature Fusion
(CRFF). Specifically, the CRFA module explicitly models the complementary features of visual
appearance and radar geometry through two attention mechanisms, enabling fine-grained
alignment and interactive enhancement between the two modalities. The MSRE module
further refines radar representations through a modality-specific down- and up-sampling
design, amplifying salient targets while suppressing background noise in sparse radar
features. The aggregated features are then fused using the CRFF module at each stage for
lateral decoding. Extensive evaluations on the nuScenes dataset demonstrate that our IECRF
framework can operate with multiple backbones and configurations, achieving higher
vehicle segmentation accuracy even when using a lightweight EfficientNet backbone, which
is six times faster than the existing state-of-the-art approach equipped with an advanced ViT
backbone. The source code and trained models are available at
[Link]
Keywords: Vehicle segmentation; Bird’s-Eye View (BEV); Radar perception; Visual-radar
fusion; Autonomous driving

Chaofeng Huang, Xiaowo Xu, Fan Fan, Shunjun Wei, Xiaoling Zhang, Dongmei Liu, Min Gu,
A low-SNR-adaptive temporal network with smart mask attention for radar signal
modulation recognition,
Digital Signal Processing,
Volume 168, Part D,
2026,
105640,
ISSN 1051-2004,
[Link]
([Link]
Abstract: The automatic modulation recognition of radar signals is a key technology in
electronic warfare and communication systems. However, traditional handcrafted features
often struggle to achieve high recognition accuracy under low signal-to-noise ratio (SNR)
conditions. With the rapid development of artificial intelligence technologies, deep learning-
based approaches have emerged as a promising alternative for modulation recognition. In
this article, a low-SNR-adaptive network architecture is proposed, which integrates a
bidirectional temporal convolutional network (Bi-TCN) and dual-channel smart mask
attention (DSMA) modules. The DSMA adaptively highlights informative features and
suppresses noise through complementary attention masks, enhancing robustness in low-SNR
conditions. Experimental results demonstrate that the autocorrelation domain outperforms
both time and frequency domains, with recognition accuracy improvements of 13.33 % and
14.71 %, respectively. Compared to state-of-the-art models, the proposed network achieves
63 % accuracy at -20 dB and more than 99 % accuracy at -6 dB, significantly enhancing radar
signal modulation recognition.
Keywords: Modulation recognition; Deep learning; Radar signal analysis,

Xiao Cui, Baisheng Nie, Hengyi He, Peng Liu, Kaidan Bai, Haowen Zhou, Jingtao Yang,
Explainable deep learning for spatiotemporal high-temperature evolution and predictive
modeling in coal seam enhanced combustion,
Process Safety and Environmental Protection,
Volume 204,
2025,
108084,
ISSN 0957-5820,
[Link]
([Link]
Abstract: Temperature monitoring during deep coal seam combustion is essential for
optimizing in-situ heat extraction and ensuring operational safety. Accurate prediction of
temperature variations under enhanced combustion thus becomes a critical tool for
maintaining both safety and efficiency. In this study, continuous-ventilation coal combustion
experiments were performed to examine the spatiotemporal evolution of temperature, gas
emissions, and mass variation. The results showed that the high-temperature zone migrated
along the airflow direction, accompanied by pronounced spatiotemporal fluctuations in gas
concentrations. Using 13 input features—Coal Weight, Cross Section, Coal Quality Loss Rate,
CO/CO2, CO/H2, C3H8, C3H6, C2H4, C2H6, CH4, H2, CO2, and CO—predictive models for
enhanced combustion temperature were developed with multiple machine learning
methods. Eleven models were assessed, including LSTM, CNN-LSTM, CNN-LSTM with
Attention, BP, RNN, CNN, GRU, Transformer, RF, RBF, and XGBoost. Among them, the CNN-
LSTM-Attention model achieved the best performance, with an R2 of 0.987, MAE of 0.41,
RMSE of 0.59, and MAPE of 0.47—substantially outperforming the other ten models. To
enhance interpretability, SHapley Additive exPlanations (SHAP) were applied, revealing that
CO2 concentration had the strongest impact on prediction (mean SHAP value: 0.0555),
followed by the CO/CO2 ratio (0.0346) and Cross Section (0.024). Overall, this study
proposes a robust and interpretable approach for high-precision temperature prediction
during underground coal combustion, offering important guidance for thermal monitoring in
in-situ heat extraction systems.
Keywords: Enhanced coal combustion; High-temperature migration; Convolutional neural
network (CNN); Long short-term memory (LSTM); Attention mechanism; SHAP
interpretability

Hongchu Yu, Chenxi Jiang, Qinglong Fang, Tianming Wei, Lei Xu,
Deep learning driven spatiotemporal prediction of global carbon emissions from container
shipping,
Transportation Research Part D: Transport and Environment,
Volume 151,
2026,
105169,
ISSN 1361-9209,
[Link]
([Link]
Abstract: Container shipping is a significant source of global CO2 emissions, making accurate
predictions essential for meeting international environmental targets. This study proposes
ConvLSTM-CBAMNet, a deep learning model integrating channel and spatial attention
mechanisms to capture complex spatiotemporal emission trends. The model significantly
outperforms four deep learning baselines, including Transformer, ConvGRU, CNN-LSTM, and
the traditional ConvLSTM. Compared to the strongest baseline, ConvLSTM, it demonstrates
marked improvements, reducing the Root Mean Square Error (RMSE) by 19.4% to 0.0914
and the Mean Absolute Error (MAE) by 16.8% to 0.0432, while increasing the Structural
Similarity Index (SSIM) by 5.8% to 0.9035. These predictions can inform targeted
environmental policies, dynamically adjust Emission Control Areas (ECAs), optimize port
scheduling to mitigate pollution peaks, and develop environmental early-warning systems,
thereby supporting the shipping industry’s transition toward sustainability.
Keywords: Container ships; Carbon emissions; AIS data; Deep learning; Spatiotemporal
prediction

Jiaxin Qian, Jie Yang, Weidong Sun, Lingli Zhao, Lei Shi, Hongtao Shi, Chaoya Dang, Qi Dou,
Application potential and spatiotemporal uncertainty assessment of multi-layer soil moisture
estimation in different climate zones using multi-source data,
Journal of Hydrology,
Volume 645, Part B,
2024,
132229,
ISSN 0022-1694,
[Link]
([Link]
Abstract: Accurately estimating multi-layer soil moisture (SM) through remote sensing
methods presents inherent challenges and limitations. Multi-layer SM provides valuable
insights into the intricate interactions within the “soil-vegetation-atmosphere” system. This
study explored the temporal dynamics of multi-layer SM in the Shandian River Basin, China,
from 2019 to 2020. Through sensitivity analysis, we demonstrated the feasibility of using
multi-source data for estimating multi-layer SM, including dual polarization radar data,
optical vegetation descriptors, terrain factors, soil parameters, and meteorological indices.
Initially, surface soil moisture (SSM) at depths of 3 cm and 5 cm was estimated using the
modified change detection (MCD) model, which reduces the impact of vegetation.
Incorporating constraints from soil parameters during the solving process improved the
estimation accuracy of multi-layer SM. Subsequently, the water balance model, involving
precipitation and evaporation, was applied to further correct the estimation results of SSM.
Based on this, the infiltration process was considered to estimate deeper SM, including near-
surface soil moisture (NSSM) at depths of 10 cm and 20 cm, and root zone soil moisture
(RZSM) at depths of 40–50 cm. Under this framework, the estimation errors for multi-layer
SM were satisfactory (RMSE = 0.041–0.045 cm3/cm3). Finally, we explored the upper limits
of multi-layer SM estimation using multi-input and multi-output machine learning regression
(MLR) algorithms. With the incorporation of multi-source data, advanced MLR algorithms
achieved higher estimation accuracy (RMSE = 0.015–0.022 cm3/cm3) and showed potential
for cross-temporal transfer (RMSE = 0.030–0.037 cm3/cm3). Moreover, spatiotemporal
robustness revalidation of multi-layer SM was conducted across 17 observation networks
distributed cross different climatic zones in China. The results shown that the MCD model
achieved satisfactory results in estimating multi-layer SM (RMSE = 0.053–0.064 cm3/cm3),
whereas the regression models displayed higher accuracy (RMSE = 0.039–0.051 cm3/cm3).
Both the MCD and MLR models yielded similar conclusions, indicating that the estimation
accuracy of NSSM and RZSM surpassed that of SSM, primarily due to the relatively lower
variability of the former and their strong coupling with vegetation productivity. This study
also specifically discussed the influence of factors such as radar incidence angles, soil texture
types, and vegetation types on the estimation accuracy of multi-layer SM. This study
introduced a novel concept and framework for regional multi-layer and profile SM
estimation and real-time prediction through multi-source data, exhibiting high potential for
practical applications.
Keywords: Multi-layer soil moisture; Dual-polarization SAR data; Multi-source data
collaboration; Various climatic zones; Spatiotemporal estimation and prediction

Jongyun Byun, Jaehoon Cha, Jeyan Thiyagalingam, Hyeon-Joon Kim, Changhyun Jun,
Enhancing rainfall prediction accuracy through image fusion of radar and numerical weather
prediction models,
Expert Systems with Applications,
Volume 303,
2026,
130516,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Precipitation is one of the most challenging atmospheric phenomena to predict
due to the complexity involved in solving dynamic and thermodynamic atmospheric
equations. To address this challenge, extensive research has been conducted to enhance the
precision of numerical weather prediction models and radar-based extrapolation data, in
conjunction with the development of various blending techniques. However, traditional
methods have proven insufficient in capturing the diversity and nonlinearity of weather
phenomena. In response to these limitations, this study introduces a novel methodology
that leverages machine-learning based image fusion models to merge radar-based
extrapolation and numerical weather prediction rainfall datasets, thereby enhancing
prediction accuracy. An image fusion model was developed using radar-based extrapolation
data and numerical weather prediction data as input datasets, with radar observation data
utilized as target dataset. To identify the most suitable image fusion model for capturing the
complex patterns of rainfall data, two experiments were conducted: 1) Impact of model
topology, and 2) Effect of model size. A systematic analysis of the model outputs was
performed using eight evaluation metrics categorized under pixel-based metrics, feature-
based metrics, structural similarity metrics, and categorical verification metrics.
Experimental results indicated that image fusion model based on a Residual Network
(ResNet) outperformed other models in terms of model topology. Regarding model size, it
was observed that the performance did not increase proportionally with the number of
residual blocks; the most suitable performance was achieved with a specific number of
residual blocks (Case 5: 8 blocks). Additionally, the metrics compared with radar observation
data indicated that the proposed model delivered superior performance, thus offering a
high-accuracy rainfall prediction methodology.
Keywords: Image fusion; Radar; Numerical weather prediction; Deep learning; Precipitation

Cunyang Zhang, Yongmao Hou, Xiaohe Xia, Jin-Jian Chen, Yue Pan,
Multimodal feature fusion deep learning for spatiotemporal prediction of deformation and
environmental impacts in pipe-roof tunnel construction,
Advanced Engineering Informatics,
Volume 69, Part C,
2026,
104022,
ISSN 1474-0346,
[Link]
([Link]
Abstract: The pre-support tunnel construction involves complex construction conditions and
multisource data, posing challenges for efficient data transfer and accurate spatiotemporal
predictions. This study proposes an attention-based multimodal feature fusion deep learning
(AMFF-DL) framework that establishes a computational link between construction activities
and geotechnical responses. The AMFF-DL framework comprises two core components: a
multisource data preprocessing (MDP) module that systematically integrates on-site
information—including construction records, structural design parameters, and geological
surveys—into a unified, structured database, and an Attention-based Multimodal Feature
Fusion (AMFF) model that enables effective feature extraction, multimodal fusion, and
predictive learning. Applied to a real-world tunnel project in Shanghai, China, AMFF-DL
demonstrates strong predictive performance, achieving a mean absolute error (MAE) of
0.80 mm for deformation forecasts and a structural similarity index measure (SSIM) of
0.8936 for deformation cloud maps. It also accurately predicts key environmental indicators
such as pore water pressure, soil pressure, and horizontal displacement. Compared to
conventional prediction approaches, AMFF-DL performs credible data-transferring and
reliable predictions through its structured multimodal database and attention-based feature
fusion. Practically, AMFF-DL provides actionable insights into tunnel-induced impacts and
supports intelligent, data-informed risk management in complex underground construction
environments.
Keywords: Multimodal feature fusion; Deep learning; Spatiotemporal prediction;
Deformation and environmental impact; Pre-support tunnel construction

Xiaofei Zhang, Zhengping Fan, Xiaojun Tan, Qunming Liu, Yanli Shi,
Spatiotemporal adaptive attention 3D multiobject tracking for autonomous driving,
Knowledge-Based Systems,
Volume 267,
2023,
110442,
ISSN 0950-7051,
[Link]
([Link]
Abstract: Three-dimensional (3D) multiobject tracking (MOT) is an essential perception task
for autonomous vehicles (AVs). Studies have indicated that multimodal data fusion can
provide more stable and efficient perception information to AVs than a single sensor.
Therefore, this paper proposes a new spatiotemporal adaptive attention 3D (3DSTAA)
tracker, which attempts to improve the tracking performance of the end-to-end 3D MOT by
adaptively correlating spatiotemporal data. The novelty of this paper includes the following.
(1) Different from nonintelligent fusion methods, this paper uses an efficiently adaptive
spatial-guided fusion (SGFus) module for multimodal feature fusion. As a result, the 3D
structural information obtained from point cloud data can provide additional spatial
information as complementary information to the 2D texture information extracted from the
image data, collaboratively facilitating and refining the perception information
representation in the margin area. (2) This paper develops a spatiotemporal object-unique
attention (STOUA) module that calculates the relational degree of each perceived object
between two adjacent frames through attentional encoding. At the same time, an adaptive
weighting strategy is used to further study the spatiotemporal correlation of unique objects,
reducing the similarity among various objects and the differences across the same object.
Experiments tested using the KITTI tracking benchmark show that the 3DSTAA tracker is
highly competitive in both inference time and tracking performance compared with state-of-
the-art (SOTA) methods. Our corresponding code will be released on the
[Link]
Keywords: 3D multiobject tracking; Multimodal data fusion; Spatiotemporal attention
mechanism; Adaptive data association

Changlong Wang, Jiawei Jiang, Chong Han, Hengyi Ren, Lijuan Sun, Jian Guo,
Through-Wall Multihuman Activity Recognition Based on MIMO Radar,
Computers, Materials and Continua,
Volume 83, Issue 3,
2025,
Pages 4537-4550,
ISSN 1546-2218,
[Link]
([Link]
Abstract: Existing through-wall human activity recognition methods often rely on Doppler
information or reflective signal characteristics of the human body. However, static
individuals, lacking prominent motion features, do not generate Doppler information.
Moreover, radar signals experience significant attenuation due to absorption and scattering
effects as they penetrate walls, limiting recognition performance. To address these
challenges, this study proposes a novel through-wall human activity recognition method
based on MIMO radar. Utilizing a MIMO radar operating at 1–2 GHz, we capture activity data
of individuals through walls and process it into range-angle maps to represent activity
features. To tackle the issue of minimal variation in reflection areas caused by static
individuals, a multi-scale activity feature extraction module is designed, capable of extracting
effective features from radar signals across multiple scales. Simultaneously, a temporal
attention mechanism is employed to extract keyframe information from sequential signals,
focusing on critical moments of activity. Furthermore, this study introduces an activity
recognition network based on a Deformable Transformer, which efficiently extracts both
global and local features from radar signals, delivering precise human posture and activity
sequences. In experimental scenarios involving 24 cm-thick brick walls, the proposed
method achieves an impressive 97.1% accuracy in activity recognition classification.
Keywords: MIMO radar; human activity; Transformer; through-wall

Vipina Valsan, A.M. Abhishek Sai, Anu G Kumar, Aryadevi Remanidevi Devidas, Maneesha
Vinodini Ramesh, Kanakasabapathy P,
Deep learning-enabled spatiotemporal sustainable energy strategy for rural multiple
microgrids,
Results in Engineering,
Volume 28,
2025,
107569,
ISSN 2590-1230,
[Link]
([Link]
Abstract: Sustainable economic development entails effective energy management. Yet
there is a dearth of schemes that jointly capture spatial and temporal energy dynamics,
while ensuring fairness in underserved regions for geographically distant microgrids through
virtual energy sharing. This work proposes a spatiotemporal framework, for microgrid
energy management, stimulating exchange of excess renewable energy. The recommended
scheme facilitates intelligent utilization of renewable energy, with moderated costs of
generation and distribution. A simulated case study of two rural microgrid locations in India,
compliant with India's National Grid, validated the propriety of the proffered approach.
Findings showed a 30.7% reduction in dependency on the main grid, with drops in annual
energy imports from (19.1 - 13.2) MWh. Besides, multi-microgrid coordination delivered cost
savings exceeding INR 15,000, validated during monsoon and winter seasons. This study
corroborates the potential of coordinated microgrids to enhance resilience, inclusivity and
clean energy, in support of UN Sustainable Development Goals 7, 9, 11, 12, and 13.
Keywords: Deep learning; Multiple microgrids; Renewable energy; Spatiotemporal;
Sustainable energy sharing

Jie Liu, Lei Xu, Nengcheng Chen,


A spatiotemporal deep learning model ST-LSTM-SA for hourly rainfall forecasting using radar
echo images,
Journal of Hydrology,
Volume 609,
2022,
127748,
ISSN 0022-1694,
[Link]
([Link]
Abstract: Accurate and timely short-term forecasting services of precipitation variable are
significant for people's lives and property security. The data-driven approaches demonstrate
promising performance in the extrapolation of precipitation. In this paper, we proposed a
spatiotemporal prediction model, namely the Spatial Temporal Long Short-Term Memory
based on the self-attentive mechanism (ST-LSTM-SA). This model enables better aggregation
of sequence features inspired by proposed improvements. The encapsulated 3D convolution
is developed to fully exploit the short-term spatiotemporal information, and the channel
correlation is modeled by self-attention mechanism to further improve representations in
the long-term interaction, the effectiveness of which is validated in the ablation study.
Comprehensive experiments have been conducted on the radar echo sequence, we
successfully predicted future radar reflectivity images for next 3 h with data for previous
three hours as inputs in Wuhan, China. Three machine learning methods: multiple linear
regression (MLR), support vector regression (SVR) and artificial neural networks (ANN) and
deep learning model—ConvLSTM have been chosen as comparative groups to corroborate
the nowcasting availability of this model for radar echo extrapolation. Moreover, the
supplementary experiment on minute dataset also illustrated the superiority of ST-LSTM-SA.
The studies analyzed the forecasting performance in terms of image quality and rainfall
error. The experimental results demonstrated the better versatility and performance of ST-
LSTM-SA. These conclusions and attempts may provide efficient guidance for precipitation
nowcasting in urban areas.
Keywords: Precipitation nowcasting; Spatiotemporal LSTM; Self-Attention; Radar echo image

Xin Bi, Caien Weng, Panpan Tong, Wei Tian, Lu Xiong,


Dual-sampling feature fusion for three-dimensional object detection using four-dimensional
radar and camera,
Engineering Applications of Artificial Intelligence,
Volume 152,
2025,
110747,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Bird’s-eye view (BEV) fusion of four-dimensional (4D) radar and camera provides a
cost-effective solution for three-dimensional (3D) object detection, ensuring all-weather
perception in autonomous vehicles. However, existing BEV-based fusion methods experience
progressive degradation of camera-generated features due to overfitting and inconsistent
generalization during training, which limits their effectiveness in multi-modal fusion.
Although complex networks estimate 3D information from images, their accuracy remains
suboptimal in end-to-end training, while their high computational and storage costs hinder
practical deployment. To address these challenges, we propose a novel approach for multi-
modal 3D object detection using 4D radar and camera fusion, termed dual-sampling feature
fusion (DSFusion). We design two modules: predefined voxel-based view transformation
(PVVT) and radar pillar-based view transformation (RPVT). Both modules aim to sample
image features through point cloud matrix transformations. The RPVT module effectively
captures more reliable image features by sampling the raw radar point cloud, mitigating the
feature degradation caused by direct view transformation. Meanwhile, the PVVT module
employs a predefined point cloud to sample additional image features and utilizes the radar
occupancy estimation network (ROENet) to enhance feature selection efficiency. Finally,
multiple BEV features are fused through convolution, and a unified detection head predicts
object classes and locations. Experimental results demonstrate that DSFusion achieves state-
of-the-art radar-camera fusion detection performance on the View-of-Delft and TJ4DRadSet
datasets, approaching the accuracy of high-resolution light detection and ranging while
maintaining all-weather capability. Furthermore, DSFusion enables real-time inference at
over 10 frames per second, meeting the real-time requirements of autonomous vehicle
applications.
Keywords: Multi-modal fusion; Three-dimensional object detection; Autonomous vehicles;
Four-dimensional radar; Deep learning

Xiaofang Sun, Meng Wang, Junbang Wang, Guicai Li, Xuehui Hou,
Deep learning classification of winter wheat from Sentinel optical-radar image time series in
smallholder farming areas,
Advances in Space Research,
Volume 75, Issue 3,
2025,
Pages 2683-2695,
ISSN 0273-1177,
[Link]
([Link]
Abstract: As crop yield stagnation, climate change, and the rising demand for agricultural
products pose increasing challenges, mapping crop systems is becoming more and more
important. Winter wheat is one of the major cereal crops cultivated in China, ranking as the
third largest crop in terms of production and harvested area. Accurately mapping winter
wheat is necessary for implementing effective farm management practices. While many
studies have successfully produced high spatiotemporal resolution land cover maps,
relatively few map products of crop types are available in China. The growing archive of
satellite image time series provides enormous opportunities to map crops more closely. This
research presents a two-step method to map winter wheat based on Sentinel-1 and
Sentinel-2 time-series data from Shandong Province using the deep learning approaches.
The winter crops were firstly mapped using time-series optical vegetation indices employing
the deep learning methods. Then winter wheat was extracted from the winter crops mask by
coupling optical and synthetic aperture radar time-series images. The results indicated that
the precision of mapping winter wheat using Temporal Convolution Neural Networks
(TempCNN) achieved the highest precision in mapping winter wheat, with an overall
accuracy of 93.7 %, a kappa coefficient of 0.907, and an F1-score of 0.989. This was followed
sequentially by the Residual 1D convolutional neural networks (ResNet), the Multi-Layer
Perceptron (MLP), and the Lightweight Temporal Self-Attention Encoder (L-TAE). The
Temporal Attention Encoder (TAE) model demonstrated the lowest precision among the
compared models. The results agree well with independent county-level official census
winter wheat area data (R2 = 0.936). The proposed framework can also be applied in other
regions to generate maps of different crops, so future work can extend the proposed model
to other agricultural regions, where an increased number of crop types and natural
vegetation types can be included and tested.
Keywords: Sentinel-1; Sentinel-2; Classification; Winter wheat mapping; Time series; Deep
learning

Yehao Wang, Zijian Liu, Yingying Jin, Xiaoliang Wang, Lingyu Xu, Lei Wang, Jie Yu, Wenjuan
Dai, Jingxia Gao, Feng Zhang,
Interpreting spatiotemporal dynamics of Ulva prolifera blooms in the southern yellow sea
using an attention-enhanced transformer framework,
Environmental Pollution,
Volume 384,
2025,
126999,
ISSN 0269-7491,
[Link]
([Link]
Abstract: Harmful algal blooms dominated by Ulva prolifera have posed recurring ecological
and economic challenges in the southern Yellow Sea. To better understand and predict the
complex spatiotemporal dynamics of these blooms, we developed an enhanced
Transformer-based deep learning framework, incorporating multi-head self-attention
mechanisms. This model dynamically captures spatial dependencies, providing a
comprehensive understanding of bloom dynamics. Utilizing twelve key marine
environmental factors, we systematically explored all possible feature combinations to
determine the optimal predictive subset. Experimental results demonstrated superior
predictive performance of the model (MAE: 0.0213, MSE: 0.0016, R2: 0.9923) compared to
conventional deep learning models and recent spatiotemporal deep learning models.
Training dynamics revealed efficient convergence, especially with comprehensive
environmental information. Spatial attention analysis revealed that offshore regions
consistently received higher attention, indicating their critical role as informative and
generalizable environmental references. Furthermore, exhaustive feature attribution
experiments identified an optimal combination of eight environmental factors—including
temperature, salinity, current velocity, precipitation, wind direction, dissolved iron,
phosphate, and silicate—were found to significantly enhance prediction accuracy. This study
highlights the capability of attention-enhanced Transformer models for interpretable and
precise ecological forecasting, providing valuable insights for targeted mitigation and
management of U. prolifera blooms.
Keywords: Harmful algal blooms; U. prolifera; Deep learning; Spatiotemporal dynamics;
Environmental factors; Southern yellow sea

Jiabing Liu, Jianhao Sun, Haiwen Wei, Qilei Li, Junzhi Shi, Mingliang Gao,
Cloud prediction via spatiotemporal-frequency differential and attentional network,
Engineering Applications of Artificial Intelligence,
Volume 166, Part A,
2026,
113476,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Cloud prediction is pivotal for meteorology, aviation safety, and renewable energy
management. A fundamental challenge in existing deep learning approaches lies in the
trade-off among prediction accuracy, computational efficiency, and long-term stability. To
bridge this gap, we introduce an end-to-end encoder–decoder architecture, termed
Spatiotemporal-Frequency Differential and Attentional Network (SFDANet). SFDANet
constructs an encoder–decoder architecture with a unique SFFE block, which integrates
spatiotemporal and frequency-domain analysis to simultaneously capture localized cloud
textures and global evolutionary dynamics. Between the encoder and decoder, an innovative
Multi-scale Differential Pyramid (MDP) module is built to selectively enhance high-frequency
details critical for rapid cloud evolution while inherently suppressing noise. To explicitly
model complex temporal dynamics, we propose a new module parallel to MDP, named
Multi-order Projection Attention (MPA). This module operates by projecting input features
into a set of parallel subspaces. Crucially, these subspaces are designed to be both linear and
non-linear. Through this architectural design, the module is capable of simultaneously
capturing predictable low-order trends and intricate high-order patterns within the data.
Comprehensive experiments on the WeatherBench dataset demonstrate that SFDANet
achieves superior accuracy and long-term stability, while it maintains remarkable efficiency
with only 0.65M parameters. The code is available at [Link]
Keywords: Cloud prediction; Spatiotemporal-frequency; Differential pyramid; Multi-order
projection

Yunting Yang, Jun Liu, Hongsi Liu, Guangfeng Jiang,


Radar M3-Net: Multi-scale, multi-layer, multi-frame network with a large receptive field for
3D object detection,
Expert Systems with Applications,
Volume 286,
2025,
127515,
ISSN 0957-4174,
[Link]
([Link]
Abstract: 4D millimeter-wave radar has demonstrated significant potential for 3D object
detection in autonomous driving due to its cost-effectiveness and robustness. However, the
inherent sparsity of radar data poses a significant challenge to achieving accurate 3D object
detection, as it limits the amount of meaningful information available for feature learning.
The key to addressing the performance degradation caused by sparsity lies in expanding the
receptive field and enriching feature representations. To this end, we propose a multi-scale,
multi-layer, multi-frame network with a large receptive field, named Radar M3-Net. First, we
design a multi-scale voxel feature encoding (MSVFE) module and a multi-layer attention
(MLA) module, both of which significantly expand the receptive field and enrich features,
effectively addressing the issue of sparsity. Then, a multi-frame fusion module is developed
to further enhance features by utilizing the accumulation of temporal information.
Simultaneously, we design a novel sparse dual-head within the sparse framework to address
the decline in detection accuracy for large object categories caused by radar sparsity.
Extensive experiments on the View-of-Delft and TJ4DRadSet datasets have confirmed the
advancement and effectiveness of our network. Specifically, our method achieves state-of-
the-art mean average precision (mAP) performance on both datasets, even outperforming
some multimodal approaches in certain metrics.
Keywords: 4D radar; 3D object detection; Autonomous driving; Receptive field

Bao-Lin Ye, Peng Wu, Lingxi Li, Weimin Wu, Bo Song, Xianchao Zhang,
Multi-intersection traffic signal control based on dynamic spatiotemporal memory enhanced
learning,
Control Engineering Practice,
Volume 165,
2025,
106606,
ISSN 0967-0661,
[Link]
([Link]
Abstract: In multi-intersection traffic signal control, spatial information contains rich traffic
state features, including intersection topology and lane associations. Effectively extracting
and integrating this spatial information is crucial for accurately characterizing the evolution
of traffic state. However, most existing methods rely on static graph structures and,
therefore, cannot dynamically model the spatiotemporal correlations of traffic states,
limiting their adaptability to real-time traffic scenarios. To address these limitations, we
propose a multi-intersection traffic signal control method based on dynamic spatiotemporal
memory enhanced learning (DSMEL). First, we develop an adaptive update mechanism for
dynamic heterogeneous graphs that analyzes correlations among heterogeneous traffic
features in real time to generate adaptive representations of spatial relationships. Second,
we introduce a dual-memory enhancement model based on spatiotemporal decoupling that
uses a multi-head attention mechanism to process spatial and temporal features separately.
Specifically, we design a temporal memory module to model periodic temporal dynamics
within the road network and a spatial memory module to track the evolving topological
relationships among traffic nodes. This enables a fine-grained capture of dynamic
spatiotemporal features within the road network. Finally, we propose an adaptive weight
learning method based on double entropy regularization that incorporates online learning of
dynamic game weight matrices and integrates entropy-constrained policy optimization with
novel reward and loss functions to enhance system stability and promote optimal
convergence in multi-agent coordination. Extensive experiments on synthetic and real-world
scenarios show that, compared with baseline methods, DSMEL reduces queue length by
27.19% to 49.89%, occupancy rate by 10.93% to 43.94%, and vehicle count by 11.39% to
43.68%. Furthermore, DSMEL demonstrated superior performance over both traditional
traffic signal control methods and reinforcement learning-based approaches in extreme
traffic scenarios, reducing the average queue length by 21.23% and the maximum queue
length by 18.20%.
Keywords: Deep reinforcement learning; Traffic signal control; Multi-agent; Dynamic graph

Pengfei Wang, Peilin Shu, MingHao Yang, Hongqiu Zhang, Jianqi Wang, Cong Wang, Hongbo
Jia,
Dual-task physiological learning for radar-based continuous blood pressure monitoring:
classification-regularized regression,
Biomedical Signal Processing and Control,
Volume 112, Part C,
2026,
108790,
ISSN 1746-8094,
[Link]
([Link]
Abstract: Continuous blood pressure (BP) monitoring is critical for hypertension
management, yet conventional non-contact radar systems face challenges such as feature
space overlap under low signal-to-noise ratio conditions and inaccurate estimation during
physiological state transitions due to the susceptibility of millimeter-level cardiovascular
vibration signals to environmental interference. To address these challenges, we propose a
dual-task learning model integrating classification-constrained regression with multi-scale
spatiotemporal feature extraction. Our framework combines: (1) A hybrid ResNet-BiGRU
backbone capturing waveform morphology and hemodynamic continuity through multi-
scale convolutions (kernels = 15/7/3) and triple-layer bidirectional gating, enabling robust 2-
second-interval predictions; (2) A physiological regularization mechanism where
classification-derived BP-range probabilities (10 mmHg bins) constrain regression outputs,
suppressing implausible fluctuations during dynamic states. Validated on 30 subjects across
resting, Valsalva, and tilt-table tests, results indicate clinically relevant accuracy (SBP:
−0.21 ± 6.74 mmHg; DBP: 0.25 ± 4.81 mmHg) at 0.5 Hz sampling rate, while demonstrating
improved dynamic-state performance versus benchmarks with DBP RMSE reductions up to
8.4 % by dual-task strategy. This work suggests dual-task learning can mitigate radar-specific
SNR constraints and physiological nonstationarity while fulfilling clinical real-time monitoring
demands (beat-to-beat resolution), contributing to practical deployment of cuffless BP
devices.
Keywords: Blood pressure; Radar; Dual-task learning; Non-contact monitoring; ResNet;
Feature combination

Hongjin Chen, Kanghui Zhou, Zhonghua Zheng, Lei Han, Yongguang Zheng,
TorViNet: A spatiotemporal deep learning network for tornado detection in user-captured
social media videos,
Expert Systems with Applications,
Volume 308,
2026,
131093,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Tornadoes are extremely destructive yet short-lived severe weather events, and
their rapid development often escapes timely confirmation by conventional meteorological
instruments. In recent years, the widespread availability of smartphones and the
proliferation of social media platforms have enabled user-captured videos to become a
valuable supplementary source for real-time severe weather monitoring. Weather radar may
indicate tornadic signatures but often cannot verify whether a tornado has touched down. In
contrast, social-media videos provide direct visual evidence of actual occurrence, making
them valuable for ground-truth validation and situational awareness. Accordingly, this study
presents TorViNet, an AI-driven spatiotemporal recognition framework designed as a
decision-support module for expert systems to detect tornadoes directly from user-captured
social media videos. To address real-world challenges such as redundant or irrelevant
frames, dynamic occlusions, and small-scale visual targets, TorViNet integrates three
domain-adapted components: a dynamic frame-level selection strategy that filters out
temporally uninformative content, a spatial-frequency attention mechanism that enhances
fine-grained vortex structures, and a contrast-aware refinement module that suppresses
background distractions. Experiments on a newly curated dataset of over 10,000 verified
tornado and non-tornado clips demonstrate that TorViNet achieves 91% accuracy and an F1-
score of 0.89, outperforming a wide range of dominant video classification models. Its
robustness under noisy, unstable, and far-visibility conditions highlights its potential for
integration into operational meteorological expert systems, providing timely situational
awareness and enhancing early warning capabilities for severe tornado events.
Keywords: Tornado detection; Social media; Meteorology systems; Deep learning; Attention
mechanism

Jiahao Liu, Yiming Zhang, Liang Song, Zheng Tong,


SCB-ADAE: An attention-based deep autoencoder for ground penetrating radar signal
denoising,
Engineering Applications of Artificial Intelligence,
Volume 160, Part A,
2025,
111902,
ISSN 0952-1976,
[Link]
([Link]
Abstract: In buried object detection, recorded signals of a ground penetrating radar (GPR)
inevitably include noise interference owing to complex underground environments. Existing
rule- and data-driven denoising methods struggle to handle non-Gaussian and real-world
noise because the rule-driven ones rely on the assumptions of simplified noise
characteristics and the data-driven ones cannot capture fine- and global-scale features of a
GPR signal well. To address the problem, this study proposes an attention-based denoising
model called the Swin-Conv Block with Attention Denoising Autoencoder (SCB-ADAE). The
model first feeds a GPR signal into a SCB module, which extracts a tensor with the fine-scale
features in the signal, such as sharp reflective interfaces and abrupt amplitude variations.
The feature tensor then passes through an ADAE module that uses encoder-decoder
structure with the self-attention to enhances the representation of the global-scale signal
features. Finally, the feature tensor from the ADAE module is decoded by another SCB
module to generate a denoised GPR signal, where the tensor includes the fine-scale and
global features of the raw signal. An experiment with three types of GPR signals
demonstrates the effectiveness of the proposed model: radar signals with Gaussian noise,
radar signals with inhomogeneous-material noise, and real-world signals. radar signals with
Gaussian noise, radar signals with inhomogeneous-material noise, and real-world signals.
Experimental results demonstrate that the proposed model outperforms other state-of-the-
art denoising methods on denosing the three types of GPR signals, where the signal-to-noise
ratio, peak signal-to-noise ratio, and structural similarity index are improved to 20.64, 14.59,
and 0.366, respectively.
Keywords: Ground penetrating radar; Denoising; Attention-based model; Autoencoder

Yong He, Zi-Long Duan, Xiang-Hong Ding, Zhao Zhang, Raud Eucaristia Mayoulou, Kao-Fei
Zhu,
Spatiotemporal prediction for groundwater heavy metal contamination using Soft-DTW-
based clustering and graph neural network framework,
Water Research,
Volume 291,
2026,
125245,
ISSN 0043-1354,
[Link]
([Link]
Abstract: Accurate prediction of groundwater heavy metal contaminant spatiotemporal
dynamics is essential for monitoring optimization and remediation decision-making at
contaminated sites. However, heterogeneous contamination distribution and complex
spatiotemporal correlations among monitoring wells pose significant prediction challenges.
In this study, a Soft Dynamic Time Warping clustering-based Graph Neural Network
(SDCGNN) was proposed for contamination zone identification, integrating multi-scale
spatiotemporal modeling. Using Soft Dynamic Time Warping (Soft-DTW) distance-based
clustering, the model partitions monitoring wells into source, plume, and attenuation zones
according to temporal contamination patterns, while a hierarchical local-global graph fusion
framework captures both zone-specific dynamics and cross-zone transport processes.
Evaluation on two-year hourly monitoring data from 25 monitoring wells at on-site
contaminated sites demonstrated that SDCGNN achieved average Mean Absolute Error
(MAE) of 0.213 mg/L and Mean Absolute Percentage Error (MAPE) of 5.51%, improving upon
baseline models by 49.9% and 61.4%, respectively. Zone-specific predictions revealed high
accuracy across source, plume, and attenuation areas despite varying concentration ranges
and spatial heterogeneity. Furthermore, spatiotemporal analysis confirmed that the model
accurately reproduced observed contamination transport patterns, including plume
migration directions and concentration gradient evolution. The proposed zone-aware
modeling approach shows promise for advancing groundwater heavy metal contamination
prediction capabilities and facilitates the optimization of site remediation strategies.
Keywords: Heavy metal contamination; Dynamic time warping; Graph clustering; Graph
neural networks
Wei Tian, Lei Yi, Xianghua Niu, Rong Fang, Lixia Zhang, Huanhuan Liu, Zhuo Xu, Shengqin
Jiang, Yonghong Zhang,
RadarNet: A parallel spatiotemporal encoder network for radar extrapolation,
Neurocomputing,
Volume 591,
2024,
127665,
ISSN 0925-2312,
[Link]
([Link]
Abstract: Radar extrapolation has been one of the most important means for nowcasting.
Most current models achieve good performance in high-frequency sequences (e.g., video,
more than 24 fps), while the temporal resolution of radar echo sequences is much lower (1
frame every 6 min) and the transforms are much more complex. The spatiotemporal
characters with some similarities would not change a lot in video sequences; however, the
radar echo sequences include more intangible changes (e.g., the echo evolution of
generation or vanish, and so on), which leads to unique distinct spatial and temporal
characters, respectively. Therefore, the singular peculiarity would be mitigated, leading to a
rapid decline in precision and sharpness during the extrapolation process. In general,
temporal feature extraction is utilized to understand the variation in pixel locations, while
spatial feature extraction is employed to capture the distribution variation of specific
regions. In this work, we propose a feature decomposition network, termed as RadarNet to
improve the extrapolation precision. The parallel independent encoders are used to enhance
multi-scale spatial feature extraction and temporal motion feature capture of radar echoes,
respectively. In addition, we design a specialized cross fusion mechanism to achieve the
inputs of the decoder which may enhance the performance of the extrapolation. The
extrapolation experiments are conducted on real radar echo datasets from Shijiazhuang and
Nanjing that demonstrate the effectiveness of our model.
Keywords: RadarNet; Radar extrapolation; Spatiotemporal prediction

Ping Zhou, Fanfan Gan, Baizhan Xia,


A multiscale separable convolution network with lightweight spatiotemporal attention for
remaining useful life prediction,
Applied Soft Computing,
Volume 186, Part A,
2026,
114122,
ISSN 1568-4946,
[Link]
([Link]
Abstract: Accurate prediction of remaining useful life (RUL) for critical equipment guides
operational strategy formulation, making it a pivotal challenge in prognostics and health
management (PHM). Existing temporal convolutional networks (TCNs) address limitations in
spatial feature representation by integrating attention mechanisms or multiscale
architecture. However, these integrations exhibit a critical limitation: prohibitive
computational burden limiting the deployment in resource-constrained conditions. Here, we
propose a multiscale separable convolution network with lightweight spatiotemporal
attention (MSSN-LST) for RUL prediction of the aero engine. First, a multiscale 1D separable
dilated causal convolutional network (MSSN) is developed to extract multiscale temporal
degradation features. By introducing varying kernel sizes and dilation factors within 1D
separable convolution, the MSSN captures multiscale degradation features while preserving
temporal directionality. Besides, the introduction of dilation factor efficiently expands the
receptive field, which helps to reduce the network depth of separable convolution.
Subsequently, a lightweight spatiotemporal attention layer (LST) is constructed to enhance
feature representation across both temporal and spatial dimensions. By utilizing non-
parametric computational structure to enhance spatiotemporal feature representation, the
LST maintains the minimal computational burden. This lightweight attention model
enhances the ability of the model to be deployed on edge devices. Finally, a regression layer
incorporating a separable convolution layer and a linear layer outputs the RUL prediction.
Results on the C-MAPSS dataset demonstrate that the proposed model achieves superior
prediction accuracy with light computational burden.
Keywords: Multiscale separable convolution network; Lightweight spatiotemporal attention;
Remaining useful life prediction; Aero engine

Sakshi Dhankhar, Stefan Wittek, Hamidreza Eivazi, Andreas Rausch,


R2RNet: A deep spatiotemporal RaintoRiverNetwork for water level prediction and flood
forecasting,
Journal of Hydrology: Regional Studies,
Volume 61,
2025,
102571,
ISSN 2214-5818,
[Link]
([Link]
Abstract: Study Region:
Goslar and Göttingen, Lower Saxony, Germany.
Study Focus:
In July 2017, the cities of Goslar and Göttingen experienced severe regional flood events
characterized by short warning time of only 20 min with a significant damage. This highlights
the critical need for a more reliable and timely flood forecasting system. This paper presents
a comprehensive study on the impact of radar-based precipitation data on forecasting river
water levels in Goslar and Göttingen. The analysis integrates radar-derived spatiotemporal
precipitation patterns with hydrological sensor data obtained from ground stations to
evaluate the effectiveness of this approach in improving flood prediction capabilities.
New Hydrological Insights for the Region:
A key innovation in this paper is the use of residual-based modeling to address the non-
linearity between precipitation images and water levels, leading to a Rain2RiverNetwork
with residual-corrections (R2RNet). The deep learning architecture integrates (2+1)D
convolutional neural networks for spatial and temporal feature extraction with LSTM for
timeseries forecasting. The results of our evaluation demonstrate the potential of R2RNet
for capturing extreme events with good prediction accuracy for 4-hours ahead as Nash–
Sutcliffe efficiency=0.93, Bravais–Pearson=0.98 and Index of Agreement=0.99. These results
are comparable to the conventional model using upstream data for predictions, with
NSE=0.95, BP=0.97 and IoA=0.98. These results quantitatively underscore the enhanced
predictive capability of R2RNet, making it independent of additional hydrological upstream
measurements.
Keywords: Flood forecasting; Radar precipitation; Extreme event prediction; Deep learning

Nana Chu, Kam K.H. Ng, Xinting Zhu, Ye Liu, Lishuai Li, Kai Kwong Hon,
Towards dynamic flight separation in final approach: A hybrid attention-based deep learning
framework for long-term spatiotemporal wake vortex prediction,
Transportation Research Part C: Emerging Technologies,
Volume 169,
2024,
104876,
ISSN 0968-090X,
[Link]
([Link]
Abstract: The conservative and distance-based static wake vortex-related separation may
restrict runway operational efficiency. Recent studies have demonstrated the potential of
wake separation reduction under the Re-categorisation scheme of Aircraft Weight (RECAT).
Furthermore, dynamic time-based flight separation considering vortex evolution with
respect to aircraft pairs and meteorological conditions will be the ultimate objective for
improving runway operational capacity without compromising safety. This paper presents a
hybrid deep learning framework for aircraft wake vortex recognition, evolution prediction,
and preliminary dynamic separation assessment in the final approach. Two-stage Deep
Convolutional Neural Networks (DCNNs) are utilised to identify vortex locations and strength
from wake images. Subsequently, we propose the Attention-based Temporal Convolutional
Networks (ATCNs) for future long-term vortex decay and transport forecasts based on initial
vortex information from DCNNs. 17,254 wake sequences generated by arrival flights at Hong
Kong International Airport (HKIA) are used in this study. The proposed ATCN models
outperform the specific benchmarks. Furthermore, the hybrid DCNN-ATCN model shows
great benefits in mining both spatial vortex characteristics and temporal dependencies in
vortex evolution, and achieves a computational speed of approximately 7 s per sequence.
The final vortex duration assessment demonstrates a significant potential for separation
reduction in the final approach when the crosswind speed exceeds 3 m/s. This study
provides important implications for online and fast-time wake behaviour monitoring and
state estimation. The results of vortex duration analysis conform to the RECAT-EU standards
and present an efficient strategy for developing dynamic flight separation systems.
Keywords: Flight separation; Aircraft wake turbulence; Recurrent neural network; Attention
mechanism; LiDAR

Wei Zhang, Xinyu Zhang, Junyu Dong, Xiaojiang Song, Renbo Pang,
CIDM: A comprehensive inpainting diffusion model for missing weather radar data with
knowledge guidance,
ISPRS Journal of Photogrammetry and Remote Sensing,
Volume 221,
2025,
Pages 299-309,
ISSN 0924-2716,
[Link]
([Link]
Abstract: Addressing data gaps in meteorological radar scan regions remains a significant
challenge. Existing radar data recovery methods tend to perform poorly under different
types of missing data scenarios, often due to over-smoothing. The actual scenarios
represented by radar data are complex and diverse, making it difficult to simulate missing
data. Recent developments in generative models have yielded new solutions for the problem
of missing data in complex scenarios. Here, we propose a comprehensive inpainting
diffusion model (CIDM) for weather radar data, which improves the sampling approach of
the original diffusion model. This method utilises prior knowledge from known regions to
guide the generation of missing information. The CIDM formalises domain knowledge into
generative models, treating the problem of weather radar completion as a generative task,
eliminating the need for complex data preprocessing. During the inference phase, prior
knowledge of known regions guides the process and incorporates domain knowledge
learned by the model to generate information for missing regions, thus supporting radar
data recovery in scenarios with arbitrary missing data. Experiments were conducted on
various missing data scenarios using Multi-Radar/MultiSensor System data sourced from the
National Oceanic and Atmospheric Administration, and the results were compared with
those of traditional and deep learning radar restoration methods. Compared with these
methods, the CIDM demonstrated superior recovery performance for various missing data
scenarios, particularly those with extreme amounts of missing data, in which the restoration
accuracy was improved by 5%–35%. These results indicate the significant potential of the
CIDM for quantitative applications. The proposed method showcases the capability of
generative models in creating fine-grained data for remote sensing applications.
Keywords: Weather radar data; Comprehensive inpainting; Diffusion models; Extreme
missing cases; Knowledge guidance

Yuankang Ye, Feng Gao, Shaoqing Zhang, Chang Liu,


Improving precipitation nowcasting via multiphysical parameter fusion in radar echo
extrapolation,
Journal of Hydrology,
Volume 668,
2026,
134947,
ISSN 0022-1694,
[Link]
([Link]
Abstract: Radar-based precipitation nowcasting plays a vital role in short-term
hydrometeorological forecasting and water resource management. Existing modeling
methodologies typically simplify precipitation nowcasting to a task of spatiotemporal
sequence prediction based on radar echo reflectivity data. However, the reliance on
unimodal reflectivity data including intensity-only information restricts the model’s ability to
characterize the phase evolution and dynamic processes of hydrometeor particles,
ultimately leading to insufficient extrapolation accuracy. This study breaks through the
conventional unimodal data paradigm, aiming to capture the complex dynamic evolutionary
features of hydrometeor particles. We integrate radar echo reflectivity and four additional
physical parameters of hydrometeor particles into a deep learning framework and propose a
novel Physics-Informed Multimodal Echo Extrapolation neural network (PIEE). Furthermore,
we systematically investigate the individual contributions of each physical parameter to the
accuracy of radar echo extrapolation. Specifically, PIEE adopts a three-stage structure. First, a
multimodal encoder with a dual-branch attention-based fusion strategy is used to capture
diverse physical signals. Second, a novel gated spatiotemporal self-attention module is
designed for deep feature extraction. Finally, the decoding stage generates the extrapolated
radar echoes. Experimental results on a real multimodal radar echo dataset show that the
proposed model demonstrates superior performance in two aspects. First, under a unimodal
baseline architecture, the PIEE model clearly outperforms the comparison model. Second,
after fusing multiple physical parameters, the PIEE achieves significant improvements in all
the evaluated metrics, especially in the CSI and HSS metrics for the high echo intensity
region (≥ 40 dBZ), with improvements of up to 24.2% and 20.3%, respectively. Furthermore,
systematic ablation experiments on physical parameters quantify the effects of different
combination methods on extrapolation accuracy, highlighting the potential of physics-
informed, multimodal deep learning approaches in improving short-term hydrological
prediction accuracy, with implications for flood forecasting, early warning systems, and
hydrometeorological risk management at catchment scales.
Keywords: Deep learning; Precipitation nowcasting; Radar echo extrapolation;
Hydrometeorological forecasting

Qiangyu Zeng, Ling Li, Hao Wang, Jianxin He, Hua Wang, Yao Gao,
MCDA-UNet: A satellite data-based model for radar composite reflectivity retrieval,
Atmospheric Research,
Volume 330,
2026,
108619,
ISSN 0169-8095,
[Link]
([Link]
Abstract: The weather radar network in China exhibits an uneven spatial distribution, with
dense coverage in the eastern regions and sparse deployment in the west, resulting in
substantial detection blind spots in areas with complex terrain. This severely limits the
continuity and precision of weather monitoring and early warning in these regions. To
address this challenge, a multi-channel deep learning model, MCDA-UNet, is proposed for
radar composite reflectivity retrieval, aiming to reconstruct and enhance radar echo patterns
in regions lacking radar coverage by leveraging the extensive spatial coverage and
continuous observation capabilities of geostationary meteorological satellites. The model
employs a multi-channel input architecture to extract features from different spectral bands,
while spatial and channel attention models are incorporated to improve the representation
of key meteorological information, thereby enhancing retrieval accuracy and regional
adaptability. Comparative experiments conducted under varying precipitation intensities
demonstrate that MCDA-UNet consistently outperforms existing models across multiple
evaluation metrics, particularly in reconstructing weather radar echo structures and edge
details. These results validate the model’s capability to adapt to the full dynamic range of
weather radar reflectivity and highlight its potential for accurate precipitation retrieval in
radar blind-spot regions.
Keywords: Weather radar composite reflectivity; Satellite data retrieval; Multi-channel
structure; Full dynamic range

Kitoshi Kawai, Bungo Konishi, Ryo Natsuaki, Akira Hirose,


Quaternion reservoir computing for spatiotemporal analysis in polarimetric synthetic
aperture radar,
Neurocomputing,
Volume 658,
2025,
131633,
ISSN 0925-2312,
[Link]
([Link]
Abstract: Quaternion neural networks possess high generalization ability in three-
dimensional (3D) information space by representing every 3D data point as a single
quaternion entity. In polarimetric synthetic aperture radar (PolSAR) applications such as land
surface classification, they are expected to deal with 3D Poincare parameters as inseparable
physical entities. With the increasing acquisition frequency, there is a growing demand also
for efficient and robust techniques to monitor temporal or spatiotemporal changes.
Reservoir computing (RC) is a variation of recurrent neural networks (RNNs) capable of
detecting changes in series data with low computational cost. In this context, we propose
quaternion reservoir computing (QRC) for spatiotemporal analysis in PolSAR. First, in a
benchmark prediction task for chaotic time-series derived from the 3D Lorenz equations, we
demonstrate that QRC achieves higher prediction accuracy than real-valued RC and
conventional RNNs. Secondly, we conduct spatiotemporal anomalous change detection for
actual PolSAR data of (1) rice fields having seasonal changes in Japan and (2) Amazon
rainforest suffering from deforestation in Brazil. Compared with real-valued RC, RNNs, one-
dimensional convolutional neural networks, Transformer, and non-adaptive methods based
on complex Wishart and Pauli RGB, QRC shows a larger area under the curve (AUC) score,
demonstrating its high efficacy in capturing spatiotemporal anomalous variations in PolSAR
data. Besides such high performance, QRC shows low training cost, which is very suitable for
real-time processing in edge computing including highly frequent satellite observations.
These experimental results indicate that combining quaternion representation with RC is a
promising approach for analyzing the ever-increasing volume of PolSAR data.
Keywords: Quaternion; Reservoir computing; Polarimetric synthetic aperture radar (PolSAR);
Spatiotemporal analysis

Pengfei Jia, Helmi Zulhaidi Mohd Shafri, Shengrui Yu, Zhi Zheng, Shiqing You, Abdul Rashid
Mohamed Shariff,
Radar-optical fusion of Sentinel-1/2 for high-resolution NDVI reconstruction and landscape-
driven carbon flux assessment in Kuala Selangor, Malaysia (2020–2024),
International Journal of Applied Earth Observation and Geoinformation,
Volume 145,
2025,
104966,
ISSN 1569-8432,
[Link]
([Link]
Abstract: Reliable quantification of carbon fluxes in humid tropical regions is constrained by
persistent cloud cover, heterogeneous land mosaics, and the limited resolution of existing
products. To address these challenges, this study developed a Cloud-Resilient Fusion
Network (CRFNet) that integrates Sentinel-1 SAR backscatter with cloud-screened Sentinel-2
NDVI using a CNN–BiLSTM–multi-head attention architecture. The framework reconstructed
10 m NDVI time series in Kuala Selangor, Malaysia (2020–2024), achieving annual R2 above
0.82 and RMSE below 0.12, thereby improving temporal continuity under heavy cloud–
rainfall interference. The reconstructed NDVI was used to drive a light-use-efficiency model
for net ecosystem productivity (NEP) estimation, supported by temperature-based
heterotrophic respiration. Results showed a 7.4 % decline in mean annual NEP across five
years, with degraded mangroves and sloping croplands emerging as hotspots of sink-to-
source transitions. Landscape analysis revealed strong structure–function coupling: stable
forests and mangroves were characterized by large cohesive patches with largest patch index
values above 40 % and edge density below 20 m ha-1, while croplands and degraded slopes
exhibited higher patch numbers, reduced patch dominance, and greater edge complexity,
which increased carbon source risk. By linking fine-scale NDVI reconstruction with process-
based carbon modeling and landscape metrics, this study provides a transferable workflow
for high-resolution carbon flux monitoring and a robust scientific basis for carbon budget
assessment, ecosystem management, and carbon-neutrality planning in tropical monsoon
regions.
Keywords: Cloud-resilient fusion network (CRFNet); NDVI reconstruction; Radar–optical
fusion; Net ecosystem productivity; Landscape metrics

Mahdi Bonyani, Maryam Soleymani, Chao Wang,


Construction workers' unsafe behavior detection through adaptive spatiotemporal sampling
and optimized attention based video monitoring,
Automation in Construction,
Volume 165,
2024,
105508,
ISSN 0926-5805,
[Link]
([Link]
Abstract: In recent years, advances in construction site image analysis faced challenges,
particularly in construction object detection and identifying unsafe actions. Challenges
involve complex backgrounds, varying object sizes, and image quality. Existing methods
address spatial and temporal features with attention mechanisms but often overlook
adaptive sampling and channel-wise adjustments, missing potential spatiotemporal
redundancies. This article introduces the Optimized Positioning (OP-Net) architectures and
an attention-based spatiotemporal sampling approach. The OP module is introduced for
object detection, which enhances channel relationships by leveraging global feature affinity
associations. Additionally, we propose an innovative spatiotemporal sampling strategy that
adapts to effectively identify unsafe actions in construction sites. We extensively evaluate
the object detection task using the SODA dataset to showcase the efficacy and effectiveness
of our approach. Furthermore, our unsafe action identification model is benchmarked on the
CMA dataset, demonstrating its ability to achieve new state-of-the-art performance in
accuracy while maintaining reasonable computational efficiency.
Keywords: Unsafe behavior detection; Construction worker safety; Adaptive spatiotemporal
sampling; Attention learning; Video analysis; Deep learning

Weidong Fang, Xibin Lin, Ji Zhang, Jiacheng Hu, Linrun Huang, Guangqian Yuan,
Spatiotemporal charging demand forecasting for EV stations via cross-attention fusion,
Applied Soft Computing,
Volume 189,
2026,
114475,
ISSN 1568-4946,
[Link]
([Link]
Abstract: With the rapid growth of electric vehicles (EVs), accurately predicting charging
demand has become crucial for intelligent transportation and energy management. Existing
deep learning methods usually neglect the integration of economic principles with complex
spatiotemporal dependencies. To this end, this paper proposes a novel framework, the Bi-
CAPNet model, which integrates a bidirectional temporal convolutional network (BiTCN), a
bidirectional gated recurrent unit (BiGRU), a crisscross attention mechanism, an economics-
informed neural network (EINN), and a sparrow search algorithm (SSA). This architecture
captures multi-scale temporal features, models sequential dependencies, integrates
spatiotemporal data, and incorporates price–demand elasticity to characterize the response
of charging demand to dynamic pricing, thereby improving economic interpretability.
Utilizing extensive real-world data collected from Shenzhen, the empirical evaluation
confirms that the proposed approach substantially outperforms representative baseline
models in terms of both predictive accuracy and economic interpretability, indicating its
strong potential for urban EV charging demand forecasting.
Keywords: Charging demand forecasting; Economics-informed neural networks;
Spatiotemporal feature fusion; Cross-attention mechanism

Cries Avian, Jenq-Shiou Leu, Hang Song, Jun-ichi Takada, Nur Achmad Sulistyo Putro,
Muhammad Izzuddin Mahali, Setya Widyawan Prakosa,
RCTrans-Net: A spatiotemporal model for fast-time human detection behind walls using
ultrawideband radar,
Computers and Electrical Engineering,
Volume 120, Part C,
2024,
109873,
ISSN 0045-7906,
[Link]
([Link]
Abstract: Ultrawideband (UWB) radar systems are becoming increasingly popular for
detecting human presence, even through walls. Recent advancements in signal processing
use deep learning techniques, which are known for their accuracy. While earlier methods
focused on spatial information using Convolutional Neural Networks (CNNs), newer research
highlights the importance of temporal information, such as how data peaks shift over time.
This study introduces RCTrans-Net, a deep-learning architecture that combines RCNet (a
Residual CNN) for spatial features with TransNet (a Transformer) for temporal features. This
fusion improves human presence classification in fast-time signal processing. Tested under
various conditions—different materials, body orientations, ranges, and radar heights—
RCTrans-Net achieved high performance with F1-scores of 0.997±0.000 for static,
0.967±0.004 for dynamic, and 0.978±0.001 for combined scenarios. The architecture
outperforms previous methods and offers real-time processing with an inference time of
about one millisecond.
Keywords: Human presence behind the wall; Residual network; Spatiotemporal'
Transformer; Ultrawideband radar system

Pengfei Yang, Feng Wu, Minyang Liu, Ting Zhong, Fan Zhou,
Beyond pillars: Advancing 3D object detection with salient voxel enhancement of liDAR-4D
radar fusion,
Pattern Recognition,
Volume 173,
2026,
112841,
ISSN 0031-3203,
[Link]
([Link]
Abstract: The fusion of LiDAR and 4D radar has emerged as a promising solution for robust
and accurate 3D object detection in complex and adverse conditions. Existing methods
typically rely on pillar-based representations, which, although computationally efficient, fail
to provide fine-grained structural details necessary for precise object localization and
recognition. In contrast, voxel-based representations offer richer spatial information but face
challenges such as background noise and data quality disparity. To address these limitations,
we propose SVEFusion, a voxel-based 3D object detection framework that integrates LiDAR
and 4D radar data using a salient voxel enhancement mechanism. Our method introduces an
adaptive feature alignment module and a novel spatial neighborhood attention module for
efficient early-stage multi-modal voxel feature integration. Furthermore, we design a salient
voxel enhancement mechanism that assigns higher weights to foreground voxels using a
multi-scale weight prediction strategy, progressively refining weight accuracy with
supervision loss. Experimental results demonstrate that SVEFusion significantly outperforms
state-of-the-art methods, establishing a new benchmark in multi-modal 3D object detection.
The source code and network weighting for reproducibility are available at
[Link]
Keywords: Object detection; Lidar; 4D Radar; Multi-modal fusion; Autonomous driving

Jinbo Fu, Hong Cao, Zhe Wang, Kuan Chang, Haitao Wang, Bo Chen, Jiuchun Sun,
WT-DANet-STAF: A spatiotemporal adaptive fusion-based denoising method for foundation
pit enclosure structure deformation data,
Measurement,
Volume 263,
2026,
120155,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Monitoring deformation in foundation pit enclosure structures is crucial, serving as
the foundation for predicting deformations, issuing early warnings for exceedances, and
implementing effective control measures. However, the collected monitoring data often
contain noise, necessitating robust denoising techniques to ensure accuracy. Existing
denoising methods for foundation pit monitoring data often struggle to effectively remove
noise while preserving key features. To address these challenges, a novel method based on
Wavelet transform (WT) and DenseNet-Attention (DANet) is proposed for denoising
deformation data of foundation pit enclosure structures, complemented by a spatiotemporal
adaptive fusion network (STAF) to optimize the results further. Initially, Wavelet transform is
applied to decompose the deformation data across spatial and temporal dimensions at
multiple scales, extracting detailed and approximation components at varying frequency
levels. Subsequently, DANet is employed to denoise and reconstruct these decomposed
components, yielding denoised data in both spatial and temporal dimensions. Finally, the
spatiotemporal adaptive fusion network integrates the denoised outputs from spatial and
temporal dimensions to generate high-quality results. The effectiveness of the proposed
method is validated through experiments using data from a road improvement project in
southern China. Comparisons with mainstream denoising techniques and evaluations via
multiple metrics demonstrate the superior performance of the proposed approach in noise
reduction.
Keywords: Foundation pit; Deformation monitoring; Wavelet transform; DenseNet-Attention
network; Spatiotemporal fusion denoising

Sasan Babaee, Mohammad Amin Khalili, Rita Chirico, Anna Sorrentino, Diego Di Martire,
Spatiotemporal characterization of the subsidence and change detection in Tehran plain
(Iran) using InSAR observations and Landsat 8 satellite imagery,
Remote Sensing Applications: Society and Environment,
Volume 36,
2024,
101290,
ISSN 2352-9385,
[Link]
([Link]
Abstract: Urban areas worldwide are increasingly facing challenges related to land
subsidence, a phenomenon exacerbated by uncontrolled groundwater extraction and urban
expansion. This research focuses on the Tehran plain, Iran's capital city, where significant
subsidence has been observed due to uncontrolled migrations influenced by various
economic and political factors. This expansion has increased demand for energy, notably
water, leading to irregular water withdrawals from underground sources and, consequently,
land subsidence. Monitoring this subsidence, particularly its effects on urban infrastructure,
has become a critical challenge. This research first reviewed the existing body of knowledge
related to subsidence measurement in the Tehran plain with an emphasis on their findings
and limitations and then used radar images to study the subsidence patterns in the Tehran
plain from 2016 to the end of 2020. Finally, the results collaborated by optical imagery
analysis to find the relationship between surface change detection and spatiotemporal
distribution of subsidence. As a result, through processing Sentinel-1A SAR images,
consistent vertical displacements (subsidence) were observed, especially in areas heavily
reliant on groundwater from wells, with some areas experiencing a rate of more than
−20 mm/year. Horizontal displacement, however, was approximately about ±8 mm/year.
Also, our results show that the subsidence rate in this plain has decreased in recent years.
Therefore, the study integrated multispectral satellite data to clarify this issue and
compensate for missing groundwater level data, specifically the Normalized-Difference
Vegetation Index (NDVI) and Normalized-Difference Moisture Index (NDMI). These datasets
were used to monitor changes in vegetation cover distribution and moisture in response to
the variations of groundwater depth over time. The results of this research can be beneficial
in adequately managing groundwater resource utilization to reduce the potential damage to
infrastructure and the environment.
Keywords: Spatiotemporal subsidence pattern; Radar interferometry; Sentinel-1A; Tehran
plain

Yongchao Zhu, Qiuling Lu, Maorong Ge, Xiaochuan Qu, Tingye Tao, Kegen Yu, Shuiping Li,
Attention enhanced ResNet for ocean surface wind speed retrieval using CYGNSS
observables,
Advances in Space Research,
2025,
,
ISSN 0273-1177,
[Link]
([Link]
Abstract: Global Navigation Satellite System Reflectometry (GNSS-R) has emerged as a
pivotal technique for ocean surface wind speed retrieval; however, establishing robust multi-
parameter retrieval models remains challenging due to the nonlinear relationships between
GNSS-R observables and geophysical variables. An Attention-enhanced Residual Network
(Att-ResNet) is proposed to address this challenge, leveraging Cyclone Global Navigation
Satellite System (CYGNSS) bistatic radar data for wind speed estimation. The CYGNSS
datasets were processed to extract multi-parameter observables, including Delay-Doppler
Maps (DDMs), normalized bistatic radar cross-section (NBRCS), and incidence angle, which
served as inputs for training wind speed retrieval models using diverse backbone
architectures (e.g., ResNet and AlexNet). Ablation experiments employing the Att-ResNet
framework were systematically conducted, with ERA5 (European Centre for Medium-Range
Weather Forecasts Reanalysis 5) and CCMP (Cross-Calibrated Multi-Platform) wind products
providing benchmark validation. Comparative analysis revealed that the Att-ResNet-retrieved
wind speeds exhibited strong spatiotemporal consistency with ERA5 and CCMP data.
Quantitative evaluations showed root mean square errors (RMSEs) of 1.379 m/s (ERA5) and
1.390 m/s (CCMP), with minimal biases (−0.069 m/s and −0.014 m/s, respectively) and
unbiased RMSEs (ubRMSEs) of 1.377 m/s and 1.390 m/s. The study demonstrates that the
Att-ResNet architecture, through its attention-driven feature selection and residual learning
mechanisms, significantly enhances spaceborne GNSS-R wind retrieval accuracy. This
artificial intelligence-driven framework establishes a new paradigm for high-resolution
spatiotemporal ocean surface wind monitoring, demonstrating the transformative potential
of deep learning in advancing GNSS-R applications.
Keywords: Residual network; GNSS-R; Wind speed; Deep learning; CYGNSS
Jiaying Li, Weidong Wang, Guangqi Chen, Zheng Han, Chongzheng Zhu, Chen Chen,
Spatiotemporal LSA modeling incorporating comprehensively the momentary effects of
rainfall and earthquake: A case study of the Liangshan Prefecture, China,
Advances in Space Research,
Volume 76, Issue 11,
2025,
Pages 6725-6740,
ISSN 0273-1177,
[Link]
([Link]
Abstract: Landslides are one of the most destructive geo-hazards, and the landslide
susceptibility assessment (LSA) can effectively reduce landslide risks and strengthen
landslide prevention. The present study explores a spatiotemporal LSA method considering
comprehensively the momentary effects of rainfall and earthquakes. Logistic regression
model, random forest model, deep belief network (DBN) model, and grey wolf optimizer
(GWO)-DBN model were used to analyze the spatial LSA, and the optimal spatial LSA
obtained using the GWO-DBN model was chosen using various evaluation metrics to analyze
the spatiotemporal LSA. Meanwhile, the historical landslide data during the year before the
study time, namely from July 5, 2020 to July 5, 2021, and the data of rainfall and earthquake
before various landslides were collected, and their effective rainfall and seismic peak ground
acceleration were calculated to construct the temporal LSA regression model. The temporal
LSA map in the study time was thus obtained and coupled with the optimal spatial LSA map
to generate the spatiotemporal LSA map. Due to dynamic changes over time of
spatiotemporal LSA, the precise landslide locations and ranges were obtained using small
baseline subset interferometric synthetic aperture radar, and the results were coupled with
spatiotemporal LSA map. There were 86.92% landslide regions with very high and high
susceptibility, and the accuracy of spatiotemporal LSA was verified, which provides a
reference for the spatiotemporal LSA verification method.
Keywords: Spatiotemporal LSA; Deep belief network model; GWO-DBN model; Temporal LSA
regression model; SBAS-InSAR

Saihan Chen, Peng Liu, Puchen Zhang, Xiaokang Ma, Ran Bao, Zixu Wang, Haixu Yang, Xiao
Ke,
Data-driven spatiotemporal fault detection in Lithium-ion batteries using isometric mapping
and modified independent component analysis,
Journal of Energy Storage,
Volume 149,
2026,
120062,
ISSN 2352-152X,
[Link]
([Link]
Abstract: Accurate detection and localization of thermal faults in lithium-ion batteries (LIBs)
are crucial for ensuring safety and preventing accidents. However, the intricate
thermodynamic behavior of large-format LIBs and battery systems, which operate as high-
dimensional distributed-parameter systems, poses significant challenges. This investigation
presents a data-driven spatiotemporal framework for detecting and localizing battery
thermal faults. Isometric mapping models the spatiotemporal dynamics of the battery's
thermal processes by decomposing high-dimensional spatiotemporal temperature data into
linear combinations of low-dimensional temporal coefficients and discrete spatial basis
functions (SBFs), with radial basis functions used to construct continuous-space SBFs.
Modified independent component analysis is further applied to temporal coefficients to
build process monitoring models and generate real-time monitoring statistics for fault
detection. Finally, spatiotemporal reconstruction yields a continuous-space contribution map
of abnormal statistics for fault localization. The proposed method is validated through
internal short circuit experiments on large-format LIBs and thermal runaway propagation
simulations of battery systems, covering 33 thermal fault scenarios under various operating
conditions. Results indicate that the method achieves high-precision fault detection and
localization using only six-dimensional temporal coefficients and corresponding SBFs, with an
average F1 score of 98.1 %, a fault detection delay of 3.4 sampling steps, and a fault
localization accuracy of 93.3 %.
Keywords: Battery thermal process; Fault detection; Fault localization; Spatial construction;
Internal short circuit; Thermal runaway

Yiwen Chen, Yuan Zhuang, Binliang Wang, Jianzhu Huai,


4D RadarPR: Context-Aware 4D Radar Place Recognition in harsh scenarios,
ISPRS Journal of Photogrammetry and Remote Sensing,
Volume 221,
2025,
Pages 210-223,
ISSN 0924-2716,
[Link]
([Link]
Abstract: Place recognition is a fundamental technology for uncrewed systems such as
robots and autonomous vehicles, enabling tasks like global localization and simultaneous
localization and mapping (SLAM). Existing Place recognition technologies based on vision or
LiDAR have made significant progress, but these sensors may degrade or fail in adverse
conditions. 4D millimeter-wave radar offers strong resistance to particles like smoke, fog,
rain, and snow, making it a promising option for robust scene perception and localization.
Therefore, we explore the characteristics of 4D radar point clouds and propose a novel
Context-Aware 4D Radar Place Recognition (4D RadarPR) method for adverse scenarios.
Specifically, we first adopt a point-based feature extraction (PFE) module to capture raw
point cloud information. On top of PFE, we propose a multi-scale context information fusion
(MCIF) module to achieve local feature extraction at different scales and adaptive fusion. To
capture global spatial relationships and integrate contextual information, the MCIF module
introduces a fusion block based on multi-head cross-attention to combine point-wise
features with local spatial features. Additionally, we explore the role of Radar Cross Section
(RCS) information in enhancing the discriminability of descriptors and propose a local RCS
relation-guided attention network to enhance local features before generating the global
descriptor. Extensive experiments are conducted on in-house datasets and public datasets,
covering various scenarios and including both long-range and short-range radar data. We
compared the proposed method with several state-of-the-art approaches, including
BevPlace++, LSP-Net, and Transloc4D, and achieved the best overall performance. Notably,
on long-range radar data, our method achieved an average Recall@1 of 89.9%,
outperforming the second-best method by 1.9%. Furthermore, our method demonstrated
acceptable generalization ability across diverse scenarios, showcasing its robustness.
Keywords: Place recognition; 4D millimeter-wave radar; Harsh scenarios; Multi-scale context
fusion; Long- and short-range radars; Relocalization

Mingyue Lu, Chuanwei Jin, Manzhu Yu, Qian Zhang, Hui Liu, Zhiyu Huang, Tongtong Dong,
MCGLN: A multimodal ConvLSTM-GAN framework for lightning nowcasting utilizing multi-
source spatiotemporal data,
Atmospheric Research,
Volume 297,
2024,
107093,
ISSN 0169-8095,
[Link]
([Link]
Abstract: Lightning phenomena can instigate a cascade of calamities, encompassing fires,
electrical infrastructure damage, and risks to human safety. Deep-learning-based lightning
nowcasting models have demonstrated significant effectiveness in disaster prevention and
mitigation. However, existing studies often neglect the impacts of surface features on
lightning activities, and conventional lightning prediction techniques based on convolutional
and recurrent networks face challenges such as the loss of feature information. Addressing
these issues, this paper presents a novel model for lightning nowcasting, the Multimodal
ConvLSTM-GAN for Lightning Nowcasting (MCGLN). This model integrates a Generative
Adversarial Network (GAN) with a Convolutional Long Short-Term Memory network
(ConvLSTM), utilizing multi-source data as inputs. It incorporates a spatiotemporal encoder-
forecaster framework within the Generator to improve the capture of multidimensional
spatiotemporal feature information, thus boosting predictive accuracy. MCGLN offers
probabilistic prediction results, allowing users to customize warning thresholds following
their specific tolerance for false and missed alarms. The performance of the MCGLN model is
evaluated through empirical analysis, utilizing real lightning datasets sourced from Zhejiang
and surrounding areas. Experimental results demonstrate that: (a) The MCGLN model
outperforms existing methods in terms of detection capability and overall performance,
showing significant improvements in the modeling process. (b) Increasing the number of
data sources improves detection capabilities, reduces the probability of false alarms, and
boosts the model performance. (c) The use of radar data enhances the recognition of high-
probability lightning occurrences, and the inclusion of surface feature data increases the
capture of terrestrial lightning genesis.
Keywords: Lightning prediction; MCGLN; Lightning nowcasting model; Ningbo

Yichen Tao, Daniel B. Wright, Abdulmuttalib M. Lokhandwala, Vitor G. Geller, Jose G.


Vasconcelos, Ben R. Hodges,
Impacts of rainfall spatiotemporal variability on pressurized flow conditions in urban
drainage systems,
Journal of Hydrology,
Volume 662, Part B,
2025,
133874,
ISSN 0022-1694,
[Link]
([Link]
Abstract: Urban expansion and the increasing frequency and intensity of extreme
precipitation events bring new challenges to stormwater collection systems. One
underrecognized issue is the occurrence of transient flow conditions that lead to adverse
multiphase flow interactions (AMFI): essentially, the formation, collapse, and uncontrolled
release of air pockets within stormwater system flows. While the fundamental physics of
AMFI have been evaluated in laboratory experiments and idealized modeling studies, much
less is known about their development in real or simulated stormwater networks, and about
the roles played by rainfall and network properties. A necessary precursor to AMFI is the
development of pressurized flow conditions within a network. The goal of this study is to
understand how spatiotemporal rainfall variability affects the occurrence of pressurized
conditions in a stormwater drainage network in the Richmond district of San Francisco,
California. High-resolution bias-corrected radar rainfall fields for 24 recent storms were used
as the independent variable of EPA-SWMM simulations. Model analyses indicate that the
incidence of pressurized flow increases with storm intensity, and is more sensitive to rainfall
temporal variability than spatial variability. This research provides a reference for analyzing
AMFI precursors in other networks and may have important implications for the
improvement of stromwater infrastructures.
Keywords: Rainfall spatiotemporal variability; Pressurized flow; EPA-SWMM; Stormwater
networks; Adverse multiphase flow interactions

Amir Ahmad Dar, Afsar Jahan Shaik,


Spatiotemporal clustering and multivariate forecasting of air quality index across Indian
cities using machine learning and deep learning models,
Franklin Open,
Volume 13,
2025,
100435,
ISSN 2773-1863,
[Link]
([Link]
Abstract: Air pollution is still a significant environmental and public health issue in India, with
heterogeneous urban systems showing unique Air Quality Index (AQI) patterns. This paper
suggests an integrated approach of combining spatiotemporal clustering and multivariate
forecasting to improve AQI prediction in ten large Indian cities. Applying K-Means clustering
on AQI, temperature, and wind speed metrics, cities were classified into three replicable
regional groups: northern industrial, coastal/eastern, and southern plateau regions. In every
cluster, machine learning (Random Forest, XGBoost) and deep learning (LSTM, GRU, and a
hybrid LSTM+GRU ensemble) models were learned in a rolling-origin cross-validation
framework to provide temporal [Link] ensemble model performed the best
consistently, with RMSE varying from 0.090 to 0.101 and R² varying from 0.71 to 0.76,
overperforming both standard baselines and naïve seasonal predictors. Seasonal analysis
indicated winter months in northern cities to be the most challenging to predict, and
monsoon periods were found to have the highest predictability. An ablation study also
measured the contribution of hybridization, clustering, and meteorological inputs
[Link] approach is scalable, interpretable, and deployable, enabling data-driven
air quality management as well as actionable insights for environmental policy and urban
planning.
Keywords: Air quality index; Spatiotemporal clustering; Multivariate forecasting; LSTM; GRU;
Random forest; Ensemble learning; Indian cities

Samaneh Bagheri, Sadra Karimzadeh, Bakhtiar Feizizadeh, Saeed Samadianfard,


An integrated data-driven approach using dual polarized SAR data for spatiotemporal
analysis of water surface changes,
Advances in Space Research,
Volume 76, Issue 11,
2025,
Pages 6623-6646,
ISSN 0273-1177,
[Link]
([Link]
Abstract: Effective monitoring of water surface changes in reservoirs is crucial for water
resource management, especially in arid and semi-arid regions. Traditional ground-based
methods, although accurate, are labor-intensive and impractical for large-scale monitoring.
Synthetic Aperture Radar (SAR) remote sensing offers a promising alternative by enabling
continuous observation of water bodies. This study utilizes 360 ascending Sentinel-1 images
in VV and VH polarizations to analyze water surface area changes at the Boukan
embankment dam in western Iran. The Support Vector Machine (SVM) algorithm was
applied for classification, while the Improved Atom Search Optimization-Extreme Learning
Machine (IASO-ELM) model was used to integrate classification results with climatic data.
The IASO-ELM model optimizes the initialization of weights and biases within the ELM
framework, enhancing predictive accuracy. The model’s performance was evaluated using
several metrics, with Scenario 8 demonstrating the best results, including the lowest RMSE
(1.8524), MAE (1.4698), and high R2 (0.9608), NSE (0.9608), and WI (0.9901), indicating
strong agreement with observed data. Scenario 3 performed the worst, with the highest
RMSE (3.0112) and MAE (2.2785), and a lower R2 (0.8983), showing weak predictive
accuracy. A comparison between water surface areas derived from Sentinel-1-based SVM
classification and those obtained using the NDWI index on Sentinel-2 imagery was also
conducted. Results showed that Sentinel-1-based SVM classification provided more accurate
results, with errors below 15 %, compared to NDWI’s error of 35 % in June 2020. This
highlights the superiority of classification-based methods over simple indices in capturing
complex variations in water surface dynamics. The IASO-ELM model’s ability to accurately
predict water surface changes offers a robust tool for water management, supporting
proactive strategies for flood prevention, drought mitigation, and sustainable water resource
planning.
Keywords: Spatiotemporal analysis; Synthetic aperture radar; IASO-ELM; Water surface
areas; Predictive modeling

Yixuan Liu, Alim Samat, Peijun Du, Jin Chen, Jilili Abuduwaili, Kaiyue Luo, Enzhao Zhu, Dana
Shokparova,
High-resolution spatiotemporal analysis and driver attribution of floods in Kazakhstan using
SHAP and remote sensing integration,
Climate Risk Management,
Volume 51,
2026,
100783,
ISSN 2212-0963,
[Link]
([Link]
Abstract: The escalating impacts of global climate change and extreme weather have
intensified flood risks worldwide, including in arid and semi-arid regions traditionally
considered low-risk. This study examines the spatiotemporal dynamics of flood events across
Kazakhstan from 2000 to 2024 by integrating remote sensing (RS) with machine learning
(ML). Using Google Earth Engine (GEE), we address data gaps and cloud interference through
spatiotemporal fusion (STARFM), denoising, smoothing, and sample transferring techniques.
In addition, this study incorporates the Time-Disaggregated Water Frequency (TWF) method,
which enables the identification of water bodies with temporal variability, eliminates
permanent water bodies, and distinguishes flood from non-flood conditions in seasonal
water bodies, thereby enhancing the accuracy of flood reconstruction and enabling precise
delineation of flood inundation areas. Landsat and MODIS imagery are combined to produce
high-resolution flood distribution maps, while spectral similarity indicators guide the transfer
of samples from the Global Flood Database. A range of spectral, texture, environmental, and
socioeconomic features is extracted, with flood classification performed using random forest
(RF) and attribution analysis conducted via XGBoost and SHAP. Results highlight a high flood
risk in northern, southwestern, and western Kazakhstan, primarily driven by changes in
precipitation (PRE), temperature (TEM), soil moisture (SM), and land use. Floods occur most
frequently in spring — especially in March and April — due to snowmelt and extreme
precipitation. The ML models achieve over 80 % classification accuracy, demonstrating their
reliability. This work improves flood monitoring and provides essential insights for climate
adaptation and targeted flood risk management in Kazakhstan.
Keywords: Flood; Spatiotemporal; SHapley additive exPlanations (SHAP); Machine learning
(ML); Kazakhstan; Remote sensing

Haoming Feng, Huaqing Li, Wenwen Zhu, Denghao Li, Yukun Huang,
Micro-motion enhanced multi-person activity recognition with millimeter-wave radar,
Measurement,
Volume 258, Part B,
2026,
119090,
ISSN 0263-2241,
[Link]
([Link]
Abstract: As a non-contact sensing device, millimeter-wave radar exhibits unique strengths
in human activity recognition (HAR). Existing methods rely on micro-Doppler signatures for
activity classification, but they often encounter feature aliasing in multi-person activity
recognition (MPAR) scenarios. Although point cloud-based approaches can distinguish
individual targets, they primarily extract static morphological features, neglecting the micro-
motion information of human joints, which is crucial for accurate activity recognition. To
address these limitations, we proposes an innovative MPAR framework that integrates
spatial point clouds and micro-motion features. First, an improved point cloud data
association algorithm is applied to achieve multi-target point cloud feature separation,
followed by a dynamic projection mechanism to construct time–Doppler feature maps.
Then, a torso micro-motion enhancement algorithm is designed to enhance the details of
human body movements. Finally, a CNN-LSTM hybrid network architecture with a temporal-
attention is constructed for action classification. Experimental results show that the
proposed micro-motion enhancement algorithm improves recognition accuracy by 27.1%
and 2.3%, compared to two traditional time–frequency analysis methods. Furthermore,
MPAR task in occlusion scenarios achieves recognition accuracy of 93.5%. In summary,
proposed framework not only retains the inherent advantages of millimeter-wave radar but
also significantly enhances multi-person activity recognition in complex scenarios.
Keywords: Human activity recognition (HAR); Multi-person activity recognition (MPAR);
Feature separability; Temporal attention; Occlusion scenarios

Yun Zhou, Yinglin Zhu, Haohao Ren, Jiahao Kang, Xuegang Wang,
Refined multi-modal feature learning framework for marine target detection using radar
sensor,
Digital Signal Processing,
Volume 170,
2026,
105816,
ISSN 1051-2004,
[Link]
([Link]
Abstract: The fusion of time and time-frequency characteristics in radar echoes offers a
novel approach for marine target detection. However, echo amplitude alone cannot fully
characterize the time-domain information, as it fails to capture the temporal correlation
between sampling points. Therefore, this article introduces the Gramian Angular Summation
Field (GASF) for processing raw radar echoes to obtain the temporal information. Concretely,
to enable the detector to utilize features from diverse signal representations of the same
target echoes, we first preprocess the echoes of radar with two signal processing methods,
GASF and STFT, which aim to reflect the temporal dependence and dynamic changes of
frequency components, respectively. Subsequently, we develop a dual-stream feature
extraction network, i.e., time-frequency self-attention learning and GASF-based spatial-
temporal correlation learning, to deeply extract the discriminative features from two
modalities of the same radar echo. Then, to overcome the heterogeneity of multimodal
features during feature fusion, we propose a cross-modal feature fusion strategy to map
multi-modal features to a unified space. Finally, the fused features are fed into the detection
module. Numerous evaluation experiments on the publicly available measured IPIX dataset
demonstrate that the proposed detector is competitive with some state-of-the-art detectors
for marine target detection.
Keywords: Radar target detection; Signal processing; Gramian angular summation field;
Deep learning; Short-time Fourier transform

Xiangyang Luo, Ying Lu, Bibo Zhang, Yadan Yang, Jiaxin Li, Wanying Fu, Xinke Bu, Cong Li,
Identification and spatiotemporal analysis of braided rivers in the Yarlung Tsangpo basin
using an enhanced U-Net approach,
Journal of Hydrology,
Volume 666,
2026,
134796,
ISSN 0022-1694,
[Link]
([Link]
Abstract: Braided river systems, characterized by their unique ecological functions, play a
vital role in maintaining biodiversity, due to the unique braided morphology of braided river
systems, the primary prerequisite for conducting related research is the accurate
identification of braided water bodies within their catchments. However, most existing
studies rely heavily on field surveys or aerial imagery to extract information on braided river
networks. Such methods are costly, operationally complex, and insufficient for applications
requiring extensive spatial and temporal coverage. To address this challenge, this study
proposes an improved U-Net model, termed MSU-Net, which incorporates a multi-scale dual
attention gate module. By integrating spatial and channel attention mechanisms, the model
enhances the extraction of water features in complex environments. This study constructs a
monthly remote sensing dataset of the Yarlung Tsangpo River Basin from 2018 to 2023 using
Sentinel-1(SAR) and Sentinel-2 (optical) imagery. A model was trained based on manually
corrected labels and data augmentation strategies, and the spatial and temporal variations
of surface water area in the basin were analyzed based on the model outputs. The research
presents a more efficient method for identifying braided river systems and analyzes the
spatiotemporal dynamics of the braided channels in the Shannan section of the Yarlung
Tsangpo River Basin and their association with climatic factors, providing a scientific basis for
watershed water resource management.
Keywords: Braided river systems; Remote sensing; Yarlung Tsangpo–Brahmaputra River;
Water surface area; River channel variability

Xingyu Wang, Zhen Yang, Jichuan Huang, Bao Zhang, Yuhe Zhang, Deyun Zhou,
Collaborative strategy for hybrid actions of radar modes and maneuver decisions under
observation errors,
Engineering Applications of Artificial Intelligence,
Volume 160, Part A,
2025,
111774,
ISSN 0952-1976,
[Link]
([Link]
Abstract: The rapid advancement of airborne avionics has driven modern air combat to rely
heavily on information-centric operations, with radar serving as a primary tool for
information acquisition and playing a critical role in air combat. However, existing research
on air combat strategies often overlooks the impact of different radar operating modes on
maneuvering strategies, as well as the challenges posed by learning strategies under
observational disturbances. To address these gaps, this study investigates the problem of
hybrid actions decision-making for radar modes and maneuver decisions in the presence of
observational errors. Specifically, the characteristics of various radar operating modes are
analyzed and modeled, followed by an exploration of the convergence process of
reinforcement learning strategies under observational disturbances. To mitigate the
instability and volatility in strategy learning caused by observation errors, Entropy-
Decoupling-Noisy-net Proximal Policy Optimization-Advanced (EDN-PPOA) algorithm is
proposed, which significantly enhances the robustness and exploratory capability of the
model. Simulation results demonstrate that the proposed algorithm effectively achieves
coordinated tactical integration of radar modes and maneuvers in complex hybrid action
spaces, producing flexible tactical strategies that outperform expert-designed heuristics.
Furthermore, compared to the existing algorithms, the proposed method exhibits superior
stability and robustness in noisy observational environments, providing a reliable technical
foundation for intelligent decision-making in complex adversarial scenarios.
Keywords: Radar mode; Deep reinforcement learning; Hybrid actions; Maneuvering
decision-making

Pengfei Ge, Mi Chen, Roberto Tomás, Hui Liu, Kailun Fan, Xi Cheng, Xingyuan Fu, Shuang
Wang,
Insights into land deformation processes in the Yellow River Delta (China) from synthetic
aperture radar interferometry time series and machine learning,
Advances in Space Research,
2025,
,
ISSN 0273-1177,
[Link]
([Link]
Abstract: The Yellow River Delta is globally recognized as a highly dynamic region due to its
continuous transformations at the land-sea interface, and it abounds in valuable natural
resources, including oil, brine-rich groundwater and natural gas. The region experiences
impacts from tectonic activity, the natural compression and consolidation of loose
sediments, and, in particular, human economic activities. These factors lead to various types
of land subsidence, posing potential risks to local communities and economic operations.
Hence, effective monitoring and acquisition of the spatiotemporal patterns of land
subsidence in the Yellow River Delta play a crucial role in minimizing geological challenges
and financial losses. In this study, surface deformation data for the Yellow River Delta were
derived by processing 70 scenes of Sentinel-1 A/B data (32 ascending, 38 descending
acquisitions) using Interferometric Synthetic Aperture Radar (InSAR) time series technique,
with the Persistent Scatterer InSAR (PS-InSAR) and Small Baseline Subset InSAR (SBAS-InSAR)
methods applied over the period from January 2020 to December 2021. Moreover,
additional datasets, such as groundwater levels, precipitation, and areas of oil field and brine
extraction, were integrated to examine the factors affecting land subsidence and analyzed
using random forest analysis and post-interpretation methods. The findings indicate land
subsidence in the Yellow River Delta region displays an uneven distribution pattern, with
areas of severe subsidence primarily concentrated in Hekou District, Kenli District, Dongying
District and Guangrao County, characterized by the mean annual subsidence rate greater
than −120 mm/year. The spatial distribution of groundwater funnels, oil fields and brine
mining areas aligns to some extent with that of the severe land subsidence zones. The
random forest model outcomes reveal that the main contributors to land subsidence in the
Yellow River Delta are brine mining and soft soil thickness. Additionally, there exists regional
variability in the influencing factors among the various typical subsidence bowls. The post-
interpretation analysis further highlights shifts in the correlations among the various impact
factors and land subsidence.
Keywords: Land subsidence; PS-InSAR; SBAS-InSAR; Random forest; Yellow River Delta

Ji Ge, Hong Zhang, Lijun Zuo, Lu Xu, Jingling Jiang, Mingyang Song, Yinhaibin Ding, Yazhe Xie,
Fan Wu, Chao Wang, Wenjiang Huang,
Large-scale rice mapping under spatiotemporal heterogeneity using multi-temporal SAR
images and explainable deep learning,
ISPRS Journal of Photogrammetry and Remote Sensing,
Volume 220,
2025,
Pages 395-412,
ISSN 0924-2716,
[Link]
([Link]
Abstract: Timely and accurate mapping of rice cultivation distribution is crucial for ensuring
global food security and achieving SDG2. From a global perspective, rice areas display high
heterogeneity in spatial pattern and SAR time-series characteristics, posing substantial
challenges to deep learning (DL) models’ performance, efficiency, and transferability.
Moreover, due to their “black box” nature, DL often lack interpretability and credibility. To
address these challenges, this paper constructs the first SAR rice dataset with
spatiotemporal heterogeneity and proposes an explainable, lightweight model for rice area
extraction, the eXplainable Mamba UNet (XM-UNet). The dataset is based on the 2023
multi-temporal Sentinel-1 data, covering diverse rice samples from the United States, Kenya,
and Vietnam. A Temporal Feature Importance Explainer (TFI-Explainer) based on the
Selective State Space Model is designed to enhance adaptability to the temporal
heterogeneity of rice and the model’s interpretability. This explainer, coupled with the DL
model, provides interpretations of the importance of SAR temporal features and facilitates
crucial time phase screening. To overcome the spatial heterogeneity of rice, an Attention
Sandglass Layer (ASL) combining CNN and self-attention mechanisms is designed to enhance
the local spatial feature extraction capabilities. Additionally, the Parallel Visual State Space
Layer (PVSSL) utilizes 2D-Selective-Scan (SS2D) cross-scanning to capture the global spatial
features of rice multi-directionally, significantly reducing computational complexity through
parallelization. Experimental results demonstrate that the XM-UNet adapts well to the
spatiotemporal heterogeneity of rice globally, with OA and F1-score of 94.26 % and 90.73 %,
respectively. The model is extremely lightweight, with only 0.190 M parameters and 0.279
GFLOPs. Mamba’s selective scanning facilitates feature screening, and its integration with
CNN effectively balances rice’s local and global spatial characteristics. The interpretability
experiments prove that the explanations of the importance of the temporal features
provided by the model are crucial for guiding rice distribution mapping and filling a gap in
the related field. The code is available in [Link]
Keywords: Synthetic aperture radar; Rice mapping; Explainable deep learning; Feature
importance

Hu Liu, Zhenghua Zhang, Jing Yang, Jörg Benndorf, Xiaofei Wang, Jiaqi Dong, Zitao Lin,
Guoliang Chen,
GhostPointNet: A deep learning-based method for ghost point noise detection in four-
dimensional (4D) millimeter-wave radar point clouds of underground mine,
Engineering Applications of Artificial Intelligence,
Volume 161, Part C,
2025,
112380,
ISSN 0952-1976,
[Link]
([Link]
Abstract: The high dust concentration, multi-metal supports, and narrow winding tunnels in
underground mines collectively lead to frequent ghost point noise in four-dimensional (4D)
millimeter-wave radar point clouds, posing serious challenges for mining perception and
localization. To address this, we propose a deep learning algorithm, named GhostPointNet,
for 4D millimeter-wave radar ghost point detection in underground mining environments.
From an artificial intelligence perspective, this model thoroughly considers the multi-modal
features of 4D millimeter-wave radar and the environmental complexity of underground
mines. It incorporates multi-parameterized spatial information inputs in both Cartesian and
Spherical coordinates, coupled with “Double T-Net” adaptive alignment correction, while
integrating non-spatial information such as radar power and Doppler data to achieve multi-
modal representation and end-to-end discrimination between ghost points and real points.
Experimental validation shows that GhostPointNet achieves excellent performance in
underground mining scenarios with 92.45 % accuracy and 95.84 % F1-score, outperforming
traditional filtering, clustering, and machine learning algorithms. From an engineering
application perspective, GhostPointNet is specifically designed for ghost noise detection in
underground mines. Even in complex scenarios such as mine tunnel intersections and turns,
it preserves critical structural points. Its end-to-end neural network simplifies post-
processing procedures, enhances operational efficiency, and provides stable and reliable
perceptual support for subsequent tasks such as autonomous mine locomotive navigation
and three-dimensional (3D) structure reconstruction. Experimental results demonstrate that
this method surpasses baseline approaches in ghost point detection, real point preservation,
and generalization capability, providing significant support for improving underground
mining safety and efficiency.
Keywords: Deep learning; Four-dimensional (4D) millimeter-wave radar; Ghost noise;
Underground mining; Point cloud segmentation

Yanwen Bai, Jibin Zheng, Hanxing Shao, Hongwei Liu,


A dual-driven hybrid tracking architecture for radar targets based on innovation,
Information Fusion,
Volume 129,
2026,
104056,
ISSN 1566-2535,
[Link]
([Link]
Abstract: Targets such as hypersonic missiles and stealth aircraft are characterized by
complex motion patterns, strong maneuverability, and anomalous radar measurement
statistics. Although model-driven radar target tracking methods offer physical
interpretability, they suffer from their dependence on explicit prior assumptions. Data-
driven methods can theoretically approximate arbitrarily complex motions through
nonlinear mappings, but suffer from poor interpretability, vulnerability to noise during
feature extraction, and loss of low-frequency maneuvering features due to sample
imbalance. Therefore, this paper proposes a Dual-Driven Hybrid Tracking Architecture Based
on Innovation (DDHTA), which fuses the advantages of both model-driven and data-driven
approaches. First, a model-driven approach is adopted for basic state estimation, and a Dual
Condition Judgment Adjustment (DCJA) method is proposed to adaptively adjust the
measurement error variance, thereby providing a high-quality baseline estimate for the
data-driven layer and reducing the interference of anomalous noise on feature extraction.
Further, in the data-driven layer, a Dual-Scale Temporal Network (DSTNet) is designed. By
learning the mapping from the innovation to the estimation errors, it combines the
strengths of causal dilated convolution and multi-head self-attention to provide dynamic
compensation, which corrects the estimation errors of the model-driven method. Numerical
simulation results demonstrate that the proposed method enhances the algorithm’s ability
to handle target maneuvers in complex environments, achieving higher tracking accuracy
and robustness.
Keywords: Radar tracking; Deep learning; Self-attention mechanism; Causal convolution

Xinmeng Zhou, Jingyi Wang, Junyi Wang, Qingfeng Guan,


Predicting air quality using a multi-scale spatiotemporal graph attention network,
Information Sciences,
Volume 680,
2024,
121072,
ISSN 0020-0255,
[Link]
([Link]
Abstract: As urbanization accelerates, air quality has become a pressing concern. Accurate
air quality prediction is essential for informed governmental decision-making and for
protecting public health. Variations in air quality are influenced by complex multi-scale
spatiotemporal processes. Existing research primarily relies on capturing single
spatiotemporal features of air quality to predict changes. Meanwhile, when constructing
spatiotemporal dynamic graphs, the inherent characteristics of the input data and the
comprehensive effects of both global and local influences are not fully considered. To
address these problems, we propose a graph-attention-based approach, named Multi-scale
Spatiotemporal Graph Attention Network (MSTGAN). MSTGAN addresses the intricate
spatiotemporal patterns of air quality across various scales through three key components:
(1) a multistation transformer to model the temporal patterns of air quality at individual
monitoring stations; (2) a bilinear spatiotemporal attention mechanism to capture the
spatiotemporal dynamic global dependencies among all stations in a region; and (3) a set of
spatiotemporal dependence graph-coupled Chebyshev graph convolution gate recurrent
units to extract and aggregate the local spatiotemporal features of interrelated stations.
Experiments conducted on three real-world datasets demonstrated that MSTGAN achieved
significant improvements of 4.2%, 3.9%, and 7.8% in the mean absolute error, root mean
square error, and R2 evaluation metrics, respectively, compared to seven state-of-the-art
time-series forecasting methods. This code is publicly available at
[Link]
Keywords: Air quality prediction; Multi-scale; Spatiotemporal graph attention; Deep learning

Junkai Liu, Xinwei Qian, Lu Peng, Dan Lou, Yiwen Li,


TEDR: A spatiotemporal attention radar extrapolation network constrained by optical flow
and distribution correction,
Atmospheric Research,
Volume 311,
2024,
107702,
ISSN 0169-8095,
[Link]
([Link]
Abstract: In recent years, deep learning has been widely applied to meteorological radar
extrapolation due to the shortcomings of traditional optical flow methods in predicting the
genesis and dissipation of radar echoes. However, it still faces challenges in addressing issues
of clarity and overall intensity attenuation caused by uncertainty. This study implemented a
dual-path spatiotemporal attention network that integrates optical flow techniques by
employing intra-frame static attention and inter-frame dynamic attention, which could
simulate motion fields and the overall intensity distribution of radar echoes separately. Our
approach effectively resolve the issues of systematic intensity attenuation and clarity
degradation introduced by deep learning methods. Through the comparisons of key metrics
such as MSE, SSIM, CSI20, CSI30, and CSI40, the results demonstrated significant
improvements over traditional approaches, particularly in CSI30 and CSI40, where the
metrics improved by more than 35 %.
Keywords: Radar extrapolation; Convective weather forecasting; Machine learning; Optical
flow

Mingqi Li, Pengxin Wang, Kevin Tansey, Fengwei Guo, Ji Zhou,


Improved leaf area index reconstruction in heavily cloudy areas: A novel deep learning
approach for SAR-Optical fusion integrating spatiotemporal features,
International Journal of Applied Earth Observation and Geoinformation,
Volume 142,
2025,
104745,
ISSN 1569-8432,
[Link]
([Link]
Abstract: The Leaf Area Index (LAI) is an essential parameter for assessing vegetation growth.
LAI derived from optical data can suffer from gaps caused by cloud cover. Synthetic Aperture
Radar (SAR) presents a solution with its all-weather observation capability. To address these
issues, this study proposes a new deep learning approach for reconstructing time series LAI
using SAR and optical data in two steps. Firstly, the two-dimensional Convolutional Neural
Network-Transformer (2D CNN-Transformer) is applied to bridge SAR and optical data.
Secondly, the 2D CNN-Transformer predicted LAI and the Sentinel-2 LAI are input into the
Enhanced Deep Convolutional Model for Spatiotemporal Image Fusion (EDCSTFN) model to
further improve the accuracy. The novelty lies in a two-step framework combining a 2D CNN-
Transformer for spatiotemporal feature extraction and a deep learning fusion algorithm
refining accurate LAI reconstruction. Results showed that the 2D CNN-Transformer achieved
a higher accuracy (R2 = 0.64, RMSE = 0.38 m2/m2) in establishing a relationship between
SAR and optical data, compared to 1D CNN, 2D CNN-LSTM, and 1D CNN-Transformer. In the
second step, the EDCSTFN reconstructed LAI achieved the highest accuracy of an R2 of 0.81
and an RMSE of 0.22 m2/m2, with an average R2 of 0.61 and RMSE of 0.37 m2/m2 across
croplands and forests in millions of pixels, further improving the accuracy based on the first
step. The approach effectively fills gaps in spatial details and achieves a more continuous
spatial distribution. The proposed approach demonstrates good generalizability in millions of
pixels under frequent cloud cover and complex surface conditions and provides a new
strategy for the fusion of optical and SAR data.
Keywords: SAR data; Optical data; Spatiotemporal features; LAI; Deep learning;
Spatiotemporal fusion

Rui Wan, Weigang Meng, Tianyun Zhao, Wei Lu,


Overcoming radar sparsity and cross-view misalignment: A sparse-to-sparse fusion paradigm
for robust 3D object detection,
Digital Signal Processing,
Volume 171,
2026,
105831,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Millimeter-wave radar-camera fusion provides a cost-effective alternative to LiDAR
for 3D perception in autonomous driving. However, its potential is constrained by two
limitations: (1) Previous bird’s-eye view fusion methods struggle to accurately align and fuse
cross-modal features, while the fusion strategy incurs substantial computational redundancy
from background processing. (2) The extreme sparsity of radar points (typically < 5 % LiDAR
density) hinders robust geometric measurement. To address these challenges, we propose a
radar-camera fusion 3D detection framework that redefines cross-modal interaction by
transitioning from dense fusion to sparse-to-sparse paradigm. This transformation is
initiated by generating spatially-aware 3D object queries from images and radar sweeps-
leveraging image-derived seed points with radar depth to anchor queries to objects via a
perspective-guided object query generator. Moreover, we introduce adaptive radar pillar
diffusion within foreground regions to mitigate radar sparsity, allowing object queries to
capture geometric information from diffused pillars. Additionally, to maximize image
semantic clues, we further refine object boxes through image keypoint feature aggregation
using a keypoint-aware object refinement module. Our framework not only circumvents
traditional fusion bottlenecks but also achieves real-time inference at 23.4 FPS. Evaluated on
nuScenes dataset, it demonstrates competitive detection performance (65.0 % NDS, 57.8 %
mAP) and tracking precision (58.3 % AMOTA and 0.687m AMOTP). By overcoming sparsity
constraints and improving cross-modal fusion, this work establishes a new paradigm for
robust perception systems.
Keywords: Autonomous driving; 3D object detection; Deep learning; Radar-camera fusion;
Cross-attention
Yunfeng Fang, Zheng Tong, Tianqing Hei, Siqi Wang, Tao Ma,
Deep learning applications in ground-penetrating radar inversion: A review,
Measurement,
Volume 258, Part D,
2026,
119399,
ISSN 0263-2241,
[Link]
([Link]
Abstract: The complex nonlinear relationship between the subsurface medium and ground-
penetrating radar signals results in the pervasive ill-posedness and non-uniqueness of
conventional inversion methods. Deep learning, with its powerful feature extraction
capabilities and advantages in modeling complex nonlinear relationships, has unique
strengths in handling complex signals and nonlinear problems, making it especially suitable
for GPR inversion tasks. This paper reviews the latest applications of deep learning in GPR
inversion, summarizing the application strategies of deep learning from two perspectives:
data-driven and data-physics hybrid-driven. Commonly used model architectures and their
performance in signal feature extraction, multi-scale information fusion, and data
preprocessing are discussed, along with the application of various loss functions in inversion
tasks. Finally, current challenges, such as limited model generalization, model dependence
on the dataset and computational efficiency constraints, are discussed, and potential future
research directions are proposed to further advance deep learning in GPR inversion.
Keywords: Ground-penetrating radar; Deep learning; Inversion

Fengrui Chen, Xi Li, Yiguo Wang, Shaoqi Pan,


A spatiotemporal-guided hybrid learning model for mapping surface air relative humidity,
International Journal of Applied Earth Observation and Geoinformation,
Volume 144,
2025,
104895,
ISSN 1569-8432,
[Link]
([Link]
Abstract: Accurately mapping the spatial distribution of surface air relative humidity (RH) is
critically important for diverse fields including climate change, public health, and
environmental conservation. Although machine learning has exhibited considerable
advantages in RH mapping, factors such as its inherent bottom-up learning strategy, the
unique spatiotemporal characteristics of geographical phenomena, and the scarcity of
ground-based observations limit model performance. To address these limitations, this study
proposes a novel spatiotemporal-guided hybrid learning (STGHL) model for precise RH
mapping. The model establishes a hybrid learning paradigm that integrates principles of
spatiotemporal autocorrelation and heterogeneity into deep learning architectures. The
paradigm organically combines top-down guided learning with bottom-up spontaneous
learning through the design of three specialized guided learning modules, which collectively
enable neural networks to deeply represent the complex spatiotemporal patterns of RH. We
conducted a comprehensive evaluation of the STGHL model by generating daily RH maps
across the Chinese mainland during the period 2014–2018. The rigorous 10-fold spatial
cross-validation results show that the STGHL model achieves exceptional performance, with
an R2 of 0.92, a mean absolute error (MAE) of 3.93%, and a root mean square error (RMSE)
of 5.24%. Compared with state-of-the-art machine learning-based mapping models, the
proposed model demonstrates significant improvements in accuracy, generalization
capability, and explainability, reducing MAE and RMSE by at least 13% and 14%, respectively.
This study contributes to the field by establishing a novel spatiotemporal-guided hybrid
learning paradigm that offering new insights and approaches for geospatial phenomena
mapping, while promoting the profound integration of machine learning and geoscience
disciplines.
Keywords: Surface air relative humidity; Deep learning; Prior knowledge; Spatiotemporal
correlation; Spatiotemporal heterogeneity

Fengzhen Sun, Luxiang Ren, Weidong Jin,


FastNet: A feature aggregation spatiotemporal network for predictive learning,
Engineering Applications of Artificial Intelligence,
Volume 130,
2024,
107785,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Spatiotemporal prediction is a challenging topic because of uncertainty. Existing
works attempt to design complicated systems to learn short-term and long-term dynamics.
However, these models suffer from heavy computational burden on spatiotemporal data,
especially for high-resolution frames. To reduce resource dependence, we propose FastNet,
a novel and light encoder–decoder for predictive learning. We stack four ConvLSTM
(Convolutional Long Short-Term Memory) based layers to construct a hierarchical
framework. Based on this architecture, the feature aggregation module first aligns the
temporal context, then decouples different frequency information, next gathers multi-level
features and last synthesizes new feature maps alternately. We aggregate various and
hierarchical features into predictions bringing two benefits: rich multi-level features and low
resource usage. As for the unit blocks, depth-wise separable convolution is used to improve
model efficiency and compress model size. Besides, we adopt perceptual loss as the cost
function between ground truths and predictions, which helps our model to get higher
similarity with true frames. In experiments, we evaluate the performances of FastNet on the
MovingMNIST (Mixed National Institute of Standards and Technology) and Radar Echo
datasets to verify its effectiveness. The quantitative metrics on the Radar Echo dataset show
that FastNet reaches a slight increase in accuracy but up to 84% decline in computation
compared with PredRNN-V2. Therefore, our FastNet achieves competitive results with lower
resource usage and fewer parameters than the state-of-the-art model.
Keywords: Spatiotemporal prediction; Predictive learning; Aggregate features; Depthwise
separable convolution; Perceptual loss

Decai Jin, Xiufang Zhu, Ying Qu, Jianbo Qi, Hanyi Wu, Yaozhong Pan,
A robust and efficient deep optimization network for spatiotemporal data fusion,
Information Fusion,
Volume 127, Part C,
2026,
103939,
ISSN 1566-2535,
[Link]
([Link]
Abstract: Satellites strive to strike a delicate balance between temporal and spatial
resolution, thereby rendering the achievement of high resolution in both aspects
challenging. Spatiotemporal fusion algorithms have emerged as a promising solution to
tackle this challenge. However, with changes in spatiotemporal conditions, existing
spatiotemporal fusion methods, particularly those based on deep learning, face challenges
such as decreased prediction accuracy and poor reconstruction accuracy in areas of abrupt
changes. This presents significant challenges for the fusion of multi-source remote sensing
data to generate cloud-free remote sensing images on a daily scale. In this context, the study
proposes a multiscale Attention-Guided deep optimization network for Spatiotemporal Data
Fusion (AGSDF) method. The algorithm is designed to generate daily fine images using
coarse image, based on historical reference fine images. Specifically, it firstly attempts to use
a physical attention mechanism to mitigate the effects of climate change in time-series
images. Implementing a continuous spatiotemporal fusion process across multiple scales
significantly enhances the model's robustness. The performance of AGSDF was evaluated
and compared to nine methods at six sites worldwide. The experimental results indicate that
AGSDF achieved a top score in the assessment. Consequently, AGSDF holds high potential to
produce accurate remote sensing products with high temporal and spatial resolution across
extensive regions.
Keywords: Spatiotemporal fusion; Data fusion; MODIS; Landsat; Super resolution

Jin Xie, Dapeng Wang, Zhimin Li,


Information perception method for highway traffic accident scenes using radar-visual fusion,
Journal of Intelligent Transportation Systems,
2025,
,
ISSN 1547-2450,
[Link]
([Link]
Abstract: To improve safety risk assessment and emergency response at highway accident
sites, this study proposes an information perception framework that integrates dynamic and
static data. It addresses challenges in identifying dynamic environmental features and fusing
multi-source data. The framework combines radar point cloud data, video streams, and
other environmental information to achieve accurate perception and real-time monitoring.
For dynamic data, a multi-modal cross-attention (MCA) mechanism aligns spatiotemporal
features of radar and video data. This extracts key parameters, such as vehicle flow, speed,
and occupancy, enhancing traffic flow characterization. For static data, a YOLOv8-based
multi-task detection model identifies road debris and obstacles, improving accuracy and
robustness. In the data fusion and validation stage, dynamic and static information is
integrated into a complete data stream. The model is tested with real highway data. Results
show high detection performance under different lighting conditions. The dynamic
perception module achieves 90.69% accuracy and 92.85% recall in both daytime and
nighttime. The static detection model reaches 96.4% accuracy and 98.8% recall, meeting
emergency response requirements. To enhance visualization, a PyQt5-based platform
displays real-time dynamic-static data. This provides traffic management with clear and
accurate decision support.
Keywords: information perception method; intelligent transportation; radar-video fusion;
traffic accidents

Chun Zou, Shenping Hu, Lijia Chen,


Environmental data driven dynamic Bayesian network: Spatiotemporal evolution of coastal
shipping risk performance in China,
Reliability Engineering & System Safety,
Volume 271,
2026,
112247,
ISSN 0951-8320,
[Link]
([Link]
Abstract: The rapid development of the global shipping industry has led to increasingly
prominent maritime traffic risks, which pose a serious threat to economic development, the
ecological environment, and public safety. In this context, this study develops an
environmental data-driven dynamic Bayesian network (DBN) model to simulate the
spatiotemporal evolution of maritime traffic risks using large-scale data from complex
shipping systems. Initially, based on the Systems Theoretic Accident Model and Processes
(STAMP) accident causation framework, risk influencing factors (RIFs) are identified through
the systematic analysis of maritime accident reports. Then, to address the dynamic nature of
maritime traffic risks, a novel transition probability matrix learning mechanism integrating
environmental data is proposed, enabling the DBN model to characterize risk performance
spatiotemporal evolution. Finally, a case study of the four major sea areas along the China
coast reveals that the evolution of risk performance exhibits significant spatiotemporal
heterogeneity across different regions, with the East Sea and the South China Sea showing
the most pronounced variations. Fluctuations in risk performance are highly correlated with
seasonal meteorological and hydrological changes, and the distribution of accident risks is
closely linked to extreme weather events. This study provides reliable quantitative tools to
support spatiotemporal risk management and cross-regional decision-making in maritime
traffic systems.
Keywords: DBN; STAMP; Maritime traffic risk; Spatiotemporal evolution; Maritime traffic
system

Xinrui Zhao, Zeng Liu, Qi Hu, Jianglong Sun, Xiaoyan Yang,


Significant wave height estimation and prediction from synthetic X-band radar data by
spatio-temporal deep neural networks,
Ocean Engineering,
Volume 339, Part 1,
2025,
122061,
ISSN 0029-8018,
[Link]
([Link]
Abstract: Estimation and prediction of real-time significant wave height (SWH) is a
fundamental requirement for the safety of offshore activities. This study utilizes a spatio-
temporal deep neural network model to effectively estimate and predict the SWH by
extracting key spatial and temporal features from synthetic X-band radar [Link] study
considers three deep neural networks: InceptionV3, ResNet and Vision Transformer (ViT) to
extract multi-scale spatial features from radar images for SWH estimation. Subsequently, a
gated recurrent unit (GRU) is employed on these spatial features to perform the time-series
SWH prediction. Irregular waves with sea state ranging from 4 to 6 were considered based
on the synthetic radar data. Results indicate that the ResNet model with deep residual
structure performs the best in both the estimation and prediction tasks, demonstrating
excellent generalization ability and adaptability to complex sea conditions. The ViT model
shows outstanding performance in scenarios without out-of-distribution data. While the
InceptionV3 model is inferior, it exhibits significant improvement for the SWH prediction
when the GRU is incorporated.
Keywords: Significant wave height(SWH); X-band marine radar; InceptionV3; ResNet; Vision
Transformer(ViT); Gated recurrent unit(GRU)

Li Wang, Baicheng Hu, Yuan Zhao, Kunlin Song, Jianmin Ma, Hong Gao, Tao Huang, Xiaoxuan
Mao,
A hybrid spatiotemporal model combining graph attention network and gated recurrent unit
for regional composite air pollution prediction and collaborative control,
Sustainable Cities and Society,
Volume 116,
2024,
105925,
ISSN 2210-6707,
[Link]
([Link]
Abstract: Machine learning (ML) models have been extensively applied in air quality
prediction. However, many of these models often failed to unveil complex mechanisms and
regional spatiotemporal variations of composite air pollution. This brings uncertainties in
using ML models for effective composite air pollution control. The present study developed a
novel hybrid spatiotemporal model framework combining Graph Attention Network (GAT)
and Gated Recurrent Unit (GRU), namely the GAT-GRU model, to foresee composite air
pollutions with a focus on PM2.5 and O3. By extracting attention matrices for PM2.5O3
composite pollution and applying the Louvain algorithm, the framework established
effective community network divisions for coordinated control of PM2.5O3 composite
pollution. The framework was applied and tested in China's “2 + 26″ cities, a city cluster with
most heavy PM2.5 and O3 pollution and precursor emission sources. The results
demonstrate that the framework successfully captured spatiotemporal evolution of
combined PM2.5 and O3 pollution. The attention matrix is autonomously generated during
course of the model learning process with the aim to interpret the complex interactions
among “2 + 26″ cities. The framework provides a new perspective for the interpretability of
artificial intelligence models and offers a methodological support and scientific evidence for
formulating regional pollution cooperative governance strategies.
Keywords: GAT-GRU; PM2.5, O3; Attention matrix; Community network
Liangang Qi, Hongzhuo Chen, Qiang Guo, Shuai Huang, Mykola Kaliuzhnyi,
GLS: A hybrid deep learning model for radar emitter signal sorting,
Digital Signal Processing,
Volume 161,
2025,
105117,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Radar emitter signal sorting is a pivotal aspect of radar reconnaissance signal
processing. The increasing density of the electromagnetic environment in modern radar
pulse streams, coupled with the growing complexity and variability of operational modes
and signal forms, results in extremely limited reference data. Consequently, most existing
sorting methods fall short of meeting the performance requirements of modern electronic
warfare. To enhance sorting performance under conditions of limited samples and labeled
data, this paper proposes a radar emitter signal sorting model based on ResGCN-BiLSTM-SE
(GLS). Firstly, we propose a novel adaptive weighted adjacency matrix construction method
that aggregates multi-scale information of local and global features. Based on this, for GLS
networks, the graph convolutional network (ResGCN) is combined with the bidirectional long
short-term memory (BiLSTM) network. The GCN is employed to extract attribute features
from interleaved radar pulse sequences, while the BiLSTM is utilized to deeply capture the
temporal dependence in interleaved pulse sequences after feature extraction. Finally, an
improved squeeze-and-excitation (SE) module is applied to perform weighted fusion of
critical channel information from both spatial and temporal features. Simulation results
demonstrate that the proposed method not only achieves higher accuracy under small
sample conditions compared to existing methods, but also exhibits strong robustness in
challenging scenarios involving measurement errors, missing pulses, and spurious pulses.
Keywords: Radar emitter signal sorting (RESS); Adaptive weighted adjacency matrix; GLS
model; Features fusion

Laura Pedretti, Pietro Teatini, Tommaso Letterio, Guadalupe Bru, Carolina Guardiola-Albert,
Roberto Tomás, María I. Navarro-Hernández, Alessandro Bondesan, Yuri Taddia, Claudia
Meisina,
Vertical land movements assessment integrating Interferometric Synthetic Aperture Radar,
in-situ data, and engineering-geological model: The case study of the reclaimed farmland of
the Po River Delta (Italy),
Engineering Geology,
Volume 363,
2026,
108544,
ISSN 0013-7952,
[Link]
([Link]
Abstract: Low-elevation reclaimed coastlands face significant challenges from land
subsidence and sea-level rise, making long-term monitoring of ground movements crucial to
ensure infrastructure safety and preserve the natural environment. This study aims to
reconstruct the long-term historical ground deformation of the reclaimed farmland in the Po
River Delta by: i) integrating nearly 30 years of multisource, multi-temporal, and multisensor
Interferometric Synthetic Aperture Radar (InSAR) satellite data (ERS-1/2, RADARSAT-1/2,
Sentinel-1); ii) combining multisource InSAR datasets generated using different algorithms
covering distinct or overlapping time periods (Sentinel-1 PSI, P-SBAS, and IPTA); and iii)
developing a 3D engineering-geological model focused on the under-consolidated fine-
grained deposits that are more prone to subsidence. By combining multiple monitoring
techniques, this multidisciplinary approach reveals that land subsidence is primarily driven
by autocompaction of under-consolidated finegrained sediments, locally accelerated by
building construction, as evidenced by InSAR data. The highest subsidence rates occur in the
youngest reclaimed areas with thicker under-consolidated fine-grained deposits. While
integrating multisensor InSAR datasets from diverse sources to reconstruct longterm ground
deformation presents challenges, it also yields valuable insights. In this work, we
demonstrate that heterogeneous datasets can still be valuable when interpreted carefully
and that the feasibility of combining legacy and modern InSAR data for long historical
deformation reconstruction is a practical challenge in real-world data integration. Moreover,
this comprehensive approach enables updating spatial and temporal records of land
movement and identifying conditioning factors for inclusion in land movement susceptibility
and risk maps supporting land planning.
Keywords: Long-term monitoring; Land subsidence spatiotemporal evolution; Multi-
sourcetemporal-sensor InSAR; Under-consolidated fine-grained sediments;
Engineeringgeological model; Land reclamation; Po River delta

Shaopeng He, Mingjun Wang, Nicola Forgione, Andrea Pucciarelli, W.X. Tian, S.Z. Qiu, G.H.
Su,
A multi-task Transformer-Mamba-Seq framework for real-time estimation of spatiotemporal
thermal stratification in passive residual heat exchanger,
International Communications in Heat and Mass Transfer,
Volume 169, Part D,
2025,
109868,
ISSN 0735-1933,
[Link]
([Link]
Abstract: Passive Residual Heat Removal Heat Exchanger (PRHR HX) is a critical component in
Generation-III nuclear power systems. Its spatiotemporal thermal stratification
characteristics directly influence residual heat removal capacity and serve as key inputs for
multiphysics coupling analyses. However, the complexity of input conditions challenges
traditional simulation and AI approaches, particularly under abnormal and accident
scenarios. To address this, we propose a multi-task Transformer-Mamba-Seq framework that
integrates multi-head attention with a selective scan mechanism. Compared to conventional
models, it demonstrates superior performance in both 5-fold cross-validation and

computational costs—cutting parameters by ∼90 % and training time by at least 59 %. Our


independent tests. Furthermore, a sequential training strategy significantly reduces

framework enables real-time prediction of the 4D temperature field and thermal


stratification characteristics in PRHR HX with high accuracy (RMSE/MAPE/R2:
1.81 K/0.41 %/0.887). It achieves a speedup of over 1500× compared to CFD simulations.
This work provides an efficient and accurate tool for real-time thermal analysis of PRHR HX,
supporting the thermal safety of Generation-III nuclear systems and could offering low-cost,
high-resolution inputs for thermal stress and flow-induced vibration analyses.
Keywords: Passive residual heat exchanger; Thermal stratification; Real-time estimation;
Transformer; Mamba

Chongxing Ji, Yuan Xu,


trajPredRNN+: A new approach for precipitation nowcasting with weather radar echo images
based on deep learning,
Heliyon,
Volume 10, Issue 18,
2024,
e36134,
ISSN 2405-8440,
[Link]
([Link]
Abstract: :Short-term rainfall prediction is a crucial and practical research area, with the
accuracy of rainfall prediction, particularly for heavy rainfall, significantly impacting people's
lives, property, and even their safety. Existing models, such as ConvLSTM, TrajGRU, and
PredRNN, exhibit limitations in capturing fine-grained appearances due to insufficient
memory units or addressing positional misalignment issues, thereby compromising the
accuracy of model predictions. In this study, we propose trajPredRNN+, an innovative
approach that integrates the trajectory segmentation model and the PredRNN deep learning
model to address both limitations in nowcasting precipitation using weather radar echo
images. By incorporating attention mechanisms, the model demonstrates an enhanced focus
on short-term and imminent heavy rainfall events. To ensure improved stability during
training, a residual network is introduced. Lastly, a more rational and effective training loss
function is proposed, encompassing weight mechanism, SSIM index, and GAN loss. To
validate the proposed model, we conducted a comparative experiment and an ablation
experiment using the radar echo map dataset obtained from the Shenzhen Meteorological
Bureau. The results of these experiments demonstrate that our model has achieved
significant improvements across multiple key performance indicators.
Keywords: Radar echo map; Deep learning; Precipitation nowcasting; PredRNN; GAN

Dunlu Peng, Meiling Chen, Yiqin Zhang, Zekun Tian,


Enhanced optic-flow extrapolation for Doppler radar nowcasting with Dynamic Weight
Attention,
Expert Systems with Applications,
Volume 267,
2025,
126168,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Doppler radar echo extrapolation is an important method for extreme weather
forecasting. However, traditional optical flow methods lack learnable components and are
not suitable for complex atmospheric changes. Consequently, researchers have turned to
deep neural networks for prediction. Yet, the predictions from this approach often suffer
from issues such as mean reversion and a lack of small- and medium-scale structure. This
paper proposes a novel approach that combines optical flow methods with deep neural
networks. By introducing an artificially defined momentum weight matrix based on prior
assumptions, predictions for any future time distance are generated from full-scale optical
flow. Additionally, we propose a full-scale advection extractor, leveraging the continuity of
distribution in mesoscale and small-scale atmospheric systems and focusing on the long- and
short-distance relationships within the contour surface distribution sequence, which
improves the prediction accuracy of fine-scale advection. The experimental results show
that, compared with other advanced methods, the proposed method demonstrates
advantages in predicting extreme radar echoes and maintaining the echo structure.
Specifically, it achieved an improvement of 24.1% and 21.3% on key indicators such as CSI
and HSS, respectively, and reached 0.948 on the SSIM. Building on this, the inference speed
of our method is comparable to other deep learning approaches, being 3.25 times faster
than flow-based methods.
Keywords: Nowcasting; Optical flow; Neural network; Doppler radar

Zhensheng Shi, Haiyong Zheng, Junyu Dong,


Spatiotemporal self-supervised predictive learning for atmospheric variable prediction via
multi-group multi-attention,
Knowledge-Based Systems,
Volume 300,
2024,
112090,
ISSN 0950-7051,
[Link]
([Link]
Abstract: Atmospheric variable prediction (AVP) is crucial for atmospheric and
environmental science, as well as practical applications related to production and daily life.
Traditional AVP commonly relies on numerical modeling methods. In recent years, deep
learning (DL) technologies have introduced new approaches for AVP. With the support of
massive high-quality data products, DL methods have achieved remarkable progress.
Simultaneously, studying and annotating complex phenomena and processes within vast
datasets require specialized researchers and come with substantial costs. The self-supervised
learning (SSL) mechanism in DL can facilitate representation learning from unlabeled data,
thereby significantly reducing the expenses associated with manual annotation. However,
there has been limited exploration of the SSL mechanism in atmospheric variable prediction
and analysis. Furthermore, the prevailing large-scale predicting applications based on AVP
mainly focus on massive data processing, multi-modal approaches, and foundation model
technologies, often neglecting the critical aspects of spatiotemporal representation learning.
In light of the above, based on the cutting-edge DL techniques but diverges from the current
popular contrastive or generative SSL frameworks, we devise a spatiotemporal (ST) self-
supervised predictive learning (SSPL) method (dubbed ST-SSPL) for AVP, which enables the
learning of predictive characteristics from unlabeled atmospheric data. Additionally, we
improve a multi-group multi-attention (MGMA) method, integrating it into our predictive
framework to enhance spatiotemporal representation learning. In experiments conducted
on the widely used ERA5 dataset, the proposed ST-SSPL method has achieved superior
performance and high efficiency using a simple convolutional neural network (CNN)
architecture. These advancements offer crucial methods and references for relevant large-
scale applications.
Keywords: Atmospheric variable prediction; Self-supervised learning; Spatiotemporal
predictive learning; Efficient deep learning model; Multi-attention

Weiliang Zhou, Dongmei Zhang, Man Xu,


SMTG-Net: A spatiotemporal deep learning model for large-scale urban land subsidence
prediction with heterogeneity awareness,
Environmental Modelling & Software,
Volume 195,
2026,
106739,
ISSN 1364-8152,
[Link]
([Link]
Abstract: Urban land subsidence is a widespread geological disaster, threatening production
and residents' lives. While synthetic aperture radar interferometry enables large-scale
monitoring, traditional models often overlook spatial-temporal heterogeneity, reducing
accuracy. To address this, we propose SMTG-Net (Spatial Mask Attention and Temporal
Granularity Network), designed to extract spatial structural features and temporal dynamic
patterns. It captures spatial local changes via a spatial factor mask matrix with a double
threshold mechanism and employs dynamic graph convolution with a gated recurrent unit to
extract dynamic spatial dependencies. The temporal decomposition mechanism decouples
data into trend and periodic components, while the multi-granularity collaborative encoder
learns global and local features. Using Sentinel-1A data from Jan 2020 to Aug 2023 in
Wuhan, China, experiments show SMTG-Net outperforms GCN, GAT, STGCN, Graph
WaveNet, and STTNs in RMSE, MAE, and MAPE. SMTG-Net effectively models spatial-
temporal heterogeneity, delivering accurate predictions and offering a novel approach to
urban subsidence monitoring.
Keywords: Urban land subsidence; Spatiotemporal prediction; Heterogeneity modeling;
InSAR

Md. Alamgir Hossain,


FED-GEM-CN: A federated dual-CNN architecture with contrastive cross-attention for
maritime radar intrusion detection,
Array,
Volume 27,
2025,
100456,
ISSN 2590-0056,
[Link]
([Link]
Abstract: The escalating complexity of maritime operations and the integration of advanced
radar systems have heightened the susceptibility of maritime infrastructures to sophisticated
cyber intrusions. Ensuring resilient and privacy-preserving intrusion detection in such
environments necessitates innovative solutions capable of learning from distributed,
heterogeneous data sources without compromising sensitive information. This study
introduces FED-GEM-CN, a novel federated learning framework designed explicitly for
maritime radar intrusion detection. The proposed architecture integrates dual parallel
convolutional neural network (CNN) pipelines to independently process network and radar
modality features, which are subsequently fused via a multi-head cross-attention
mechanism to capture intricate inter-modal dependencies. To enhance feature
discriminability, a supervised contrastive learning paradigm is incorporated, while a gradient
episodic memory (GEM) buffer strategically retains challenging instances to bolster model
robustness against hard-to-detect intrusions. Operating under a federated learning scheme,
FED-GEM-CN facilitates collaborative model optimization across distributed radar nodes,
preserving data locality and mitigating privacy risks inherent in centralized approaches.
Experimental evaluations conducted on a comprehensive real-world maritime radar dataset
reveal that FED-GEM-CN achieves superior performance, attaining an overall accuracy
exceeding 99 % and macro F1-scores above 0.97 across federated rounds, with convergence
typically observed within 15 communication iterations. These findings substantiate the
efficacy of the proposed system in delivering robust, energy-efficient, and privacy-aware
intrusion detection tailored to the constraints of maritime radar networks. The approach
underscores a significant advancement toward deploying intelligent, distributed
cybersecurity solutions within critical maritime infrastructures.
Keywords: Federated learning for maritime cybersecurity; Dual-CNN architecture; Cross-
attention mechanism; Intrusion detection system; Privacy-preserving learning; Multi-head
attention networks; Marine radar security

Jun Liu, Zhibo Kong, Xiaoying Wang, Li Wu, Guojing Zhang,


Enhancing multivariate weather forecasting via temporal attention and spatiotemporal
fusion,
Engineering Applications of Artificial Intelligence,
Volume 163, Part 4,
2026,
113133,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Accurate multivariate weather forecasting underpins agriculture, transportation,
and hazard preparedness, yet numerical weather prediction (NWP) is computationally
intensive and purely data-driven artificial intelligence (AI) models often struggle with scale-
coupled spatiotemporal dependencies. We introduce the Cross-Spatiotemporal Fusion
Model (C-STFM), a lightweight deep-learning framework based on a U-shaped encoder–
decoder convolutional neural network (U-Net) and built from stacked fusion blocks that
combine Sequential Temporal Attention Fusion (STAF) and a Cross-Spatiotemporal Fusion
Unit (CSFU). STAF performs ordered, segment-wise temporal fusion with carry-over, while
CSFU enables mutually modulated temporal and spatial attention to realize explicit cross-
variable interaction. Using the Fifth-Generation reanalysis from the European Centre for
Medium-Range Weather Forecasts (ERA5) over Qinghai Province, we evaluate C-STFM
against 17 representative convolutional, recurrent, video-prediction, and transformer
baselines under a unified split and preprocessing pipeline. With a 12-hour input window, we
report both 12-hour-ahead and 24-hour-ahead forecasts. Across six variables and four
metrics — root mean square error (RMSE), mean squared error (MSE), mean absolute error
(MAE), and anomaly correlation coefficient (ACC) — C-STFM achieves the lowest aggregate
MSE at both horizons and shows statistically significant improvements over the best single
baseline on all metrics (paired permutation tests with 95% bootstrap confidence intervals,
false discovery rate (FDR) controlled; p<1×10−4). Ablation studies isolate the contributions
of STAF and CSFU (on/off), the number of stacked fusion blocks (depth), and STAF temporal
segmentation, revealing that both modules are necessary and that moderate depth with
four temporal segments provides the best accuracy–efficiency trade-off. Efficiency
measurements (parameter counts, wall-clock training time, and inference throughput)
indicate that gains are not obtained simply by increasing model capacity. Overall, C-STFM
improves the stability of lead-time skill and the fidelity of multivariate spatial patterns,
offering a practical path toward fast and accurate short-term weather prediction. Source
code is available at [Link]
Keywords: Short-term multivariate weather forecasting; Artificial intelligence and deep
learning; Spatiotemporal attention; U-shaped encoder–decoder convolutional neural
network

Hongyan Dui, Huanqi Zhang, Shaomin Wu, Min Xie,


Spatiotemporal Resilience of IoT-Enabled Unmanned System of Systems,
Engineering,
Volume 54,
2025,
Pages 355-369,
ISSN 2095-8099,
[Link]
([Link]
Abstract: As advancements in the Internet of Things (IoT) and unmanned technologies
continues to progress, the development of unmanned system of systems (USS) has reached
unprecedented levels. While prior research has predominantly examined temporal
variations in USS resilience, spatial changes remain underexplored. However, USS may
involve kinetic engagements and frequent spatial changes during mission execution,
affecting signal interference in data layer communications. Although time-dependent factors
primarily govern mission effectiveness of the USS, spatial factors influence the transmission
stability of the data layer. Consequently, assessing spatiotemporal variations in USS
performance is critical. To address these challenges, this study introduces a spatiotemporal
resilience assessment framework, which evaluates USS resilience across both temporal and
spatial dimensions. Furthermore, we propose a spatiotemporal resilience optimization
scheme that enhances system adaptability throughout the mission lifecycle, with a particular
emphasis on prevention and recovery strategies. Finally, we validate the validity of the
proposed concepts and methods with a case study featuring a regular hexagonal
deployment of USS. The results show that the spatiotemporal resilience can better reflect
the spatial change characteristics of USS, and the proposed optimization strategy improves
the prevention spatiotemporal resilience, recovery spatiotemporal resilience, and entire-
process spatiotemporal resilience of USS by 0.22%, 8.39%, and 11.29%, respectively.
Keywords: Spatiotemporal performance; Spatiotemporal resilience; Unmanned equipment;
Importance measure
Shengyu Li, Shuolong Chen, Xingxing Li, Yuxuan Zhou, Shiwen Wang,
Accurate and automatic spatiotemporal calibration for multi-modal sensor system based on
continuous-time optimization,
Information Fusion,
Volume 120,
2025,
103071,
ISSN 1566-2535,
[Link]
([Link]
Abstract: Current intelligent robotic applications, such as unmanned aerial vehicles (UAV)
and autonomous driving, generally rely on multi-modal sensor fusion to continuously strive
towards higher levels of autonomy. To achieve this goal, accurate and consistent inter-sensor
spatiotemporal relationship is a fundamental prerequisite for fusing heterogeneous sensor
information. Nevertheless, current calibration frameworks typically necessitate specialized
tools or additional infrastructures, rendering them labor-intensive and only applicable to
certain sensor combinations. To address this issue, we propose an accurate and easy-to-use
spatiotemporal calibration framework tailored to current primary sensors, including inertial
measurement unit (IMU), LiDAR, camera and Radar. This calibration framework can be
seamlessly extended to other sensors that could independently recover ego-motion or ego-
velocity, such as wheel odometry and GPS devices. A rigorous multistage initialization
approach is first developed to obtain reasonable initial guesses of spatiotemporal
parameters without relying on prior knowledge of environmental information or specialized
movements. Leveraging the IMU-centric principle, the spatiotemporal parameters of other
sensors relative to IMU can be jointly optimized and refined via continuous-time batch
estimation without sharing the overlapping field-of-views (FoVs) among exteroceptive
sensors. A comprehensive series of experiments is carried out to quantitatively evaluate the
proposed method in both simulation and real-world scenarios. The results demonstrate that
the proposed method could achieve comparable calibration accuracy against state-of-the-art
target-based calibration methods and outperform targetless calibration methods in terms of
consistency and repeatability.
Keywords: Spatiotemporal calibration; Continuous-time optimization; Targetless; LiDAR;
Radar

Yash Soni, Malhaar Goswami, Nishit Prabhakar Shetty, Dhiraj,


Millimeter-wave radar for intelligent sensing: A comprehensive review of techniques,
applications, and challenges,
Computers and Electrical Engineering,
Volume 128, Part A,
2025,
110696,
ISSN 0045-7906,
[Link]
([Link]
Abstract: Millimeter-wave (mmWave) radar sensing has established itself as a robust
technology across diverse applications, such as automotive, healthcare, security, and smart
homes. Its exceptional capacity to function effectively in varying environmental conditions,
detect concealed objects, sense physiological signals, and facilitate precise target detection
positions it as a pivotal enabler for next-generation sensing solutions. The survey employs
bibliometric analysis to critically evaluate the existing literature surrounding mmWave radar,
highlighting key research trends, notable publications, and the challenges faced within the
field. This work presents a comprehensive examination of mmWave radar-based sensing,
detailing its fundamental operating principles, signal processing methodologies,
advancements in hardware, and the latest developments in machine learning applications. It
also addreses the key challenges in signal processing, including resolution enhancement,
environmental adaptability, and data fusion with complementary sensors such as LiDAR and
cameras. Furthermore, explored the potential of deep learning techniques to enhance target
classification, activity recognition, gesture identification, and healthcare applications while
addressing concerns related to accuracy and precision. This survey also sheds light on
emerging trends by assessing the strengths, limitations, and prospects of mmWave radar
technology. This review aims to provide insightful guidance for researchers and practitioners
committed to advancing radar-based sensing and its real-world implementations.
Keywords: mmWave radar; FMCW radar; Wireless sensing; Machine learning; Deep learning;
Application taxonomy

Cheng-Xi Tu, Yi-Jia Sun, Hu Gong, Jun-Chao Zhu,


In-situ acoustic characterization method for micro-tool ultrasonic vibration based on on-axis
sound pressure,
Measurement,
Volume 258, Part C,
2026,
119252,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Resonance identification and vibration amplitude quantification of the tool–
transducer ultrasonic vibration unit are critical for maximizing efficiency and improving
process performance in rotary ultrasonic–assisted machining (RUM) systems. However, due
to the complex geometry of the bottom edge surfaces, existing ultrasonic amplitude
measurement techniques struggle to perform accurate in-situ pre-calibration on most types
of micro tools, such as micro-mill and micro-twist drill. Acoustic measurement offers a
geometry-insensitive alternative, yet strategies tailored for micro-tool characterization
remain underexplored. This study proposes a sound-pressure-based method to characterize
micro-tool ultrasonic vibration in RUM systems. The near-field acoustic distribution of
cutting tool and its coupling with the transducer are investigated via theoretical modeling
and finite element simulation. A robust signal-processing strategy is introduced to extract
vibration features under practical conditions. The method is validated in RUM micro-drilling
of Ti-6Al-4V, demonstrating reliable resonance identification and relative amplitude
quantification. This provides a practical and low-cost tool for ultrasonic process optimization
in high-quality micromachining.
Keywords: Ultrasonic vibration measurement; Rotary ultrasonic machining; Micro tool
vibration calibration; Resonance identification; Sound pressure field

K. Venkata rao,
A study on performance characteristics and multi response optimization of process
parameters to maximize performance of micro milling for Ti-6Al-4V,
Journal of Alloys and Compounds,
Volume 781,
2019,
Pages 773-782,
ISSN 0925-8388,
[Link]
([Link]
Abstract: In machining of hard metals, surface roughness, tool vibration and tool wear are
used as performance characteristics to estimate overall performance of process. This work is
aimed to maximize overall performance in micromachining of Ti-6Al-4V and investigate
effect of process parameters on performance characteristics. As per orthogonal array of L27,
twenty seven experiments are carried out on the proposed metal with cemented carbide
tools at three levels of cutting speed, feed and depth of cuts. According to user's preference
rating, graph theory and matrix approach is used to estimate weights and preference scales
for the performance characteristics. Utility concept is used to calculate overall performance
of the process using weights and preference scales. Responses surface methodology is used
to optimize process parameters to maximize overall performance of the process. In this
study, maximum performance of the process is found at cutting speed of 19.78 m/min, feed
of 75 μm/tooth and depth of cut of 50 μm. In addition to that, metal recovery in machining
due to elastic recovery is also estimated theoretically and measured practically for different
uncut chip thicknesses.
Keywords: Micro milling; Optimization; Utility concept; Taguchi; Graph theory and matrix
approach (GTMA)

Martin Brenner, Napoleon H. Reyes, Teo Susnjak, Andre L.C. Barczak,


MM5: Multimodal image capture and dataset generation for RGB, depth, thermal, UV, and
NIR,
Information Fusion,
Volume 126, Part A,
2026,
103516,
ISSN 1566-2535,
[Link]
([Link]
Abstract: Existing multimodal datasets often lack sufficient modality diversity, raw data
preservation, and flexible annotation strategies, seldom addressing modality-specific cues
across multiple spectral channels. Current annotations typically concentrate on pre-aligned
images, neglecting unaligned data and overlooking crucial cross-modal alignment
challenges. These constraints significantly impede advanced multimodal fusion research,
especially when exploring modality-specific features or adaptable fusion methodologies. To
address these limitations, we introduce MM5, a comprehensive dataset integrating RGB,
depth, thermal (T), ultraviolet (UV), and near-infrared (NIR) modalities. Our capturing system
utilises off-the-shelf components, incorporating stereo RGB-D imaging to provide additional
depth and intensity (I) information, enhancing spatial perception and facilitating robust
cross-modal learning. MM5 preserves depth and thermal measurements in raw, 16-bit
formats, enabling researchers to explore advanced preprocessing and enhancement
techniques. Additionally, we propose a novel label re-projection algorithm that generates
ground-truth annotations directly for distorted thermal and UV modalities, supporting
complex fusion strategies beyond strictly aligned data. Dataset scenes encompass varied
lighting conditions (e.g. shadows, dim lighting, overexposure) and diverse objects, including
real fruits, plastic replicas, and partially rotten produce, creating challenging scenarios for
robust multimodal analysis. We evaluate the effects of multi-bit representations, adaptive
gain control (AGC), and depth preprocessing on a transformer-based segmentation network.
Our preprocessing improved mean IoU from 70.66% to 76.33% for depth data and from
72.67% to 79.08% for thermal encoding, using our novel preprocessing techniques,
validating MM5’s efficacy in supporting comprehensive multimodal fusion research.
Keywords: Multimodal dataset; Thermal imaging; UV imaging; Preprocessing; Sensor fusion;
Dataset annotation

YongTeng Sun, HongZhong Ma,


Research progress on oil-immersed transformer mechanical condition identification based
on vibration signals,
Renewable and Sustainable Energy Reviews,
Volume 196,
2024,
114327,
ISSN 1364-0321,
[Link]
([Link]
Abstract: In recent years, vibration signals have been widely applied for the identification of
mechanical states in oil-immersed transformers. This paper, following the framework of
‘vibration generation – sensing – processing – recognition – evaluation – solution,’
introduces the progress in mechanical state recognition of oil-immersed transformers based
on vibration signals from a novel sensor-oriented perspective, which covers sensor
deployment, sensor specialization, and equipment integration. The advancements in signal
processing and feature selection are also discussed and compared with the identification of
states in rotating machinery. To Address challenges like limited rule transferability and the
weakness in vibration characteristics and models, some emerging technologies such as
Operational Modal Analysis and multisource data fusion are introduced, which may bring
new prospects. This paper aims to provide scholars engaged in research on the mechanical
state identification of transformers and other electrical equipment with some technical
references.
Keywords: Oil-immersed transformer; Mechanical condition identification; Vibration signals;
Feature extraction

Bachina Harish Babu, Sujith Bobba, T.C.H. Anil Kumar, NB. Prakash Tiruveedula, Talluri
Srinivasarao,
Optimization of dead metal zone to reduce cutting forces in micro milling of Inconel 718
using RSM,
Materials Today: Proceedings,
2023,
,
ISSN 2214-7853,
[Link]
([Link]
Abstract: This study focuses on the mechanism of DMZ (dead metal zone) creation, as well
as the impact of cutting edge geometries (sharp, chamfered, double chamfered, and blunt
edges), cutting speed, and coefficient of friction on DMZ formation while milling Inconel 718
material (FEM). A non-contact type sensor called a laser doppler vibrometer (LDV) is used to
monitor the vibration of rotating surfaces. In current research work, the LDV is used to
measure the mill cutter vibration in micro-milling of Inconel 718 in terms of acoustic optic
emission signals. A FFT (fast fourier transformer) is used for signals processing in to
frequency domain. Design of experiments as per Taguchi, experiments were performed on
the alloy at three levels of spindle speeds, depth of cuts, feed rates. Experimental results on
the amplitude o vibration of tool along X and Y directions, surface roughness were measured
and analysed using response surface methodology. Analysis obtained from the variance was
used to recognize the significant parameters which effect the vibration of tool and roughness
of surface. RSM was implemented and optimized process parameters for the minimum
vibration amplitude and surface roughness.
Keywords: Laser Doppler vibrometer; Inconel 718; Surface roughness; Tool vibration; DMZ,
Response surface methodology (RSM)

Nayeemul Islam Nayeem, Shirin Mahbuba, Sanjida Islam Disha, Md Rifat Hossain Buiyan,
Shakila Rahman, M. Abdullah-Al-Wadud, Jia Uddin,
A YOLOv11-Based Deep Learning Framework for Multi-Class Human Action Recognition,
Computers, Materials and Continua,
Volume 85, Issue 1,
2025,
Pages 1541-1557,
ISSN 1546-2218,
[Link]
([Link]
Abstract: Human activity recognition is a significant area of research in artificial intelligence
for surveillance, healthcare, sports, and human-computer interaction applications. The
article benchmarks the performance of You Only Look Once version 11-based (YOLOv11-
based) architecture for multi-class human activity recognition. The article benchmarks the
performance of You Only Look Once version 11-based (YOLOv11-based) architecture for
multi-class human activity recognition. The dataset consists of 14,186 images across 19
activity classes, from dynamic activities such as running and swimming to static activities
such as sitting and sleeping. Preprocessing included resizing all images to 512 × 512 pixels,
annotating them in YOLO’s bounding box format, and applying data augmentation methods
such as flipping, rotation, and cropping to enhance model generalization. The proposed
model was trained for 100 epochs with adaptive learning rate methods and hyperparameter
optimization for performance improvement, with a mAP@0.5 of 74.93% and a mAP@0.5-
0.95 of 64.11%, outperforming previous versions of YOLO (v10, v9, and v8) and general-
purpose architectures like ResNet50 and EfficientNet. It exhibited improved precision and
recall for all activity classes with high precision values of 0.76 for running, 0.79 for
swimming, 0.80 for sitting, and 0.81 for sleeping, and was tested for real-time deployment
with an inference time of 8.9 ms per image, being computationally light. Proposed
YOLOv11’s improvements are attributed to architectural advancements like a more complex
feature extraction process, better attention modules, and an anchor-free detection
mechanism. While YOLOv10 was extremely stable in static activity recognition, YOLOv9
performed well in dynamic environments but suffered from overfitting, and YOLOv8, while
being a decent baseline, failed to differentiate between overlapping static activities. The
experimental results determine proposed YOLOv11 to be the most appropriate model,
providing an ideal balance between accuracy, computational efficiency, and robustness for
real-world deployment. Nevertheless, there exist certain issues to be addressed, particularly
in discriminating against visually similar activities and the use of publicly available datasets.
Future research will entail the inclusion of 3D data and multimodal sensor inputs, such as
depth and motion information, for enhancing recognition accuracy and generalizability to
challenging real-world environments.
Keywords: Human activity recognition; YOLOv11; deep learning; real-time detection; anchor-
free detection; attention mechanisms; object detection; image classification; multi-class
recognition; surveillance applications

Md Shohag Mollik, Tanveer Saleh, Khairul Affendy Bin Md Nor, Mohamed Sultan Mohamed
Ali,
A machine learning-based classification model to identify the effectiveness of vibration for
μEDM,
Alexandria Engineering Journal,
Volume 61, Issue 9,
2022,
Pages 6979-6989,
ISSN 1110-0168,
[Link]
([Link]
Abstract: Micro electro-discharge machining (μEDM) uses electro-thermal energy from
repetitive sparks generated between the tool and workpiece to remove material from the
latter. However, one of the bottlenecks of μEDM is the phenomenon of short circuits due to
the physical contact between the tool and debris (formed during the erosion of the
workpiece). Adequate flushing of the debris can be achieved by applying low amplitude
high-frequency vibration to the workpiece. This study, however, shows that the application
of vibration does not yield beneficial results for the μEDM for all the parametric conditions.
This research used an off-the-shelf piezo vibrator as the high-frequency, low amplitude
vibration source to the workpiece during the μEDM process. The experiments were
conducted with and without vibration with the variation of applied discharge energy and
μEDM speed. The samples were characterized using scanning electron microscopes to gather
various data related to μEDM outputs. The results of this study revealed that vibration-
assisted μEDM becomes less effective as the discharge energy is increased (primarily by
increasing the capacitor value of the RC pulse generator). Similarly, the reduction of the
occurrence of the short circuit was profound when the low discharge energy level with low
voltage and low capacitor setting of the RC Pulse generator was used. The overall scale of
the overcut with various discharge energy and μEDM speed varied from 15.5 μm to 42 μm
for the conventional μEDM process. However, the scale above slightly reduced to 14.5 μm to
39 μm using an ultrasonic vibration device. Also, the taperness of the machined hole was
of ∼7%).
slightly reduced by applying the vibration device during the μEDM operation (overall average

Keywords: Machine learning; Micro electro discharge machining; Ultrasonic vibration; MRR;
Tool wear; EDM

K. Venkata Rao,
Power consumption optimization strategy in micro ball-end milling of D2 steel via TLBO
coupled with 3D FEM simulation,
Measurement,
Volume 132,
2019,
Pages 68-78,
ISSN 0263-2241,
[Link]
([Link]
Abstract: The present challenge in the manufacturing industry is to improve efficiency of
production activities while reducing wastage of power consumption. Past research focused
on multi response optimization of process parameters to improve performance of the
process. The present study proposed an optimization-based strategy to reduce power
consumption in micro ball end milling of D2 steel. As the power consumption is directly
proportional to cutting forces, the process parameters such as cutting speed, feed and depth
of cut were optimized to reduce cutting forces using teaching learning based optimization
(TLBO) technique coupled with 3D finite element method (FEM) simulation. During the
optimization, amplitude of cutter vibration and surface roughness were taken as constraints
as 60 µm (ISO 10816) and 2 µm (ISO 1302) respectively. Three best combinations of cutting
speed, feed and depth of cut were obtained for minimum cutting force. Among them,
combination of cutting speed of 15 m/min, feed of 112.5 µm/tooth and depth of cut of
85.25 µm has low power consumption of 67 W with tool vibration of 36.5 µm. However,
remaining two combinations were also considered to be the next best optimal cutting
conditions. Numerical simulation was carried out for the three best solutions and the cutting
forces and amplitude of cutter vibration were predicted. There was good agreement
between simulation results and experimental results that verified the acceptance of the
simulation. It was also found that the three best candidate solutions were having same the
cutting speed of 15 m/min (minimum cutting speed). Hence, the induced stresses in the
work piece were found to be with low values around 350 Mpa.
Keywords: Power consumption; Ball end milling; Simulation; Optimization; TLBO; Tool
vibration

Ayesha Ibrahim, Muhammad Zakir Khan, Muhammad Imran, Hadi Larijani, Qammer H.
Abbasi, Muhammad Usman,
RadSpecFusion: Dynamic attention weighting for multi-radar human activity recognition,
Internet of Things,
Volume 33,
2025,
101682,
ISSN 2542-6605,
[Link]
([Link]
Abstract: This paper presents RadSpecFusion, a novel dynamic attention-based fusion
architecture for multi-radar human activity recognition (HAR). Our method learns activity-
specific importance weights for each radar modality (24 GHz, 77 GHz, and Xethru sensors).
Unlike existing concatenation or averaging approaches, our method dynamically adapts
radar contributions based on motion characteristics. This addresses cross-frequency
generalization challenges, where transfer learning methods achieve only 11%–34% accuracy.
Using the CI4R dataset with spectrograms from 11 activities, our approach achieves 99.21%
accuracy, representing a 15.8% improvement over existing fusion methods (83.4%). This
demonstrates that different radar frequencies capture complementary information about
human motion. Ablation studies show that while the three-radar system optimizes
performance, dual-radar combinations achieve comparable accuracy (24GHz+77GHz: 96.1%,
24GHz+Xethru: 95.8%, 77GHz+Xethru: 97.2%), enabling flexible deployment for resource-
constrained applications. The attention mechanism reveals interpretable patterns: 77 GHz
radar receives higher weights for fine movements (superior Doppler resolution), while 24
GHz dominates gross body movements (better range resolution). The system maintains
71.4% accuracy at 10 dB SNR, demonstrating environmental robustness. This research
establishes a new paradigm for multimodal radar fusion, moving from cross-frequency
transfer learning to adaptive fusion with implications for healthcare monitoring, smart
environments, and security applications.
Keywords: Human activity recognition; Multi-modal fusion; Attention mechanisms; Cross-
frequency transfer learning

Nadav Cohen, Itzik Klein,


Adaptive Kalman-Informed Transformer,
Engineering Applications of Artificial Intelligence,
Volume 146,
2025,
110221,
ISSN 0952-1976,
[Link]
([Link]
Abstract: The extended Kalman filter (EKF) is a widely adopted method for sensor fusion in
navigation applications. A crucial aspect of the EKF is the online determination of the
process noise covariance matrix reflecting the model uncertainty. While common EKF
implementation assumes a constant process noise, in real-world scenarios, the process noise
varies, leading to inaccuracies in the estimated state and potentially causing the filter to
diverge. Model-based adaptive EKF methods were proposed and demonstrated performance
improvements to cope with such situations, highlighting the need for a robust adaptive
approach. In this paper, we derive an adaptive Kalman-informed transformer (A-KIT)
designed to learn the varying process noise covariance online. Built upon the foundations of
the EKF, A-KIT utilizes the well-known capabilities of set transformers, including inherent
noise reduction and the ability to capture nonlinear behavior in the data. This approach is
suitable for any application involving the EKF. In a case study, we demonstrate the
effectiveness of A-KIT in nonlinear fusion between a Doppler velocity log and inertial
sensors. This is accomplished using real data recorded from sensors mounted on an
autonomous underwater vehicle operating in the Mediterranean Sea. We show that A-KIT
outperforms the conventional EKF by more than 49.5% and model-based adaptive EKF by an
average of 35.4% in terms of position accuracy.
Keywords: Inertial sensing; Navigation; Set transformer; Kalman filter; Sensor fusion;
Autonomous underwater vehicle

Sui Tan, Li Zhu, Jia-Huan Li, Zhi-Ruo Cui, Guan-Yuan Zhao, Jia-Yang Tu,
Non-contact vibration displacement measurement method for railway bridges based on
computer vision,
Structures,
Volume 80,
2025,
109803,
ISSN 2352-0124,
[Link]
([Link]
Abstract: To address the demand for efficient, non-contact vibration measurement of railway
bridges, this study proposes a novel structural displacement measurement method based on
Recurrent All-pairs Field Transforms (RAFT) optical flow estimation. It utilizes a pre-trained
RAFT network to infer inter-frame optical flows, extracting the average within a selected
region of interest (ROI) as the pixel displacement, and then converting it into real-world via
coordinate transformation. An indoor cable-stayed bridge model vibration test
demonstrated that all three methods exhibit increasing EMSE with increasing amplitude, but
the proposed method shows the slowest increase in error, with a relative RMSE still below
20 % at 0.297 mm (RMSE: 0.057 mm), while the LightTrack-based method (RMSE: 0.067 mm)
and the GMFlow-based method (RMSE: 0.097 mm) both have relative RMSE exceeding 20 %.
Under large amplitude vibrations (13.222–30.469 mm), all three methods exhibit
comparable accuracy, with relative RMSE below 10 %. Frequency identification results show
a relative error of less than 2.9 % for the first-order frequency across all cases. To further
validate its applicability in practical engineering, a Linear Variable Differential Transformer
(LVDT) was installed at the midspan of a railway simply-supported beam bridge, and visual
measurement devices were arranged both under and on the bridge, with vibration images
and LVDT data captured during train passing. Results show the proposed method’s midspan
deflection closely matches LVDT data, with a peak absolute error of 0.087 mm and RMSE of
0.037 mm, lower than the LightTrack - based method (0.258 mm peak error, 0.080 mm
RMSE). Ultimately, the proposed method was used to identify the deflection at the quarter
points, the relative displacement between the beam end and pier, and the relative
displacement between the track and track slab as well, which provides a reference for
similar engineering applications.
Keywords: Computer vision; Non-contact detection; Railway bridge health monitoring; Deep
learning; Vibration test

Archana Mathur, Abbas Mufaddal Dudhiyawala, Sudeepa Roy Dey, Snehanshu Saha,
Toward accurate breast cancer classification: A review of multi-modal machine learning
approaches,
Methods,
Volume 246,
2026,
Pages 48-61,
ISSN 1046-2023,
[Link]
([Link]
Abstract: The innovations in classifying breast cancer into malignant and benign categories
and further categorizing it into molecular subtypes have reshaped healthcare services,
enabling accurate diagnosis of these complex conditions. Identification of molecular
subtypes of breast cancer is one of the most important treatment challenges, as these
subtypes can have an enormous effect on the prognosis and treatment approaches. Data
integration from various modalities, such as transcriptomics, imaging, and genomics, has
been crucial in leveraging new opportunities to increase classification accuracy and improve
individualized treatment plans. These heterogeneous data sources are examined by applying
deep learning algorithms, which provide further insights into the complex patterns that
traditional approaches often overlook. In this paper, we explore the various modalities
researchers use to investigate breast cancer and the intriguing fusion techniques employed
to combine these modalities. We also review the most recent models (traditional, machine
learning, and deep learning), emphasizing their improvements over traditional classification
methods and the molecular subtype categorization of breast cancer. Furthermore, the
emphasis of this review is to examine techniques to process the entire image of the breast
tissue slide, which is challenging, particularly due to its size. We explore recent advances in
multiple instance learning tasks and the use of attention-based transformers and similar
architectures for annotating the WSI slides before using them for cancer classification. We
additionally discuss the interpretability tools—attention maps, saliency maps and model
explainability— in the context of transformers. In a nutshell, we aim to provide an in-depth
look at the revolutionary capabilities of deep learning models in precision oncology and
guide future research paths in this crucial field by synthesizing existing studies.
Keywords: Multimodality; Molecular subtype classification; Breast cancer prediction; Feature
fusion; Multiple instance learning; Whole slide imaging

Tiantian Wang, Nan Yan, Chaosan Yang, Zeliang An, Gongjing Zhang, Yuqing Xu,
Electromagnetic signal recognition using multimodal tri-branch semantic fusion network in
the UAV-assist integrated sensing and communication systems,
Digital Signal Processing,
Volume 171,
2026,
105820,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Driven by the proliferation of integrated sensing and communication (ISAC)
systems, the accurate recognition of unauthorized unmanned aerial vehicle (UAV) signals in
dynamic electromagnetic environments has emerged as a critical challenge for spectrum
security and cognitive radio applications. Conventional automatic modulation recognition
(AMR) frameworks suffer from significant performance degradation in low signal-to-noise
ratio (SNR) regimes and exhibit limited adaptability to resource-constrained edge computing
platforms. To address these limitations, we propose a novel Multimodal Tri-branch Fusion
Network (MTF-Net) architecture that synergistically integrates time-frequency analysis with
statistical feature learning. The framework systematically processes binarized time-
frequency images (B-TFIs) and higher-order cumulant vectors through three collaboratively
operating branches: (1) A primary temporal feature extractor employing dilated convolution-
residual blocks (DCRBlocks) with hierarchical dilatation factors, incorporating channel
attention mechanisms to dynamically emphasize discriminative temporal patterns; (2) Dual
auxiliary branches based on Edge-Transformer modules (ETFormers), which achieve efficient
spatial-structural learning through depthwise separable convolutions (DSC) while capturing
long-range spectral dependencies via additive attention mechanisms with linear complexity;
(3) A hierarchical fusion module implementing cross-branch feature recalibration through
learnable parameter matrices. Extensive Monte Carlo experiments demonstrate that our
MTF-Net significantly outperforms traditional methods in recognition accuracy for radar and
communication signals under low SNR conditions, establishing a new benchmark for
lightweight AMR solutions in ISAC systems.
Keywords: Multi-modal feature fusion; Unmanned aerial vehicle(UAV); Integrated sensing
and communication (ISAC); Lightweight neural network; Transformer

Vladislav Semenyuk, Ildar Kurmashev, Alberto Lupidi, Dmitriy Alyoshin, Liliya Kurmasheva,
Alessandro Cantelli-Forti,
Advances in UAV detection: integrating multi-sensor systems and AI for enhanced accuracy
and efficiency,
International Journal of Critical Infrastructure Protection,
Volume 49,
2025,
100744,
ISSN 1874-5482,
[Link]
([Link]
Abstract: This review critically examines the progress in unmanned aerial vehicle (UAV)
detection and classification technologies from 2020 to the present. It highlights a range of
detection methods, including radar, radio frequency (RF), optical, and acoustic sensors, with
particular emphasis on the integration of these technologies through advanced sensor
fusion techniques. The paper explores the core technologies driving improvements in
detection accuracy, range, and reliability, with a special focus on the transformative role of
artificial intelligence and machine learning. These innovations have significantly enhanced
system performance, enabling more precise and efficient UAV detection. The review
concludes with insights into emerging trends and future developments that promise to
further refine UAV detection technologies, ensuring greater security and operational
reliability.
Keywords: UAV Detection; UAV Classification; Radar Technology; Sensor Fusion; Optical
Sensor; Acoustic Sensor

Jitao Zhang, Xingkui Mu, Qingfang Zhang, Natallia Poddubnaya, Dmitry Filippov, Jiagui Tao,
Fang Wang, Liying Jiang, Lingzhi Cao,
Structural, micro-structure, magnetic and dynamic magneto-elastic properties of samarium-
doped nickel-zinc spinel ferrites for efficient power conversion applications,
Journal of Magnetism and Magnetic Materials,
Volume 602,
2024,
172176,
ISSN 0304-8853,
[Link]
([Link]
Abstract: Development of high-quality ferrites behaved enhanced properties via inclusion of
ions in 3dn and 4fn series are desirable for efficient power conversion solid-state devices.
Nevertheless, inadequate microscopic behaviors with complex chemistry in materials
enables researchers to design and produce the macroscopic target device that remained
rudimentary and mindless. In this work, the microscopic properties in nickel-zinc spinel
ferrite series of Ni0.8Zn0.2SmxFe2-xO4 (x = 0, 0.02, 0.04, 0.06, 0.08) encompassing XRD,
SEM, EDS, VSM, FTIR and ESR were systemically characterized, and the evolutionary
mechanism of the corresponding structural, micro-structure, elemental composition and
distribution, magnetic, cations exchange, and micro-magnetic behavior was profoundly
revealed. Under microscopic examinations, the optimum composition at x = 0.02 with well-
arranged spinel structure, dense texture with expected elemental composition, favorable
soft magnetic properties and micro-magnetic behaviors is evident. Fortunately, this optimum
is in coincidence with the achievable maximum magneto-mechanical coefficient, even the
macroscopic electric properties with stronger magnetoelectric (ME) interactions and higher
power conversion efficiency (PE) in tri-layered ME samples as expected. Experimental results
show that the eventual PE reaches its maximum of 75.67 % under Ropt = 33kΩ for samples
at x = 0.02, and exhibit a 2.24 times higher PE than that of sample without samarium
substitution. These findings provide a holographic perspective to connect the microscopic
beneficial effects of materials to bulk device that are promising for efficient power
conversion solid-state electronics.
Keywords: Samarium-doped spinel ferrites; Dynamic magneto-elastic properties; Power
conversion devices

Ruige Yang, Peng Shan, Yang He, Hongming Xiao, Lin Zhang, Yuliang Zhao, Qiang Fu,
A lightweight bionic flapping wing drone recognition network based on data enhancement,
Measurement,
Volume 239,
2025,
115476,
ISSN 0263-2241,
[Link]
([Link]
Abstract: In recent times, there has been a growing focus on research into bionic drones,
which seek to mimic biological behavior and structure, thus overcoming the limitations of
conventional drones. The ability of bionic drones to blend into their surroundings presents a
significant challenge for identification. This study presents a dataset of bionic drones and
introduces the Bionic Drone Identification Network (BDRNet). The dataset was enriched
using data augmentation techniques to improve model recognition. Moreover, an
Aggregated Attention Mechanism (AAM) captures input feature correlation. Furthermore, a
Merged and Integrated Detector Head (MIDHead) and Multi-scale Lightweight Convolution
(MLWConv) have been proposed to lessen computational costs. The findings reveal that
BDRNet achieves an AP0.5 of 94.4 %, a Params of 2.6 M, and a Flops of 5.5G, surpassing
mainstream object detection models such as Faster R-CNN and YoloV5. This suggests that
BDRNet demonstrates robust recognition capabilities and holds potential for deployment in
resource-constrained embedded devices.
Keywords: Bionic flapping-wing drone Identification; Data augmentation; Lightweight;
Engineering application

Benjamin Rise, Murat Uney, Xiaowei Huang,


Two-stage transfer learning for airborne multi-spectral image classifiers,
Signal Processing,
Volume 240,
2026,
110358,
ISSN 0165-1684,
[Link]
([Link]
Abstract: In this work, we propose a novel training paradigm designed to support transfer
learning for more effective classification in multispectral airborne imagery. Current state-of-
the-art approaches typically rely on either leveraging solely RGB (red-green-blue) pretraining
or applying in-domain transfer learning for multispectral imagery classification. Instead, our
approach constructs and trains two separate neural network models (backbones): one
specifically for wavelengths with available pretrained data (like visible bands) and another
trained from scratch on all-bands available in the dataset. These models are then integrated
with a fully-connected layer or multi-layered perceptron, which is trained on the features
from both networks. This allows us to exploit the significant benefits of generalizable
features learned from RGB datasets and the information provided by the full spectrum of
multispectral bands. We employ the BigEarthNet and EuroSAT datasets, encompassing
Sentinel-2 satellite imagery in the visual and infrared bands. This approach yields
considerable performance gains in comparison with other training strategies across every
evaluation metric we utilized for these datasets. The results are also consistent across a
variety of backbone architectures, underlining the efficacy of our transfer learning technique
in the analysis of multispectral data.
Keywords: Multispectral imagery; Remote sensing; Deep learning; Transfer learning; Scene
classification

J.R.J. Bennett, G.P. Škoro, John Back, S.J. Brooks, T.R. Edgecock, S.A. Gray, A.J. McFarland, K.J.
Rodgers, C.N. Booth,
Lifetime and strength tests of tantalum and tungsten under thermal shock for a Neutrino
Factory target,
Nuclear Instruments and Methods in Physics Research Section A: Accelerators,
Spectrometers, Detectors and Associated Equipment,
Volume 646, Issue 1,
2011,
Pages 1-6,
ISSN 0168-9002,
[Link]
([Link]
Abstract: A description is given of tests on tantalum and tungsten wires to evaluate their
lifetime and strength under the thermal shock that will be experienced when a solid target is
bombarded with short pulses of high energy protons in a Neutrino Factory. The results of
lifetime tests and measurements of dynamic strength characteristics at high temperatures,
stresses and strain rates using a laser Doppler vibrometer are given. The tests show that a
solid tungsten target will have a life of at least 3 years, which, with other beneficial
characteristics, make it an excellent candidate for the Neutrino Factory.
Keywords: Tantalum; Tungsten; Thermal shock; Material strength; Target lifetime; Neutrino
Factory

Pieter G.G. Muyshondt, Lukas Prochazka, Merlin Schär, Michail Chatzimichalis, Bastian
Baselt, Guy Fierens, Flurin Pfiffner,
Finite-element modelling of the 3D motion of the malleus-incus complex validated with 3D
laser Doppler vibrometry,
Hearing Research,
Volume 469,
2026,
109477,
ISSN 0378-5955,
[Link]
([Link]
Abstract: Three-dimensional (3D) motions of the middle ear (ME) are investigated with
finite-element (FE) modelling by comparison with 3D laser Doppler vibrometer (LDV)
measurements of the malleus-incus complex. 3D point velocity measurements are converted
to 3D rigid-body motion (RBM) components of the malleus and incus under acoustic
excitation of the ME from 0.2 kHz to 8 kHz. The parameters in the FE model are adjusted to
provide qualitative agreement with the 3D motion measurements for three separate model
geometries. The results show a dominant hinge-like motion for malleus and incus across the
frequency range, but with an increase of other components at high frequencies to yield a
more complex motion. Incudomallear joint flexibility increases the relative motion between
malleus and incus and is shown to contribute most to the ME transformer ratio at low and
especially high frequencies, including the phase delay across the two ossicles. The dominant
motion direction of the umbo coincides with the medial-lateral axis across the frequency
range. The malleus head, incus head and incus long process show a deviation from this
motion direction between 1.5 kHz and 5 kHz, associated with dips in the corresponding
velocity magnitude. Motion trajectories at these points follow a line below 1.5 kHz but
alternate between a line and ellipse at higher frequencies. While the tympanic membrane
influences the 3D motion of malleus and incus in a similar way, the ME suspensory ligaments
affect the motion components of the ossicles to varying degrees depending on the location
on the ossicles.
Keywords: Middle ear mechanics; Middle ear modelling; 3D laser Doppler vibrometry; Finite
element modelling

Dongfang Wang, Yufeng Yang, Jilin Lei, Baojian Wang, Qiming Ouyang, Penghao Yin,
Frontier exploration of image recognition in fuel spray diagnostics: Hybrid deep learning
models and multimodal data fusion,
Journal of the Energy Institute,
Volume 123,
2025,
102274,
ISSN 1743-9671,
[Link]
([Link]
Abstract: Owing to its high detection accuracy and real-time processing capabilities, image
recognition technology has become an indispensable tool for extracting spray morphological
characteristics and analyzing dynamic evolution processes in combustion systems. This
review systematically summarizes recent advances in image recognition technology, with a
focus on its applications in multiphase flow coupling and detailed feature extraction within
spray environments, while also providing a forward-looking discussion of current challenges
and future trends. Conventional methods are limited by weak anti-interference capability,
low feature extraction efficiency, and poor generalization. Deep learning techniques have
been increasingly adopted to enhance boundary segmentation precision and quantitative
feature parameter extraction. However, a major challenge remains in adapting these
technologies to complex environments, as most existing models struggle to balance
lightweight design with measurement accuracy—a critical barrier to real-time engineering
applications. Emerging approaches, including hybrid CNN–Transformer architectures and
novel Mamba-based models such as UltraLight_VM_UNet, have demonstrated significant
potential. The model achieves a segmentation accuracy of up to 95.43 % mIoU for complex
sprays, while reducing computational costs to just 0.05M parameters and 0.33 GFLOPs.
These advancements significantly improve robustness and generalization under noisy and
dynamic spray conditions. Future developments are expected to focus on computational
efficiency, robustness in extreme scenarios, and more effective global–local feature fusion,
thereby paving the way for real-time diagnostic applications in combustion systems.
Keywords: Spray combustion diagnostics; Image recognition technology; Image acquisition;
Parameter extraction; Deep learning

Zhiyan Lin, Minming Gu, Keyu Pan, Wei-Ping Zhu,


Adaptive temporal convolutional network with multi-head EMA-gated attention for
continuous radar-based human activity recognition,
Biomedical Signal Processing and Control,
Volume 117,
2026,
109667,
ISSN 1746-8094,
[Link]
([Link]
Abstract: Continuous human activity recognition (HAR) using radar signals offers strong
potential for privacy-preserving clinical health monitoring. However, its performance is
limited by challenges such as multi-scale temporal variation, signal noise, and unstable
activity transitions. To address these issues, this study introduces a radar-based HAR
framework with three tailored components. First, an adaptive temporal convolutional
network (ATCN) uses learnable dilation rates and sampling offsets to flexibly capture both
abrupt and periodic motion patterns over time. Second, an exponential moving average
(EMA)-gated attention (EDGA) module integrates linear attention with exponential moving
average smoothing through a dynamic gating mechanism, effectively suppressing noise
while preserving temporal continuity. Third, an attention-guided multi-stage refinement
(AMSR) module refines coarse predictions using global attention-driven residual corrections,
thereby reducing segmentation noise and improving boundary precision. Experiments on a
77 GHz frequency modulated continuous wave (FMCW) radar dataset show that the
proposed model achieves 96.09% accuracy, demonstrating its strong potential for
continuous and unobtrusive activity monitoring in healthcare applications.
Keywords: ATCN; EDGA; AMSR; Continuous HAR; FMCW radar

Juan José Villamarín Marrugo, Juan Manuel Naranjo Piñeros, Erwin Hernando Hernandez
Rincon,
Evidence on the Utility of Artificial Intelligence in the Interpretation of Diagnostic
Radiological Images in Low and Middle-Income Countries: A Scoping Review,
Academic Radiology,
2025,
,
ISSN 1076-6332,
[Link]
([Link]
Abstract: Rationale and Objectives
Access to diagnostic imaging in low- and middle-income countries (LMICs) is limited by
scarce equipment, geographic barriers, weak digital infrastructure, and shortages of trained
personnel. Artificial intelligence (AI) has emerged as a promising tool to mitigate these gaps
by improving diagnostic accuracy, assisting non-specialist health workers, and optimizing
workflows. This scoping review aimed to synthesize current evidence on the use of AI for
interpreting radiological diagnostic images in LMICs.
Materials and Methods
A scoping review was conducted in July 2025 following Arksey and O’Malley’s framework
and PRISMA-ScR guidelines. Searches were performed in PubMed, Scopus, and Clinical Key
for studies published between 2000 and July 2025 in English and Spanish. Eligible studies
included clinical applications of AI in radiological imaging within LMICs, reporting relevant
outcomes.
Results
From 620 records, 51 studies conducted across 33 LMICs were included. Most were
published between 2022 and 2025 and focused on ultrasound, X-ray, and computed
tomography. AI consistently improved diagnostic sensitivity, specificity, and applicability,
particularly for tuberculosis, pneumonia, obstetric care, and oncologic screening. Magnetic
resonance imaging showed promising yet mostly experimental evidence, while
mammography research remained scarce. Frequent limitations included small sample sizes,
single-center designs, reliance on public datasets, and limited multicenter validation.
Conclusion
AI demonstrates significant potential to enhance the interpretation of diagnostic radiological
images in LMICs, with consistent gains in sensitivity, specificity, and applicability across
modalities such as ultrasound, X-ray, and computed tomography. Several studies also
reported improvements in workflow efficiency and support for non-specialist providers,
underscoring AI’s dual role as a diagnostic and operational tool. Nonetheless,
methodological heterogeneity and infrastructural challenges highlight the need for
multicenter validation and context-adapted implementation strategies to ensure sustainable
integration.
Keywords: artificial intelligence; diagnostic imaging; low- and middle-income countries;
radiology; scoping review

Debashis Dey, Ashutosh Mahanta, Sukanta K. Dash,


Flow visualization of nanoparticles in a natural convection cavity,
Materials Today: Proceedings,
2023,
,
ISSN 2214-7853,
[Link]
([Link]
Abstract: A two-litre natural convection cavity containing transformer oil is used to visualize
the flow of nanoparticles. Transparency-wise, transformer oil resembles water. Kinematic
viscosity, however, is nineteen times greater than water. Its thermophysical characteristic
aids in keeping the nanoparticles suspended for a longer period. Transparency gives equal
clarity like demineralized water. As a result, the flow of nanoparticles may be seen more
clearly in transformer oil than in demineralized water. The following cases are seen in the
motions of the natural convection nanoparticles that resemble cavities in the transformer oil
due to the effects of a) gravity, b) mechanical agitation, c) temperature, d) rotational
magnetic field, e) cavity tilting, etc. The random movement of nanoparticles,
thermophoresis due to temperature gradient, and improvement in settling time of
nanoparticles due to rotary magnetic field, shifting of vortex towards the upper tilting wall,
and path line of few particles are also observed.
Keywords: Natural convection; Nanoparticles; Flow Visualization; Transformer Oil; Macro
Lens; DSLR camera

Yu Han, Panpan Wen, Zhuoying Liu, Rui Yi, Yinuo Chen, Sheng Cao,
From Recognition to Action: Integrating Deep Learning and Robotic Control in Transthoracic
Echocardiography,
Ultrasound in Medicine & Biology,
2026,
,
ISSN 0301-5629,
[Link]
([Link]
Abstract: Population aging has driven a rise in heart failure cases, increasing the clinical
burden on cardiac diagnostics. As a first-line imaging method, transthoracic
echocardiography (TTE) faces limitations due to operator dependence, patient variability,
and workflow inefficiencies. Meanwhile, advances in artificial intelligence (AI) and robotic
ultrasound systems offer new potential pathways toward automated diagnosis. This review
examines the current landscape of AI-based image analysis and robotic-assisted
echocardiography. It presents a detailed analysis of advancements in artificial intelligence
(AI) applied to echocardiography and the evolution of robotic ultrasound systems, aiming to
introduce a discussion on semantic-to-motion mapping. By synthesizing recent progress and
outlining future directions, we can correctly recognize the current maturity level of artificial
intelligence development in the field of ultrasound examination and prepare well for the
subsequent work.
Keywords: Transthoracic echocardiography; Deep learning for medical imaging; Robotic
ultrasound scanning; Semantic-guided control; Multimodal intelligence

D. Brahmeswara Rao, K. Venkata Rao, A. Gopala Krishna,


A hybrid approach to multi response optimization of micro milling process parameters using
Taguchi method based graph theory and matrix approach (GTMA) and utility concept,
Measurement,
Volume 120,
2018,
Pages 43-51,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Nowadays, it is required to produce micro products with high dimensional accuracy
to use them in different applications like aerospace, electronic and optics. The objective of
this study is to investigate influence of process parameters on surface roughness (Ra and
Rq), tool wear and cutter vibration in micro milling of AISI304 stainless steel. According to
orthogonal array of L27, twenty-seven experiments were conducted on the workpiece with
carbide end mill cutter at different levels of spindle speeds, feeds and depth of cuts. A hybrid
approach of Taguchi method based graph theory and matrix approach (GTMA) and utility
concept was used for multi response optimization of process parameters. The GTMA was
used to calculate weightage of four responses as per user’s opinion or preference. The utility
concept was used to calculate utility value of four responses using preference scale. Mean
utility values of responses are analyzed with Taguchi method and analysis of variance. The
optimum process parameters for the minimization responses were found to be 6000 rpm of
spindle speed, 95 µm/teeth of feed and 50 µm of depth of cut. The predicted responses at
optimal process parameters are Ra = 0.534 µm, VB = 70.861 µm, Amp = 54.395 µm and
Rq = 0.894 µm. A confirmation test was also carried out to verify the results.
Keywords: Micro milling; GTMA; Utility concept; Taguchi method; Tool vibration;
Optimization

Taofeng Gu, Yang Liang, Yangtian Yan, Wenjun Jiang, Haiyan Yue, Gang Hu, Jize Zhang,
Towards high-fidelity urban wind profiles for the built environment: a neural field to fuse
multi-source observational data in Guangzhou, China,
Building and Environment,
Volume 288,
2026,
114009,
ISSN 0360-1323,
[Link]
([Link]
Abstract: Accurate urban wind analysis is critically hampered by sparse and heterogeneous
observational data. This work presents a solution through NF-MW (stands for Neural Field
for Multi-source Winds), a model that fuses data from Doppler LiDAR and wind profiler radar
into a continuous high-resolution wind field. By learning a direct mapping from spatio-
temporal coordinates to wind values, NF-MW can reconstruct wind speed and direction at
any arbitrary height and time. The framework uniquely handles the 360∘ periodicity of wind
direction and uses Fourier-enriched features to capture high-frequency gusts and turbulence
often missed by other models. In a Guangzhou case study, NF-MW achieved a Mean
Absolute Error of 0.55 m/s for wind speed and 8.95∘ for wind direction, demonstrating
superior accuracy over traditional methods. This approach provides the building and
environment community with a robust method to generate the realistic dynamic wind data
essential for applications ranging from pedestrian comfort assessments to urban air quality
modeling.
Keywords: Urban wind environment; Neural fields; Doppler LiDAR; Wind profiler radar; Data
fusion; Deep learning

Jun Teng, Yuchao Wang, Yong Xia, Weihua Hu,


Robust discriminative correlation-based full-field motion estimation of large-scale structures
using a single video camera,
Engineering Structures,
Volume 334,
2025,
120224,
ISSN 0141-0296,
[Link]
([Link]
Abstract: Full-field motion, which reflects the health state of large-scale structures, is
difficult to capture using traditional structural health monitoring (SHM) systems due to
limited measurement points. Moreover, numerous structures lack SHM systems capable of
accurately monitoring motions. Readily available videos, with their numerous pixels acting as
an array of sensors, are promising in estimating full-field motion with a high resolution.
However, existing vision-based motion estimation methods fail to achieve good accuracy and
robustness. Accordingly, a novel motion estimation method is proposed to measure the full-
field motion of large-scale structures with a single video camera. This approach adopts
robust discriminative correlation to detect targets of various shapes by adaptively filtering
the textures of the target and the background. The accuracy of the estimated motion
reaches subpixel levels by introducing continuous convolution, which transforms discrete
pixels into a continuous function. A factorized convolution operator and a Gaussian mixture
model are used to compact the number of model parameters and training samples. This
approach estimates the accurate displacement and high-resolution mode shape in an
experimental study. Moreover, the displacement of the antenna on the high-rise Saige
Building is estimated with a portable camera, yielding an error of 1.75 % compared with the
laser Doppler vibrometer result. The high-resolution mode shape of the antenna is further
visualized with a modal assurance criterion (MAC) value of over 0.98 compared with the
simulated result. The full-field vortex-excited resonance of the long-span Humen Bridge is
estimated using a surveillance video, yielding accurate mode shapes with an MAC value over
0.97 compared with the accelerometer result.
Keywords: Full-field motion; Discriminative correlation; Large-scale structure; Single camera

Lixing Shi, Xueling Liang, Wenchao Chen, Yaoqiang Liu, Tong Ding, Kun Qin, Bo Chen,
Hongwei Liu,
Masked variational transformer for complex clutter modeling and target detection,
Signal Processing,
Volume 239,
2026,
110236,
ISSN 0165-1684,
[Link]
([Link]
Abstract: Weak target detection commonly encounters intense clutter interference, which
overshadows weak signals and complicates the task. Taking advantage of the powerful data
mining capability of neural networks, more and more deep learning-based methods are
applied to radar target detection. Among the approaches, those founded upon unsupervised
learning methodologies exhibit remarkable merit because they dispense with the
requirement for target samples within the training step, making them highly applicable in
practical target detecting scenarios. However, existing methods suffer from limitations in
leveraging the range-Doppler (R-D) two-dimensional correlation and finely modeling in
multiple clutter scenarios. In this paper, an unsupervised Transformer-based detector (TrDet)
is proposed to break through the boundary of modeling capability. First, with the designed
two-dimensional position embedding (2-DPE) and global query embedding (GQE)
techniques, an unsupervised training strategy for R-D spectrum based on Transformer
framework is utilized to achieve refined clutter modeling. Then, radar target detection is
formulated as an out-of-distribution (OOD) detection task to mitigate clutter interference.
Moreover, the masked variational Transformer-based detector (MVTrDet) is further
proposed to prevent target information leakage when the target is in close proximity to the
clutter in Doppler domain. Compared with several relative algorithms, our proposed
methods are better suited for radar target detection in complex clutter environments. The
experimental results derived from both measured data and simulated data verify the
effectiveness of our proposed methods.
Keywords: Radar target detection; Clutter modeling; Range-Doppler (R-D) spectrum;
Unsupervised learning; Out-of-distribution detection; Transformer

Changlong Wang, Jiawei Jiang, Chong Han, Hengyi Ren, Lijuan Sun, Jian Guo,
Through-Wall Multihuman Activity Recognition Based on MIMO Radar,
Computers, Materials and Continua,
Volume 83, Issue 3,
2025,
Pages 4537-4550,
ISSN 1546-2218,
[Link]
([Link]
Abstract: Existing through-wall human activity recognition methods often rely on Doppler
information or reflective signal characteristics of the human body. However, static
individuals, lacking prominent motion features, do not generate Doppler information.
Moreover, radar signals experience significant attenuation due to absorption and scattering
effects as they penetrate walls, limiting recognition performance. To address these
challenges, this study proposes a novel through-wall human activity recognition method
based on MIMO radar. Utilizing a MIMO radar operating at 1–2 GHz, we capture activity data
of individuals through walls and process it into range-angle maps to represent activity
features. To tackle the issue of minimal variation in reflection areas caused by static
individuals, a multi-scale activity feature extraction module is designed, capable of extracting
effective features from radar signals across multiple scales. Simultaneously, a temporal
attention mechanism is employed to extract keyframe information from sequential signals,
focusing on critical moments of activity. Furthermore, this study introduces an activity
recognition network based on a Deformable Transformer, which efficiently extracts both
global and local features from radar signals, delivering precise human posture and activity
sequences. In experimental scenarios involving 24 cm-thick brick walls, the proposed
method achieves an impressive 97.1% accuracy in activity recognition classification.
Keywords: MIMO radar; human activity; Transformer; through-wall

Hana Sebia, Thomas Guyet, Mickaël Pereira, Marco Valdebenito, Hugues Berry, Benjamin
Vidal,
Vascular segmentation of functional ultrasound images using deep learning,
Computers in Biology and Medicine,
Volume 194,
2025,
110377,
ISSN 0010-4825,
[Link]
([Link]
Abstract: Segmentation of medical images is a fundamental task with numerous
applications. While MRI, CT, and PET modalities have significantly benefited from deep
learning segmentation techniques, more recent modalities, like functional ultrasound (fUS),
have seen limited progress. fUS is a non invasive imaging method that measures changes in
cerebral blood volume (CBV) with high spatio-temporal resolution. However, distinguishing
arterioles from venules in fUS is challenging due to opposing blood flow directions within
the same pixel. Ultrasound localization microscopy (ULM) can enhance resolution by tracking
microbubble contrast agents but is invasive, and lacks dynamic CBV quantification. In this
paper, we introduce the first deep learning-based application for fUS image segmentation,
capable of differentiating signals based on vertical flow direction (upward vs. downward),
using ULM-based automatic annotation, and enabling dynamic CBV quantification. In the
cortical vasculature, this distinction in flow direction provides a proxy for differentiating
arteries from veins. We evaluate various UNet architectures on fUS images of rat brains,
achieving competitive segmentation performance, with 90% accuracy, a 71% F1 score, and
an IoU of 0.59, using only 100 temporal frames from a fUS stack. These results are
comparable to those from tubular structure segmentation in other imaging modalities.
Additionally, models trained on resting-state data generalize well to images captured during
visual stimulation, highlighting robustness. Although it does not reach the full granularity of
ULM, the proposed method provides a practical, non-invasive and cost-effective solution for
inferring flow direction—particularly valuable in scenarios where ULM is not available or
feasible. Our pipeline shows high linear correlation coefficients between signals from
predicted and actual compartments, showcasing its ability to accurately capture blood flow
dynamics.
Keywords: Functional ultrasound; Segmentation; Ultrafast ultrasound localization
microscopy; Preclinical; Medical images; Neuroscience
Pengfei Yan, Wushuang Gong, Minglei Li, Jiusi Zhang, Xiang Li, Yuchen Jiang, Hao Luo, Hang
Zhou,
TDF-Net: Trusted Dynamic Feature Fusion Network for breast cancer diagnosis using
incomplete multimodal ultrasound,
Information Fusion,
Volume 112,
2024,
102592,
ISSN 1566-2535,
[Link]
([Link]
Abstract: Ultrasound is a critical imaging technique for diagnosing breast cancer. However,
the multimodal breast ultrasound diagnostic process is time-consuming and labor-intensive,
heavily dependent on the physician’s extensive expertise. Therefore, developing a computer-
aided diagnosis system for breast cancer is essential. Existing diagnostic systems fail to
consider the varying impacts of different ultrasound modalities on diagnostic results and
struggle to address the issue of missing modalities in clinical practice. Consequently, this
paper proposes the Trusted Dynamic Feature Fusion Network (TDF-Net) for diagnosing
breast cancer using incomplete multimodal ultrasound data. Initially, this method introduces
a dual-branch feature extraction module to capture modality-specific information.
Meanwhile, a contrastive clustering loss is designed to enforce the consistency constraint,
ensuring the coherence of different modal features within the semantic space for each
sample. Additionally, an invertible neural network-based recovery method is suggested to
establish mappings between different modalities, enabling the recovery of missing
modalities. Finally, a trusted dynamic feature fusion module based on the Dirichlet
distribution is proposed to quantify each modality’s contribution to the diagnostic result by
considering uncertainty, thereby achieving the dynamic fusion of each modality’s features
across different samples. The proposed method is validated on an established multimodal
breast ultrasound dataset, demonstrating superior diagnostic performance compared to
existing methods, with an average AUC of 98.31%, a 95% confidence interval of [96.98%,
99.64%], and a p-value < 0.05. A pilot study is planned to assess the effectiveness and
usability of TDF-Net in clinical settings.
Keywords: Breast cancer; Transformer; Ultrasound; Multimodality; Missing modality

Mao Li, Sen Wang, Tao Liu, Xiaoqin Liu, Chang Liu,
Rotating box multi-objective visual tracking algorithm for vibration displacement
measurement of large-span flexible bridges,
Mechanical Systems and Signal Processing,
Volume 200,
2023,
110595,
ISSN 0888-3270,
[Link]
([Link]
Abstract: Visual displacement measurement methods for flexible structural bodies like large-
span bridges has gained wide popularity in recent years, but practical applications still have
some limitations. For instance, when acquiring images of large-span flexible bridges at a
distance, the slight angular tilt of the detection target due to irregular vibrations can cause
extremely serious misfit errors in the displacement curves returned by the vision
measurement algorithm. To improve the reliability of vibration displacement measurement
of flexible structural bodies, this paper takes the bridge subjected to external excitation in
the acquired image sequence as the object of vibration displacement measurement and
uses a designed high-precision displacement measurement algorithm for a single-stage
rotating target tracking anchor-free box to track the vibration displacement of the target in
the flexible structural body. We first extract multi-scale feature information of bridge model
image sequences using the improved YOLOv5-s backbone network and combine the
Transformer self-attention mechanism with PANet to perform a top-down and bottom-up bi-
directional fusion of target feature maps at three different scales to achieve semantic feature
fusion of shallow and deep information. Second, the improved Efficient Decoupled Head
performs the detection of rotating target centroid offset and bounding box size. Finally, the
detected results are passed into the multi-objective tracking algorithm ByteTrack, which
strengthens the spatio-temporal correlation between frames and obtains a better-fitting
vibration displacement curve. The validation and comparison of traditional visual
measurement methods and deep learning measurement methods on cable-stayed bridge
models, small arch bridges, and large span bridges show that the vibration displacement
trajectories regressed by the algorithm in this paper have the best fit with the actual
vibration displacement trajectories, which also verifies that the algorithm in this paper has
good potential for engineering applications and implementation space in the field of
condition monitoring of flexible structural bodies.
Keywords: Flexible structure; Visual vibration measurement; Tilted targets; Rotating box;
Multi-target visual tracking; Deep convolutional neural network

Pieter G.G. Muyshondt, Peter Aerts, Joris J.J. Dirckx,


Acoustic input impedance of the avian inner ear measured in ostrich (Struthio camelus),
Hearing Research,
Volume 339,
2016,
Pages 175-183,
ISSN 0378-5955,
[Link]
([Link]
Abstract: In both mammals and birds, the mechanical behavior of the middle ear structures
is affected by the mechanical impedance of the inner ear. In this study, the aim was to
quantify the acoustic impedance of the avian inner ear in the ostrich, which allows us to
determine the effect on columellar vibrations and middle ear power flow in future studies.
To determine the inner ear impedance, vibrations of the columella were measured for both
the quasi-static and acoustic stimulus frequencies. In the frequency range of 0.3–4 kHz, we
used electromagnetic stimulation of the ossicle and a laser Doppler vibrometer to measure
the vibration response. At low frequencies, harmonic displacements were imposed on the
columella using piezo stimulation and the resulting force response was measured with a
force sensor. From these measurement data, the acoustic impedance of the inner ear could
be determined. A simple RLC model in series of the impedance measurements resulted in a
stiffness reactance of KIE = 0.20·1012 Pa/m³, an inertial impedance of
MIE = 0.652·106 Pa s2/m³, and a resistance of RIE = 1.57·109 Pa s/m. We found that values
of the inner ear impedance in the ostrich are one to two orders in magnitude smaller than
what is found in mammal ears.
Keywords: Avian ear; Acoustic impedance; Laser Doppler vibrometry; Electromagnetic
stimulation; Piezo stimulation

K. Venkata Rao, Bachina Harish Babu, V. Umasai Vara Prasad,


A study on effect of dead metal zone on tool vibration, cutting and thrust forces in micro
milling of Inconel 718,
Journal of Alloys and Compounds,
Volume 793,
2019,
Pages 343-351,
ISSN 0925-8388,
[Link]
([Link]
Abstract: Inconel 718 is a nickel-based super alloy, its characteristics pose more complex in
micro machining. Aim of the present study is to investigate effect of process parameters and
tool parameters on the dead metal zone (DMZ) geometry and length of shear plane and
further on cutting force, thrust force and amplitude of cutter vibration. In the present study,
mechanistic models were integrated with finite element method (FEM) simulation to predict
cutting force and thrust force in micro milling of Inconel 718. Numerical simulation was
carried out to estimate size of DMZ and its angles for all the experiments. Simulation results
were used in the mechanistic models to estimated cutting force and thrust force. Series of
experiments were carried out on the Inconel 718 at different levels of process and tool
parameters and cutting forces were measured. Comparison between the predicted values of
cutting forces and thrust forces with experimental results shown a good agreement. Effect of
process and tool parameters on size of DMZ and length of shear plane was studied. Further,
effect of DMZ and length of shear plane on cutting force, thrust force and amplitude cutter
vibration was studied. Feed per tooth and cutter nose radius have significant effect on length
of shear plane and length of DMZ side respectively. The parameters were optimized as
4500 rpm of cutting speed, 10 μm per tooth of feed, 0.12 mm of cutter nose radius and
6.8485° of rake angle for minimum side length of DMZ and length of shear plane in order to
reduce cutting force, thrust force and amplitude of cutter vibration.
Keywords: Inconel 718; Dead metal zone; Shear zone; Cutting force; FEM; Tool vibration

Alan S. Morris, Reza Langari,


Chapter 13 - Sensor technologies,
Editor(s): Alan S. Morris, Reza Langari,
Measurement and Instrumentation (Third Edition),
Academic Press,
2021,
Pages 381-411,
ISBN 9780128171417,
[Link]
([Link]
Abstract: This chapter explains the different physical principles and technologies that are
used in various types of measurement sensors. As well as summarizing the background
theory, the different measurement sensors that use each principle are reviewed. The recent
development of micro-electro-mechanical-systems (MEMS) devices and nano-electro-
mechanical-system (NEMS) devices is also discussed in some detail. The chapter starts with a
discussion on capacitive sensors and then moves on to resistive sensors. Following this, the
various magnetic sensors in the form of variable-inductance, variable-reluctance, eddy-
current, and hall-effect devices are reviewed. This is followed by three sections that cover
piezoelectric transducers, strain gauges, and piezoresistive sensors, respectively. The next
subject of discussion is optical sensors, which can exist in both open air-path and fiber-optic
forms. In the case of the latter, both intrinsic and extrinsic fiber-optic sensors are discussed
and the fundamental difference between the two types is explained. Following this,
ultrasonic sensors are considered in both transit-time and Doppler-shift forms. Then, after a
brief look at nuclear sensors, the chapter finishes with a discussion on microsensors (MEMS
and NEMS devices).
Keywords: Capacitive sensor; Hall-effect sensor; Magnetic sensor; Microsensor; Nuclear
sensor; Optical sensor; Piezoelectric sensor; Piezoresistive sensor; Resistive sensor; Sensor
technologies; Strain gauge; Ultrasonic sensor

Brian T. Morris, Michael D. Seidman, Eric M. Kraus,


Ossicular Chain Reconstruction Using Glass Ionomer Cement: Cement Alone and Cement
with a Titanium Prosthesis Having a Helical Coil,
Otolaryngologic Clinics of North America,
2025,
,
ISSN 0030-6665,
[Link]
([Link]
Keywords: Cement ossiculoplasty; Endoskeletal ossicular chain reconstruction; Glass
ionomer cement; Ossicular chain reconstruction; Middle ear reconstruction; Middle ear
prostheses; Total & partial ossicular replacement prostheses; PORPs; TORPs; Long-term
hearing results

O.M. Hemeda, K.R. Mahmoud, T. Sharshar, M. Elsheshtawy, Mahmoud A. Hamad,


ESR, thermoelectrical and positron annihilation Doppler broadening studies of CuZnFe2O4-
BaTiO3 composite,
Journal of Magnetism and Magnetic Materials,
Volume 429,
2017,
Pages 124-128,
ISSN 0304-8853,
[Link]
([Link]
Abstract: Composite materials of Cu0.6Zn0.4Fe2O4 (CZF) and barium titanate (BT) with
different concentrations were prepared by high energy ball milling method. The composite
samples of CZF and BT were studied using Infrared, ESR and positron annihilation Doppler
broadening (PADB) spectroscopy techniques as well as thermo-electric power
measurements. The results confirm formation of the composite, and presence of two
ferrimagnetic and ferroelectric phases, simultaneously. In addition, Fe–O bond for both
tetrahedral and octahedral sites, population and distribution of cations at A and B sites are
varied with BT content. The values of resonance field, line width of ESR spectrum and charge
carrier concentration increase by increasing BT content. The value of the g factor for our
samples with low BT content is greater than g-factor value of the isolated free electron. On
the contrary, the g-factor values for samples with high BT content are smaller than the free
isolated electron. PADB line-shape S-parameter suggests that there are increases of the
density of the delocalized electrons, defect size and concentration caused by highly adding
BT phase. In addition, PADB results confirm the homogeneity of composite phases and same
structure of defects in BT-CZF composite samples.
Keywords: Ferrite; Composite; Barium titanate; ESR; IR spectroscopy; Doppler broadening
spectroscopy; Thermoelectric power

Ab Waheed Lone, Ahmet Elbir, Nizamettin Aydin,


A comprehensive review on cerebral emboli detection algorithms,
WFUMB Ultrasound Open,
Volume 2, Issue 1,
2024,
100030,
ISSN 2949-6683,
[Link]
([Link]
Abstract: The increasing availability of biomedical data has attracted the interest of many
researchers to understand and perform analysis on extracted patterns from data. Stroke is
considered as one of the main causes of deaths worldwide. A considerable amount of work
has been performed related to the cause of stroke and other physiological effects. Cerebral
emboli is considered as one of the main sources of stroke. Algorithms from one of the
traditional subjects called signal processing have been used in cerebral emboli detection and
lot of researchers have performed emboli detection and classification using Fourier
transform based algorithms and different filtering approaches. In this paper, we discuss the
physics of Doppler ultrasound and perform review of cerebral emboli detection algorithms
and some animal models used in understanding the behaviour, size, and composition of
emboli development. Ranging from Fourier transform, wavelet transform based emboli
detection to neural network architectures trained with Doppler signal spectra, we
performed comprehensive review of signal processing based cerebral emboli detection
works and provide some basic understanding of related terms in emboli detection. With
natural arterial structural relation between humans and some animals, we highlight some of
the animal models used for understanding the nature of emboli development process.
Keywords: Cerebral embolism; Fourier transform; Transcranial Doppler ultrasound; Wavelet
transform; Correlation; Wigner distribution

Haiqiao Wang, Hong Wu, Zhuoyuan Wang, Peiyan Yue, Dong Ni, Pheng-Ann Heng, Yi Wang,
A Narrative Review of Image Processing Techniques Related to Prostate Ultrasound,
Ultrasound in Medicine & Biology,
Volume 51, Issue 2,
2025,
Pages 189-209,
ISSN 0301-5629,
[Link]
([Link]
Abstract: Prostate cancer (PCa) poses a significant threat to men's health, with early
diagnosis being crucial for improving prognosis and reducing mortality rates. Transrectal
ultrasound (TRUS) plays a vital role in the diagnosis and image-guided intervention of PCa.
To facilitate physicians with more accurate and efficient computer-assisted diagnosis and
interventions, many image processing algorithms in TRUS have been proposed and achieved
state-of-the-art performance in several tasks, including prostate gland segmentation,
prostate image registration, PCa classification and detection and interventional needle
detection. The rapid development of these algorithms over the past 2 decades necessitates
a comprehensive summary. As a consequence, this survey provides a narrative review of this
field, outlining the evolution of image processing methods in the context of TRUS image
analysis and meanwhile highlighting their relevant contributions. Furthermore, this survey
discusses current challenges and suggests future research directions to possibly advance this
field further.
Keywords: Transrectal ultrasound; Prostate cancer; Medical image processing; Deep
learning; Machine learning; Medical image segmentation; Medical image registration;
Classification; Computer-assisted detection; Computer-assisted diagnosis

Table of Content,
Chinese Journal of Aeronautics,
Volume 38, Issue 8,
2025,
103685,
ISSN 1000-9361,
[Link]
([Link]

Xiang Li, PengTao Guo, Yuan Ding, Zhiwei Chen, Xu Wang, Qibao Lv,
A generalized electromechanical coupled model of standing-wave linear ultrasonic motors
and its nonlinear version,
Mechanical Systems and Signal Processing,
Volume 186,
2023,
109870,
ISSN 0888-3270,
[Link]
([Link]
Abstract: For the systematization of research on standing-wave linear ultrasonic motors
(SWLUMs), this work develops a generalized electromechanical coupled model for
characterizing SWLUMs. The proposed model focuses on dealing with modeling the
generalized two-stage energy conversion in SWLUMs. The first-stage energy conversion is
modeled by a four-terminal equivalent circuit model, which with a phase shifter and a
couple of electromechanical transformers based on the electromechanical analogy method.
The second-stage energy conversion is modeled by a physics-based, friction-driven system,
involving contact nonlinearities between the stator and the mover. Furthermore, the
effectiveness of this model and its further extension considering nonlinear vibrations of
piezoelectric transducer (stator) are exemplified and discussed by a classical SWLUM with V-
configuration stator, thereby indicating that the presented generalized model is valuable and
pragmatic in simulating and characterizing SWLUMs both in electrical and mechanical
domains.
Keywords: Linear ultrasonic motor; Piezoelectric transducer; Generalized model; Contact
nonlinearities; Nonlinear vibration

Mohammad Hossein Shirazi, Sira Yongchareon, Anuradha Singh, Jing Ma,


A survey on machine learning approaches for vital sign monitoring using radar,
Measurement,
Volume 253, Part D,
2025,
117707,
ISSN 0263-2241,
[Link]
([Link]
Abstract: The integration of machine learning methodologies with radar-based vital sign
monitoring represents a significant advancement in non-contact healthcare surveillance
systems. This systematic literature review synthesizes and critically analyzes research from
2020 to 2025, addressing substantive theoretical and methodological gaps in extant
literature. Our comprehensive taxonomic classification of machine learning paradigms
employed in this domain elucidates the progressive refinement from conventional
algorithmic approaches to sophisticated deep learning architectures, with particular
emphasis on hybrid neural network configurations optimized for physiological signal
extraction in non-stationary environments. Methodologically, this survey contributes a
rigorous evaluation framework comprising standardized assessment protocols, quantifiable
performance metrics, and cross-validation methodologies—elements conspicuously absent
in previous reviews. Empirical analysis demonstrates substantial correlations between
dataset demographic characteristics and algorithmic generalizability, with heterogeneous
participant cohorts yielding markedly enhanced performance across cardiac, respiratory, and
hemodynamic parameter estimation tasks. The review delineates four distinct
developmental phases in the field’s chronological evolution and provides analytical insight
into persistent technical challenges: motion artifact compensation, multi-subject
disambiguation, and the translation of laboratory efficacy to clinical utility. This
comprehensive examination of computational approaches for radar-based vital sign
monitoring establishes a theoretical foundation and methodological framework to guide
future research towards physiologically robust and clinically viable implementations.
Keywords: Non-intrusive vital sign monitoring; Machine learning; Radar

Bofeng Liang, Li-Yun Fu, Mian Lin, Tobias Müller, Wubing Deng, Tongcheng Han,
Seismic efficiency: From hydraulic fracturing-acoustic emission laboratory experiments of
shale based on energy budget,
Geoenergy Science and Engineering,
Volume 252,
2025,
213917,
ISSN 2949-8910,
[Link]
([Link]
Abstract: Fluid injection-triggered earthquakes have been documented worldwide and quite
a number of events have significant moment magnitudes (Mw ≥ 3). Seismic efficiency (η),
defined as the ratio of injection volume to net seismic moment release in hydraulic
fracturing operations, is a crucial parameter to evaluate seismic hazard. However, a
quantitative assessment of seismic and non-seismic (aseismic) energy release is a key aspect
of understanding the intrinsic properties of cracking rocks. Therefore, we develop a novel η
model based on the hydraulic-fracturing-propagation energy budget and performed
laboratory experiments on hydraulic fracturing in shale by injecting distilled water at
different rates under pseudo-triaxial stress conditions with simultaneous monitoring of
acoustic emission (AE). We estimate AE energy accurately with absolute value correction of
sensors using a laser Doppler vibrometer, and the dissipation of the potential energy using
displacement and pressure sensors. The results show that the proposed η model can
evaluate the induced seismic characteristics effectively compared with field data and the
injection rate controls the change of η to some extent. Moreover, there is a log-linear
relationship between seismic efficiency and injection efficiency (ratio of AE energy and
injection energy), which may provide an experiential method for evaluating seismicity during
the early phase of hydraulic fracturing.

K. Van Der Molen, H.R.E. Van Maanen,


Laser-Doppler measurements of the turbulent flow in stirred vessels to establish scaling
rules,
Chemical Engineering Science,
Volume 33, Issue 9,
1978,
Pages 1161-1168,
ISSN 0009-2509,
[Link]
([Link]
Abstract: A laser-Doppler velocimeter, equipped with a frequency shift so as to eliminate
directional ambiguity, has been used to measure the turbulent flow in stirred vessels with
diameters of 0.12, 0.29 and 0.90 m of the same geometry. The vessels contained water and
measurements were done in the impeller stream region. Scaling rules have been derived for
average velocity, the periodic component, turbulent intensities and turbulence power
spectra. It appears that close to the impeller the flow is dominated by the periodically
fluctuating flow of the trailing vortices behind the impeller blades. The normalized mean
velocity in the trailing vortices, and therefore the turbulence intensity close to the impeller,
is very sensitive to impeller geometry and shows a slight increase with size of the vessel. In
the greater part of the impeller stream region the power spectra have a section with a −52
slope on a log-log scale and consequently the energy of the small eddies decreases with
increasing scale. At the vessel wall the vortices have decayed completely to random
turbulence and the spectrum shows a − 53 slope.
Alexandr Neftissov, Lalita Kirichenko, Ilyas Kazambayev, Aliya Aubakirova,
Development of an Energy-Efficient Security Control Installation Based on Fiber-Optic
Sensors for The Water Resources Monitoring System,
Procedia Computer Science,
Volume 272,
2025,
Pages 92-99,
ISSN 1877-0509,
[Link]
([Link]
Abstract: This article presents the architecture of an autonomous security system designed
to protect automated monitoring equipment for water resources located near open water
bodies. The basis of the proposed solution is a sensor platform based on the use of a single-
mode optical fiber as a sensor, capable of registering the presence of people in the
protected area. The system is implemented on a Raspberry Pi single-board microcomputer.
The key element of the architecture is a multi-level structure of interaction between
components: the optical fiber registers external influences, changes in the optical signal are
recorded by a photomatrix, then the information is analysed using built-in algorithms, and if
an anomaly is detected, video surveillance and alarm notification modules are activated. The
architecture is complemented by an IoT module that provides data transmission and
centralized control in a distributed system. The developed software is an integral part of the
system and provides real-time control of all components of the complex. The proposed
system operates in a completely autonomous mode thanks to the use of a solar panel as the
main source of power supply, which eliminates the need to connect to a centralized power
grid and the constant presence of maintenance personnel.
Keywords: security system; optical fiber; optical sensors; water monitoring systems

Qiming Zhang, Yang Li, Zhi Zhang, Shibo Yin, Lin Ma,
Marine target detection for PPI images based on YOLO-SWFormer,
Alexandria Engineering Journal,
Volume 82,
2023,
Pages 396-403,
ISSN 1110-0168,
[Link]
([Link]
Abstract: For the task of detecting marine targets, numerous machine-learning methods
have been suggested, which can achieve comparable accuracy. Despite the successful
application of deep learning in the field of marine target detection in recent years, existing
detection methods face challenges due to the significant interference of sea clutter. As
artificial intelligence technology advances, the Swin Transformer can serve as an effective
backbone for extracting discriminative features. However, it has not yet been utilized for
target detection, and the combination of Swin Transformer and YOLO architecture has not
been applied to similar missions. In light of this, we propose a novel method, YOLO-
SWFormer, which combines the Swin Transformer and YOLO framework for target detection.
Our method can extract discriminative features from plan-position indicator (PPI) images
despite the interference of sea clutter, thereby reducing computational complexity and
enhancing target detection accuracy. Experimental results on a Sea Clutter Database
demonstrates that our method surpasses existing methods in terms of accuracy, indicating
its potential as a promising solution for marine target detection tasks.
Keywords: PPI images; Target detection; Sea clutter suppression; Swin transformer; YOLO

Zhi Liu, Dexiang Le, Tianyu Zhang, Qingrong Lai, Jiansheng Zhang, Bin Li, Yunfeng Song, Nan
Chen,
Detection of apple moldy core disease by fusing vibration and Vis/NIR spectroscopy data
with dual-input MLP-Transformer,
Journal of Food Engineering,
Volume 382,
2024,
112219,
ISSN 0260-8774,
[Link]
([Link]
Abstract: Moldy core is a highly contagious internal disease of apples, and even a small
number of diseased apples can trigger large-scale infections during the storage stage. In this
study, a combined acoustic vibration and Vis/NIR spectroscopy method for moldy core apple
identification was proposed to improve the accuracy of moldy core identification. The
vibration signals and Vis/NIR spectroscopy of apples were collected using a self-designed
micro-LDV detection device and Vis/NIR spectroscopy online detection device respectively
for constructing multiple moldy core apple classification models. The results showed that
the classification model combining vibration spectrum and Vis/NIR spectral data had
significant advantages in apple moldy core identification accuracy compared to the
classification model using a single vibration spectrum or Vis/NIR spectral data. Ultimately,
dual-input MLP-Transformer (DMLPT) demonstrated the best recognition performance with
an overall classification accuracy of 99.31% for the model, with 100%, 97.56%, 100.00% and
100% accuracy for normal, mild moldy core disease, moderate moldy core disease and
severe moldy core disease apples, respectively. This study demonstrated the excellent
performance and great potential of acoustic vibration and visible/near-infrared spectral data
fusion for fruit internal quality detection.
Keywords: Laser Doppler vibrometer; Visible near-infrared spectroscopy; Dual-input MLP-
Transformer; Non-destructive detection; Moldy core apple

Bin Zhang, Xinru Ma, Xiaoping Ma, Xiaohong Jia, Xuejun Zhang,
The impacts of rainfall on MEMS Lidar SNR and detecting ability,
Infrared Physics & Technology,
Volume 150,
2025,
106009,
ISSN 1350-4495,
[Link]
([Link]
Abstract: This study investigates the impact of rainfall on performance degradation of Lidar
through a combination of experiments and analytical modeling. A Micro-electronic-
mechanical system (MEMS) Lidar is employed for single-point detection to analyze how
rainfall affects the range-dependent signal-to-noise ratio (SNR). Consequently, we establish
an ’exponential plus linear’ decay model of the reciprocal of target distance square.
Comparative outdoor and indoor trials are conducted to assess the impacts of raindrops on
Lidar range detectability and scanning angle. The results demonstrate the raindrop-induced
attenuation has a greater impact on the Lidar point cloud than on SNR. In addition, droplets
changing laser direction causes outlier points in the point cloud. The findings underscore the
importance of rain-resilient perception strategies for reliable Lidar performance in adverse
weather conditions.
Keywords: Lidar; MEMS; SNR; Raindrop; Point cloud

Yuhan Liu, Jinlin Ye, Zecheng He, Mingyue Wang, Changjun Wang, Jie Lang, Yidong Zhou, Wei
Zhang,
Deep learning assisted non-invasive lymph node burden evaluation and CDK4/6i
administration in luminal breast cancer,
iScience,
Volume 28, Issue 7,
2025,
112849,
ISSN 2589-0042,
[Link]
([Link]
Abstract: Summary
Precise lymph node evaluation is fundamental to optimize CDK4/6 inhibitor therapy in
luminal breast cancer, particularly given contemporary trends toward axillary surgery de-
escalation that may compromise traditional lymph node staging for recurrence risk
evaluation. The lymph node prediction network (LNPN) was developed as a multi-modal
model incorporating both clinicopathological parameters and ultrasonographic
characteristics for lymph node burden differentiation. In a multicenter cohort of 411
patients, LNPN demonstrated robust performance, achieving an AUC of 0.92 for binary
lymph node burden classification (N0 vs. N+) and 0.82 for ternary lymph node burden
classification (N0/N1–3/N ≥ 4). Notably, among patients undergoing sentinel lymph node
biopsy (SLNB) with confirmed 1–2 metastatic lymph nodes, LNPN predicted high-burden
metastases (N ≥ 4) with an AUC of 0.77. LNPN provided a non-invasive method to assess
lymph node metastasis and recurrence risk, potentially reducing unnecessary axillary lymph
node dissection (ALND), and facilitating decision-making regarding the intervention of
CDK4/6i in luminal breast cancer patients.
Keywords: Cancer; Machine learning

Roy W. Martin, David A. Gilbert, Fred E. Silverstein, Michele Deltenre, Guido Tytgat,
Rhealond K. Gange, John Myers,
An endoscopic Doppler probe for assessing intestinal vasculature,
Ultrasound in Medicine & Biology,
Volume 11, Issue 1,
1985,
Pages 61-69,
ISSN 0301-5629,
[Link]
([Link]
Abstract: Flexible fiberoptic endoscopes permit the physician to inspect the mucosal surface
of the upper gastrointestinal tract and colon. However, this visual inspection provides little
information about the underlying vascular supply to the intestinal wall. We tested the
hypothesis that a Doppler probe could be constructed small enough to pass through the
biopsy channel of a fiber endoscope and be used with it while performing endoscopy. The
purpose would be to determine the location of patent arteries or veins, determine the
magnitude and waveforms of the velocity in them, and estimate their contribution or
potential contribution to intestinal bleeding. For this purpose, a miniature catheter probe
(1.8 mm O.D. and 2 m in length) and an electronic range limited pulsed Doppler unit were
developed. This probe and unit were studied in a series of 13 dogs to determine efficacy of
detecting arterial and enous flow and to test the safety of the device. The duodenum was
surgically exposed and opened in the region of the common bile duct (CBD). Arteries and
veins surrounding the CBD were studied. Particular attention was directed to arterial
structures which clinically pose a risk of bleeding when performing endoscopic papillotomy,
a therapeutic technique in which the papilla of Vater is cut to release bile duct stones. The
results of the study revealed that the probe could indeed detect arterial and venous
structures accurately. There was no evidence that the probe produced any injury to the
common bile duct or pancreas by histological or serum amylase studies and the device was
determined safe and suitable for clinical evaluation.
Keywords: Ultrasound; Intestinal blood flow; Papillotomy; Varices; Blood velocity; Doppler;
Endoscope; Catheter probe

Zhaochun Ding, Xiang Li, Jiang Wu, Jinshuo Liu, Lipeng Wang, Yu Tian, Yanhu Zhang, Xuewen
Rong, Yibin Li,
External-pipe-climbing piezoelectric actuator with high climbing/towing capability and
untethered movement,
International Journal of Mechanical Sciences,
Volume 306,
2025,
110840,
ISSN 0020-7403,
[Link]
([Link]
Abstract: To accomplish high climbing/towing capability and untethered movement, a
miniature external-pipe-climbing piezoelectric actuator (MEPCPA) is developed by
integrating a pair of wing-shaped transducers driven by piezoelectric stack plates and an
onboard circuit. Here, the transducers provide the climbing and clamping functions with the
driving feet and the spring, respectively; these interestingly imitate the propelling and
hugging functions of the sloth’s lower and upper limbs. The micro controller, boost module,
and transistors arranged in the H-bridge shape form the minimum system of a lightweight
onboard circuit. To verify our proposal, first, by constructing a vibration model, the
transducer was designed to enhance the driving force without excessively increasing the
weight. Meanwhile, the friction coefficient was modified by considering the surface
roughness to predict the climbing/towing performance. Then, a prototype whose
mechanical part had the size of 52 × 35 × 72 mm3 and the weight of 20.5 g was fabricated
for performance assessment. In a tethered manner, the MEPCPA climbed up the glass tube
vertically, towed the maximal weight of 120 g (equal to 5.9 times the mechanical part’s
weight), and yielded the maximal speed of 103.8 mm/s. Installed with a 12-V 300-mAh
battery, the MEPCPA successfully climbed up the tube having the tilting angle of 45° with the
ground and it produced the maximal towing weight, the maximal climbing speed, and the
minimal stepwise displacement of 20 g, 18 mm/s, and 0.36 μm, respectively, at the tilting
angle of 30° To the best of our knowledge, this study is an initial report regarding
piezoelectric actuators climbable in an untethered manner, and provides fundamental
technique for designing miniature piezoelectric actuators potentially applicable to narrow
environments particularly lacking external power source.
Keywords: Piezoelectric driving; Miniature actuator; Bioinspired actuation; Climbing
piezoelectric actuator; Towing capability; Untethered movement

Jebin Francis, Prawin Angel Michael,


Investigation of micro/nano motors based on renewable energy sources,
Materials Today: Proceedings,
Volume 27, Part 1,
2020,
Pages 150-157,
ISSN 2214-7853,
[Link]
([Link]
Abstract: Nano/micro motors are machines specially designed to accomplish special
functions as they drive themselves with respect to particular stimuli. Design construction of
such synthetic motors experience heavy challenge due to their minute structural complexity.
However, design of micro/nanomotors get wider attraction in recent world because of their
minute structural construction. Many of them requires electrical or chemical propulsion for
their mechanical mobility. But, propulsion by renewable energy sources like solar, wind and
light/sound energy are become trending techniques in recent world. Such kind of renewable
energy based nano/micro motor propulsion does not require any physical attachment with
motor design also leads to reduction in waste products with easy control. This review
investigates the highlights of various studies dealing with renewable energy sources based
micro/nano motor design. Also, this review summarizes different propulsion mechanisms
along with application of micro/nano motors in different fields.
Keywords: Micro motors; Nano motors; Solar; Wind; Light and renewable energy sources

Liu Zhi, Chen Nan, Le Dexiang, Lai Qingrong, Li Bin, Wu Jian, Song Yunfeng, Liu Yande,
Acoustic vibration multi-domain images vision transformer (AVMDI-ViT) to the detection of
moldy apple core: Using a novel device based on micro-LDV and resonance speaker,
Postharvest Biology and Technology,
Volume 211,
2024,
112838,
ISSN 0925-5214,
[Link]
([Link]
Abstract: Moldy-core is a common internal disease in apples, and apples infected with this
disease cannot be directly identified according to their external characteristics. In this study,
a novel acoustic vibration device based on micro-LDV, resonance speaker and microphone
was employed to detect moldy-core in apples, the acoustic vibration signals of healthy
apples and apples with different degrees of moldy-core were converted into acoustic
vibration multi-domain images (AVMDI), which consisted of time-frequency images
generated through continuous wavelet transform (CWT), as well as time-domain and
frequency-domain images generated through Gramian Angular Field (GAF). Subsequently
the combination of AVMDI and Vision Transformer (ViT) was applied to the identification of
internal defects in fruits. The outcomes evince that the classification efficacy of the model,
amalgamating sound and vibration signals, surpasses that of models reliant on solitary
sound or vibration signals. The AVMDI-ViT model achieved an overall classification accuracy
of 97.96 %. Specifically, it achieved 100 % accuracy in identifying healthy apples, 94.74 %
accuracy in identifying mild moldy-core apples (≤ 7 %), 97.50 % accuracy in identifying
moderate moldy-core apples (> 7 % and ≤ 15 %), and 100 % accuracy in identifying severe
moldy-core apples (> 15 %). The proposed method demonstrates a high level of accuracy in
the identification of moldy-core apples, while also offering advantages in terms of simplicity,
speed, cost-effectiveness and non-contact detection.
Keywords: Moldy core; Apple; Micro-LDV; Resonance speaker; Acoustic vibration multi-
domain images; Vision Transformer

Mohammed R. Abdulwahab, Khaled A. Al-attab, Irfan Anjum Badruddin, Muhammad Nasir


Bashir, Joon Sang Lee,
Biofuels spray and combustion characteristics in a new micro gas turbine combustion
chamber design with internal exhaust recycling,
Case Studies in Thermal Engineering,
Volume 65,
2025,
105595,
ISSN 2214-157X,
[Link]
([Link]
Abstract: The characteristics of atomization and combustion of biodiesel and palm oil were
evaluated in this study. A new combustor design with internal exhaust gas recycling (iEGR)
and internal fuel pre-evaporation was investigated numerically and then verified
experimentally using micro gas turbine (MGT) test rig. CFD evaluation of hydrodynamics flow
of 8 iEGR mechanisms geometries showed that simple connection between the exhaust and
recycle tube resulted in 0 % gas recycling due to the pressure difference. Low recycling <1 %
can be obtained by adding gas guiding channels, while increasing mass recycling from 3 % to
8 % was achieved by adding annular tubes with careful control of differential pressure using
pressure relief holes. Experimental cold-fuel-flow spray atomization quality was investigated
using high-shutter-speed camera. Increasing palm oil flow from 60 ml/min to 120 ml/min
significantly increased the spray angle from 1.8° to 21° while average droplet diameter
reduced from 665 μm to 148 μm. Minimum CO emissions in the range of 132–135 ppm for
diesel and biodiesel were achieved due to their better atomization compared to palm oil
which resulted in slightly higher value of 207 ppm. The opposite effect was observed for NOx
emissions where it elevated at the higher combustion temperature, where all the fuels
showed comparable values in the range of 32–39 ppm. On the other hand, diesel suffered
from its higher TIT value that reached 800 °C, compared to 785 °C and 762 °C for palm oil
and biodiesel, respectively.
Keywords: Liquid biofuels; CFD; Micro gas turbine; Combustion; Spray atomization; Exhaust
gas recycling

Chenglin Yao, Jianfeng Ren, Ruibin Bai, Heshan Du, Jiang Liu, Xudong Jiang,
Progressively-orthogonally-mapped EfficientNet for action recognition on time-range-
Doppler signature,
Expert Systems with Applications,
Volume 255, Part D,
2024,
124824,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Although 2D radar signal representations, such as spectrograms and range-Doppler
maps have been widely used for target recognition, 3D time-range-Doppler (TRD) has been
less studied, partially because of the difficulties in extracting features from the TRD
representation, i.e., shallow 3D neural networks have limited discriminant power, but
repeatedly applying 3D convolutions will lead to an oversized 3D network. A hybrid 3D–2D
network architecture, Progressively-Orthogonally-Mapped EfficientNet (POMEN), is
proposed to address these challenges. More specifically, the proposed POMEN utilizes 3D
convolutions in the earlier stages to capture the information embedded in the sparse 3D
TRD representation, and to avoid the oversized feature map caused by excessively applying
3D convolutions, we propose to progressively map the 3D features into three sets of 2D
features corresponding to the range-time signature, range-Doppler map and time-Doppler
signature (spectrogram), respectively. Subsequently, 2D EfficientNet blocks were designed to
extract discriminant information from the three sets of 2D feature maps. This hybrid 3D–2D
network design effectively extracts features from the 3D TRD representation, thereby
avoiding oversized features from full-sized 3D networks and the information loss of 2D
networks on 2D representations. Finally, a homogeneous gated fusion network was designed
to fuse the three sets of 2D features. The proposed method was evaluated on the UGRS,
MIMOGR, and mmWRWD datasets. The experimental results for all datasets demonstrate
that the proposed POMEN significantly and consistently outperforms the state-of-the-art
models in both 2D and 3D representations.
Keywords: Radar activity recognition; Progressively-orthogonally-mapped EfficientNet;
Homogeneous gated fusion; Time-range-Doppler representation

Yongchao Zhu, Qiuling Lu, Maorong Ge, Xiaochuan Qu, Tingye Tao, Kegen Yu, Shuiping Li,
Attention enhanced ResNet for ocean surface wind speed retrieval using CYGNSS
observables,
Advances in Space Research,
2025,
,
ISSN 0273-1177,
[Link]
([Link]
Abstract: Global Navigation Satellite System Reflectometry (GNSS-R) has emerged as a
pivotal technique for ocean surface wind speed retrieval; however, establishing robust multi-
parameter retrieval models remains challenging due to the nonlinear relationships between
GNSS-R observables and geophysical variables. An Attention-enhanced Residual Network
(Att-ResNet) is proposed to address this challenge, leveraging Cyclone Global Navigation
Satellite System (CYGNSS) bistatic radar data for wind speed estimation. The CYGNSS
datasets were processed to extract multi-parameter observables, including Delay-Doppler
Maps (DDMs), normalized bistatic radar cross-section (NBRCS), and incidence angle, which
served as inputs for training wind speed retrieval models using diverse backbone
architectures (e.g., ResNet and AlexNet). Ablation experiments employing the Att-ResNet
framework were systematically conducted, with ERA5 (European Centre for Medium-Range
Weather Forecasts Reanalysis 5) and CCMP (Cross-Calibrated Multi-Platform) wind products
providing benchmark validation. Comparative analysis revealed that the Att-ResNet-retrieved
wind speeds exhibited strong spatiotemporal consistency with ERA5 and CCMP data.
Quantitative evaluations showed root mean square errors (RMSEs) of 1.379 m/s (ERA5) and
1.390 m/s (CCMP), with minimal biases (−0.069 m/s and −0.014 m/s, respectively) and
unbiased RMSEs (ubRMSEs) of 1.377 m/s and 1.390 m/s. The study demonstrates that the
Att-ResNet architecture, through its attention-driven feature selection and residual learning
mechanisms, significantly enhances spaceborne GNSS-R wind retrieval accuracy. This
artificial intelligence-driven framework establishes a new paradigm for high-resolution
spatiotemporal ocean surface wind monitoring, demonstrating the transformative potential
of deep learning in advancing GNSS-R applications.
Keywords: Residual network; GNSS-R; Wind speed; Deep learning; CYGNSS

Fnu Neha, Deepshikha Bhati, Deepak Kumar Shukla, Sonavi Makarand Dalvi, Nikolaos
Mantzou, Safa Shubbar,
An analytics-driven review of U-Net for medical image segmentation,
Healthcare Analytics,
Volume 8,
2025,
100416,
ISSN 2772-4425,
[Link]
([Link]
Abstract: Medical imaging (MI) plays a vital role in healthcare by providing detailed insights
into anatomical structures and pathological conditions, supporting accurate diagnosis and
treatment planning. Noninvasive modalities, such as X-ray, magnetic resonance imaging
(MRI), computed tomography (CT), and ultrasound (US), produce high-resolution images of
internal organs and tissues. The effective interpretation of these images relies on the precise
segmentation of the regions of interest (ROI), including organs and lesions. Traditional
methods based on manual feature extraction are time-consuming, inconsistent, and not
scalable. This review explores recent advances in artificial intelligence (AI)-driven
segmentation, focusing on Convolutional Neural Network (CNN) architectures, particularly
the U-Net family and its variants—U-Net++, and U-Net 3+. These models enable automated,
pixel-wise classification across modalities and have improved segmentation accuracy and
efficiency. The review outlines the evolution of U-Net architectures, their clinical integration,
and offers a modality-wise comparison. It also addresses challenges such as data
heterogeneity, limited generalizability, and model interpretability, proposing solutions
including attention mechanisms and Transformer-based designs. Emphasizing clinical
applicability, this work bridges the gap between algorithmic development and real-world
implementation.
Keywords: Medical image analysis; Deep learning models; Image segmentation; Healthcare
imaging; Pattern recognition; Artificial intelligence

Sebastian Wandelt, Ming Zhou, Shuhua Song, Xiaoqian Sun,


On the construction of effective DRONEWALLS: Technologies and challenges for Counter-UAV
systems,
Journal of the Air Transport Research Society,
Volume 6,
2026,
100093,
ISSN 2941-198X,
[Link]
([Link]
Abstract: The proliferation of unmanned aerial vehicles (UAVs) has exposed various
security/safety vulnerabilities in our infrastructure systems. The increased use of drones for
military purposes, especially in the ongoing conflict between Russia and Ukraine, together
with geopolitical tensions, further exacerbate the situation. In September 2025, the
European Union has called for a “drone wall” initiative to build up a robust Counter-UAV (C-
UAV) system along the east flank of Europe. Existing research on C-UAV often remains
fragmented, with detection, analysis, and neutralization studied in isolation or reviews
covering a limited number of papers in community-specific outlets only. Our study
introduces the so-called DRONEWALLS framework, a unified approach structured in three
phases: Detection (radio frequency/spectrum sensing, radar/LiDAR, optical/thermal imaging,
and acoustic/vibration-based), analysis/decision-making (weighted sensor fusion, edge
AI/real-time processing, and simulation/benchmarking), as well as action/compliance
(neutralization, legal/ethical adherence, and litigation). By synthesizing over 400 studies,
nearly three times as many as in existing reviews, we provide an accurate depiction of the
state of the art on C-UAV systems and their technological challenges, including systemic
integration and adaptive, scalable solutions tailored to both civilian and strategic security
needs.
Keywords: Counter-UAV; Drone; Detection; Decision-making; Compliance

Lingwei Xu, Haiyang Sun, Kai Wang, Gaofeng Nie, Zhe Chen, T. Aaron Gulliver,
An intelligent wireless sensing algorithm for complex cross-domain scenarios based on DB-
FA-YoLov6,
Expert Systems with Applications,
Volume 296, Part A,
2026,
128912,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Wireless sensing technology can identify human motion via feature information
from WiFi signals. The popularity of smartphones, wearable devices, and other smart
devices has increased the use of wireless sensing in fields such as smart homes, smart
healthcare, human–computer interaction, and autonomous vehicles. However, the mobile
communication environment is complex and dynamic which makes wireless sensing
challenging. The issues include low model sensing accuracy, poor scene generalization
ability, and high environmental dependence. Therefore, this paper proposes a cross-domain
intelligent wireless sensing algorithm based on a double branch frequency attention
mechanism Yolov6 network called DB-FA-YoLov6. This integrates a Yolov6 neural network,
frequency attention module, and residual module to provide efficient extraction of signal
features and enhance model generalization. The goal is to reduce the effect of the
environment on sensing tasks and improve robustness, portability, and cross-domain
accuracy. The DB-FA-YOLOv6 model integrates two types of residual modules, BasicBlock and
Bottleneck. It replaces the large modules in the Yolov6 network model with lightweight
structures, which can decrease the number of parameters, improve the efficiency of model
training and testing, and reduce the complexity. Compared with current sensing algorithms
such as Vision Transformer Network for Multiple Vision Tasks (ViT-MVT), Environment
Independent (EI), and Joint Adversarial Domain Adaptation (JADA), the proposed DB-FA-
YOLOv6 algorithm has better sensing accuracy, sensing efficiency, and cross-domain
performance. For the in-domain scenario, the proposed algorithm achieves improvements of
10.0 % in sensing accuracy and 10.1 % in sensing efficiency. The sensing accuracy of the
proposed algorithm in cross-domain scenarios, namely location and orientation, is improved
by 10.5 % and 9.7 %, and the sensing efficiency is improved by 7.0 % and 52.1 %,
respectively.
Keywords: Intelligent wireless sensing; Cross-domain sensing; Attention mechanism; Double-
branch Yolov6 neural network

Adrito Das, Danyal Z. Khan, Dimitrios Psychogyios, Yitong Zhang, John G. Hanrahan,
Francisco Vasconcelos, You Pang, Zhen Chen, Jinlin Wu, Xiaoyang Zou, Guoyan Zheng, Abdul
Qayyum, Moona Mazher, Imran Razzak, Tianbin Li, Jin Ye, Junjun He, Szymon Płotka, Joanna
Kaleta, Amine Yamlahi, Antoine Jund, Patrick Godau, Satoshi Kondo, Satoshi Kasai, Kousuke
Hirasawa, Dominik Rivoir, Stefanie Speidel, Alejandra Pérez, Santiago Rodriguez, Pablo
Arbeláez, Danail Stoyanov, Hani J. Marcus, Sophia Bano,
PitVis-2023 challenge: Workflow recognition in videos of endoscopic pituitary surgery,
Medical Image Analysis,
Volume 106,
2025,
103716,
ISSN 1361-8415,
[Link]
([Link]
Abstract: The field of computer vision applied to videos of minimally invasive surgery is ever-
growing. Workflow recognition pertains to the automated recognition of various aspects of a
surgery, including: which surgical steps are performed; and which surgical instruments are
used. This information can later be used to assist clinicians when learning the surgery or
during live surgery. The Pituitary Vision (PitVis) 2023 Challenge tasks the community to step
and instrument recognition in videos of endoscopic pituitary surgery. This is a particularly
challenging task when compared to other minimally invasive surgeries due to: the smaller
working space, which limits and distorts vision; and higher frequency of instrument and step
switching, which requires more precise model predictions. Participants were provided with
25-videos, with results presented at the MICCAI-2023 conference as part of the Endoscopic
Vision 2023 Challenge in Vancouver, Canada, on 08-Oct-2023. There were 18-submissions
from 9-teams across 6-countries, using a variety of deep learning models. The top
performing model for step recognition utilised a transformer based architecture, uniquely
using an autoregressive decoder with a positional encoding input. The top performing model
for instrument recognition utilised a spatial encoder followed by a temporal encoder, which
uniquely used a 2-layer temporal architecture. In both cases, these models outperformed
purely spatial based models, illustrating the importance of sequential and temporal
information. This PitVis-2023 therefore demonstrates state-of-the-art computer vision
models in minimally invasive surgery are transferable to a new dataset. Benchmark results
are provided in the paper, and the dataset is publicly available at:
[Link]
Keywords: Endoscopic vision; Instrument recognition; Step recognition; Surgical AI; Surgical
vision; Workflow analysis

Haoming Feng, Huaqing Li, Wenwen Zhu, Denghao Li, Yukun Huang,
Micro-motion enhanced multi-person activity recognition with millimeter-wave radar,
Measurement,
Volume 258, Part B,
2026,
119090,
ISSN 0263-2241,
[Link]
([Link]
Abstract: As a non-contact sensing device, millimeter-wave radar exhibits unique strengths
in human activity recognition (HAR). Existing methods rely on micro-Doppler signatures for
activity classification, but they often encounter feature aliasing in multi-person activity
recognition (MPAR) scenarios. Although point cloud-based approaches can distinguish
individual targets, they primarily extract static morphological features, neglecting the micro-
motion information of human joints, which is crucial for accurate activity recognition. To
address these limitations, we proposes an innovative MPAR framework that integrates
spatial point clouds and micro-motion features. First, an improved point cloud data
association algorithm is applied to achieve multi-target point cloud feature separation,
followed by a dynamic projection mechanism to construct time–Doppler feature maps.
Then, a torso micro-motion enhancement algorithm is designed to enhance the details of
human body movements. Finally, a CNN-LSTM hybrid network architecture with a temporal-
attention is constructed for action classification. Experimental results show that the
proposed micro-motion enhancement algorithm improves recognition accuracy by 27.1%
and 2.3%, compared to two traditional time–frequency analysis methods. Furthermore,
MPAR task in occlusion scenarios achieves recognition accuracy of 93.5%. In summary,
proposed framework not only retains the inherent advantages of millimeter-wave radar but
also significantly enhances multi-person activity recognition in complex scenarios.
Keywords: Human activity recognition (HAR); Multi-person activity recognition (MPAR);
Feature separability; Temporal attention; Occlusion scenarios
Vera Lucia Da Silveira Nantes Button,
Chapter 5 - Displacement, Velocity, and Acceleration Transducers,
Editor(s): Vera Lucia Da Silveira Nantes Button,
Principles of Measurement and Transduction of Biomedical Variables,
Academic Press,
2015,
Pages 155-219,
ISBN 9780128007747,
[Link]
([Link]
Abstract: This chapter presents the main methods for measuring displacement, velocity, and
acceleration, commonly used in biomedical determination of other quantities, such as
pressure, flow, and force. This chapter will discuss resistive transducers with particular
emphasis on strain gauges, both metallic and semiconductor, capacitive, piezoelectric and
inductive transducers, with special emphasis on the functioning of linear variable differential
transformer. The functioning principle of transducers for velocity and acceleration
measurements, tachometers, and accelerometers, respectively, will be presented and their
biomedical applications will be exemplified.
Keywords: Resistive transducer; strain gage; capacitive transducer; LVDT; accelerometer

M.E. Pleydell,
Laser Doppler vibration measurement using a polarisation-based device,
Measurement,
Volume 6, Issue 1,
1988,
Pages 10-18,
ISSN 0263-2241,
[Link]
([Link]
Abstract: A non-contacting laser-based displacement and vibration measuring device is
presented. It is based on the coherent detection of the Doppler shift introduced into light
scattered by an optically rough target object. Geometric manipulation of the beams within
the system compensates for the speckle in the scattered light. The sense of the target
motion is derived using a passive technique based on controlled polarisation of the
reference beam. In its simplest mode of operation it has a resolution of half of the laser
wavelength, but this may be increased by a factor of four. Continuous measurements of
displacement of up to 10 cm have been made with an experimental system, with less than
1% error; vibrations of smaller amplitude have been measured with greater accuracy. In the
existing configuration the maximum target velocity in the direction of the measuring beam is
limited by the 100 kHz analogue signal processing roll-off to 32 mm/s; however, this may
easily be extended by using more sophisticated electronics.
Keywords: Laser doppler; vibration; speckle; polarisation

Jiachen Yang, Zhuo Zhang, Wei Mao, Yue Yang,


Identification and micro-motion parameter estimation of non-cooperative UAV targets,
Physical Communication,
Volume 46,
2021,
101314,
ISSN 1874-4907,
[Link]
([Link]
Abstract: With the wide application of unmanned aerial vehicles (UAV) in industrial
production, transportation, and entertainment, it is urgent to identify UAVs in time.
Traditional UAV recognition mainly depends on wireless communication, which puts forward
high requirements for a communication environment and has no way to deal with non-
cooperative targets. Therefore, it is urgent to explore a UAV target recognition scheme based
on perception. In this paper, aiming at the time series preprocessing method, a coding-based
sequence preprocessing method is proposed. This method effectively improves the effect of
the Deep Learning method in the identification task. In order to verify the ability of Deep
Learning in radar time series data processing and the effectiveness of the proposed method,
the Deep Learning method is used to analyze the radar signal time series of the target to
realize the target recognition. Finally,considering the influence of micro-motion factors on
UAV targets, the neural network is used to estimate UAV’s micro-motion parameters to
enhance the ability of target recognition with the help of micro-motion information.
Keywords: Radar cross section; UAV identification; Micro-motion; Deep learning

Federico Carlos Gallardo, Jorge Luis Bustamante, Clara Martin, Cristian Marcelo Orellana,
Mauricio Rojas Caviglia, Guillermo Garcia Oriola, Agustin Ignacio Diaz, Pablo Augusto Rubino,
Vicent Quilis Quesada,
Novel Simulation Model with Pulsatile Flow System for Microvascular Training, Research, and
Improving Patient Surgical Outcomes,
World Neurosurgery,
Volume 143,
2020,
Pages 11-16,
ISSN 1878-8750,
[Link]
([Link]
Abstract: Background
Simulation allows surgical trainees to acquire surgical skills in a safe environment. With the
aim of reducing the use of animal experimentation, different alternative nonliving models
have been pursued. However, one of the main disadvantages of these nonliving models has
been the absence of arterial flow, pulsation, and the ability to integrate both during a
procedure on a blood vessel. In the present report, we have introduced a microvascular
surgery simulation training model that uses a fiscally responsible and replicable pulsatile
flow system.
Methods
We connected 30 human placentas to a pulsatile flow system and used them to simulate
aneurysm clipping and vascular anastomosis.
Results
The presence of the pulsatile flow system allowed for the simulation of a hydrodynamic
mechanism similar to that found in real life. In the aneurysm simulation, the arterial flow
could be evaluated before and after clipping the aneurysm using a Doppler ultrasound
system. When practicing anastomosis, the use of the pulsatile flow system allowed us to
assess the vascular flow through the anastomosis, with verification using the Doppler
ultrasound system. Leaks were manifested as “blood” pulsatile ejections and were more
frequent at the beginning of the surgical practice, showing a learning curve.
Conclusions
We have provided a step-by-step guide for the assembly of a replicable and inexpensive
pulsatile flow system and its use in placentas for the simulation of, and training in,
performing different types of anastomoses and intracranial aneurysms surgery.
Keywords: Anastomosis; Aneurysm; Microsurgery; Neurosurgery; Placenta; Training
simulation

Wentao He, Jianfeng Ren, Ruibin Bai, Xudong Jiang,


Radar gait recognition using Dual-branch Swin Transformer with Asymmetric Attention
Fusion,
Pattern Recognition,
Volume 159,
2025,
111101,
ISSN 0031-3203,
[Link]
([Link]
Abstract: Video-based gait recognition suffers from potential privacy issues and performance
degradation due to dim environments, partial occlusions, or camera view changes. Radar has
recently become increasingly popular and overcome various challenges presented by vision
sensors. To capture tiny differences in radar gait signatures of different people, a dual-branch
Swin Transformer is proposed, where one branch captures the time variations of the radar
micro-Doppler signature and the other captures the repetitive frequency patterns in the
spectrogram. Unlike natural images where objects can be translated, rotated, or scaled, the
spatial coordinates of spectrograms and CVDs have unique physical meanings, and there is
no affine transformation for radar targets in these synthetic images. The patch splitting
mechanism in Vision Transformer makes it ideal to extract discriminant information from
patches, and learn the attentive information across patches, as each patch carries some
unique physical properties of radar targets. Swin Transformer consists of a set of cascaded
Swin blocks to extract semantic features from shallow to deep representations, further
improving the classification performance. Lastly, to highlight the branch with larger
discriminant power, an Asymmetric Attention Fusion is proposed to optimally fuse the
discriminant features from the two branches. To enrich the research on radar gait
recognition, a large-scale NTU-RGR dataset is constructed, containing 45,768 radar frames of
98 subjects. The proposed method is evaluated on the NTU-RGR dataset and the MMRGait-
1.0 database. It consistently and significantly outperforms all the compared methods on
both datasets. The codes are available at: [Link]
Keywords: Micro-Doppler signature; Radar gait recognition; Spectrogram; Cadence velocity
diagram; Asymmetric Attention Fusion

Berke Cansiz, M. Ikbal Karadeli, Nizamettin Aydin, Gorkem Serbes,


A comprehensive approach for high intensity blood flow signal classification using multi-
scale spectrogram fusion and tunable Q-factor wavelet analysis,
Biomedical Signal Processing and Control,
Volume 109,
2025,
107991,
ISSN 1746-8094,
[Link]
([Link]
Abstract: The proposed manuscript presents a novel automated approach for the
classification of high intensity blood flow signals by combining the features obtained from
the hybrid multi-scale spectrogram image samples with the tunable Q-factor wavelet
transform based statistical features. Micro embolic signals, whose presence in the blood
flow are early indicators of stroke condition, have transient nature. The existence of two
other similar high-intensity signals (artifacts and Doppler speckle) make the traditional micro
embolic signal detection a challenging procedure. The proposed method utilizes the
convolutional neural network (CNN) architecture DenseNet201 to extract features from
spectrogram images generated at different scales to mimic the human ear’s basilar
membrane for better understanding of blood flow signals. Additionally, the tunable Q-factor
wavelet transform (TQWT) was employed to extract statistical features from the
decomposed sub-bands, which were obtained by using an adjustable time-frequency
domain analysis. This adjustable time-frequency tuning property of the TQWT has provided
a better representation of the non-stationary characteristics of the Doppler ultrasound
signals. A feature-level fusion technique was also applied to combine the strengths of both
CNN-based and TQWT-based features to enhance the discriminative power of the proposed
automated classification approach. Experimental results demonstrate the effectiveness of
the proposed approach, in which the hybrid spectrogram image based features and
optimum parameter set tuned TQWT based features were fused. A notable performance
was achieved in distinguishment of micro embolic signals from artifacts and speckle signals
by obtaining 94.42% accuracy and 94.36% F1-Score.
Keywords: Embolic signals; Artifacts; Doppler speckles; Tunable Q-factor wavelet transform;
DenseNet201; Spectrogram; Feature level fusion; Ensemble learning

Peng Zhu, Kai Chen, Cong Xu, Shuangfei Zhao, Ruiqi Shen, Yinghua Ye,
Development of a monolithic micro chip exploding foil initiator based on low temperature
co-fired ceramic,
Sensors and Actuators A: Physical,
Volume 276,
2018,
Pages 278-283,
ISSN 0924-4247,
[Link]
([Link]
Abstract: The performances of exploding foil initiator has been improved in terms of high
safety and high reliability over the other electrical initiators. This work develops a monolithic
micro-chip exploding foil initiator (McEFI) based on low temperature co-fired ceramic. This
McEFI has a really monolithic construction without any adhesive or bonding parts, which
endows it with the inherent benefits of large volume/low-cost production and high-
reproducibility. Using this method, it can eliminate complicated fabrication processes such as
precise machining, aligning and bonding that are inevitable in conventional manufacturing
process of EFIs. Its primary characteristics for electrical burst, flyer acceleration and
detonating capability are presented and discussed. Results show that McEFI could reliably
detonate hexanitrostilbene at 2.5 kV/0.22 μF, and its performance could be improved by
optimizing design parameters.
Keywords: Exploding foil initiator; Low temperature co-fired ceramic; Monolithic
construction; Firing characteristics

Pengfei Xue, Peng Xiong, Heng Hu, Tao Wang, Mingyu Li, Qingxuan Zeng,
Integration of the exploding foil initiator with capacitor discharge unit and its performance
characterization,
Measurement,
Volume 242, Part C,
2025,
116069,
ISSN 0263-2241,
[Link]
([Link]
Abstract: This study presents a novel integrated exploding foil initiator system (EFIs),
consisting of a planar switch, an EFI chip, and a capacitor. Additionally, we propose an
evaluation method for integrated EFIs that combines numerical simulations with
experimental testing, enabling a thorough assessment of their performance. Under test
conditions of 900 V/0.22 μF, the EFI demonstrated an inductance of 10.2 nH and a resistance
of 107.5 mΩ, measured via the sampling resistance method. The operational behavior of the
EFIs was analyzed through a two-dimensional metal electrical explosion model and a one-
dimensional flyer propulsion model, with the results compared to experimental data. Results
indicated that under 900–1200 V/0.22 μF, the deviations between measured and calculated
values for electrical explosion performance and flyer velocity were both within 5 %,
validating the accuracy of the evaluation approach. Firing tests further confirmed that the
EFIs successfully detonated HNS-IV at 1000 V/0.22 μF.
Keywords: Exploding foil initiator system; Flexible Printed Circuit; Micro-Electro-Mechanical-
Systems; Calculation model

Yanfang Yu, Dadian Wang, Huibo Meng, Jinyu Guo, Zhiying Han,
Spatiotemporal evolution analysis of bubble swarms in gas-liquid static mixers based on an
improved CNN,
International Journal of Heat and Mass Transfer,
Volume 255, Part 1,
2026,
127731,
ISSN 0017-9310,
[Link]
([Link]
Abstract: Static mixers (SM) are significant in multiphase mixing due to their high efficiency,
energy saving and reliability. Komax static mixer (Komax) effectively promotes bubble
breakup and liquid turbulence that leads to the formation of numerous small and medium-
sized bubbles. Three-dimensional features of bubbles with high deformation, overlap and
offset angle under turbulent conditions in SM are difficult to be extracted. In this paper, a
lightweight multidimensional backbone neural network incorporating a deformable
attention transformer is proposed for parallelled and targeted extracting the
multidimensional information of bubbles. The deviation between convolutional neural
network-predicted void fractions (Vf) and experimentally measured values was below 10%.
Compared with the empty pipe, the flow pattern could maintain bubble flow even if the Vf is
up to 0.7 in Komax. The bubble spatiotemporal characteristics and radial Sauter mean
diameter (d32) distribution demonstrated the local slug flow occurred when average velocity
differences between large and small bubble swarms exceeded 0.37 m/s. This phenomenon is
effectively restrained when the bubble velocity uniformity factor remains below 0.1 in
Komax. Chaotic characteristics analysis combined with longitudinal bubble distribution
revealed the bubble flow was stable when the superficial liquid velocity (UL) = 0.0283–
0.0424 m/s and the slip velocity (US) = 0.15–0.25 m/s. Additionally, the comprehensive
breakup efficiency (H) of Komax is 7.9%‒11.8% higher than that of Quatro static mixers at US
= 0.193‒0.259 m/s. Finally, the generalization of the improved model for predicting flow
patterns and extracting bubble information was excellent with the precision exceeding 0.96.
Keywords: Komax static mixer; Neural network; Bubble dynamics; Deformable attention
transformer; Chaos analysis

Yuanzhao Yang, Qi Jiang,


A novel phase-based video motion magnification method for non-contact measurement of
micro-amplitude vibration,
Mechanical Systems and Signal Processing,
Volume 215,
2024,
111429,
ISSN 0888-3270,
[Link]
([Link]
Abstract: Vibration analysis is crucial for structural health monitoring and fault diagnosis.
Conventional contact sensors present limitations, prompting the adoption of non-contact
methods such as laser Doppler vibration measurements and computer vision-based
techniques. Among these, phase-based video motion magnification has gained prominence
for its high resolution and ability to capture comprehensive vibration data across the entire
field. However, traditional video motion magnification methods are often affected by noise
and artifacts when facing micro vibration measurements. Especially at high magnification,
artifacts such as ”double edges” often appear, which seriously affects the accuracy of
vibration analysis. In addition, for complex structures, the existing methods still have some
difficulties in extracting vibration modes and maintaining fine motion details. Therefore, we
propose a novel method for micro-amplitude vibration magnification and a combination of
the Horn–Schunck method and the motion intensity averaging method for non-contact
vibration displacement measurement. The proposed method uses a complex steerable
pyramid to decompose the video into multi-scale and multi-directional sub-bands, extract
the phase variations of interest and magnify them, and then an optimized one-dimensional
row-gradient domain-guided image filter is used to finely eliminate the double edges,
artifacts, and other noises in the high-frequency sub-bands, and finally synthesize the
motion-magnified video. Experimental validations on a vibration platform and real structural
elements, including a three-story building and an aluminum cantilever beam, demonstrate
the method’s superiority in preserving structural clarity and minimizing information loss. Our
approach significantly enhances structural vibration analysis, offering accurate frequency
identification and mode shape extraction.
Keywords: Vibration signal analysis; Non-contact measurement; Phase-based motion
magnification; Video processing

Yuxi Qin, Su Pan, Weiwei Zhou, Duowei Pan, Zibo Li,


WiASL: American Sign Language writing recognition system using commercial WiFi devices,
Measurement,
Volume 218,
2023,
113125,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Contactless human–computer interaction has been widely used in a variety of
scenarios. WiFi-based approaches can effectively protect user privacy compared to vision-
based and wearable sensor-based approaches. However, existing WiFi-based input systems
require long movement trajectories, making input inefficient and fatiguing users quickly. This
paper presents WiASL, a commercial WiFi device-based micro-motion input system. WiASL
uses American Sign Language (ASL) to represent letters, which requires only finger
movements for most letters. In WiASL, we merge the amplitudes and phases to detect the
signal segments containing micro-motions, extract features using the Attention-based
Weighted Linear Discriminant Analysis (AWLDA) algorithm and a spatiotemporal deep neural
network, and utilize a fully connected neural network for classification. The experimental
results demonstrate that WiASL achieves a 3.3% false rejection rate for detection and
95.10% accuracy for recognition. WiASL can significantly improve input efficiency and
maintain a high recognition accuracy compared with existing systems.
Keywords: Channel state information; American sign language; Micro-motion recognition;
Attention mechanism

Ting-Ruen Wei, Yuling Yan,


Multimodal medical imaging AI for breast cancer diagnosis: A comprehensive review,
Intelligent Oncology,
Volume 2, Issue 1,
2026,
100037,
ISSN 2950-2616,
[Link]
([Link]
Abstract: Traditional artificial intelligence (AI)-based methods for breast cancer diagnosis
often rely on a single modality, such as ultrasound images. With the rise of multimodal
approaches, multiple data sources, including imaging from diverse medical modalities,
structured clinical information, and unstructured medical reports, are increasingly integrated
to provide richer and more informative signals for model training. This survey reviews the
data modalities employed in AI-based breast cancer research, examines common
multimodal combinations and fusion strategies, and discusses their applications across
clinical tasks such as diagnosis, treatment planning, and outcome prediction. By
consolidating current literature and identifying critical gaps, this survey aims to guide future
research toward the development of reliable, clinically relevant multimodal AI systems for
use in breast cancer management.
Keywords: Breast cancer; Artificial intelligence; Machine learning; Deep learning; Multimodal

Yinshen Wang, Zhengxuan Hu, Ping Zhang, Zhihua Fan, Wenming Li, Xuejun An, Xiaochun Ye,
A real-time edge SAR imaging acceleration architecture utilizing multi-level dataflow
parallelism,
Journal of Systems Architecture,
Volume 170,
2026,
103635,
ISSN 1383-7621,
[Link]
([Link]
Abstract: Synthetic Aperture Radar (SAR), a key radar signal processing technology, is widely
deployed on edge devices due to its high resolution, long-range detection, and all-weather,
all-day imaging. However, achieving real-time SAR imaging on resource-constrained edge
platforms is challenging because SAR algorithms involve complex workflows and diverse
operators. Prior works focused on DSP, FPGA, and GPU platforms have struggled to balance
performance and power efficiency. Furthermore, frequent kernel switching necessitates
repeated reconfigurations and memory accesses, increasing latency. To address these
challenges, we propose a SAR-tailored dataflow model enabling multi-level dataflow
parallelism across various operators. First, we introduce a reconfigurable architecture that
integrates customized processing elements optimized for SAR. Second, we propose a multi-
level dataflow model exploiting parallelism at the task, instruction, and node levels.
Additionally, we present an instruction switching mechanism and a preprocessing method
for matrix transposition to reduce kernel switching overhead. Experimental results
demonstrate that for an 8K × 8K image, our approach achieves a processing time of 0.66 s,
with a 37.1× performance improvement over a CPU (i5-9500) and a 1.42× improvement over
a GPU (NVIDIA Orin). Evaluations of SAR operators across diverse scales indicate a 1.45×
performance gain over state-of-the-art reconfigurable architectures featuring dataflow
modeling.
Keywords: Synthetic aperture radar (SAR); Reconfigurable architecture; Dataflow model;
Multi-level parallelism; Real-time imaging

Abdulrhman H. Al-Jebrni, Saba Ghazanfar Ali, Bin Sheng, Huating Li, Xiao Lin, Ping Li,
Younhyun Jung, Jinman Kim, Li Xu, Lixin Jiang, Jing Du,
SynTaskNet: A synergistic multi-task network for joint segmentation and classification of
small anatomical structures in ultrasound imaging,
Computer Vision and Image Understanding,
Volume 263,
2026,
104616,
ISSN 1077-3142,
[Link]
([Link]
Abstract: Segmenting small, low-contrast anatomical structures and classifying their
pathological status in ultrasound (US) images remain challenging tasks in computer vision,
especially under the noise and ambiguity inherent in real-world clinical data. Papillary
thyroid microcarcinoma (PTMC), characterized by nodules ≤1.0 cm, exemplifies these
challenges where both precise segmentation and accurate lymph node metastasis (LNM)
prediction are essential for informed clinical decisions. We propose SynTaskNet, a synergistic
multi-task learning (MTL) architecture that jointly performs PTMC nodule segmentation and
LNM classification from US images. Built upon a DenseNet201 backbone, SynTaskNet
incorporates several specialized modules: a Coordinated Depth-wise Convolution (CDC) layer
for enhancing spatial features, an Adaptive Context Block (ACB) for embedding contextual
dependencies, and a Multi-scale Contextual Boundary Attention (MCBA) module to improve
boundary localization in low-contrast regions. To strengthen task interaction, we introduce a
Selective Enhancement Fusion (SEF) mechanism that hierarchically integrates features across
three semantic levels, enabling effective information exchange between segmentation and
classification branches. On top of this, we formulate a synergistic learning scheme wherein
an Auxiliary Segmentation Map (ASM) generated by the segmentation decoder is injected
into SEF’s third class-specific fusion path to guide LNM classification. In parallel, the
predicted LNM label is concatenated with the third-path SEF output to refine the Final
Segmentation Map (FSM), enabling bidirectional task reinforcement. Extensive evaluations
on a dedicated PTMC US dataset demonstrate that SynTaskNet achieves state-of-the-art
performance, with a Dice score of 93.0% for segmentation and a classification accuracy of
94.2% for LNM prediction, validating its clinical relevance and technical efficacy.
Keywords: Medical image classification; Medical image segmentation; Multi-task learning;
Synergistic learning; Small anatomical structures

Li Qiusheng, Zhu Huajuan,


Target classification with low-resolution radars based on cyclic bispectrum and improved
ACGAN,
Measurement,
Volume 259, Part B,
2026,
119715,
ISSN 0263-2241,
[Link]
([Link]
Abstract: To address the challenges of insufficient generalization and high noise sensitivity in
low-resolution radar target recognition under limited-sample conditions, this paper
proposes a joint optimization framework integrating cyclic bispectral analysis and an
improved Auxiliary Classifier Generative Adversarial Network (ACGAN). First, a third-order
cyclic cumulant spectral model is designed to extract modulation-specific signatures of
aircraft targets in the cyclostationary domain, effectively suppressing both Gaussian and
non-Gaussian noise while preserving discriminative features that are robust to low SNR
conditions (maintaining 92.7 % accuracy at 0 dB). Second, an enhanced ACGAN architecture
is developed by incorporating self-attention mechanisms and Wasserstein distance
optimization with gradient penalty, with spectral normalization and dynamic gradient
penalties introduced to stabilize training dynamics and improve synthetic sample fidelity.
Extensive experiments on a real-world dataset collected by a certain Chinese-made VHF-
band radar demonstrate that the proposed method achieves state-of-the-art performance,
with average recognition accuracies of 98.46 % and 98.52 % for approaching and departing
targets in complex noise environments, respectively, alongside a Kappa coefficient exceeding
0.97. Comprehensive comparisons with traditional methods (wavelet, HOS, FrFT), GAN
variants (WGAN-GP, AFGAN + ResNet, Diffusion-GAN), VAE-based approaches, few-shot
learning models (ProtoNet, MatchingNet), and a modern Vision Transformer (ViT) baseline
show consistent improvements of 1.38–12.19 % in accuracy. Ablation studies validate the
contributions of key components, where the self-attention module and Wasserstein
optimization improve accuracy by 1.27 % and 0.97 %, respectively. Furthermore, embedded
platform tests confirm the framework’s feasibility for real-time deployment (inference
time < 15 ms/sample), offering a robust solution for resource-constrained radar systems.
This work highlights the efficacy of unifying physics-inspired feature extraction with
stabilized deep generative models for practical radar recognition.
Keywords: Radar target recognition; Cyclic bispectral analysis; Auxiliary classifier generative
adversarial network (ACGAN); Self-attention mechanism; Few-shot learning

Wentao Zhao, Guoxiang Tong,


A deep learning based heart rate estimation method for millimeter wave radar,
Measurement,
Volume 255,
2025,
117923,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Contact-based vital sign detection technology has been widely used in the medical
field. However, contact-based devices may cause discomfort to users and suffer from user
dependency issues. Frequency Modulated Continuous Wave (FMCW) millimeter-wave radar
provides an efficient and accurate solution for heart rate and respiratory rate monitoring.
Nevertheless, due to the small amplitude of heartbeat micro-motion signals, they are
susceptible to noise interference such as respiratory harmonics, making accurate
measurement challenging. Aiming at the problem that traditional methods are difficult to
adapt to different environmental noises, we propose a heart rate estimation method based
on a convolutional neural network. The heart rate estimation accuracy is significantly
improved by recognizing the phase change patterns in radar signals. We reduce the
interference of respiratory harmonics on the heart rate micromotion signal by decomposing
the extracted phase signal into multiple frequency components using the Empirical Wavelet
Transform (EWT) algorithm. The proposed deep learning model is used by means of depth
convolution in order to balance the model size and accuracy. Additionally, we introduce
time–frequency channels into the convolutional neural network to further enhance its
feature extraction capability. Comparisons with various related works on heart rate
distribution, sensing range, and subjects demonstrate the lightweight nature and higher
accuracy of the proposed model. Extensive experiments are conducted on a dataset based
on the TI AWR1642BOOST radar. The results show that the proposed method achieves
outstanding performance, with an accuracy of ±5 BPM in heart rate monitoring.
Keywords: Deep learning; Heart rate (HR); Millimeter-wave (mmW) radar; Vital signs
monitoring; Moving target indication (MTI)

Yuchen Liu, Mahshid Hafezi, Andrew Feeney,


Dynamic stability of a nitinol Langevin ultrasonic transducer under power and current
tracking conditions,
Applied Acoustics,
Volume 225,
2024,
110188,
ISSN 0003-682X,
[Link]
([Link]
Abstract: The Langevin transducer is widely used in medical and industrial ultrasonics,
typically comprising a piezoelectric ceramic stack preloaded between two end-masses. A
principal operational challenge is that, depending on the target application or operational
environment, piezoelectric ceramics commonly experience elevated temperatures
promoting nonlinear dynamic behaviours, often reducing operational performance.
Transducers can be operated via burst excitation to mitigate this, but such methods can have
limited efficacy, particularly at elevated temperatures. A novel emergent approach is the
incorporation of shape memory Nitinol into the transducer, where its temperature-
dependent microstructural phase transformations can compensate for changes in the
mechanical properties of the piezoelectric ceramics. However, the ability to maintain
dynamic stability under power and current tracking, common methods used to ensure
stability of Langevin transducers in practical applications, remains unknown. In this research,
a cascaded Nitinol Langevin transducer is used to demonstrate its practical application
potential. Dynamic stability is assessed whilst power and current levels are tracked in
operation. Stable operating frequency, electrical impedance, and vibration amplitude are
shown, using a combination of methods including electrical impedance analysis and laser
Doppler vibrometry. This research demonstrates the practical application of Langevin
transducers incorporating shape memory materials, including with dynamic stability at
elevated temperatures.
Keywords: Nitinol; Langevin transducer; Dynamic stability; Tracking methods

Heng Zhao, Zhili Long, Shuyuan Ye, Jianzhong Ju, Yuxiang Li,
Ultrasonic tool shank with multiple vibration mode for micro-nano drilling: Design,
optimization and experiment,
Sensors and Actuators A: Physical,
Volume 366,
2024,
114987,
ISSN 0924-4247,
[Link]
([Link]
Abstract: The machining effects of hole drilling is limited by merely changing the amplitude.
Adjusting ultrasonic vibration frequency to expand the matching range of machining
parameters optimization is necessary to improve hole drilling effects. Therefore, the design,
optimization, and experiment evaluation for a novel integrated, fined, low-cost dual-
frequency ultrasonic tool shank is proposed, which can work in three vibration modes for
micro-nano precision drilling. The principle of the dual-frequency ultrasonic transducer is
introduced. The electro-mechanical equivalent circuit model with step horn is applied to
study the dual-frequency characteristics. The transducer displacement equation model
based on the wave equation theory is deduced to find the shared vibration node of the first
and third frequency. Sensitivity analysis by equivalent circuit model and Finite element
method (FEM) simulation of the key parameter is both conducted to optimize ultrasonic
transducer geometry dimension. The vibration node of mounting flange is specified in an
identical position for the dual vibration modes. A prototype of ultrasonic tool shank is
fabricated and evaluated in experiments. It demonstrates that the dual working frequency of
the ultrasonic transducer is 34.9 kHz and 104.5 kHz in the first and third mode resonant
vibration, which is benefit to the low and high frequency drilling. Three working modes of
the low, high, and coupling frequency vibration is measured by a self-developed ultrasonic
generator. When 40 V voltage is excited, the vibration amplitudes of the low and high
frequency are 17.8 µm, and 9.3 µm, respectively, which can meet with the drilling
requirement. Moreover, the coupling vibration with the low and high mode are excited
simultaneously, and the motion trajectory of the coupling vibration is observed in a “M”
shape, which is a novel trajectory for the precision micro-nano drilling. Micro drilling
experiments show that the 35 kHz ultrasonic vibration drilling can achieve better machining
effect than conventional drilling (CD). The 105 kHz ultrasonic drilling outperforms the CD,
35 kHz and coupling ultrasonic drilling at cutting force reduction. The coupling ultrasonic
vibration drilling achieves the lowest chipping diameter. It is to be noted that the 105 kHz-
8 µm produce the fined smoothly surface. The 35 kHz, 105 kHz and coupling ultrasonic
drilling can produce better surface micro morphology with less quantity and size of craters
than that of CD.
Keywords: Ultrasonic tool shank; Dual-frequency; Multiple vibration mode; Ultrasonic
transducer

Dingcheng Ji, Jing Lin, Fei Gao, Jiadong Hua, Wenhao Li,
A deep learning-based spatial gradient reconstruction method for efficient damage
identification in composite with high-sparsity Lamb wavefield,
Mechanical Systems and Signal Processing,
Volume 224,
2025,
112018,
ISSN 0888-3270,
[Link]
([Link]
Abstract: The structural integrity and safety of carbon fiber reinforced plastics (CFRP) are
vulnerable to delamination, which is often imperceptible to the naked eye. Although the
Scanning Laser Doppler Vibrometer (SLDV) has shown promise in damage quantification of
CFRP, its time-consuming measurement process limits its application in engineering
scenarios. To address this, we introduce a novel damage index, the spatial gradient, which
captures the interaction between delamination and the wavefield. We have also developed a
neural network capable of reconstructing the spatial gradient directly from high-sparsity
Lamb wavefield data obtained at an extremely low spatial sampling rate, thereby
significantly reducing measurement time. To enhance the network’s capability to detect
wavefield anomalies, we employ the cross-attention technique, allowing for the direct
injection of shallow features representing local wavefield distortions caused by damage into
the decoder. Additionally, we integrate multiple reconstruction layers to guide the wavefield
reconstruction process, ensuring meaningful information is captured at each stage. Our
method achieves substantial improvements in reconstruction accuracy, increasing from 70 %
to 92 % in single-damage scenario and from 14 % to 72 % in multi-damage scenario
compared to the previous state-of-the-art techniques. By using the reconstructed spatial
gradient field for damage imaging through spatial covariance analysis, our approach
demonstrates its feasibility and generalizability across various damage locations. This
suggests its potential as a reliable solution for fast and accurate damage characterization,
reducing the measurement burden and enhancing practical applicability.
Keywords: High-sparsity wavefield reconstruction; Lamb waves; Deep learning; Spatial
gradient imaging

Binglei Yue, Aili Jiang, Chun Yang, Junwei Lei, Heng Liu, Yin Zhang,
Deep Learning-Enhanced Human Sensing with Channel State Information: A Survey,
Computers, Materials and Continua,
Volume 86, Issue 1,
2025,
Pages 1-28,
ISSN 1546-2218,
[Link]
([Link]
Abstract: With the growing advancement of wireless communication technologies, WiFi-
based human sensing has gained increasing attention as a non-intrusive and device-free
solution. Among the available signal types, Channel State Information (CSI) offers fine-
grained temporal, frequency, and spatial insights into multipath propagation, making it a
crucial data source for human-centric sensing. Recently, the integration of deep learning has
significantly improved the robustness and automation of feature extraction from CSI in
complex environments. This paper provides a comprehensive review of deep learning-
enhanced human sensing based on CSI. We first outline mainstream CSI acquisition tools
and their hardware specifications, then provide a detailed discussion of preprocessing
methods such as denoising, time–frequency transformation, data segmentation, and
augmentation. Subsequently, we categorize deep learning approaches according to sensing
tasks—namely detection, localization, and recognition—and highlight representative models
across application scenarios. Finally, we examine key challenges including domain
generalization, multi-user interference, and limited data availability, and we propose future
research directions involving lightweight model deployment, multimodal data fusion, and
semantic-level sensing.
Keywords: Channel State Information (CSI); human sensing; human activity recognition;
deep learning
Chen Nan, Liu Zhi, Le Dexiang, Lai Qingrong, Jiang Bingnian, Li Bin, Wu Jian, Song Yunfeng,
Liu Yande,
Prediction of yellow flesh peach firmness using a novel device and data augmentation
acoustic vibration multi-domain images array Swin Transformer (DA-AVMDIA-SwinT),
Computers and Electronics in Agriculture,
Volume 235,
2025,
110402,
ISSN 0168-1699,
[Link]
([Link]
Abstract: Firmness is an important indicator closely related to the ripeness of yellow flesh
peaches. Non-destructive detection of yellow flesh peach firmness is beneficial for managing
yellow flesh peaches during storage and for quality grading. In this study, a novel acoustic
vibration device based on micro-LDV, a microphone, and a resonance loudspeaker was used
to predict the firmness of yellow flesh peaches. The acoustic vibration response signals of
yellow flesh peaches with different firmness were transformed into acoustic vibration multi-
domain images arrays (AVMDIA), and Swin Transformer (SwinT) model was employed to
extract the acoustic vibration response features of the fruits. Firstly, the prediction
performance of the models built with acoustic multi-domain images array (AMDIA),
vibration multi-domain images array (VMDIA) and AVMDIA as inputs were compared.
Subsequently, the effects of using data augmentation (DA) and not using DA on model
performance were compared. Finally, the prediction performance of the SwinT model, the
Vision Transformer (ViT) model, and the Resnet50 model (CNN-based) for yellow flesh peach
firmness was compared. The results showed that the SwinT model based on data
augmentation acoustic vibration multidomain data array (DA-AVMDIA-SwinT) gave the best
prediction of firmness for yellow flesh peaches, with RP2 = 0.951, RMSEP = 0.515 N/mm,
RPDP = 4.524. In this paper, a method is reported for converting fruit acoustic vibration
response signals into two-dimensional images for fruit firmness prediction is reported for
the first time, which can accurately predict fruit firmness in a simple, fast and cost-effective
manner.
Keywords: Firmness prediction; Yellow flesh peach; Acoustic vibration multi-domain images
array; Data augmentation; Swin transformer

Yuxiang Li, MARIIA KIREEVA, Xicheng Liu, Zhili Long, Shuyuan Ye, Jianzhong Ju,
Design and performance analysis of bidirectional vibration ultrasonic transducer for wire
bonding,
Applied Acoustics,
Volume 238,
2025,
110791,
ISSN 0003-682X,
[Link]
([Link]
Abstract: Ultrasonic frequency, amplitude and vibration mode are the key factors affecting
the stability and reliability of ultrasonic wire bonding. Conventional wire bonding is realized
by ultrasonic transducer (UT) with single frequency and longitudinal vibration. To optimize
the bonding process and achieve high-performance bonding joints, we propose an UT that
utilizes two longitudinal vibration and one bending vibration mode, which realized by the full
PZT and regional polarization PZT. By employing a step-type flexible structure, the flanges for
three vibration modes are optimized at the same node position. To drive the UT, the
amplifier module circuit with 39.98 times amplification and 130 kHz bandwidth is designed.
Under the 9 kg·cm torque for assembling UT, the longitudinal vibration frequencies of the UT
are 71.2 kHz and 122.9 kHz, and the bending vibration frequency is 43.1 kHz. The self-
developed amplifier module is designed to drive the UT for both single and coupled
vibration, and the test results show that the amplitude of the bending vibration is 7.49 µm at
a driving voltage of 20 V, and the amplitudes of the 1st longitudinal vibration and the 2nd
longitudinal vibration are 5.72 µm and 4.44 µm, respectively, which satisfy the amplitude
requirements for wire bonding. Furthermore, the coupled vibration trajectory of the UT
forms a parallelogram, and the vibration area can be adjusted by changing the amplitude.
Experimental results have shown that the bending mode is excited by regionally polarized
PZT, achieving sufficient vibration in both x/y directions while ensuring a small and
lightweight structure. This provides a potential application solution for the new bonding
method of wire bonding.
Keywords: Wire bonding; Ultrasonic transducer; Bidirectional vibration; Regional polarization
PZT; Amplifier module

Van Ngoc Dang, Ngoc Chau Hoang, Quoc Cuong Nguyen, Minh Thuy Le,
Advancing robust human activity recognition via informative mmWave radar characteristics
and a lightweight spatio-spectro-temporal network,
Measurement,
Volume 256, Part A,
2025,
118056,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Human activity recognition (HAR) is increasingly important in aiding our daily life,
with millimeter-wave (mmWave) radar sensors emerging as a promising noninvasive solution
thanks to their excellent spatial and velocity resolution. Although existing radar-based
systems have shown strong performance, they primarily focus on micro-Doppler signatures
while neglecting angle information, which can hinder practical deployment in real-world
scenarios. Moreover, current state-of-the-art recognition models using mmWave radar often
require substantial computational resources, making integration into resource-constrained
devices challenging. This work proposes an efficient radar-based HAR system that leverages
angle and spectro-temporal information from micro-Doppler signatures. Our system utilizes
a multi-channel micro-Doppler representation corresponding to the number of virtual
antenna receivers as input. Then, a lightweight dilated convolutional network, namely SST-
DCN, extracts spatial-aware multi-scale spectro-temporal information through time-
frequency dilated convolutions. Experimental results on our real-world dataset demonstrate
the superiority of our approach compared to conventional features and other state-of-the-
art radar-based HAR systems.
Keywords: Human activity recognition; Millimeter-wave radar; Deep learning; Lightweight
network; Dilated convolution
Yu Dou, Yongjian Li, Shuaichao Yue, Yang Li, Jianguo Zhu,
Measurement of alternating and rotational magnetostrictions of Non-oriented silicon steel
sheets,
Journal of Magnetism and Magnetic Materials,
Volume 571,
2023,
170566,
ISSN 0304-8853,
[Link]
([Link]
Abstract: Extensive numerical analyses and experimental tests have shown that while major
regions of magnetic cores in electromagnetic devices, such as transformers and rotating
electric machines, are dominated by the alternating magnetic fields, the rotating magnetic
fields exist in the corner joints of multi-phase transformers and the yoke area behind the
teeth of rotating electric machines. These magnetic fields cause core losses due to hysteresis
and eddy currents, mechanical vibration and acoustic noises due to magnetostriction. Many
studies have been reported in the literature on alternating and rotational core losses and
alternating magnetostriction, but not so much on rotational magnetostriction, especially in
non-oriented (NO) silicon steel sheets. The vibration and noise caused by rotational
magnetostriction are much higher than those caused by alternating magnetostriction. This
paper reports the development of a magnetostriction measurement system consisting of a
2D symmetrical single sheet tester (SST) and resistance strain gauges. The selection
considerations of resistance strain gauges are summarized. The alternating and rotational
magnetostriction characteristics of a NO silicon steel, B35A300 (0.35 mm), are measured.
Because the rotating magnetic field in the stator is not always ideally circular but more often
elliptical, the magnetostriction under elliptical magnetizations with different axis ratios is
also measured. The magnetostriction anisotropy of the NO steel sheet is observed and
analyzed. The elongation and contraction under alternating, circular and elliptical rotating
magnetizations with different axis ratios are discussed. The results can provide data support
for calculating and designing multi-phase transformers and rotating electric machines under
different types of magnetizations. The research can provide theoretical and practical
guidance for mitigating mechanical vibration and acoustic noise caused by magnetostriction.
Keywords: Rotational magnetostriction; Non-oriented silicon steel sheets; Strain gauge

Donepudi Jagadish, A.V. NageswaraRao, M. Sreenivasa Kumar,


Engine combustion and emission analysis using optical methods: An overview,
Materials Today: Proceedings,
2023,
,
ISSN 2214-7853,
[Link]
([Link]
Abstract: Measurements in combustion places important role in design of reciprocating and
turbo hot engines. The available methods of combustion analysis of engines have gained
much attention and are capable of giving good results in determining the fuel characteristics
and engine design optimizations. However, the full understanding of engine combustion can
be made through optical image processing in addition to the existing methods. Optical tools
are used efficiently to predict the combustion phenomena in internal combustion engine. In
order to observe typical regimes of combustion most widely used techniques are Schlieren
and shadowgraph. The devices used in optical measurements are high speed CCD camera
and Constant Volume Spray Chambers (CVSCs). The analysis is carried to know the spray
characteristics at atmospheric and elevated pressure conditions. The luminosity techniques
are widely used to capture the combustion flames with high speed camera in optically
accessible engines. The present engines have to be redesigned for making optical access for
the desired measurements. Several endoscopy techniques are used for combustion
visualization. Phase Doppler Interferometer (PDI) is used for the quantitative spray
measurements. Particle Image Velocimetry is an intrusive technique which is capable of
characterizing instantaneous velocity fields in a fluid flow. This paper presents the works on
different optical techniques available and is in use to diagnose the engine combustion
especially for reciprocating engines.
Keywords: Optical methods; Engine combustion analysis; Pollution control

Jitao Zhang, Han Qiao, Qingfang Zhang, Bingfeng Ge, D.A. Filippov, Jie Wu, Fang Wang, Jiagui
Tao, Jing Chen, Liying Jiang, Lingzhi Cao,
Compact magnetoelectric power splitter with high isolation using ferrite/piezoelectric
transformer composite,
Journal of Magnetism and Magnetic Materials,
Volume 574,
2023,
170691,
ISSN 0304-8853,
[Link]
([Link]
Abstract: A compact, passive magnetoelectric (ME) power splitter, with a
ferrite/piezoelectric transformer bilayer composite with a coil wound around it, as well as a
structure-constructing strategy, was presented and developed in this research. The
impedance differences in a Rosen-type PT induced by integrated transverse/longitudinal
polarizations facilitated the device realization with higher isolation rather than elementary
LCR lumped elements. Furthermore, the measurements of the material properties and
electrical resonance behaviors for the presented ME splitter were implemented, and the
power-dividing capabilities were verified by measuring the output and corresponding power
for each port of the device. The experimental results demonstrated that the power for Port I
in the ME power splitter reached its maximum values of 154.26 nW at R = 6 kΩ and 475.95
nW at R = 5.5 kΩ. Correspondingly, the power for Port II reached its maximum values of
12.73 nW at R = 30 kΩ and 27.85 nW at R = 12 kΩ. Consequently, a power division ratio of
3:1 was obtained under optimum load resistance and EMR conditions with constant input
power. These results provided an innovative approach to constructing novel functional
power electronics with solid-state materials while suggesting the promising applications for
some special scenarios such as powering and controlling multi-channel LCD strings.
Keywords: Power splitter; Magnetoelectric composite; Piezoelectric transformer

Dapeng Zhang, Yifan Xie, Yining Zhang, Zhengjie Liang, Yutao Tian,
Experimental Advances in Airfoil Dynamic Stall and Transition Phenomena,
Fluid Dynamics and Materials Processing,
Volume 21, Issue 4,
2025,
Pages 697-739,
ISSN 1555-256X,
[Link]
([Link]
Abstract: Airfoil structures play a crucial role across numerous scientific and technological
disciplines, with the transition to turbulence and stall onset remaining key challenges in
aerodynamic research. While experimental techniques often surpass numerical simulations
in accuracy, they still present notable limitations. This paper begins by elucidating the
fundamental principles of transition, dynamic stall, and airfoil behavior. It then provides a
systematic review of six major experimental methodologies and examines the emerging role
of artificial intelligence in this domain. By identifying key challenges and limitations, the
study proposes strategic advancements to address these issues, offering a foundational
framework to guide future research in airfoil structures and related fields.
Keywords: Airfoil; dynamic stall; transition; experimental methodologies; artificial
intelligence

Abid Hussain, Xiaoqiang Zhu, Zhang Sihai, Fujiang Lin,


Spectrum-based anomaly detection using channel state information and attention
mechanisms for elderly health monitoring,
Engineering Applications of Artificial Intelligence,
Volume 166, Part A,
2026,
113583,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Detecting abnormal human activities is essential in environments such as multi
resident homes and elderly care facilities, where continuous visual monitoring is limited.
Early identification of unsafe events, including falls and abrupt movements, can significantly
reduce health risks for older adults. However, reliable anomaly detection using wireless
signals remains challenging due to the scarcity of datasets that represent diverse real world
behavioral deviations. This study proposes a spectrum-aware encoder architecture for
anomaly detection that integrates wireless channel state information with frequency domain
sensing. The method applies wavelet-based denoising, median filtering, and feature
normalization to refine the channel measurements. It then extracts spectral descriptors
including power spectral density, skewness, and kurtosis to capture irregular frequency
signatures associated with abnormal activities. To address class imbalance in real world data,
the training pipeline incorporates the Synthetic Minority Oversampling Technique. The
proposed encoder employs positional encoding and multi head self attention to model long
range temporal relationships in the processed sequences, forming an Artificial Intelligence
framework tailored for human activity anomaly detection. Experimental results demonstrate
that the spectrum-aware encoder achieves higher precision, recall, and overall robustness
compared with deep learning baselines such as convolutional neural networks, long short-
term memory networks, gated recurrent units, spectral temporal Transformer and attention
based methods . The encoder-only design also offers reduced memory usage and faster
training relative to traditional encoder decoder architectures, highlighting its suitability for
real-time deployment in resource-constrained elderly-care environments.
Keywords: Spectrum detection; Anomaly detection; Channel state information; Elderly
health monitoring; Wireless communication systems; Internet of things monitoring systems;
Deep learning

Jiale Ren, Hengyi Li, Aihui Wang, Kenshi Saho, Lin Meng,
Radar-based gait analysis by Transformer-liked network for dementia diagnosis,
Biomedical Signal Processing and Control,
Volume 91,
2024,
105986,
ISSN 1746-8094,
[Link]
([Link]
Abstract: Providing reliable diagnostic evidence to doctors while minimizing financial and
physical burdens on patients is a prominent focus of current research. Gait features show
potential as a clinical marker for dementia diagnosis. Radar is capable of efficient,
contactless collecting human motion. This paper proposes a novel radar-based gait analysis
strategy with a Transformer-liked network for dementia diagnosis. The gait data are collected
by Micro-Doppler radar. Then Welch’s power spectral density estimation is adopted to
obtain the frequency features while unifying and compressing data size. The network is
designed to explore the relationship between dementia and gait features. In the network,
1D convolution is crucial in extracting local features and encoding features into deeper
dimensions. The attention-based module, inspired by the encoder of the Transformer,
possesses an edge in capturing long-sequence dependencies. The dual-stage gating
mechanism enhances the discriminative power of the learned representations by fine-tuning
the weights of extracted features. To validate the effectiveness of the proposed strategy,
comparative experiments are performed with prevailing networks in both time and
frequency domains. Experimental results demonstrate the superiority of the frequency
domain processing method over the time domain processing method and fusion time-
frequency processing method. Notably, the proposed model outperforms others, achieving
the highest accuracy of 94.93% in frequency domain processing-based experiments — 5.91%
and 4.25% higher than the highest accuracies in time domain processing-based experiments
and fusion time-frequency processing-based experiments respectively. The overall findings
illustrate that our proposal can provide a reliable reference for dementia diagnosis
effectively.
Keywords: Dementia diagnosis; Gait analysis; Power spectral density estimation; Attention
mechanism; Convolutional neural network; Gating mechanism

Michael A. Rothfuss, Jignesh V. Unadkat, Michael L. Gimbel, Marlin H. Mickle, Ervin Sejdić,
Totally Implantable Wireless Ultrasonic Doppler Blood Flowmeters: Toward Accurate
Miniaturized Chronic Monitors,
Ultrasound in Medicine & Biology,
Volume 43, Issue 3,
2017,
Pages 561-578,
ISSN 0301-5629,
[Link]
([Link]
Abstract: Totally implantable wireless ultrasonic blood flowmeters provide direct-access
chronic vessel monitoring in hard-to-reach places without using wired bedside monitors or
imaging equipment. Although wireless implantable Doppler devices are accurate for most
applications, device size and implant lifetime remain vastly underdeveloped. We review past
and current approaches to miniaturization and implant lifetime extension for wireless
implantable Doppler devices and propose approaches to reduce device size and maximize
implant lifetime for the next generation of devices. Additionally, we review current and past
approaches to accurate blood flow measurements. This review points toward relying on
increased levels of monolithic customization and integration to reduce size. Meanwhile,
recommendations to maximize implant lifetime should include alternative sources of power,
such as transcutaneous wireless power, that stand to extend lifetime indefinitely. Coupling
together the results will pave the way for ultra-miniaturized totally implantable wireless
blood flow monitors for truly chronic implantation.
Keywords: Battery-less; Blood flow monitor; Flowmeter; Free flap; Wireless power;
Transcutaneous wireless power

Jens Schwarz, Brian Hutsel, Thomas Awe, Bruno Bauer, Jacob Banasek, Eric Breden, Joe
Chen, Michael Cuneo, Katherine Chandler, Karen DeZetter, Mark Gilmore, Matthew Gomez,
Hannah Hasson, Maren Hatch, Nathan Hines, Trevor Hutchinson, Deanna Jaramillo, Christine
Kalogeras Loney, Ian Kern, Derek Lamppa, Diego Lucero, Larry Lucero, Keith LeChien, Mike
Mazarakis, Thomas Mulville, Robert Obregon, John Porter, Pablo Reyes, Alex Sarracino,
Daniel Scoglietti, Gabriel Shipley, Trevor Smith, Brian Stoltzfus, William Stygar, Adam Steiner,
David Yager-Elorriaga, Kevin Yates,
Mykonos: A pulsed power driver for science and innovation,
High Energy Density Physics,
Volume 53,
2024,
101144,
ISSN 1574-1818,
[Link]
([Link]
Abstract: Sandia National Laboratories has been operating the Mykonos linear transformer
driver (LTD) in a five-cavity configuration since 2014. The machine operates at 1MA output
current, 500kV output voltage, with a 10–90% current rise time of 85ns, which enables small
scale physics and engineering pulsed power experiments. Mykonos provides hands-on
pulsed power experimental training for students and staff alongside senior Sandia scientists
in an environment that is more accessible than the Z Facility. Over the years, we have fielded
and accumulated a wide variety of optical, x-ray and electrical diagnostics and we are
preparing to open this facility to outside users. Here, we are presenting the pulsed power
and diagnostic capability of Mykonos as well as some recent experiments that have been
performed on the facility. The goal of this publication is to attract researchers across the
pulsed power and high energy density (HED) community to collaborate with Sandia on
exciting, innovative science and to train the next generation of researchers for the National
Nuclear Security Agency (NNSA) and the nation. As such, we have established a Mykonos
Academic Access Program (MAAP) as part of ZNetUS to enable academic utilization of the
Mykonos Pulsed Power Facility.
Keywords: Pulsed power; ZNetUS; Linear transformer driver

Linqi Zhao, Pedro Cheong,


Open-set pedestrian identification via transformer-based neural architecture with extreme
value theory,
Engineering Applications of Artificial Intelligence,
Volume 164, Part A,
2026,
113253,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Real-world radio frequency (RF) sensing deployments frequently encounter
identities not observed during training. However, most existing methods are designed under
closed-set assumptions. To address this limitation, we introduce an open-set pedestrian
identification system that leverages an attention neural network (ANN) together with an
extreme value theory (EVT)-based algorithm. Initially, pedestrian features are captured via a
radio frequency identification (RFID) system and processed by a meticulously crafted ANN
with a transformer architecture to synthesize and analyze features from two-channel RFID
signals. To counteract the final logit layer’s tendency to generate low-dimensional
embeddings overly specific to known classes, we employ a decision-driven EVT approach on
intermediate layers to establish robust class-specific acceptance zones. Experimental results
demonstrate a closed-set identification accuracy of 99%, while maintaining over 90%
accuracy even when half of the test identities are unknown. These findings highlight a
practical RF-sensing solution for large-scale access control and privacy-preserving, ambient-
assisted Internet of Things (IoT) applications.
Keywords: Attention mechanism; Extreme value theory; Open-set recognition; Pedestrian
identification; Radio frequency identification

Derek Ka-Hei Lai, Li-Wen Zha, Tommy Yau-Nam Leung, Andy Yiu-Chau Tam, Bryan Pak-Hei So,
Hyo-Jung Lim, Daphne Sze Ki Cheung, Duo Wai-Chi Wong, James Chung-Wai Cheung,
Dual ultra-wideband (UWB) radar-based sleep posture recognition system: Towards
ubiquitous sleep monitoring,
Engineered Regeneration,
Volume 4, Issue 1,
2023,
Pages 36-43,
ISSN 2666-1381,
[Link]
([Link]
Abstract: Sleep posture monitoring is an essential assessment for obstructive sleep apnea
(OSA) patients. The objective of this study is to develop a machine learning-based sleep
posture recognition system using a dual ultra-wideband radar system. We collected
radiofrequency data from two radars positioned over and at the side of the bed for 16
patients performing four sleep postures (supine, left and right lateral, and prone). We
proposed and evaluated deep learning approaches that streamlined feature extraction and
classification, and the traditional machine learning approaches that involved different
combinations of feature extractors and classifiers. Our results showed that the dual radar
system performed better than either single radar. Predetermined statistical features with
random forest classifier yielded the best accuracy (0.887), which could be further improved
via an ablation study (0.938). Deep learning approach using transformer yielded accuracy of
0.713.
Keywords: Obstructive sleep apnea; Deep learning; Sleep monitoring; Feature extraction;
Ablation study

Dunlu Peng, Meiling Chen, Yiqin Zhang, Zekun Tian,


Enhanced optic-flow extrapolation for Doppler radar nowcasting with Dynamic Weight
Attention,
Expert Systems with Applications,
Volume 267,
2025,
126168,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Doppler radar echo extrapolation is an important method for extreme weather
forecasting. However, traditional optical flow methods lack learnable components and are
not suitable for complex atmospheric changes. Consequently, researchers have turned to
deep neural networks for prediction. Yet, the predictions from this approach often suffer
from issues such as mean reversion and a lack of small- and medium-scale structure. This
paper proposes a novel approach that combines optical flow methods with deep neural
networks. By introducing an artificially defined momentum weight matrix based on prior
assumptions, predictions for any future time distance are generated from full-scale optical
flow. Additionally, we propose a full-scale advection extractor, leveraging the continuity of
distribution in mesoscale and small-scale atmospheric systems and focusing on the long- and
short-distance relationships within the contour surface distribution sequence, which
improves the prediction accuracy of fine-scale advection. The experimental results show
that, compared with other advanced methods, the proposed method demonstrates
advantages in predicting extreme radar echoes and maintaining the echo structure.
Specifically, it achieved an improvement of 24.1% and 21.3% on key indicators such as CSI
and HSS, respectively, and reached 0.948 on the SSIM. Building on this, the inference speed
of our method is comparable to other deep learning approaches, being 3.25 times faster
than flow-based methods.
Keywords: Nowcasting; Optical flow; Neural network; Doppler radar

Zhongrui Bai, Fanglin Geng, Hao Zhang, Xianxiang Chen, Lidong Du, Peng Wang, Pang Wu,
Gang Cheng, Zhen Fang, Yirong Wu,
Non-contact blood pressure estimation using FMCW radar: A two-stream approach focused
on central arterial activity,
Biomedical Signal Processing and Control,
Volume 106,
2025,
107718,
ISSN 1746-8094,
[Link]
([Link]
Abstract: This paper proposes a radar-based two-stream blood pressure (BP) estimation
framework (R2S-BP), focusing on central arterial activity. It separately analyzes central-
arterial pulse transit time (caPTT) and pulse wave morphology using multi-location Doppler
Cardiogram (DCG) data from millimeter wave FMCW radar. Specifically, phase information at
harmonic heart rate frequencies is used to compute time delay arrays, representing caPTT-
related features. Additionally, k-Shape clustering is employed to select optimal DCGs from
the neck and chest regions that contain BP-related morphological features. These features
are processed through a two-stream neural network combining BiLSTM, ResNet, and multi-
head attention modules. Subject-independent 9-fold cross-validation results show that the
standard deviations of the errors for systolic and diastolic BP are 7.33 and 5.36 mmHg,
respectively. The intra-subject correlation coefficient for both systolic and diastolic BP
averages 0.82. Comparative and ablation studies demonstrate the superiority of the two-
stream approach and the critical importance of its components. This approach integrates
physiologically guided manual feature construction with a deep learning model, fully
leveraging the capabilities of FMCW radar data.
Keywords: Non-contact blood pressure estimation; Two-stream neural network; Doppler
Cardiogram; Central-artery pulse transit time

Akira Nakagawa, Helena Delmotte, Stefan Hallström, Yi Wan, Jun Takahashi,


Topology-constrained U-Net for delamination mapping in ballistic-damaged thin GFRP,
Composites Part B: Engineering,
Volume 311,
2026,
113240,
ISSN 1359-8368,
[Link]
([Link]
Abstract: Delamination severely degrades the residual strength of glass-fibre-reinforced
plastic (GFRP) laminates impacted by ballistic projectiles. This study introduces TopoPrior-
UNet, a U-Net variant named for its novel loss function that embeds Topological Prior
knowledge (specifically, a persistent-homology shape-prior loss to enforce the global
topological structure of the damage, and a level-set active-contour loss to refine fuzzy
boundaries) to segment multi-layer delamination images from a digital single lens reflex
camera (DSLR). An initial set of 24 raw panel images was expanded to a 480-image set
through three augmentation schemes, enabling learning under extreme data scarcity.
Compared with a vanilla U-Net, TopoPrior-UNet improved mean Intersection-over-Union to
0.9832 (+9.9pp) and mean Dice to 0.9915 (+5.3pp), achieving the best scores across six
metrics. These gains stem from jointly refining fuzzy boundaries and enforcing global ply
topology, as corroborated by an ablation study. The trained model operates in quasi-real-
time (i.e., inference in seconds per image) and requires only surface imagery, offering a cost-
effective and rapid alternative to traditional NDT methods like X-ray or ultrasonic inspection,
which often require minutes to hours for acquisition. Such rapid, non-destructive
quantification of delamination extent lays the foundation for on-site residual-strength
estimation and timely structural maintenance.
Keywords: GFRP delamination; TopoPrior-UNet; Semantic segmentation; Persistent
homology; Level-set active contour; Non-destructive testing

Aiguo Li, Bowen Li,


CFPNet: Multivariate time series classification based on frequency domain reconstruction,
Digital Signal Processing,
Volume 168, Part C,
2026,
105598,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Transformer-based deep learning models have significantly advanced the field of
multivariate time series classification. However, the intrinsic self-attention mechanism
renders existing Transformer methods prone to frequency bias and inadequate in extracting
local features, which ultimately limits their representational capacity. To address these
issues, we propose CFPNet, a novel network that enhances representation learning by
reconstructing a Crucial Frequency Patch in the frequency domain, thereby effectively
mitigating frequency bias. Additionally, we introduce the Wav-KAN encoder, which integrates
wavelet transforms with the Kolmogorov-Arnold Network to accurately capture local
dependencies. Extensive experiments on fourteen public datasets from the UEA(Multivariate
Time Series Classification Archive), as well as on a custom dataset of ultrasonic signals from
metallic materials, demonstrate that CFPNet achieves superior classification accuracy
compared to state-of-the-art methods.
Keywords: Multivariate time series classification; Transformer; Frequency domain feature
reconstruction; Kolmogorov–Arnold network

Marco Antonacci, Emanuele Riva, Attilio Frangi, Alberto Corigliano, Valentina Zega,
Planar GRIN lenses: Numerical modeling and experimental validation,
Journal of Sound and Vibration,
Volume 537,
2022,
117217,
ISSN 0022-460X,
[Link]
([Link]
Abstract: Phononic Crystal (PnC) Gradient Index (GRIN) lenses have been intensively
investigated in recent years for their promising applications in energy harvesting. Here we
propose and verify, both numerically and experimentally, three designs of PnC GRIN lenses
with amplification factors at the focal points of 4.33x, 7.35x and 7.49x. A design procedure
based on the combination of two mechanisms of refraction index gradient formation is
employed to boost the efficiency of the lens. Moreover, thanks to the planarity of the design
and to the single-phase constitutive material, the proposed lenses are fully compatible with
microfabrication processes and can be therefore employed as micro energy harvesters after
proper miniaturization.
Keywords: GRIN lenses; Phononic crystals; Energy harvesting; Numerical modeling;
Experiments

You might also like