Transformer-Based Radar Signal Processing
Transformer-Based Radar Signal Processing
Yiming Zhang, Jiaqi Li, Zheng Tong, Weiguang Zhang, Xiyuan Shen,
A direction-aware and expert-inspired network for internal crack size detection using on-site
ground penetrating radar data,
Engineering Applications of Artificial Intelligence,
Volume 165, Part A,
2026,
113414,
ISSN 0952-1976,
[Link]
([Link]
Abstract: The size of internal cracks is a key basis for determining maintenance measures.
Existing methods primarily utilize Ground Penetrating Radar (GPR) signals or images to
detect crack size, but still face two challenges: limited robustness in interpreting on-site data
based on GPR signals and the inability to directly characterize the crack size based on raw
GPR image features. To address these issues, this study has proposed a novel internal crack
size detection network, which was trained by a dataset with on-site GPR B-Scans and
interpreted crack size labels. In the proposed model, a deformable Cross Stage Partial (CSP)
block is first used to extract the irregular hyperbolic features of crack reflected waves from
B-Scans. Then, a directional fusion attention module is designed to construct direction-
aware channel attention and generate spatial interaction weights. Finally, a bipartite graph
matching detection head is proposed to emulate the expert behavior to analyze crack
reflected waves from a global B-Scan perspective, outputting trapezoidal-sized boxes to
detect internal crack size. The experimental results demonstrate that the proposed model
exceeds other state-of-the-art models on the tasks thanks to the channel-spatial weight
aggregation and global output strategy via bipartite graph matching. Additionally, the model
exhibits good stability across various antenna frequencies and pavement structures. The on-
site testing indicates that the predicted crack sizes sufficiently meet engineering
requirements in most scenarios, though challenges remain in detecting the bottom width of
small and water-saturated cracks.
Keywords: Asphalt pavement; Non-destructive testing; Ground penetrating radar; Internal
crack; Size detection
Long He, Kun Zheng, Huihua Ruan, Shuo Yang, Jinbiao Zhang, Cong Luo, Siyu Tang, Yunlei Yi,
Yugang Tian, Jianmei Cheng,
A spatiotemporal mixed-enhanced generative adversarial network for radar-based
precipitation nowcasting,
Computers & Geosciences,
Volume 200,
2025,
105919,
ISSN 0098-3004,
[Link]
([Link]
Abstract: Skillful precipitation nowcasting with high resolution and detailed information
holds promise for providing reliable alerts about severe weather events to society. Radar
echo extrapolation is an essential method for precipitation nowcasting, but traditional
methods struggle to capture rapidly changing regions. Deep learning (DL)-based methods
exhibit superior performance. However, existing DL-based methods face challenges such as
low accuracy, particularly in producing clear forecasts over longer lead times and accurately
forecasting moderate to heavy rainfall events. To address these challenges, we developed a
novel radar-based precipitation nowcasting model, STMixGAN, which can be described as a
nonlinear proximity forecasting model. This model effectively aggregates global-to-local
information and imposes constraints to represent the complex evolution of rainfall
efficiently. Consequently, STMixGAN produces realistic and spatiotemporally consistent
predictions. Using radar observations from South China, STMixGAN successfully forecasted
radar maps for the next 1 h using 24 min of input data. Two traditional methods (Persistence
and Optical flow) and five DL-based methods (ConvLSTM, Rainformer, IAM4VP, REMNet, and
GAN-argcPredNet) were employed as benchmarks to validate STMixGAN’s forecasting
capabilities. The experimental results demonstrate STMixGAN’s superior performance and
provide valuable insights for enhancing heavy rainfall forecasting.
Keywords: Precipitation nowcasting; Spatiotemporal mixed enhancement; Generative
adversarial networks; Self-attention
Wenju Zhao, Shijie Xing, Futao Ni, Yongding Tian, Qiang Liu,
FMCW radar-based high-precision range estimation with generalized eigenvalue
decomposition algorithm,
Measurement,
Volume 255,
2025,
117957,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Frequency-modulated Continuous Wave (FMCW) radar has been widely applied in
defense systems, automotive collision avoidance, intelligent traffic monitoring, and precision
level measurement due to its excellent long-range detection capabilities and improved range
resolution. However, traditional FMCW radar techniques are inherently prone to spectral
leakage artifacts and picket fence effects caused by non-integer period sampling and limited
frequency resolution of discrete Fourier transform (DFT) processing, which seriously
compromise ranging accuracy. To overcome these challenges, this paper proposes a novel
high-precision range estimation algorithm based on generalized eigenvalue decomposition
(GEVD). The key contributions of this study are: (1) the derivation of the analytical
relationship between generalized eigenvalues and beat frequencies through formulating a
generalized eigenvalue equation that includes both the beat frequency signal and its first-
order derivative; and (2) the effective reduction of discretization errors and noise
interference using frequency-shifting techniques combined with singular value
decomposition (SVD)-based signal enhancement. A thorough parametric analysis has been
performed to evaluate the impact of sampling frequency, matrix dimension, signal-to-noise
ratio, and the number of targets on range precision. Extensive numerical simulations and
controlled laboratory experiments validate the theoretical framework and operational
effectiveness of the proposed methodology. Comparative results demonstrate that the
GEVD-based method achieves superior resolution compared to conventional techniques,
even in noisy environments. Field validation using single-target measurement trials confirms
exceptional measurement stability, with empirical data showing maximum range deviation
within 0.2 mm at sampling frequencies exceeding 400 kHz.
Keywords: Frequency-modulated continuous wave radar; Range estimation; Generalized
eigenvalue decomposition; Civil engineering
Wei Quan, Wenjing Cheng, Yike Yang, Haiquan Zhao, Zhaoyu Chen, Yunfan Luo,
A signal fingerprint feature extraction method based on decomposition and fusion for radar
emitter individual identification,
Digital Signal Processing,
Volume 164,
2025,
105257,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Radar emitter individual identification is one of the key technologies of modern
electronic countermeasure reconnaissance and electronic intelligence. With the
advancement of radar technology and the increasingly complex electromagnetic
environment, existing methods for identifying emitter are gradually becoming unable to
meet the performance requirements of modern radar individual identification. Aiming at
improving the adaptability of feature extraction for non-cooperative radar emitter signals
and the robustness of individual identification in the complex modern electronic warfare
environment, a signal fingerprint feature extraction method based on decomposition and
fusion is proposed. It firstly integrates signal decomposition and scattering convolution
networks (SCN) to adaptively extract the multi-scale intra-pulse feature of the signal, while
removing the potential noise of the redundant component by energy proportion. And then a
deep feature fusion model based on multi-head self-attention and residual connection is
proposed to fuse the multi-scale features and the time domain features to further extract
signal fingerprint of radar emitter. Experimental results based on the real radar emitter
signals demonstrate that the identification method proposed in this paper can more
effectively extract signal fingerprint features and the identification accuracy reaches 96.45%,
which outperforms other existing identification methods.
Keywords: Radar emitter individual identification; Signal fingerprint feature; Signal
decomposition; Scattering convolution networks (SCN); Fusion
Haoyuan Ding, Tingsong Zhang, Zhangting Wang, Guangran Bai, Liujun Han, Yujia Dai, Ziyuan
Liu,
Classification and quantification of sodium metabisulfite in goji berry powder: Applications
of hyperspectral technology and transformer-based hybrid models,
Spectrochimica Acta Part A: Molecular and Biomolecular Spectroscopy,
Volume 349,
2026,
127379,
ISSN 1386-1425,
[Link]
([Link]
Abstract: Sodium metabisulfite (Na2S2O5) is widely used as an antioxidant and preservative
in food products, but excessive residues pose health risks and are strictly regulated. Rapid
and non-destructive detection methods are therefore essential, particularly for goji berry
powder where sulfite addition is common during processing. In this study, near-infrared
hyperspectral imaging (HSI, 900–1700 nm) was employed in combination with advanced
deep learning models to classify and quantify sodium metabisulfite concentrations. A total
of 360 samples (nine concentration levels including control, 40 replicates each) were
prepared, with 70 % allocated for training and 30 % for testing. For classification, a hybrid
ResLocalformer model that integrates local attention and residual paths within a
Transformer framework achieved an accuracy of 97.22 %. For regression, the Resformer
model, combining Transformer's global attention with dense layers, yielded an R2 of 0.9945
and RMSE of 0.0343. Spectral preprocessing using first-derivative Gaussian smoothing
(1DER-GS) significantly enhanced feature quality, improving overall model performance by
more than 40 %. Compared with conventional approaches such as CNN and LSTM, the
proposed models demonstrated superior robustness and predictive accuracy. These results
indicate that HSI combined with transformer-based hybrid models provides an effective and
non-destructive approach for monitoring sodium metabisulfite in goji berry powder, with
strong potential for extension to the detection of other food additives and matrices.
Keywords: Hyperspectral imaging; Deep learning; Classification; Quantification; Goji berry;
Transformer-Based Hybrid Models.
Siyu Chen, Xiaoyan Zhang, Hongjun Xue, Xiang Fang, Xueren Li,
Knowledge-guided graph transformer for gaze-based pilot operation recognition in dynamic
flight operations,
Aerospace Science and Technology,
Volume 171,
2026,
111610,
ISSN 1270-9638,
[Link]
([Link]
Abstract: Real-time and accurate recognition of pilots’ operational intent in dynamic flight
tasks is critical for enhancing aviation safety, preventing human error, enabling intelligent
decision support, and supporting online human-reliability assessment. Eye-tracking data, an
objective physiological signal obtainable in real time in the cockpit, provides crucial evidence
for inferring intent. However, the volatility of time-series gaze data and its tight, nonlinear
coupling with complex operational responses limit the accuracy of approaches that rely on
gaze features alone. Methods based on physical parameters capture action outcomes rather
than intent, while purely data-driven models often generalize poorly and offer limited
interpretability due to the absence of task-logic guidance. This study introduces a
Knowledge-enhanced Graph Transformer Network (KGTN) that integrates structured domain
knowledge as a computable guidance signal with dynamic eye-movement behavior, and
explicitly models their nonlinear interactions. A task knowledge graph grounded in a
cognitive-activity taxonomy is constructed and encoded by a Relational Graph Neural
Network, alongside a multi-scale Patch-Transformer and a Hierarchical Task-Context Aware
Operation Recognition Module (HTCA-ORM) for knowledge-guided fusion and precise
operation recognition. On a high-fidelity flight-simulation dataset, KGTN attains 87.63%
accuracy and 87.48% weighted F1, outperforming a range of baselines. Ablation and
interpretability analyses indicate strong potential for high-precision, reliable aviation human-
factors analytics.
Keywords: Human-machine systems; Intelligent cockpit; Aviation human factors; Eye
tracking; Cognitive activity modeling
Tiantian Wang, Nan Yan, Chaosan Yang, Zeliang An, Gongjing Zhang, Yuqing Xu,
Electromagnetic signal recognition using multimodal tri-branch semantic fusion network in
the UAV-assist integrated sensing and communication systems,
Digital Signal Processing,
Volume 171,
2026,
105820,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Driven by the proliferation of integrated sensing and communication (ISAC)
systems, the accurate recognition of unauthorized unmanned aerial vehicle (UAV) signals in
dynamic electromagnetic environments has emerged as a critical challenge for spectrum
security and cognitive radio applications. Conventional automatic modulation recognition
(AMR) frameworks suffer from significant performance degradation in low signal-to-noise
ratio (SNR) regimes and exhibit limited adaptability to resource-constrained edge computing
platforms. To address these limitations, we propose a novel Multimodal Tri-branch Fusion
Network (MTF-Net) architecture that synergistically integrates time-frequency analysis with
statistical feature learning. The framework systematically processes binarized time-
frequency images (B-TFIs) and higher-order cumulant vectors through three collaboratively
operating branches: (1) A primary temporal feature extractor employing dilated convolution-
residual blocks (DCRBlocks) with hierarchical dilatation factors, incorporating channel
attention mechanisms to dynamically emphasize discriminative temporal patterns; (2) Dual
auxiliary branches based on Edge-Transformer modules (ETFormers), which achieve efficient
spatial-structural learning through depthwise separable convolutions (DSC) while capturing
long-range spectral dependencies via additive attention mechanisms with linear complexity;
(3) A hierarchical fusion module implementing cross-branch feature recalibration through
learnable parameter matrices. Extensive Monte Carlo experiments demonstrate that our
MTF-Net significantly outperforms traditional methods in recognition accuracy for radar and
communication signals under low SNR conditions, establishing a new benchmark for
lightweight AMR solutions in ISAC systems.
Keywords: Multi-modal feature fusion; Unmanned aerial vehicle(UAV); Integrated sensing
and communication (ISAC); Lightweight neural network; Transformer
Sidra Ghayour Bhatti, Imtiaz Ahmad Taj, Mohsin Ullah, Aamer Iqbal Bhatti,
Transformer-based models for intrapulse modulation recognition of radar waveforms,
Engineering Applications of Artificial Intelligence,
Volume 136, Part B,
2024,
108989,
ISSN 0952-1976,
[Link]
([Link]
Abstract: The increasing prevalence of low probability of intercept (LPI) radars in electronic
warfare (EW) systems highlights the need to effectively recognize phase-coded radar
waveforms intercepted at radar warning receivers (RWRs) from various threat emitters. The
complexities of the electromagnetic (EM) spectrum necessitate the implementation of an
automatic modulation recognition system (AMRS) within the RWR. However, a major
challenge is accurately identifying phase-coded waveforms with high accuracy at low signal-
to-noise ratios (SNRs). This research addresses the challenge by exploring three artificial
intelligence (AI)-driven AMRS architectures for identifying phase-coded waveforms using
short-time Fourier transform (STFT): vision transformer (ViT), vicinity vision transformer
(VViT), and deep convolutional neural network (DCNN). Unlike recent methods focusing on
amplitude spectra, our research delves into the phase spectra for the feature extraction of
phase-coded waveforms. We leverage phase-based features extracted from intercepted
phase-coded waveforms to classify six types of phase-coded signals using these AMRS
architectures across SNR levels ranging from −16 dB to 8 dB. The simulation experiments
show that these methods are effective at an SNR of −16 dB, with VViT and ViT achieving
recognition accuracies of 93% and 92.7%, respectively. Both outperform the DCNN, which
achieves an RA of 89% at the same SNR. This approach promises to enhance situational
awareness and decision-making in EW operations by improving phase-coded radar
waveform recognition and enabling appropriate countermeasure deployment.
Keywords: Automatic modulation recognition system; Feature extraction; Low probability of
intercept; Short time Fourier transform
Shuai Guo, Ting Chen, Penghui Wang, Jun Ding, Junkun Yan, Hongwei Liu,
Knowledge embedding fusion based on language model for enhanced radar target
recognition,
Signal Processing,
Volume 238,
2026,
110199,
ISSN 0165-1684,
[Link]
([Link]
Abstract: Traditional radar target recognition methods typically model only single echoes,
neglecting the crucial information that domain knowledge can provide for understanding
data. In this paper, we propose a knowledge embedding fusion (KEF) method for enhanced
high-resolution range profile (HRRP) recognition, which utilizes the target state descriptions
available during radar detection. KEF leverages a language model (LM) to integrate textual
knowledge with echo features for fusion recognition. It consists of three components: HRRP
feature extraction, measurement-based knowledge construction, and knowledge embedding
fusion module. First, we perform feature extraction on the HRRP to obtain echo tokens.
Next, in the knowledge construction module, the measurement statuses are standardized to
a natural language format, and the LM is utilized to extract semantic information, resulting in
text tokens. Finally, in the knowledge embedding fusion module, a cross-attention HRRP-text
fusion strategy is employed to facilitate interaction between echo tokens and textual tokens.
We also design a combination of HRRP-text matching loss and fusion classification loss to
guide model training. Experiments are conducted on a real measured dataset, and the
results indicate that KEF effectively enhances recognition performance across multiple
scenarios compared with approaches that only utilize echoes.
Keywords: High-resolution range profile (HRRP); Knowledge embedding fusion; Language
model (LM); Radar target recognition
Mingyue Lu, Menglong Wang, Qian Zhang, Manzhu Yu, Caifen He, Yadong Zhang, Yuchen Li,
A vision transformer for lightning intensity estimation using 3D weather radar,
Science of The Total Environment,
Volume 853,
2022,
158496,
ISSN 0048-9697,
[Link]
([Link]
Abstract: Lightning has strong destructive powers; its blast wave, high temperature, and high
voltage can pose a great threat to human production, life, and personal safety. The
destructive power of high-intensity lightning is much greater than that of low-intensity
lightning. The estimation of lightning intensity can provide an important reference for
determining the lightning protection level and lightning disaster risk assessment. Lightning is
a type of small-scale severe convective weather phenomenon. Weather radar is one of the
best monitoring systems that can frequently sample the detailed three-dimensional (3D)
structures of convective storms, with a small spatial scale and short lifetime at high temporal
and spatial resolutions. Therefore, it is possible to extract the 3D spatial feature strongly
correlated with lightning from 3D weather radar for estimating lightning intensity. This paper
proposes a Vision Transformer model for lightning intensity estimation that can
automatically estimate lightning intensity from 3D weather radar data. In an experiment, we
transferred the task of estimating lightning intensity into a multicategory classification task.
A framework was designed to produce lightning feature samples for model input from 3D
weather radar and lightning location data. Then, the Synthetic Minority Over-Sampling
Technique (SMOTE) algorithm was used to balance and optimize the sample distribution.
Finally, samples were input into the proposed lightning intensity estimation model based on
Vision Transformer for training and evaluation. Experimental results show that the proposed
model based on Vision Transformers performs well with lightning intensity estimation.
Keywords: Lightning intensity estimation; 3D weather radar; Vision transformer; SMOTE;
Multicategory classification
Shenghua Lv, Xiaowei Zhang, Xuan Zhao, Meng Li, Jianghao Zhang, Chen Lin, Jian Wen,
Rapid and accurate assessment of filed scale soil moisture using ground-penetrating radar
deep learning-based inversion,
Measurement,
2026,
120594,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Accurate quantification of soil moisture content (SMC) is essential for sustaining
plant growth and maintaining ecosystem stability. Current SMC monitoring approaches are
subject to several constraints: satellite-based remote sensing frequently suffers from
inadequate spatial resolution for large-scale precision, while point-scale techniques are
incapable of effectively capturing soil moisture spatiotemporal variations and are unsuitable
for large-area monitoring. Ground Penetrating Radar (GPR), as an efficient and non-
destructive subsurface detection technique, has been widely employed for estimating soil
water content. However, existing methods based on full-waveform inversion exhibit strong
dependency and involve computationally expensive processes, resulting in insufficient
efficiency when handling large-scale GPR data. To overcome these challenges, this study
proposes a GPR deep learning-based inversion framework for rapid and precise SMC
estimation. A 900 MHz GPR system was deployed to survey an experimental site equipped
with pre-installed moisture sensors. The results demonstrated strong agreement between
GPR-derived SMC values and measurements obtained via Time Domain Reflectometry (TDR).
Furthermore, continuous high-temporal-resolution data collection verified the capability of
the method to characterize the spatiotemporal dynamics of SMC. Notably, for regional soil
moisture content (RSMC) estimation, the proposed method achieved a substantially lower
error (0.0019 m3) compared to conventional point-based measurements (0.0146 m3) and
stratified estimation (0.0122 m3). This methodology provides a robust technical foundation
for accurate field-scale SMC monitoring and exhibits significant potential for use in ecological
surveillance, precision agriculture irrigation, and sustainable water resource management.
Keywords: Ground-penetrating radar; Non-destructive testing; Soil moisture content; Deep
learning-based inversion; Spatio-temporal evolution
Zhuoning Hao, Shaojuan Luo, Wei Meng, Huapan Xiao, Heng Wu, Chunhua He,
SAR ship detection in complex coastal areas: A multi-scale residual fusion transformer with
spatial-channel attention for high detection accuracy,
Measurement,
Volume 264,
2026,
120257,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Synthetic Aperture Radar (SAR) provides critical all-weather measurement
capabilities for microwave imaging in remote sensing. Accurate SAR ship detection is pivotal
for measurement-driven applications such as maritime surveillance, where reliable
identification and localization of ships are essential. Current SAR target detection methods
enable accurate ship detection in open waters but face challenges in detecting ships near
complex coastlines and small targets. To address these problems, we propose a SAR ship
detection network based on a multi-scale query refinement transformer and spatial-channel
attention fusion module, MRFSA-Net. Specifically, a lightweight backbone network with a
SAR enhancement module is designed to reduce model complexity while preserving feature
extraction capabilities, simultaneously enhancing multi-scale feature extraction and feature
fusion. A spatial-channel attention fusion module is developed to strengthen contextual
awareness and suppress background interference. Cross-scale interaction branches are
employed to achieve effective fusion of information across different receptive fields.
Additionally, a transformer with biased attention and query optimization mechanisms is
constructed to enhance the robustness of multi-scale object detection. Quantitative and
qualitative experimental results show that MRFSA-Net achieves competitive performance
compared to the previous state-of-the-art object detection methods, significantly improving
detection accuracy in complex coastal environments.
Keywords: Synthetic Aperture Radar; Deep Learning; Object Detection
Kangle Song, Jingbin Li, Yang Li, Jing Nie, Yuntao Sun, Wujun Zhang, Xiaojie Hou, Qiang
Wang, Pengxiang Song,
Long-distance soil moisture monitoring via Helmholtz resonator–enhanced acoustic
transmission and Swin-Transformer modeling,
Computers and Electronics in Agriculture,
Volume 242,
2026,
111344,
ISSN 0168-1699,
[Link]
([Link]
Abstract: To address the severe energy attenuation of acoustic waves during propagation
through soil, which restricts signal reception and modeling in large scale soil moisture
monitoring, this study designs a resonance enhancement structure to overcome the short
detection range of conventional acoustic methods and proposes a large-scale transmission
method for soil moisture detection based on a Helmholtz resonator. The response
characteristics and consistency of the Helmholtz resonator under identical excitation
conditions were first evaluated. The enhancement of acoustic signal reception and the
extension of detection capability were then verified by comparing system responses over
various propagation distances. Furthermore, the influence of excitation periodicity on
resonance behavior was analyzed, leading to optimized excitation parameters for improved
system performance. A mapping model was subsequently constructed between acoustic
time–frequency spectrograms and soil water content through feature extraction and model
development, with a visualization mechanism introduced to interpret model decision
making processes. System validation was conducted under real field conditions.
Experimental results demonstrated that the Helmholtz resonator exhibited highly consistent
response characteristics and strong frequency selectivity. The effective detection range of
the system was extended from 4 m to 60 m, significantly enhancing the reception of acoustic
signals in large-scale soil environments. In the test dataset, the Swin-Transformer–based
regression model achieved a mean absolute error (MAE) of 0.207 %, root mean square error
(RMSE) of 0.244 %, and coefficient of determination (R2) of 0.992. In field trials, the
corresponding metrics were 1.046 %, 0.851 %, and 0.735, respectively, confirming the
effectiveness of the feature extraction and modeling approach. This study demonstrates that
the proposed Helmholtz resonator–based method enables effective large-scale soil moisture
detection and offers a novel, low cost, transmission-based solution for wide-area soil
moisture monitoring and water management.
Keywords: Soil moisture monitoring; Helmholtz resonator; Acoustic wave; EMD; Swin-
transformer
Xiaosong Tang, Feng Yang, Xu Qiao, Jialin Liu, Haitao Zuo, Liang Gao, Jianshe Zhao, Suping
Peng,
GPR-HIDiff: A diffusion-based model for horizontal interference suppression in urban
underground detection radar profiles,
Underground Space,
Volume 26,
2026,
Pages 458-478,
ISSN 2467-9674,
[Link]
([Link]
Abstract: Automated subsurface utility detection systems in construction rely heavily on the
quality of ground-penetrating radar (GPR) profiles, which are often degraded by high-
amplitude horizontal interference. Existing low-rank decomposition methods lack the
intelligence and flexibility required for multi-site data processing and involve labor-intensive
parameter tuning, impeding their integration into intelligent construction workflows. To
address these challenges, this paper proposes a horizontal interference suppression
algorithm based on a diffusion model, termed GPR-HIDiff. The proposed model replaces
conventional sequential convolutional operators with ResBlocks throughout the encoder,
intermediate layer, and decoder of the UNet architecture, enhancing training stability.
Lightweight agent attention modules are embedded between ResBlocks at each level to
improve global information modeling capability. A spatial attention mechanism is deployed
between the encoder and decoder to achieve adaptive spatial feature optimization.
Furthermore, the forward diffusion phase adopts a cosθ schedule-based strategy to ensure a
smooth temporal variation of noise variance. A standardized dataset comprising real-world
measured samples and finite difference time domain simulation samples of urban road
models has also been constructed. The effectiveness of the hybrid dataset, the introduced
modules, the robustness analysis, and the cosθ schedule is validated through training with
single/mixed datasets, ablation studies, evaluation of metric variations before and after the
introduction of different noise levels, and comparative experiments with constant, linear,
and cosθ schedules. Experimental results demonstrate that GPR-HIDiff significantly
outperforms both traditional methods and state-of-the-art deep learning models on both
simulated and real-world test samples. It effectively suppresses horizontal artifacts,
preserves target hyperbolic contours, and avoids excessive reduction of target scattering,
showcasing its exceptional performance. This method provides a powerful algorithmic
foundation for high-resolution GPR imaging and target detection.
Keywords: Ground-penetrating radar; Horizontal interference; Diffusion model; Agent
attention module; Spatial attention; Hybrid dataset
Zhiyan Lin, Minming Gu, Keyu Pan, Wei-Ping Zhu,
Adaptive temporal convolutional network with multi-head EMA-gated attention for
continuous radar-based human activity recognition,
Biomedical Signal Processing and Control,
Volume 117,
2026,
109667,
ISSN 1746-8094,
[Link]
([Link]
Abstract: Continuous human activity recognition (HAR) using radar signals offers strong
potential for privacy-preserving clinical health monitoring. However, its performance is
limited by challenges such as multi-scale temporal variation, signal noise, and unstable
activity transitions. To address these issues, this study introduces a radar-based HAR
framework with three tailored components. First, an adaptive temporal convolutional
network (ATCN) uses learnable dilation rates and sampling offsets to flexibly capture both
abrupt and periodic motion patterns over time. Second, an exponential moving average
(EMA)-gated attention (EDGA) module integrates linear attention with exponential moving
average smoothing through a dynamic gating mechanism, effectively suppressing noise
while preserving temporal continuity. Third, an attention-guided multi-stage refinement
(AMSR) module refines coarse predictions using global attention-driven residual corrections,
thereby reducing segmentation noise and improving boundary precision. Experiments on a
77 GHz frequency modulated continuous wave (FMCW) radar dataset show that the
proposed model achieves 96.09% accuracy, demonstrating its strong potential for
continuous and unobtrusive activity monitoring in healthcare applications.
Keywords: ATCN; EDGA; AMSR; Continuous HAR; FMCW radar
Lijie Yang, Yu Wang, Zhaohui Yang, Tongkai Xu, Lizhi Dang, Zhongyue Chen, Weipeng Mao,
mmPPT: Hierarchical-serialization-enhanced point transformer for mmWave pedestrian
reconstruction,
Information Fusion,
Volume 127, Part B,
2026,
103835,
ISSN 1566-2535,
[Link]
([Link]
Abstract: In autonomous driving, redundant perception is critical for robust environmental
understanding, especially under adverse conditions where traditional sensors like RGB-D fail.
This paper introduces a novel Point Transformer for mmWave-radar-based pedestrian 3D
reconstruction (mmPPT), addressing the challenges of sparse, noisy 4D mmWave radar point
clouds. mmPPT employs a novel Hierarchical Serialization Strategy that combines Hybrid
Local Serialization to preserve fine-grained body details via adaptive space-filling curves and
Anchor-based Global Serialization to model long-range structural relationships. Additionally,
it integrates temporal information and Doppler velocity features to leverage radar’s unique
capabilities, while a Bone Length Loss is adopted to enforce geometric constraints on joint
position prediction. Experiments on 4D radar dataset demonstrate that mmPPT outperforms
state-of-the-art methods, achieving the lowest total reconstruction error (12.76 cm with
multi-frame input) and superior robustness in harsh environments like smoke, occlusion, and
poor lighting scenarios where RGB-D sensors fail. Ablation studies validate the effectiveness
of the Hybrid Local Serialization Strategy and the Anchor-based Global Serialization Strategy,
highlighting mmPPT’s potential to enhance redundant perception in autonomous driving.
The source code will be available at [Link]
Keywords: Redundant perception; 4D mmWave radar; 3D reconstruction; Hierarchical
serialization strategy; Hybrid local serialization; Anchor-based global serialization; Temporal
and velocity embedding; Bone length loss
Chaofeng Huang, Xiaowo Xu, Fan Fan, Shunjun Wei, Xiaoling Zhang, Dongmei Liu, Min Gu,
A low-SNR-adaptive temporal network with smart mask attention for radar signal
modulation recognition,
Digital Signal Processing,
Volume 168, Part D,
2026,
105640,
ISSN 1051-2004,
[Link]
([Link]
Abstract: The automatic modulation recognition of radar signals is a key technology in
electronic warfare and communication systems. However, traditional handcrafted features
often struggle to achieve high recognition accuracy under low signal-to-noise ratio (SNR)
conditions. With the rapid development of artificial intelligence technologies, deep learning-
based approaches have emerged as a promising alternative for modulation recognition. In
this article, a low-SNR-adaptive network architecture is proposed, which integrates a
bidirectional temporal convolutional network (Bi-TCN) and dual-channel smart mask
attention (DSMA) modules. The DSMA adaptively highlights informative features and
suppresses noise through complementary attention masks, enhancing robustness in low-SNR
conditions. Experimental results demonstrate that the autocorrelation domain outperforms
both time and frequency domains, with recognition accuracy improvements of 13.33 % and
14.71 %, respectively. Compared to state-of-the-art models, the proposed network achieves
63 % accuracy at -20 dB and more than 99 % accuracy at -6 dB, significantly enhancing radar
signal modulation recognition.
Keywords: Modulation recognition; Deep learning; Radar signal analysis,
Jiaquan Wan, Junchao Wang, Wei Zhang, Hao Song, Congyi Nai, Fengchang Xue, Tao Yang,
Chunxiang Shi, Quan J. Wang, Baoxiang Pan,
RadarDiT: An advanced radar echo extrapolation model for three gorges reservoir area via
diffusion transformer,
Journal of Hydrology: Regional Studies,
Volume 61,
2025,
102703,
ISSN 2214-5818,
[Link]
([Link]
Abstract: Study region
The Three Gorges Reservoir Area (TGRA)
Study focus
TGRA faces increasing vulnerability to extreme precipitation events driven by complex
convective weather systems. Radar echo extrapolation—predicting future precipitation
patterns from current radar data—is essential for early warning systems but faces significant
challenges in this topographically complex region. While data-driven approaches have
advanced the field, current convolutional neural network-based diffusion models struggle
with the TGRA's dynamic meteorological conditions due to their reliance on translational
invariance, which often fails to capture rapid weather transitions in complex terrain.
New hydrogeological insights from the region
To address these limitations, we introduce RadarDiT, a Vision Transformer-based diffusion
model specifically engineered for radar extrapolation in the TGRA. First, we develop a five-
year radar dataset capturing diverse convective weather phenomena unique to this region.
Then, leveraging this dataset, RadarDiT employs multi-layer Vision Transformers that
effectively model global dependencies and complex spatial relationships, enabling accurate
prediction of convective cell evolution. Our model demonstrates superior performance in
maintaining strong echo and spatial coherence over longer forecast horizons. Quantitative
evaluations across multiple metrics and thresholds confirm RadarDiT's enhanced skill in
forecasting heavy precipitation events, with particular improvements in Critical Success
Index at higher radar echo values. This work establishes a foundation for more reliable
nowcasting systems in regions with complex terrain and dynamic weather patterns, directly
supporting enhanced disaster preparedness and response strategies.
Keywords: Radar Echo Extrapolation; Three Gorges Reservoir Area; Diffusion Model; Vision
Transformer; Nowcasting
Xiaole Han, Jintao Liu, Jian Ye, Zihe Wang, Pengfei Wu, Hai Yang,
Deep Learning-Based GPR interpretation of soil thickness in headwater hillslopes,
Geoderma,
Volume 462,
2025,
117530,
ISSN 0016-7061,
[Link]
([Link]
Abstract: Soil thickness strongly influences eco-hydrological and geomorphic processes, yet
conventional measurements such as auger drilling are invasive, labor-intensive, and
unsuitable for large-scale surveys. Ground-penetrating radar (GPR) provides a non-invasive
alternative, but its manual interpretation remains slow and prone to observer bias. To
address this challenge, we developed a fully automated framework that couples a hybrid
CNN-Transformer deep learning architecture with optimized signal filtering to predict soil
thickness directly from GPR profiles. The convolutional layers extract local waveform
features, while the attention mechanism captures long-range dependencies. Using field data
from a steep headwater hillslope (H1) in the Taihu Basin, China, we compared five filtering
strategies—median, Savitzky-Golay, Gaussian, moving average, and none—and found that
median filtering yielded the most accurate results (R2 up to 0.92, CCC of 0.96, RMSE near
10 cm). We further identified optimal filter window sizes (61–101 samples) and a training
duration threshold (≥500 epochs) that ensured stable and accurate predictions. Cross-site
validation on an independent hillslope (H2) without retraining showed that the pretrained
CNN-Transformer model achieved the highest R2 (0.80), CCC (0.89), and lowest RMSE
(11.3 cm), outperforming traditional machine learning models (CNN, MLP, RF, SVM) in
transferability. These findings demonstrate that integrating CNN-Transformer architectures
with appropriate signal filtering enables scalable, accurate, and objective soil thickness
mapping in complex terrain. The proposed approach also holds promise for broader GPR-
based subsurface applications, including soil horizon delineation and root system detection.
Keywords: Ground-penetrating radar; Soil thickness; Transformer; Headwater hillslopes;
Median filtering
Ruizhe Feng, Shuzhao Zhu, Ruixin Jiang, Xin Cai, Rui Lin,
Degradation prediction of the low-Pt loading proton exchange membrane fuel cell based on
spatio-temporal Transformer network,
Energy,
Volume 342,
2026,
139625,
ISSN 0360-5442,
[Link]
([Link]
Abstract: Owing to the high energy density and low pollutant emissions, proton exchange
membrane fuel cells (PEMFCs) have emerged as promising solutions for sustainable energy
applications. Low-Pt loading proton exchange membrane fuel cells (low-Pt PEMFCs) hold
great promise in the field of sustainable energy for reducing the use of precious metals and
enhancing cost-efficiency. However, the long-term operational stability of low-Pt PEMFCs
remains a major barrier to large-scale commercialization. Accurate prediction of the
degradation process is essential for achieving system health management and ensuring
operational reliability. Under dynamic operating conditions, the increased nonlinearity and
uncertainty in the degradation process of low-Pt PEMFCs make accurate prediction more
difficult. To address these issues, a novel Transformer model named Adaptive Cross-
Dimensional Transformer (ACD-Transformer) is proposed. The model integrates a Dual
Deformation Attention Block (DDAB) and an Adaptive Filtering Block (AFB), effectively
extracting spatio-temporal dependencies in time-series data through deformable attention
mechanisms and learnable filtering thresholds. To evaluate the effectiveness and
generalizability of the proposed model, two degradation tests were designed for low-Pt
PEMFCs: Steady-State Conditions (SSC) and Multi-Temperature Region Conditions (MTRC),
the latter simulating real-world onboard temperature distributions. The results show that
the ACD-Transformer achieves superior performance on low-Pt PEMFCs compared to the
baseline models. Compared to the LSTM model, MAE is reduced by 63.64 % and 47.62 %
across two datasets. This work contributes to the development of prognostic methodologies,
facilitating effective health monitoring and durability enhancement of low-Pt PEMFCs.
Keywords: PEMFCs; Low-Pt loading; Transformer; Degradation prediction; Deep learning
Changlong Wang, Jiawei Jiang, Chong Han, Hengyi Ren, Lijuan Sun, Jian Guo,
Through-Wall Multihuman Activity Recognition Based on MIMO Radar,
Computers, Materials and Continua,
Volume 83, Issue 3,
2025,
Pages 4537-4550,
ISSN 1546-2218,
[Link]
([Link]
Abstract: Existing through-wall human activity recognition methods often rely on Doppler
information or reflective signal characteristics of the human body. However, static
individuals, lacking prominent motion features, do not generate Doppler information.
Moreover, radar signals experience significant attenuation due to absorption and scattering
effects as they penetrate walls, limiting recognition performance. To address these
challenges, this study proposes a novel through-wall human activity recognition method
based on MIMO radar. Utilizing a MIMO radar operating at 1–2 GHz, we capture activity data
of individuals through walls and process it into range-angle maps to represent activity
features. To tackle the issue of minimal variation in reflection areas caused by static
individuals, a multi-scale activity feature extraction module is designed, capable of extracting
effective features from radar signals across multiple scales. Simultaneously, a temporal
attention mechanism is employed to extract keyframe information from sequential signals,
focusing on critical moments of activity. Furthermore, this study introduces an activity
recognition network based on a Deformable Transformer, which efficiently extracts both
global and local features from radar signals, delivering precise human posture and activity
sequences. In experimental scenarios involving 24 cm-thick brick walls, the proposed
method achieves an impressive 97.1% accuracy in activity recognition classification.
Keywords: MIMO radar; human activity; Transformer; through-wall
Lixing Shi, Xueling Liang, Wenchao Chen, Yaoqiang Liu, Tong Ding, Kun Qin, Bo Chen,
Hongwei Liu,
Masked variational transformer for complex clutter modeling and target detection,
Signal Processing,
Volume 239,
2026,
110236,
ISSN 0165-1684,
[Link]
([Link]
Abstract: Weak target detection commonly encounters intense clutter interference, which
overshadows weak signals and complicates the task. Taking advantage of the powerful data
mining capability of neural networks, more and more deep learning-based methods are
applied to radar target detection. Among the approaches, those founded upon unsupervised
learning methodologies exhibit remarkable merit because they dispense with the
requirement for target samples within the training step, making them highly applicable in
practical target detecting scenarios. However, existing methods suffer from limitations in
leveraging the range-Doppler (R-D) two-dimensional correlation and finely modeling in
multiple clutter scenarios. In this paper, an unsupervised Transformer-based detector (TrDet)
is proposed to break through the boundary of modeling capability. First, with the designed
two-dimensional position embedding (2-DPE) and global query embedding (GQE)
techniques, an unsupervised training strategy for R-D spectrum based on Transformer
framework is utilized to achieve refined clutter modeling. Then, radar target detection is
formulated as an out-of-distribution (OOD) detection task to mitigate clutter interference.
Moreover, the masked variational Transformer-based detector (MVTrDet) is further
proposed to prevent target information leakage when the target is in close proximity to the
clutter in Doppler domain. Compared with several relative algorithms, our proposed
methods are better suited for radar target detection in complex clutter environments. The
experimental results derived from both measured data and simulated data verify the
effectiveness of our proposed methods.
Keywords: Radar target detection; Clutter modeling; Range-Doppler (R-D) spectrum;
Unsupervised learning; Out-of-distribution detection; Transformer
Han Zhang, Shengheng Liu, Hao Chi Zhang, Le Peng Zhang, Tong Chen,
Globally fused hierarchical transformer for nonuniform frequency diverse arrays,
Digital Signal Processing,
Volume 159,
2025,
105009,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Nonuniform frequency diverse array (FDA) enables significant advantages in joint
range-angle measurement for radar localization tasks. In this article, we propose a global-
context two-stage transformer (GC-TSformer) specifically tailored for high-resolution target
localization using a nonuniform FDA transmitter and a single-channel receiver. The global
context network aggregates learnable array-wide features, which enhances resilience to
amplitude-phase errors and improves frequency accuracy. Built on the Transformer
backbone, the hierarchical model facilitates high-order estimation through multi-scale
interactions among array elements. This enables effective processing within the size
constraints of the covariance matrix. Additionally, the hybrid model refines element-specific
attention weights to ensure balanced representation. Extensive simulations and field
experiments verify that GC-TSformer achieves sub-meter localization accuracy, with
attention map analysis confirming robustness across varied sensor configurations.
Keywords: Target localization; Frequency diverse array; Nonuniform frequency offsets;
Parameter estimation; Transformer architecture
Jinyang Xie, Kanghui Zhou, Lei Han, Liang Guan, Maoyu Wang, Yongguang Zheng, Hongjin
Chen, Jiaqi Mao,
Enhancing multi-task learning-based Tornado identification using spatial and temporal
information from weather radar images,
Applied Soft Computing,
Volume 184, Part B,
2025,
113834,
ISSN 1568-4946,
[Link]
([Link]
Abstract: Tornadoes, as dynamic weather phenomena, exhibit unique spatial and temporal
evolution characteristics that reflect their formation and development. Existing tornado
detection algorithms often struggle with high false alarm rates, primarily due to insufficient
capture of temporal correlations in tornado development. As an improvement, we propose a
multi-task tornado identification network with three-dimensional temporal and spatial
information (TS-MTINet). Taking continuous three-frame radar data as input, the Multi-
frame Temporal Interaction Block (MTIB) utilizes multi-head attention to model the dynamic
interaction information between the radar data, thus exploring in-depth the temporal
features during tornado development. Further, we design a Spatial-Temporal Enhancement
Module (STEM), which analyzes the difference information between continuous data to
extract local and global spatial and temporal feature variations about tornadoes. Based on
this architecture, TS-MTINet incorporates a multi-task learning framework to perform
tornado detection and number estimation tasks simultaneously, thus extracting
comprehensive information related to tornadoes. To validate the performance of the
proposed model, we construct the first Chinese tornado identification dataset with fine
radar features. The experimental results show that the proposed method shows significant
advantages in several evaluation metrics, especially in reducing false alarms. In practical case
studies, compared to the traditional TVS method, TS-MTINet achieves an increase in POD of
approximately 30% and a decrease in FAR of about 20% in several typical tornado events.
Particularly in environments with strong interference, TS-MTINet demonstrates higher
detection accuracy, reflecting greater robustness and practical value.
Keywords: Deep learning; Multi-task learning; Tornado identification; Weather radar;
Attention mechanisms
Hao Huang, Xueli Hao, Lili Pei, Jiangang Ding, Yujiao Hu, Wei Li,
Automated detection of through-cracks in pavement using three-instantaneous attributes
fusion and Swin Transformer network,
Automation in Construction,
Volume 158,
2024,
105179,
ISSN 0926-5805,
[Link]
([Link]
Abstract: To improve the performance of existing through-crack detection networks by
solving the problem in which through-cracks are misidentified as simple surface cracks due
to limited feature extraction, this study proposes an automated detection method based on
the fusion of three instantaneous attributes and the Swin Transformer network. First, a
900MHZ ground-coupled radar system was used to collect data and construct the original
dataset (Origin). Then, the Hilbert-Huang transform was used to extract three instantaneous
attributes (instantaneous amplitude (IA), instantaneous phase (IP) and instantaneous
frequency (IF)). Second, four fusion-feature datasets, i.e., IA + IP, IA + IF, IP + IF and
IA + IP + IF, were constructed using the spectral weighting of individual features. Finally, the
Swin Transformer network was proposed to detect through-cracks. The results show that the
IA + IF dataset exhibited the best performance. The improved network achieved a 5.4%
increase in the mean average precision compared with the initial network, reaching 87.78%.
Keywords: Through-cracks detection; Ground penetrating radar (GPR); Hilbert-Huang
Transform; Three instantaneous attribute; Feature fusion; Swin Transformer
Hongxi Zhao, Yiran Shi, Wenchao He, Hewei Sun, Haoran Wang, Jiahao Liu, Lin Gui,
Novel graph neural network and GNN-C-Transformer model construction for direction of
arrival estimation,
Digital Signal Processing,
Volume 168, Part D,
2026,
105619,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Direction of Arrival (DOA) estimation is essential in radar, sonar, wireless
communications, and speech processing. Traditional methods like MUSIC and ESPRIT provide
high resolution but suffer from high computational complexity and poor performance in low
signal-to-noise ratio (SNR) environments. Recent advances in neural networks, particularly
Convolutional Neural Networks (CNN), improve accuracy and robustness; however, CNNs’
ability to reduce time complexity and improving robustness under low SNR conditions
remains insufficient. This paper presents a novel framework for DOA estimation in sparse
arrays based on Graph Neural Networks (GNN) and proposes an entirely new array-based
graph connectivity structure. By modeling the array geometry as a graph, our GNN approach
captures spatial relationships effectively, addressing the challenges of time complexity and
low SNR. We further integrate Transformer layers to capture both spatial and temporal
dependencies, enhancing the model’s performance. Experimental results demonstrate that,
at SNRs ≤5dB, our GNN-based framework and the GNN-C-Transformer model developed
thereon achieve superior accuracy compared to existing methods, while exhibiting lower
computational complexity than all other algorithms except ESPRIT. This work advances the
application of GNN-based DOA estimation by providing a scalable solution for large-scale,
multi-dimensional signal processing in both dense and sparse array configurations.
Keywords: Direction of arrival estimation; Deep learning; Parameter estimation; Graph
neural networks; GNN-C-Transformer; Sparse array
Bixuan Gao, Riwei Zhang, Xiangyu Kong, Gaohua Liu, Kaijie Fang, Meimei Duan,
A Novel Carbon Emission Calculation Method for Power System Based on Personalized
Transformer with Two-Stage Training,
Engineering,
2026,
,
ISSN 2095-8099,
[Link]
([Link]
Abstract: Accurately and comprehensively calculating carbon emissions in the power system
is a fundamental prerequisite for achieving low-carbon energy transitions. Existing carbon
flow theory-based methods primarily concentrate on emissions from the grid and load sides,
while the methods for generation-side emissions often rely on costly continuous monitoring
systems or imprecise default emission factors. However, in practice, most power generation
units cannot achieve real-time and precise carbon emission measurement, leading to
generation side data deviations that affect overall computational accuracy. To address the
limitations, this paper proposes a novel carbon emission measurement method based on
heterogeneous data and personalized Transformer, applicable to various power generation
units. This method has several key innovations: ① extends traditional total electricity
production based models are extended to a multi-feature framework to capture similarities
and differences in time-series data among various generator units; ② an improved
Transformer is designed that integrates short-term relationship extraction, long-term
differential identification and long-term and short-term feature fusion modules to enhance
the multi-feature based emission mapping process, and ③ a two stage training protocol is
adopted, with self-supervised pretraining of feature extractors followed by fine tuning, to
accelerate convergence and improve accuracy. Experiments on real-world generation-unit
data show that the proposed method reduces average RMSE by 22.3% relative to a standard
Transformer and by 15.9% relative to Informer. Further validation utilizing the Institute of
Electrical and Electronics Engineers (IEEE) 30-bus test case confirms the effectiveness and
applicability of the model for carbon emission measurement across all segments of the
power system.
Keywords: Power system carbon emissions; Carbon emission calculation; Transformer; Multi-
frequency feature mapping
Walter Brescia, Pedro Gomes, Laura Toni, Saverio Mascolo, Luca De Cicco,
GT-MilliNoise: Graph transformer for point-wise denoising of indoor millimetre-wave point
clouds,
Signal Processing: Image Communication,
Volume 142,
2026,
117453,
ISSN 0923-5965,
[Link]
([Link]
Abstract: Millimetre-wave (mmWave) radars are gaining popularity thanks to their low cost
and robustness in low-visibility conditions. However, the 3D point clouds they produce are
sparser and noisier than those from LiDARs and depth cameras. These differences create
challenges when applying existing methods, originally designed for dense point clouds, to
mmWave data. Specifically, there is a gap in point-level precision tasks, such as full point
cloud denoising for mmWave data, partly due to the lack of fully annotated datasets. In this
work, we employ the MilliNoise dataset, a fully annotated indoor mmWave point clouds
dataset, to advance the understanding of mmWave point clouds denoising via two main
steps: (i) we carry out an experimental analysis of the most common point cloud processing
approaches and show their limitations in exploring the local-to-global structures in sparse
and noisy point clouds; (ii) in light of the identified limitations, we propose a graph-based
transformer architecture, denoted as GT-MilliNoise, composed of two main blocks to
effectively leverage both the temporal and geometric structures of the data: a Temporal
block leverages the sparsity of data to learn the dynamic behaviour of the points; a
Geometric block, uses a point-wise attention mechanism to form representative
neighbourhoods for feature extraction. The experimental results obtained in the MilliNoise
dataset show that our proposed GT-MilliNoise architecture outperforms the state-of-the-art
both qualitatively and quantitatively. Specifically, it achieves 75% accuracy (5% gain
compared to the state-of-the-art), and a significantly low Earth Mover’s distance value of
0.193.
Keywords: Point cloud; mmWave; Denoising; Deep learning
Adil Ali Saleem, Hafeez Ur Rehman Siddiqui, Muhammad Amjad Raza, Sandra Dudley, Julio
César Martínez Espinosa, Luis Alonso Dzul López, Isabel de la Torre Díez,
Ultra Wideband radar-based gait analysis for gender classification using artificial intelligence,
Array,
Volume 27,
2025,
100477,
ISSN 2590-0056,
[Link]
([Link]
Abstract: Gender classification plays a vital role in various applications, particularly in
security and healthcare. While several biometric methods such as facial recognition, voice
analysis, activity monitoring, and gait recognition are commonly used, their accuracy and
reliability often suffer due to challenges like body part occlusion, high computational costs,
and recognition errors. This study investigates gender classification using gait data captured
by Ultra-Wideband radar, offering a non-intrusive and occlusion-resilient alternative to
traditional biometric methods. A dataset comprising 163 participants was collected, and the
radar signals underwent preprocessing, including clutter suppression and peak detection, to
isolate meaningful gait cycles. Spectral features extracted from these cycles were
transformed using a novel integration of Feedforward Artificial Neural Networks and
Random Forests , enhancing discriminative power. Among the models evaluated, the
Random Forest classifier demonstrated superior performance, achieving 94.68% accuracy
and a cross-validation score of 0.93. The study highlights the effectiveness of Ultra-wideband
radar and the proposed transformation framework in advancing robust gender classification.
Keywords: Gait; Ultra-wide band radar; Gender classification; Spectral features; Feed
forward artificial neural network; Ridge classifier; Hist gradient boosting
Md. Jalil Piran, Xiaoding Wang, Ho Jun Kim, Hyun Han Kwon,
Precipitation nowcasting using transformer-based generative models and transfer learning
for improved disaster preparedness,
International Journal of Applied Earth Observation and Geoinformation,
Volume 132,
2024,
103962,
ISSN 1569-8432,
[Link]
([Link]
Abstract: Due to the rapidly changing climate conditions, precipitation nowcasting poses a
daunting challenge because it is impossible to make accurate short-term forecasts due to the
rapid fluctuations in weather conditions. There are limitations to traditional methods of
forecasting precipitation, such as the use of numerical models and radar extrapolation, when
it comes to providing highly detailed and timely forecasts. With the help of contemporary
machine learning (ML) models, including deep neural networks, transformers and generative
models, complex precipitation nowcasting tasks can be performed in an efficient way. To
address this critical task and enhance proactive emergency disaster management, we
propose an innovative method based on transformer-based generative models for
precipitation nowcasting. Our study area is the Soyang Dam basin in South Korea, located
upstream of the Han River, characterized by a monsoon climate with approximately 1200
mm of annual precipitation. To develop a precipitation nowcasting model, radar composite
data from 10 weather radars across South Korea is used. By utilizing radar reflective data in
order to train our model, we are able to effectively predict future precipitation patterns, thus
mitigating the risk of catastrophic weather conditions caused by heavy rainfalls. This dataset
covers reflectivity data from 2018 to 2022, with a spatial resolution of 1km over a 960 ×
1200 grid. Normalization using the min–max scaler method is applied to this reflectivity
data, which is then transformed into grayscale images for uniform comparison. We enhance
performance effectively by employing transfer learning with pre-trained Transformer
models. Initially, we train the model using a comprehensive dataset. Subsequently, we fine-
tune it for precipitation nowcasting using radar reflective data. This adaptation improves the
accuracy of rainfall forecasting by capturing crucial features. Leveraging prior task knowledge
through transfer learning not only enhances prediction accuracy but also increases overall
efficiency. In terms of predictive accuracy, extensive experimental results demonstrate that
our transformer-based nowcasting model outperforms related approaches, including
conditional generative adversarial networks (cGANs), U-Net, convolutional long short-term
memory (ConvLSTM), pySTEP. As a result of this research, disaster preparedness and
response will be greatly improved through improved weather prediction.
Keywords: Precipitation nowcasting; Transformer-based generative model; Radar reflective
data; ConvLSTM; cGAN; U-net
Dikun Hu, Weidong Gao, Kai Keng Ang, Mengjiao Hu, Rong Huang, Yingying Shao,
FECT-OSA: A transformer-enhanced multimodal system for non-contact sleep apnea
monitoring,
Alexandria Engineering Journal,
Volume 128,
2025,
Pages 628-641,
ISSN 1110-0168,
[Link]
([Link]
Abstract: Obstructive sleep apnea (OSA) is a common sleep disorder linked to an increased
risk of cardiovascular and neurocognitive disorders. While portable wearables provide a low-
cost alternative for OSA detection compared to polysomnography (PSG), their prolonged
wear burden, limited accuracy, and susceptibility to motion artifacts hinder practical
application. This study introduces FECT-OSA, a non-contact framework that extracts reliable
OSA indicators from piezoelectric physiological signals (PPS). FECT-OSA combines three key
components: adaptive signal processing, TransUnet, and ConvTransLSTM. Adaptive signal
processing utilizes entropy-based selection to determine the most stable input, improving
signal robustness. TransUnet improves feature extraction by integrating U-Net’s localization
with Transformer-based global modeling. It segments critical spectrogram regions, mitigating
motion artifacts and sidelobe interference. ConvTransLSTM integrates CNNs for local feature
extraction, a Multimodal Transformer for global alignment, and LSTM for temporal
modeling, enhancing OSA detection accuracy. In a study of 34 patients, FECT-OSA achieved
84.79% accuracy, 74.80% sensitivity, and 88.53% specificity, surpassing state-of-the-art
portable devices. For severe OSA patients (AHI≥30), it attains 94.1% accuracy and 90.9%
sensitivity/specificity, exceeding traditional methods by 5%–15% in accuracy and 5%–22% in
sensitivity. FECT-OSA’s non-contact design minimizes patient discomfort while ensuring high
accuracy across varying OSA severities, providing a practical home-based alternative.
Keywords: Obstructive sleep apnea; Smart sleep monitoring; Non-contact detection;
Enhanced feature extraction; Cross-attention fusion; Time-frequency analysis
Kai Yue, Zemeng Huang, Yubing Li, Yujia Chen, Tao Tan, Tao He, Yu Wang, Peng Ke, Xiuping Li,
A parallel Class-E power amplifier with doubly-tuned transformer-based load network and
high-efficiency cascode in 110-nm CMOS,
AEU - International Journal of Electronics and Communications,
Volume 205,
2026,
156117,
ISSN 1434-8411,
[Link]
([Link]
Abstract: This article presents a doubly tuned (DT) transformer-based parallel Class-E power
amplifier (PA). A compact parallel Class-E load network consisting of only one DT transformer
and a pair of capacitors is proposed to enhance output power and efficiency. Compared with
the traditional DT transformer-based series Class-E load, the proposed DT transformer-based
parallel Class-E load can further mitigate the constraints placed on device size and reduce
the impedance transformation ratio of the load. Besides, a cascode structure with
neutralization and charging acceleration capacitor (CX) is used as active core to enhance gain
and efficiency. The gain and stability of the active core are quantitatively analyzed based on
the transistor small-signal model, and it can be concluded that the gain of the active core
exhibits an increasing trend with the growth of CX while ensuring stability. As a proof of the
design, a 12 GHz Class-E PA is fabricated using 110-nm CMOS process. The measurement
results show that the proposed PA realizes a peak power-added-efficiency (PAE) of 30.9%, a
maximum saturated output power (Psat) of 18.2 dBm and a peak gain of 19.0 dB. The core
area of the circuit is only 990 μm ×260μm.
Keywords: Class-E power amplifier; Doubly-tuned transformer; Harmonic impedance;
Charging acceleration capacitor; Stability
Zhipeng Qing, Kecheng Ge, Shunsheng Zhang, Jing Yang, Zhijin Wen, Youlei Pu,
An inverse synthetic aperture radar imaging framework based on multi-layer networks and
heat conduction attention,
Engineering Applications of Artificial Intelligence,
Volume 167, Part 1,
2026,
113708,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Accurate compensation is essential for achieving high-resolution inverse synthetic
aperture radar (ISAR) imaging. Traditional parametric methods usually rely on iterative
optimization of the objective function to compensate for target motion in radar echoes.
However, the iteration process is often computationally intensive, difficult to integrate into
deep learning frameworks, and may discard sufficiently acceptable intermediate solutions.
To address these challenges, this study proposes a deep unfolding-based translational
compensation network that combines unsupervised learning with gradient back-
propagation. A prototype network is incorporated to monitor the imaging process, enabling
early termination of iterations. Moreover, a U-shaped network architecture based on a heat
conduction attention is employed to enhance ISAR image resolution and focusing
performance. To solve the problem of offset or splitting in the imaging results caused by
residual motion errors, a learnable affine transformation is employed for automatic
centering. These modules are integrated into an echo-to-image ISAR imaging framework.
Experimental results on both simulated and real radar data demonstrate the framework’s
effectiveness and robustness.
Keywords: Inverse synthetic aperture radar imaging; Translational compensation; Heat
conduction; Deep unfolding network; Affine transformation
Xinwen Yi, Dongyu He, Jiachang Liu, Xiaoling Zhu, Zhifang Pan,
HaFeiT: A fetal hypoxia diagnosis model using health status and fetal heart rate based on
vision transformer,
Biomedical Signal Processing and Control,
Volume 116,
2026,
109503,
ISSN 1746-8094,
[Link]
([Link]
Abstract: Electronic fetal monitoring is widely employed during pregnancy and labor periods
to detect fetal hypoxia. Due to internal observer differences, visual inspection of
cardiotocography based on clinical guidelines exhibits a high false-positive rate. Therefore,
AI-based cardiotocography classification using fetal heart rate signals is essential and
challenging, as it assists clinicians in objectively and accurately assessing fetal health status.
Most existing methods either focus on single-modal cardiotocography classification based on
signals or employ multimodal modeling with signals combined with natural language or
other statistical features. However, these methods do not consider the health status of the
gravida or fetus. This study is the first to incorporate maternal and fetal health status into
the fetal heart rate classification task. The fetal heart rate is transformed into a two-
dimensional image, and health status is extracted from the database. Both of them are fed
to the network we propose for classification. To address the lack of generalization and
robustness in existing methods, we propose a fetal hypoxia diagnosis model based on the
vision transformer. Compared to the related works, the proposed method demonstrates
strong generalization, achieving 95.582% accuracy with a 97.222% AUC on the public
database, and 97.658% accuracy with an 79.910% AUC on the private database. Compared
to related works, our proposed model demonstrates the most balanced performance across
all metrics. Moreover, experiments conducted in various scenarios demonstrate our model’s
strong robustness.
Keywords: Cardiotocography; Fetal hypoxia; Feature fusion network; Vision transformer
Farhana Ahmed Chowdhury, Md Kamal Hosain, Md Sakib Bin Islam, Md Shafayet Hossain,
Promit Basak, Sakib Mahmud, M. Murugappan, Muhammad E.H. Chowdhury,
ECG waveform generation from radar signals: A deep learning perspective,
Computers in Biology and Medicine,
Volume 176,
2024,
108555,
ISSN 0010-4825,
[Link]
([Link]
Abstract: Cardiovascular diagnostics relies heavily on the ECG (ECG), which reveals significant
information about heart rhythm and function. Despite their significance, traditional ECG
measures employing electrodes have limitations. As a result of extended electrode
attachments, patients may experience skin irritation or pain, and motion artifacts may
interfere with signal accuracy. Additionally, ECG monitoring usually requires highly trained
professionals and specialized equipment, which increases the treatment's complexity and
cost. In critical care scenarios, such as continuous monitoring of hospitalized patients,
wearable sensors for collecting ECG data may be difficult to use. Although there are issues
with ECG, it remains a valuable tool for diagnosing and monitoring cardiac disorders due to
its non-invasive nature and the detailed information it provides about the heart. The goal of
this study is to present an innovative method for generating continuous ECG waveforms
from non-contact radar data by using Deep Learning. The method can eliminate the need for
invasive or wearable biosensors and expensive equipment to collect ECGs. In this paper, we
propose the MultiResLinkNet, a one-dimensional convolutional neural network (1D CNN)
model for generating ECG signals from radar waveforms. With the help of a publicly
accessible radar benchmark dataset, an end-to-end DL architecture is trained and assessed.
There are six ports of raw radar data in this dataset, along with ground truth physiological
signals collected from 30 participants in five distinct scenarios: Resting, Valsalva, Apnea, Tilt-
up, and Tilt-down. By using strong temporal and spectral measurements, we assessed our
proposed framework's ability to convert ECG data from Radar signals in three distinct
scenarios, namely Resting, Valsalva, and Apnea (RVA). ECG segmentation performed better
by MultiResLinkNet than by state-of-the-art networks in both combined and individual
cases. As a result of the simulations, the resting, valsalva, and RVA scenarios showed the
highest average temporal values, respectively: 66.09523 ± 19.33, 60.13625 ± 21.92, and
61.86265 ± 21.37. In addition, it exhibited the highest spectral correlation values
(82.4388 ± 18.42 (Resting), 77.05186 ± 23.26 (Valsalva), 74.65785 ± 23.17 (Apnea), and
79.96201 ± 20.82 (RVA)), along with minimal temporal and spectral errors in almost every
case. The qualitative evaluation revealed strong similarities between generated and actual
ECG waveforms. As a result of our method of forecasting ECG patterns from remote radar
data, we can monitor high-risk patients, especially those undergoing surgery.
Keywords: ECG; Raw radar data; MultiResLinkNet; CNN; Deep learning
Cries Avian, Jenq-Shiou Leu, Hang Song, Jun-ichi Takada, Nur Achmad Sulistyo Putro,
Muhammad Izzuddin Mahali, Setya Widyawan Prakosa,
RCTrans-Net: A spatiotemporal model for fast-time human detection behind walls using
ultrawideband radar,
Computers and Electrical Engineering,
Volume 120, Part C,
2024,
109873,
ISSN 0045-7906,
[Link]
([Link]
Abstract: Ultrawideband (UWB) radar systems are becoming increasingly popular for
detecting human presence, even through walls. Recent advancements in signal processing
use deep learning techniques, which are known for their accuracy. While earlier methods
focused on spatial information using Convolutional Neural Networks (CNNs), newer research
highlights the importance of temporal information, such as how data peaks shift over time.
This study introduces RCTrans-Net, a deep-learning architecture that combines RCNet (a
Residual CNN) for spatial features with TransNet (a Transformer) for temporal features. This
fusion improves human presence classification in fast-time signal processing. Tested under
various conditions—different materials, body orientations, ranges, and radar heights—
RCTrans-Net achieved high performance with F1-scores of 0.997±0.000 for static,
0.967±0.004 for dynamic, and 0.978±0.001 for combined scenarios. The architecture
outperforms previous methods and offers real-time processing with an inference time of
about one millisecond.
Keywords: Human presence behind the wall; Residual network; Spatiotemporal'
Transformer; Ultrawideband radar system
Ziwei Zhang, Mengtao Zhu, Yunjie Li, Yan Li, Shafei Wang,
Joint recognition and parameter estimation of cognitive radar work modes with LSTM-
transformer,
Digital Signal Processing,
Volume 140,
2023,
104081,
ISSN 1051-2004,
[Link]
([Link]
Abstract: The recent developed cognitive radars can implement flexible work modes with
programmable modulation types and optimized modulating values for each mode definition
parameter. Automatic analysis of these work modes is a significant challenge for modern
electromagnetic reconnaissance receivers. In this paper, a Multi-Output Multi-Structure
(MOMS) learning-based processing framework is proposed for Joint inter-pulse automatic
Modulation Recognition and Parameter Estimation (JMRPE-MOMS). We propose a label
construction method as a feature interpretation method of the network to facilitate MOMS
learning and utilize the correlations between labels for performance gain. Moreover, an
LSTM-Transformer is designed to mine deep time-series characteristics, which can model
local and global relationships and reduce quantization loss. The proposed framework can
perform joint modulation recognition and parameter estimation (JMRPE) tasks
simultaneously with flexible output structures including scalar output and vector output
with fixed or variable sizes. Extensive simulations are performed based on the simulated
radar work modes defined with pulse repetition interval (PRI) sequences. The simulation
results validate the effectiveness and superiority of the proposed method especially under
non-ideal electromagnetic environments.
Keywords: Radar work mode; Automatic modulation recognition; Modulation parameter
estimation; Multi-output learning; Transformer
Binyue Cao, Mi He, Meiyun Zhao, Qinwen Ping, Chang He, Xiangyu Zhou, Yushun Gong,
Non-contact detection of electrocardiogram fiducial points with millimeter-wave radar,
Biomedical Signal Processing and Control,
Volume 111,
2026,
108273,
ISSN 1746-8094,
[Link]
([Link]
Abstract: Purpose:
Accurate detection of fiducial points in electrocardiograms (ECG) is crucial for diagnosing
heart diseases. However, traditional ECG devices require direct contact with the skin, which
can cause discomfort for patients. Therefore, there is an increasing need for non-contact
detection of ECG fiducial points, while research in this area remains limited.
Methods:
In response to this need, we proposed a millimeter-wave radar-based method for non-
contact detection of ECG fiducial points. Initially, fiducial points from synchronously acquired
ECG signals were annotated to serve as labels for network training. Subsequently, a radar
cardiogram (RCG) was derived from radar radio frequency signals using a second-order
differentiator. To enhance the detection process, we developed a one-dimensional semantic
segmentation model by integrating a deep residual shrinkage network block (DRSN) into the
U-Net architecture, which we term DRSN-Unet.
Results:
The results from the testing dataset demonstrate that the proposed method accurately
detects the five fiducial points, achieving an average precision of 0.986, an average recall of
0.986, and an average F1-score of 0.986. Notably, the network’s parameter count is only
2.042 Mega, with a computation amount of 0.086 GFLOPs.
Conclusion:
This method effectively detects the onsets of the P-wave and QRS complex, the peak of the
R-wave, and the offsets of the QRS complex and T-wave without direct contact with the
human body, laying a solid foundation for the automated diagnosis of complex cardiac
diseases.
Keywords: Millimeter-wave radar; Non-contact; Electrocardiogram; Fiducial point detection;
Semantic segmentation
Jun Li, Yihui Wang, Qinghong Sheng, Zhaocong Wu, Bo Wang, Xiao Ling, Xiang Liu, Yang Du,
Fan Gao, Gustau Camps-Valls, Matthieu Molinier,
CloudRuler: Rule-based transformer for cloud removal in Landsat images,
Remote Sensing of Environment,
Volume 328,
2025,
114913,
ISSN 0034-4257,
[Link]
([Link]
Abstract: Clouds are a key factor influencing transmission of the radiance signal in optical
remote sensing images. For mapping or monitoring the Earth's surface, it is inevitable to
mask or remove clouds before applying optical remote sensing images. Nowadays, deep
learning (DL) based thin cloud removal methods far outperform traditional methods. Yet
these DL-based methods often overlook position information or the physical cloud model in
thermal bands. Moreover, most existing cloud physical models for cloud removal overlook
the down-transmittance of the cloud in optical bands and do not account for the radiance of
thermal bands. This work proposes a novel transformer network, CloudRuler, coupled with
three rules in remote sensing domain for cloud removal. The proposed CloudRuler can
distinguish the semantic meanings between similar features in different pixel positions by
utilizing the Half-Spherical Coordinate System, aggregating features from local neighborhood
windows with remote sensing mosaicking, and solving the parameters of the cloud physical
model without limitations. Experimental results on 20 paired Landsat 8 and 9 images
demonstrate that CloudRuler outperforms seven baseline methods, based on GAN, CNN,
and transformer, both visually and quantitatively. Ablation experiments demonstrate that
the proposed rule-based modules are highly effective in improving CloudRuler's
performance for thin cloud removal. This work demonstrates that the joint use of Landsat 8
and 9 images for cloud removal is effective, producing more reliable data for downstream
applications than methods that utilize only one satellite with a longer revisit period. For
future research of the field, the code and dataset for reproducing the reported results are
available on: [Link]
Keywords: Cloud removal; Transformer; Cloud physical model; Landsat imagery; Deep
learning
Jiale Ren, Hengyi Li, Aihui Wang, Kenshi Saho, Lin Meng,
Radar-based gait analysis by Transformer-liked network for dementia diagnosis,
Biomedical Signal Processing and Control,
Volume 91,
2024,
105986,
ISSN 1746-8094,
[Link]
([Link]
Abstract: Providing reliable diagnostic evidence to doctors while minimizing financial and
physical burdens on patients is a prominent focus of current research. Gait features show
potential as a clinical marker for dementia diagnosis. Radar is capable of efficient,
contactless collecting human motion. This paper proposes a novel radar-based gait analysis
strategy with a Transformer-liked network for dementia diagnosis. The gait data are collected
by Micro-Doppler radar. Then Welch’s power spectral density estimation is adopted to
obtain the frequency features while unifying and compressing data size. The network is
designed to explore the relationship between dementia and gait features. In the network,
1D convolution is crucial in extracting local features and encoding features into deeper
dimensions. The attention-based module, inspired by the encoder of the Transformer,
possesses an edge in capturing long-sequence dependencies. The dual-stage gating
mechanism enhances the discriminative power of the learned representations by fine-tuning
the weights of extracted features. To validate the effectiveness of the proposed strategy,
comparative experiments are performed with prevailing networks in both time and
frequency domains. Experimental results demonstrate the superiority of the frequency
domain processing method over the time domain processing method and fusion time-
frequency processing method. Notably, the proposed model outperforms others, achieving
the highest accuracy of 94.93% in frequency domain processing-based experiments — 5.91%
and 4.25% higher than the highest accuracies in time domain processing-based experiments
and fusion time-frequency processing-based experiments respectively. The overall findings
illustrate that our proposal can provide a reliable reference for dementia diagnosis
effectively.
Keywords: Dementia diagnosis; Gait analysis; Power spectral density estimation; Attention
mechanism; Convolutional neural network; Gating mechanism
Yangxiaoyue Liu, Yuan Tian, Ying Xin, Yizhuo Yang, Jiangyuan Zeng, Min Feng, Chunqiao Song,
Transformer-based soil moisture simulation for understanding future drying trend globally,
Journal of Hydrology,
Volume 665,
2026,
134709,
ISSN 0022-1694,
[Link]
([Link]
Abstract: As a crucial element of the terrestrial water cycle, multiple future scenario soil
moisture (SM) datasets are widely applied to investigating Earth surface processes using
ensemble averages. However, they may run the risk of vague variation trend resulted from
averaging multiple models, which are characterized by different land surface models on
hydrological process simulation. To improve spatiotemporal pattern reliability, this study
innovatively designs a Transformer SM Simulation Net (TSMSNet), to conduct global SM
simulation of SSP1-2.6, SSP2-4.5, and SSP5-8.5 during 2016–2099. Nine qualified future SM
datasets, along with their spatial distribution of error parameters, and geographic data are
selected as model inputs. The learning target is calculated through merging merits from Soil
Moisture Active Passive and European Centre for Medium-Range Weather Forecast
Reanalysis v5-Land SM. The TSMSNet SM (R = 0.68, ubRMSE = 0.045 m3/m3) achieves good
matching degree against in situ measurements compared to the Convolutional Long Short
Term Memory (CSMSNet) simulated SM (R = 0.65, ubRMSE = 0.047 m3/m3). The TSMSNet
SM could favorably match the long-term trend of learning target, which exhibits advantage
over ensemble averages and CSMSNet SM. TSMSNet SM presents an overwhelming drying
trend. The decline magnitude rises accompanied by SSP changing from sustainable pathway
to fossil-fueled development. In terms of land cover types, evident drying trends are found
in cropland and forest. SM shows faster descent rate at habitable areas than inhabitable
areas. This paper develops a reliable TSMSNet SM dataset, which is expected to be a
valuable reference for understanding future SM variations.
Keywords: Soil moisture; Simulation; Global scale; Transformer; Trend analysis
Muhammad Yasir, Liu Shanwei, Xu Mingming, Wan Jianhua, Shah Nazir, Qamar Ul Islam, Kinh
Bac Dang,
SwinYOLOv7: Robust ship detection in complex synthetic aperture radar images,
Applied Soft Computing,
Volume 160,
2024,
111704,
ISSN 1568-4946,
[Link]
([Link]
Abstract: Using satellite-based SAR (Synthetic Aperture Radar) imagery to detect and track
ships is a formidable challenge. However, accurate analysis is hampered by inherent
difficulties such as obscured edges, multiple targets, varying dimensions and complex
backgrounds. These factors contribute to a poor signal-to-noise ratio. Various artificial
intelligence models have been developed to improve these problems, especially with YOLO-
based models. To overcome the above challenges and achieve robust performance, this
study presents a model called SwinYOLOv7, a novel fusion of the YOLOv7 framework with
the Features Pyramid Network and the Swin Transformer using ground-breaking anchor-free
detection algorithms. This innovative method aims to increase the accuracy of vessel
detection while reducing the impact of background clutter. The proposed model improved
by YOLOv7 examined three different datasets in detail: SRSDD-v1.0, HRSID, and SSDD to
optimize the performance of the model. The training process was consistently verified using
superior recall, precision, and F1-score values, which can be easily compared with previous
studies. The results show that the model using the Swin Transformer attention mechanism
and using an image size of 640×640 achieves the highest accuracy of 96.59%. Alternative
attention mechanisms, including the Squeeze-and-Excitation Network (SEnet), the
Convolutional Block Attention Module (CBAM), Channel Attention (CA) and Efficient Channel
Attention (ECA), deliver poorer accuracy rates. The combination of YOLOv7 and Swin
Transformer yielded encouraging results that enabled the proposed model to outperform
the current benchmarking models. Therefore, the proposed model provides a compelling
solution to ensure accurate vessel identification in complex search and rescue scenarios.
Keywords: Robust Ship detection; YOLO; Swin transformer; Complex synthetic aperture radar
images
Wenxu Zhang, Kang Luo, Fuli Sun, Zhongkai Zhao, Yunxiao Fu, Feiran Liu,
Radar working mode recognition for small samples based on the DBA-CIB-IMP method,
Physical Communication,
Volume 72,
2025,
102718,
ISSN 1874-4907,
[Link]
([Link]
Abstract: To address the issue that the reliability of radar working mode recognition
decreases as detected radar pulses decrease, a Dynamic Time Warping (DTW) Barycenter
Averaging-Circularly Integrated Bispectrum-Informer Multilayer Perceptron (DBA-CIB-IMP)
recognition method is proposed. This method uses CIB feature extraction to reduce the
complexity of the input without losing the feature information, and computes the global
feature similarity by DTW Barycenter Averaging (DBA). Generating samples based on the
original data set and expanding the data by generating a supplementary database through a
weighted average algorithm. The supplementary database is then fused with the original
database to complete the radar working mode recognition work by intelligent network
model. Higher recognition accuracies are achieved in scenarios with a limited training
samples. Recognition accuracy exceeds 90% at 0 dB SNR with low time spent.
Keywords: Radar working mode recognition; Multi-functional radar; Informer; Circularly
integrated bispectrum
Bin Zhang, En-Cheng Liou, Yi-Chih Tung, Muhammad Usman, Chiung-An Chen, Chao-Shun
Yang,
Cross-Site Map-Free Indoor Localization for 6G ISAC Systems Using Low-Frequency Radio and
Transformer Networks,
CMES - Computer Modeling in Engineering and Sciences,
Volume 145, Issue 2,
2025,
Pages 2551-2571,
ISSN 1526-1492,
[Link]
([Link]
Abstract: Indoor localization is a fundamental requirement for future 6G Intelligent Sensing
and Communication (ISAC) systems, enabling precise navigation in environments where
Global Positioning System (GPS) signals are unavailable. Existing methods, such as map-
based navigation or site-specific fingerprinting, often require intensive data collection and
lack generalization capability across different buildings, thereby limiting scalability. This
study proposes a cross-site, map-free indoor localization framework that uses low-frequency
sub-1 GHz radio signals and a Transformer-based neural network for robust positioning
without prior environmental knowledge. The Transformer’s self-attention mechanisms allow
it to capture spatial correlations among anchor nodes, facilitating accurate localization in
unseen environments. Evaluation across two validation sites demonstrates the framework’s
effectiveness. In cross-site testing (Site-A), the Transformer achieved a mean localization
error of 9.44 m, outperforming the Deep Neural Network (DNN) (10.76 m) and
Convolutional Neural Network (CNN) (12.02 m) baselines. In a real-time deployment (Site-B)
spanning three floors, the Transformer maintained an overall mean error of 9.81 m,
compared with 13.45 m for DNN, 12.88 m for CNN, and 53.08 m for conventional
trilateration. For vertical positioning, the Transformer delivered a mean error of 4.52 m,
exceeding the performance of DNN (4.59 m), CNN (4.87 m), and trilateration (>45 m). The
results confirm that the Transformer-based framework generalizes across heterogeneous
indoor environments without requiring site-specific calibration, providing stable, sub-12 m
horizontal accuracy and reliable vertical estimation. This capability makes the framework
suitable for real-time applications in smart buildings, emergency response, and autonomous
systems. By utilizing multipath reflections as an informative structure rather than treating
them as noise, this work advances artificial intelligence (AI)-native indoor localization as a
scalable and efficient component of future 6G ISAC networks.
Keywords: Indoor localization; 6G; ISAC; transformer; deep learning; map-free; cross-site;
wireless sensing
Bingzhe Fu, Wei Wang, Yihuan Li, Guorui Ren, Kang Li,
A frequency loss function based dynamic convolutional transformer model with data
denoising for short-term wind speed forecasting,
Expert Systems with Applications,
Volume 299, Part B,
2026,
130166,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Accurate and efficient wind speed forecasting is essential for power grid
management in balancing supply and demand, and support cost-effective transition to
renewable energy. This study develops a novel wind power forecasting system designed to
reduce deployment training computational time and support periodic updates in practical
applications. The system first applies Stagewise Orthogonal Matching Pursuit (StOMP) to
suppress noise in the raw wind speed sequences and performs correlation-based feature
selection to determine the most informative historical inputs. Subsequently, DCFDTFDD, the
forecasting model is constructed, in which multi-layer Dynamic Convolution (DC) modules
are employed for multi-scale feature fusion, and a frequency debiased Transformer is
introduced to directly learn multiple frequency components of wind speed sequences
through frequency-domain modeling and patching operations, without the need for prior
signal decomposition. Additionally, a Frequency-Domain Decorrelation (FDD) loss function is
incorporated to mitigate the inherent autocorrelation of label sequences in Transformer.
Experiments conducted on three real-world wind speed datasets demonstrate that the
proposed system delivers both accurate and efficient predictions, reducing training
computational time by 36.20 %, 34.24 % and 41.46 % compared with state-of-the-art
decomposition-based hybrid methods, while maintaining accuracy within 0.54 m/s RMSE.
These results indicate that the developed system offers a practical solution for wind farm
applications considering engineering time.
Keywords: Wind speed forecasting; Stagewise orthogonal matching pursuit denoising;
Transformer; Frequency-domain decorrelation loss function
Hua Wang, Qiangyu Zeng, Hao Wang, Jianxin He, Tiantian Yu, Guangpu Liu,
Temporal super-resolution reconstruction of weather radar echoes using a deep learning
approach,
Expert Systems with Applications,
Volume 300,
2026,
130189,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Severe convective weather events are characterised by rapid evolution and high
destructive potential, requiring weather radars to provide observations with high temporal
resolution. However, current S-band weather radar systems, constrained by their volumetric
scanning strategies, often fail to capture the rapidly changing features of these systems
promptly. To address this limitation, we propose EMAIRA-VFI, a deep learning–based
method for temporal super-resolution reconstruction of radar echoes, which enhances the
temporal resolution of radar data to meet the demands of severe convective weather
monitoring. By introducing an inter-frame attention mechanism, the proposed method
effectively fuses spatiotemporal features from sequential radar echoes, enabling accurate
modelling of dynamic weather evolution and the generation of continuous, high-temporal-
resolution radar echoes. Compared with conventional temporal interpolation methods,
EMAIRA-VFI demonstrates significant improvements in both interpolation accuracy and the
preservation of fine-scale meteorological structures. Experimental results show that the
model not only enhances the capability of S-band radars in monitoring rapidly evolving
weather events but also provides a new perspective for spatiotemporal fusion and the
intelligent application of radar data. We have open-sourced the code for this work at
[Link]
Keywords: Temporal super-resolution; Radar echo; Inter-frame attention mechanism
Yunlong Zhou, Chen Zhao, Fanfan Ji, Renlong Hang, Qingshan Liu, Xiao-Tong Yuan,
More realistic and accurate precipitation nowcasting with Conditional Rectified Flow
Transformers,
Engineering Applications of Artificial Intelligence,
Volume 165, Part A,
2026,
113402,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Precipitation nowcasting plays a critical role in disaster prevention and daily life but
remains challenging due to the intricate spatiotemporal dynamics of atmospheric processes.
In response to these challenges, recent research has shown that diffusion models can
generate visually realistic precipitation results. However, challenges such as accurately
predicting precipitation positions and improving inference speed remain unresolved. To
address these issues, we propose a novel Conditional Rectified Flow Transformers (CRFT)
architecture to improve precipitation nowcasting, which is designed to deliver both
predictive accuracy and visual realism. At its core, CRFT features an efficient latent space
predictor powered by OmniFormer blocks, which integrate spatial, temporal, and
spatiotemporal Transformers to holistically capture the atmosphere dynamics. We explore
five variants of spatiotemporal dynamic information interactions for OmniFormer and
demonstrate that integrating triple Transformers achieves the best performance.
Additionally, we significantly reduce inference time by employing a rectified flow approach,
achieving a reduction in inference steps by 98.4% compared to existing methods, enabling
high-quality 20-frame predictions within 2 s. Evaluated on three benchmark datasets, CRFT
outperforms state-of-the-art (SOTA) models in both accuracy and quality across multiple
metrics, offering an accurate and efficient solution for real-world nowcasting. The code is
publicly available at [Link]
Keywords: Precipitation nowcasting; Rectified flow; Transformer; Variational autoencoder
Liuyu Yang, Yuan An, Gang Zhang, Tuo Xie, Mengxin Liu,
Day-ahead electricity price forecasting method integrating multi-scale hypergraph features
and dual-layer transformer,
Applied Energy,
Volume 407,
2026,
127396,
ISSN 0306-2619,
[Link]
([Link]
Abstract: Accurate forecasting of spot electricity prices is critical yet challenging due to the
multi-scale temporal coupling, nonlinear volatility, and complex spatial dependencies
influenced by supply-demand fluctuations, extreme weather, and transmission topology.
This study proposes a novel day-ahead price forecasting model integrating multi-scale
hypergraph features with a dual-layer Transformer. A hypergraph is constructed based on
price trend similarity to capture spatial dependencies at local, global, and full-fusion levels.
High-relevance exogenous variables are selected using the maximum information coefficient
(MIC), and a two-tier Transformer separately models temporal and spatial dynamics.
Spectral hypergraph convolution is introduced to generate dynamic spatial representations.
The model is evaluated on real-world data from the Guangdong electricity market using both
single-day and rolling forecast tasks. Compared with the second-best model, RMSE, MAE,
and MAPE are reduced by 9.23%, 12.00%, and 21.74%, respectively, with R2 improved by
2.25%. Additionally, SHAP analysis quantifies feature contributions, forming a closed-loop
feature selection and validation process with MIC. The results demonstrate that
incorporating multi-scale dynamic modeling and spatiotemporal feature fusion can
significantly enhance forecasting accuracy.
Keywords: Multi-scale analysis; Spatial-temporal feature; Hypergraph; Transformer;
Electricity price forecasting; SHAP
Yunlin Ma, Tengfei Bao, Yangtao Li, Mengfan Zhao, Zhenhao Wu, Chengbo Fan,
GANFormerNet: A UAV-based Concrete Crack Segmentation Model for Water-related
Structures Using Vision Transformer and Graph Attention Network,
Advanced Engineering Informatics,
Volume 68, Part B,
2025,
103725,
ISSN 1474-0346,
[Link]
([Link]
Abstract: In this study, we propose a novel concrete crack segmentation model for water-
related structures, GANFormerNet, which realizes multi-feature fusion, dynamic attention
mechanism, and lightweight design. Specifically, the combination of ViT and Graph Attention
Network improves the crack recognition effect through topological modeling; the innovative
design of spatial-channel synergetic attention (SCSA) and semantic interaction module
(MSAF) work in concert to achieve multi-scale feature extraction and dynamic focusing; the
use of Depth-Separable Convolution reduces the amount of parameters, and the
introduction of the Focal Tversky loss function is introduced to solve the crack imbalance
problem at the boundary. The experiments show that the mIoU, Recall, Precision, and F1-
score of the GANFormerNet model reach 0.92178, 0.94375, 0.93531, and 0.94289,
respectively, which are significantly better than those of existing methods. The final model
was integrated into a GUI system to realize offline real-time detection of dam cracks.
Keywords: Water-related Structures; UAV; Crack Segmentation; GANFormerNet; Vision
Transformer; Graph Attention Network
Zhu Li, Xu Wanru, Zhu Chunqiang, Gao Jingkai, Mi Lugema, Deng Fan, Qu Jinqi,
AFMT:Adaptive frequency decomposition and multi-scale transformer for time series
forecasting,
Information Sciences,
Volume 726,
2026,
122735,
ISSN 0020-0255,
[Link]
([Link]
Abstract: Time series forecasting is essential in various fields, including power systems,
transportation, and meteorology. Although many existing methods improve predictive
performance by decomposing sequences into trend and seasonal components, traditional
techniques, such as moving averages, often suffer from spectral aliasing between high- and
low-frequency features, thereby impairing forecasting precision. Furthermore, current multi-
scale frameworks frequently neglect high-frequency patterns, restricting their capacity to
model complex nonlinear dependencies. To overcome these limitations, we propose AFMT, a
novel time series forecasting framework that integrates Adaptive Frequency-Domain
Decomposition with a multi-scale patch-wise Transformer architecture. Specifically, we
design a dynamic filter capable of adaptively isolating high- and low-frequency components
based on their spectral distributions, effectively mitigating feature entanglement.
Subsequently, a multi-scale patching strategy enables independent modeling of these
components through Transformer blocks, followed by a learnable frequency-aware fusion
mechanism, thereby enhancing feature independence and boosting predictive accuracy.
Extensive empirical studies on eight public datasets demonstrate that AFMT consistently
surpasses state-of-the-art approaches in both forecasting accuracy and generalization ability,
validating its effectiveness in frequency-domain decomposition and multi-scale temporal
modeling.
Keywords: Time series forecasting; Adaptive frequency domain decomposition; Multi-scale
patches; Transformer; Deeply separable convolution
Hebin Liu, Qizhi Xu, Xiaolin Han, Biao Wang, Xiaojian Yi,
Attention on the key modes: Machinery fault diagnosis transformers through variational
mode decomposition,
Knowledge-Based Systems,
Volume 289,
2024,
111479,
ISSN 0950-7051,
[Link]
([Link]
Abstract: Machinery signals typically consist of multiple sub-signals in different frequency
bands, while existing Transformer-based fault diagnosis methods often lack attention to key
fault frequencies, causing interference in fault diagnosis. Therefore, an innovative
Transformer structure for fault diagnosis based on variational mode decomposition (VMD) is
proposed. First, to address the difficulty in identifying signal features arising from coupling of
multiple frequency bands, a mode encoder based on VMD is proposed to decompose the
coupled modes and calculate the key modes. Second, a position encoding method based on
central frequency is proposed to address the lack of attention to signal’s frequency in
existing position encoding methods. Third, fault characteristic frequency is used to verify the
frequency band attention scores, improving the interpretability and reliability of the network
in response to the lack of internal interpretability in fault diagnosis methods based on deep
learning. Finally, the proposed method was validated on bearing vibration dataset and motor
sound dataset. The results showed that the method has high diagnostic accuracy, and could
capture the intrinsic modes of different faults.
Keywords: Fault diagnosis; Transformer; Mode decomposition
Tong Hou, Hongqing Zhu, Ziying Wang, Zhong Zheng, Kai Chen, Ying Wang, Bingcang Huang,
Low-rank fused modality assisted magnetic resonance imaging reconstruction via an
anatomical variation adaptive transformer,
Pattern Recognition,
Volume 175,
2026,
113044,
ISSN 0031-3203,
[Link]
([Link]
Abstract: In clinical practice, precise magnetic resonance imaging (MRI) reconstruction from
undersampled data is crucial. While multi-modal approaches can enhance reconstruction
quality, acquiring fully sampled auxiliary information is often time-consuming. Given that
computed tomography (CT) images are routinely obtained during clinical examinations, this
paper utilizes an anatomical variation adaptive transformer (AVAT) assisted by low-rank
fused CT-MRI modality to propose an MRI reconstruction network (ATLF-Net). Specifically,
this method leverages a CT-MRI fused modality to assist in MRI reconstruction. The ATLF-Net
encompasses fusion and reconstruction processes. The fusion process aims to generate a CT-
MRI fused modality that minimizes the gap between it and the MRI modality, serving as
auxiliary information for reconstruction. The proposed ATLF-Net comprises the global
feature-aware block (GFAB), the local feature-aware block (LFAB), and the low-rank fusion
module (LRFM). GFAB and LFAB extract global and local information from shallow features,
respectively. LRFM fuses CT and undersampled MRI through modality-specific low-rank
factors. During the reconstruction process, an AVAT is developed to extract complex and
elongated pathological features. Extensive experiments show that the proposed ATLF-Net
achieves robust performance and high-quality reconstructed images with few parameters
compared to benchmarks, across various acceleration rates on public datasets. The code for
this project is available at [Link]
Keywords: MRI Reconstruction; Multi-modality; Low-rank; Anatomical variation adaptive
transformer
Jiming Lv, Daiyin Zhu, Zhe Geng, Shengliang Han, Yu Wang, Zheng Ye, Tao Zhou, Hongren
Chen, Jiawei Huang,
Recognition for SAR deformation military target from a new MiniSAR dataset using multi-
view joint transformer approach,
ISPRS Journal of Photogrammetry and Remote Sensing,
Volume 210,
2024,
Pages 180-197,
ISSN 0924-2716,
[Link]
([Link]
Abstract: Accurately detecting ground armored weapons is crucial for achieving initiative
advantages in military operations. Generally, satellite or airborne synthetic aperture radar
(SAR) systems face limitations due to their revisit cycles and fixed flight trajectories, resulting
in single-view imaging of targets, thereby hampering the recognition of small SAR ground
targets. In contrast, MiniSAR possesses the capability to capture the multi-view of a target by
acquiring images from different azimuth angles. In this research, our team utilizes a self-
developed MiniSAR system to generate multi-view SAR images of real ground armored
targets and recognize targets. However, the recognition of small targets in SAR images
encounters two significant difficulties. First, small targets in SAR images are prone to
interference from background noise. Second, SAR target deformation arises from variations
in depression angles and imaging processes. To tackle these difficulties, this paper proposes
a novel SAR ground deformation target recognition approach based on a joint multi-view
transformer model. The method first preprocesses SAR images based on a low-frequency
priori SAR image denoising method. Next, it obtains multi-view joint information through a
self-attentive mechanism, inputs joint features to the transformer structure. The outputs are
jointly updated by a multi-way averaging adaptive loss function to improve the recognition
accuracy of deformed targets. The experimental results demonstrate the superiority of the
proposed method in SAR ground deformation target recognition, outperforming other
representative approaches such as information fusion of target and shadow (IFTS) and Vision
Transformer (ViT). It is concluded that the proposed method has high recognition accuracies
of 98.37% and 93.86 % on the moving and stationary target acquisition and recognition
(Mstar) and our SAR images dataset, respectively, in the field of SAR ground deformation
target recognition. We have included links to the code and data in the abstract of this paper
for ease of access. The source code and sample dataset are available at
[Link]
Keywords: MiniSAR; Transformer network; Denoising with low-frequency prior information
(LFPD); Recognition of deformation targets; Cross attention mechanism
Ayesha Ibrahim, Muhammad Zakir Khan, Muhammad Imran, Hadi Larijani, Qammer H.
Abbasi, Muhammad Usman,
RadSpecFusion: Dynamic attention weighting for multi-radar human activity recognition,
Internet of Things,
Volume 33,
2025,
101682,
ISSN 2542-6605,
[Link]
([Link]
Abstract: This paper presents RadSpecFusion, a novel dynamic attention-based fusion
architecture for multi-radar human activity recognition (HAR). Our method learns activity-
specific importance weights for each radar modality (24 GHz, 77 GHz, and Xethru sensors).
Unlike existing concatenation or averaging approaches, our method dynamically adapts
radar contributions based on motion characteristics. This addresses cross-frequency
generalization challenges, where transfer learning methods achieve only 11%–34% accuracy.
Using the CI4R dataset with spectrograms from 11 activities, our approach achieves 99.21%
accuracy, representing a 15.8% improvement over existing fusion methods (83.4%). This
demonstrates that different radar frequencies capture complementary information about
human motion. Ablation studies show that while the three-radar system optimizes
performance, dual-radar combinations achieve comparable accuracy (24GHz+77GHz: 96.1%,
24GHz+Xethru: 95.8%, 77GHz+Xethru: 97.2%), enabling flexible deployment for resource-
constrained applications. The attention mechanism reveals interpretable patterns: 77 GHz
radar receives higher weights for fine movements (superior Doppler resolution), while 24
GHz dominates gross body movements (better range resolution). The system maintains
71.4% accuracy at 10 dB SNR, demonstrating environmental robustness. This research
establishes a new paradigm for multimodal radar fusion, moving from cross-frequency
transfer learning to adaptive fusion with implications for healthcare monitoring, smart
environments, and security applications.
Keywords: Human activity recognition; Multi-modal fusion; Attention mechanisms; Cross-
frequency transfer learning
Mohammed A.A. Al-qaness, Abdelghani Dahou, Mohamed Abd Elaziz, Ahmed M. Helmi,
Human activity recognition and fall detection using convolutional neural network and
transformer-based architecture,
Biomedical Signal Processing and Control,
Volume 95, Part B,
2024,
106412,
ISSN 1746-8094,
[Link]
([Link]
Abstract: Human Activity Recognition (HAR) and fall detection, as applications within the
field of biomedical signal processing, are increasingly pivotal in enhancing patient care,
preventive healthcare, and rehabilitation. Fall therapy is one of the most expensive
treatments, usually taking a long time to complete. A single fall accident might result in
serious injuries, long-term incapacity, or even death. As a result, a reliable and cost-effective
fall detection system is essential. Wearable sensors have received wide attention due to
their availability and capability to capture different human motions. Thus, in the current
study, we develop a comprehensive HAR system for multi-classification tasks to recognize
several human actions, such as walking, sitting, standing, falling, and others. At the same
time, a binary classification of this model is developed to recognize fall and non-fall actions,
which can be used to track elderly actions and send an alert in case of falling to do necessary
rescue actions. The developed system is built using a Parallel Convolutional Neural Network
and Transformer-based architecture (PCNN-Transformer). PCNN-Transformer benefits from
the parallel architecture and the residual mapping mechanism to learn temporal feature
representations from the sensors’ data. The CNN blocks are aligned in parallel alongside
several Transformer-based encoders, followed by a concatenation operation to sum up the
extracted features from the input data ( sensors data). Moreover, the CNN blocks implement
a residual mapping mechanism to reduce the model complexity and training time. The
proposed model is tested using several open-source datasets: SisFall, UniMib-SHAR, and
MobiAct. It recorded high accuracy rates compared to several deep learning models. For
instance, in the binary classification (fall detection), the proposed model achieved an
average accuracy of 99.95%, 98.68%, and 99.71% for SisFall, UniMib-SHAR, and MobiAct,
respectively.
Keywords: Fall detection; Human activity recognition (HAR); Pattern recognition; Wearable
sensors; Deep learning
Yang Qin, Huiming Xie, Shuxue Ding, Yujie Li, Benying Tan,
Enhancing vision-and-language transformers through two-stage generative alignment pre-
training,
Engineering Applications of Artificial Intelligence,
Volume 163, Part 3,
2026,
113076,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Vision-language pre-training (VLP) models based on transformer architectures have
achieved significant success in bridging the gap between natural language processing and
computer vision. However, the alignment between visual and textual semantic objects
remains a major challenge, particularly when noisy image–text pairs are used for training,
which can lead to incorrect semantic associations and degrade model performance. In this
paper, we propose a novel two-stage generative alignment pre-training framework, called
VL-GAP (Vision-Language Generative-Alignment Pre-training), designed to improve object
alignment quality in vision-language models. In the first stage, we utilize high-quality
annotated datasets to perform supervised learning, introducing a center-point strategy for
automated object alignment and optimizing multiple loss functions to achieve precise visual-
text alignment. The second stage leverages large-scale noisy datasets for self-supervised
learning, where momentum models and confidence-based pseudo-label filtering are
employed to enhance the model’s robustness to noise. Experimental results demonstrate
that VL-GAP outperforms state-of-the-art models in various downstream tasks, highlighting
the importance of object alignment quality over data scale in improving VLP model
performance. Our approach provides new insights into effective handling of noisy data and
advances the capability of vision-language models to understand and generate coherent
multimodal descriptions.
Keywords: Vision-and-language; Transformer; Generative-alignment pre-training; Object
alignment
Yu Si, Zhaofeng He, Fan Zhang, Xiaoyun Sun, Yong Chen, Haiqing Zheng,
Cost-effective and real-time landslide monitoring method based on ultra-wideband using
ultra-wideband transformer neural network,
Engineering Applications of Artificial Intelligence,
Volume 160, Part A,
2025,
111851,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Landslides rank among the most destructive natural phenomena, posing
substantial risks to human safety, infrastructure, and ecological systems. Their frequent
occurrence in topographically complex regions demands urgent development in real-time
monitoring solutions. Current monitoring methodologies, however, are constrained by
prohibitive costs, limited temporal resolution, and high-power consumption. These factors
create substantial implementation barriers to implementing landslide monitoring systems.
To address these limitations, this study proposes an economical real-time monitoring
method leveraging ultra-wideband (UWB) technology for landslide detection. The
implementation of a dual-Microcontroller Unit (MCU) distributed hardware architecture
enables high-accuracy ranging capabilities and high real-time performance. To enhance the
spatial resolution of UWB systems in landslide monitoring, we propose an optimized sensor
deployment structure and a novel deep learning architecture called Ultra-wideband
Transformer (UWBformer). This network utilizes differential UWB-ranging data to predict
spatial displacement at monitoring locations, specifically the displacement distance,
horizontal angle, and pitch angle. UWBformer incorporates a spatial multi-head attention
mechanism and a dual-channel architecture processing both time-domain and frequency-
domain features. It is specifically designed to mitigate ranging error propagation and
enhance prediction stability by focusing on relative distance changes rather than absolute
ranging accuracy. Empirical results demonstrate UWBformer's superior performance in
predicting displacement distance, horizontal angle, and pitch angle, outperforming the
conventional Caffery-Taylor (C-T) localization approach and established deep learning
benchmarks. Field tests incorporated 3σ criterion and Kalman filtering alongside to pre-
process raw measurements, thereby enhancing data stability. Comprehensive validation
across field tests demonstrates UWBformer's capability to maintain accurate spatial
displacement estimation under harsh environments.
Keywords: Landslide monitoring; Ultra-wideband sensors; Ultra-wideband Transformer;
Deep learning; Cost-effective monitoring; Real-time monitoring
Yajing Wang, Xiuchen Wang, Miaomiao Kang, Zhihui Zhang, Yichen Yang, Wei Zeng, Zhe Liu,
Recent advances in graphene-based materials for radar and infrared stealth application,
Composites Part A: Applied Science and Manufacturing,
Volume 192,
2025,
108807,
ISSN 1359-835X,
[Link]
([Link]
Abstract: Graphene has attracted attention in the field of electromagnetic stealth due to its
excellent electrical properties. In this work, we first describe how microwave absorption
properties can be enhanced by using graphene to construct heterogeneous interfaces and
structural defects through conventional composite preparation methods. Meanwhile, the
infrared stealth performance is discussed by tuning the charge density and Fermi energy
levels of graphene as well as by constructing three-dimensional porous structures. Then, the
current state of research on microwave and infrared stealth based on metamaterials and
metasurface structure design strategies is reviewed. In addition, the research progress of
radar-infrared compatible stealth technology using three strategies of material composite,
metamaterial and metasurface structure design is summarized. Finally, the advantages and
limitations of graphene-based stealth materials prepared using different strategies are
analyzed, as well as the current challenges.
Keywords: Graphene-based composites; Metamaterials; Metasurfaces; Stealth
Bo Xu, Jialin Yan, Hao Tang, Yonghui Zhang, Maozheng Wang, Jinxiong Gao,
An efficient compressive CNN and transformer hybrid framework for long-term dissolved
oxygen prediction in aquaculture,
Information Processing in Agriculture,
2026,
,
ISSN 2214-3173,
[Link]
([Link]
Abstract: Current methods for predicting water quality do not fully consider the nonlinear
coupling relationship between multiple parameters, given the multifaceted nature of water
quality factor evolution in aquaculture systems. This results in limited prediction accuracy
and generalization ability, making it impossible to predict water quality parameters over an
extended period. This paper proposes a multi-factor related long-term dissolved oxygen
concentration forecasting framework, leveraging the synergistic integration of efficient
compression CNN and Transformer. First, to further improve the short-term feature fusion of
multi-sensor data, a CNN is used to learn localized spatial representations from multi-sensor
measurement inputs; then, a compressed pooling strategy with binary coding is introduced
to mitigate network accuracy degradation during pooling and achieve efficient and accurate
feature coding; finally, the feature fusion results are inputted into Transformer to improve
the global temporal capture ability of the model, achieving long-term water quality
parameter prediction. Empirical evidence demonstrates the superior performance of this
approach in long-term dissolved oxygen concentration forecasting. Compared with
traditional max and average pooling methods, the proposed compressed pooling strategy
improves model convergence speed by 13.59%. Compared with LSTM, GRU, Transformer,
CNN-Transformer (Max-Pool), Strided Convolution-Transformer, iTransformer, and Informer
models, the Coefficient of Determination (R2) can reach 0.969 ± 0.004 and 0.963 ± 0.003 at
prediction step lengths of 96 (4 days) and 168 (7 days). The Root Mean Squared Error (RMSE)
and Mean Absolute Error (MAE) were reduced to 0.022 ± 0.004 mg/L and 0.017 ± 0.002
mg/L for the 96-step prediction horizon, and to 0.025 ± 0.001 mg/L and 0.020 ± 0.001 mg/L
for the 168-step horizon, respectively. The novel predictive framework engineered in this
paper can optimize faster during the parameter learning, enhancing the model training
efficiency and improving the accuracy of long-term predictions.
Keywords: Long-term water quality prediction; Compressed pooling strategy; Binary coding;
Multi-factor correlation
Hui Bi, Chengjie Cai, Jiawei Sun, Shihao Ge, Huazhong Shu, Xinye Ni,
DRTNet: Dual-route transformer network for thyroid ultrasound segmentation based on
Bbox-supervised learning,
Knowledge-Based Systems,
Volume 324,
2025,
113781,
ISSN 0950-7051,
[Link]
([Link]
Abstract: Background and Objective:
Accurate nodule delineation plays a significant role in the intelligent diagnosis of thyroid
disease. However, the labels accessing is difficult since it is time-consuming and laborious. To
mitigate the over-dependence of the segmentation accuracy on the labels, we proposed a
dual-route transformer network (DRTNet) based on Bbox-supervised learning that only
requires the rough bounding rectangle as labels instead of a precise boundary for thyroid
ultrasound segmentation.
Methods:
DRTNet incorporates double-branch foreground class activation mappings (CAMs) into
Transformers to combine key areas. Meanwhile, double-branch architecture dynamically
adjusts the feature distribution of nodules of different sizes in frequency channels and
spatial dimensions, effectively addressing the localization of nodules of different sizes.
Moreover, ultrasound prior background-aware pooling (UPBAP) is proposed in both
branches to deal with the ambiguous boundary of thyroid nodules. Finally, adaptive
uncertainty estimate multi-scale consistency (AUEMC) is proposed to help mitigate the risk
of excessive over-fitting because of pseudo annotations, which further guarantees
consistency among nodules with diverse resolutions.
Results:
Substantial improvement of segmentation accuracy is shown on the public thyroid dataset of
TN3k and DDTI dataset with Dice similarity coefficient (DSC) of 84.94% and 83.98%, with
95% of the asymmetric Hausdorff distance (HD95) of 27.69 and 29.18, respectively. And our
private dataset has a DSC of 84.39% and HD95 of 14.53.
Conclusions:
The proposed DRTNet used rectangular box labeled for thyroid ultrasound images based on
Bbox-supervised learning. The experimental results show that the DRTNet is comparable to
these fully supervised methods. Code is available at [Link]
Keywords: Ultrasound image segmentation; Thyroid nodule segmentation; Transformer-
based network; Bbox-supervised segmentation; DRTNet
Viet Dinh Le, Gyu-Hyun Go, Sayali Pangavhane, Chang Kyoon Yoo,
Developing fully convolutional networks with permittivity-based class mapping for tunnel
lining defects detection in ground penetration radar scan data,
Engineering Applications of Artificial Intelligence,
Volume 166, Part B,
2026,
113670,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Non-destructive testing using ground penetrating radar (GPR) plays a vital role in
identifying defects in tunnel linings, such as cavities, delamination, and interlayers, to ensure
the safety and longevity of underground structures in mega-cities. However, the absence of
automated tools for generating realistic simulation datasets and the limitations of existing
methods in handling complex defect scenarios, such as cavity and delamination, pose
significant challenges for accurate detection. This study addresses these gaps by introducing
KIT-GPR, a numerical simulation tool, and a fully convolutional network (FCN) for automated
defect detection. KIT-GPR has employed the Finite-Difference Time-Domain method to
simulate electromagnetic wave propagation through multi-layered tunnel structures,
resulting in high-fidelity B-scan data. Its innovative application of the Perlin noise algorithm,
commonly used in game design, generates realistic interlayers between grout and rock
layers, thereby improving the accuracy of the simulation model. A modified FCN, designed
with a 320x320 square input size to accommodate versatile data and a 256-class output
corresponding to dielectric constant ranges, was trained using 1000 paired B-scan and
permittivity datasets. This approach achieves a Root Mean Square Error of approximately 0.8
in most scenarios, including interlayers with delamination. Furthermore, the predicted
delamination locations using FCN from B-scan data of a real tunnel closely align with the
results from endoscopic imaging. This demonstrates that the FCN prediction model has
significant promise as a scalable and efficient solution for early defect detection, thereby
greatly enhancing the safety of underground infrastructure.
Keywords: Ground penetrating radar; Fully convolutional network; Local defects; Tunnel
lining
Jongyun Byun, Jaehoon Cha, Jeyan Thiyagalingam, Hyeon-Joon Kim, Changhyun Jun,
Enhancing rainfall prediction accuracy through image fusion of radar and numerical weather
prediction models,
Expert Systems with Applications,
Volume 303,
2026,
130516,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Precipitation is one of the most challenging atmospheric phenomena to predict
due to the complexity involved in solving dynamic and thermodynamic atmospheric
equations. To address this challenge, extensive research has been conducted to enhance the
precision of numerical weather prediction models and radar-based extrapolation data, in
conjunction with the development of various blending techniques. However, traditional
methods have proven insufficient in capturing the diversity and nonlinearity of weather
phenomena. In response to these limitations, this study introduces a novel methodology
that leverages machine-learning based image fusion models to merge radar-based
extrapolation and numerical weather prediction rainfall datasets, thereby enhancing
prediction accuracy. An image fusion model was developed using radar-based extrapolation
data and numerical weather prediction data as input datasets, with radar observation data
utilized as target dataset. To identify the most suitable image fusion model for capturing the
complex patterns of rainfall data, two experiments were conducted: 1) Impact of model
topology, and 2) Effect of model size. A systematic analysis of the model outputs was
performed using eight evaluation metrics categorized under pixel-based metrics, feature-
based metrics, structural similarity metrics, and categorical verification metrics.
Experimental results indicated that image fusion model based on a Residual Network
(ResNet) outperformed other models in terms of model topology. Regarding model size, it
was observed that the performance did not increase proportionally with the number of
residual blocks; the most suitable performance was achieved with a specific number of
residual blocks (Case 5: 8 blocks). Additionally, the metrics compared with radar observation
data indicated that the proposed model delivered superior performance, thus offering a
high-accuracy rainfall prediction methodology.
Keywords: Image fusion; Radar; Numerical weather prediction; Deep learning; Precipitation
Guomin Xie, Zijian Zhang, Sen Xie, Zhaowei Yuan, Hao Liu,
CPWformer-DEC: improved Transformer with class-priority weather attention and dynamic
error compensation for photovoltaic power forecasting,
Expert Systems with Applications,
Volume 301,
2026,
130580,
ISSN 0957-4174,
[Link]
([Link]
Abstract: With the rapid expansion of photovoltaic (PV) capacity, accurate short-term and
medium-term PV power forecasting has become critical for grid dispatch and energy storage
coordination. However, the sudden and drastic fluctuations of meteorological variables over
a short period of time, as well as the abrupt transitions between contrasting weather states,
severely hinder the forecasting accuracy. Therefore, improved Transformer with class-
priority weather attention and dynamic error compensation (CPWformer-DEC) for
photovoltaic power forecasting is developed. In detail, based on the Transformer network,
sample-level Gaussian mixture model (GMM) categorization is firstly used to finely encode
the weather patterns. Then, class-priority weather attention and multi-scale convolution
module are added to balance intra-class and inter-class information. Moreover, a lightweight
long short-term memory (LSTM) is utilized for online residual correction. The experimental
results indicate that CPWformer-DEC achieves substantially lower errors under diverse
conditions. Compared with the suboptimal model, MAE is reduced by 35.4%, 64.8%, and
70.1% respectively in three typical cases, thereby supporting day-ahead and intra-day
dispatch and improving reserve allocation and storage scheduling in PV integrated power
systems.
Keywords: Photovoltaic power forecasting; Transformer; Class-priority weather attention;
Multi-scale fusion; Dynamic error compensation
Haoyu Zhang, Stephen Wu, Xiangyun Luo, Yong Huang, Hui Li,
Efficient matching of Transformer-enhanced features for accurate vision-based displacement
measurement,
Automation in Construction,
Volume 171,
2025,
105962,
ISSN 0926-5805,
[Link]
([Link]
Abstract: Computer vision technology and monitoring videos have been employed to obtain
structural displacement measurements. Noniterative algorithms are mainly designed for
rapid tracking of the motions of individual image points, rather than dense motion fields.
Iterative algorithms are limited to estimating motion fields with small amplitudes and
require high computation cost to achieve high accuracy. This paper introduces a noniterative
method for vision-based measurements that balances speed and density. The method
employs an attention-based matching strategy applied to Transformer-enhanced image
features. Motion priors and a physics-informed denoising approach are integrated to
improve measurement accuracy. Tested on challenging truss and cable-stayed bridge
vibration videos, the method demonstrated superior displacement measurement
performance compared to conventional approaches. It also achieved greater robustness to
brightness changes and partial occlusions while requiring minimal human intervention. This
method supports the development of automated and affordable vibration monitoring
systems.
Keywords: Attention mechanism; Subpixel feature matching; Transformer; Computer vision;
Displacement measurement; Vibration monitoring
Guowen Li, Zihang Huang, Teng Fei, Dunxin Jia, Meng Bian,
Listen to the road: acoustic traffic monitoring on edge platforms via Lightweight Noise
Spectrogram Transformer (LNST),
Pervasive and Mobile Computing,
Volume 115,
2026,
102132,
ISSN 1574-1192,
[Link]
([Link]
Abstract: Accurate real-time traffic flow monitoring is crucial for intelligent transportation
systems (ITS), enabling optimized traffic management, urban planning, and policy-making.
However, conventional methods face cost, deployment, weather, and privacy challenges.
Addressing these shortcomings, this study investigates the potential of utilizing ubiquitous
traffic noise, an inherently accessible, cost-efficient, non-intrusive, and privacy-preserving
signal, as a viable data source. We propose the Lightweight Noise Spectrogram Transformer
(LNST), a novel deep learning model for analyzing traffic noise spectrograms as a Proof of
Concept. LNST leverages the Transformer architecture's self-attention mechanism to
effectively capture long-range temporal and spectral dependencies crucial for interpreting
complex traffic acoustics. Trained and evaluated on diverse urban traffic scenarios, LNST
demonstrates significant advantages. Experimental results show it consistently outperforms
baseline models, achieving superior prediction accuracy (MSE, MAE, R²). Furthermore,
through transfer learning and model pruning, LNST achieves high computational efficiency
with substantially fewer parameters and faster inference speeds. Its lighter design also
ensures its feasibility for deployment on resource-constrained edge computing platforms.
This work validates the practicality of acoustic sensing for traffic monitoring and presents an
accurate, computationally efficient, and LNST as a cost-effective, easily deployable, and
privacy-respecting solution, offering a valuable supplementary tool for advancing ITS.
Keywords: Traffic flow monitoring; Traffic noise; Transformer; Intelligent transportation
system; Edge computing
Canfeng Liu, Binhui Wang, Hui Dong, Yihan Pan, Jiawen Lin, Jintian Yang, Yihui Tao, Hao Sun,
Time series analysis of nucleic acid reactions via a generalized transformer model,
Chemometrics and Intelligent Laboratory Systems,
Volume 267,
2025,
105522,
ISSN 0169-7439,
[Link]
([Link]
Abstract: The contemporary landscape of medical diagnostics and therapeutic interventions
has witnessed a remarkable surge in the production of time series data. Artificial intelligence
(AI), particularly the deep learning, has presented promising values in investigating the high-
dimension and meaningful significance hidden behind these diagnostic data. In this work,
we propose a novel analytics for intelligent nucleic acid amplification tests (NAAT) based on
deep learning and paper microfluidics. On-chip amplification data were straightforwardly fed
to a deep learning model derived from Transformer neural network. To facilitate the
development and deployment of the approach, we conducted a lightweight processing of
the Transformer model. Then, the capacity of the model for accurately predicting the
reaction trend and end-point value was validated. We also employed ablation experiments
to evaluate the effects of various parameters on prediction performance followed by
optimizing the model. Then, three clinical datasets including 706 positive and 205 negative
samples obtained from Fujian Provincial Hospital were used to verify the generalization of
the approach. Without any modification of the model structure and hyperparameters,
accuracy, sensitivity, and specificity by the presented approach were 98.28 %, 97.52 % and
99.02 %. Further comparison studies based on the nine different AI algorithms including
recurrent neural network and long-short term memory were performed. The presented
study holds potential to facilitating routine diagnostic tasks for preventing pandemic and
propelling the development of smart portable instruments.
Keywords: NAAT; Deep learning; Generalized transformer model; Predictive analytics
Shuai Liu, Tong Yu, Jun Zhou, Guanglong Xing, Huanfa Chen,
Cross-source transformer-based neighborhood contrastive learning for joint classification of
hyperspectral and LiDAR Data,
Information Fusion,
Volume 124,
2025,
103225,
ISSN 1566-2535,
[Link]
([Link]
Abstract: The fusion of the hyperspectral image (HSI) and light detection and ranging (LiDAR)
data has demonstrated significant potential in the land cover classification task. Although
deep learning has shown remarkable success in the joint classification of HSI and LiDAR data,
the large amount of unlabeled multisource remote sensing data is not fully utilized.
Additionally, effectively integrating HSI and LiDAR data remains a challenging task, and the
semantic relationship of neighborhood regions needs to be further exploited. In this paper,
we propose a cross-source transformer-based neighborhood contrastive learning model
(CTNCLM), which acquires a more discriminative feature representation from unlabeled data
through the pre-training stage. Considering the semantic correlation between neighboring
image patches, CTMCLM achieves the joint classification of HSI and LiDAR data at a more
precise level. A cross-patch contrastive learning (CPCL) module is proposed to calculate the
similarity between original patches and neighborhood patches. Furthermore, a cross-source
transformer (CST) with cross-source attention is proposed to fuse the multi-source data,
which exploits the intermodal information interaction between the HSI and LiDAR data.
Extensive experiments on three public datasets demonstrate the superior classification
performance of the proposed method compared with several state-of-the-art methods.
Keywords: HSI-LiDAR classification; Self-supervised learning; Momentum contrast;
Neighborhood contrastive learning; Cross-source transformer
Meng Zhang, Jilong Liu, Bing Han, Shengli Dong, Tong Cui, Yan Ren,
Adaptive processing of non-stationary ship lubrication signals under sporadic impulsive
noise,
Ocean Engineering,
Volume 342, Part 2,
2025,
122693,
ISSN 0029-8018,
[Link]
([Link]
Abstract: During the 2022 transpacific voyage of the tanker “Yuan Fu Yang”, we encountered
a challenging problem in machinery monitoring: sensor signals were severely corrupted by
sporadic impulsive noise superimposed on already non-stationary patterns. The harsh
marine environment–with intense mechanical vibrations and salt-induced sensor
degradation–created signal characteristics that deviated significantly from the Gaussian
assumptions of classical methods. To tackle this real-world challenge, we developed an
adaptive three-stage processing framework. Our first innovation, Adaptive Spectral
Enhancement Signal Extension (ASE-SE), addresses the dyadic length requirements of
wavelet transforms through context-aware autoregressive modeling that maintains signal
dynamics while satisfying computational constraints. Recognizing that conventional
denoising approaches often learn padding artifacts, we then designed a Learnable Wavelet
Packet Transform with Focus mechanism (LWPT-Focus), which cleverly masks synthetic
regions during training to avoid this pitfall. Finally, given the time-varying nature of ship
signals, we created the Frequency-Adaptive Decomposition Linear Model (FADLinear) that
dynamically adjusts trend-seasonal separation according to local spectral features. Testing
on 84,951 samples spanning 12 lubrication system parameters revealed that our approach
achieved an MSE of 0.0386 for 96-step predictions—outperforming various state-of-the-art
baselines including Transformer variants, yet maintaining the computational efficiency
essential for shipboard deployment. This work offers a practical solution for maritime
predictive maintenance under conditions where conventional methods struggle. The code
and dataset for this study are publicly available at: [Link]
padd-foucs-lwpt-fadlinear.
Keywords: Frequency-adaptive modeling; Non-stationary signals; Ship lubrication signals;
Sporadic noise; Time series forecasting; Unsupervised denoising
Pengkai Wang, Jonghoek Kim, Mitra Ghergherehchi, Mingxuan Zhang, Estrella Montero,
Luwei Liao, Zhong Yang, Hongyu Xu,
Transformer-based aerial robot tracking system in environments with wind disturbances,
Robotics and Autonomous Systems,
Volume 193,
2025,
105104,
ISSN 0921-8890,
[Link]
([Link]
Abstract: Unmanned aerial vehicles (UAVs) are increasingly used in agriculture, surveillance,
and search and rescue. However, maintaining stable flight and accurate navigation in
dynamic environments, especially with wind disturbances, remains a challenge. Traditional
navigation systems often struggle with unreliable sensor data, complicating pose estimation
and tracking. This article proposes an advanced master–slave UAV system combining a
transformer-based model with YOLO for enhanced tracking in wind-affected environments.
YOLO performs real-time object detection, extracting feature points matched with known
landmarks to estimate the UAV’s position. To address the challenges of wind disturbances,
we simulate various wind conditions and train the model under different wind disturbance
environments. Using transformer-based trajectory and pose predictions, we provide control
compensation to counteract the effects of wind disturbances, ensuring stable flight in
dynamic conditions. The pose estimation is refined by integrating visual data with inertial
measurement unit (IMU) data using transformer architectures. A vision-based formation
control strategy is introduced for precise relative positioning in multi-UAV formations.
Initially designed for three UAVs, this strategy is extended to handle larger formations and
complex geometric shapes, focusing on maintaining a triangle formation. A graph-based
dynamic formation control framework enables real-time adaptation to formation changes
and environmental conditions. The approach improves MPC control with a transformer
model, enhancing adaptability to wind disturbances. The system’s effectiveness is validated
using webots simulations, demonstrating its ability to track UAVs and adapt to challenging
environmental conditions. A theorem proves the convergence of the control law using
Lyapunov’s direct method, ensuring that formation errors decay over time. Comparative
experiments and webots simulations confirm the approach’s feasibility, validating its
robustness in maintaining precise formation control under dynamic environmental factors.
Finally, we validate the reliability of our method in real-world environments, confirming its
practical applicability.
Keywords: Transformer-based model; YOLO algorithm; Image processing; Target tracking;
Wind disturbances; Unmanned aerial vehicles; Formation control
Xinyu Zhang, Chu Zhang, Rui He, Changwen Ma, Junhao Yao, Muhammad Shahzad Nazir,
Tian Peng,
A pyramidal attention-based transformer model based on improved differential innovation
search algorithm and feature extraction for solar radiation prediction considering relevant
factors,
Renewable Energy,
Volume 253,
2025,
123666,
ISSN 0960-1481,
[Link]
([Link]
Abstract: Accurate prediction of solar radiation intensity (SI) is crucial for power system
scheduling and site selection of photovoltaic power stations. This paper proposes a
multivariable solar radiation prediction model based on Time-varying Filter-based Empirical
Mode Decomposition (TVFEMD), Fuzzy Entropy (FE), Random Forest (RF), Triangular
Wandering Strategy Improved Differential Creative Search Algorithm (TDCS), and Pyramidal
Attention-based Transformer (Pyraformer). First, the Random Forest is used for feature
extraction of solar radiation data; then, TVFEMD is employed to break down solar radiation
data into constituent sub-modes, thereby mitigating the non-stationarity present in the data
sequence. Fuzzy Entropy-based aggregation is utilized to diminish the quantity of data
sequences, which, together with various features, forms a multivariable input feature
matrix. The Differential Creative Search Algorithm (DCS) is enhanced with a Triangular
Wandering Strategy to obtain Improved TDCS, which optimizes the hyperparameters of
Pyraformer for solar radiation prediction, thereby enhancing the model's predictive
performance. This paper analyzes the prediction metrics of the TDCS-RF-TVFEMD-FE-
Pyraformer multivariate model compared with nine other multivariate benchmark models.
The results show TVFEMD, FE, and RF boost models accuracy. Post TDCS optimization,
Pyraformer's RMSE and MAE outperform the baseline models by 10 %–50 %, with R and
SMAPE also outperforming the baseline models.
Keywords: Solar radiation prediction; Multivariate prediction; Pyraformer; Differential
creative search algorithm optimization
Shengchun Wang, Haowen Li, Lianye Liu, Ronghui Cai, Zhonghai Yin, Huijie Zhu,
TSFI-Fusion: A dual-branch decoupled infrared and visible image fusion network based on
transformer and spatial-frequency interaction,
Optics and Lasers in Engineering,
Volume 195,
2025,
109287,
ISSN 0143-8166,
[Link]
([Link]
Abstract: Infrared and visible image fusion (IVIF) aims to generate high-quality images by
combining detailed textures from visible images with the target-highlight capabilities of
infrared images. However, many existing methods struggle to capture both shared and
unique features of each modality. They often focus only on spatial domain fusion, such as
pixel averaging, while overlooking valuable frequency domain information. This makes it
hard to retain fine details. To overcome these limitations, we propose TSFI-Fusion, a dual-
branch network that combines Transformer-based global understanding with spatial-
frequency detail enhancement. The two branches include a Transformer-based semantic
construction branch for capturing global features and a detail enhancement branch utilizing
an invertible neural network (INN) and a frequency domain compensation module (FDCM)
to integrate spatial and frequency information. We also design a dual-domain interaction
module (DDIM) to improve feature correlation across domains and a collaborative
information integration module (CIIM) to effectively merge features from both branches.
Additionally, we introduce a focal frequency loss to guide the model in learning important
frequency information. Experimental results demonstrate that TSFI-Fusion outperforms
existing methods across multiple datasets and metrics on the IVIF task. In downstream
applications such as object detection, it effectively enhances performance. Furthermore,
extended experiments on the MIF task reveal the robust generalization ability of the
proposed mechanism across diverse fusion scenarios. Our code will be available at
[Link]
Keywords: Infrared and visible image fusion; Transformer; Spatial-frequency information;
Medical image fusion
Murat Hi̇çyılmaz,
Bayesian optimization-enhanced vision transformer for damage detection in steel truss
structures using continuous wavelet transform analysis,
Structures,
Volume 85,
2026,
111059,
ISSN 2352-0124,
[Link]
([Link]
Abstract: The main contribution of this paper is to optimize the hyperparameters of the
Vision transformer (ViT) architecture using an adaptive Bayesian optimization process that
provides superior performance compared to state-of-the-art models in terms of accuracy
and efficiency while remaining within acceptable limits in terms of computational efficiency.
For this purpose, a Bayesian-optimized Image Transformer model (CwBot) is proposed for
damage detection in steel truss structures using continuous wavelet transform (CWT)
analysis. CWT spectrograms generated from a 72-bar steel truss model and measurements
of a real steel truss bridge are considered to better reveal the complex dynamic behavior of
steel truss structures. The ViT hyperparameters including patch size, embedding size,
transformer depth, multi-head attention heads, MLP size, learning rate, and batch size are
optimized and compared with seven different state-of-the-art transfer learning models. The
proposed model achieves 93.29 % accuracy for the 72-bar steel truss and 98.43 % accuracy
for the real steel truss bridge, outperforming other models.
Keywords: ViT; Bayesian optimization; CWT, Steel trusses; Damage detection
Jing Jiao, Hu Wang, Yamin Dang, Yingying Ren, Caiya Yue, Xiujuan Wu, Haomeng Cui, Xinlin
Wang,
Noise-resilient GNSS coordinate time series prediction using AVMD-sLSTM-transformer
hybrid model,
Advances in Space Research,
Volume 76, Issue 11,
2025,
Pages 6863-6881,
ISSN 0273-1177,
[Link]
([Link]
Abstract: High-precision modeling and analysis of Global Navigation Satellite System (GNSS)
coordinate time series constitute a fundamental basis for geophysical and crustal
deformation studies. To overcome the accuracy limitations of traditional methods under
noisy GNSS observations, we introduce a hybrid AVMD-sLSTM-Transformer prediction
model. The approach is validated using over a decade of observations from 584 global GNSS
stations. Initially, an improved Adaptive Variational Mode Decomposition (AVMD) method is
applied to decompose the preprocessed time series, with power spectral density analysis
guiding signal reconstruction and effectively suppressing noise. The reconstructed signals are
then input into a hybrid architecture combining scalar Long Short-Term Memory (sLSTM)
networks and Transformer models. This architecture leverages the sLSTM’s ability to capture
local temporal dynamics and the Transformer’s strength in modeling global spatiotemporal
dependencies, enabling robust long-term forecasting. The results indicate that significant
advantages in the north, east, vertical components, particularly achieving sub-millimeter
accuracy in the east direction. The proposed model reduces prediction errors by 65 %
compared to sLSTM only implementations and demonstrates 32 % improvement over
standard alone Transformer models. When compared to conventional Support Vector
Regression (SVM) and Autoregressive Integrated Moving Average (ARIMA) approaches, it
achieves over 80 % enhancement in mean absolute error (MAE) metrics. For horizontal
components, the model exhibits exceptional stability, validating the efficacy of the hybrid
architecture in synergizing sLSTM’s temporal modeling with Transformer’s global attention
mechanism. Notably, the model maintains residual errors within ±1 mm for stations affected
by coseismic offset in plate boundary zones, demonstrating robust adaptability to non-
stationary crustal deformation patterns.
Keywords: GNSS time series prediction; AVMD; sLSTM; Transformer
Jiajing Xie, Ying Chen, Shijie Luo, Wenxian Yang, Yuxiang Lin, Liansheng Wang, Xin Ding,
Mengsha Tong, Rongshan Yu,
Tracing unknown tumor origins with a biological-pathway-based transformer model,
Cell Reports Methods,
Volume 4, Issue 6,
2024,
100797,
ISSN 2667-2375,
[Link]
([Link]
Abstract: Summary
Cancer of unknown primary (CUP) represents metastatic cancer where the primary site
remains unidentified despite standard diagnostic procedures. To determine the tumor origin
in such cases, we developed BPformer, a deep learning method integrating the transformer
model with prior knowledge of biological pathways. Trained on transcriptomes from 10,410
primary tumors across 32 cancer types, BPformer achieved remarkable accuracy rates of
94%, 92%, and 89% in primary tumors and primary and metastatic sites of metastatic
tumors, respectively, surpassing existing methods. Additionally, BPformer was validated in a
retrospective study, demonstrating consistency with tumor sites diagnosed through
immunohistochemistry and histopathology. Furthermore, BPformer was able to rank
pathways based on their contribution to tumor origin identification, which helped to classify
oncogenic signaling pathways into those that are highly conservative among different
cancers versus those that are highly variable depending on their origins.
Keywords: cancer of unknown primary; CUP; biological pathway; transformer; tracing the
origin of cancer
Jiayi Cai, Zhaocheng Yang, Ping Chu, Juntao Guo, Jianhua Zhou,
Robust hand gesture detection and recognition using 4D millimeter-wave radar in a
ubiquitous scene,
Measurement,
Volume 253, Part C,
2025,
117545,
ISSN 0263-2241,
[Link]
([Link]
Abstract: In current research on HGR using radar sensors, hand gestures are typically
confined to a smaller region. However, in ubiquitous scenarios, unrestricted human body
movements and unexpected hand gesture motions usually occur, which results in a large
false alarms and recognition performance degradation. To address this issue, we propose a
robust hand gesture detection and recognition method in ubiquitous scenarios using
Frequency-Modulated Continuous Wave (FMCW) Multiple-Input Multiple-Output (MIMO)
radar. The core idea is to progressively define and classify motions in a cascaded manner,
gradually filtering out non-specific movements, reducing false positives, and enhancing the
applicability of HGR. Specifically, we first propose a suspected hand gesture motion
detection method to help identify suspicious hand gestures. Then, the velocity and position
features of the mutated signal and the stable signal are extracted. A mutated signal motion
recognition method based on a single-layer long short-term memory (LSTM) network is used
to effectively distinguish non-hand gesture motions from hand gestures. Finally, the two-
dimensional trajectory features are extracted, and cascaded with a LSTM network combined
a Gaussian probability model is developed to enhance the ability of open-set recognition.
Experimental results show that the proposed method can achieve the recognition accuracy
of 99.53% for designed hand gestures, the false alarm rate of 1.5% for unexpected hand
gestures and 0.11% for non-hand gesture motions.
Keywords: Hand gesture recognition; Non-hand gesture motion; Feature extraction;
Probability models; Ubiquitous scene
Liang Li, Guochu Chen, Haiyan Wang, Baojiang Li, Bin Wang, Zizhen Yi, Chunbo Zhao,
VHTformer: A joint query perception method for visual-haptic-textual information based on
Transformer,
Applied Soft Computing,
Volume 181,
2025,
113529,
ISSN 1568-4946,
[Link]
([Link]
Abstract: Multimodal information fusion research struggles with aligning heterogeneous
modalities and addressing data imbalance, especially when integrating visual, haptic, and
text—three modalities offering complementary perceptual and semantic features. Current
research focuses on Transformers for unimodal and vision-haptics bimodal tasks, neglecting
tri-modal integration. Leveraging text's semantic bridging capacity could address this
limitation in cross-sensory learning. We propose VHTformer, a Transformer-based
framework designed to unify visual, haptic, and textual modalities via joint query learning.
The model leverages hierarchical attention mechanisms: self-attention refines intra-modal
features (e.g., extracting texture from haptic signals or contextual semantics from text).
Meanwhile, cross-attention aligns spatial-semantic patterns across modalities through
learnable joint queries. This enables synergistic fusion of geometric shapes (vision), material
properties (haptics), and descriptive attributes (text). Experiments were conducted on three
multimodal datasets—ObjectFolder 2.0, Touch and Go, and ObjectFolder Real—covering a
total of 100 + object categories with diverse material and shape properties. To mitigate class
imbalance and ensure statistical reliability, we adopted stratified 5-fold cross-validation. In
addition, we conducted robustness evaluations under Gaussian noise injection to verify the
model's robustness. VHTformer achieves up to 99.55 % recognition accuracy and
demonstrates strong robustness, highlighting the value of tri-modal integration for
comprehensive object understanding.
Keywords: Multimodal information recognition; Cross-modal fusion; Transformer; Attention
mechanism; Semantic alignment
Yongxin Li, Yukun Xue, Zhihui Xin, Guisheng Liao, Penghui Huang,
Multi-modal cross Swin transformer network for multi-label classification landslide detection
with optical and SAR images of Luding,
International Journal of Applied Earth Observation and Geoinformation,
Volume 145,
2025,
104954,
ISSN 1569-8432,
[Link]
([Link]
Abstract: Natural phenomena such as earthquakes and heavy rainfall can trigger landslide
events hundreds or even thousands of times in a given area. Therefore, rapid and efficient
intervention in affected regions is essential. Existing studies have demonstrated good
detection performance using optical remote sensing images; however, optical data has
significant limitations under cloud cover. Moreover, current multi-modal fusion methods
struggle to effectively capture the nonlinear interactions between different modes when
processing data with large informational disparities. To address these issues, we developed
the China Luding multi-modal landslide dataset, which includes optical and polarimetric
synthetic aperture radar (PolSAR) images. We also propose a multi-modal network for
landslide detection based on multi-label classification, called multi-modal cross Swin
Transformer network (MCSTNet). Additionally, a new weighted asymmetric loss function
(WASL) is proposed to improve multi-label classification tasks. Our proposed network
consists of two stages: feature extraction and feature fusion. In the first stage, two
independent branches extract high-level semantic features from optical and PolSAR images.
In the second stage, a multi-modal cross multi-head self-attention (MCMSA) mechanism
fuses the high-level semantic features from the multi-modal information. Therefore, the
model enhances the learning and feature representation capabilities through parallel
processing the input data information in different representation subspaces. Compared to
other recent landslide detection methods, our results indicate that the proposed method is
more effective, achieving an F1 score of 88.24% for landslide detection, with greater
accuracy in identifying landslide-prone areas. Therefore, our model can provide reliable and
effective decision support for disaster emergency responses.
Keywords: Landslide detection; Multi-label classification; Luding multi-modal landslide
dataset; MCSTNet; WASL
Bochao Zou, Zizheng Guo, Jiansheng Chen, Junbao Zhuo, Weiran Huang, Huimin Ma,
RhythmFormer: Extracting patterned rPPG signals based on periodic sparse attention,
Pattern Recognition,
Volume 164,
2025,
111511,
ISSN 0031-3203,
[Link]
([Link]
Abstract: Remote photoplethysmography (rPPG) is a non-contact method for detecting
physiological signals based on facial videos, holding high potential in various applications.
Due to the periodicity nature of rPPG signals, the long-range dependency capturing capacity
of the transformer was assumed to be advantageous for such signals. However, existing
methods have not conclusively demonstrated the superior performance of transformers
over traditional convolutional neural networks. This may be attributed to the quadratic
scaling exhibited by transformer with sequence length, resulting in coarse-grained feature
extraction, which in turn affects robustness and generalization. To address that, this paper
proposes a periodic sparse attention mechanism based on temporal attention sparsity
induced by periodicity. A pre-attention stage is introduced before the conventional attention
mechanism. This stage learns periodic patterns to filter out a large number of irrelevant
attention computations, thus enabling fine-grained feature extraction. Moreover, to address
the issue of fine-grained features being more susceptible to noise interference, a fusion stem
is proposed to effectively guide self-attention towards rPPG features. It can be easily
integrated into existing methods to enhance their performance. Extensive experiments show
that the proposed method achieves state-of-the-art performance in both intra-dataset and
cross-dataset evaluations. The codes are available at
[Link]
Keywords: Remote physiological measurement; Periodic sparse attention
Yiheng Xie, Xiaoping Rui, Yarong Zou, Heng Tang, Ninglei Ouyang, Yingchao Ren,
STDPNet: supervised transformer-driven network for high-precision oil spill segmentation in
SAR imagery,
International Journal of Applied Earth Observation and Geoinformation,
Volume 143,
2025,
104812,
ISSN 1569-8432,
[Link]
([Link]
Abstract: Oil spill incidents are one of the major factors damaging marine ecosystems, and
there is an urgent need for effective detection and identification technologies to quickly
locate oil spill contamination areas. Synthetic Aperture Radar (SAR) is capable of monitoring
the ocean surface under various weather and lighting conditions, but the SAR images often
contain dense speckle noise, and popular SAR oil spill image datasets typically lack sufficient
polarization information. To overcome these issues, this study introduces a novel
polarimetric decomposition method to generate synthetic color image datasets that
integrate multiple polarization features, thereby enhancing image texture and contrast. An
image denoising module is designed, which reduces noise interference in the color images
through an adaptive sampling approach. Furthermore, a novel Transformer-CNN
architecture model is proposed, integrating two modules: the Super Visual Attention
Transformer and the Directional Multi-Branch Scale Self-Calibration Module. The
segmentation performance of the model is comprehensively evaluated on three datasets,
and compared with state-of-the-art segmentation methods, demonstrating superior
classification accuracy and stability. This research provides an effective technical support for
accurate oil spill detection and marine ecosystem protection.
Keywords: Oil spill segmentation; SAR imagery; Proximity pixel compensation method;
Transformer-CNN model; Super visual attention transformer; Directional multi-branch scale
self-calibration method
Yuhao Wu, Bin Li, Jun Li, Yonglou Liang, Naiqiang Zhang, Anlai Sun,
Enhancing nighttime cloud detection for moderate resolution imagers using a transformer
based deep learning network,
Remote Sensing of Environment,
Volume 332,
2026,
115067,
ISSN 0034-4257,
[Link]
([Link]
Abstract: Accurate cloud detection is essential for the quantitative applications of satellite
imager observations, but nighttime cloud detection has challenges due to limited spectral
bands, for example, physical methods using only infrared (IR) bands without using spatial
textures as input for cloud detection often result in high uncertainties, especially in some
situations such as cryosphere surface. Although numerous segmentation-style deep learning
cloud detection algorithms have proposed in previous studies, they are inadequate for
nighttime due to the difficulty in acquiring two-dimensional truth data for training and
validation. To overcome these challenges, the Transformer based Nighttime Cloud Detection
(TNCD) framework, which integrates spatial features and utilizes an advanced Transformer
architecture with relative position encoding, layer scaling, and channel attention
mechanisms, is proposed and investigated for nighttime cloud detection. The model was
trained on labels derived from CALIOP data, utilizing a dataset comprising nearly one
hundred million segments from MODIS. Independent validation indicates that TNCD
achieves robust and consistent performance across various scenarios, with an overall
accuracy (OA) of 93.26 % and over 90 % in cryosphere regions. The proposed algorithm
avoids the pattern noise appeared in the traditional physical methodology due to the
utilization of auxiliary data at coarser resolutions, it also mitigates the negative impact of
stripes in IR images for cloud detection. Moreover, TNCD shows high transferable
practicability across sensors, with over 90 % OA for MERSI. More importantly, our research
underscores the importance of water vapor absorption bands for nighttime cloud detection
over the cryosphere. TNCD's high accuracy and robustness provide unique methodology that
could be used operationally for nighttime cloud detection.
Keywords: Nighttime cloud detection; Transformer; Deep learning; MODIS; MERSI; CALIPSO
Yizhen Jia, Hui Chen, Bang Huang, WenKai Jia, Wen-Qin Wang,
Riemannian gradient deep network for joint waveform and filter optimization in MIMO radar
against chopping forwarding jamming and clutter,
Signal Processing,
Volume 239,
2026,
110257,
ISSN 0165-1684,
[Link]
([Link]
Abstract: With the rise of digital radio frequency memory technology, active deception
jamming poses a significant threat to radar systems, especially in detecting targets amid
mainlobe jamming and non-Gaussian clutter. Traditional methods like space–time matched
filtering struggle in such scenarios. This study introduces the Riemannian gradient deep
network (RGDN), a framework for joint optimization of transmit waveforms and receive
filters to improve target detection. Unlike conventional signal-to-clutter noise ratio (SCNR)
maximization, RGDN leverages information geometry to maximize the Kullback–Leibler
Divergence (KLD) between targets and clutter. By modeling non-Gaussian data with a
Gaussian mixture distribution and constructing a Riemannian manifold, the framework
achieves effective jamming suppression through receive filter term in the loss function,
minimizing jamming effects while enhancing target-clutter distinguishability. To address non-
convex optimization, Riemannian gradient descent is integrated into a deep network.
Numerical experiments show that RGDN achieves superior detection performance compared
to SCNR maximization method.
Keywords: Riemannian gradient; KL divergence; Waveform design; MIMO radar; Mainlobe
deception jamming; Deep learning
Jing Wang, Chao Li, Lu Li, Zhihua Huang, Chao Wang, Hong Zhang, Zhengjia Zhang,
InSAR time-series deformation forecasting surrounding Salt Lake using deep transformer
models,
Science of The Total Environment,
Volume 858, Part 2,
2023,
159744,
ISSN 0048-9697,
[Link]
([Link]
Abstract: The free and open data policy of Sentinel-1 SAR images enables Radar
interferometry (InSAR) to perform time series surface deformation monitoring over large
areas. InSAR deformation monitoring and prediction can investigate the freeze-thaw cycles
of permafrost on the Qinghai-Tibet Plateau. However, the convolutional and recurrent
neural networks cannot accurately model long-term and complex relations in multivariate
time series data, it is challenging to implement time series deformation prediction with high
spatial resolution. In this paper, an innovative InSAR deformation prediction integrated
algorithm based on the transformer models is proposed to predict time series deformation
more accurately surrounding Salt Lake. Compared with the other solutions, the unique
feature of the proposed method is that: 1) this method takes advantage of the self-attention
mechanism to study complicated dynamic deformation features of permafrost caused by
temperature and other variables from InSAR time series deformation. 2) The transformer-
based model can more accurately simulate seasonal and non-seasonal deformation signals,
and is effective for short-term prediction of surface deformation in permafrost areas. The
InSAR deformation prediction results demonstrate that the InSAR deformation prediction
method achieves better prediction performance in predicting the deformation trends of
permafrost with a point scale compared with the prediction results of other models. Based
on the predicted deformation and the water extraction results, the expansion trends
surrounding Salt Lake are discussed and evaluated. The total area of Salt Lake increased by
57.32 km2 during the period 2015–2019. And Salt Lake maintained slowing expansion trend
from 2019 to 2022. The time series deformation forecasting method can be used as a
generic framework for modeling nonlinear deformation processes in complex permafrost
areas, and it reveals the potential impact of the Salt Lake outburst event on the deformation
processes and the degradation of permafrost.
Keywords: InSAR; Qinghai-Tibet Plateau; Deformation prediction; Transformer; Salt Lake;
Permafrost
Menghao Du, Zhenfeng Shao, Xiongwu Xiao, Jindou Zhang, Duowang Zhu, Jinyang Wang,
Timo Balz, Deren Li,
High-precision flood change detection with lightweight SAR transformer network and
context-aware attention for enriched-diverse and complex flooding scenarios,
ISPRS Journal of Photogrammetry and Remote Sensing,
Volume 231,
2026,
Pages 507-531,
ISSN 0924-2716,
[Link]
([Link]
Abstract: Floods are highly destructive natural disasters that threaten both society and the
environment. Given the all-weather, all-time imaging capability of synthetic aperture radar
(SAR), analyzing flood events using SAR imagery across diverse scenarios is essential for
developing high-precision and robust detection models. However, existing transformer-
based change detection methods achieve high precision, but their high computational cost
and large parameter sizes necessitate lightweight design while maintaining detection
accuracy. Moreover, existing studies focus on a few specific scenarios without thorough
validation and in-depth analysis of model strengths across diverse flood conditions with
imbalanced inundation ratios. To address these challenges, this paper proposes an adaptive
window and context-aware attention network (AWCA-Net) for high-precision SAR-based
flood change detection under diverse flooding scenarios, achieving a lightweight model
while maintaining the highest detection accuracy. AWCA-Net has three key advantages:
Firstly, the neighborhood feature enhancement module with contextual information (NECM)
strengthens the discrimination of subtle and heterogeneous flood changes. Secondly, the
large kernel grouping attention gate module based on high-low layer feature difference
(LGDM) leverages difference-weighted attention to effectively guide the selection of flood-
relevant features. Thirdly, the multi-scale convolutional attention module based on adaptive
window selection (MSAWM) dynamically adjusts kernel sizes to capture diverse flood change
patterns. To better train and evaluate AWCA-Net, we constructed VarFloods, the first large-
scale and enriched-diverse benchmark dataset for flood change detection that spans five
continents and includes diverse regions, scenarios, land cover types, causes, years, and
inundation ratios, which includes both GRD and preprocessed versions. We evaluated
AWCA-Net’s performance on three datasets. We found that: (1) AWCA-Net achieves the
highest-precision while maintaining significantly lower computational cost, outperforming
other state-of-the-art (SOTA) methods. On the two representative public datasets and the
enriched-diverse benchmark dataset (VarFloods-G and VarFloods-P), AWCA-Net improves
the IoU by 11.59 % to 42.57 % over the basic model, 2.16 % to 11.48 % over an advanced
transformer-based model, and 1.31 % to 1.68 % over the best comparative model, while
maintaining a computational cost of only 17.53G, which is just 8.7 % to 60 % of existing SOTA
models. (2) Difference-guided attention enhances detection in complex background regions,
neighborhood-based fusion improves performance in irregular terrains, and adaptive
convolution contributes to stable results across diverse flood scenarios. And the proposed
AWCA-Net demonstrates strong generalization and stability under diverse flood scenarios
with imbalanced inundation ratios. The dataset and code of AWCA-Net will be released at:
[Link]
Keywords: High-precision flood change detection; Context-aware attention; Adaptive
window selection; Diverse flood scenarios; VarFloods dataset; Synthetic aperture radar (SAR)
Yihao Wu, Zhenrong Li, Shize Duan, Xuanzhang He, Dawei Dong, Liyan Yu,
A 87.5 – 104.4 GHz broadband power amplifier with RF switch using transmission line
connected in parallel with MCR transformer in 130 nm SiGe BiCMOS,
Microelectronics Journal,
Volume 168,
2026,
106979,
ISSN 1879-2391,
[Link]
([Link]
Abstract: This paper presents a power amplifier (PA) and RF switch architecture designed for
W-band phased array transmit channels. To achieve wideband performance at high
frequencies, an output matching network based on a transmission line connected in parallel
with transformer is proposed, along with a method to improve the flatness of in-band
impedance matching. This approach significantly enhances the current-handling capacity of
the output matching network. To ensure adequate gain, the performance of various active
circuit topologies is analyzed. According to post-layout simulation results, the proposed
circuit achieves a peak gain of 19.8 dB and a 3 dB bandwidth ranging from 87.5 GHz to 104.4
GHz. At 93 GHz, despite the insertion loss introduced by the RF switch, the power amplifier
maintains a saturated output power(Psat) of 16 dBm and achieves a power-added efficiency
(PAE) of 8.4%. S11 remains below −13.2 dB throughout the entire 3 dB bandwidth, reaching
a minimum of −20.9 dB.
Keywords: Power amplifier; RF switch; Broadband; Transformer; Transmission line
Haoming Feng, Huaqing Li, Wenwen Zhu, Denghao Li, Yukun Huang,
Micro-motion enhanced multi-person activity recognition with millimeter-wave radar,
Measurement,
Volume 258, Part B,
2026,
119090,
ISSN 0263-2241,
[Link]
([Link]
Abstract: As a non-contact sensing device, millimeter-wave radar exhibits unique strengths
in human activity recognition (HAR). Existing methods rely on micro-Doppler signatures for
activity classification, but they often encounter feature aliasing in multi-person activity
recognition (MPAR) scenarios. Although point cloud-based approaches can distinguish
individual targets, they primarily extract static morphological features, neglecting the micro-
motion information of human joints, which is crucial for accurate activity recognition. To
address these limitations, we proposes an innovative MPAR framework that integrates
spatial point clouds and micro-motion features. First, an improved point cloud data
association algorithm is applied to achieve multi-target point cloud feature separation,
followed by a dynamic projection mechanism to construct time–Doppler feature maps.
Then, a torso micro-motion enhancement algorithm is designed to enhance the details of
human body movements. Finally, a CNN-LSTM hybrid network architecture with a temporal-
attention is constructed for action classification. Experimental results show that the
proposed micro-motion enhancement algorithm improves recognition accuracy by 27.1%
and 2.3%, compared to two traditional time–frequency analysis methods. Furthermore,
MPAR task in occlusion scenarios achieves recognition accuracy of 93.5%. In summary,
proposed framework not only retains the inherent advantages of millimeter-wave radar but
also significantly enhances multi-person activity recognition in complex scenarios.
Keywords: Human activity recognition (HAR); Multi-person activity recognition (MPAR);
Feature separability; Temporal attention; Occlusion scenarios
Rui He, Tian Peng, Xinyu Zhang, Zhigang Chen, Junhao Yao, Muhammad Shahzad Nazir, Chu
Zhang,
A novel hybrid model for state of health prediction in lithium batteries based on non-
stationary transformers optimized by tree-structured Parzen estimator considering health
factors,
Applied Energy,
Volume 402, Part C,
2026,
127030,
ISSN 0306-2619,
[Link]
([Link]
Abstract: Accurate prediction of State of Health (SOH) in lithium batteries is crucial for
improving the performance, prolonging the service life, preventing failures, and ensuring the
safe use of lithium batteries. This paper proposes a multivariate predictive correction model
for lithium battery SOH based on Time-Varying Filter Empirical Mode Decomposition
(TVFEMD), Pearson Correlation Coefficient (PCC), Kernel Principal Component Analysis
(KPCA), improved Bayesian algorithm, Non-stationary Transformers (NSTransformers), and
Regularized Online Sequential Extreme Learning Machine (ReOSELM). In order to reduce the
complexity of lithium battery data and health factors and to fully extract the features,
multiple methods are used for processing. Firstly, TVFEMD is used for the initial
decomposition of lithium battery health state data, then KPCA is applied to downsize the
decomposed data, and then PCC is selected for correlation analysis of the health factors to
select features with high correlation. Next, the NSTransformers model is employed for
predicting the lithium battery SOH, and a tree-structured Bayesian optimization algorithm,
namely, Tree-structured Parzen Estimator (TPE) is used to optimize the important
parameters of the NSTransformers model, enhancing the model's predictive performance.
Finally, the ReOSELM model is used to correct the initial prediction errors, and the initial
predicted values and error-corrected predicted values are summed to obtain the final
lithium battery SOH prediction values. This paper compares the prediction results of the
multivariate and univariate models. Compared with the other eight multivariate benchmark
models, the MAE and RMSE of the TVFEMD-PCC-KPCA-TPE-NSTransformers-ReOSELM
multivariate model proposed in this paper are reduced by about 0.1 %, and the R are
increased by more than 1 %, which verifies the superiority of the multivariate model
proposed in this paper in the prediction of lithium battery SOH.
Keywords: State of health of Lithium battery; TVFEMD; Tree-structured Parzen estimator;
Non-stationary transformers; ReOSELM
Runwei Guan, Shanliang Yao, Lulu Liu, Xiaohui Zhu, Ka Lok Man, Yong Yue, Jeremy Smith, Eng
Gee Lim, Yutao Yue,
Mask-VRDet: A robust riverway panoptic perception model based on dual graph fusion of
vision and 4D mmWave radar,
Robotics and Autonomous Systems,
Volume 171,
2024,
104572,
ISSN 0921-8890,
[Link]
([Link]
Abstract: With the development of Unmanned Surface Vehicles (USVs), the perception of
inland waterways has become significant to autonomous navigation. RGB cameras can
capture images with rich semantic features, but they would fail in adverse weather and at
night. As a perception sensor that has initially emerged in recent years, 4D millimeter-wave
radar (4D mmWave radar) can work in all weather and has more abundant point-cloud
features than ordinary radar, but it also suffers from water-surface clutter seriously.
Furthermore, the shape and outline of dense point cloud captured by 4D mmWave radar are
irregular. CNN-based neural networks treat features as 2D rectangle grids, which excessively
favor image modality and are unfriendly to radar modality. Therefore, we transform both
features of image and radar into non-Euclidean space as graph structures. In this paper, we
focus on robust panoptic perception in inland waterways. Firstly, we propose the first
Clutter-Point-Removal (CPR) algorithm for 4D mmWave radar, removing water-surface clutter
and improving the recall of radar targets. Secondly, we propose a high-performance
panoptic perception model based on the graph neural network called Mask-VRDet, fusing
features of vision and radar to simultaneously perform object detection and semantic
segmentation. To the best of our knowledge, Mask-VRDet is the first riverway panoptic
perception model based on vision-radar graphical fusion. It outperforms other single-modal
and fusion models, and achieves state-of-the-art performance on our collected dataset. We
release our code at [Link]
Keywords: Riverway panoptic perception; Fusion of vision and radar; Graph convolution
network; Radar clutter removal
Pengcheng Hu, Kai Yang, Heng Wang, Jiadui Chen, Haisong Huang, Jingwei Yang,
Laser welding penetration states recognition based on Gramian Angular Difference Field
images of spectral signals and ConvNeXt model,
Measurement,
Volume 260,
2026,
119786,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Methods using spectral signals for penetration recognition offer significant
research potential. Traditional approaches rely on manual feature extraction, which requires
expertise and is inherently subjective. This research proposes a novel method for visualizing
spectral information to recognize penetration states. In this study, spectral signals were
collected during the laser welding process, and the correlation between the spectral data
and the penetration states was established based on elemental content, boiling point and
atomic transition coefficient. A method based on the Gramian Angular Field was proposed,
converting one-dimensional data into two-dimensional images while largely preserving the
original features. A ConvNeXt-based model was developed and optimized, achieving an
accuracy of 95.34% after 100 rounds of training, which is an increase of 0.46% in accuracy
compared with the method of directly using spectral data for penetration recognition. In
addition, high-precision prediction of weld depth was also achieved based on the spectral
data.
Keywords: Laser welding; Penetration state recognition; Spectral signal; Gramian Angular
Field; ConvNeXt
Kun Qian, Dingwei Zhu, Yutong Wu, Jian Shen, Shoujin Zhang,
TransIST: Transformer based infrared small target tracking using multi-scale feature and
exponential moving average learning,
Infrared Physics & Technology,
Volume 145,
2025,
105674,
ISSN 1350-4495,
[Link]
([Link]
Abstract: Small Unmanned Aerial Vehicle (UAV) tracking against complex sky backgrounds is
of significant importance in both the military and civilian domains, with traditional
correlation filters imposing high requirements on feature models. Consequently, a deep
learning model is designed to achieve effective infrared tracking of a small UAV target.
Specifically, a transformer model is used as the backbone, and a multi-scale attention is
introduced to obtain the perceptual features referring to small targets. Then, an edge
suppression model named side window filter is embedded to suppress the negative effect of
edge on small target tracking. Furthermore, the proposed model is trained using an
exponential moving average learning strategy, which results in more precise network
parameters. Experimental results demonstrate that the proposed Transformer-based
Infrared Small Target Tracking (TransIST) algorithm exhibits superior performance in public
near-infrared videos compared to current related algorithms, with enhanced stability over
correlation filtering-based trackers. The code will be available at [Link]
ayan/TransIST, contributing to the remote sensing community.
Keywords: Infrared tracking; Small objects; Transformer; Multi-scale dilated attention;
Exponential moving average; Edge suppression
Zhuangji Wang, Dennis Timlin, Xiaofei Gong, Yuki Kojima, Shan Hua, David Fleisher,
Wenguang Sun, Sahila Beegum, Vangimalla R. Reddy, Katherine Tully, Robert Horton,
TDR-Transformer: A transformer neural network model to determine soil relative
permittivity variations along a time domain reflectometry sensor waveguide,
Computers and Electronics in Agriculture,
Volume 237, Part C,
2025,
110730,
ISSN 0168-1699,
[Link]
([Link]
Abstract: Interpreting soil relative permittivity (εr) variations along a time domain
reflectometry (TDR) waveguide provides an opportunity to determine soil water content at
multiple depths using a vertically installed TDR sensor. Compared to placing sensors at
different depths, vertical sensor installation reduces measurement efforts and enhances
data-use-efficiency. Revealing εr variations includes two aspects: identifying εr change
positions and determining εr values. Traditional inverse analyses are not widely applied due
to their high computational demands. Machine learning-based methods, e.g., TDR-CNN,
provide a forward computational workflow to track εr change positions and reduce
computational load, but errors in εr values are relatively large. In this study, TDR-
Transformer is developed as a new waveform interpretation model to improve εr estimation
accuracy. Modified from the standard transformer architecture, an encoder with
convolutional neural layers is used to extract waveform geometric features, and a decoder
generates a sequence of εr values to represent εr variations. Attention is a mechanism that
can dynamically extract and process the relevant information within the data, which
processes and integrates the waveform geometric information in the encoder, ensures the
causality (time-order) of the waveform data in the decoder, and transfers information from
the encoder to the decoder. TDR-Transformer was trained and tested using simulated
waveforms where εr changes along the waveguides, but soil electrical conductivity (EC) was
assumed to be small and stable. The RMSE for εr values was within 0.5–1.6 % and the RMSE
of εr change positions was within 5–8 %. A soil infiltration experiment and a precipitation-
evaporation experiment illustrated applications of TDR-Transformer to observed waveforms.
Consequently, TDR-Transformer is a promising artificial intelligence model to interpret TDR
waveforms in soils with nonuniform εr, and fine-tuning TDR-Transformer is recommended
for specific commercial TDR sensor designs.
Keywords: Time Domain Reflectometry (TDR); TDR waveform interpretation; Nonuniform
relative permittivity (εr); Transformer neural network; Machine learning
Caiyi Sun, Dawei Wang, Mingming Xu, Shiqing Wei, Shanwei Liu, Zhongwei Li,
MAF-UFormer: Oil spill detection in SAR images using multi-scale alignment and fusion U-
shaped transformer network,
Regional Studies in Marine Science,
Volume 94,
2026,
104773,
ISSN 2352-4855,
[Link]
([Link]
Abstract: Synthetic aperture radar (SAR) has emerged as a vital technology for detecting oil
spills, even in challenging weather conditions. Deep learning models have demonstrated
significant potential in leveraging SAR images for oil spill detection, owing to their robust
feature extraction capabilities. However, considering the diversity of oil spill target scales
and the extraction of global and local information, there are still particular challenges in
accurately extracting oil spill areas from SAR images. Additionally, polarimetric information
can significantly enhance the separability of oil films and seawater. To overcome these
challenges, a Multi-scale Alignment and Fusion U-Shape Transformer Network (MAF-
UFormer) is proposed, which enhances feature representation by integrating multi-scale
fusion and agent attention mechanisms. To evaluate the effectiveness of MAF-UFormer, we
perform experiments on the publicly available Deep-SAR Oil Spill Detection (SOS) dataset.
The results demonstrate that MAF-UFormer achieves F1-Scores of 87.51 % and 83.03 % on
the Sentinel-1 and PALSAR subsets of SOS, respectively. To further validate the robustness of
MAF-UFormer, we create a new dataset, the Sentinel-1 Oil Spill Detection Dataset Part 1
(S1OSD-1). Experiments on S1OSD-1 demonstrate MAF-UFormer’s superior accuracy in oil
spill detection, outperforming existing methods. Given SAR’s capability to extract
polarimetric features that aid in distinguishing oil spills from seawater, we enhance S1OSD-1
by incorporating polarimetric data to construct Part 2 (S1OSD-2). On S1OSD-2, MAF-
UFormer achieves an additional 1.68 % improvement in F1-Score over S1OSD-1. These
results highlight the potential of MAF-UFormer for oil spill detection, offering vital technical
support for oil spill emergency response and marine environmental protection.
Keywords: Oil spill detection; SAR; Sentinel-1; Polarization feature
Yi Peng, Kui Wang, Chaoyang Wu, Lingyun Kong, Jianming Wu, Zhengyu Ren, Fei Yu, Yiyuan
Duan, Bo Wang, Jiaojiao Wei,
Framework for viscoelastic pavement layer moduli back-calculation from FWD data:
Integrating the spectral element method and Transformer-MLP network,
Construction and Building Materials,
Volume 503,
2025,
144454,
ISSN 0950-0618,
[Link]
([Link]
Abstract: Asphalt pavements experience progressive layer modulus degradation and
frequent distresses as the in-service time increases, leading to higher maintenance costs and
compromising operational safety and service life. Therefore, accurate assessment of layer
moduli and monitoring of structural bearing capacity are essential for achieving long-life
pavements. However, conventional back-calculation methods often assume linear elasticity,
which fails to capture the true viscoelastic pavement structure and often results in limited
accuracy. To address these limitations, this study develops a viscoelastic modulus back-
calculation framework. First, a forward model based on the Spectral Element Method (SEM)
was established to simulate the mechanical response of viscoelastic pavement structures
under Falling Weight Deflectometer (FWD) loading. After model validation, a parameter
sensitivity analysis was conducted to determine the reasonable value ranges for key
parameters, and a comprehensive database of 10,000 simulation-derived pavement cases
was generated. Subsequently, a Transformer-MLP hybrid network model was developed for
back-calculation. The results demonstrate that the model possesses excellent generalization
ability, with the difference in the average R2 between the cross-validation and independent
test sets being less than 1 %. The model achieved high prediction accuracy on the
independent test set, with R2 values of 0.998 for the subgrade modulus, 0.966 for the base
modulus, and 0.887 for the upper surface layer modulus. Furthermore, the model shows
strong robustness, maintaining reliable predictive performance when subjected to simulated
field FWD sensor noise and load fluctuations. By integrating the SEM-based forward model
with the Transformer-MLP network, this framework enables the accurate and efficient
evaluation of pavement structural bearing capacity, providing significant practical value for
construction quality control, structural monitoring, and maintenance decision-making.
Keywords: Spectral element method; Falling weight deflectometer; Layer modulus back-
calculation; Transformer-MLP; Hybrid network
Abel Corrêa Dias, Viviane Pereira Moreira, João Luiz Dihl Comba,
RoBIn: A Transformer-based model for risk of bias inference with machine reading
comprehension,
Journal of Biomedical Informatics,
Volume 166,
2025,
104819,
ISSN 1532-0464,
[Link]
([Link]
Abstract: Objective:
Scientific publications are essential for uncovering insights, testing new drugs, and informing
healthcare policies. Evaluating the quality of these publications often involves assessing their
Risk of Bias (RoB), a task traditionally performed by human reviewers. The goal of this work
is to create a dataset and develop models that allow automated RoB assessment in clinical
trials.
Methods:
We use data from the Cochrane Database of Systematic Reviews (CDSR) as ground truth to
label open-access clinical trial publications from PubMed. This process enabled us to
develop training and test datasets specifically for machine reading comprehension and RoB
inference. Additionally, we created extractive (RoBInExt) and generative (RoBInGen)
Transformer-based approaches to extract relevant evidence and classify the RoB effectively.
Results:
RoBIn was evaluated across various settings and benchmarked against state-of-the-art
methods, including large language models (LLMs). In most cases, the best-performing RoBIn
variant surpasses traditional machine learning and LLM-based approaches, achieving a
AUROC of 0.83.
Conclusion:
This work addresses RoB assessment in clinical trials by introducing RoBIn, two Transformer-
based models for RoB inference and evidence retrieval, which outperform traditional models
and LLMs, demonstrating its potential to improve efficiency and scalability in clinical
research evaluation. We also introduce a public dataset that is automatically annotated and
can be used to enable future research to enhance automated RoB assessment.
Keywords: Evidence-based medicine; Systematic reviews; Risk of bias; Deep learning; Natural
language processing; Machine reading comprehension; Classification models
Baiju Yan, Hao Zhang, Yicheng Yao, Changyu Liu, Pu Jian, Peng Wang, Lidong Du, Xianxiang
Chen, Zhen Fang, Yirong Wu,
Heart signatures: Open-set person identification based on cardiac radar signals,
Biomedical Signal Processing and Control,
Volume 72, Part A,
2022,
103306,
ISSN 1746-8094,
[Link]
([Link]
Abstract: Objective
Non-contact continuous biometric identification system based on the heartbeat signals has
attracted more attention due to its privacy-friendly properties. Current methods, however,
mostly focused on the traditional machine learning methods under the close-set condition.
This paper aims to investigate the feasibility of using the cardiac radar heartbeat signals and
deep learning techniques to identify person in the open-set environment.
Methods
A novel dipole deep learning model (DDLM) was proposed for the open-set person
identification problem without heartbeat segmentation and feature engineering. The
normalized heartbeat samples with time duration of 5 s were used as input and encoded
into the feature space, where the encoded features of the same person cluster closely
around the corresponding negative pole and repel far from positive pole, and those of
different persons separate loosely from each other. Finally, threshold on the distance from
the features to the dipoles in the feature space was set for each known identity.
Results
Extensive experiments conducted on a public dataset of clinically recorded vital signs
indicate that:(1) The proposed model shows high stability under close-set condition with an
accuracy higher than 99 % with 30 subjects.(2) The accuracy and the F1-score attain 93.42 %
and 93.57 % under the open-set condition with the maximum openness of 29.3 %,
respectively.
Conclusion
The proposed model shows high effectiveness in person identification using heartbeat
signals. The DDLM outperforms most of the existing methods under the close-set condition.
And the DDLM shows a promising future for person identification in open-set environment.
Keywords: Cardiac radar heartbeat signatures; Dipole deep learning model (DDLM); Machine
learning; Non-contact biometric identification; Open set identification
Yunfeng Fang, Zheng Tong, Tianqing Hei, Siqi Wang, Tao Ma,
Deep learning applications in ground-penetrating radar inversion: A review,
Measurement,
Volume 258, Part D,
2026,
119399,
ISSN 0263-2241,
[Link]
([Link]
Abstract: The complex nonlinear relationship between the subsurface medium and ground-
penetrating radar signals results in the pervasive ill-posedness and non-uniqueness of
conventional inversion methods. Deep learning, with its powerful feature extraction
capabilities and advantages in modeling complex nonlinear relationships, has unique
strengths in handling complex signals and nonlinear problems, making it especially suitable
for GPR inversion tasks. This paper reviews the latest applications of deep learning in GPR
inversion, summarizing the application strategies of deep learning from two perspectives:
data-driven and data-physics hybrid-driven. Commonly used model architectures and their
performance in signal feature extraction, multi-scale information fusion, and data
preprocessing are discussed, along with the application of various loss functions in inversion
tasks. Finally, current challenges, such as limited model generalization, model dependence
on the dataset and computational efficiency constraints, are discussed, and potential future
research directions are proposed to further advance deep learning in GPR inversion.
Keywords: Ground-penetrating radar; Deep learning; Inversion
Tianyang Li, Chao Wang, Sirui Tian, Bo Zhang, Fan Wu, Yixian Tang, Hong Zhang,
TACMT: Text-aware cross-modal transformer for visual grounding on high-resolution SAR
images,
ISPRS Journal of Photogrammetry and Remote Sensing,
Volume 222,
2025,
Pages 152-166,
ISSN 0924-2716,
[Link]
([Link]
Abstract: This paper introduces a novel task of visual grounding for high-resolution synthetic
aperture radar images (SARVG). SARVG aims to identify the referred object in images
through natural language instructions. While object detection on SAR images has been
extensively investigated, identifying objects based on natural language remains under-
explored. Due to the unique satellite view and side-look geometry, substantial expertise is
often required to interpret objects, making it challenging to generalize across different
sensors. Therefore, we propose to construct a dataset and develop multimodal deep
learning models for the SARVG task. Our contributions can be summarized as follows. Using
power transmission tower detection as an example, we have built a new benchmark of
SARVG based on images from different SAR sensors to fully promote SARVG research.
Subsequently, a novel text-aware cross-modal Transformer (TACMT) is proposed which
follows DETR’s architecture. We develop a cross-modal encoder to enhance the visual
features associated with the textual descriptions. Next, a text-aware query selection module
is devised to select relevant context features as the decoder query. To retrieve the object
from various scenes, we further design a cross-scale fusion module to fuse features from
different levels for accurate target localization. Finally, extensive experiments on our dataset
and widely used public datasets have demonstrated the effectiveness of our proposed
model. This work provides valuable insights for SAR image interpretation. The code and
dataset are available at [Link]
Keywords: Synthetic aperture radar (SAR); Power transmission tower; Visual grounding;
Multimodal; Deep learning
Raihan Ahamed Rifat, Fuyad Hasan Bhoyan, Md Humaion Kabir Mehedi, Md Kaviul Hossain,
Md. Jakir Hossen, M.F. Mridha,
ConMatFormer: A multi-attention and transformer integrated ConvNext based deep learning
model for enhanced diabetic foot ulcer classification,
Results in Engineering,
Volume 28,
2025,
108248,
ISSN 2590-1230,
[Link]
([Link]
Abstract: Diabetic foot ulcer (DFU) detection is a clinically significant yet challenging task due
to the scarcity and variability of publicly available datasets. Limited annotated samples
restrict the ability of conventional deep learning models to achieve robust generalization in
real-world clinical scenarios. To solve these problems, we propose ConMatFormer, a new
hybrid deep learning architecture that combines ConvNeXt blocks, multiple attention
mechanisms convolutional block attention module (CBAM) and dual attention network
(DANet), and transformer modules in a way that works together. This design facilitates the
extraction of better local features and understanding of the global context, which allows us
to model small skin patterns across different types of DFU very accurately. To address the
class imbalance, we used data augmentation methods. A ConvNeXt block was used to obtain
detailed local features in the initial stages. Subsequently, we compiled the model by adding a
transformer module to enhance long-range dependency. This enabled us to pinpoint the
DFU classes that were underrepresented or constituted minorities. Tests on the DS1
(DFUC2021) and DS2 (diabetic foot ulcer (DFU)) datasets showed that ConMatFormer
outperformed state-of-the-art (SOTA) convolutional neural network (CNN) and Vision
Transformer (ViT) models in terms of accuracy, reliability, and flexibility. The proposed
method achieved an accuracy of 0.8961 and a precision of 0.9160 in a single experiment,
which is a significant improvement over the current standards for classifying DFUs. In
addition, by 4-fold cross-validation, the proposed model achieved an accuracy of 0.9755
with a standard deviation of only 0.0031. We further applied explainable artificial
intelligence (XAI) methods, such as Grad-CAM, Grad-CAM++, and LIME, to consistently
monitor the transparency and trustworthiness of the decision-making process. These
human-readable tools enhance the comprehension of the explanations and can substantially
increase the practical use of our methodology. Our findings set a new benchmark for DFU
classification and provide a hybrid attention transformer framework for medical image
analysis.
Keywords: Diabetic foot ulcers classification; Multi-attention; Transformer; GradCam;
Explainable AI; LIME
Van Ngoc Dang, Ngoc Chau Hoang, Quoc Cuong Nguyen, Minh Thuy Le,
Advancing robust human activity recognition via informative mmWave radar characteristics
and a lightweight spatio-spectro-temporal network,
Measurement,
Volume 256, Part A,
2025,
118056,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Human activity recognition (HAR) is increasingly important in aiding our daily life,
with millimeter-wave (mmWave) radar sensors emerging as a promising noninvasive solution
thanks to their excellent spatial and velocity resolution. Although existing radar-based
systems have shown strong performance, they primarily focus on micro-Doppler signatures
while neglecting angle information, which can hinder practical deployment in real-world
scenarios. Moreover, current state-of-the-art recognition models using mmWave radar often
require substantial computational resources, making integration into resource-constrained
devices challenging. This work proposes an efficient radar-based HAR system that leverages
angle and spectro-temporal information from micro-Doppler signatures. Our system utilizes
a multi-channel micro-Doppler representation corresponding to the number of virtual
antenna receivers as input. Then, a lightweight dilated convolutional network, namely SST-
DCN, extracts spatial-aware multi-scale spectro-temporal information through time-
frequency dilated convolutions. Experimental results on our real-world dataset demonstrate
the superiority of our approach compared to conventional features and other state-of-the-
art radar-based HAR systems.
Keywords: Human activity recognition; Millimeter-wave radar; Deep learning; Lightweight
network; Dilated convolution
Haoyu Jiang, Xiaoliang Chen, Duoqian Miao, Hongyun Zhang, Xiaolin Qin, Shangyi Du, Peng
Lu,
3WD-DRT: A three-way decision enhanced dynamic routing transformer for cost-sensitive
multimodal sentiment analysis,
Information Sciences,
Volume 725,
2026,
122704,
ISSN 0020-0255,
[Link]
([Link]
Abstract: Accurately interpreting human emotion from language, facial expressions, and
vocal tones remains a fundamental challenge in artificial intelligence. Current Multimodal
Sentiment Analysis (MSA) models often struggle with two key issues. First, their static fusion
strategies fail to handle conflicting modalities, such as sarcasm. Second, their standard loss
functions ignore the asymmetric risks of severe misjudgments. To address these limitations,
we propose the Three-Way Decision Enhanced Dynamic Routing Transformer (3WD-DRT), a
framework operating on a "quality-aware, decision-driven" principle. It dynamically assesses
each modality’s quality using a three-way decision gate, implemented via a dedicated MLP,
to partition information into acceptance, deferment, or rejection pathways. This enables the
model to amplify informative signals, moderately scale uncertain ones (deferment), and
attenuate noisy or misleading ones. We also introduce a novel cost-sensitive loss function
that imposes greater penalties on major semantic errors, such as polarity misclassifications.
This approach better aligns the model’s training objective with human perception. Extensive
experiments on CH-SIMS, CH-SIMSv2, MOSI, and MOSEI datasets show that 3WD-DRT
consistently outperforms state-of-the-art methods, setting new benchmarks with F1-scores
of 87.08 % on MOSI and 88.26 % on MOSEI. This work provides a robust solution for MSA,
fostering more nuanced and reliable emotionally-aware AI systems.
Keywords: Multimodal sentiment analysis (MSA); Three-way decision theory; Dynamic
routing transformer; Emotion-aware fusion
Wei Wang, Ruobing Song, Yunxiao Wu, Li Zheng, Wenyu Zhang, Zhaoxi Chen, Gang Li, Zhifei
Xu,
Deep learning-based automated diagnosis of obstructive sleep apnea and sleep stage
classification in children using millimeter-wave radar and pulse oximeter,
Sleep Health,
Volume 11, Issue 6,
2025,
Pages 859-867,
ISSN 2352-7218,
[Link]
([Link]
Abstract: Study objectives
Due to the high cost, complexity, and workload of polysomnography, a radar-based sleep
monitoring device, QSA600, has been developed as a more simplified alternative for
children. This study evaluates its agreement with polysomnography for obstructive sleep
apnea diagnosis and sleep staging.
Methods
This diagnostic accuracy study included 281 children (1-18 years) who underwent
simultaneous polysomnography and QSA600 monitoring at Beijing Children's Hospital from
September-November 2023. QSA600 recordings were automatically analyzed using a deep
learning model, while polysomnography data were manually scored.
Results
The obstructive apnea-hypopnea index (OAHI) obtained from QSA600 and polysomnography
demonstrates a high level of agreement with an intraclass correlation coefficient of 0.945
(95% CI: 0.93-0.96). Bland-Altman analysis indicated that the mean difference of obstructive
apnea-hypopnea index between QSA600 and polysomnography was −0.10 events/h (95% CI:
−11.15 to 10.96). The deep learning model evaluated through cross-validation showed good
sensitivity (81.8%, 84.3%, and 89.7%) and specificity (90.5%, 95.3%, and 97.1%) values for
diagnosing children with OAHI >1, OAHI >5, and OAHI >10. The area under the receiver
operating characteristic curve was 0.923, 0.955, and 0.988, respectively. For sleep stage
classification, the model achieved Kappa coefficients of 0.854, 0.781, and 0.734, with
corresponding overall accuracies of 95.0%, 84.8%, and 79.7% for Wake-Sleep classification,
Wake-REM-Light-Deep classification, and Wake-REM-N1-N2-N3 classification, respectively.
Conclusions
QSA600 has demonstrated high agreement with polysomnography in diagnosing obstructive
sleep apnea and performing sleep staging in children. The device is portable, low-burden,
and suitable for follow-up and long-term pediatric sleep assessment.
Keywords: Obstructive sleep apnea; Children; Deep learning; Millimeter-wave radar;
Portable sleep monitoring device; Polysomnography
Chen Jiang, Shuxia Lu, Xianghu Zhou, Tingting Ma, Junhai Zhai,
KNN improved Transformer for 3D object detection,
Signal Processing: Image Communication,
Volume 142,
2026,
117488,
ISSN 0923-5965,
[Link]
([Link]
Abstract: In recent years, 3D object detection in autonomous driving perception has gained
significant attention in the industry. Due to its characteristics, LiDAR has become the most
commonly used and essential sensor. However, voxel-based networks often lose context
information during the voxelization process, which negatively impacts the detection of small
objects. In this paper, we address the challenge of low accuracy in LiDAR-based detection,
especially for small object categories, by proposing an improved Transformer structure.
Transformers are a type of deep learning model known for their ability to capture long-range
dependencies and contextual relationships in data. In our approach, we incorporate a k-
Nearest Neighbors (KNN) algorithm, which is a method for identifying the closest points in
space, to enhance the spatial relationships between point clouds. This combination allows
the model to better capture context information, strengthen feature extraction, and
significantly reduce both missed and false detections. Our method is designed to be plug-
and-play, allowing it to be directly applied to existing point cloud detectors. We evaluate our
approach on the public KITTI and Astyx datasets. Experimental results show significant
improvements, especially in detecting small object categories, even in challenging
conditions.
Keywords: Autonomous vehicle; Transformer; Lidar; 3D object detection; Point cloud
Tamer Saleh, Xingxing Weng, Shimaa Holail, Chen Hao, Gui-Song Xia,
DAM-Net: Flood detection from SAR imagery using differential attention metric-based vision
transformers,
ISPRS Journal of Photogrammetry and Remote Sensing,
Volume 212,
2024,
Pages 440-453,
ISSN 0924-2716,
[Link]
([Link]
Abstract: Flood detection from synthetic aperture radar (SAR) imagery plays an important
role in crisis and disaster management. Based on pre- and post-flood SAR images, flooded
areas can be extracted by detecting changes of water bodies. Existing state-of-the-art
change detection methods primarily target optical image pairs. The nature of SAR images,
such as scarce visual information, similar backscatter signals, and ubiquitous speckle noise,
pose great challenges to identifying water bodies and mining change features, thus resulting
in unsatisfactory performance. Besides, the lack of large-scale annotated datasets hinders
the development of accurate flood detection methods. In this paper, we focus on the
difference between SAR image pairs and present a differential attention metric-based
network (DAM-Net), to achieve flood detection. By introducing feature interaction during
temporal-wise feature representation, we guide the model to focus on changes of interest
rather than fully understanding the scene of the image. On the other hand, we devise a class
token to capture high-level semantic information about water body changes, increasing the
ability to distinguish water body changes and pseudo changes caused by similar signals or
speckle noise. To better train and evaluate DAM-Net, we create a large-scale flood detection
dataset using Sentinel-1 SAR imagery, namely S1GFloods. This dataset consists of 5,360
image pairs, covering 46 flood events during 2015–2022, and spanning 6 continents of the
world. The experimental results on this dataset demonstrate that our method outperforms
several advanced change detection methods. DAM-Net achieves 97.8% overall accuracy,
96.5% F1, and 93.2% IoU on the test set. Our dataset and code are available at
[Link]
Keywords: Flood detection; SAR imagery; S1GFloods dataset; Vision transformers
Jia Liu, Hang Gu, Fangmei Liu, Hao Chen, Zuhe Li, Gang Xu, Qidong Liu, Wei Wang,
CE-CDNet: A Transformer-Based Channel Optimization Approach for Change Detection in
Remote Sensing,
Computers, Materials and Continua,
Volume 83, Issue 1,
2025,
Pages 803-822,
ISSN 1546-2218,
[Link]
([Link]
Abstract: In recent years, convolutional neural networks (CNN) and Transformer
architectures have made significant progress in the field of remote sensing (RS) change
detection (CD). Most of the existing methods directly stack multiple layers of Transformer
blocks, which achieves considerable improvement in capturing variations, but at a rather
high computational cost. We propose a channel-Efficient Change Detection Network (CE-
CDNet) to address the problems of high computational cost and imbalanced detection
accuracy in remote sensing building change detection. The adaptive multi-scale feature
fusion module (CAMSF) and lightweight Transformer decoder (LTD) are introduced to
improve the change detection effect. The CAMSF module can adaptively fuse multi-scale
features to improve the model’s ability to detect building changes in complex scenes. In
addition, the LTD module reduces computational costs and maintains high detection
accuracy through an optimized self-attention mechanism and dimensionality reduction
operation. Experimental test results on three commonly used remote sensing building
change detection data sets show that CE-CDNet can reduce a certain amount of
computational overhead while maintaining detection accuracy comparable to existing
mainstream models, showing good performance advantages.
Keywords: Remote sensing; change detection; attention mechanism; channel optimization;
multi-scale feature fusion
Derek Ka-Hei Lai, Li-Wen Zha, Tommy Yau-Nam Leung, Andy Yiu-Chau Tam, Bryan Pak-Hei So,
Hyo-Jung Lim, Daphne Sze Ki Cheung, Duo Wai-Chi Wong, James Chung-Wai Cheung,
Dual ultra-wideband (UWB) radar-based sleep posture recognition system: Towards
ubiquitous sleep monitoring,
Engineered Regeneration,
Volume 4, Issue 1,
2023,
Pages 36-43,
ISSN 2666-1381,
[Link]
([Link]
Abstract: Sleep posture monitoring is an essential assessment for obstructive sleep apnea
(OSA) patients. The objective of this study is to develop a machine learning-based sleep
posture recognition system using a dual ultra-wideband radar system. We collected
radiofrequency data from two radars positioned over and at the side of the bed for 16
patients performing four sleep postures (supine, left and right lateral, and prone). We
proposed and evaluated deep learning approaches that streamlined feature extraction and
classification, and the traditional machine learning approaches that involved different
combinations of feature extractors and classifiers. Our results showed that the dual radar
system performed better than either single radar. Predetermined statistical features with
random forest classifier yielded the best accuracy (0.887), which could be further improved
via an ablation study (0.938). Deep learning approach using transformer yielded accuracy of
0.713.
Keywords: Obstructive sleep apnea; Deep learning; Sleep monitoring; Feature extraction;
Ablation study
Shuochen Han, Zhonghao Wang, Guochang Zhang, Chengyang Li, Haitao Zhu, Yanyan Wang,
ROV Trajectory Prediction Algorithm Based on Transformer-LSTM,
IFAC-PapersOnLine,
Volume 59, Issue 35,
2025,
Pages 454-459,
ISSN 2405-8963,
[Link]
([Link]
Abstract: This study proposes a Transformer-LSTM hybrid algorithm to address the
degradation in trajectory prediction accuracy for Remotely Operated Vehicles (ROVs) caused
by umbilical cable communication delays during collaborative operations with Unmanned
Surface Vehicles (USVs). The methodology integrates historical USV observation data with
ROV motion characteristics to construct multimodal feature vectors, employing the
Transformer’s multi-head self-attention mechanisms for global trajectory feature extraction
and LSTM networks for local temporal dependency optimization. A compensation
mechanism utilizing historical predictions ensures stability during USV observation failures.
Experimental results demonstrate that the proposed approach significantly outperforms
baseline methods in both normal and failure scenarios.
Keywords: ROV; trajectory prediction; Transformer; LSTM; cooperative operation
Wang Zhang, Tingting Li, Yuntian Zhang, Gensheng Pei, Xiruo Jiang, Yazhou Yao,
LTFormer: A light-weight transformer-based self-supervised matching network for
heterogeneous remote sensing images,
Information Fusion,
Volume 109,
2024,
102425,
ISSN 1566-2535,
[Link]
([Link]
Abstract: Matching visible and near-infrared (NIR) images is a major challenge in remote
sensing image fusion due to nonlinear radiometric differences. Deep learning has shown
promise in computer vision, but most methods rely on supervised learning with limited
annotated data in remote sensing. To address this, we propose a novel keypoint descriptor
approach that obtains robust feature descriptors via a self-supervised matching network.
Our light-weight transformer network, LTFormer, generates deep-level feature descriptors.
Furthermore, we implement an innovative triplet loss function, LT Loss, to enhance the
matching performance further. Our approach outperforms conventional hand-crafted local
feature descriptors and proves equally competitive compared to state-of-the-art deep
learning-based methods, even amidst the shortage of annotated data. Code and pre-trained
model are available at [Link]
Keywords: Image matching; Transformer; Light-weight; Heterogeneous remote sensing
images; Self-supervised learning
Wenhao Dong, Yueyang Li, Weiming Zeng, Lei Chen, Hongjie Yan, Wai Ting Siok, Nizhuan
Wang,
STARFormer: A novel spatio-temporal aggregation reorganization transformer of FMRI for
brain disorder diagnosis,
Neural Networks,
Volume 192,
2025,
107927,
ISSN 0893-6080,
[Link]
([Link]
Abstract: Many existing methods that use functional magnetic resonance imaging (fMRI) to
classify brain disorders, such as autism spectrum disorder (ASD) and attention deficit
hyperactivity disorder (ADHD), often overlook the integration of spatial and temporal
dependencies of the blood oxygen level-dependent (BOLD) signals, which may lead to
inaccurate or imprecise classification results. To solve this problem, we propose a spatio-
temporal aggregation reorganization transformer (STARFormer) that effectively captures
both spatial and temporal features of BOLD signals by incorporating three key modules. The
region of interest (ROI) spatial structure analysis module uses eigenvector centrality (EC) to
reorganize brain regions based on effective connectivity, highlighting critical spatial
relationships relevant to the brain disorder. The temporal feature reorganization module
systematically segments the time series into equal-dimensional window tokens and captures
multiscale features through variable window and cross-window attention. The spatio-
temporal feature fusion module employs a parallel transformer architecture with dedicated
temporal and spatial branches to extract integrated features. The proposed STARFormer has
been rigorously evaluated on two publicly available datasets for the classification of ASD and
ADHD. The experimental results confirm that STARFormer achieves state-of-the-art
performance across multiple evaluation metrics, providing a more accurate and reliable tool
for the diagnosis of brain disorders and biomedical research. The official implementation
codes are available at: [Link]
Keywords: Brain disorder diagnosis; fMRI; Eigenvector centrality; Spatio-temporal
information integration; Transformer
Ibrahim Fayad, Philippe Ciais, Martin Schwartz, Jean-Pierre Wigneron, Nicolas Baghdadi,
Aurélien de Truchis, Alexandre d'Aspremont, Frederic Frappart, Sassan Saatchi, Ewan Sean,
Agnes Pellissier-Tanon, Hassan Bazzi,
Hy-TeC: a hybrid vision transformer model for high-resolution and large-scale mapping of
canopy height,
Remote Sensing of Environment,
Volume 302,
2024,
113945,
ISSN 0034-4257,
[Link]
([Link]
Abstract: Accurate and timely monitoring of forest canopy height is critical for assessing
forest dynamics, biodiversity, carbon sequestration as well as forest degradation and
deforestation. Recent advances in deep learning techniques, coupled with the vast amount
of spaceborne remote sensing data offer an unprecedented opportunity to map canopy
height at high spatial and temporal resolutions. Current techniques for wall-to-wall canopy
height mapping correlate remotely sensed information from optical and radar sensors in the
2D space to the vertical structure of trees using lidar's 3D measurement abilities serving as
height proxies. While studies making use of deep learning algorithms have shown promising
performances for the accurate mapping of canopy height, they have limitations due to the
type of architectures and loss functions employed. Moreover, mapping canopy height over
tropical forests remains poorly studied, and the accurate height estimation of tall canopies is
a challenge due to signal saturation from optical and radar sensors, persistent cloud cover,
and sometimes limited penetration capabilities of lidar instruments. In this study, we map
heights at 10 m resolution across the diverse landscape of Ghana with a new vision
transformer (ViT) model, dubbed Hy-TeC, optimized concurrently with a classification
(discrete) and a regression (continuous) loss function. This model achieves significantly
higher accuracy than previously employed convolutional-based approaches (ConvNets)
optimized with only a continuous loss function. Hy-TeC results show that our proposed
discrete/continuous loss formulation significantly increases the sensitivity for very tall trees
(i.e., > 35 m). Overall, Hy-TeC has significantly reduced bias (0.8 m) and higher accuracy
(RMSE = 6.6 m) over tropical forests for which other approaches show poorer performance
and oftentimes a saturation effect. The height maps generated by Hy-TeC also have better
ground sampling distance and better sensitivity to sparse vegetation. Over these areas, Hy-
TeC showed an RMSE of 3.1 m in comparison to a reference dataset while the baseline
ConvNet model had an RMSE of 4.3 m. Hy-TeC, which was used to generate a height map of
Ghana using free and open access remotely sensed data with Sentinel-2 and Sentinel-1
images as predictors and GEDI height measurements as calibration data, has the potential to
be used globally.
Keywords: Canopy height; GEDI; Sentinel-1; Sentinel-2; Vision transformers, deep learning,
knowledge distillation
Tao Tan, Xiuping Li, Yubing Li, Shuai Wu, Bohai Fang, Changkai Zhang, Yujian Qin,
A transformer-based current reuse CMOS Armstrong VCO using gm-boosted technique,
Microelectronics Journal,
Volume 163,
2025,
106753,
ISSN 1879-2391,
[Link]
([Link]
Abstract: This paper presents a low power LC Armstrong VCO based on current reuse
structure and gm-boosted technique with a 5-port 3-coil transformer. The applied current
reuse structure converts the traditional parallel cross-coupled transistor pair into a series
PMOS-NMOS pair, reducing the overall power consumption. Moreover, the phase-shifted
between ids and Vds of the transistors reduces power dissipated in the devices by leveraging
the large language model optimization results. The proposed gm-boosted technique is
employed by two transistors in series and capacitors in parallel. With the proposed topology,
the negative conductance, oscillation amplitude, and the Perturbation Projection Vector
(PPV) are improved, thus better phase noise performance. The 5-port 3-coil transformer
integrates all inductors to minimizing the area. Fabricated in GlobalFoundries 0.11-μm CMOS
technology, the VCO achieves a power consumption of 2 mW at a 1.2 V supply voltage, with
a phase noise of −114.5 dBc/Hz @1MHz offset at 11.2 GHz.
Keywords: Armstrong VCO; gm-boosted; Current reuse; 5-port transformer; Voltage-
controlled oscillator (VCO)
Senguo Cao, Congde Lu, Xiao Wang, Peng Zhang, Guanglai Jin, Wenlong Cai,
ME-YOLO: A novel real-time detection network for pavement interlayer distress using
ground-penetrating radar,
Journal of Applied Geophysics,
Volume 245,
2026,
106057,
ISSN 0926-9851,
[Link]
([Link]
Abstract: Interlayer distress detection in asphalt pavement is critical for highway
maintenance, as timely identification of pavement distress can ensure operational safety,
reliability, and extended service life. However, the problems of feature information loss and
the substantial confusable backgrounds significantly hinder detection accuracy. To address
these limitations, we propose an enhanced network specifically designed for automated
interlayer distress detection named ME-YOLO. Firstly, we design a Multiscale Adaptive
Feature Fusion (MAFF) module, which aggregates more scale information by Adaptive
Spatial Feature Fusion (ASFF). This design links all feature scales to make discriminative
features in each scale propagate directly to subsequent modules, enriching semantic
representations and mitigating the risk of feature loss, while leveraging shallow-layer
features to strengthen spatial localization. Furthermore, the Efficient Partial Self-Attention
(EPSA) module is introduced to suppress background interference in complex environments.
Unlike conventional transformers, EPSA adopts partial self-attention operations with multi-
path fusion, which can enable the network to acquire global representation capability with
low computational overhead. Extensive experiments indicate that the ME-YOLO network
outperforms the given state-of-the-art models, including Faster-RCNN, RT-DETR, YOLOv8s,
and YOLOv11s, on the interlayer distress dataset. Compared to YOLOv5s, ME-YOLO achieves
improvements of 2.2% in mAP0.5 and 3.5% in mAP0.5:0.95, while maintaining an inference
speed of 6.7 ms per image. The source code will be available at
[Link]
Keywords: Asphalt pavement; Interlayer distress; Ground penetrating radar; Transformer;
Multiscale feature fusion
Zhen Wen, Zongxuan Li, Shuping Tao, Yu Zhao, Yifan Li, Xinlong Wang,
Hybrid Mamba-Transformer network for phase unwrapping in optical interferometry,
Optics Communications,
Volume 601,
2026,
132726,
ISSN 0030-4018,
[Link]
([Link]
Abstract: Phase unwrapping (PU) is crucial in optical interferometry, as accurate phase
information directly affects quantitative analysis and precise reconstruction quality.
Conventional PU methods suffer from performance degradation under severe noise or
undersampling. With the rise of deep learning, recent advancements in PU have been
improved upon CNN and Transformer-based frameworks. Nonetheless, CNNs lack sufficient
capability to model spatial dependencies of wrapped phase, while Transformer architectures
suffer from quadratic computational complexity. State space models like Mamba have
recently become attractive solutions due to efficient linear complexity in capturing long-
range dependencies. Inspired by this, we propose HMTPU, an innovative hybrid architecture
for PU that integrates the strengths of Mamba and Transformer to achieve high performance
with computational efficiency. Specifically, we integrate Transformer layers after Mamba
layers to strengthens the model’s capability in modeling long-range spatial relationships and
improves its effectiveness in processing local wrapped phase information. Within the
Mamba, a geometric transformable selective scan module is designed to enhance the
acquisition of global spatial context through efficient state space modeling. And a
deformable local enhanced window attention is introduced to refine local representations
and handle structural variations in the Transformer. Additionally, we employ an enhanced
feedforward network that leverages context broadcasting and the gating mechanism to
facilitate efficient cross-channel interaction. Extensive experiments demonstrate that
HMTPU outperforms current advanced PU techniques. Testing on real-world datasets of
dynamic candle flames and holographic tomography shows the generalization capability of
our PU method.
Keywords: Phase unwrapping; Mamba; Transformer; Optical interferometry
M. Mortazavi, Z. Moravej, G.B. Gharehpetian,
Detection and localization of LV winding radial deformation in transformers using
electromagnetic waves - a feasibility study,
International Journal of Electrical Power & Energy Systems,
Volume 155, Part B,
2024,
109602,
ISSN 0142-0615,
[Link]
([Link]
Abstract: Recently, online methods based on electromagnetic waves have been proposed to
detect the mechanical defects of high voltage (HV) windings in power transformers. In this
article, for the first time, the possibility of online detecting and locating the radial
deformation (RD) of low voltage (LV) windings using electromagnetic waves is presented. In
the proposed method, a high frequency and wideband signal is sent to a simplified model of
transformer winding via small antennas. The antennas are connected to the inner side of the
oil tank cover through Radio Frequency (RF) cables, so that online monitoring can be realized
with minimal changes in the transformer structure. The data of the reflected signals in
different states is recorded in a database and sorted as primary data bank. By comparing the
results of the sound state with measured signals, the possible faults can be detected. In this
paper, the detection and location of radial deformation is estimated by using regression tree
based on three features. Several cases are simulated by Computer Simulation Technology
(CST) software, and their verification are conducted by a setup in laboratory. Based on
comparison results, it can be claimed that the proposed method has LV windings radial
deformation online detecting and locating ability with an acceptable accuracy.
Keywords: Power transformers; LV winding; Electromagnetic Waves; Regression; Radial
deformation; Wideband antenna
Yukai Kong, Xianxiang Yu, Jiachen Li, Kui Xiong, Guolong Cui,
Non-uniform pulse intervals based intra-pulse forwarding jamming detection and
recognition in clutter circumstance,
Signal Processing,
Volume 238,
2026,
110193,
ISSN 0165-1684,
[Link]
([Link]
Abstract: The detection and identification of jamming is the prerequisite and key to the
implementation of anti-jamming measures in radar. In the target detection scenario of
airborne radar, strong clutter causes great difficulty in the detection and identification of
intra-pulse forwarding jamming. This paper proposes a jamming detection and recognition
method based on non-uniform pulse interval coupled with encoder–decoder network.
Specifically, the emission mechanism with non-uniform pulse interval is utilized to disrupt
the echo order of clutter and target, which ensure that only can the jamming gain full
coherent accumulation gain. Subsequently, the jamming signal is recovered using pulse
selection and inverse Fourier transform. Eventually, the combination of multiple loss
functions based-encoder–decoder network is utilized to learn both useful information from
the labels and valid semantic information from the time-frequency feature of the recovered
jamming signal. This can improve the accuracy of jamming recognition. The experimental
results shows that the proposed algorithm achieves more than 90% jamming detection
accuracy and over 94% jamming identification accuracy at JCNR>-15 dB even under the
limitation of insufficient training data.
Keywords: Intra-pulse forwarding jamming; Clutter; Non-uniform pulse interval; Encoder–
decoder network; Combination of multiple loss functions
Chaojie Fan, Shuxiang Lin, Baoquan Cheng, Diya Xu, Kui Wang, Yong Peng, Sam Kwong,
EEG-TransMTL: A transformer-based multi-task learning network for thermal comfort
evaluation of railway passenger from EEG,
Information Sciences,
Volume 657,
2024,
119908,
ISSN 0020-0255,
[Link]
([Link]
Abstract: The evaluation of thermal comfort for railway passengers holds considerable
importance, not only in reducing energy consumption but also in enhancing the passengers'
experience. This paper presents a Transformer-based multi-task learning network
(TransMTL) designed for railway passenger thermal comfort evaluation using EEG. We
utilized manual features to extract temporal and frequency information, while a Transformer
encoder distilled spatial information. The multi-task learning structure enhances model
robustness by leveraging thermal comfort task correlations. We conducted experiments
during winter and summer with high-speed railway passengers, establishing a
comprehensive EEG dataset. The results demonstrated that our proposed EEG-TransMTL
model outperformed classical machine learning and deep learning models in all four thermal
comfort evaluation tasks, achieving accuracy rates of 65.00%, 66.70%, 80.38%, and 71.01%,
respectively. We enhanced model interpretability by visualizing attention weights from the
Transformer encoder, identifying key EEG channels. A simplified model utilizing only eight
crucial channels also delivered notable performance. This research provides a practical and
neuro-mechanism interpretable solution for thermal comfort evaluation.
Keywords: Electroencephalogram; Railway passenger; Thermal comfort evaluation; Deep
learning; Interpretable neural network
Pengfei Zheng, Anxue Zhang, Zhensheng Shi, Sen Wang, Yi'an Ma, Zhaodan Liu,
TLAD-YOLO: Lightweight network for intelligent detection of railway tunnel lining anomalies
using ground penetrating radar,
Journal of Applied Geophysics,
Volume 241,
2025,
105869,
ISSN 0926-9851,
[Link]
([Link]
Abstract: Ground Penetrating Radar (GPR) B-scan images and the you only look once (YOLO)
series are widely used for tunnel lining intelligent inspections to ensure quality. However, in
practical applications, lightweight YOLO detection networks fail to meet the requirements of
accuracy and robustness. In view of this, a tunnel lining anomalies detection YOLO (TLAD-
YOLO) is proposed for the intelligent detection of railway tunnel lining anomalies based on
GPR B-scan images. TLAD-YOLO introduces lightweight spatial and channel synergistic multi-
shape attention (SCSMSA) to enhance the detection accuracy of complex scenes and multi-
size abnormal objects, while ghost convolution is used to reduce parameters and
computation. The experiments are conducted on a dataset consisting of 47 railway tunnels.
Furthermore, we propose a multi-scale data augmentation to further expand the dataset,
which improves the detection accuracy. The experimental results demonstrate that TLAD-
YOLO is an accurate and lightweight detection network, outperforming SOTA detection
networks in non-destructive testing of railway tunnels. On the tunnel engineering
verification platform and newly built railway tunnels, TLAD-YOLO demonstrates remarkable
robustness.
Keywords: Tunnel lining; Anomaly detection; Ground penetrating radar (GPR); You only look
once (YOLO); B-scan images
Chenglong Liu, Yuchuan Du, Guanghua Yue, Yishun Li, Difei Wu, Feng Li,
Advances in automatic identification of road subsurface distress using ground penetrating
radar: State of the art and future trends,
Automation in Construction,
Volume 158,
2024,
105185,
ISSN 0926-5805,
[Link]
([Link]
Abstract: Affected by soil erosion and material deterioration, road subsurface is prone to
distress such as cavities, water-rich, and cracks. Ground penetrating radar (GPR), as a real-
time geophysical survey method that uses electromagnetic radiation to image the
subsurface, offers promising non-destructive solutions to road subsurface health monitoring.
However, the interpretation of GPR signals is non-intuitive and obscure in terms of distress
identification, whose performance is also limited by the heterogeneous road condition. In
conjunction with knowledge diagram analysis, a state-of-the-art review is applied to
summarize the advances in the automatic identification of road subsurface distress (RSD).
The algorithms based on the single-channel waveform (A-scan), two-dimensional profile (B-
scan), and three-dimensional data (C-scan) are elaborated from the perspectives of rule-
based recognition algorithm, machine learning algorithm, and deep learning algorithm. In
comparison to analytical methods, the emerging deep learning models have a powerful
ability to extract complex features from multi-dimensional GPR radargrams, enhancing the
efficiency and accuracy of road subsurface distress detection. Recommendations for model
selection are compiled from existing literature together with empirical evidence. The most
significant variables that influence the model selections are thought to be the type of
identified RSD, training sample quality and quantity, prior knowledge, and computational
cost. Some challenges, such as insufficient training samples and diverse road structures, are
presented. Future trends are concluded to draw the implications for GPR research.
Keywords: Road subsurface distress detection; GPR; Automatic identification; Machine
learning; Deep learning
Haiyan Yao, Yuefei Xu, Qiang Guo, Shizhe Chen, Bin Lu, Yuanjun Huang,
Study on transformer fault diagnosisbased on improved deep residual shrinkage network
and optimized residual variational autoencoder,
Energy Reports,
Volume 13,
2025,
Pages 1608-1619,
ISSN 2352-4847,
[Link]
([Link]
Abstract: The transformer as the core equipment in the power system, its fault diagnosis has
a vital role in ensuring the safe and stable operation of the power grid. However, traditional
transformer fault diagnosis methods often rely on manual experience or simple models,
which are difficult to meet the demand for efficient and accurate diagnosis when faced with
complex and evolving fault patterns. In this study, a new method for transformer fault
diagnosis based on improved deep residual shrinkage network (DRSN) and optimized
residual variational autoencoders (ORVAE) is proposed. Firstly, this study improves the DRSN
to enhance its feature extraction capability. By designing a specific shrinkage mechanism,
the improved DRSN can reduce the information loss in the face of complex data, greatly
improve the extraction ability of the key features of the transformer operating state, and
thus improve the accuracy of fault recognition. Secondly, in view of the difficulty and high
cost of transformer fault sample data collection, this study introduces a residual connection
structure based on the traditional variational autoencoder (VAE), and constructs the ORVAE
method to effectively address the challenge of insufficient data. The results show that the
fault recognition rate of the proposed method on the real transformer fault dataset reaches
97.14 %, which is better than the traditional method, showing excellent diagnostic
performance and strong practical application potential. Compared with the existing
technologies, this method not only improves the accuracy of transformer fault diagnosis, but
also provides new ideas and technical support for the intelligent development of power
system. This study offers an innovative solution for the field of fault diagnosis of power
equipment, and providing a strong technical guarantee for fault prediction and maintenance
in future smart grids.
Keywords: Transformer; Fault diagnosis; Improved DRSN; Shrinkage mechanism; Feature
extraction; ORVAE; Recognition rate
Zixuan Wang, Gang Liu, Hanlin Xu, Yao Qian, Rui Chang, Durga Prasad Bavirisetti,
Transformer architecture with illumination aware mechanisms for low-light image
enhancement via Retinex decomposition,
Engineering Applications of Artificial Intelligence,
Volume 162, Part B,
2025,
112414,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Enhancing low-light images is a complex task that involves not only restoring
brightness but also preserving color fidelity and reducing noise interference. In this paper,
we propose a novel Retinex-based Transformer Model with Illumination Aware Mechanisms
(TIMRetinex-Net), which achieves physically interpretable modeling through a
decomposition network guided by Retinex theory. To adapt to light variations in different
regions, we randomly apply gamma transformations to several subregions of the
illumination component and use a Color Estimation Module to capture the color global
distribution of the natural scene in the reflection component. By modeling the color global
distribution and repairing the degraded regions collaboratively, we alleviate the issue of
being highly sensitive to data usage during training and improve the model’s ability to
handle unknown scenes. The Illumination and Reflection Adjustment Transformer Network
(IRAT-Net) produces enhanced images, achieving a balanced enhancement of detail and
color. In addition, IRAT-Net incorporates an attention mechanism into the feature extraction
layer and introduces the Illumination-Guided Information Aggregation Module to adaptively
estimate lighting conditions. In the field of image processing, our method based on artificial
intelligence was evaluated on five datasets and compared with twelve state-of-the-art
methods. The results demonstrated strong alignment with the ground truth, with our
method achieving superior performance in both subjective and objective assessments.
Keywords: Low-light image enhancement; Retinex decomposition; Transformer; Image
restoration; Deep learning
Shuai Wang, Dehao Zhang, Ammar Belatreche, Yichen Xiao, Hongyu Qing, Wenjie Wei, Malu
Zhang, Yang Yang,
Ternary spike-based neuromorphic signal processing system,
Neural Networks,
Volume 187,
2025,
107333,
ISSN 0893-6080,
[Link]
([Link]
Abstract: Deep Neural Networks (DNNs) have been successfully implemented across various
signal processing fields, resulting in significant enhancements in performance. However,
DNNs generally require substantial computational resources, leading to significant economic
costs and posing challenges for their deployment on resource-constrained edge devices. In
this study, we take advantage of spiking neural networks (SNNs) and quantization
technologies to develop an energy-efficient and lightweight neuromorphic signal processing
system. Our system is characterized by two principal innovations: a threshold-adaptive
encoding (TAE) method and a quantized ternary SNN (QT-SNN). The TAE method can
efficiently encode time-varying analog signals into sparse ternary spike trains, thereby
reducing energy and memory demands for signal processing. QT-SNN, compatible with
ternary spike trains from the TAE method, quantifies both membrane potentials and synaptic
weights to reduce memory requirements while maintaining performance. Extensive
experiments are conducted on two typical signal-processing tasks: speech and
electroencephalogram recognition. The results demonstrate that our neuromorphic signal
processing system achieves state-of-the-art (SOTA) performance with a 94% reduced
memory requirement. Furthermore, through theoretical energy consumption analysis, our
system shows 7.5× energy saving compared to other SNN works. The efficiency and efficacy
of the proposed system highlight its potential as a promising avenue for energy-efficient
signal processing.
Keywords: Quantization spiking neural networks; Neural encoding for signals; Neuritic signal
processing; Ternary spiking neural networks; Keyword spotting and EEG
Chunyu Zhu, Tinghao Zhang, Qiong Wu, Yachao Li, Qin Zhong,
An Implicit Transformer-based Fusion Method for Hyperspectral and Multispectral Remote
Sensing Image,
International Journal of Applied Earth Observation and Geoinformation,
Volume 131,
2024,
103955,
ISSN 1569-8432,
[Link]
([Link]
Abstract: There is an effective way to enhance the spatial resolution of hyperspectral remote
sensing images by fusing them with multispectral remote sensing images. However, most of
the existing deep fusion techniques adopt discretized explicit models to approximate the
complex continuous nonlinear mapping in the fusion process, leading to limitations in
enhancing the fidelity of spatial details. Additionally, existing algorithms commonly utilize
discrete methods such as bilinear or bicubic interpolation during the hyperspectral
upsampling process, leading to the loss of crucial spatial-spectral features. To this end, this
study proposes a novel Implicit Transformer Fusion Generative Adversarial Network (ITF-
GAN), which incorporates the continuity perception mechanism of implicit neural
representation with the powerful self-attention mechanism of the Transformer architecture,
which uses point-to-point implicit functions aiming to efficiently process information in both
spatial and spectral dimensions. Besides, a guided implicit neural sampling module is
introduced in the hyperspectral image up-sampling process to enhance the coordinated
expression of features in the spatial and spectral domains, which improves the spatial
resolution and spectral fidelity of the fused image during the upsampling process. A series of
fusion experiments including 4x, 8x, and 16x scale factors have shown that ITF-GAN has
significant advantages over current popular fusion algorithms in both objective evaluation
indicators and subjective visual evaluation.
Keywords: Implicit Neural Repersentation; Image fusion; Implicit Transformer; ITF-GAN
Sajid Ullah, Xi Chen, Han Han, Junhao Wu, Jinghan Dong, Ruiqing Liu, Weijie Ding, Min Liu,
Qingli Li, Honggang Qi, Yonggui Huang, Philip Lh Yu,
A novel hybrid ensemble approach for wind speed forecasting with dual-stage
decomposition strategy using optimized GRU and transformer models,
Energy,
Volume 329,
2025,
136739,
ISSN 0360-5442,
[Link]
([Link]
Abstract: Wind energy has attracted global interest owing to its sustainable and
environmentally friendly characteristics. Nevertheless, precisely forecasting wind speed can
be challenging due to its volatile and unpredictable nature. This paper presents a new hybrid
forecasting approach based on dual stage decomposition mechanism, namely TMQGDT for
wind speed prediction. At first, a decomposition technique called time-varying filtered based
empirical mode decomposition (TVFEMD) is utilized to decompose the original wind speed
data into several intrinsic mode functions (IMFs). Afterwards, multi-scale permutation
entropy (MPE) is used to assess the complexity of each IMF. Based on the entropy values,
the IMFs are further classified into high-frequency and low-frequency IMFs. To address the
significant volatility observed in the high-frequency IMFs, discrete wavelet transform (DWT)
method is employed to perform secondary decomposition. The low-frequency IMFs are
forecasted using gated recurrent unit (GRU) model optimized with quantum particle swarm
optimization (QPSO) algorithm, while the high-frequency IMFs are forecasted with the
Transformer model. The proposed model is trained and validated using four wind speed time
series datasets collected from Germany and China. Five individual models and six hybrid
models are compared against the proposed model to validate the forecasting performance
of the proposed TMQGDT model. The prediction outcomes reveals that the R2 of the model
is 0.973, 0.968, 0.956, and 0.996 on the four dataset test sets, which has improved by
3.39 %, 3.93 %, 5.53 %, and 0.50 %, respectively, compared to the TVFEMD-MPE-QPSO-GRU-
DWT-Autoformer model. The excellent accuracy performance of the TMQGDT model
indicates that developing a hybrid model based on deep learning techniques using
secondary decomposition mechanism and optimization algorithm can enhance the precision
of wind speed prediction.
Keywords: Wind speed prediction; Time-varying filtered based empirical mode
decomposition; Discrete wavelet transform; Quantum particle swarm optimization; Multi-
scale permutation entropy
Yu-Jin Jeon, Min Jeong Hong, Chan Seop Ko, So Jin Park, Hyein Lee, Won-Gyeong Lee, Dae-
Hyun Jung,
A hybrid CNN-Transformer model for identification of wheat varieties and growth stages
using high-throughput phenotyping,
Computers and Electronics in Agriculture,
Volume 230,
2025,
109882,
ISSN 0168-1699,
[Link]
([Link]
Abstract: Wheat (Triticum aestivum L.) is a major crop consumed and cultivated worldwide,
with various varieties bred to meet the growth characteristics and resistances required by
the climate conditions of each cultivation region. However, as global climate change
accelerates, rapid environmental shifts have led to crossover interaction, where previously
superior varieties undergo changes, making variety selection increasingly challenging. In
particular, there is a lack of research on methods for rapidly assessing growth rate, a key
characteristic of crossover interactions, during the variety selection process. This study
proposes a deep learning-based model and method for identifying wheat varieties and
growth stages using hyperspectral data obtained from six wheat varieties cultivated on a
high-throughput phenotyping platform. The proposed model, which combines a CNN and
Transformer, achieved 94.05% accuracy in variety detection, surpassing the performance of
related studies, and 99.24% accuracy in growth stage detection. This model enables high-
throughput monitoring of wheat variety and growth information effectively. Furthermore, if
applied to the wheat breeding process, the proposed model is expected to contribute to the
rapid selection of superior varieties suited to specific climate conditions.
Keywords: Phenotyping platform; Hyperspectral imaging; Self-attention mechanism;
Convolutional neural networks
Fei Xiong, Weili Kou, Yuhan Xun, Yinuo He, Bo Hu, Xinchen Ye, Yongke Sun,
A unified Vision Transformer (ViT) backbone with Penalty Outside Point Loss for monocular
body measurement of Binglangjiang buffaloes,
Computers and Electronics in Agriculture,
Volume 240,
2026,
111166,
ISSN 0168-1699,
[Link]
([Link]
Abstract: Binglangjiang buffalo (Bubalus bubalis), a China’s national-level protected livestock
genetic resource, lacks comprehensive body measurement data due to its nervous
temperament, which impedes conservation and research. However, existing non-contact
measurement schemes often rely on expensive multi-camera setups (depth or point cloud),
leading to high equipment costs and complex deployment. To overcome these limitations,
we propose a cost-effective monocular camera-based body measurement method. It
employs a unified Vision Transformer-based backbone network for feature extraction,
seamlessly integrating 2D image keypoints detection along with depth estimation via a
dedicated head, and introduces a novel Penalty Outside Point Loss. This loss function
enhances keypoints localization accuracy by penalizing predictions outside the body region,
outperforming conventional loss functions in boundary-sensitive scenarios. Experimental
results show that the depth estimation achieves an absolute relative error of 0.159, an
average precision of 94.83% for keypoints detection and an interior point ratio of 95.14%.
The mean absolute percentage errors for body height, hip height, oblique body length, chest
circumference, and abdominal circumference are 7.37%, 6.79%, 13.67%, 8.39%, and 7.35%,
respectively. By applying this method, we have successfully completed body measurements
for 424 Binglangjiang buffaloes, effectively filling the long-standing gap in comprehensive
body measurement data for this breed. This study establishes a reliable, low-cost framework
for buffalo body measurement data, offering crucial technical support for efficient
conservation and advancing precision livestock management practices.
Keywords: Binglangjiang buffalo; Body measurement; Vision Transformer; Monocular depth
estimation; Keypoints detection
Xinyue Xin, Ming Li, Yan Wu, Peng Zhang, Dazhi Xu,
DCDLNet: A label-noise tolerant classification algorithm for polsar images based on dual-
band consistency and difference,
Knowledge-Based Systems,
Volume 334,
2026,
115120,
ISSN 0950-7051,
[Link]
([Link]
Abstract: With the advancement of technology, PolSAR systems can acquire multiple signals
by transmitting and receiving electromagnetic waves in different frequency bands, thereby
enabling the collection of richer ground observation information. However, due to the lack
of consideration for the concepts of dual-band consistency and dual-band difference,
existing fusion methods still encounter problems of incomplete semantic information and
low computational efficiency. Moreover, in practice, the process of sample labeling often
involves manual intervention, which inevitably introduces labeling errors. To tackle these
problems, we propose a novel label-noise tolerant classification framework called DCDLNet:
dual-band consistency and difference learning network. Specifically, to extract the rich
information contained in dual-band PolSAR data, the DCDLNet comprises two principal
parts. The first part is an inter-band difference acquisition module (IDAM), which learns
dual-band complementary information based on the concept of dual-band difference. The
second part is a spatial-domain and frequency-domain feature extraction (SFFE) module. It
acquires more discriminative information by capturing local spatial information in the
spatial-domain and global spatial information in the frequency-domain. Furthermore, by
integrating the concept of dual-band consistency and the fitting capabilities of neural
networks, DCDLNet adopts a cross-band and bidirectional supervised (CBS) strategy to
mitigate the impact of label noise during the training process. Experiments on measured
PolSAR datasets demonstrate that our method outperforms several existing approaches in
terms of dual-band fusion and noisy label processing.
Keywords: Polarimetric synthetic aperture radar; Image classification; Dual-band fusion;
Label noise
Haoyu Wang, Chuanjiang Li, Peng Ding, Shaobo Li, Tandong Li, Chenyu Liu, Xiangjie Zhang,
Zejian Hong,
A novel transformer-based few-shot learning method for intelligent fault diagnosis with
noisy labels under varying working conditions,
Reliability Engineering & System Safety,
Volume 251,
2024,
110400,
ISSN 0951-8320,
[Link]
([Link]
Abstract: Recent years have witnessed the success of Few-shot Learning (FSL) methods in
equipment reliability enhancement and fault diagnosis, by virtue of learning from limited
data and adapting to new operating conditions. However, due to sensor bias, manual
collection, and mislabeling, label noise is inevitably introduced into the dataset, which
further reduces the quality of supervised information contained in the few-shot dataset,
posing significant challenges for accurate fault diagnosis. In this paper, the problem of Few-
shot Fault Diagnosis with Noisy Labels (FFDNL) is studied for the first time, and a novel
method named Enhanced Transformer with Asymmetric Loss Function (ETALF) is proposed.
ETALF leverages the self-attention mechanism of the transformer to dynamically measure
the similarity between fault samples in the support set to enhance the model's robustness
against label noise, then naturally aggregates the similar samples into corresponding correct
prototypes. Furthermore, an asymmetric loss function is designed, which adaptively assigns
the model with larger penalties for incorrect category predictions and smaller penalties for
correct category predictions, thereby enhancing fault diagnostic performance through
inherent asymmetry. Comprehensive experiments are conducted on two benchmark
datasets, and the compared results with representative approaches validate the
effectiveness of our proposed ETALF in performing intelligent fault diagnosis using limited
and noise-labeled data under varying working conditions, which achieves accuracies of
97.77% and 95.78% with 0.2 noisy-level labels during meta-training and meta-testing on the
CWRU and KAIST datasets, respectively.
Keywords: Few-shot learning; Noisy label; Intelligent fault diagnosis; Transformer;
Asymmetric loss function
Tao Zhou, Dechen Yao, Jianwei Yang, Chang Meng, Ankang Li, Xi Li,
DRSwin-ST: An intelligent fault diagnosis framework based on dynamic threshold noise
reduction and sparse transformer with Shifted Windows,
Reliability Engineering & System Safety,
Volume 250,
2024,
110327,
ISSN 0951-8320,
[Link]
([Link]
Abstract: In real industrial environments, acquiring vibration data from bearings is often
challenging due to noise, resulting in network models that excel when trained on datasets
with sufficient samples but struggle with accurate fault identification in real-world scenarios,
inevitably threatening the reliability of fault diagnosis. To address this problem, this paper
proposes an end-to-end fault diagnosis framework (DRSwin-ST) based on sparse transformer
with a shift window and dynamic threshold noise reduction. The Swin-Transformer serves as
the backbone, leveraging a multi-head self-attention mechanism with a shift window to
capture global information. The 1.5-Entmax replaces Softmax in the self-attention
mechanism, sparsifying irrelevant information and allowing the model to focus on essential
details. The self-attention mechanism, combined with a multi-scale structure, forms a
forward feedback network to obtain rich fault feature information. In addition, the paper
integrates a large convolutional kernel and a dynamic soft-threshold noise reduction module
to construct a convolutional network in front of the transformer structure. This configuration
extracts fault feature information and removes the noise, enhancing the fault recognition
accuracy of the model. Experimental results on three diverse datasets demonstrate that
DRSwin-ST exhibits robustness and high accuracy even in scenarios with limited samples and
high noise, validating its exceptional performance.
Keywords: Few samples; High noise; DRSwin-ST; Fault diagnosis
Sheng Kuang, Jie Shi, Kiki van der Heijden, Siamak Mehrkanoon,
BAST-Mamba: Binaural Audio Spectrogram Mamba Transformer for binaural sound
localization,
Neurocomputing,
Volume 650,
2025,
130804,
ISSN 0925-2312,
[Link]
([Link]
Abstract: Accurate sound localization in reverberant environments is essential for human
auditory perception. Recently, Convolutional Neural Networks (CNNs) have been used to
model the binaural human auditory pathway. However, CNNs face limitations in capturing
global acoustic features. To address this issue, we propose a novel end-to-end Binaural
Audio Spectrogram Mamba Transformer (BAST-Mamba) model to predict sound azimuth in
both anechoic and reverberant conditions. We explore two implementation modes: BAST-
Mamba-SP and BAST-Mamba-NSP, which correspond to shared and non-shared parameter
configurations, respectively. Our best model BAST-Mamba-SP, equipped with subtraction-
based interaural integration and a hybrid loss function, achieves a state-of-the-art angular
distance (AD) error of 0.89°and mean squared error of 0.0004, significantly outperforming
baseline models. The model demonstrates generalization across acoustic environments,
robust hemifield symmetry and high accurate real-time localization performance (<4°AD at
300 ms). Moderate noise augmentation at 30 dB SNR yields the strongest noise resilience.
Explainability analyses highlight consistent frequency focus in the 2–3 kHz and 5.5–6.5 kHz
bands, aligning with known neurophysiological cues. These results validate the potential of
neurobiologically inspired Transformer for robust, high-precision sound localization and offer
new insights into human sound localization.
Keywords: Transformer; Sound localization; Binaural integration
Ge Junkai, Sun Huaifeng, Shao Wei, Liu Dong, Yao Yuhong, Zhang Yi, Liu Rui, Liu Shangbin,
GPR-TransUNet: An improved TransUNet based on self-attention mechanism for ground
penetrating radar inversion,
Journal of Applied Geophysics,
Volume 222,
2024,
105333,
ISSN 0926-9851,
[Link]
([Link]
Abstract: Convolutional Neural Networks (CNN) are widely applied to Ground Penetrating
Radar (GPR) inversion because they have strong data-driven capabilities and are suitable for
the data structure form of GPR. For CNN, the computation increases with the distance that
the convolutional block moves from one region to another when it calculates the
relationship between two regions. For GPR data, the target reflection exists in the
surrounding traces and full time-window of the target, which leads to high degree of remote
relationship. In this paper, we propose GPR-TransUNet, a deep-learning based inversion
network which use self-attention mechanism. According to the characteristics of GPR data,
regression network and GPR-Loss mechanism were used. Both numerical and model
experiments were arranged to test the performance of the network, and the result as well as
comparative analysis demonstrate the superiority of GPR-TransUNet. Finally, we applied this
method to the field GPR data of Guangxi as an attempt.
Keywords: GPR; Inversion; Deep learning
Sheng Li, J.C. Ji, Yadong Xu, Ke Feng, Ke Zhang, Jingchun Feng, Michael Beer, Qing Ni, Yuling
Wang,
Dconformer: A denoising convolutional transformer with joint learning strategy for
intelligent diagnosis of bearing faults,
Mechanical Systems and Signal Processing,
Volume 210,
2024,
111142,
ISSN 0888-3270,
[Link]
([Link]
Abstract: Rolling bearings are the core components of rotating machinery, and their normal
operation is crucial to entire industrial applications. Most existing condition monitoring
methods have been devoted to extracting discriminative features from vibration signals that
reflect bearing health status. However, the complex working conditions of rolling bearings
often make the fault-related information easily buried in noise and other interference.
Therefore, it is challenging for existing approaches to extract sufficient critical features in
these scenarios. To address this issue, this paper proposes a novel CNN-Transformer
network, referred to as Dconformer, capable of extracting both local and global
discriminative features from noisy vibration signals. The main contributions of this research
include: (1) Developing a novel joint-learning strategy that simultaneously enhances the
performance of signal denoising and fault diagnosis, leading to robust and accurate
diagnostic results; (2) Constructing a novel CNN-transformer network with a multi-branch
cross-cascaded architecture, which inherits the strengths of CNNs and transformers and
demonstrates superior anti-interference capability. Extensive experimental results reveal
that the proposed Dconformer outperforms five state-of-the-art approaches, particularly in
strong noisy scenarios.
Keywords: Rolling bearing; Fault diagnosis; Vibration signal; Dconformer; Complex working
conditions; Noisy scenarios
Yanming Gu, Zhuhua Hu, Yaochi Zhao, Jianglin Liao, Weidong Zhang,
MFGTN: A multi-modal fast gated transformer for identifying single trawl marine fishing
vessel,
Ocean Engineering,
Volume 303,
2024,
117711,
ISSN 0029-8018,
[Link]
([Link]
Abstract: In order to achieve sustainable development of marine fishery resources, effective
supervise of trawl fishing during forbidden fishing period is of great significance. This paper
addresses the challenges of poor generalization and the lack of unstructured information in
the precise identification of single trawler fishing behavior. We propose a Transformer
network with multi-source information fusion processing (MFGTN), which accurately
classifies fishing vessels as single trawl or non-single trawl vessels. Firstly, a private fishing
dataset of single trawl behavior is constructed by integrating AIS data with radar data,
named HaiNan_SingleTrawlVessel(HN_STV). Subsequently, as fused data lacks unstructured
information, it undergoes transformation into trajectory point images and recurrence plot
images to reveal the internal structure of the fused data. As such, a visual module is
introduced to handle the trajectory point images and recurrence plot images as a branch.
Simultaneously, the fused data are input into a Double-Tower Transformer with Dual-gate
structures to extract information in different dimensions of the time series and feature space
as two separate branches. The Fast Attention module replaces the traditional Attention
module to improve network speed and reduce memory consumption. Ultimately, the output
of the three branches are fused and controlled by a Dual-gate structure that can
autonomously learn to determine the network output. Experimental results show that
compared to the current best-performing methods, the method discussed herein on the
HN_STV dataset has improved the accuracy, recall, precision, and F1-score performance
indicators by 2.34%, 2.46%, 0.97%, and 1.39%, respectively. The AUC area on the ROC curve
increased by 4%. In a public dataset including three fishing activities, the proposed method
improved accuracy, recall, precision, and F1-score by 2.95%, 2.59%, 2.25%, and 2.70%,
respectively, and the AUC area on the ROC curve increased by 3%. And in all experiments,
our network incurs the lowest time cost. Therefore, the method proposed herein
demonstrates its advanced performance.
Keywords: Deep learning; Data fusion; Automatic identification system; Ship trajectory
classification; Recurrence plot image
Yanrong Wang, Zihan Wang, Wanqing Zeng, Jingbao Wang, Zhiqiang Wang, Yubin Lan,
Identification of the geographical origin of wolfberry by synergetic application of electronic
eye and near-infrared spectroscopy combined with a Swin Transformer multi-scale fusion
model,
Microchemical Journal,
Volume 213,
2025,
113800,
ISSN 0026-265X,
[Link]
([Link]
Abstract: The nutritional effects and commercial value of wolfberry largely depend on its
geographical origin. This study proposed a novel method to identify the origin of wolfberry
by applying an electronic eye (EE) and near-infrared (NIR) spectroscopy combined with a
Swin Transformer multi-scale fusion model (STMIFNet). First, the exterior image and internal
quality information of wolfberry samples are collected by EE and NIR spectroscopy,
respectively. Subsequently, the Continuous Wavelet Transform (CWT) is implemented to
convert the NIR spectra into a two-dimensional (2D) spectrogram, thereby enhancing the
analysis and interpretation of spectral information. A multi-scale fusion model is further
proposed to perform feature extraction and pattern recognition based on the obtained EE
images and NIR spectrograms. This model utilizes the Swin Transformer to extract local and
global multi-scale features and incorporates multiple Information Interactive Fusion (IIF)
modules to facilitate the interaction of information between the EE images and NIR
spectrograms. The experimental results indicate that the proposed method yields more
comprehensive and accurate identification compared to using EE or NIR spectroscopy
individually. Compared to traditional machine learning methods and deep learning models,
the proposed STMIFNet demonstrates superior recognition accuracy and stronger
generalization ability. On the test set, the model achieves an accuracy, precision, recall, and
F1-score of 99.00%, 99.02%, 99.00%, and 0.9899, respectively. This study provides a rapid,
efficient, and environmentally friendly method for identifying the geographical origin of
wolfberries, which has great potential for applications in traceability detection of other food
types.
Keywords: Origin of wolfberry; Electronic eye; Near-infrared Spectroscopy; Swin
transformer; Information interactive fusion
Wandi Wang, Mahdi Motagh, Zhuge Xia, Simon Plank, Zhe Li, Aiym Orynbaikyzy, Chao Zhou,
Sigrid Roessner,
A framework for automated landslide dating utilizing SAR-Derived Parameters Time-Series,
An Enhanced Transformer Model, and Dynamic Thresholding,
International Journal of Applied Earth Observation and Geoinformation,
Volume 129,
2024,
103795,
ISSN 1569-8432,
[Link]
([Link]
Abstract: Determining the timing of landslide occurrence is crucial for establishing an
accurate, comprehensive and systematic landslide inventory while assessing the potential for
reducing landslide risk. Unfortunately, many existing landslide inventories lack temporal
information such as the precise time of landslide events. Optical and Synthetic Aperture
Radar (SAR) sensors are the most commonly used remote sensing technologies for landslide
detection. Unlike optical sensors, SAR sensors are not affected by cloudy conditions and
provide valuable imagery regardless of sunlight availability. Therefore, SAR-derived
parameters, i.e., SAR amplitude, interferometric coherence, and polarimetric features (alpha
and entropy), offer a higher temporal resolution for detecting landslide occurrence times
compared to optical data. Despite the advantages, there is currently no universally accepted
automatic method for determining the time of landslide events using SAR data. This is due to
the lack of anomaly labels and the high time-series volatility in detecting landslide
occurrence times. Despite advances in deep-learning methods for anomaly detection in
time-series, only a few of them can address these challenges in our case. In this paper, we
propose an unsupervised multivariate transformed-based deep-learning model to
automatically and efficiently estimate landslide occurrence times using multivariate SAR-
derived parameters time-series analysis. The designed gated relative position can increase
robustness and temporal context information, by learning global temporal trends in the
time-series. Subsequently, the time-series of the anomaly score derived from the proposed
Transformer model is analyzed using an adaptive thresholding strategy to dynamically and
automatically mark anomalies related to the landslide occurrence. Our research focuses on
collapsed landslides characterized by dramatic changes in ground surface topography, with a
particular attention for the need of a prior knowledge about landslide boundaries. We assess
the performance of the proposed methodology for several collapsed landslides including the
July 21, 2020 Shaziba and 23 July, 2019 Shuicheng landslides in China, March 19, 2019 Takht
landslide in Iran, June 15, 2018 Jalgyz-Jangak and May 25, 2018 Kugart landslides in
Kyrgyzstan, July 7, 2018 Hitardalur landslide in Iceland, and January 25, 2019 Brumadinho
landslide in Brazil. In comparison to commonly used neural networks like the LSTM
algorithm, our proposed framework leads to a more accurate estimate for the time of
landslide failure using time-series of SAR-derived parameters. Furthermore, our results
suggest the great potential of SAR data to narrow the time period detected from optical data
when used in conjunction with them.
Keywords: Landslide; SAR; Anomaly detection; Deep-learning
Qiming Cheng, Yihong Su, Yang He, Yang Wu, Fei Liu, Ye Rao, Yunsong Chao, Kaifeng Wang,
Zhen Liu, Jun Liu, Yao Chen,
Enhanced radar echo extrapolation for precipitation nowcasting quality using the
convolutional Kolmogorov–Arnold networks,
Journal of Hydrology,
Volume 663, Part A,
2025,
134134,
ISSN 0022-1694,
[Link]
([Link]
Abstract: With the ongoing climate warming, recurrent extreme rainfall events have become
a pervasive global challenge. The integration of disaster warnings with precipitation
nowcasting can effectively mitigate both human casualties and economic losses. Currently,
deep learning techniques are widely employed for radar echo extrapolation as a primary
approach to precipitation nowcasting. However, the lack of physical constraints often leads
to blurriness in predicted images as the forecast time increases, ultimately resulting in a
decline in forecast quality. In this study, we proposed an evolution network that incorporates
Kolmogorov-Arnold networks (KANs) to extract physical motion information and enhance
the network’s ability to learn such information. The results demonstrated that employing
convolutional KANs (ConvKANs) as the fundamental module significantly reduced the
number of model parameters while achieving superior performance. ConvKAN proved to be
a highly effective foundational module for the evolution network, not only substantially
reducing the number of model parameters and effectively improving specific meteorological
metrics (CSI, POD) and the clarity metric of predicted images (Tenengrad), but also achieving
the best performance in precipitation forecasting. Notably, our CUX2&evOnet (K2) model
combining convolutional modules exhibited optimal performance. Compared to the CNN-
based model named CUX2&evOnet(C32), it improved CSI, POD, and Tenengrad by 2.02%,
5.17%, and 2.52%, respectively, while requiring only 2.37% of its evolution network
parameters. These findings confirm that deep learning models driven by the Kolmogorov-
Arnold theorem (KAT) exhibit enhanced proficiency in precipitation nowcasting.
Keywords: Precipitation nowcasting; Deep learning; Kolmogorov-Arnold networks; Physical
constraints; Radar echo extrapolation
Yongni Shao, Dan Chen, Binggan Wang, Chen Zhao, Yun Tang, Yan Peng, Huiping Zhang,
Yiming Zhu, Wenchao Tang,
Identification of Pueraria lobata origin using terahertz precision spectroscopy and CNN-
transformer hybrid network algorithm,
Spectrochimica Acta Part A: Molecular and Biomolecular Spectroscopy,
Volume 348, Part 1,
2026,
127212,
ISSN 1386-1425,
[Link]
([Link]
Abstract: This study addresses the underexplored potential of terahertz (THz) spectroscopy
for geographical origin authentication of Pueraria lobata. We developed a novel non-
destructive approach integrating THz spectroscopy with a CNN-Transformer hybrid network
to classify samples from eight Chinese regions. High-Performance Liquid Chromatography
(HPLC) validated the correlation between THz spectral features and bioactive components
(puerarin, daidzein, daidzin). Comparative analysis with Raman spectroscopy and five
machine learning algorithms demonstrated THz spectroscopy's superiority. Under optimal
conditions, the hybrid model reached an accuracy of 91.67%, which is significantly higher
than that of traditional methods (60.42% to 64.58%) and the standard CNN
architecture(85.42%). Additionally, it achieved perfect classification (F1-score = 1.000) for
the Jiangxi/Shaanxi [Link] results establish THz spectroscopy coupled with deep
learning as a robust, accurate tool for origin traceability in traditional medicine quality
control.
Keywords: Terahertz spectroscopy; Raman spectroscopy; Pueraria lobata; Origin
identification; CNN-transformer hybrid network
Sumanta Das, Bhagyasree Chatterjee, Malini Roy Choudhury, Suman Dutta, Bhabani Prasad
Mondal, Amit Awasthi,
Synthetic aperture radar for a changing planet: A 25-year global synthesis in hazard
assessment, urban development, and ecological applications,
Ecological Informatics,
Volume 92,
2025,
103477,
ISSN 1574-9541,
[Link]
([Link]
Abstract: The increasing frequency and intensity of natural disasters, rapid urbanization, and
accelerating ecological degradation underscore the urgent need for robust monitoring tools.
Over the past two and a half decades, Synthetic Aperture Radar (SAR) technology has
transformed the landscape of remote sensing, offering unique capabilities for all-weather,
day-and-night imaging. SAR applications have expanded dramatically into critical domains
such as hazard assessment, urban development, and ecological management. Despite
significant progress, fragmented knowledge, uneven adoption across geographies, and
several technical, methodological, and application-specific challenges continue to constrain
its full potential. This systematic review synthesizes the global research trends, technological
evolution, application domains, and persisting bottlenecks in SAR utilization across three key
thematic areas: hazard assessment, urban development, and ecological management over
the past 25 years (2000–2024), supported by a comprehensive bibliometric analysis of
11,201 peer-reviewed publications indexed in the Scopus database. Unlike previous reviews
that often focus on narrow applications or specific sensor types, this review offers a holistic,
cross-sectoral, and longitudinal perspective, identifying emerging trends, underexplored
geographies, and growing intersections with novel computational frameworks. The study
employed a systematic review protocol based on PRISMA guidelines, combined with
bibliometric techniques using VOSviewer and Bibliometrix (R-package). The articles were
retrieved using a well-defined keyword query related to SAR, hazards, urban studies, and
ecosystems. Analyses included publication trends, co-authorship networks, keyword co-
occurrence, source impact, and thematic evolution. Cluster analysis identified four major
research themes and three temporal development phases. Results indicate that publications
increased nearly tenfold from 2000 to 2024, with a peak after 2015 due to the launch of
Sentinel-1 and the rise of open data. Major contributing countries include China, USA, Italy,
Germany, and India, with strong international collaborations. Interferometric SAR (InSAR)
and Polarimetric SAR (PolSAR) dominated hazard and urban studies, while multi-temporal-
based approaches emerged in ecological monitoring. Notably, integration with AI and cloud-
based geospatial platforms remains limited (<15 % of publications). Urban development
applications, especially subsidence and infrastructure monitoring, show the fastest growth,
while ecological applications lag, indicating a critical research gap. Furthermore, this review
underscores that while SAR has made substantial strides, methodological integration,
capacity building in the Global South, and translation into policy-oriented tools remain key
challenges. The findings advocate for multi-sensor synergy, open-data initiatives, and
interdisciplinary collaborations as pathways to expand SAR's impact. Overall, this review
serves as a reference point for researchers, practitioners, and policymakers aiming to
harness SAR for sustainable development and disaster resilience in the coming decades.
Keywords: Remote sensing; Disaster resilience; Interferometric SAR; Polarimetric SAR; Multi-
temporal analysis; Radar imaging
Dongjie Liu, Dawei Li, Hongliang Ding, Yang Cao, Kun Gao,
Beyond vision: A unified transformer with bidirectional attention for predicting driver
perceived risk from multi-modal data,
Transportation Research Part C: Emerging Technologies,
Volume 179,
2025,
105270,
ISSN 0968-090X,
[Link]
([Link]
Abstract: Modeling driver perceived risk (or subjective risk) plays a critical role in improving
driving safety, as different drivers often perceive varying levels of risk under identical
conditions, prompting adjustments in their driving behavior. Driving is a complex activity
involving multiple cognitive and perceptual processes, such as visual information, driver
feedback, vehicle dynamics, and traffic and environmental conditions. However, existing
models for subjective risk perception have yet to fully address the need for integrating multi-
modal data. To address this gap, we present a Transformer-based model aimed at processing
multimodal inputs in a unified manner to enhance the prediction of subjective risk
perception. Unlike existing methodologies that extract features specific to each modality, it
employs embedding layers to transform images, unstructured, and structured fields into
visual and text tokens. Subsequently, bi-directional multimodal attention blocks with inter-
modal and intra-modal attention mechanisms capture comprehensive representations of
traffic scene images, unstructured traffic scene descriptions, structured traffic data,
environmental statistics, and demographics. Experimental results show that the proposed
unified model achieves superior predictive performance over existing benchmarks while
maintaining reasonable interpretability. Furthermore, the model is generalizable, making it
applicable to various multi-modal prediction tasks across different transportation contexts.
Keywords: Traffic safety; Driving behavior modeling; Subjective risk perception; Multi-modal
data fusion
Lu Wang, Bailiang Sun, Chunhui Zhao, Suleman Mazhar, Tomoaki Ohtsuki, P. Takis
Mathiopoulos, Fumiyuki Adachi,
SAR image change detection based on saliency region guidance and SIFT keypoint extraction,
Pattern Recognition,
Volume 172, Part B,
2026,
112471,
ISSN 0031-3203,
[Link]
([Link]
Abstract: Synthetic Aperture Radar (SAR) can operate under all-weather, all-day conditions,
playing a crucial role in regional change detection (CD). However, due to its unique imaging
principles, SAR images contain significant speckle noise and blurred boundary and detail
features, which reduces the detection accuracy and leads to missed detection and false
detection. To address these issues, this paper proposes a SAR image CD method based on
saliency region guidance and Scale-Invariant Feature Transform (SIFT) keypoint extraction to
reduce the interference of speckle noise. First, a saliency region guidance method is
introduced to analyze the saliency of local features in SAR images, extracting potentially
changed regions and reducing the interference of speckle noise. Second, the SIFT is
employed to extract keypoints in regions significantly different from the background in the
difference map, leveraging its robustness to speckle noise. By extracting keypoints, the
approximate location and extent of the changed regions are determined. These are, then,
fused with the saliency region information, enhancing the saliency weights of pixels around
keypoints for more extraction of change regions. Finally, a Vision Transformer (ViT) detection
network is used for SAR image CD, utilizing the combined saliency information from the
original saliency map and SIFT keypoints. This approach effectively integrates SIFT’s stable
description of local features with ViT’s modeling capability for global features, improving the
model’s accuracy and robustness.
Keywords: Saliency region guidance; Scale-invariant feature transform; Change detection;
Vision transformer; Synthetic aperture radar image
Hongyang Liang, Jiajun Wang, Jun Zhang, Xiaoling Wang, Shiwei Guan, Hao Yu,
Improved BOTSORT multi-object tracking algorithm for robotic rollers using feature-level
fusion of millimeter-wave radar and camera sensors,
Information Fusion,
Volume 123,
2025,
103294,
ISSN 1566-2535,
[Link]
([Link]
Abstract: The operational safety of robotic rollers is of paramount importance, particularly in
the challenging construction environment of dam construction sites. However, factors like
low-illumination and intense vehicle vibrations can critically impair obstacle tracking and
decision-making processes. To address this issue, this study proposes an improved BOTSORT
multi-object tracking algorithm using feature-level fusion of millimeter-wave radar and
camera sensors. Initially, by utilizing convolutional and PS-ROI align networks, radar and
camera data are merged into feature maps, which are then processed by the improved
BOTSORT algorithm using YOLOv8 instead of YOLOX for precise obstacle detection in low-
illumination conditions. Additionally, an unscented Kalman filter module is employed to
predict nonlinear motion of objects within the image during vibrations, while radar data
refines the target association process, improving tracking accuracy under severe vibration
conditions. A case study of a large-scale hydropower project demonstrates that the
proposed method achieves 61.7 % mAP and 76.5 % MOTA, outperforming other obstacle
detection and multi-object tracking algorithms. The proposed method improves the safety
and reliability of robotic rollers under low-illumination and severe vibration working
conditions.
Keywords: BOTSORT; Feature-level fusion; Multi-object tracking; Millimeter-wave radar and
camera; Low-illumination; Robotic roller; Severe vibration
Yehao Wang, Zijian Liu, Yingying Jin, Xiaoliang Wang, Lingyu Xu, Lei Wang, Jie Yu, Wenjuan
Dai, Jingxia Gao, Feng Zhang,
Interpreting spatiotemporal dynamics of Ulva prolifera blooms in the southern yellow sea
using an attention-enhanced transformer framework,
Environmental Pollution,
Volume 384,
2025,
126999,
ISSN 0269-7491,
[Link]
([Link]
Abstract: Harmful algal blooms dominated by Ulva prolifera have posed recurring ecological
and economic challenges in the southern Yellow Sea. To better understand and predict the
complex spatiotemporal dynamics of these blooms, we developed an enhanced
Transformer-based deep learning framework, incorporating multi-head self-attention
mechanisms. This model dynamically captures spatial dependencies, providing a
comprehensive understanding of bloom dynamics. Utilizing twelve key marine
environmental factors, we systematically explored all possible feature combinations to
determine the optimal predictive subset. Experimental results demonstrated superior
predictive performance of the model (MAE: 0.0213, MSE: 0.0016, R2: 0.9923) compared to
conventional deep learning models and recent spatiotemporal deep learning models.
Training dynamics revealed efficient convergence, especially with comprehensive
environmental information. Spatial attention analysis revealed that offshore regions
consistently received higher attention, indicating their critical role as informative and
generalizable environmental references. Furthermore, exhaustive feature attribution
experiments identified an optimal combination of eight environmental factors—including
temperature, salinity, current velocity, precipitation, wind direction, dissolved iron,
phosphate, and silicate—were found to significantly enhance prediction accuracy. This study
highlights the capability of attention-enhanced Transformer models for interpretable and
precise ecological forecasting, providing valuable insights for targeted mitigation and
management of U. prolifera blooms.
Keywords: Harmful algal blooms; U. prolifera; Deep learning; Spatiotemporal dynamics;
Environmental factors; Southern yellow sea
Quan Zhou, Mingwei Wen, Bin Yu, Cuijuan Lou, Mingyue Ding, Xuming Zhang,
Self-supervised transformer based non-local means despeckling of optical coherence
tomography images,
Biomedical Signal Processing and Control,
Volume 80, Part 2,
2023,
104348,
ISSN 1746-8094,
[Link]
([Link]
Abstract: Optical coherence tomography (OCT) depends on the coherence characteristics of
scattered light to reveal tissue morphology. Therefore, OCT images are inevitably corrupted
by speckle noise. The non-local means (NLM) method is a popular method for OCT image
despeckling. However, the existing NLM algorithm cannot preserve image details well or
deliver sufficient noise reduction because they calculate the similarity weights based on the
grayscale information or human-designed features of an image. This letter presents a self-
supervised transformer based NLM method for despeckling OCT images. The presented
method computes the weight of the NLM using the deep features extracted by the self-
supervised transformer and adopts the boosting strategy to realize the effective OCT image
despeckling. The experiments on two OCT image datasets demonstrate that our algorithm
performs better than other compared denoising algorithms in terms of the quantitative
metrics and visual evaluation.
Keywords: Optical coherence tomography; Non-local means; Self-supervised transformer;
Boosting strategy; OCT image despeckling
Pahati Tuxunjiang, Chencui Huang, Zhen Zhou, Wenyi Zhao, Bingyan Han, Weixiong Tan,
Jingru Wang, HuangHanjiaerbieke Kukun, Wei Zhao, Rui Xu, Ainikaerjiang Aihemaiti, Yimuran
Subi, Jingyang Zou, Chao Xie, Yifan Chang, Yunling Wang,
Prediction of NIHSS Scores and Acute Ischemic Stroke Severity Using a Cross-attention Vision
Transformer Model with Multimodal MRI,
Academic Radiology,
Volume 32, Issue 9,
2025,
Pages 5453-5467,
ISSN 1076-6332,
[Link]
([Link]
Abstract: Rationale and Objectives
This study aimed to develop and evaluate models for classifying the severity of neurological
impairment in acute ischemic stroke (AIS) patients using multimodal MRI data.
Methods
A retrospective cohort of 1227 AIS patients was collected and categorized into mild
(NIHSS<5) and moderate-to-severe (NIHSS≥5) stroke groups based on NIHSS scores. Eight
baseline models were constructed for performance comparison, including a clinical model,
radiomics models using DWI or multiple MRI sequences, and deep learning (DL) models with
varying fusion strategies (early fusion, later fusion, full cross-fusion, and DWI-centered cross-
fusion). All DL models were based on the Vision Transformer (ViT) framework. Model
performance was evaluated using metrics such as AUC and ACC, and robustness was
assessed through subgroup analyses and visualization using Grad-CAM.
Results
Among the eight models, the DL model using DWI as the primary sequence with cross-fusion
of other MRI sequences (Model 8) achieved the best performance. In the test cohort, Model
8 demonstrated an AUC of 0.914, ACC of 0.830, and high specificity (0.818) and sensitivity
(0.853). Subgroup analysis shows that model 8 is robust in most subgroups with no
significant prediction difference (p > 0.05), and the AUC value consistently exceeds 0.900. A
significant predictive difference was observed in the BMI group (p < 0.001). The results of
external validation showed that the AUC values of the model 8 in center 2 and center 3
reached 0.910 and 0.912, respectively. Visualization using Grad-CAM emphasized the infarct
core as the most critical region contributing to predictions, with consistent feature attention
across DWI, T1WI, T2WI, and FLAIR sequences, further validating the interpretability of the
model.
Conclusion
A ViT-based DL model with cross-modal fusion strategies provides a non-invasive and
efficient tool for classifying AIS severity. Its robust performance across subgroups and
interpretability make it a promising tool for personalized management and decision-making
in clinical practice.
Keywords: Acute ischemic stroke; Vision transformer; Deep learning; Prediction; Multimodal
MRI Fusion; Cross-Attention
Hao Xu, Zhenhao Zhu, Hongbing Liu, Enrico Zio, Xiaolong Qiu, Yuchen Lu, Xianqiang Qu,
Spectral dynamic aggregation transformer and fitted swing-door algorithm for wind power
monitoring,
Energy,
Volume 341,
2025,
139396,
ISSN 0360-5442,
[Link]
([Link]
Abstract: Current research on wind energy monitoring predominantly focuses on power
prediction while often overlooking the advanced warning of sudden operational anomalies.
To this end, we propose a wind power monitoring model based on a spectral dynamic
aggregation transformer integrated with a fitted swing gate algorithm. First, the integration
of spectral and dynamic aggregation blocks within the Transformer framework yields an
accurate wind power prediction model that effectively alleviates the impact of data
fluctuations. On this basis, the MI method is utilized to quantify the nonlinear relationships
between multi-source meteorological variables and wind power output. By integrating STL
for residual analysis to extract salient features, the proposed approach not only enhances
the input quality of the prediction model but also provides a physically interpretable
foundation for the early warning module. This facilitates seamless integration of prediction
and early warning at the feature level. Furthermore, the prediction results and key features
jointly drive a two-tier early warning framework: a wind power early warning system is
constructed based on the random forest algorithm and the swing door algorithm, followed
by joint calibration with the predictive model. By leveraging multi-source data, the model is
capable of detecting anomalous power changes and ramp events, thereby ensuring efficient
anomaly identification and advanced warning. Through the case study, it is demonstrated
that the proposed model can achieve wind power prediction and advanced warning
functions, thereby providing robust support for flexible grid scheduling and efficient wind
power integration.
Keywords: Wind power; Power prediction; Spectral dynamic aggregation; Dual early warning
with fitted swing-door algorithm; Transformer
Wei Zhang, Xinyu Zhang, Junyu Dong, Xiaojiang Song, Renbo Pang,
CIDM: A comprehensive inpainting diffusion model for missing weather radar data with
knowledge guidance,
ISPRS Journal of Photogrammetry and Remote Sensing,
Volume 221,
2025,
Pages 299-309,
ISSN 0924-2716,
[Link]
([Link]
Abstract: Addressing data gaps in meteorological radar scan regions remains a significant
challenge. Existing radar data recovery methods tend to perform poorly under different
types of missing data scenarios, often due to over-smoothing. The actual scenarios
represented by radar data are complex and diverse, making it difficult to simulate missing
data. Recent developments in generative models have yielded new solutions for the problem
of missing data in complex scenarios. Here, we propose a comprehensive inpainting
diffusion model (CIDM) for weather radar data, which improves the sampling approach of
the original diffusion model. This method utilises prior knowledge from known regions to
guide the generation of missing information. The CIDM formalises domain knowledge into
generative models, treating the problem of weather radar completion as a generative task,
eliminating the need for complex data preprocessing. During the inference phase, prior
knowledge of known regions guides the process and incorporates domain knowledge
learned by the model to generate information for missing regions, thus supporting radar
data recovery in scenarios with arbitrary missing data. Experiments were conducted on
various missing data scenarios using Multi-Radar/MultiSensor System data sourced from the
National Oceanic and Atmospheric Administration, and the results were compared with
those of traditional and deep learning radar restoration methods. Compared with these
methods, the CIDM demonstrated superior recovery performance for various missing data
scenarios, particularly those with extreme amounts of missing data, in which the restoration
accuracy was improved by 5%–35%. These results indicate the significant potential of the
CIDM for quantitative applications. The proposed method showcases the capability of
generative models in creating fine-grained data for remote sensing applications.
Keywords: Weather radar data; Comprehensive inpainting; Diffusion models; Extreme
missing cases; Knowledge guidance
Pengfei Jia, Helmi Zulhaidi Mohd Shafri, Shengrui Yu, Zhi Zheng, Shiqing You, Abdul Rashid
Mohamed Shariff,
Radar-optical fusion of Sentinel-1/2 for high-resolution NDVI reconstruction and landscape-
driven carbon flux assessment in Kuala Selangor, Malaysia (2020–2024),
International Journal of Applied Earth Observation and Geoinformation,
Volume 145,
2025,
104966,
ISSN 1569-8432,
[Link]
([Link]
Abstract: Reliable quantification of carbon fluxes in humid tropical regions is constrained by
persistent cloud cover, heterogeneous land mosaics, and the limited resolution of existing
products. To address these challenges, this study developed a Cloud-Resilient Fusion
Network (CRFNet) that integrates Sentinel-1 SAR backscatter with cloud-screened Sentinel-2
NDVI using a CNN–BiLSTM–multi-head attention architecture. The framework reconstructed
10 m NDVI time series in Kuala Selangor, Malaysia (2020–2024), achieving annual R2 above
0.82 and RMSE below 0.12, thereby improving temporal continuity under heavy cloud–
rainfall interference. The reconstructed NDVI was used to drive a light-use-efficiency model
for net ecosystem productivity (NEP) estimation, supported by temperature-based
heterotrophic respiration. Results showed a 7.4 % decline in mean annual NEP across five
years, with degraded mangroves and sloping croplands emerging as hotspots of sink-to-
source transitions. Landscape analysis revealed strong structure–function coupling: stable
forests and mangroves were characterized by large cohesive patches with largest patch index
values above 40 % and edge density below 20 m ha-1, while croplands and degraded slopes
exhibited higher patch numbers, reduced patch dominance, and greater edge complexity,
which increased carbon source risk. By linking fine-scale NDVI reconstruction with process-
based carbon modeling and landscape metrics, this study provides a transferable workflow
for high-resolution carbon flux monitoring and a robust scientific basis for carbon budget
assessment, ecosystem management, and carbon-neutrality planning in tropical monsoon
regions.
Keywords: Cloud-resilient fusion network (CRFNet); NDVI reconstruction; Radar–optical
fusion; Net ecosystem productivity; Landscape metrics
Ke Wang, Bingyang Zhu, Banteng Liu, Jingyao Liang, Tan Lv, Jianfeng Wu,
Automatic sleep staging based on single-channel ballistocardiogram signals and multiple
scales temporal feature analysis,
Engineering Applications of Artificial Intelligence,
Volume 156, Part B,
2025,
111249,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Traditional sleep staging studies often rely on contact sensors for signal data
acquisition, which may compromise the integrity of sleep data. Our research presents an
automatic sleep staging method utilizing non-contact single-channel Ballistocardiogram
(BCG) signals to address this limitation. This study proposes a multiple scale(multi-scale)
time window feature extraction method based on heart rate variability (HRV) and respiratory
rate variability (RRV) to establish a more precise correlation between BCG signals and sleep
stages. Additionally, we present an advanced two-layer stacked ensemble model designed to
enhance the accuracy and robustness of sleep staging. The innovative sleep staging model is
subjected to a rigorous 5-fold cross-validation on 10 diverse recordings, encompassing
10,614 sleep segments. Experimental results indicate that the proposed feature extraction
method constructs a Top-50 feature set with an average weight of 0.2082, representing a
195.8 % improvement compared to the 0.6158 of the traditional 30s feature set.
Additionally, the proposed classification model achieves an accuracy of 89.15 %,
outperforming traditional sleep staging models by 2 percentage points. It provides valuable
references for research based on HRV, RRV, and other biological information features. This
advancement enhances sleep monitoring, especially for home and mobile healthcare,
offering a more user-friendly experience and practical medical tools. "The BCG signal sleep
data source code is available at [[Link]
Sleepstaging.]."
Keywords: Sleep staging; Ballistocardiogram signal; Heart rate variability; Respiratory rate
variability; Multiple scale time window; Two-layer stacked model
Xiuxin Xia, Yatao Cheng, Zhuo Zhang, Zhijie Hua, Qun Wang, Yan Shi, Hong Men,
Advancing research on odor-induced sweetness enhancement: A EEG local-global fusion
transformer network for sweetness quantification combined with EEG technology,
Food Chemistry,
Volume 463, Part 4,
2025,
141533,
ISSN 0308-8146,
[Link]
([Link]
Abstract: Reducing sugar intake is crucial for health, and odor sweetening enhances food
enjoyment and quality perception. Current research relies on subjective manual sensory
evaluations, which are poorly reproducible. Traditional methods also fail to capture dynamic
neural responses to odor-induced sweetness. We propose an electroencephalogram local-
global fusion transformer network (EEG-LGFNet) model to decode this impact objectively.
Electroencephalogram data were collected from 16 subjects under different odor and
sucrose stimuli. The model captures complex neural signals by integrating local and global
feature extraction mechanisms. Its performance was validated across three-time windows,
demonstrating efficacy over various temporal ranges. Analysis of the coefficient of
determination across brain regions confirmed the importance of the frontal, central, and
parietal areas of sweetness perception. The EEG-LGFNet model excelled in quantifying odor-
enhanced sweetness, significantly outperforming state-of-the-art models. This research
offers new insights into odor sweetening, with applications in food development,
personalized nutrition, and neuroscience.
Keywords: Reducing sugar; Odor sweetening; Sensory prediction; Neural responses; Artificial
intelligence
Niantang Liu, Qunshan Zhao, Richard Williams, Si-Bo Duan, Yingwei Sun, Brian Barrett,
Ensemble modelling based on transfer learning for enhancing crop mapping through
synergistic integration of InSAR coherence and multispectral satellite data,
Computers and Electronics in Agriculture,
Volume 242,
2026,
111332,
ISSN 0168-1699,
[Link]
([Link]
Abstract: Recent advancements in remote sensing have enabled the integration of multi-
temporal and multi-modal data for agricultural applications, such as crop mapping. This
study proposes an innovative framework that explores the synergistic use of multi-temporal
Sentinel-1 Interferometric Synthetic Aperture Radar (InSAR) coherence alongside Sentinel-2
and RapidEye multispectral data to enhance crop mapping in smallholder croplands in Bei’an
county, China. Various deep learning models were evaluated, including the 3-Dimensional U-
Net (3D U-Net), Transformer, Attention-based Long Short-Term Memory (AtLSTM), and a
baseline machine learning Random Forest (RF) model, focusing on their transfer learning
capabilities in complex intercropping patterns. Our new architecture, Transformer-AtLSTM-
RF, uses ensemble learning to fuse features from different classifiers with a rule-based
strategy, facilitating multi-source feature fusion for enhanced crop classification
performance. Fine-tuning with region-specific data yielded high overall accuracy (OA), mean
F1 score, and mean intersection over union (mIoU) for two test sites: site A (OA: 96.2%,
mean F1: 92.7%, mIoU: 86.9%) and site B (OA: 90.7%, mean F1: 88.6%, mIoU: 79.7%).
Additionally, we assessed feature importance by visualizing critical temporal features during
the model inference process to improve an in-depth understanding of underlying patterns in
the feature learning process. Our findings demonstrate the effectiveness of integrating time
series SAR-derived and optical data with advanced models for mapping intercropping
systems.
Keywords: Crop mapping; InSAR; Coherence; Deep learning; Transfer learning; Feature
importance
Yuanbo Li, Wenwu Zhang, Songtao Lv, Jing Yu, Dongdong Ge, Jiawei Guo, Lin Li,
YOLOv11-CAFM model in ground penetrating radar image for pavement distress detection
and optimization study,
Construction and Building Materials,
Volume 485,
2025,
141907,
ISSN 0950-0618,
[Link]
([Link]
Abstract: Ground Penetrating Radar (GPR) is an effective technology for detecting
underground structures and has been widely utilized for monitoring road damage.
Traditional B-scan-based one-dimensional images often fail to preserve continuous spatial
information, thus inadequately reflecting the nuances of damage patterns. This paper
investigates the accurate recognition of hidden internal road damage using 3D-sliced C-scan
images. While YOLO is one of the most effective and rapid neural network models for object
detection, it still suffers from low recognition accuracy and a high rate of missed detections.
To address these issues, this study proposes an improved Convolution and Attention Fusion
Module (CAFM) fusion network model for YOLOv11, which combines the CAFM with the
C2PSA global-local feature extraction mechanism to significantly enhance the recognition
performance for complex road damage. Experimental comparisons between the YOLOv11m-
CAFM and the YOLOv11 model reveal that the combined metrics for the small (n/s) and large
(l/x) models are lower than those for the medium model (m). The YOLOv11m-CAFM
demonstrates strong performance in key metrics such as precision, recall, mAP50, and
mAP50:95, achieving values of 0.840, 0.850, 0.881, and 0.584, respectively, representing
improvements of 0 %, 4.6 %, 1.8 %, and 2.0 % over the baseline model. The confidence level
in detecting standardized targets (e.g., pipelines and well covers) exceeds 0.89. Borehole
validation confirms that the model's localization error is less than 0.15 m, and the detection
frame rate reaches 71 FPS, satisfying the requirements for rapid road assessment. This study
offers a novel method for the intelligent interpretation of GPR images, considering both
detection accuracy and real-time performance, which holds significant engineering
applications in identifying hidden road damages.
Keywords: Pavement disease detection; Ground penetrating radar; Neural network; Object
detection; Deep learning algorithm
Feng Liu, Kunde Yang, Guohui Li, Zipeng Li, Guangyu Gong,
Feature extraction of underwater acoustic signal based on variational mode decomposition
and fractional-order chaotic oscillator,
Measurement,
Volume 253, Part B,
2025,
117542,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Underwater acoustic signal (UAS) feature extraction plays a crucial role in marine
target recognition. As human activities and research in marine environments continue to
increase, the task of extracting meaningful features from UAS has become more challenging.
To address this issue, this paper proposed an enhanced feature extraction method that
improves the accuracy of marine target recognition. The method combines an improved
variational mode decomposition, fractional order Duffing (FOD) oscillator, improved
multiscale amplitude-aware permutation entropy (IMAAPE), and an improved least squares
support vector machine by turbulent flow of water-based optimization (TFWO-LSSVM). By
decomposing UAS into intrinsic mode functions (IMFs), the method selects the IMF with the
smallest IMAAPE value for feature extraction. The FOD oscillator is then used to detect the
line spectrum of the selected IMF and determine the frequency range. The frequency
corresponding to the smallest IMAAPE value is identified as the line spectrum frequency.
Finally, the extracted features are input into the TFWO-LSSVM for recognition, achieving a
recognition rate of 97.4%. This method demonstrates high accuracy, and the success rate of
actual marine biological signal extraction reaches 95.1%. The proposed method offers
significant advancements in marine target recognition, with potential applications in
environmental monitoring and marine biology research.
Keywords: Underwater acoustic signal; Mode decomposition; fractional order Duffing;
Feature extraction; Entropy
Zhijun Xiao, Maarten De Vos, Christos Chatzichristos, Kejun Dong, Yunyi Jiang, Zhongyu
Wang, Yuwei Zhang, Fei Ding, Chenxi Yang, Jianqing Li, Chengyu Liu,
Noncontact capacitive coupling ECG-Derived respiratory signals using the conformer based
time–frequency domain generative adversarial network,
Expert Systems with Applications,
Volume 289,
2025,
128360,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Respiratory monitoring and analysis is a key method for detecting sleep-related
diseases. This paper presents a novel approach for respiratory monitoring that utilizes
noncontact capacitive coupling electrocardiograms-derived respiration (cEDR) method. We
propose a Time-Frequency Domain Generative Adversarial Network (TF-GAN) method for
generating respiratory signals, and successfully apply it to capacitive coupling
electrocardiograms(cECG). First, we analyze the mechanism of respiratory coupling with
cECG and verify the feasibility of the theory. Then, using the developed device, we collect
cECG data from 16 subjects during the night and simultaneously collect respiratory signals as
a reference, to validate the feasibility of our approach. Next, we convert the collected cECG
data into time–frequency domain features using Short-Time Fourier Transform (STFT) and
input these features into a Convolution-augmented transformer (Conformer) based
Generative Adversarial Network(GAN) to generate the cEDR. The network architecture
integrates self-attention mechanisms and time–frequency domain enhancement
mechanisms to effectively extract the respiratory energy components. Finally, we compare
the generated respiratory signals with the reference signals. The experimental results show
that the generated respiratory signals exhibit a high correlation with the reference signals.
Specifically, 86.3 % of the signals have a absolute waveform correlation coefficient greater
than 0.5, indicating good reproduction of real breathing waveforms. Our proposed model
demonstrates superior performance in respiratory signal extraction, achieving a low Root
Mean Square Error (RMSE) of 0.96 ± 0.12 bpm and a high agreement rate of 94.83 % ±
0.30 % within the Bland–Altman limits. Additionally, the model maintains an effective
respiratory segment ratio of 67.56 % ± 8.89 %, even under poor cECG signal conditions,
showcasing its robustness and reliability.
Keywords: Bedside Respiratory Monitoring; Capacitive coupling electrocardiogram;
Generative Adversarial Network; Time-Frequency Domain Enhancement
Jay Shen Teoh, Chee Keong Tan, Vishnu Monn Baskaran, Wai Peng Wong,
Spiking the transformer: NeuViT with single-step attention for efficient high-resolution UAV
vision,
Results in Engineering,
Volume 29,
2026,
108955,
ISSN 2590-1230,
[Link]
([Link]
Abstract: High-resolution aerial imagery from unmanned aerial vehicles (UAVs) provides the
fine-grained spatial detail required for reliable perception during field missions, yet onboard
processing of such data is constrained by high computational, latency, and energy demands.
We present Neuromorphic Vision Transformer (NeuViT), a spiking vision transformer tailored
for real-time high-resolution object detection under the strict power and latency constraints
of UAV platforms. NeuViT integrates a single-step stateless spiking paradigm that removes
temporal dependencies and membrane state tracking, and a spike-driven multi-head self-
attention mechanism that replaces dense attention with binary spike coincidences,
eliminating all multiply-accumulate operations. Evaluated on the VisDrone2019 dataset at
1500 × 2000 resolution, NeuViT achieves 88.8 % lower energy consumption and 31 % lower
latency than Swin Transformer, while retaining 79.4 % of its detection accuracy. NeuViT
remains fully operational within a 10 W power budget and sustains more than 22 frames per
second, whereas baseline vision transformers, efficient convolutional neural networks, and
spiking neural networks fail to meet UAV deployment requirements. Through per-class
precision analysis, confusion matrix characterization, and spike distribution correlation, we
identify that accuracy loss concentrates on small, low-contrast objects where spike
thresholds suppress weak activations, informing an enhancement roadmap. Energy
projections across 45 nm, 16 nm, and 7 nm technology nodes confirm that NeuViT maintains
sub-10 W operation regardless of fabrication process. These results demonstrate that
neuromorphic principles, when pragmatically adapted to silicon constraints, can enable
high-resolution vision capabilities in resource-constrained aerial platforms. The code and
models are publicly available at [Link]
Keywords: Spiking neural networks; Neuromorphic computing; Unmanned aerial vehicles;
Energy-efficient neural network; Edge AI; Real-time object detection; High-resolution image
processing
Junyu Zhou, Yuting Fu, Sihan Dong, Yuemeng Liu, Han Sun, Yanmin Li, Xunbin Wei,
Multi-model deep learning on photoacoustic flow cytometry signals for real-time melanoma
circulating tumor cells detection and biological characterization,
Expert Systems with Applications,
Volume 308,
2026,
131123,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Background
Melanoma remains one of the most aggressive forms of skin cancer, with early detection
being critical for patient outcomes. This study introduces a novel photoacoustic
fingerprinting approach integrated with advanced machine learning for non-invasive
melanoma detection and circulating tumor cell (CTC) identification.
Methods
We developed a three-tiered photoacoustic fingerprinting system combining photoacoustic
flow cytometry (PAFC) with machine learning algorithms. A uniform PAFC configuration
employed a 532 nm laser for vascular localization followed by a 1064 nm laser targeting
melanin-rich melanoma cells. The study included 50 melanoma patients and healthy
controls, analyzing spectral features across multiple wavelengths. We compared self-
supervised learning architectures (PAFCMamba vs. Transformer) and developed a hybrid
CNN-Transformer model for simultaneous CTC identification, staging, and metastatic
dissemination prediction.
Results
The photoacoustic fingerprinting system achieved exceptional diagnostic discrimination
between melanoma patients and healthy controls. Random Forest achieved area under
curve (AUC) values up to 0.97. The PAFCMamba model outperformed the Transformer
architecture (accuracy 0.75 vs. 0.62, AUC 0.785 vs. 0.730). The hybrid CNN-Transformer
architecture achieved exceptional performance with AUCs up to 0.974 and precision > 94 %
in simultaneous CTC detection, staging, and metastasis prediction. High-immunogenicity
genes including MLANA, GPR89B/A, and PIGF were identified as potential immunotherapy
targets, with photoacoustic signatures serving as non-invasive surrogate biomarkers for
underlying molecular characteristics.
Conclusions
This study establishes photoacoustic fingerprinting as a clinically viable, non-invasive
approach for melanoma detection and CTC monitoring, achieving performance comparable
to conventional methods. The integration of machine learning with photoacoustic
biomarkers provides a scalable framework with interpretable features that facilitates clinical
translation.
Keywords: Photoacoustic fingerprinting; Circulating tumor cells; Deep learning; Non-invasive
diagnosis; Biomarkers; Precision medicine
Yiming Chen, Zhen Zhang, Hao Li, Shuyuan Yang, Xiangyi Wang, Lei Wang,
GTDEKAN: Graph-aware transformer and enhanced Kolmogorov-Arnold Network for
microbe-drug association prediction,
Expert Systems with Applications,
Volume 285,
2025,
127968,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Microbes play a crucial role in human health, influencing immune responses and
disease progression. Predicting microbe-drug associations is essential for advancing drug
discovery and enhancing clinical treatments. However, current predictive methods face
challenges in processing features, preserving information, and modeling complex
relationships within intricate network structures. To address these limitations, we propose
GTDEKAN, a novel prediction model that combines an Graph-aware Transformer with a Dual
Cross-Attention (DCA) module. This integration effectively resolves issues related to long-
term dependencies and local structural information loss by extracting both global and local
features from the heterogeneous microbe-drug network. The DCA module, featuring
Channel and Spatial Cross-Attention, enhances feature processing, minimizing information
loss across complex structures. Additionally, we incorporate the Enhanced Kolmogorov-
Arnold Network (EKAN) to generate predictions of microbe-drug associations. EKAN
improves the model’s ability to capture relationships between nodes, avoiding catastrophic
forgetting and addressing the limitations of traditional deep learning models. This
integration significantly boosts the model’s accuracy. Extensive experiments show that
GTDEKAN consistently outperforms existing methods, offering superior predictive
capabilities. Case studies of four drugs across multiple databases further validate the
model’s effectiveness, revealing previously unknown microbe-drug associations and
providing new avenues for therapeutic development.
Keywords: Channel Cross-Attention (CCA); Spatial Cross-Attention (SCA); Dual Cross-
Attention (DCA); Heterogeneous-network; Kolmogorov-Arnold networks; Microbe-drug
association prediction
Seunghyun Kim, Seunghwan Shin, Sangwon Lee, Kaewon Choi, Yusung Kim,
Learning Visual Clue for UWB-based multi-person pose estimation,
Knowledge-Based Systems,
Volume 284,
2024,
111289,
ISSN 0950-7051,
[Link]
([Link]
Abstract: Compared to camera image-based methods, radio frequency (RF) based pose
estimation has great potential for use in situations where the field of view is obstructed. In
this paper, we present a novel RF-based Pose Estimation framework with Transformer (RPET)
that operates in a fully end-to-end fashion and uses an easy-to-install portable radar. RPET
eliminates the need for complex preprocessing and hand-crafted post-processing modules,
such as region-of-interest (RoI) cropping, non-maximum suppression (NMS), and keypoint
grouping. We also introduce a novel concept called Visual Clue (VC), which mimics a pose
feature represented in image-based methods and improves the learning performance of
multi-person pose estimation from RF signals. Our experimental results demonstrate the
effectiveness of VC and the generalizability of our model to different environmental
conditions, including changes in location and obstructed views.
Keywords: RF-based Pose Estimation; Multi-person pose estimation; End-to-end learning
Shahid Akbar, Ali Raza, Wajdi Alghamdi, Hashim Ali, Quan Zou, Ximei Luo,
Identifying protein succinylation sites using generative transformer and a two-dimensional
representation with a deep capsule network,
iScience,
Volume 28, Issue 12,
2025,
114137,
ISSN 2589-0042,
[Link]
([Link]
Abstract: Summary
Protein succinylation is a vital post-translational modification that regulates diverse cellular
processes. Accurate identification of succinylation sites is crucial for understanding protein
function and development of targeted drugs. In this study, we propose an intelligent
computational model, iSucc-SnCNs, which encodes protein sequences using the ProtGPT2-
based protein language model. Structural representations are derived from SMR and PSSM
matrices to extract SMR-HOG, SMR-DCT, and PSSM-DWT features. The BTGA+KNN algorithm
selects top-ranked features from the hybrid feature vector. Finally, a self-normalized capsule
neural network (Sn-CapsNet) is trained using a BTGA-based optimal feature set. The
proposed iSucc-SnCNs achieved an accuracy of 92.92% and an AUC of 0.96, outperforming
traditional models by 17%. The generalization of the iSucc-SnCNs model on two independent
datasets (Ind-I and Ind-II) demonstrated improved performance by approximately 13% and
2%, respectively. These results highlight iSucc-SnCNs as a robust and efficient framework for
large-scale succinylation site prediction and protein function analyses in drug discovery.
Keywords: Structural biology; Bioinformatics
Kang Ni, Pengcheng Wang, Zhizhong Zheng, Yanfei Zhong,
Complex-valued mix transformer for SAR ship detection,
ISPRS Journal of Photogrammetry and Remote Sensing,
Volume 231,
2026,
Pages 1-16,
ISSN 0924-2716,
[Link]
([Link]
Abstract: Transformer-based object detection algorithms, known for their powerful global
modeling capabilities and unified modeling framework, have now been widely applied to
SAR ship target detection. However, most existing approaches still rely solely on SAR
amplitude data, neglecting the importance of phase data in capturing fine structural details
of SAR targets. This reliance not only makes existing methods vulnerable to speckle noise,
but also fails to capture the high-frequency structural information in SAR phase data and
inadequate capacity to integrate fine-grained structural details and contextual semantic
information of SAR data. To address this limitation, this paper proposes a complex-valued
SAR target detection framework based on a “dual-branch feature learning–feature fusion–
feature selection” strategy, named Complex-valued Mix Transformer (CVMT), which
incorporates both SAR amplitude and phase information. The network incorporates a super
feature learning module and a detail enhancement module, which respectively mitigate
speckle noise in amplitude features and enhance the extraction of high-frequency phase
details. Moreover, the integration of a high-level feature fusion mechanism and a query
selection strategy significantly enhances the interaction between amplitude and phase
information, leading to more accurate ship target localization. Experiments conducted on
the OpenSARShip and FAIR-CSAR datasets validate the necessity of incorporating phase data
for SAR ship target detection. Furthermore, CVMT not only effectively suppresses
background clutter and preserves target structural integrity, but also achieves superior
detection robustness in both offshore and nearshore scenarios, outperforming other related
algorithms. The source code is available at [Link]
Keywords: SAR ship detection; Transformer; Complex-valued data; Feature fusion
Nadiane Nguekeu Metepong Lagpong, Joseph Mvogo Ngono, Auguste Vigny Noumsi
Woguia, Pierre Ele, Adrien Arnaud Kemche Ghomsi,
Automatic Detection of Flooded Areas in Polarimetric Radar Images From the Sentinel-1
Satellite,
International Journal of Applied Geospatial Research,
Volume 16, Issue 1,
2025,
,
ISSN 1947-9654,
[Link]
([Link]
Abstract: ABSTRACT
The free availability of Synthetic Aperture Radar (SAR) data from the sentinel satellite offers
a unique opportunity for developing countries. The research work focuses on the floods in
the town of Yagoua. The choice of this area is based on the multitude of floods causing
enormous damage. Existing methods, primarily based on machine learning and deep
learning algorithms, present major limitations such as sensitivity to radar noise, algorithmic
complexity, and dependency on training data. The methodology proposed here uses the
Kolmogorov algebraic method algorithm, which will be applied to the pre-processed images.
The Fuzzy C-Means algorithm will then be used to generate a change map consisting of two
output classes (water and not water). coupling these two methods gives good results and
analysis of pre- and post-flood images resulted in an average improvement of 12% compared
to state-of-the-art methods. This approach enhances rapid and reliable flood monitoring.
Keywords: Radar; Polarimetric Radar; Algebraic Method; Climate Change; Flooded Areas;
Fuzzy C-Means
Yanjiao Song, Linyi Li, Yun Chen, Junjie Li, Zhe Wang, Zhen Zhang, Xi Wang, Wen Zhang,
Lingkui Meng,
GCT-GF: A generative CNN-transformer for multi-modal multi-temporal gap-filling of surface
water probability,
International Journal of Applied Earth Observation and Geoinformation,
Volume 141,
2025,
104596,
ISSN 1569-8432,
[Link]
([Link]
Abstract: Spatial and temporal data gaps present a significant challenge to high-frequency
surface water mapping using satellite imagery. Utilizing observations from temporally close
periods and multi-modal sensors for gap-filling is of critical importance. However,
discontinuous pixel values inherent to conventional water maps hinder the application of
deep learning methods, which are effective and popular for relevant studies. In this study, a
novel approach, termed “gap-filling of surface water probability”, is introduced to achieve
seamless surface water mapping. A new fused dataset tailored for this purpose was
constructed, consisting of paired synthetic aperture radar (SAR) and surface water
probability data with a 10-meter spatial resolution at a 10-day interval. A Generative CNN-
Transformer (GCT) for Gap-Filling (GF) of surface water probability, GCT-GF, was then
proposed to integrate the strengths of convolutional neural networks (CNNs) and
transformers to reconstruct gapless water probability images from multi-modal and multi-
temporal data. The GCT-GF employs a coarse-to-fine structure: information from different
time points is initially aggregated using a branched gated inpainting module, followed by
refinement and alignment of the coarse output under target SAR guidance. For adversarial
learning, a branched SN-PatchGAN discriminator is introduced to adapt to the multi-
temporal input. The results show that the GCT-GF surpasses the state-of-the-art relevant
methods in quantitative metrics and visual perception. The fusion of multi-modal, multi-
temporal inputs obvious enhance the gap-filling performance across varying gap ratios.
Applied to Baiyangdian, Poyang Lake Basin and Qinghai Lake, GCT-GF demonstrates its high
reliability on large scale scenes.
Keywords: Gap-filling; Surface water mapping; Data fusion; CNN-transformer; Generative
adversarial network (GAN)
Kaiyuan Li, Wei Chen, Yanyan Zou, Zhigang Wang, Xianzhong Zhou, Jihao Shi,
Optimized PSOMV-VMD combined with ConvFormer model: A novel gas pipeline leakage
detection method based on low sensitivity acoustic signals,
Measurement,
Volume 247,
2025,
116804,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Traditional acoustic leak detection methods rely on artificial sensor systems and
are expensive to implement. The signals collected by pipeline leak detection robots based on
low-cost microphone arrays have low signal-to-noise ratios and are difficult to capture signal
details, which affects the detection results. Therefore, this paper introduces a cost-effective
method for gas pipeline leakage detection using a combination of optimized Variational
Mode Decomposition (VMD) and the ConvFormer model. The optimized VMD reduces noise
in low-sensitivity acoustic signals, enhancing feature extraction. The ConvFormer model then
processes the spectrogram to detect leaks. Leakage experiments conducted on a 100 m gas
pipeline validated the method’s effectiveness. Results demonstrate a significant
improvement in noise reduction, with reductions in Mean Squared Error (MSE) and Mean
Absolute Error (MAE) by 20 %–30 % and 18 %–24 %, respectively. The method achieved a
high detection accuracy of 99.31 %, offering a reliable solution for intelligent pipeline
inspection robots.
Keywords: Leakage detection; Microphone arrays; VMD; ConvFormer
Tianjiao Liu, Si-Bo Duan, Niantang Liu, Baoan Wei, Juntao Yang, Jiankui Chen, Li Zhang,
Estimation of crop leaf area index based on Sentinel-2 images and PROSAIL-Transformer
coupling model,
Computers and Electronics in Agriculture,
Volume 227, Part 2,
2024,
109663,
ISSN 0168-1699,
[Link]
([Link]
Abstract: Accurate estimation of leaf area index (LAI) is hindered by challenges in capturing
crop-specific spectral variability and integrating complex model-data relationships. To
address these issues, this study proposes a novel framework based on Sentinel-2 images,
coupling the PROSAIL physical model with a Transformer-based deep learning model. This
framework incorporates three key features contributing to its effectiveness. Firstly, Sentinel-
2 reflectance was generated using the PROSAIL model and refined through sample matching
to ensure optimal alignment with Sentinel-2 imagery specific to each crop type. Secondly,
the Maximum Information Coefficient (MIC) and Recursive Feature Elimination (RFE) were
employed to identify the most relevant spectral feature combinations for different crop
categories. Thirdly, a PROSAIL-Transformer coupling model was constructed based on
selected feature combinations to generate accurate Sentinel-2 LAI products. To validate the
proposed approach, field crop LAI measurements were collected at five plots within the
study area. Quantitative assessments demonstrate a coefficient of determination (R2) of
0.87, root mean square error (RMSE) of 0.48, and mean absolute error (MAE) of 0.36. The
proposed framework enables the production of time-series LAI maps at fine resolution,
facilitating dynamic crop monitoring and management in areas of high spatial heterogeneity.
Keywords: Leaf area index; Sentinel-2; PROSAIL-Transformer coupling model; Spectral
feature combinations
Sujin Jin, Homin Song, Jungoo Kang, Byoungjoon Yu, Seunghee Park,
Concrete crack reasoning: Explainable defect diagnosis incorporating generative Pretrained
transformer 4 and multimodal nondestructive testing data,
Automation in Construction,
Volume 181, Part C,
2026,
106661,
ISSN 0926-5805,
[Link]
([Link]
Abstract: Timely detection of aging concrete deterioration requires diagnostic methods
combining laboratory-level accuracy with field robustness. Existing models suffer from
opaque decision-making and limited integration of multimodal sensor data. This paper
presents Concrete Crack Reasoning (CCR), a field-validated diagnostic framework for
concrete bridge defects that overcomes these limitations via a two-module pipeline: (i)
sensor-specific convolutional neural networks distill features from ground penetrating radar,
impact echo, and ultrasonic testing into concise, human-readable sentences; (ii) a memory-
guided GPT-4 reasoning stage, operable in Baseline, Feedback, and Adaptive Refiner modes,
the last of which uses prompt retrieval for self-correction. On a 1088-sample, four-class
bridge-deck dataset, CCR raises top-1 accuracy from 49.4 % to 87.8 % and 97.9 %, reducing
the 95 % bootstrap confidence interval to ±1.8 %. Explanations in Adaptive Refiner mode
employ threshold-aware, cluster-referenced language consistent with expert practice,
enhancing auditability. With lightweight in-model retrieval and no external databases, CCR is
deployable on resource-constrained units, offering a practical path toward explainable AI-
assisted structural health monitoring.
Keywords: Structural health monitoring; Concrete crack detection; Multimodal non-
destructive testing; Explainable AI; GPT-4; Machine learning
Saidul Islam, Hanae Elmekki, Ahmed Elsebai, Jamal Bentahar, Nagat Drawel, Gaith Rjoub,
Witold Pedrycz,
A comprehensive survey on applications of transformers for deep learning tasks,
Expert Systems with Applications,
Volume 241,
2024,
122666,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Transformers are Deep Neural Networks (DNN) that utilize a self-attention
mechanism to capture contextual relationships within sequential data. Unlike traditional
neural networks and variants of Recurrent Neural Networks (RNNs), such as Long Short-Term
Memory (LSTM), Transformer models excel at managing long dependencies among input
sequence elements and facilitate parallel processing. Consequently, Transformer-based
models have garnered significant attention from researchers in the field of artificial
intelligence. This is due to their tremendous potential and impressive accomplishments,
which extend beyond Natural Language Processing (NLP) tasks to encompass various
domains, including Computer Vision (CV), audio and speech processing, healthcare, and the
Internet of Things (IoT). Although several survey papers have been published, spotlighting
the Transformer’s contributions in specific fields, architectural disparities, or performance
assessments, there remains a notable absence of a comprehensive survey paper that
encompasses its major applications across diverse domains. Therefore, this paper addresses
this gap by conducting an extensive survey of proposed Transformer models spanning from
2017 to 2022. Our survey encompasses the identification of the top five application domains
for Transformer-based models, namely: NLP, CV, multi-modality, audio and speech
processing, and signal processing. We analyze the influence of highly impactful Transformer-
based models within these domains and subsequently categorize them according to their
respective tasks, employing a novel taxonomy. Our primary objective is to illuminate the
existing potential and future prospects of Transformers for researchers who are passionate
about this area, thereby contributing to a more comprehensive understanding of this
groundbreaking technology.
Keywords: Transformer; Self-attention; Deep learning; Natural language processing (NLP);
Computer vision (CV); Multi-modality
Farhatullah, Xin Chen, Deze Zeng, Rahmat Ullah, Rab Nawaz, Jiafeng Xu, Tughrul Arslan,
A deep learning approach for non-invasive Alzheimer’s monitoring using microwave radar
data,
Neural Networks,
Volume 181,
2025,
106778,
ISSN 0893-6080,
[Link]
([Link]
Abstract: Over 50 million people globally suffer from Alzheimer’s disease (AD), emphasizing
the need for efficient, early diagnostic tools. Traditional methods like Magnetic Resonance
Imaging (MRI) and Computed Tomography (CT) scans are expensive, bulky, and slow.
Microwave-based techniques offer a cost-effective, non-invasive, and portable solution,
diverging from conventional neuroimaging practices. This article introduces a deep learning
approach for monitoring AD , using realistic numerical brain phantoms to simulate scattered
signals via the CST Studio Suite. The obtained data is preprocessed using normalization,
standardization, and outlier removal to ensure data integrity. Furthermore, we propose a
novel data augmentation technique to enrich the dataset across various AD stages. Our deep
learning approach combines Recursive Feature Elimination (RFE) with Principal Component
Analysis (PCA) and Autoencoders (AE) for optimal feature selection. Convolution Neural
Network (CNN) is combined with Gated Recurrent Unit (GRU), Bidirectional Long Short Term
Memory (Bidirectional-LSTM), and Long Short-Term Memory (LSTM) to improve
classification performance. The integration of RFE-PCA-AE significantly elevates
performance, with the CNN+GRU model achieving an 87% accuracy rate, thus outperforming
existing studies.
Keywords: Alzheimer’s disease; Classification; Deep learning; Data augmentation; Microwave
scattering; Signal processing
Mingyang Du, Ping Zhong, Xiaohao Cai, Daping Bi, Aiqi Jing,
Robust Bayesian attention belief network for radar work mode recognition,
Digital Signal Processing,
Volume 133,
2023,
103874,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Understanding and analyzing radar work modes play a key role in electronic
support measure system. Many classifiers, for example those based on convolutional neural
network (CNN) and recurrent neural network (RNN), are available for recognizing radar work
modes as well as emitter types from their waveform parameters. However, the performance
of these methods may suffer significantly when confronting different types of signal
degradation, e.g., measurement error, lost pulse and spurious pulse. To tackle this issue, we
in this paper develop a Bayesian attention belief network (BABNet) based on Bayesian neural
networks in which the probability distribution over weights can help to enhance the model
robustness for corrupted data. In particular, we adopt pre-trained CNN as the Bayesian
inference prior. This not only accelerates the convergence speed, but also avoids the training
process getting stuck in bad local minima. Meanwhile, instead of using RNNs which are
difficult to be implemented in parallel, the combination of padding operation and attention
module in the proposed BABNet enables CNN, as the backbone, to process sequential data
with variable length. Extensive experiments are conducted to demonstrate the recognition
capability and robustness of the BABNet in different environments.
Keywords: Radar work mode; Pulse descriptor word; Attention mechanism; Bayesian neural
network; Robustness; Recognition
Qidi Shu, Xiaolin Zhu, Shuai Xu, Yan Wang, Denghong Liu,
RESTORE-DiT: Reliable satellite image time series reconstruction by multimodal sequential
diffusion transformer,
Remote Sensing of Environment,
Volume 328,
2025,
114872,
ISSN 0034-4257,
[Link]
([Link]
Abstract: Repetitive optical observations from satellites are crucial for monitoring earth
surface dynamics over time. However, optical satellite image time series is severely affected
by frequent data gaps due to clouds and shadows. While synthetic aperture radar (SAR)
provides cloud-penetrating capabilities to complement missing optical data, recent
advancements in time series reconstruction have shifted focus from incorporating single SAR
image to exploiting SAR time series. However, current methods still struggle for challenging
scenarios like highly dynamic surface, persistent data gaps, and exhibit poor resilience to
inaccurate cloud masks. In this research, we approach the time series reconstruction
problem from the perspective of conditional generation. We propose a multimodal diffusion
framework termed RESTORE-DiT, which firstly promotes the sequence-level optical-SAR
fusion through a diffusion framework. Specifically, date-matched SAR time series provide
under-cloud surface dynamics to guide the denoising process of cloudy areas, and date
information is embedded to account for irregular observation intervals and periodic
patterns. Extensive experiments on three regions have shown the proposed method
achieves state-of-the-art performance. RESTORE-DiT outperforms comparison methods by
2.87 dB in PSNR and a 27.2 % reduction in RMSE on France site. SAR and date information
together increase PSNR by 2.41 dB. The reconstructed optical image time series is verified to
accurately reflect the crop growth condition and support for long-term vegetation
observations. In addition, RESTORE-DiT can be easily extended to other conditional
reconstruction or prediction tasks for arbitrary time series image data, thus facilitating
spatiotemporal analysis research. The codes will be public available at:
[Link]
Keywords: Time series reconstruction; Diffusion model; Data fusion; Optical-SAR fusion;
Cloud removal
Marco Martino Rosso, Giulia Marasco, Salvatore Aiello, Angelo Aloisio, Bernardino Chiaia,
Giuseppe Carlo Marano,
Convolutional networks and transformers for intelligent road tunnel investigations,
Computers & Structures,
Volume 275,
2023,
106918,
ISSN 0045-7949,
[Link]
([Link]
Abstract: Visual inspections do not provide a reliable and objective assessment of the
conservation state of road tunnels. Although direct tests might represent a valid survey
approach, they would often lead to prohibitive costs if performed extensively. Therefore,
indirect techniques, such as ground-penetrating radar (GPR), have become fundamental to
supporting limited direct tests. The analysis of the GPR tunnel linings profiles is mainly hand-
operated. It permits the detection of various tunnel linings defects, characterizing a tunnel’s
global health state. In the present work, the authors developed an artificial intelligence (AI)
based automatic road tunnel defects hierarchical classification framework to improve the
efficiency of this powerful indirect surveying method. Adopting the most recent tools in
image processing provided by the deep learning (DL) community, the authors proposed a
convolutional neural network (CNN) with the acknowledged ResNet-50 architecture,
initialized through the transfer learning method. For the sake of comparisons, the authors
also adopted the state-of-art convolutional EfficientNet architecture. To further improve the
proposed framework, the authors investigated how the bidimensional Fourier transform
applied as a preprocessing procedure could affect the classification performances of the
ResNet-50 model. Finally, to further enhance the classification performance, the state-of-art
neural vision transformer (ViT) architecture has been adopted with the transfer learning
approach to the currently proposed defects classification framework.
Keywords: Deep learning; Vision Transformers; Road tunnels; Fourier transform;
Convolutional Neural Network; Structural Health Monitoring; Ground Penetrating Radar
Omar Elharrouss, Yassine Himeur, Yasir Mahmood, Saed Alrabaee, Abdelmalik Ouamane,
Faycal Bensaali, Yassine Bechqito, Ammar Chouchane,
ViTs as backbones: Leveraging vision transformers for feature extraction,
Information Fusion,
Volume 118,
2025,
102951,
ISSN 1566-2535,
[Link]
([Link]
Abstract: The emergence of Vision Transformers (ViTs) has marked a significant shift in the
field of computer vision, presenting new methodologies that challenge traditional
convolutional neural networks (CNNs). This review offers a thorough exploration of ViTs,
unpacking their foundational principles, including the self-attention mechanism and multi-
head attention, while examining their diverse applications. We delve into the core mechanics
of ViTs, such as image patching, positional encoding, and the datasets that underpin their
training. By categorizing and comparing ViTs, CNNs, and hybrid models, we shed light on
their respective strengths and limitations, offering a nuanced perspective on their roles in
advancing computer vision. A critical evaluation of notable ViT architectures—including DeiT,
DeepViT, and Swin-Transformer—highlights their efficacy in feature extraction and domain-
specific tasks. The review extends its scope to illustrate the versatility of ViTs in applications
like image classification, medical imaging, object detection, and visual question answering,
supported by case studies on benchmark datasets such as ImageNet and COCO. While ViTs
demonstrate remarkable potential, they are not without challenges, including high
computational demands, extensive data requirements, and generalization difficulties. To
address these limitations, we propose future research directions aimed at improving
scalability, efficiency, and adaptability, especially in resource-constrained settings. By
providing a comprehensive overview and actionable insights, this review serves as an
essential guide for researchers and practitioners navigating the evolving field of vision-based
deep learning.
Keywords: Vision transformers; Transformers; Deep learning; Computer vision; Attention
Zhenhua Li, Jiuxi Cui, Heping Lu, Feng Zhou, Yinglong Diao, Zhenxing Li,
Prediction method for instrument transformer measurement error: Adaptive decomposition
and hybrid deep learning models,
Measurement,
Volume 253, Part D,
2025,
117592,
ISSN 0263-2241,
[Link]
([Link]
Abstract: The measurement accuracy of current transformers is crucial for power system
protection and trade fairness. The high penetration of renewable energy into the power grid
has affected the transient performance of power systems, posing significant challenges for
accurate current transformer measurement. To address this issue, this paper proposes a
prediction model for transformer measurement accuracy based on an adaptive dual-modal
decomposition strategy and a hybrid deep learning architecture. The framework integrates
an enhanced Adaptive Time-Varying Filter (A-TVF), an enhanced Adaptive Variational Mode
Decomposition (A-VMD), the Residual Error Index (REI), and the Maximum Information
Coefficient (MIC). First, A-TVF preprocesses the collected data by setting REI as the
optimization objective to adaptively adjust filter construction parameters, including the B-
spline order, bandwidth threshold, and decomposition number, and decomposes the
collected ratio error sequence to reduce the non-stationarity of the original sequence.
Subsequently, indices such as PE and Kurt are used to screen the decomposed sub-
sequences and reconstruct the complex components. Then, A-VMD is applied to further
decompose the complex components, minimizing MIC by adaptively determining the
decomposition number, penalty factor, convergence accuracy, and fidelity parameters.
Afterward, the complexity of the subcomponents obtained from the secondary
decomposition is calculated, and the entire sequence is reconstructed. Finally, a hierarchical
prediction model integrating Temporal Convolutional Networks (TCN), Bidirectional Gated
Recurrent Units (BiGRU), and a Multi-Head Attention mechanism (MHA) is employed to
predict the reconstructed components and generate the final results. Experimental results
demonstrate that the proposed adaptive dual-modal decomposition method significantly
improves prediction performance: compared with non-decomposition models, RMSE, MAE,
and SMAPE were reduced by an average of 50.12%, 46.09%, and 37.70% in global
decomposition scenarios, and by 25.92%, 23.69%, and 19.96% in rolling decomposition
scenarios, respectively. These results validate the effectiveness of the proposed method in
reducing data complexity and improving the accuracy and stability of Ratio Error predictions.
Keywords: ECT; Ratio error prediction; Adaptive dual-modal decomposition; Decomposition
and combination strategy; Hybrid deep model; Measurement accuracy
Seyedeh Leili Mirtaheri, Ali Kafi Tafti, Hamid Heidari Soureshjani, Andrea Pugliese,
GreenBERT: A lightweight green transformer for automated prediction of software
vulnerability scores,
Array,
Volume 28,
2025,
100536,
ISSN 2590-0056,
[Link]
([Link]
Abstract: Timely assessment of software vulnerabilities is critical for effective patch
prioritization, yet manual Common Vulnerability Scoring System (CVSS) scoring remains slow
and resource-intensive. While transformer-based models such as BERT have advanced
automated scoring, their substantial computational demands conflict with sustainable, green
computing objectives. This paper introduces GreenBERT, a tailored ensemble of lightweight
student Transformers, each specialized on individual CVSS metrics through a targeted multi-
head knowledge distillation framework. By jointly optimizing alignment with ground-truth
labels and softened outputs from a fine-tuned BERT teacher, GreenBERT efficiently captures
complex vulnerability patterns while significantly reducing computational overhead.
Extensive experiments on the National Vulnerability Database (NVD) and a more challenging
COMBINED dataset demonstrate that GreenBERT achieves an average F1-score
improvement exceeding 6% over the BERT baseline, while simultaneously reducing inference
time by approximately 80% and cutting energy usage and CO2 emissions by about 70%.
These results position GreenBERT as a robust, scalable, and environmentally conscious
solution for high-performance vulnerability scoring, effectively reconciling the traditionally
conflicting goals of predictive accuracy and sustainable AI.
Keywords: Green computing; Software vulnerability assessment; Deep learning; Knowledge
distillation; Transformers
John Atanbori, Christos A. Frantzidis, Mohammed Al-Khafajiy, Aliyu Aliyu, Behnaz Sohani,
Kofi Appiah, Harriet Moore, Catherine Sanders, Alastair I. Ward,
Learning with noisy labels for classifying biological echoes in polarimetric weather radar
observations using artificial neural networks,
Neurocomputing,
Volume 634,
2025,
129892,
ISSN 0925-2312,
[Link]
([Link]
Abstract: The identification of biological echoes in radar data has revolutionized research
into airborne migratory species. Deep learning applied to polarimetric weather radar
observations can reveal signature patterns of mass movement by bio-scatterers such as
birds, bats, and insects. However, due to the difficulties in labelling bio-scatterers in these
data, threshold approaches have been proposed in the literature. In this research, we used
the depolarization ratio (DR) based on differential reflectivity (zDR) and the cross-correlation
coefficient (pHV), along with citizen scientist-reported data, to label bio-scatterers for deep
learning. This method of labelling biological echoes in radar signatures is prone to noise,
which impacts the accuracy of any model that relies on it. We introduce a novel semi-
supervised co-training approach that uses a bootstrap ensemble with a confidence
threshold. Our ensemble consists of the newly proposed STNet and two modified FNet
models, which incorporate co-learning through bootstrap sampling for label correction. This
innovative method significantly improves classification accuracy across all three multivariate
numerical datasets compared to baseline models that lack co-learning with bootstrap-based
label correction.
Keywords: Artificial neural networks (ANN); Ensemble classifiers; Radar bio-scatterer
classification; Semi-supervised co-training
Cencen Liu, Dongyang Zhang, Guoming Lu, Wen Yin, Jielei Wang, Guangchun Luo,
SRMamba-T: Exploring the hybrid Mamba-Transformer network for Single Image Super-
Resolution,
Neurocomputing,
Volume 624,
2025,
129488,
ISSN 0925-2312,
[Link]
([Link]
Abstract: Single Image Super-Resolution (SISR) has made significant advancements with both
CNN-based and Transformer-based models. However, CNNs often struggle to capture long-
range dependencies effectively, and Transformers, though powerful, are hindered by
quadratic computational complexity. In recent years, state space models (SSMs), such as
Mamba, have emerged as promising alternatives due to their ability to model long-range
dependencies with linear time complexity. In light of this, we propose SRMamba-T, a hybrid
model that strategically combines Mamba and Transformer architectures to balance
computational efficiency with high performance. Specifically, we employ Transformer layers
following Mamba layers to further enhance the model’s capability to handle long-range
spatial information and expand its effective receptive fields. To design a lightweight network,
we propose a multi-directional selective scanning module to reduce parameter count and
improve computational efficiency. Additionally, a feature fusion module serves as a
bottleneck to effectively integrate hierarchical features and enhance the model’s
representational ability. Comprehensive experimental evaluations across five widely
recognized benchmarks underscore our model’s effectiveness, demonstrating its superiority
over existing state-of-the-art (SOTA) methods. For example, our model achieved a significant
PSNR improvement of 0.28 dB for ×2 lightweight super-resolution on Urban100, with a
reduction in MACs by 38.7% (from 198.1G to 121.5G) compared to the SOTA model,
MambaIR.
Keywords: Single image super-resolution; Mamba; Transformer; Lightweight model
Yifan Chen, Haibin Zhang, Xiang Shen, Xiangsheng Chen, Dong Su, Jiuqi Wu,
A novel deep learning-based identification technology of cutting pile states during super-
large diameter shield tunnelling,
Tunnelling and Underground Space Technology,
Volume 164,
2025,
106836,
ISSN 0886-7798,
[Link]
([Link]
Abstract: Accurately identifying the position and quantity of piles is critical for ensuring the
safe tunnelling process in shield cutting pile projects. The vibration signals generated during
the shield cutting pile process contain abundant information. To address the challenge of
determining pile positions and quantities, this study proposes a method for the
identification of strata based on vibration characteristics, integrating the dual advantages of
knowledge-driven and data-driven approaches. The method includes a data processing
module, a knowledge-driven module, a transformer-based model (MT), and a
comprehensive evaluation module, and it has been validated in the Guangzhou Haizhu Bay
shield tunnel project. The results show that the developed method achieves an accuracy of
99.56% in the identification of strata types, improving by 1.33%, 1.11%, and 14.16%
compared to the MLP, RF, and LSTM models, respectively. As the number of cutting piles
increases, the frequency of vibration signals gradually rises, while the amplitude shows no
significant change. Based on this finding, the top five frequencies were used as input.
Position encoding was employed to effectively learn the positional information of the
frequency, enabling the MT model to achieve an accuracy of 65.71% in identifying multiple
piles, improving by 10.51%, 19.14%, and 23.46% compared to the MLP, RF, and LSTM
models, respectively. Comprehensive evaluation analysis indicates that this method
demonstrates superior recall and weighted accuracy, highlighting its strong flexibility and
applicability in engineering contexts.
Keywords: Shield tunnel; Cutting pile; Deep learning; Ground identification; Transformer
Lifu He, Zhongchu Huang, Haidong Shao, Zhangbo Hu, Yuting Wang, Jie Mei, Xiaofei Zhang,
Fault Diagnosis of Wind Turbine Blades Based on Multi-Sensor Weighted Alignment Fusion in
Noisy Environments,
Computers, Materials and Continua,
Volume 86, Issue 3,
2026,
,
ISSN 1546-2218,
[Link]
([Link]
Abstract: Deep learning-based wind turbine blade fault diagnosis has been widely applied
due to its advantages in end-to-end feature extraction. However, several challenges remain.
First, signal noise collected during blade operation masks fault features, severely impairing
the fault diagnosis performance of deep learning models. Second, current blade fault
diagnosis often relies on single-sensor data, resulting in limited monitoring dimensions and
ability to comprehensively capture complex fault states. To address these issues, a multi-
sensor fusion-based wind turbine blade fault diagnosis method is proposed. Specifically, a
CNN-Transformer Coupled Feature Learning Architecture is constructed to enhance the
ability to learn complex features under noisy conditions, while a Weight-Aligned Data Fusion
Module is designed to comprehensively and effectively utilize multi-sensor fault information.
Experimental results of wind turbine blade fault diagnosis under different noise
interferences show that higher accuracy is achieved by the proposed method compared to
models with single-source data input, enabling comprehensive and effective fault diagnosis.
Keywords: Wind turbine blade; multi-sensor fusion; fault diagnosis; CNN-transformer
coupled architecture
Zongbin Zhang, Xiaoqiao Huang, Chengli Li, Feiyan Cheng, Yonghang Tai,
CRAformer: A cross-residual attention transformer for solar irradiation multistep forecasting,
Energy,
Volume 320,
2025,
135214,
ISSN 0360-5442,
[Link]
([Link]
Abstract: In recent years, solar energy has gained widespread adoption in smart grids due to
its safety, environmental friendliness, abundance, and other advantages, driving the
application of photovoltaic (PV) power generation technology. Accurately predicting solar
irradiance is essential for ensuring the operational stability of PV power systems, making it a
critical challenge for maintaining grid security and stability. Although Transformer models in
deep learning have achieved significant advancements in solar irradiance forecasting,
existing studies often treat cross-batch time-series data (TSD) as independent. By
overlooking the complex coupling relationships between different data batches, they fail to
fully capture the underlying patterns in TSD under varying conditions. Moreover, handling
the long-term dependencies and short-term weather-induced fluctuations inherent in TSD
remains difficult. To address these issues, this paper proposes an efficient Transformer
model (CRAformer) based on Cross-Residual Attention (CRA) for multi-step solar irradiance
forecasting. CRAformer effectively captures the deep coupling relationships within TSD
through a residual scoring mechanism, which can dynamically adjust feature weights and
balance long-term dependencies with short-term variations. Furthermore, by incorporating
a dual-output mode and dual-attention strategy, the model can deconstruct complex data
structures and guide the prediction process with greater accuracy. Additionally, the newly
designed Convolutional Weighted Fusion Module (CWFM) enhances the model's capability
to recognize diverse patterns and characteristics in TSD. By dynamically regulating the
information transfer process, the CWFM improves the model's generalization, fitting
accuracy, and robustness. To evaluate CRAformer's performance, four prediction tasks with
varying time steps (24 h, 48 h, 72 h, 96 h) were designed using irradiance datasets from
different locations: Denver, Clark, and Folsom. The experimental results demonstrate that,
compared to the second-best model, iTransformer, CRAformer reduces the RMSE by an
average of 5.6 %, 3.9 %, and 5.6 % across the four prediction steps for the datasets from
Denver, Clark, and Folsom, respectively. These results indicate that CRAformer offers
significant advantages in multi-step solar irradiance forecasting, providing a valuable
reference for future model optimization.
Keywords: Multi-step irradiance forecasting; Cross-residual; Convolutional weighting; Dual-
output mode; Photovoltaic power generation
Shaopeng He, Mingjun Wang, Nicola Forgione, Andrea Pucciarelli, W.X. Tian, S.Z. Qiu, G.H.
Su,
A multi-task Transformer-Mamba-Seq framework for real-time estimation of spatiotemporal
thermal stratification in passive residual heat exchanger,
International Communications in Heat and Mass Transfer,
Volume 169, Part D,
2025,
109868,
ISSN 0735-1933,
[Link]
([Link]
Abstract: Passive Residual Heat Removal Heat Exchanger (PRHR HX) is a critical component in
Generation-III nuclear power systems. Its spatiotemporal thermal stratification
characteristics directly influence residual heat removal capacity and serve as key inputs for
multiphysics coupling analyses. However, the complexity of input conditions challenges
traditional simulation and AI approaches, particularly under abnormal and accident
scenarios. To address this, we propose a multi-task Transformer-Mamba-Seq framework that
integrates multi-head attention with a selective scan mechanism. Compared to conventional
models, it demonstrates superior performance in both 5-fold cross-validation and
Zhi Tang, Zikang Feng, Zuqiang Su, Maolin Luo, Guo Wu, Lin Bo,
Fault diagnosis for the gas-path system of an engine test bed based on multi-modal signals,
Neurocomputing,
Volume 670,
2026,
132526,
ISSN 0925-2312,
[Link]
([Link]
Abstract: The aerospace engine test bed, as a pivotal equipment for evaluating engine
reliability, necessitates rigorous monitoring of its health state to ensure the safe operation of
the engine. The multi-point, multi-modal sensor information in the test bed's gas-path
system is characterized by complexity and variability. Traditional fault diagnosis methods
often struggle with issues such as labor-intensive threshold setting, frequent false alarms,
and missed detections. Inspired by the self-attention mechanism, a dynamic aware diagnosis
network (DADN) is proposed based on multi-modal sensor fusion. Firstly, DADN employs
shift-aware attention to focus on a small number of critical points (offset sampling) in the
sensor signal so that the discriminative local temporal features are extracted. Secondly,
dynamic pooling tricks are employed to score the channel-wise features over the duration
and generate fixed-length abstract representations. Shift-aware attention and dynamic
pooling constitute a temporal feature learning pipeline from “fine-grained localization” to
“global dynamic pooling”. Finally, multi-point, multi-modal sensor signals are fed into the
DADN to capture inter-sensor interactions and fault-related patterns, thereby achieving
accurate fault identification in the gas-path system. Experimental results demonstrate that
DADN exhibits superior diagnostic accuracy and robustness, highlighting its potential for
advanced fault diagnosis in an engine test bed's gas-path system.
Keywords: Fault diagnosis; Multi-modal; Engine test bed; Gas-path system
Zhongrui Bai, Fanglin Geng, Hao Zhang, Xianxiang Chen, Lidong Du, Peng Wang, Pang Wu,
Gang Cheng, Zhen Fang, Yirong Wu,
Non-contact blood pressure estimation using FMCW radar: A two-stream approach focused
on central arterial activity,
Biomedical Signal Processing and Control,
Volume 106,
2025,
107718,
ISSN 1746-8094,
[Link]
([Link]
Abstract: This paper proposes a radar-based two-stream blood pressure (BP) estimation
framework (R2S-BP), focusing on central arterial activity. It separately analyzes central-
arterial pulse transit time (caPTT) and pulse wave morphology using multi-location Doppler
Cardiogram (DCG) data from millimeter wave FMCW radar. Specifically, phase information at
harmonic heart rate frequencies is used to compute time delay arrays, representing caPTT-
related features. Additionally, k-Shape clustering is employed to select optimal DCGs from
the neck and chest regions that contain BP-related morphological features. These features
are processed through a two-stream neural network combining BiLSTM, ResNet, and multi-
head attention modules. Subject-independent 9-fold cross-validation results show that the
standard deviations of the errors for systolic and diastolic BP are 7.33 and 5.36 mmHg,
respectively. The intra-subject correlation coefficient for both systolic and diastolic BP
averages 0.82. Comparative and ablation studies demonstrate the superiority of the two-
stream approach and the critical importance of its components. This approach integrates
physiologically guided manual feature construction with a deep learning model, fully
leveraging the capabilities of FMCW radar data.
Keywords: Non-contact blood pressure estimation; Two-stream neural network; Doppler
Cardiogram; Central-artery pulse transit time
Yonggang Qian, Yinghua Wang, Hongwei Liu, Zelong Wang, Feipeng Yu, Chunhui Qu,
MPRANet: Multi-scale perception and reference attention network for lightweight SAR target
recognition,
Neurocomputing,
Volume 668,
2026,
132310,
ISSN 0925-2312,
[Link]
([Link]
Abstract: Deep learning methods have been widely used in Synthetic Aperture Radar
Automatic Target Recognition (SAR ATR). However, challenges remain due to limited SAR
data and computational constraints on mobile devices, which hinder model training and
deployment. In this paper, we propose a Multi-scale Perception and Reference Attention
Network (MPRANet) for lightweight SAR ATR, which is a hybrid structure combining
convolutional networks and transformers, built upon the ShuffleNetV2 network. Specifically,
MPRANet introduces two key improvements compared to the CNN-based ShuffleNetV2.
Firstly, we replace the depthwise convolutions (DWConv) in the downsampling and basic
units of ShuffleNetV2 with the Multi-scale Parameter-Shared Convolution (MPConv) module.
MPConv enables the extraction of multi-scale features of SAR targets with almost no
additional parameters, thereby enhancing the network’s feature extraction capabilities.
Secondly, we propose a lightweight Reference Attention Transformer (RAformer) to capture
global information, addressing the issue of insufficient channel feature interaction in
ShuffleNetV2. In RAformer, a Local Linear Mapping Unit (LMU) is designed to perform linear
mappings, reducing the introduction of redundant features while ensuring its lightweight
and efficient nature. RAformer contains two modules: the Reference Vector Attention (RVA)
module, which efficiently models attention relationships, and the Lightweight Feedforward
Neural Network (LW-FFN) module, which enhances the network’s ability to capture
nonlinear representations. We evaluated the performance of MPRANet using publicly
available SAR datasets, including the MSTAR dataset, OpenSARShip dataset, and SAR-
AIRcraft-1.0 dataset. The experimental results demonstrate that MPRANet consistently
achieves superior recognition performance compared to other lightweight networks of
similar complexity.
Keywords: Synthetic aperture radar (SAR); Automatic target recognition (ATR); Convolutional
neural networks (CNN); Transformer; Lightweight
Yuheng Chen, Decheng Feng, Zhongshi Pei, Xiaoxuan Mao, Lulu Fan, Meng Xu, Yang Li,
Dongsheng Wang, Junyan Yi,
Identification and information acquisition of high-value construction solid waste combined
millimeter-wave radar and convolutional neural networks,
Waste Management,
Volume 194,
2025,
Pages 390-400,
ISSN 0956-053X,
[Link]
([Link]
Abstract: The accumulation of construction solid waste (CSW) leads to the waste of land
resources and environmental pollution, becoming a significant social problem. Identifying
the amount of high-value CSW is essential for assessing the value of accumulated CSW and
formulating appropriate recycling strategies. With the development of machine learning
technology, CSW recognition techniques combining image acquisition devices and
convolutional neural networks have been widely applied. However, most technologies are
based on 2D images, making it difficult to recognize high-value CSW in accumulated CSW.
This study proposes a new method to identify high-value CSW using millimeter wave radar
based on penetration properties of electromagnetic waves. First, efficient imaging of CSW
was achieved by optimizing the imaging algorithm. Then, CSW were classified in
combination with the selected convolutional neural network (CNN) method based on the
dataset constructed in the lab. At last, high-value CSW was screened by normalizing the
imaging algorithm. The findings indicate that the balance between imaging effect and
efficiency is achieved when the step speed and height are 200 mm/s and 4 mm. The
optimized imaging approach effectively captures images of CSW. Compared with SegNet and
PSPNet, DeepLabv3+ can identify complete bricks and reinforcing bars precisely. The
accuracy can reach 85.18 %. Moreover, the millimeter-wave radar can determine the
location and size of waste and can potentially acquire three-dimensional and buried
information about waste.
Keywords: Construction solid waste; Millimeter wave radar; Electromagnetic scattering;
Convolutional neural networks
Yukai Kong, Xianxiang Yu, Jiachen Li, Kui Xiong, Guolong Cui,
Non-uniform pulse intervals based intra-pulse forwarding jamming detection and
recognition in clutter circumstance,
Signal Processing,
Volume 238,
2026,
110193,
ISSN 0165-1684,
[Link]
([Link]
Abstract: The detection and identification of jamming is the prerequisite and key to the
implementation of anti-jamming measures in radar. In the target detection scenario of
airborne radar, strong clutter causes great difficulty in the detection and identification of
intra-pulse forwarding jamming. This paper proposes a jamming detection and recognition
method based on non-uniform pulse interval coupled with encoder–decoder network.
Specifically, the emission mechanism with non-uniform pulse interval is utilized to disrupt
the echo order of clutter and target, which ensure that only can the jamming gain full
coherent accumulation gain. Subsequently, the jamming signal is recovered using pulse
selection and inverse Fourier transform. Eventually, the combination of multiple loss
functions based-encoder–decoder network is utilized to learn both useful information from
the labels and valid semantic information from the time-frequency feature of the recovered
jamming signal. This can improve the accuracy of jamming recognition. The experimental
results shows that the proposed algorithm achieves more than 90% jamming detection
accuracy and over 94% jamming identification accuracy at JCNR>-15 dB even under the
limitation of insufficient training data.
Keywords: Intra-pulse forwarding jamming; Clutter; Non-uniform pulse interval; Encoder–
decoder network; Combination of multiple loss functions
Ting Dai, Liye Mei, Yue Zhang, Biao Tian, Rui Guo, Teng Wang, Shan Du, Shiyou Xu,
UAVs and birds classification using robust coordinate attention synergy residual split-
attention network based on micro-Doppler signature measurement by using L-band staring
radar,
Measurement,
Volume 222,
2023,
113692,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Developing unmanned aerial vehicles (UAVs) and birds surveillance technologies to
produce accurate descriptions and achieve high classification accuracy is critical in the field
of radar automatic target recognition (RATR). This article proposes a grayscale spectrogram
image-based UAVs and birds classification method using a robust coordinate attention
synergy residual Split-Attention network (RCA-ResNeSt) under the holographic staring radar
system. Specifically, the ResNet structure with Split-Attention is used as an m-D feature
extractor. The CrossNorm and SelfNorm (CNSN) mechanism is then incorporated into the
network to advance generalization robustness. After that, to consider the spatial direction of
the m-D signature, a coordinated attention (CA) mechanism is introduced at the tail end of
the network to enable fine-grained mining of potential m-D features. Experiments are
carried out using a designed radar system. The results show the superiority of the proposed
method over existing approaches in classification accuracy and noise robustness.
Keywords: Micro-doppler (m-D) signature; Time–frequency representation (TFR); Robust
coordinate attention synergy residual split-attention network (RCA-resNeSt); Holographic
staring radar; Radar automatic target recognition (RATR)
Jiachen Li, Jiaxian Hao, Yukai Kong, Xianxiang Yu, Zhaoyin Xiang, Guolong Cui, Wenmin Wang,
Few-shot jamming recognition based on NMF combined with multi-dimensional fusion
network,
Signal Processing,
Volume 237,
2025,
110089,
ISSN 0165-1684,
[Link]
([Link]
Abstract: Accurately identifying specific types of active jamming is essential for optimizing
radar resources and enhancing anti-jamming efficiency, particularly in the context of few-
shot sample sizes, as discussed in this study. We first employ non-negative matrix
factorization (NMF) to pre-process the radar signal. NMF enhances the feature
representation of data while simultaneously augmenting the sample size. Subsequently, we
propose a multi-dimensional fusion network (MDFN) designed to integrate high-dimensional
features and classify jamming signals effectively. The proposed method demonstrates
superior performance compared to existing approaches across twelve categories of jamming
in few-shot scenario. Experimental results are presented to validate the reliability and
effectiveness of the proposed method.
Keywords: Active jamming recognition; Non-negative matrix factorization (NMF); Multi-
dimensional fusion network (MDFN); Efficient channel attention (ECA); Few-shot samples
Wei Quan, Wenjing Cheng, Yike Yang, Haiquan Zhao, Zhaoyu Chen, Yunfan Luo,
A signal fingerprint feature extraction method based on decomposition and fusion for radar
emitter individual identification,
Digital Signal Processing,
Volume 164,
2025,
105257,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Radar emitter individual identification is one of the key technologies of modern
electronic countermeasure reconnaissance and electronic intelligence. With the
advancement of radar technology and the increasingly complex electromagnetic
environment, existing methods for identifying emitter are gradually becoming unable to
meet the performance requirements of modern radar individual identification. Aiming at
improving the adaptability of feature extraction for non-cooperative radar emitter signals
and the robustness of individual identification in the complex modern electronic warfare
environment, a signal fingerprint feature extraction method based on decomposition and
fusion is proposed. It firstly integrates signal decomposition and scattering convolution
networks (SCN) to adaptively extract the multi-scale intra-pulse feature of the signal, while
removing the potential noise of the redundant component by energy proportion. And then a
deep feature fusion model based on multi-head self-attention and residual connection is
proposed to fuse the multi-scale features and the time domain features to further extract
signal fingerprint of radar emitter. Experimental results based on the real radar emitter
signals demonstrate that the identification method proposed in this paper can more
effectively extract signal fingerprint features and the identification accuracy reaches 96.45%,
which outperforms other existing identification methods.
Keywords: Radar emitter individual identification; Signal fingerprint feature; Signal
decomposition; Scattering convolution networks (SCN); Fusion
Xuezhong Wang,
Electronic radar signal recognition based on wavelet transform and convolution neural
network,
Alexandria Engineering Journal,
Volume 61, Issue 5,
2022,
Pages 3559-3569,
ISSN 1110-0168,
[Link]
([Link]
Abstract: With the continuous use of various new radar systems and complex radar systems,
the electromagnetic environment is extremely deteriorated. The traditional emitter
recognition methods have been difficult to meet the requirements of recognition
performance in the rapidly changing battlefield environment. To solve this problem, a radar
electronic signal recognition algorithm based on wavelet transform and deep learning is
proposed in this paper. Starting from the radar reconnaissance system, the causes of signal
preprocessing are analyzed, and the methods of signal denoising, signal normalization, signal
intra pulse modulation recognition, multipath signal detection and suppression are deeply
studied. In particular, the denoising algorithm based on threshold wavelet transform is
proposed, which significantly improves the reliability of the algorithm. Aiming at the
individual feature extraction of emitter signal, the extraction methods of emitter signal time
domain feature, frequency domain feature, fuzzy function slice feature and cyclic spectrum
feature based on wavelet transform are studied and analyzed, which provides stable and
reliable classification features for emitter signal recognition. According to the characteristics
of radar emitter signal, an optimized convolution neural network is designed, and the
feature fusion processing is carried out at the decision-making level, which greatly improves
the recognition effect and enhances the robustness of the recognition system. Experiments
on radar data show that the fusion recognition rate is higher than any single feature and has
strong robustness. In addition, compared with the traditional SVM and elm networks, the
CNN network proposed in this paper can extract detailed features more effectively and
improve the recognition rate of electronic radar.
Keywords: Radar signal recognition; Wavelet transform; Convolution neural network;
Feature fusion; Noise reduction
Teng Huang, Yongfeng Chen, Bingjian Yao, Bifen Yang, Xianmin Wang, Ya Li,
Adversarial attacks on deep-learning-based radar range profile target recognition,
Information Sciences,
Volume 531,
2020,
Pages 159-176,
ISSN 0020-0255,
[Link]
([Link]
Abstract: Target recognition based on a high-resolution range profile (HRRP) has always been
a research hotspot in the radar signal interpretation field. Deep learning has been an
important method for HRRP target recognition. However, recent research has shown that
optical image target recognition methods based on deep learning are vulnerable to
adversarial samples. Whether HRRP target recognition methods based on deep learning can
be attacked remains an open question. In this paper, four methods of generating adversarial
perturbations are proposed. Algorithm 1 generates the nontargeted fine-grained
perturbation based on the binary search method. Algorithm 2 generates the targeted fine-
grained perturbation based on the multiple-iteration method. Algorithm 3 generates the
nontargeted universal adversarial perturbation (UAP) based on aggregating some fine-
grained perturbations. Algorithm 4 generates the targeted universal perturbation based on
scaling one fine-grained perturbation. These perturbations are used to generate adversarial
samples to attack HRRP target recognition methods based on deep learning under white-box
and black-box attacks. The experiments are conducted with actual radar data and show that
the HRRP adversarial samples have certain aggressiveness. Therefore, HRRP target
recognition methods based on deep learning have potential security risks.
Keywords: Adversarial attacks; Deep neural networks; Radar images; Target recognition
Shuai Guo, Ting Chen, Penghui Wang, Jun Ding, Junkun Yan, Hongwei Liu,
Knowledge embedding fusion based on language model for enhanced radar target
recognition,
Signal Processing,
Volume 238,
2026,
110199,
ISSN 0165-1684,
[Link]
([Link]
Abstract: Traditional radar target recognition methods typically model only single echoes,
neglecting the crucial information that domain knowledge can provide for understanding
data. In this paper, we propose a knowledge embedding fusion (KEF) method for enhanced
high-resolution range profile (HRRP) recognition, which utilizes the target state descriptions
available during radar detection. KEF leverages a language model (LM) to integrate textual
knowledge with echo features for fusion recognition. It consists of three components: HRRP
feature extraction, measurement-based knowledge construction, and knowledge embedding
fusion module. First, we perform feature extraction on the HRRP to obtain echo tokens.
Next, in the knowledge construction module, the measurement statuses are standardized to
a natural language format, and the LM is utilized to extract semantic information, resulting in
text tokens. Finally, in the knowledge embedding fusion module, a cross-attention HRRP-text
fusion strategy is employed to facilitate interaction between echo tokens and textual tokens.
We also design a combination of HRRP-text matching loss and fusion classification loss to
guide model training. Experiments are conducted on a real measured dataset, and the
results indicate that KEF effectively enhances recognition performance across multiple
scenarios compared with approaches that only utilize echoes.
Keywords: High-resolution range profile (HRRP); Knowledge embedding fusion; Language
model (LM); Radar target recognition
Tingpei Huang, Rongyu Gao, Haotian Wang, Jianhang Liu, Shibao Li,
mBox: 3D object detection based on millimeter-wave radar,
Measurement,
Volume 246,
2025,
116568,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Millimeter-wave radar is utilized for 3D object detection in autonomous driving
due to its advantage of not being affected by lighting conditions. Previous millimeter-wave
radar object detection algorithms have not sufficiently used the distributional and statistical
properties of sparse point clouds. This paper introduces mBox, a 3D object detection
framework using only millimeter-wave radar. To eliminate unnecessary information, we
propose a background filtering algorithm based on subtraction(BGFS) that matches and
differentiates the point cloud frame by frame. We propose a voting-based algorithm for
generating object centers(CPGV), which expands the valid data to generate more accurate
initial anchor boxes. To utilize feature information across different scales and capture the
structure of each granularity, we propose a multi-scale feature fusion network based on the
attention mechanism(GLFF-Net). We conduct experiments using the Pointillism and Astyx
datasets. The results show that the mBox outperforms the comparative methods in terms of
3D mean average precision(mAP).
Keywords: Object detection; Point cloud; Millimeter-wave radar
Zhuangzhuang Tian, Wei Wang, Fengchuan Wu, Kai Zhou, Shengqi Liu, Huiqiang Zhang,
SAR target recognition based on CNN with 2-D dual-tree complex wavelet transform
decomposition,
Pattern Recognition,
Volume 172, Part C,
2026,
112585,
ISSN 0031-3203,
[Link]
([Link]
Abstract: Synthetic aperture radar (SAR) is an effective imaging and observation sensor that
has been widely applied in both military and civilian fields. Deep learning approaches have
gained prominence in SAR target recognition and received extensive attention. However,
these methods often struggle when data is scarce, leading to insufficient training and
challenges in effective feature extraction. To address this limitation, we propose a less data-
dependent feature extraction framework. Specifically, we introduce the dual-tree complex
wavelet transform (DTCWT) to capture multi-frequency feature details of SAR images,
integrated with convolutional neural network. This approach enables effective extraction of
high- and low-frequency information. By leveraging the characteristics of these frequency
features, low-frequency subbands are used to emphasize the global structural features in
the images, and high-frequency subbands are employed to identify the significance of
different regions in the images. In response to the aforementioned characteristics, we
introduced an attention mechanism to effectively incorporate high-frequency local
information into low-frequency global information, thereby enhancing feature
representation and recognition efficiency. Moreover, we propose an adaptive rotational
convolution, and apply it to the high-frequency feature extraction. The adaptive rotational
convolution can adapt to the directionally selective subbands with a single convolution
kernel. Experiments conducted on the MSTAR and SAR car datasets demonstrate that the
proposed method can achieve better recognition performance with fewer parameters,
especially on small-scale datasets. The ablation study also confirms the effectiveness of the
introduced DTCWT and rotational convolution.
Keywords: Synthetic aperture radar; Target recognition; Dual-tree complex wavelet
transform; Convolutional neural network
Wenxu Zhang, Xian Lei, Zhongkai Zhao, Fuli Sun,
A dual-decision-maker frequency domain cooperative jamming method against multi-
function radar based on PPO,
Digital Signal Processing,
Volume 169,
2026,
105709,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Optimizing jamming strategies is crucial for coping with complex electromagnetic
countermeasures in dynamic spectrum environments, Among them, the frequency agility
characteristic of the multi-function radar enables them to exhibit strong anti-jamming
capabilities by quickly changing the carrier frequency. Aiming at the problem that traditional
electronic countermeasure strategies show insufficient adaptability to this, a dual-decision-
maker collaborative jamming method based on the proximal policy optimization (PPO)
framework is proposed in this paper. The method achieves dynamic adaptive jamming
against radars through a cascaded collaborative mechanism involving the frequency band
decision maker and the bandwidth decision maker. The confrontation scenario between the
radar network and multiple jammers is established, and the penetration confrontation
process is abstracted and modeled as a markov decision process. The simulation results
show that compared with classic reinforcement learning algorithms such as deep Q network,
the proposed dual-decision-maker collaborative jamming method based on PPO exhibits
superior performance in action estimation and jamming success rate. Specifically, the
convergence speed of the average jamming gain is improved by more than 65 % compared
with the sub-optimal algorithm, while its final stable reward is approximately 10 % higher.
Meanwhile, the key performance indicator values fluctuate minimally under different
experimental conditions, showcasing good robustness and scalability, and achieving
intelligent optimization of frequency domain decision-making.
Keywords: Frequency agility; Deep reinforcement learning; Intelligent jamming decision;
Proximal policy optimization
Yong Liu, Chenyang Lu, Liang Li, Xiangchao Meng, Qiuping Jiang, Feng Shao,
Interactive feature fusion for camera-radar-based vehicle segmentation in bird’s-eye view,
Pattern Recognition,
Volume 172, Part D,
2026,
112698,
ISSN 0031-3203,
[Link]
([Link]
Abstract: Vehicle segmentation in Bird’s-Eye View (BEV) is a fundamental task for
autonomous driving, and integrating multi-modal sensory inputs, e.g., cameras and radars,
could enhance perception capability by leveraging their complementary strengths. However,
cross-modal feature fusion raises additional challenges due to the sparse and noisy
characteristics of radar data and the inherent misalignment between radar and camera
features. While existing fusion methods frequently leverage powerful attention mechanisms,
they often overlook the aforementioned heterogeneities and their impact on achieving
consistent, fine-grained alignment across modalities. We introduce the Interactively
Enhanced Camera-Radar Fusion (IECRF) framework, a novel approach that effectively bridges
cross-modal discrepancies in two stages through three new modules: Camera-Radar Feature
Aggregation (CRFA), Multi-Scale Radar Enhancer (MSRE), and Camera-Radar Feature Fusion
(CRFF). Specifically, the CRFA module explicitly models the complementary features of visual
appearance and radar geometry through two attention mechanisms, enabling fine-grained
alignment and interactive enhancement between the two modalities. The MSRE module
further refines radar representations through a modality-specific down- and up-sampling
design, amplifying salient targets while suppressing background noise in sparse radar
features. The aggregated features are then fused using the CRFF module at each stage for
lateral decoding. Extensive evaluations on the nuScenes dataset demonstrate that our IECRF
framework can operate with multiple backbones and configurations, achieving higher
vehicle segmentation accuracy even when using a lightweight EfficientNet backbone, which
is six times faster than the existing state-of-the-art approach equipped with an advanced ViT
backbone. The source code and trained models are available at
[Link]
Keywords: Vehicle segmentation; Bird’s-Eye View (BEV); Radar perception; Visual-radar
fusion; Autonomous driving
Chaofeng Huang, Xiaowo Xu, Fan Fan, Shunjun Wei, Xiaoling Zhang, Dongmei Liu, Min Gu,
A low-SNR-adaptive temporal network with smart mask attention for radar signal
modulation recognition,
Digital Signal Processing,
Volume 168, Part D,
2026,
105640,
ISSN 1051-2004,
[Link]
([Link]
Abstract: The automatic modulation recognition of radar signals is a key technology in
electronic warfare and communication systems. However, traditional handcrafted features
often struggle to achieve high recognition accuracy under low signal-to-noise ratio (SNR)
conditions. With the rapid development of artificial intelligence technologies, deep learning-
based approaches have emerged as a promising alternative for modulation recognition. In
this article, a low-SNR-adaptive network architecture is proposed, which integrates a
bidirectional temporal convolutional network (Bi-TCN) and dual-channel smart mask
attention (DSMA) modules. The DSMA adaptively highlights informative features and
suppresses noise through complementary attention masks, enhancing robustness in low-SNR
conditions. Experimental results demonstrate that the autocorrelation domain outperforms
both time and frequency domains, with recognition accuracy improvements of 13.33 % and
14.71 %, respectively. Compared to state-of-the-art models, the proposed network achieves
63 % accuracy at -20 dB and more than 99 % accuracy at -6 dB, significantly enhancing radar
signal modulation recognition.
Keywords: Modulation recognition; Deep learning; Radar signal analysis,
Futai Liang, Xin Chen, Song He, Zihao Song, Hao Lu,
An Aerial Target Recognition Algorithm Based on Self-Attention and LSTM,
Computers, Materials and Continua,
Volume 81, Issue 1,
2024,
Pages 1101-1121,
ISSN 1546-2218,
[Link]
([Link]
Abstract: In the application of aerial target recognition, on the one hand, the recognition
error produced by the single measurement of the sensor is relatively large due to the impact
of noise. On the other hand, it is difficult to apply machine learning methods to improve the
intelligence and recognition effect due to few or no actual measurement samples. Aiming at
these problems, an aerial target recognition algorithm based on self-attention and Long
Short-Term Memory Network (LSTM) is proposed. LSTM can effectively extract temporal
dependencies. The attention mechanism calculates the weight of each input element and
applies the weight to the hidden state of the LSTM, thereby adjusting the LSTM’s attention
to the input. This combination retains the learning ability of LSTM and introduces the
advantages of the attention mechanism, making the model have stronger feature extraction
ability and adaptability when processing sequence data. In addition, based on the prior
information of the multi-dimensional characteristics of the target, the three-point estimation
method is adopted to simulate an aerial target recognition dataset to train the recognition
model. The experimental results show that the proposed algorithm achieves more than 91%
recognition accuracy, lower false alarm rate and higher robustness compared with the multi-
attribute decision-making (MADM) based on fuzzy numbers.
Keywords: Aerial target recognition; long short-term memory network; self-attention; three-
point estimation
Jongyun Byun, Jaehoon Cha, Jeyan Thiyagalingam, Hyeon-Joon Kim, Changhyun Jun,
Enhancing rainfall prediction accuracy through image fusion of radar and numerical weather
prediction models,
Expert Systems with Applications,
Volume 303,
2026,
130516,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Precipitation is one of the most challenging atmospheric phenomena to predict
due to the complexity involved in solving dynamic and thermodynamic atmospheric
equations. To address this challenge, extensive research has been conducted to enhance the
precision of numerical weather prediction models and radar-based extrapolation data, in
conjunction with the development of various blending techniques. However, traditional
methods have proven insufficient in capturing the diversity and nonlinearity of weather
phenomena. In response to these limitations, this study introduces a novel methodology
that leverages machine-learning based image fusion models to merge radar-based
extrapolation and numerical weather prediction rainfall datasets, thereby enhancing
prediction accuracy. An image fusion model was developed using radar-based extrapolation
data and numerical weather prediction data as input datasets, with radar observation data
utilized as target dataset. To identify the most suitable image fusion model for capturing the
complex patterns of rainfall data, two experiments were conducted: 1) Impact of model
topology, and 2) Effect of model size. A systematic analysis of the model outputs was
performed using eight evaluation metrics categorized under pixel-based metrics, feature-
based metrics, structural similarity metrics, and categorical verification metrics.
Experimental results indicated that image fusion model based on a Residual Network
(ResNet) outperformed other models in terms of model topology. Regarding model size, it
was observed that the performance did not increase proportionally with the number of
residual blocks; the most suitable performance was achieved with a specific number of
residual blocks (Case 5: 8 blocks). Additionally, the metrics compared with radar observation
data indicated that the proposed model delivered superior performance, thus offering a
high-accuracy rainfall prediction methodology.
Keywords: Image fusion; Radar; Numerical weather prediction; Deep learning; Precipitation
Jackson S. Zaunegger, Paul G. Singerman, Ram M. Narayanan, Muralidhar Rangaswamy,
RadarTD: A Radar Text Dataset for multi-parameter optimization,
Natural Language Processing Journal,
Volume 12,
2025,
100178,
ISSN 2949-7191,
[Link]
([Link]
Abstract: This paper introduces the radar text dataset (RadarTD) for technical language
modeling. This dataset is comprised of sentences containing radar parameters, values, and
units determined from published radar literature. Additionally, each statement is assigned a
sentiment, goal priority, and goal direction label. In this work, we show how RadarTD may be
used to train simple Natural Language Processing (NLP) models to identify the attributes of
each sentence listed in RadarTD. Once the NLP models have identified these attributes from
text, we can use this information to develop Language Based Cost Functions (LBCF). Our
study shows that the proposed text classification model achieves a classification accuracy
between 96.7% and 97.8%, while the proposed named entity recognition model achieves an
F1 score of 99.7. These findings suggest that the developed models are capable of achieving
good performance for both text classification and named entity recognition for autonomous
radar applications. We then illustrate an example of how these models could be used with
Language Based Cost Functions to develop multi-parameter radar optimization schemes. We
also provide a method of providing scalarization weights for each parameter, to improve the
results of the optimization process.
Keywords: Text classification; Named entity recognition; Language modeling; Language-
based cost functions; Multi-parameter optimization; Cognitive radar
Changlong Wang, Jiawei Jiang, Chong Han, Hengyi Ren, Lijuan Sun, Jian Guo,
Through-Wall Multihuman Activity Recognition Based on MIMO Radar,
Computers, Materials and Continua,
Volume 83, Issue 3,
2025,
Pages 4537-4550,
ISSN 1546-2218,
[Link]
([Link]
Abstract: Existing through-wall human activity recognition methods often rely on Doppler
information or reflective signal characteristics of the human body. However, static
individuals, lacking prominent motion features, do not generate Doppler information.
Moreover, radar signals experience significant attenuation due to absorption and scattering
effects as they penetrate walls, limiting recognition performance. To address these
challenges, this study proposes a novel through-wall human activity recognition method
based on MIMO radar. Utilizing a MIMO radar operating at 1–2 GHz, we capture activity data
of individuals through walls and process it into range-angle maps to represent activity
features. To tackle the issue of minimal variation in reflection areas caused by static
individuals, a multi-scale activity feature extraction module is designed, capable of extracting
effective features from radar signals across multiple scales. Simultaneously, a temporal
attention mechanism is employed to extract keyframe information from sequential signals,
focusing on critical moments of activity. Furthermore, this study introduces an activity
recognition network based on a Deformable Transformer, which efficiently extracts both
global and local features from radar signals, delivering precise human posture and activity
sequences. In experimental scenarios involving 24 cm-thick brick walls, the proposed
method achieves an impressive 97.1% accuracy in activity recognition classification.
Keywords: MIMO radar; human activity; Transformer; through-wall
Jinyang Xie, Kanghui Zhou, Lei Han, Liang Guan, Maoyu Wang, Yongguang Zheng, Hongjin
Chen, Jiaqi Mao,
Enhancing multi-task learning-based Tornado identification using spatial and temporal
information from weather radar images,
Applied Soft Computing,
Volume 184, Part B,
2025,
113834,
ISSN 1568-4946,
[Link]
([Link]
Abstract: Tornadoes, as dynamic weather phenomena, exhibit unique spatial and temporal
evolution characteristics that reflect their formation and development. Existing tornado
detection algorithms often struggle with high false alarm rates, primarily due to insufficient
capture of temporal correlations in tornado development. As an improvement, we propose a
multi-task tornado identification network with three-dimensional temporal and spatial
information (TS-MTINet). Taking continuous three-frame radar data as input, the Multi-
frame Temporal Interaction Block (MTIB) utilizes multi-head attention to model the dynamic
interaction information between the radar data, thus exploring in-depth the temporal
features during tornado development. Further, we design a Spatial-Temporal Enhancement
Module (STEM), which analyzes the difference information between continuous data to
extract local and global spatial and temporal feature variations about tornadoes. Based on
this architecture, TS-MTINet incorporates a multi-task learning framework to perform
tornado detection and number estimation tasks simultaneously, thus extracting
comprehensive information related to tornadoes. To validate the performance of the
proposed model, we construct the first Chinese tornado identification dataset with fine
radar features. The experimental results show that the proposed method shows significant
advantages in several evaluation metrics, especially in reducing false alarms. In practical case
studies, compared to the traditional TVS method, TS-MTINet achieves an increase in POD of
approximately 30% and a decrease in FAR of about 20% in several typical tornado events.
Particularly in environments with strong interference, TS-MTINet demonstrates higher
detection accuracy, reflecting greater robustness and practical value.
Keywords: Deep learning; Multi-task learning; Tornado identification; Weather radar;
Attention mechanisms
Jiaxiang Zhang, Bo Wang, Xinrui Han, Min Zhao, Zhennan Liang, Xinliang Chen, Quanhua Liu,
A multi-radar emitter sorting and recognition method based on hierarchical clustering and
TFCN,
Digital Signal Processing,
Volume 160,
2025,
105005,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Radar emitter recognition is an important part of the electronic attack and defense
system, and its key task is to sort and identify various information of emitters from mixed
pulse streams. Pulse descriptor word, as an easily obtainable feature, is often used for
sorting and recognition. Aiming at the problems of high complexity and difficulty in handling
intra-class clustering and inter-class aliasing of existing sorting methods, a hierarchical
clustering method based on Kernel Density Estimation-Kullback-Leibler Divergence-Template
Matching (KDE-KLD-TM) is proposed. Firstly, the down-sampled data is used to construct
central clusters, greatly improving the processing speed. Then, with probability theory as the
theoretical support, inter-class clustering on all samples is performed based on maximum
posterior probability. Finally, based on cluster distribution similarity and periodic template
matching, intra-class merging and inter-class deinterleaving are completed. After sorting,
considering the periodic differences in pulse repetition intervals among different types of
emitter and the insufficient attention paid by existing recognition methods to this feature, a
time-frequency convolution network (TFCN) based emitter recognition method is proposed
for the first time in terms of pulse description word. Using time-frequency analysis (TFA) to
extract periodic features and using convolutional neural networks (CNN) for classification,
the one-dimensional sequence classification problem is treated as a two-dimensional image
classification problem. The proposed method is simulated in a typical scenario with intra-
class clustering and inter-class aliasing. The results show that the proposed method can sort
a total of 2.06 million aliased samples composed of 10 classes within 58.61 s, and the
recognition accuracy reaches 96.33%. The comparison with the baseline method proves the
effectiveness and progressiveness of the proposed method. Finally, the sorting and
recognition performance of the proposed algorithm in complex scenarios in the presence of
pulse loss is discussed.
Keywords: Radar emitter sorting and recognition; Hierarchical clustering; Time-frequency
convolution network; Intra-class clustering; Inter-class aliasing
Xiaodan Wang, Rui Li, Jian Wang, Lei Lei, Yafei Song,
One-dimension hierarchical local receptive fields based extreme learning machine for radar
target HRRP recognition,
Neurocomputing,
Volume 418,
2020,
Pages 314-325,
ISSN 0925-2312,
[Link]
([Link]
Abstract: Radar automatic target recognition (RATR) aims at extracting meaningful target
features from the electromagnetic echo signal and utilizing the features to automatically
recognize the target types. The high-resolution range profile (HRRP) plays an important role
in RATR field, HRRP is the amplitude of the echo summation for target scattering centers in
each range cell of wideband radar. Using deep neural networks for HRRP radar target
recognition encounters the problem of storage overhead and slow convergence rate, to
resolve those issues, we propose a one-dimension local receptive fields based extreme
learning auto-encoder (1D ELM-LRF-AE) network for HRRP local structures and meaningful
representations learning. ELM-LRF-AE consists of an input layer, a random convolution layer,
a pooling layer, several local connected layers and an output layer, it reconstructs the input
with a greedy strategy that the input feature vectors are divided into several subgroups and
the i-th pooling feature vector is used to reconstruct the i-th grouping input feature vector.
Then we use the learned pooling feature vectors to replace the random pooling feature
vectors as the learned representations. We also stack several 1D ELM-LRF-AEs to build 1D
hierarchical local receptive fields based extreme learning machine (1D H-ELM-LRF) for high
level HRRP abstract representations learning and recognition. Experimental results on
simulated HRRP data set demonstrate the superior high recognition performance and high
computational efficiency of our algorithm.
Keywords: Local receptive fields based extreme learning machine; Auto-encoder; HRRP
recognition; Deep learning
Guanhong Lu, Lei Kou, Pei Niu, Gaohang Lv, Xiao Zhang, Jian Liu, Quanyi Xie,
GPRTransNet: A deep learning–based ground-penetrating radar translation network,
Tunnelling and Underground Space Technology,
Volume 161,
2025,
106557,
ISSN 0886-7798,
[Link]
([Link]
Abstract: Ground-penetrating radar (GPR) is an essential nondestructive testing tool widely
used in tunnel and road defect detection, underground object detection, and unstructured
terrain perception. Currently, most GPR inversion methods based on deep learning use
convolutional neural networks, which have limitations such as incomplete feature extraction
and low accuracy. Inspired by advancements in the NLP field, this paper proposes a novel
deep learning framework for GPR called GPRTransNet. The algorithm introduces a “wave to
permittivity” translation architecture, leveraging the “memory” function of recurrent neural
networks to translate GPR data into permittivity model images, similar to language
translation. In addition, the inclusion of the attention mechanism significantly enhances the
network’s ability to represent defects in complex scenarios, resulting in outstanding
translation performance. GPRTransNet has been validated on a simulated dataset, with the
newly proposed G-SSIM evaluation metric showing that the permittivity similarity of
GPRTransNet-Lite reaches 98.47% and that of GPRTransNet-Pro is 99.26%. Furthermore,
GPRTransNet demonstrates good translation performance on real data. Experimental results
show that the results of GPRTransNet’s translation of GPR are satisfactorily matched.
Keywords: Ground-Penetrating Radar Data Translation; Recurrent Neural Networks; Deep
Learning; Gated Recurrent Unit
Adil Ali Saleem, Hafeez Ur Rehman Siddiqui, Muhammad Amjad Raza, Sandra Dudley, Julio
César Martínez Espinosa, Luis Alonso Dzul López, Isabel de la Torre Díez,
Ultra Wideband radar-based gait analysis for gender classification using artificial intelligence,
Array,
Volume 27,
2025,
100477,
ISSN 2590-0056,
[Link]
([Link]
Abstract: Gender classification plays a vital role in various applications, particularly in
security and healthcare. While several biometric methods such as facial recognition, voice
analysis, activity monitoring, and gait recognition are commonly used, their accuracy and
reliability often suffer due to challenges like body part occlusion, high computational costs,
and recognition errors. This study investigates gender classification using gait data captured
by Ultra-Wideband radar, offering a non-intrusive and occlusion-resilient alternative to
traditional biometric methods. A dataset comprising 163 participants was collected, and the
radar signals underwent preprocessing, including clutter suppression and peak detection, to
isolate meaningful gait cycles. Spectral features extracted from these cycles were
transformed using a novel integration of Feedforward Artificial Neural Networks and
Random Forests , enhancing discriminative power. Among the models evaluated, the
Random Forest classifier demonstrated superior performance, achieving 94.68% accuracy
and a cross-validation score of 0.93. The study highlights the effectiveness of Ultra-wideband
radar and the proposed transformation framework in advancing robust gender classification.
Keywords: Gait; Ultra-wide band radar; Gender classification; Spectral features; Feed
forward artificial neural network; Ridge classifier; Hist gradient boosting
Xueqing Zhao, Ren Xu, Yutao Zhang, Andrew Ty Lau, Ruitian Xu, Xingyu Wang, Andrzej
Cichocki, Jing Jin,
A novel paradigm based on radar-like scanning for directional recognition in event-related
potentials based brain-computer interfaces,
Journal of Neuroscience Methods,
Volume 423,
2025,
110546,
ISSN 0165-0270,
[Link]
([Link]
Abstract: Background
Event-related potentials (ERPs) based brain-computer interface (BCI) systems have shown
significant potential for directional control applications. Existing paradigms are constrained
by the limited scalability of directional commands that demand interface reconfiguration for
varying target numbers.
New method
We propose a novel radar-like scanning (RS) paradigm for 32-directional recognition tasks to
address these limitations. This paradigm continuously scans through directions using a
sector-shaped visual stimulus, naturally evoking ERP responses without discrete directional
indicators. During the online experiments, an early-stopping strategy is employed to
enhance efficiency. Additionally, this study analyzes subjects' directional recognition
performance using EEGNet under three sector rotation periods. Thirteen subjects
participated in the experiments.
Results
The grand-averaged ERP amplitudes exhibited a stronger negative deflection in the parietal,
occipital, and temporoparietal regions. The results demonstrated that, with a 2 s rotation
period and early-stopping strategy, the best subject achieved an accuracy of 87.50 % with a
mean absolute angle error of 1.64°. When the directional error tolerance was set to 11.25°,
the subject-averaged accuracy reached 91.83 % under the same conditions. Longer rotation
periods led to better subject-averaged recognition performance. When the rotation period
was short (1 s), targets close to the scanning center were challenging to recognize.
Comparison with existing methods
Compared with others, the RS paradigm enables more fine-grained directional target
recognition and is unaffected by the target numbers.
Conclusions
The proposed paradigm demonstrates significant potential for applications in ERP-BCI
systems.
Keywords: Brain-computer interface (BCI); Electroencephalography (EEG); Event-related
potential (ERP); Radar-like scanning; Directional recognition
Yuwen Wu,
Fusion-based modeling of an intelligent algorithm for enhanced object detection using a
Deep Learning Approach on radar and camera data,
Information Fusion,
Volume 113,
2025,
102647,
ISSN 1566-2535,
[Link]
([Link]
Abstract: Object detection, the process of detecting and classifying objects within a given
environment, forms the foundational element. Multisensory fusion incorporates data from
diverse sensors, like radar and cameras, to refine the reliability and accuracy of detection.
Further, Radar and camera data fusion refine this process by integrating the unique strength
of both technologies, which leverage the radar's proficiency in adverse weather conditions
and the camera's high-resolution imaging. This incorporation enhances the object detection
systems, which enables them to effectively operate across the spectrum of scenarios, from
autonomous vehicles navigating challenging weather to surveillance systems monitoring
critical infrastructure. Deep learning (DL), a branch of machine learning (ML), empowers this
system with the capability to learn complex representations and patterns directly from the
data, which enables them to generalize and adapt to new situations. By integrating the
advanced methodology, we can develop strong perception system capable of interpreting
and detecting objects accurately in dynamic and diverse environments, from autonomous
vehicles navigating urban landscapes to surveillance systems monitoring complex
environments. This study designs an Intelligent Algorithm for Enhanced Object Detection
Using Deep Learning Approach on the Radar and Camera Data Fusion (IAEOD-DLRCDF)
technique. The presented IAEOD-DLRCDF technique uses multi-angle joint calibration where
the spatial sparse alignment of the heterogeneous data of the camera and Radar is realized
with image falsification disregarded. Besides, the IAEOD-DLRCDF technique applies YOLOv8
object detector for radar and camera target detection individually which are then integrated
with the image plane. Moreover, the detected objects are then classified via the
bidirectional long short-term memory (BiLSTM) model. Furthermore, the Adam optimizer is
used for the optimum hyperparameter selection of the BiLSTM network which results in a
better recognition rate. The performance assessment of the IAEOD-DLRCDF method is tested
under benchmark dataset. The empirical analysis stated that the IAEOD-DLRCDF method
gains better performance over other models.
Keywords: Object detection; Deep learning; Data fusion; Radar; YOLOv8; Adam optimizer;
Machine learning
Haojie Wei, Min Fang, Haixiang Li, Yinan Wang, Zhanpeng Zheng,
Prototype-based dual-alignment of multi-source domain adaptation for radar emitter
recognition,
Signal Processing,
Volume 230,
2025,
109853,
ISSN 0165-1684,
[Link]
([Link]
Abstract: Recent research on radar emitter recognition has primarily focused on single-
source domain adaptation, neglecting the potential benefits of leveraging multiple source
domains. While multi-source domain adaptation (MDA) models have been extensively
developed for image data, they often fail to address the unique challenges presented by
radar emitter data, such as domain class hierarchy discrepancies and intra-class
compactness. To address these challenges, this study proposes a prototype-based dual-
alignment (PBDA) method specifically designed for multi-source radar emitter recognition
tasks. The PBDA method incorporates both domain-level and class-level alignment. For
domain-level alignment, a reconstruction loss is introduced to preserve classification-
relevant features during adversarial learning. For class-level alignment, a prototype
alignment loss is proposed to minimize distributional discrepancies of the same class across
different domains, reducing fine-grained distribution gaps between each source domain and
the target domain. To ensure compact sample distribution within the target domain and
avoid negative transfer effects, information entropy loss and classification consistency loss
are applied, guiding target domain samples toward the correct class prototypes.
Experimental validation on both image benchmark datasets and radar emitter datasets
demonstrates the effectiveness of the proposed PBDA method.
Keywords: Domain adaptation; Radar emitter recognition; Multi-source; Adversarial learning
Wenxu Zhang, Yajie Wang, Xiuming Zhou, Zhongkai Zhao, Feiran Liu,
An interference power allocation method against multi-objective radars based on optimized
proximal policy optimization,
Signal Processing,
Volume 230,
2025,
109785,
ISSN 0165-1684,
[Link]
([Link]
Abstract: Aiming at the problem of interference resource scheduling in cognitive electronic
warfare, a multi-objective interference power allocation method based on the proximal
policy optimization (PPO) framework is proposed in this paper. Firstly, the confrontation
between jammers and multi-objective radar networks is mapped as the interaction between
the agent and the environment, and the radar target detection model under suppression
interference is established. On this basis, an interference power allocation model against
multi-objective radars based on PPO framework is constructed. Moreover, a reward
normalization mechanism is introduced to optimize the reward setting, and an interference
power allocation method based on optimized PPO is proposed. Meanwhile, this paper
constructs a confrontation scenario in which the jammer covers the target aircraft to break
through the multi-objective radar network. Simulation experiments are conducted based on
this scenario to verify the effectiveness of the method proposed in this paper. The
interference power allocation method proposed in this paper can intelligently adjust the
power allocation scheme of the jammer according to the electromagnetic situation on the
battlefield, optimize the resource utilization of the jammer, and occupy the initiative on the
battlefield.
Keywords: Interference power allocation; Proximal policy optimization; Limited interference
resources; Reward normalization
Jiayi Cai, Zhaocheng Yang, Ping Chu, Juntao Guo, Jianhua Zhou,
Robust hand gesture detection and recognition using 4D millimeter-wave radar in a
ubiquitous scene,
Measurement,
Volume 253, Part C,
2025,
117545,
ISSN 0263-2241,
[Link]
([Link]
Abstract: In current research on HGR using radar sensors, hand gestures are typically
confined to a smaller region. However, in ubiquitous scenarios, unrestricted human body
movements and unexpected hand gesture motions usually occur, which results in a large
false alarms and recognition performance degradation. To address this issue, we propose a
robust hand gesture detection and recognition method in ubiquitous scenarios using
Frequency-Modulated Continuous Wave (FMCW) Multiple-Input Multiple-Output (MIMO)
radar. The core idea is to progressively define and classify motions in a cascaded manner,
gradually filtering out non-specific movements, reducing false positives, and enhancing the
applicability of HGR. Specifically, we first propose a suspected hand gesture motion
detection method to help identify suspicious hand gestures. Then, the velocity and position
features of the mutated signal and the stable signal are extracted. A mutated signal motion
recognition method based on a single-layer long short-term memory (LSTM) network is used
to effectively distinguish non-hand gesture motions from hand gestures. Finally, the two-
dimensional trajectory features are extracted, and cascaded with a LSTM network combined
a Gaussian probability model is developed to enhance the ability of open-set recognition.
Experimental results show that the proposed method can achieve the recognition accuracy
of 99.53% for designed hand gestures, the false alarm rate of 1.5% for unexpected hand
gestures and 0.11% for non-hand gesture motions.
Keywords: Hand gesture recognition; Non-hand gesture motion; Feature extraction;
Probability models; Ubiquitous scene
Yue XU, Quan PAN, Zengfu WANG, Hua LAN, Shuling JIN,
A self-learning refined model and tracking for near space hypersonic vehicle by space-based
radar,
Chinese Journal of Aeronautics,
2025,
103840,
ISSN 1000-9361,
[Link]
([Link]
Abstract: The Near Space Hypersonic Vehicle (NSHV) features a unique design and
propulsion system, achieving exceptional speed, range, and maneuverability, which
challenge ground-based radars. Space-Based Radar (SBR) offers a breakthrough for tracking
NSHV targets, with all-weather operation and freedom from Earth’s curvature, but faces
complex coordinate transformations. Traditional models often overlook the NSHV’s dynamic
gliding trajectory, especially the impact of hidden control variables on maneuvering, causing
mismatches during rapid motion changes. This paper proposes a refined tracking model
unified in the ECEF coordinate frame, incorporating model parameters that implicitly encode
control laws, and presents an Expectation-Maximization Multi-swarm Cooperative Particle
Swarm Optimization (EM-MCPSO) framework for both NSHV tracking and model parameter
estimation to address this problem. To minimize conversion errors, a transformation matrix
directly represented by the state in the Earth-Centered Earth-Fixed (ECEF) coordinate is
derived. Then the hybrid aerodynamic acceleration coefficients are introduced to precisely
describe the dynamic behaviors, formulating target tracking as a joint estimation problem of
state and parameters within EM framework. Finally, a self-learning algorithm based on a
master-slave structured PSO is proposed to solve the optimization of the conditional
expectations of EM under strong nonlinearity, with a Proportional-Derivative (PD) controller
accelerating convergence, and updating the population structure with historical data.
Simulations of vertical gliding and horizontal maneuvers validate the algorithm’s
effectiveness.
Keywords: Dynamics modeling; Expectation Maximization (EM); Maneuvering target
tracking; Near Space Hypersonic Vehicle (NSHV); Particle Swarm Optimization (PSO)
Dongming Wu, Junpeng Shi, Zhiyuan Zhang, Zhihui Li, Fangling Zeng,
Generative-contrastive learning for open set radar emitter identification,
Signal Processing,
Volume 239,
2026,
110295,
ISSN 0165-1684,
[Link]
([Link]
Abstract: In traditional radar emitter identification (REI) tasks, both the training and testing
samples share the same distribution, and the model is trained solely to recognize known
targets. However, in non-cooperative electromagnetic environments, unknown classes are
often absent from the training data, which may be incorrectly classified as known classes. To
address this issue, we propose an innovative Generative-contrastive Learning method for
Open Set REI (GLOSE) from the perspective of feature space optimization. We first introduce
a conditional generative model derived from diffusion to generate stable interpolated
samples within the feature space, which are defined as an additional class to compress the
coverage of known classes, thereby enhancing the capability to handle unknown space.
Subsequently, we employ contrastive learning with an adaptive contrastive loss to further
optimize the discriminative power of the feature space, which applies varying levels of intra-
class similarity for different types of samples. Extensive experiments are conducted on a
simulated radar emitter dataset based on intra-pulse unintentional modulation and a real-
world automatic dependent surveillance-broadcast (ADS-B) dataset. The results
demonstrate that the proposed method significantly improves the detection capability of
unknown class samples while maintaining high classification accuracy for known classes.
Keywords: Radar emitter identification; Open set recognition; Generative-contrastive
learning; Intra-pulse unintentional modulation
Qihang Zhai, Xiongkui Zhang, Zilin Zhang, Jiabin Liu, Shafei Wang,
Online few-shot learning for multi-function radars mode recognition based on backtracking
contextual prototypical memory,
Digital Signal Processing,
Volume 141,
2023,
104189,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Modern multi-function radars (MFRs) can flexibly generate multiple fine-grained
working modes for different missions through programmable parameters. Precise
recognition of these working modes by analyzing the parameter combinations lays
foundation for situational awareness. This recognition becomes more challenging to
electromagnetic reconnaissance system when prior information of radiation source is
unavailable and effective labeled signal data is not sufficient in adversarial scenarios. The
presence of incremental mode involved in radar data flow requires that the recognition
model enables to dynamic adjust and online perceive current data. This issue incorporates
online learning to few-shot learning to accomplish efficiently recognition to a novel mode
with the support of small amount of data, and still retain model's ability to recognize existing
modes. This paper designed a backtracking contextual prototypical memory (BCPM) network
for online MFR mode recognition. The proposed method learns temporal and spatial
information from data stream as a reference predicting the signal segments as existing
working modes or a novel mode. The BCPM network also designs a backtracking module and
an alignment regularization term to utilize the data without annotation adequately and
obtain a more reliable category representation. The experimental results verified the
proposed method's good performance and high robustness to non-ideal factors and
distractors.
Keywords: Signal modulation recognition; Radar signal hierarchical modeling; Few-shot
learning; Online learning; Incremental learning
Zhipeng Qing, Kecheng Ge, Shunsheng Zhang, Jing Yang, Zhijin Wen, Youlei Pu,
An inverse synthetic aperture radar imaging framework based on multi-layer networks and
heat conduction attention,
Engineering Applications of Artificial Intelligence,
Volume 167, Part 1,
2026,
113708,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Accurate compensation is essential for achieving high-resolution inverse synthetic
aperture radar (ISAR) imaging. Traditional parametric methods usually rely on iterative
optimization of the objective function to compensate for target motion in radar echoes.
However, the iteration process is often computationally intensive, difficult to integrate into
deep learning frameworks, and may discard sufficiently acceptable intermediate solutions.
To address these challenges, this study proposes a deep unfolding-based translational
compensation network that combines unsupervised learning with gradient back-
propagation. A prototype network is incorporated to monitor the imaging process, enabling
early termination of iterations. Moreover, a U-shaped network architecture based on a heat
conduction attention is employed to enhance ISAR image resolution and focusing
performance. To solve the problem of offset or splitting in the imaging results caused by
residual motion errors, a learnable affine transformation is employed for automatic
centering. These modules are integrated into an echo-to-image ISAR imaging framework.
Experimental results on both simulated and real radar data demonstrate the framework’s
effectiveness and robustness.
Keywords: Inverse synthetic aperture radar imaging; Translational compensation; Heat
conduction; Deep unfolding network; Affine transformation
Yuanbo Li, Wenwu Zhang, Songtao Lv, Jing Yu, Dongdong Ge, Jiawei Guo, Lin Li,
YOLOv11-CAFM model in ground penetrating radar image for pavement distress detection
and optimization study,
Construction and Building Materials,
Volume 485,
2025,
141907,
ISSN 0950-0618,
[Link]
([Link]
Abstract: Ground Penetrating Radar (GPR) is an effective technology for detecting
underground structures and has been widely utilized for monitoring road damage.
Traditional B-scan-based one-dimensional images often fail to preserve continuous spatial
information, thus inadequately reflecting the nuances of damage patterns. This paper
investigates the accurate recognition of hidden internal road damage using 3D-sliced C-scan
images. While YOLO is one of the most effective and rapid neural network models for object
detection, it still suffers from low recognition accuracy and a high rate of missed detections.
To address these issues, this study proposes an improved Convolution and Attention Fusion
Module (CAFM) fusion network model for YOLOv11, which combines the CAFM with the
C2PSA global-local feature extraction mechanism to significantly enhance the recognition
performance for complex road damage. Experimental comparisons between the YOLOv11m-
CAFM and the YOLOv11 model reveal that the combined metrics for the small (n/s) and large
(l/x) models are lower than those for the medium model (m). The YOLOv11m-CAFM
demonstrates strong performance in key metrics such as precision, recall, mAP50, and
mAP50:95, achieving values of 0.840, 0.850, 0.881, and 0.584, respectively, representing
improvements of 0 %, 4.6 %, 1.8 %, and 2.0 % over the baseline model. The confidence level
in detecting standardized targets (e.g., pipelines and well covers) exceeds 0.89. Borehole
validation confirms that the model's localization error is less than 0.15 m, and the detection
frame rate reaches 71 FPS, satisfying the requirements for rapid road assessment. This study
offers a novel method for the intelligent interpretation of GPR images, considering both
detection accuracy and real-time performance, which holds significant engineering
applications in identifying hidden road damages.
Keywords: Pavement disease detection; Ground penetrating radar; Neural network; Object
detection; Deep learning algorithm
Haoming Feng, Huaqing Li, Wenwen Zhu, Denghao Li, Yukun Huang,
Micro-motion enhanced multi-person activity recognition with millimeter-wave radar,
Measurement,
Volume 258, Part B,
2026,
119090,
ISSN 0263-2241,
[Link]
([Link]
Abstract: As a non-contact sensing device, millimeter-wave radar exhibits unique strengths
in human activity recognition (HAR). Existing methods rely on micro-Doppler signatures for
activity classification, but they often encounter feature aliasing in multi-person activity
recognition (MPAR) scenarios. Although point cloud-based approaches can distinguish
individual targets, they primarily extract static morphological features, neglecting the micro-
motion information of human joints, which is crucial for accurate activity recognition. To
address these limitations, we proposes an innovative MPAR framework that integrates
spatial point clouds and micro-motion features. First, an improved point cloud data
association algorithm is applied to achieve multi-target point cloud feature separation,
followed by a dynamic projection mechanism to construct time–Doppler feature maps.
Then, a torso micro-motion enhancement algorithm is designed to enhance the details of
human body movements. Finally, a CNN-LSTM hybrid network architecture with a temporal-
attention is constructed for action classification. Experimental results show that the
proposed micro-motion enhancement algorithm improves recognition accuracy by 27.1%
and 2.3%, compared to two traditional time–frequency analysis methods. Furthermore,
MPAR task in occlusion scenarios achieves recognition accuracy of 93.5%. In summary,
proposed framework not only retains the inherent advantages of millimeter-wave radar but
also significantly enhances multi-person activity recognition in complex scenarios.
Keywords: Human activity recognition (HAR); Multi-person activity recognition (MPAR);
Feature separability; Temporal attention; Occlusion scenarios
Guanliang Liu, Wenchao Chen, Bo Chen, Bo Feng, Penghui Wang, Hongwei Liu,
Supervised contrastive deep Q-Network for imbalanced radar automatic target recognition,
Pattern Recognition,
Volume 161,
2025,
111264,
ISSN 0031-3203,
[Link]
([Link]
Abstract: In the presence of limited and extremely imbalanced data, deep learning methods
for radar automatic target recognition (RATR) often suffer from significant performance
degradation and overfitting. To tackle this issue, we propose Supervised Contrastive Deep Q-
network (SCDQ), a novel end-to-end reinforcement learning method, for multi-class
imbalanced RATR. SCDQ formulates the imbalanced recognition problem as a Markov
decision process (MDP) and optimizes the classifier through an enhanced Q-learning
paradigm. In order to augment the model’s feature extraction capabilities under the
constraint of limited samples, we tightly integrate reinforcement learning (RL) with
supervised contrastive learning, introducing an innovative feature enhancement module. To
further enhance the model’s adaptability to challenging samples, we integrate a
meticulously designed priority sampling into the proposed SCDQ framework, denoted as
SCDQ-P. Experimental results on both simulated and real datasets demonstrate the reliability
and effectiveness of the proposed method.
Keywords: Imbalanced RATR; Deep learning; Deep reinforcement learning (DRL); Supervised
contrastive learning; Priority sampling
Ligen Chen, Nannan Zhu, Hongbo Chen, Yonghao Dong, Yue Zhang, Nian Cai,
A causality-inspired single-source domain generalized method for low-slow-small threat
target recognition through holographic Doppler radar,
Expert Systems with Applications,
Volume 287,
2025,
128104,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Low-slow-small (LSS) target monitoring is critical for airport safety management,
particularly when LSS objects such as unmanned aerial vehicles (UAVs) and birds
unexpectedly enter airport airspace, posing significant risks to regular flights and airport
operations. Recent advances have taken advantage of deep learning for the recognition of
LSS radar targets, achieving promising classification accuracy. However, existing LSS radar
target recognition approaches often rely on statistical correlations, including unstable
spurious correlations, which can undermine the generalization performance of classification
networks, limiting their effectiveness in all-time radar recognition. To address this, we
propose a causality-inspired single-source domain generalization method for radar LSS target
recognition. Our method introduces a Causal-Symmetric Transformation (CST) module for
data augmentation, combining Non-Causal Augmentation for global perturbations and
Symmetric Transformation for local motion reversal, enhancing data diversity and reducing
bias. Additionally, we propose a Causal Mining (CM) module with a Causal Consistency loss
to extract causal features that boost generalization. A Fourier-Aware Attention (FAA) module
leverages frequency-domain information to strengthen feature representation and preserve
causal information. Extensive experiments on four real-world datasets validate the
effectiveness of our approach.
Keywords: Radar target recognition; Single-source domain generalization; Causility-inspired
model; Low-slow-small target; Holographic Doppler radar
Ziwei Zhang, Mengtao Zhu, Yunjie Li, Yan Li, Shafei Wang,
Joint recognition and parameter estimation of cognitive radar work modes with LSTM-
transformer,
Digital Signal Processing,
Volume 140,
2023,
104081,
ISSN 1051-2004,
[Link]
([Link]
Abstract: The recent developed cognitive radars can implement flexible work modes with
programmable modulation types and optimized modulating values for each mode definition
parameter. Automatic analysis of these work modes is a significant challenge for modern
electromagnetic reconnaissance receivers. In this paper, a Multi-Output Multi-Structure
(MOMS) learning-based processing framework is proposed for Joint inter-pulse automatic
Modulation Recognition and Parameter Estimation (JMRPE-MOMS). We propose a label
construction method as a feature interpretation method of the network to facilitate MOMS
learning and utilize the correlations between labels for performance gain. Moreover, an
LSTM-Transformer is designed to mine deep time-series characteristics, which can model
local and global relationships and reduce quantization loss. The proposed framework can
perform joint modulation recognition and parameter estimation (JMRPE) tasks
simultaneously with flexible output structures including scalar output and vector output
with fixed or variable sizes. Extensive simulations are performed based on the simulated
radar work modes defined with pulse repetition interval (PRI) sequences. The simulation
results validate the effectiveness and superiority of the proposed method especially under
non-ideal electromagnetic environments.
Keywords: Radar work mode; Automatic modulation recognition; Modulation parameter
estimation; Multi-output learning; Transformer
Xiaoyuan Zhang, Shaohang Jing, Jingshu Li, Yechao Bai, Feng Yan,
Cognitive radar recognition with Kolmogorov-Smirnov test and momentum gradient descent,
Digital Signal Processing,
Volume 163,
2025,
105212,
ISSN 1051-2004,
[Link]
([Link]
Abstract: The emission parameters of cognitive radars can adaptively change according to
the environment, which poses a challenge to radar electronic countermeasures (ECM). To
counter cognitive radars, it is essential to identify the cognitive characteristics. In this paper,
a method is proposed to recognize cognitive radars with power allocation function. The
signal-to-interference-plus-noise ratio (SINR) distribution of cognitive radars is derived
through feature functions, and hypothesis test is used to identify whether the target radar
has cognitive function by designing a Kolmogorov-Smirnov (K-S) detector to recognize
adaptive optimization power allocation. Subsequently, a momentum gradient descent
algorithm is used to optimize the signal of the jamming machine to reduce type II error
probability of radar recognition. K-S detector is simulated and compared with Afriat
detector, SVM and MLP detector. Results demonstrate that the K-S detector outperforms
both the Afriat and MLP detectors in identifying cognitive radars with dynamic power
allocation functionality. At the same detection probability, the K-S detector achieves a 2 dB
improvement over the MLP detector and a 4 dB improvement over the Afriat detector.
Keywords: Cognitive radar; Electronic countermeasures (ECM); Kolmogorov–Smirnov test;
Momentum gradient descent algorithm; Afriat theorem
Baiju Yan, Peng Wang, Lidong Du, Xianxiang Chen, Zhen Fang, Yirong Wu,
mmGesture: Semi-supervised gesture recognition system using mmWave radar,
Expert Systems with Applications,
Volume 213, Part B,
2023,
119042,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Gesture recognition has found versatile applications in natural human–computer
interaction (HCI). Compared with traditional camera-based or wearable sensors-based
solutions, gesture recognition using the millimeter wave (mmWave) radar has attracted
growing attention for its characteristics of contact-free, privacy-preserving and less
environment-dependence. Recently, most of studies adopted one of the Range Doppler
Image (RDI), Range Angle Image (RAI), Doppler Angle Image (DAI) or Micro-Doppler
Spectrogram extracted from the raw radar signal as the input of a deep neural network to
realize gesture recognition. However, the effectiveness of these four inputs in gesture
recognition has attracted little attention so far. Moreover, the lack of large amounts of
labeled data restricts the performance of traditional supervised learning network. In this
paper, we first conducted extensive experiments to compare the effectiveness of these four
inputs in the gesture recognition, respectively. Then we proposed a semi-supervised leaning
framework by utilizing few labeled data in the source domain and large amounts of
unlabeled data in the target domain. Specially, we combine the ∏-model and some specific
data augmentation tricks on the mmWave signal to realize the domain-independent gesture
recognition. Extensive experiments on a public mmWave gesture dataset demonstrate the
superior effectiveness of the proposed system.
Keywords: Gesture recognition; mmWave radar; Semi-supervised learning; ∏-model
Nanyu Jiang, Yuyuan Fang, Lei Zhang, Chao He, Zhenhua Wu,
IFM-PointNet++: Achieving efficient radar signal waveform recognition with instantaneous
frequency measurement,
Digital Signal Processing,
Volume 168, Part B,
2026,
105507,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Efficient and robust identification of radar signal waveforms is an essential task in
electronic reconnaissance. Current deep learning-based methods can obtain satisfying
accuracy, but they are usually with high computational burden. To address the issue, this
article develops an efficient algorithm IFM-PointNet++. This algorithm transfers the radar
signal waveform recognition into the point cloud recognition task by adopting the
instantaneous frequency measurement (IFM). By integrating IFM with PointNet++ network,
our method achieves superior efficiency and accuracy. To demonstrate this effectiveness, we
conduct comprehensive comparison experiments with YOLOv8 waveform recognition on the
time-frequency images. The results demonstrate that our proposed method significantly
accelerates signal waveform recognition while maintaining high accuracy.
Keywords: Radar signal waveform recognition; Signal recognition; PointNet++; Instantaneous
frequency measurement; Deep learning
Liheng Dong, Chengyang Tao, Zhaoxiang Zhang, Guiqing He, Yuelei Xu, Dong Liu,
An attitude-centric and cross-band infrared framework for aerial target intention
recognition,
Aerospace Science and Technology,
Volume 170,
2026,
111510,
ISSN 1270-9638,
[Link]
([Link]
Abstract: In terminal attack-defence scenarios, aerial target intention can be accurately
recognized based on trajectory and attitude information. Currently, radar-based intention
recognition methods hold a dominant position. However, radar is not adept at capturing
target attitude information (e.g., pitch, roll, and yaw) and can only perform intention
recognition based on trajectory information. To bridge this gap, this paper proposes an
attitude-centric and cross-band infrared aerial target intention recognition framework that
simultaneously extracts trajectory and attitude representations from multi-band infrared
images. The framework comprises two steps: pose estimation for predicting keypoints from
infrared images, followed by intention recognition based on these keypoints. For pose
estimation, a cross-band invariant representation learning method is proposed to reduce the
impact of band-bias, thereby improving the model’s generalization on multi-band infrared
images. For intention recognition, regularized adaptive adjacency matrix and parameter
fusion mechanisms are designed to effectively capture aerial target attitude representations,
forming an attitude-centric approach. Experiments demonstrate that the proposed method
significantly enhances the generalization of various pose estimation models on multi-band
infrared images. Additionally, with the introduction of attitude representations, the intention
recognition accuracy increases from 90.12 % to 96.64 %.
Keywords: Aerial target; Intention recognition; Pose estimation; Multi-band infrared;
Attitude
Hu Liu, Zhenghua Zhang, Jing Yang, Jörg Benndorf, Xiaofei Wang, Jiaqi Dong, Zitao Lin,
Guoliang Chen,
GhostPointNet: A deep learning-based method for ghost point noise detection in four-
dimensional (4D) millimeter-wave radar point clouds of underground mine,
Engineering Applications of Artificial Intelligence,
Volume 161, Part C,
2025,
112380,
ISSN 0952-1976,
[Link]
([Link]
Abstract: The high dust concentration, multi-metal supports, and narrow winding tunnels in
underground mines collectively lead to frequent ghost point noise in four-dimensional (4D)
millimeter-wave radar point clouds, posing serious challenges for mining perception and
localization. To address this, we propose a deep learning algorithm, named GhostPointNet,
for 4D millimeter-wave radar ghost point detection in underground mining environments.
From an artificial intelligence perspective, this model thoroughly considers the multi-modal
features of 4D millimeter-wave radar and the environmental complexity of underground
mines. It incorporates multi-parameterized spatial information inputs in both Cartesian and
Spherical coordinates, coupled with “Double T-Net” adaptive alignment correction, while
integrating non-spatial information such as radar power and Doppler data to achieve multi-
modal representation and end-to-end discrimination between ghost points and real points.
Experimental validation shows that GhostPointNet achieves excellent performance in
underground mining scenarios with 92.45 % accuracy and 95.84 % F1-score, outperforming
traditional filtering, clustering, and machine learning algorithms. From an engineering
application perspective, GhostPointNet is specifically designed for ghost noise detection in
underground mines. Even in complex scenarios such as mine tunnel intersections and turns,
it preserves critical structural points. Its end-to-end neural network simplifies post-
processing procedures, enhances operational efficiency, and provides stable and reliable
perceptual support for subsequent tasks such as autonomous mine locomotive navigation
and three-dimensional (3D) structure reconstruction. Experimental results demonstrate that
this method surpasses baseline approaches in ghost point detection, real point preservation,
and generalization capability, providing significant support for improving underground
mining safety and efficiency.
Keywords: Deep learning; Four-dimensional (4D) millimeter-wave radar; Ghost noise;
Underground mining; Point cloud segmentation
Liying Wang, Zongyong Cui, Yiming Pi, Changjie Cao, Zongjie Cao,
Low personality-sensitive feature learning for radar-based gesture recognition,
Neurocomputing,
Volume 493,
2022,
Pages 373-384,
ISSN 0925-2312,
[Link]
([Link]
Abstract: Radar-based sensing of gestures has gained tremendous attention with the recent
advancements in radar technologies. However, evident discrepancies exist in the gesture
samples due to hand flexibility and individual habits. It is challenging for traditional methods
to identify the gestures from unknown data sources. Cross-person (Cross-scenario)
recognition refers to a recognition where the training and test samples are from different
people (scenarios), respectively. To explore how the recognition performance is affected by
the individual habits, the reasons are analyzed and visualized through the experiments. On
this basis, HandNet is targeted proposed for the low personality-sensitive feature learning
and it has two main contributions. First, a Stepped Data Augmentation (SDA) is proposed to
reduce the sample interferences by non-coherent accumulating, and capture the inter-frame
dependencies. Second, a Focus on Generalization loss (FoG loss) is proposed to highlight the
generalized feature learning by res tricting the distances of inter-source features. Extensive
experiments demonstrate that HandNet effectively reduces the classifier’s sensitivity to the
personalized habits, and outperforms the existing state-of-the-art methods on the cross-
person and cross-scenario gesture recognition. To the best of our knowledge, it is the first
time to dedicate to addressing the radar-based gesture recognition with low personal
sensitivity, which is more suitable for practical scenarios.
Keywords: Convolutional neural network; Gesture recognition; Feature learning
Mingyang Du, Ping Zhong, Xiaohao Cai, Daping Bi, Aiqi Jing,
Robust Bayesian attention belief network for radar work mode recognition,
Digital Signal Processing,
Volume 133,
2023,
103874,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Understanding and analyzing radar work modes play a key role in electronic
support measure system. Many classifiers, for example those based on convolutional neural
network (CNN) and recurrent neural network (RNN), are available for recognizing radar work
modes as well as emitter types from their waveform parameters. However, the performance
of these methods may suffer significantly when confronting different types of signal
degradation, e.g., measurement error, lost pulse and spurious pulse. To tackle this issue, we
in this paper develop a Bayesian attention belief network (BABNet) based on Bayesian neural
networks in which the probability distribution over weights can help to enhance the model
robustness for corrupted data. In particular, we adopt pre-trained CNN as the Bayesian
inference prior. This not only accelerates the convergence speed, but also avoids the training
process getting stuck in bad local minima. Meanwhile, instead of using RNNs which are
difficult to be implemented in parallel, the combination of padding operation and attention
module in the proposed BABNet enables CNN, as the backbone, to process sequential data
with variable length. Extensive experiments are conducted to demonstrate the recognition
capability and robustness of the BABNet in different environments.
Keywords: Radar work mode; Pulse descriptor word; Attention mechanism; Bayesian neural
network; Robustness; Recognition
Van Ngoc Dang, Ngoc Chau Hoang, Quoc Cuong Nguyen, Minh Thuy Le,
Advancing robust human activity recognition via informative mmWave radar characteristics
and a lightweight spatio-spectro-temporal network,
Measurement,
Volume 256, Part A,
2025,
118056,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Human activity recognition (HAR) is increasingly important in aiding our daily life,
with millimeter-wave (mmWave) radar sensors emerging as a promising noninvasive solution
thanks to their excellent spatial and velocity resolution. Although existing radar-based
systems have shown strong performance, they primarily focus on micro-Doppler signatures
while neglecting angle information, which can hinder practical deployment in real-world
scenarios. Moreover, current state-of-the-art recognition models using mmWave radar often
require substantial computational resources, making integration into resource-constrained
devices challenging. This work proposes an efficient radar-based HAR system that leverages
angle and spectro-temporal information from micro-Doppler signatures. Our system utilizes
a multi-channel micro-Doppler representation corresponding to the number of virtual
antenna receivers as input. Then, a lightweight dilated convolutional network, namely SST-
DCN, extracts spatial-aware multi-scale spectro-temporal information through time-
frequency dilated convolutions. Experimental results on our real-world dataset demonstrate
the superiority of our approach compared to conventional features and other state-of-the-
art radar-based HAR systems.
Keywords: Human activity recognition; Millimeter-wave radar; Deep learning; Lightweight
network; Dilated convolution
Anand Dubey, Avik Santra, Jonas Fuchs, Maximilian Lübke, Robert Weigel, Fabian Lurz,
HARadNet: Anchor-free target detection for radar point clouds using hierarchical attention
and multi-task learning,
Machine Learning with Applications,
Volume 8,
2022,
100275,
ISSN 2666-8270,
[Link]
([Link]
Abstract: Target localization and classification from radar point clouds is a challenging task
due to the inherently sparse nature of the data with highly non-uniform target distribution.
This work presents HARadNet, a novel attention based anchor free target detection and
classification network architecture in a multi-task learning framework for radar point clouds
data. A direction field vector is used as motion modality to achieve attention inside the
network. The attention operates at different hierarchy of the feature abstraction layer with
each point sampled according to a conditional direction field vector, allowing the network to
exploit and learn a joint feature representation and correlation to its neighborhood. This
leads to a significant improvement in the performance of the classification. Additionally, a
parameter-free target localization is proposed using Bayesian sampling conditioned on a pre-
trained direction field vector. The extensive evaluation on a public radar dataset shows an
substantial increase in localization and classification performance.
Keywords: Multi-task learning; Radar detection; Scene understanding
Zilu Ying, Wenyu Ke, Yikui Zhai, Xinglin Liu, Jianhong Zhou, Pasquale Coscia, Angelo
Genovese,
Diffusion-augmented direct classification: A few-shot learning framework for Synthetic
Aperture Radar image automatic target recognition,
Engineering Applications of Artificial Intelligence,
Volume 166, Part B,
2026,
113648,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Deep learning-based (DL-based) synthetic aperture radar automatic target
recognition technology (SAR-ATR) has undergone extensive development, demonstrating
superiority over other competitive methods. However, the intrinsic requirement of deep
learning for a large labeled dataset restricts its practical application. Moreover, some DL-
based few-shot SAR-ATR methods are overly complex, hindering their deployment in real-
world applications. In addressing these obstacles, our solution introduces a straightforward
yet efficient few-shot learning approach titled Diffusion-Augmented Direct Classification for
few-shot SAR-ATR applications. The proposed method adopts a two-stage paradigm, where a
diffusion model first learns from unlabeled data and then produces synthetic samples to
train a recognition model. In the upstream stage, a lightweight diffusion-based image
generator build upon the shuffle-residual network structure is trained on a limited number
of annotated SAR images to generate artificial training samples for the downstream
recognition model. In the downstream stage, a Siamese network-based recognition model
and a similarity training procedure are proposed to train the model on a combination of real-
world and artificial samples, thereby improving recognition accuracy. A projection expansion
layer is proposed to improve the efficiency of cosine similarity loss in the downstream.
Experiments conducted on the Moving and Stationary Target Acquisition and Recognition
dataset demonstrated that our method outperformed other few-shot learning methods
concerning recognition accuracy in SAR-ATR tasks. Specifically, our method achieves over
73% accuracy in a 5-sample-per-class scenario and over 85% accuracy in a 10-samples-per-
class scenario. Source code of our paper is available at [Link]
Keywords: Synthetic Aperture Radar; Automatic Target Recognition; Denoising Diffusion
Probability model; Few-shot learning
Chuan Du, Long Tian, Bo Chen, Lei Zhang, Wenchao Chen, Hongwei Liu,
Region-factorized recurrent attentional network with deep clustering for radar HRRP target
recognition,
Signal Processing,
Volume 183,
2021,
108010,
ISSN 0165-1684,
[Link]
([Link]
Abstract: Feature extraction plays an essential role in radar automatic target recognition
(RATR) with high-resolution range profiles (HRRPs). Traditional feature extraction algorithms
usually ignore that different regions in HRRP contain the information with different
importance, resulting in their inadequacy in characterizing HRRP data. In this work, we
propose a region factorized recurrent attentional network (RFRAN) for HRRP-RATR by making
use of the temporal dependence through recurrent neural network (RNN) and automatically
finding the informative regions by a deep clustering mechanism in HRRP samples, which
reflects the distribution of scatterers in target along range dimension. Specifically, we
represent the temporal RNN hidden state using a region factorized encoder whose
parameters are conditioned on the HRRP region cluster centers. Moreover an attention
mechanism is used to weight up the different recognition contribution of each time step’s
hidden state. The aim of all the above modules is to achieve a more informative and
discriminative feature. Crucially, the loss function of RFRAN is differentiable, so all
components can be jointly trained with a gradient-based optimization. Compared with
traditional methods, besides the competitive recognition performance, RFRAN has a
promising interpretability thanks to the sequential region-specific hidden states.
Keywords: Region factorization; Clustering strategy; Attention mechanism; HRRP-RATR; RNN
Shuyu ZHENG, Libing JIANG, Qingwei YANG, Yingjian ZHAO, Zhuang WANG,
GS-orthogonalization OMP method for space target detection via bistatic space-based radar,
Chinese Journal of Aeronautics,
Volume 37, Issue 7,
2024,
Pages 333-351,
ISSN 1000-9361,
[Link]
([Link]
Abstract: A space-based bistatic radar system composed of two space-based radars as the
transmitter and the receiver respectively has a wider surveillance region and a better early
warning capability for high-speed targets, and it can detect focused space targets more
flexibly than the monostatic radar system or the ground-based radar system. However, the
target echo signal is more difficult to process due to the high-speed motion of both space-
based radars and space targets. To be specific, it will encounter the problems of Range Cell
Migration (RCM) and Doppler Frequency Migration (DFM), which degrade the long-time
coherent integration performance for target detection and localization inevitably. To solve
this problem, a novel target detection method based on an improved Gram Schmidt (GS)-
orthogonalization Orthogonal Matching Pursuit (OMP) algorithm is proposed in this paper.
First, the echo model for bistatic space-based radar is constructed and the conditions for
RCM and DFM are analyzed. Then, the proposed GS-orthogonalization OMP method is
applied to estimate the equivalent motion parameters of space targets. Thereafter, the RCM
and DFM are corrected by the compensation function correlated with the estimated motion
parameters. Finally, coherent integration can be achieved by performing the Fast Fourier
Transform (FFT) operation along the slow time direction on compensated echo signal.
Numerical simulations and real raw data results validate that the proposed GS-
orthogonalization OMP algorithm achieves better motion parameter estimation
performance and higher detection probability for space targets detection.
Keywords: Bistatic space-based radar; High-speed maneuvering space targets detection;
Range Cell Migration (RCM); Doppler Frequency Migration (DFM); Gram Schmidt (GS)-
orthogonalization Orthogonal Matching Pursuit (OMP) algorithm
Yonggang Qian, Yinghua Wang, Hongwei Liu, Zelong Wang, Feipeng Yu, Chunhui Qu,
MPRANet: Multi-scale perception and reference attention network for lightweight SAR target
recognition,
Neurocomputing,
Volume 668,
2026,
132310,
ISSN 0925-2312,
[Link]
([Link]
Abstract: Deep learning methods have been widely used in Synthetic Aperture Radar
Automatic Target Recognition (SAR ATR). However, challenges remain due to limited SAR
data and computational constraints on mobile devices, which hinder model training and
deployment. In this paper, we propose a Multi-scale Perception and Reference Attention
Network (MPRANet) for lightweight SAR ATR, which is a hybrid structure combining
convolutional networks and transformers, built upon the ShuffleNetV2 network. Specifically,
MPRANet introduces two key improvements compared to the CNN-based ShuffleNetV2.
Firstly, we replace the depthwise convolutions (DWConv) in the downsampling and basic
units of ShuffleNetV2 with the Multi-scale Parameter-Shared Convolution (MPConv) module.
MPConv enables the extraction of multi-scale features of SAR targets with almost no
additional parameters, thereby enhancing the network’s feature extraction capabilities.
Secondly, we propose a lightweight Reference Attention Transformer (RAformer) to capture
global information, addressing the issue of insufficient channel feature interaction in
ShuffleNetV2. In RAformer, a Local Linear Mapping Unit (LMU) is designed to perform linear
mappings, reducing the introduction of redundant features while ensuring its lightweight
and efficient nature. RAformer contains two modules: the Reference Vector Attention (RVA)
module, which efficiently models attention relationships, and the Lightweight Feedforward
Neural Network (LW-FFN) module, which enhances the network’s ability to capture
nonlinear representations. We evaluated the performance of MPRANet using publicly
available SAR datasets, including the MSTAR dataset, OpenSARShip dataset, and SAR-
AIRcraft-1.0 dataset. The experimental results demonstrate that MPRANet consistently
achieves superior recognition performance compared to other lightweight networks of
similar complexity.
Keywords: Synthetic aperture radar (SAR); Automatic target recognition (ATR); Convolutional
neural networks (CNN); Transformer; Lightweight
Xiaofang Pei, Yan Hu, Jun Zhu, Yun Dong, Peng Wang, Yinghua Ye, Ruiqi Shen,
Sustainable carbon-based materials for radar-infrared compatible stealth: Progress and
prospects,
Composites Part B: Engineering,
Volume 311,
2026,
113252,
ISSN 1359-8368,
[Link]
([Link]
Abstract: Stealth technology, as a cornerstone of modern defense and aerospace systems, is
increasingly challenged by multi-modal detection spanning radar, infrared, and emerging
sensor platforms. Carbon-based materials, owing to their lightweight nature, structural
tunability, and superior electromagnetic and thermal management capabilities, have
emerged as ideal candidates for radar-infrared compatible stealth applications. However,
traditional carbon sources demand energy-intensive processing and impose environmental
burdens that highlight the need for sustainable alternatives. This review provides a
comprehensive overview of recent advances in sustainable and environmentally friendly
carbon-based stealth materials, focusing on biomass (plant-, animal-, and microorganism-
derived), industrial by-products, and municipal or consumer wastes. Particular emphasis is
placed on the underlying mechanisms, highlighting that radar stealth originates from
dielectric and magnetic losses, while infrared stealth depends on temperature regulation
and emissivity control within the 3–5 μm and 8–14 μm atmospheric windows.
Microstructural engineering, heteroatom doping, and dielectric-magnetic synergy enable
broadband absorption at low filler loadings while providing tunable emissivity for infrared
suppression. Despite remarkable progress, key challenges remain in achieving simultaneous
broadband microwave absorption and low emissivity in critical infrared windows, as well as
ensuring structural uniformity and stability from complex renewable precursors. Looking
ahead, we propose a roadmap toward performance-sustainability-engineering synergy,
involving green precursor selection, scalable processing, multifunctional integration and
adaptive intelligent design. This Review thus bridges stealth performance with sustainability
imperatives, which provides strategic insights into the next generation of radar-infrared
compatible stealth materials.
Keywords: Carbon-based materials; Radar-infrared compatible stealth; Biomass; Waste
resources; Emissivity regulation
Nima Roshandel, Constantin Scholz, Hoang-Long Cao, Milan Amighi, Hamed Firouzipouyaei,
Aleksander Burkiewicz, Sebastien Menet, Felipe Ballen-Moreno, Dylan Warawout Sisavath,
Emil Imrith, Antonio Paolillo, Jan Genoe, Bram Vanderborght,
mmPrivPose3D: A dataset for pose estimation and gesture command recognition in human-
robot collaboration using frequency modulated continuous wave 60Hhz RaDAR,
Data in Brief,
Volume 59,
2025,
111316,
ISSN 2352-3409,
[Link]
([Link]
Abstract: 3D pose estimation and gesture command recognition are crucial for ensuring
safety and improving human-robot interaction. While RGB-D cameras are commonly used
for these tasks, they often raise privacy concerns due to their ability to capture detailed
visual data of human operators. In contrast, using RaDAR sensors offers a privacy-preserving
alternative, as they can output point-cloud data rather than images. We introduce
mmPrivPose3D, a dataset of 3D RaDAR point-cloud data that captures human movements
and gestures using a single IWR6843AOPEVM RaDAR sensor with a frequency of 10 Hz
synchronized with 19 corresponding 3D skeleton keypoints as the ground truth. These
keypoints were extracted from RGB-D images captured by an Intel RealSense camera
recorded at 30 frames per second using the Nuitrack SDK, and labeled with gestures. The
dataset was collected from n = 15 participants. Our dataset serves as a fundamental
resource for developing machine learning algorithms to improve the accuracy of pose
estimation and gesture recognition using RaDAR data.
Keywords: Human-robot collaboration; IWR6843AOPEVM; RaDAR; Pose estimation; Gesture
command recognition
Hua Wang, Qiangyu Zeng, Hao Wang, Jianxin He, Tiantian Yu, Guangpu Liu,
Temporal super-resolution reconstruction of weather radar echoes using a deep learning
approach,
Expert Systems with Applications,
Volume 300,
2026,
130189,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Severe convective weather events are characterised by rapid evolution and high
destructive potential, requiring weather radars to provide observations with high temporal
resolution. However, current S-band weather radar systems, constrained by their volumetric
scanning strategies, often fail to capture the rapidly changing features of these systems
promptly. To address this limitation, we propose EMAIRA-VFI, a deep learning–based
method for temporal super-resolution reconstruction of radar echoes, which enhances the
temporal resolution of radar data to meet the demands of severe convective weather
monitoring. By introducing an inter-frame attention mechanism, the proposed method
effectively fuses spatiotemporal features from sequential radar echoes, enabling accurate
modelling of dynamic weather evolution and the generation of continuous, high-temporal-
resolution radar echoes. Compared with conventional temporal interpolation methods,
EMAIRA-VFI demonstrates significant improvements in both interpolation accuracy and the
preservation of fine-scale meteorological structures. Experimental results show that the
model not only enhances the capability of S-band radars in monitoring rapidly evolving
weather events but also provides a new perspective for spatiotemporal fusion and the
intelligent application of radar data. We have open-sourced the code for this work at
[Link]
Keywords: Temporal super-resolution; Radar echo; Inter-frame attention mechanism
Xiaolin Zhu, Dongli Wang, Yan Zhou, Zixin Zhang, Jianxun Li, Rui Su, Yongcan Weng, Tao Zhu,
Deep learning-based group activity recognition in videos: A survey,
Neurocomputing,
Volume 661,
2026,
131150,
ISSN 0925-2312,
[Link]
([Link]
Abstract: Group activity recognition (GAR), which aims to identify the activity performed by
a group of people in a given video, is one of the representative tasks for video
understanding. With the advancement of deep neural networks, deep learning-based
methods have emerged as a dominant alternative to handcrafted feature engineering, which
can automatically mine the feature representations from the input data. In this survey, we
present a comprehensive review with in-depth analysis of deep learning-based group
activity recognition from 2016 to 2024. Specifically, we first briefly outline the definition and
several major challenges of group activity recognition. Then, a detailed taxonomy is
introduced in terms of different supervision types, network types, modeling mechanism
types, and input types, which can better classify the existing state-of-the-art methods from
different perspectives. To the best of our knowledge, we are the first to develop such a
taxonomy for GAR. For a better understanding of the pros and cons of each type, the
detailed discussions under each type are also presented with further categorization.
Moreover, we also provide the datasets, evaluation metrics, and performance comparisons
for group activity recognition. In the end, we conclude the survey by suggesting future
potential research directions in this rapidly growing GAR field to facilitate new research
ideas.
Keywords: Group activity recognition (GAR); Deep learning; Modeling mechanism; Self-
supervised learning; Multimodal fusion; Graph convolutional network (GCN); Transformer
Yifan Chen, Meng Jiang, Chao Xia, Hang Zhao, Panpan Ke, Sheng Chen, Heng Ge, Keran Li, Xu
Wang, Yufei Wang, Yezi Chai, Qiming Liu, Zhengyu Tao, Yuyan Lyu, Yani Wu, Ao Shi, Yang Liu,
Hongyi Xin, Yu Zhong, Wei Zhang, Fuhua Yan, Weiwei Quan, Yingjia Xu, Dan Liu, Yumin Sun,
Xinli Li, Yuanyuan Tian, Lianming Wu, Shengxian Tu, Hongwei Ji, Bin Sheng, Jun Pu,
A novel deep learning system for STEMI prognostic prediction from multi-sequence cardiac
magnetic resonance,
Science Bulletin,
Volume 70, Issue 24,
2025,
Pages 4241-4252,
ISSN 2095-9273,
[Link]
([Link]
Abstract: ST-elevation myocardial infarction (STEMI) remains a leading cause of
cardiovascular morbidity and mortality worldwide, and accurate early risk stratification is
critical for implementing precision therapies in clinical practice. However, existing clinical risk
scores and manually derived imaging biomarkers have limited accuracy in predicting post-
STEMI outcomes. To address this gap, we developed DeepSTEMI, an end-to-end deep
learning system that integrates multi-sequence cardiac magnetic resonance (CMR) images
with clinical parameters for predicting 2-year major adverse cardiovascular events (MACE).
The system comprised two key algorithmic modules: a U-Net module that automatically
segments heart regions from raw CMR images and a Transformer-based module that
predicted future cardiovascular events. DeepSTEMI was developed using a multicenter
dataset (n = 610; 20,618 images) from STEMI patients enrolled in the EARLY-MYO-CMR
registry (NCT03768453), with external validation performed in 334 patients (9944 images)
from three independent cardiac centers. In external validation, DeepSTEMI demonstrated
superior predictive performance compared to conventional clinical risk scores and manual
CMR parameters (AUC 0.894, 95% CI: 0.823–0.965; overall accuracy 94.3%). The model
identified high-risk patients who exhibited a 20-fold MACE risk compared to low-risk
counterparts (HR 20.43, log-rank P < 0.001). SHapley Additive exPlanations (SHAP) analysis
revealed that DeepSTEMI’s predictive power stems from clinical-imaging synergy, enabling it
to capture complex pathological patterns. DeepSTEMI achieved consistently superior
performance over the Eitel score across all subgroups, with the greatest benefit observed in
women (NRI 1.597) and in patients imaged 4–7 d post-STEMI (NRI 1.442). Overall,
DeepSTEMI serves as an automated, scalable, and interpretable clinical copilot, which
advances post-STEMI risk stratification beyond the limitations of current paradigms.
Keywords: Myocardial infarction; Deep learning; Transformer; Prognostic prediction; Cardiac
magnetic resonance
Mingjun Cheng, Hong Jin, Qinfeng Zhao, Yurun Wang, Yanxi Wu, Shan Huang, Wenze Yue,
Deep learning for optimizing urban governance by "sensing-processing-responding" cycle:
Recent advances, future prospects and challenges,
Sustainable Cities and Society,
Volume 135,
2025,
106994,
ISSN 2210-6707,
[Link]
([Link]
Abstract: With accelerating urbanization, traditional governance models are increasingly
strained. Deep learning (DL) offers powerful solutions, but its application in urban
governance lacks a systematic framework and faces significant hurdles. This paper addresses
these gaps through a systematic review of 329 articles published from 2016 to 2025. We
introduce a novel Sensing-Processing-Responding framework to classify the technological
pathways of DL in urban governance. This framework organizes applications into three core
stages: (1) Sensing technologies (e.g., CNNs) for dynamic data acquisition; (2) Processing
technologies (e.g., RNNs, Transformers) for predictive modeling and analysis; and (3)
Responding technologies (e.g., LLMs) for automated decision support. Our analysis reveals
that while DL is widely applied in traffic forecasting, environmental monitoring, and disaster
response, its deployment is constrained by key challenges. We found a significant gap
between research and practice, with only 7.6% of studies demonstrating real-world
application. Furthermore, it concerns data privacy and model interpretability limit public
acceptance, although our review indicates that fewer than 10% of studies involve high-risk
personal data. Future progress depends on integrating emerging technologies like
multimodal large models and multi-agent systems. We conclude by advocating for a
paradigm shift from focusing purely on accuracy to prioritizing public value, fairness, and
transparency. This study provides a comprehensive roadmap for developing more intelligent,
resilient, and sustainable urban governance systems.
Keywords: Deep learning; Urban governance; Large models; Sensing-Processing-Responding
framework; Sustainable development; Literature review
Haiyan Yao, Yuefei Xu, Qiang Guo, Shizhe Chen, Bin Lu, Yuanjun Huang,
Study on transformer fault diagnosisbased on improved deep residual shrinkage network
and optimized residual variational autoencoder,
Energy Reports,
Volume 13,
2025,
Pages 1608-1619,
ISSN 2352-4847,
[Link]
([Link]
Abstract: The transformer as the core equipment in the power system, its fault diagnosis has
a vital role in ensuring the safe and stable operation of the power grid. However, traditional
transformer fault diagnosis methods often rely on manual experience or simple models,
which are difficult to meet the demand for efficient and accurate diagnosis when faced with
complex and evolving fault patterns. In this study, a new method for transformer fault
diagnosis based on improved deep residual shrinkage network (DRSN) and optimized
residual variational autoencoders (ORVAE) is proposed. Firstly, this study improves the DRSN
to enhance its feature extraction capability. By designing a specific shrinkage mechanism,
the improved DRSN can reduce the information loss in the face of complex data, greatly
improve the extraction ability of the key features of the transformer operating state, and
thus improve the accuracy of fault recognition. Secondly, in view of the difficulty and high
cost of transformer fault sample data collection, this study introduces a residual connection
structure based on the traditional variational autoencoder (VAE), and constructs the ORVAE
method to effectively address the challenge of insufficient data. The results show that the
fault recognition rate of the proposed method on the real transformer fault dataset reaches
97.14 %, which is better than the traditional method, showing excellent diagnostic
performance and strong practical application potential. Compared with the existing
technologies, this method not only improves the accuracy of transformer fault diagnosis, but
also provides new ideas and technical support for the intelligent development of power
system. This study offers an innovative solution for the field of fault diagnosis of power
equipment, and providing a strong technical guarantee for fault prediction and maintenance
in future smart grids.
Keywords: Transformer; Fault diagnosis; Improved DRSN; Shrinkage mechanism; Feature
extraction; ORVAE; Recognition rate
Qinzhong Hou, Yonghao Yang, Jiatong Liang, Xiaoyan Huo, Junqiang Leng,
A deep transfer learning approach for Real-Time traffic conflict prediction with trajectory
data,
Accident Analysis & Prevention,
Volume 214,
2025,
107966,
ISSN 0001-4575,
[Link]
([Link]
Abstract: Recently, real-time traffic conflict prediction has drawn increasing attention due to
its significant potential in proactive traffic safety systems. While various statistical and
machine learning models have been developed for conflict prediction, transferability
remains a fundamental issue across these models. Specifically, the predictive performance of
a real-time conflict prediction model developed for a specific location can significantly
decline when directly applied to a new location without any modifications, primarily due to
substantial differences in traffic environments between these areas. To address this gap, this
study proposed a novel deep transfer learning approach aimed at enhancing the
transferability of real-time conflict prediction models. Initially, a real-time conflict prediction
framework was designed utilizing trajectory data for merging areas with consideration of
temporal variations in traffic flow characteristics. Subsequently, the Gated-Transformer, Fully
Convolutional Networks (FCN), Long Short-Term Memory Fully Convolutional Networks
(LSTM-FCN), and Multivariate Long Short-Term Memory Fully Convolutional Networks
(MLSTM-FCN) were employed as backbone feature extraction networks to capture the
hidden correlations between time-varying traffic flow characteristics and traffic conflicts.
After that, an independent transfer learning architecture was established to assess the
similarity of the distribution of traffic flow characteristics at different locations, based on the
maximum mean discrepancy criteria. For empirical evaluation, merging areas from the exiD
dataset were differentiated into source and target domains. The results demonstrated that
the Gated-Transformer model outperforms other baseline models (FCN, LSTM–FCN and
MLSTM–FCN) in both feature extraction and predictive performance, achieving an F1 score
of 0.864 and an area under the curve (AUC) of 0.980. Furthermore, the transfer learning
architecture can substantially enhance the predictive performance of a model trained in the
source domain when applied to the target domain. In particular, the F1 score and AUC for
the Gated-Transformer model improved by 11.9% and 10.2%, respectively, after
incorporating the transfer learning architecture. Finally, the optimal values of key model
parameters, including the sliding time window (6 s) and the prewarning time (5 s), were
recommended for practical applications through sensitivity analysis. This study illustrates the
potential of the deep transfer learning approach as a reliable and effective alternative to
improve the transferability of real-time conflict prediction models. Additionally, results from
this study can offer valuable insights for practical applications in traffic safety warning
systems, particularly in vehicle-to-infrastructure traffic environments.
Keywords: Real-time conflict prediction; Deep transfer learning; Gated-Transformer; Merging
area; Trajectory data
Sidra Ghayour Bhatti, Imtiaz Ahmad Taj, Mohsin Ullah, Aamer Iqbal Bhatti,
Transformer-based models for intrapulse modulation recognition of radar waveforms,
Engineering Applications of Artificial Intelligence,
Volume 136, Part B,
2024,
108989,
ISSN 0952-1976,
[Link]
([Link]
Abstract: The increasing prevalence of low probability of intercept (LPI) radars in electronic
warfare (EW) systems highlights the need to effectively recognize phase-coded radar
waveforms intercepted at radar warning receivers (RWRs) from various threat emitters. The
complexities of the electromagnetic (EM) spectrum necessitate the implementation of an
automatic modulation recognition system (AMRS) within the RWR. However, a major
challenge is accurately identifying phase-coded waveforms with high accuracy at low signal-
to-noise ratios (SNRs). This research addresses the challenge by exploring three artificial
intelligence (AI)-driven AMRS architectures for identifying phase-coded waveforms using
short-time Fourier transform (STFT): vision transformer (ViT), vicinity vision transformer
(VViT), and deep convolutional neural network (DCNN). Unlike recent methods focusing on
amplitude spectra, our research delves into the phase spectra for the feature extraction of
phase-coded waveforms. We leverage phase-based features extracted from intercepted
phase-coded waveforms to classify six types of phase-coded signals using these AMRS
architectures across SNR levels ranging from −16 dB to 8 dB. The simulation experiments
show that these methods are effective at an SNR of −16 dB, with VViT and ViT achieving
recognition accuracies of 93% and 92.7%, respectively. Both outperform the DCNN, which
achieves an RA of 89% at the same SNR. This approach promises to enhance situational
awareness and decision-making in EW operations by improving phase-coded radar
waveform recognition and enabling appropriate countermeasure deployment.
Keywords: Automatic modulation recognition system; Feature extraction; Low probability of
intercept; Short time Fourier transform
Dawei Li, Jingnan Wang, Kefeng Deng, Di Zhang, Chengwu Zhao, Hongze Leng, Yingfang Wen,
Yudi Liu, Kaijun Ren, Junqiang Song,
Review on deep learning quantitative precipitation nowcasting: Advances and challenges,
Expert Systems with Applications,
Volume 305,
2026,
130775,
ISSN 0957-4174,
[Link]
([Link]
Abstract: In recent decades, extreme precipitation events have increased dramatically due to
global warming, resulting in significant casualties and economic losses. Quantitative
Precipitation Nowcasting (QPN), which predicts precipitation intensity within a six-hour
timeframe, plays a critical role in public safety, infrastructure protection, transportation
management, outdoor event planning, and flood prevention systems. Building on successful
innovations in computer vision, deep learning advancements have substantially improved
prediction accuracy and transformed QPN methodologies. However, despite the
proliferation of research in this rapidly advancing field, comprehensive surveys that
systematically examine mainstream techniques and identify key challenges remain limited.
This paper provides a thorough review of current deep learning approaches in precipitation
nowcasting, examining important challenges, analyzing methodologies across the QPN
development lifecycle, and exploring promising research directions. Through systematic
synthesis of emerging developments, we aim to foster interdisciplinary collaboration and
stimulate continued innovation in this essential field.
Keywords: Quantitative precipitation nowcasting; Computer vision; Multidisciplinary
cooperation
Shenghua Lv, Xiaowei Zhang, Xuan Zhao, Meng Li, Jianghao Zhang, Chen Lin, Jian Wen,
Rapid and accurate assessment of filed scale soil moisture using ground-penetrating radar
deep learning-based inversion,
Measurement,
2026,
120594,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Accurate quantification of soil moisture content (SMC) is essential for sustaining
plant growth and maintaining ecosystem stability. Current SMC monitoring approaches are
subject to several constraints: satellite-based remote sensing frequently suffers from
inadequate spatial resolution for large-scale precision, while point-scale techniques are
incapable of effectively capturing soil moisture spatiotemporal variations and are unsuitable
for large-area monitoring. Ground Penetrating Radar (GPR), as an efficient and non-
destructive subsurface detection technique, has been widely employed for estimating soil
water content. However, existing methods based on full-waveform inversion exhibit strong
dependency and involve computationally expensive processes, resulting in insufficient
efficiency when handling large-scale GPR data. To overcome these challenges, this study
proposes a GPR deep learning-based inversion framework for rapid and precise SMC
estimation. A 900 MHz GPR system was deployed to survey an experimental site equipped
with pre-installed moisture sensors. The results demonstrated strong agreement between
GPR-derived SMC values and measurements obtained via Time Domain Reflectometry (TDR).
Furthermore, continuous high-temporal-resolution data collection verified the capability of
the method to characterize the spatiotemporal dynamics of SMC. Notably, for regional soil
moisture content (RSMC) estimation, the proposed method achieved a substantially lower
error (0.0019 m3) compared to conventional point-based measurements (0.0146 m3) and
stratified estimation (0.0122 m3). This methodology provides a robust technical foundation
for accurate field-scale SMC monitoring and exhibits significant potential for use in ecological
surveillance, precision agriculture irrigation, and sustainable water resource management.
Keywords: Ground-penetrating radar; Non-destructive testing; Soil moisture content; Deep
learning-based inversion; Spatio-temporal evolution
Bo Da, Xianyou Chang, Lilin Zhu, Hao Huang, Jiayang Ma, Weili Song, Da Chen,
Flexural behavior prediction of reinforced seawater sea sand concrete beams based on deep
learning techniques,
Construction and Building Materials,
Volume 495,
2025,
143576,
ISSN 0950-0618,
[Link]
([Link]
Abstract: To accurately assess the flexural behavior of reinforced seawater sea sand concrete
beams (RSSSCB), the influence of input features was investigated using an experimental
dataset with 163 ultimate bending moment (Mu) sets. For optimal model selection, deep
learning algorithms (CNN, Transformer, LSTM) were compared with traditional machine
learning methods (XGBoost, Random Forest) via comprehensive metrics: mean absolute
error, root mean square error, and coefficient of determination. This integrated methodology
enabled precise characterization of flexural behavior and quantitative interpretation of
feature importance correlations. The results indicate that: The CNN exhibited superior
predictive performance on both training and testing datasets compared to Transformer,
LSTM, XGBoost, and Random Forest algorithms, increased by 9.2 % and 29.33 %, 13.1 % and
11.49 %, −2.1 % and 16.87 %, 0 % and 24.36 %, respectively. This finding indicates that the
CNN model demonstrates a notable advantage in terms of prediction accuracy and
generalization capability. Furthermore, relative to the empirical formulations in GB 50010–
2010, JGJ 12–2006, and Da et al., the CNN model exhibited accuracy enhancements of 21 %,
25 %, and 19 %, respectively. Through the interpretation of the CNN model, it is found that
the cross-sectional area of longitudinal reinforcement in the tension zone and height of
concrete in the compression zone have a significant influence on its predictive performance.
Additionally, a noteworthy positive correlation was observed between rectangular cross-
section height and concrete height in the compression zone.
Keywords: Reinforced seawater sea sand concrete beams; Flexural behavior; Deep learning;
CNN; Predictive modeling
Kai Zhao, Zhongqi Sun, Hao Jiang, Zhixuan Zou, Qiong Wu, Yanjie Xin, Xiangru Liu, Huijie
Jiang,
Transformer-based integration of radiomics and deep learning for differentiating lipid-poor
adrenal adenomas from malignant tumors,
Meta-Radiology,
Volume 3, Issue 4,
2025,
100183,
ISSN 2950-1628,
[Link]
([Link]
Abstract: Purpose
To evaluate the effectiveness of a Transformer model based on contrast-enhanced computed
tomography (CECT) that integrates radiomics and deep learning features in differentiating
adrenal lipid-poor adenomas (LPA) and malignant tumors (MT).
Methods
This retrospective study included 282 patients with adrenal tumors from two medical
centers between October 2018 and October 2024. The patients were classified into adrenal
(LPA) and adrenal (MT) groups. Radiomics and deep learning features were extracted from
CECT images. A total of 240 patients from the first center were randomly divided into
Training Set and Test Set at a 7:3 ratio, while 42 patients from the second center served as an
External Validation Set. A Transformer algorithm was employed to integrate radiomics and
deep learning features for building predictive models. Its self-attention mechanism was
utilized to capture intrinsic associations within each feature type and to uncover hidden
information related to clinical outcomes. Additionally, a Radiomics model, a Deep Learning
model (DL_model), and a Traditional Combined model integrating radiomics and deep
learning features were constructed. Model performance was assessed using the area under
the receiver operating characteristic (ROC) curve (AUC) and radar chart. Calibration curves
and decision curve analysis (DCA) were employed to assess the predictive accuracy and
clinical net benefit of the models. Furthermore, radiomics feature activation maps and
gradient-weighted class activation mapping (Grad-CAM) were utilized to visualize radiomics
and deep learning features, respectively.
Results
The Transformer model achieved the best predictive performance in the training, test, and
external validation sets, with AUCs of 0.949, 0.917, and 0.852, respectively. The DeLong test
indicated that the performance differences between this model and the other models were
statistically significant. Furthermore, the radar chart illustrated that the Transformer model
achieved superior overall performance, and DCA confirmed its higher clinical net benefit
compared with the other models.
Conclusion
The Transformer model that integrates radiomics and deep learning features can accurately
distinguish between LPA and MT. Furthermore, the visual analysis of radiomics feature
activation maps and Grad-CAM intuitively illustrates the distribution of radiomics and deep
learning features, enhancing their potential for clinical application in preoperative
assessment of adrenal tumors.
Keywords: Lipid-poor adrenal adenomas; Computed tomography; Radiomics; Deep learning;
Transformer
Gabriela Czibula, Andrei Mihai, Paul-Dumitru Orăşan, Istvan Gergely Czibula, Eugen Mihuleţ,
Sorin Burcea,
SepConv-ens: An ensemble of separable convolution-based deep learning models for
weather radar echo temporal extrapolation,
Procedia Computer Science,
Volume 246,
2024,
Pages 666-675,
ISSN 1877-0509,
[Link]
([Link]
Abstract: The paper addresses the topic of radar echo temporal extrapolation which is of
major interest in both operational and research meteorology. Weather radar measurements
are an important data source used by operational meteorologists for weather analysis, radar
refectivity having a significant influence on short-term heavy rainfall prediction. Thus,
extrapolating radar products’ values is important for early storm evolution assessment. The
paper proposes SepConv-ens approach for temporal extrapolation of radar observations
using an ensemble of three separable convolution-based deep learning models. Experiments
performed on real radar data from the Romanian National Meteorological Administration
(NMA) highlight a good performance of SepConv-ens in predicting radar data up to more
than 40 minutes ahead and a good correlation between the radar measurements and the
predictions in terms of spatial and intensity evolution of the radar echoes. SepConv-ens is
integrated in the operational visualization software utilized by the Romanian NMA and is the
first attempt, at the national level, to offer an artificial intelligence-based automated
assistance for operational meteorologists.
Keywords: deep learning; convolutional neural network; separable convolution; nowcasting;
weather radar 2000 MSC: 68T07; 68T10
Kundan Meshram, Aryan Saurabh, Vinay Kharole, Chatrabhuj, Umank Mishra, Kennedy C.
Onyelowe, Viroon Kamchoom, Krishna Prakash Arunachalam,
Design of an integrated model for pothole detection and repair optimization using
multimodal transformers and hybrid deep learning,
Case Studies in Construction Materials,
Volume 23,
2025,
e05431,
ISSN 2214-5095,
[Link]
([Link]
Abstract: The detection and timely repair of potholes are crucial for maintaining road safety
and minimizing vehicle damage. However, existing methods often suffer from limitations
such as reliance on single-modal data, poor generalization across diverse environments, and
suboptimal resource management. To address these challenges, we propose a
comprehensive framework for enhanced pothole detection and repair optimization using
advanced deep learning techniques. Our approach integrates four key methodologies:
Multimodal Enhanced Pothole Detection with Person-Level Data (M-E-Pot holeNet), Hybrid
Machine Learning-Deep Learning for Classification (Hybrid-Pot holeNet), Deep
Reinforcement Learning for Pot hole Detection and Repair Optimization (DRL-Pot holeOpt),
and Transfer Learning for Pothole Detection in Diverse Environments (TL-Pot holeAdaptNet).
M-E-Pot holeNet employs a Self-Supervised Multimodal Transformer (SSMT) to fuse camera,
accelerometer, and crowdsourced smartphone data, achieving robust detection with a 97 %
accuracy and under 2 % false positive rate. Hybrid-Pot holeNet combines Graph Attention
Networks (GAT) and XGBoost, modeling spatial road features to classify potholes with 95 %
accuracy and an F1-Score of 0.92. DRL-Pot holeOpt uses Soft Actor-Critic (SAC) with Bayesian
Optimization to efficiently schedule repair tasks, reducing repair costs by up to 20 % and
crew travel time by 15–25 %. Finally, TL-Pot holeAdaptNet leverages Domain-Adversarial
Neural Networks (DANN) to ensure cross-domain adaptability, with 90 % accuracy in new
environments and a 40–50 % reduction in domain discrepancy. This multi-faceted approach
addresses the limitations of previous work by providing scalable, real-time, and resource-
optimized solutions for pothole detection and maintenance, offering significant
improvements in accuracy, cost efficiency, and adaptability.
Keywords: Pothole detection; Multimodal data; Hybrid deep learning; Reinforcement
learning; Transfer learning; Scenarios
Jiaquan Wan, Junchao Wang, Wei Zhang, Hao Song, Congyi Nai, Fengchang Xue, Tao Yang,
Chunxiang Shi, Quan J. Wang, Baoxiang Pan,
RadarDiT: An advanced radar echo extrapolation model for three gorges reservoir area via
diffusion transformer,
Journal of Hydrology: Regional Studies,
Volume 61,
2025,
102703,
ISSN 2214-5818,
[Link]
([Link]
Abstract: Study region
The Three Gorges Reservoir Area (TGRA)
Study focus
TGRA faces increasing vulnerability to extreme precipitation events driven by complex
convective weather systems. Radar echo extrapolation—predicting future precipitation
patterns from current radar data—is essential for early warning systems but faces significant
challenges in this topographically complex region. While data-driven approaches have
advanced the field, current convolutional neural network-based diffusion models struggle
with the TGRA's dynamic meteorological conditions due to their reliance on translational
invariance, which often fails to capture rapid weather transitions in complex terrain.
New hydrogeological insights from the region
To address these limitations, we introduce RadarDiT, a Vision Transformer-based diffusion
model specifically engineered for radar extrapolation in the TGRA. First, we develop a five-
year radar dataset capturing diverse convective weather phenomena unique to this region.
Then, leveraging this dataset, RadarDiT employs multi-layer Vision Transformers that
effectively model global dependencies and complex spatial relationships, enabling accurate
prediction of convective cell evolution. Our model demonstrates superior performance in
maintaining strong echo and spatial coherence over longer forecast horizons. Quantitative
evaluations across multiple metrics and thresholds confirm RadarDiT's enhanced skill in
forecasting heavy precipitation events, with particular improvements in Critical Success
Index at higher radar echo values. This work establishes a foundation for more reliable
nowcasting systems in regions with complex terrain and dynamic weather patterns, directly
supporting enhanced disaster preparedness and response strategies.
Keywords: Radar Echo Extrapolation; Three Gorges Reservoir Area; Diffusion Model; Vision
Transformer; Nowcasting
Xiaole Han, Jintao Liu, Jian Ye, Zihe Wang, Pengfei Wu, Hai Yang,
Deep Learning-Based GPR interpretation of soil thickness in headwater hillslopes,
Geoderma,
Volume 462,
2025,
117530,
ISSN 0016-7061,
[Link]
([Link]
Abstract: Soil thickness strongly influences eco-hydrological and geomorphic processes, yet
conventional measurements such as auger drilling are invasive, labor-intensive, and
unsuitable for large-scale surveys. Ground-penetrating radar (GPR) provides a non-invasive
alternative, but its manual interpretation remains slow and prone to observer bias. To
address this challenge, we developed a fully automated framework that couples a hybrid
CNN-Transformer deep learning architecture with optimized signal filtering to predict soil
thickness directly from GPR profiles. The convolutional layers extract local waveform
features, while the attention mechanism captures long-range dependencies. Using field data
from a steep headwater hillslope (H1) in the Taihu Basin, China, we compared five filtering
strategies—median, Savitzky-Golay, Gaussian, moving average, and none—and found that
median filtering yielded the most accurate results (R2 up to 0.92, CCC of 0.96, RMSE near
10 cm). We further identified optimal filter window sizes (61–101 samples) and a training
duration threshold (≥500 epochs) that ensured stable and accurate predictions. Cross-site
validation on an independent hillslope (H2) without retraining showed that the pretrained
CNN-Transformer model achieved the highest R2 (0.80), CCC (0.89), and lowest RMSE
(11.3 cm), outperforming traditional machine learning models (CNN, MLP, RF, SVM) in
transferability. These findings demonstrate that integrating CNN-Transformer architectures
with appropriate signal filtering enables scalable, accurate, and objective soil thickness
mapping in complex terrain. The proposed approach also holds promise for broader GPR-
based subsurface applications, including soil horizon delineation and root system detection.
Keywords: Ground-penetrating radar; Soil thickness; Transformer; Headwater hillslopes;
Median filtering
Lixing Shi, Xueling Liang, Wenchao Chen, Yaoqiang Liu, Tong Ding, Kun Qin, Bo Chen,
Hongwei Liu,
Masked variational transformer for complex clutter modeling and target detection,
Signal Processing,
Volume 239,
2026,
110236,
ISSN 0165-1684,
[Link]
([Link]
Abstract: Weak target detection commonly encounters intense clutter interference, which
overshadows weak signals and complicates the task. Taking advantage of the powerful data
mining capability of neural networks, more and more deep learning-based methods are
applied to radar target detection. Among the approaches, those founded upon unsupervised
learning methodologies exhibit remarkable merit because they dispense with the
requirement for target samples within the training step, making them highly applicable in
practical target detecting scenarios. However, existing methods suffer from limitations in
leveraging the range-Doppler (R-D) two-dimensional correlation and finely modeling in
multiple clutter scenarios. In this paper, an unsupervised Transformer-based detector (TrDet)
is proposed to break through the boundary of modeling capability. First, with the designed
two-dimensional position embedding (2-DPE) and global query embedding (GQE)
techniques, an unsupervised training strategy for R-D spectrum based on Transformer
framework is utilized to achieve refined clutter modeling. Then, radar target detection is
formulated as an out-of-distribution (OOD) detection task to mitigate clutter interference.
Moreover, the masked variational Transformer-based detector (MVTrDet) is further
proposed to prevent target information leakage when the target is in close proximity to the
clutter in Doppler domain. Compared with several relative algorithms, our proposed
methods are better suited for radar target detection in complex clutter environments. The
experimental results derived from both measured data and simulated data verify the
effectiveness of our proposed methods.
Keywords: Radar target detection; Clutter modeling; Range-Doppler (R-D) spectrum;
Unsupervised learning; Out-of-distribution detection; Transformer
Binyu Xiong, Yuntian Chen, Dali Chen, Jun Fu, Dongxiao Zhang,
Deep probabilistic solar power forecasting with Transformer and Gaussian process
approximation,
Applied Energy,
Volume 382,
2025,
125294,
ISSN 0306-2619,
[Link]
([Link]
Abstract: Solar power generation encounters instability and unpredictability issues due to
the uncertainty of weather changes. Consequently, probabilistic forecasting of solar power is
essential for the effective management and integration of solar energy into the power grid,
substantially enhancing the reliability and efficiency of the electrical system. Among various
methods, time series analysis for probabilistic forecasting, which leverages historical data to
predict future solar power generation, has become a significant area of research due to
advancements in deep learning. However, existing methods often fall short in accuracy and
operational efficiency. This paper introduces an innovative deep learning framework tailored
for probabilistic forecasting of solar power generation. Considering the unique distribution
characteristics of solar power data, a novel data preprocessing method integrating Box–Cox
and Z-score transformations is applied to the input time series data. Subsequently, a novel
probabilistic time series forecasting method, leveraging a Transformer network enhanced
with Gaussian process approximation, predicts solar power generation for the forthcoming
24 h. The delta method is then employed to reverse transform the forecasts into actual
predicted values. Comparative analyses using a real-world solar power dataset demonstrate
that the proposed model outperforms existing probabilistic forecasting networks in
deterministic, probabilistic, and interval forecasting tasks. Compared to the commonly used
probabilistic forecasting method MC Dropout, our method decreases the CRPS index by
22.6% on the Shenzhen dataset and 39.7% on the Xingtai dataset. Furthermore, the
proposed model exhibits superior computational efficiency, reflecting an optimal balance
between accuracy and computational demands.
Keywords: Probabilistic forecasting; Solar power; Transformer network; Gaussian process
approximation
Xiaofang Sun, Meng Wang, Junbang Wang, Guicai Li, Xuehui Hou,
Deep learning classification of winter wheat from Sentinel optical-radar image time series in
smallholder farming areas,
Advances in Space Research,
Volume 75, Issue 3,
2025,
Pages 2683-2695,
ISSN 0273-1177,
[Link]
([Link]
Abstract: As crop yield stagnation, climate change, and the rising demand for agricultural
products pose increasing challenges, mapping crop systems is becoming more and more
important. Winter wheat is one of the major cereal crops cultivated in China, ranking as the
third largest crop in terms of production and harvested area. Accurately mapping winter
wheat is necessary for implementing effective farm management practices. While many
studies have successfully produced high spatiotemporal resolution land cover maps,
relatively few map products of crop types are available in China. The growing archive of
satellite image time series provides enormous opportunities to map crops more closely. This
research presents a two-step method to map winter wheat based on Sentinel-1 and
Sentinel-2 time-series data from Shandong Province using the deep learning approaches.
The winter crops were firstly mapped using time-series optical vegetation indices employing
the deep learning methods. Then winter wheat was extracted from the winter crops mask by
coupling optical and synthetic aperture radar time-series images. The results indicated that
the precision of mapping winter wheat using Temporal Convolution Neural Networks
(TempCNN) achieved the highest precision in mapping winter wheat, with an overall
accuracy of 93.7 %, a kappa coefficient of 0.907, and an F1-score of 0.989. This was followed
sequentially by the Residual 1D convolutional neural networks (ResNet), the Multi-Layer
Perceptron (MLP), and the Lightweight Temporal Self-Attention Encoder (L-TAE). The
Temporal Attention Encoder (TAE) model demonstrated the lowest precision among the
compared models. The results agree well with independent county-level official census
winter wheat area data (R2 = 0.936). The proposed framework can also be applied in other
regions to generate maps of different crops, so future work can extend the proposed model
to other agricultural regions, where an increased number of crop types and natural
vegetation types can be included and tested.
Keywords: Sentinel-1; Sentinel-2; Classification; Winter wheat mapping; Time series; Deep
learning
Wenyu Wang, Chenyang Wang, Libo Zhang, Yuchen Yan, Linxiu Wang, Jin Guo,
Monitoring mining-induced subsidence from satellite imagery using transformer-based deep
learning trained on gridded subsidence measurements,
Journal of Environmental Management,
Volume 394,
2025,
127536,
ISSN 0301-4797,
[Link]
([Link]
Abstract: The inherent concealment of underground coal mining makes it difficult for
environmental protection authorities to detect and regulate illicit activities. These mining
activities are only identified after severe environmental damage has occurred, such as
farmland flooding or structural cracks in residential buildings. By the time enforcement
actions are taken, the opportunity for early intervention is lost, and ecological restoration
becomes nearly impossible. Using artificial intelligence (AI) to analyse satellite imagery for
monitoring land subsidence in coal mining-affected areas is considered a promising solution.
However, two major research gaps remain unresolved. First, the lack of ground-truth
subsidence measurements limits the amount of training data available for AI models.
Second, traditional convolutional neural network (CNN) architectures, such as VGGNet and
ResNet, often fail to achieve satisfactory classification accuracy in this context. In this study,
a Vision Transformer (ViT-Base) model was trained using 191,630 land subsidence grid
measurements paired with high-resolution satellite images. The model achieved an overall
accuracy of 94 % in identifying land subsidence in the region corresponding to the training
data. To further evaluate its generalizability, ten representative mining-affected areas were
selected from China’s top ten coal-producing provinces, each providing 250 independent
subsidence grid measurements paired with high-resolution satellite imagery. The overall
accuracies obtained were ranging from 77.2 % to 84.8 %. These results demonstrate that ViT-
Base consistently outperforms conventional models in identifying mining-induced land
subsidence from satellite imagery, maintaining high accuracy across diverse geographic and
geological settings while requiring less training data. The proposed model thus addresses key
research gaps and provides a practical tool for the monitoring and management of mining-
induced land subsidence.
Keywords: Mining-induced land subsidence; Vision transformer (ViT); Satellite imagery
analysis
Yuwen Wu,
Fusion-based modeling of an intelligent algorithm for enhanced object detection using a
Deep Learning Approach on radar and camera data,
Information Fusion,
Volume 113,
2025,
102647,
ISSN 1566-2535,
[Link]
([Link]
Abstract: Object detection, the process of detecting and classifying objects within a given
environment, forms the foundational element. Multisensory fusion incorporates data from
diverse sensors, like radar and cameras, to refine the reliability and accuracy of detection.
Further, Radar and camera data fusion refine this process by integrating the unique strength
of both technologies, which leverage the radar's proficiency in adverse weather conditions
and the camera's high-resolution imaging. This incorporation enhances the object detection
systems, which enables them to effectively operate across the spectrum of scenarios, from
autonomous vehicles navigating challenging weather to surveillance systems monitoring
critical infrastructure. Deep learning (DL), a branch of machine learning (ML), empowers this
system with the capability to learn complex representations and patterns directly from the
data, which enables them to generalize and adapt to new situations. By integrating the
advanced methodology, we can develop strong perception system capable of interpreting
and detecting objects accurately in dynamic and diverse environments, from autonomous
vehicles navigating urban landscapes to surveillance systems monitoring complex
environments. This study designs an Intelligent Algorithm for Enhanced Object Detection
Using Deep Learning Approach on the Radar and Camera Data Fusion (IAEOD-DLRCDF)
technique. The presented IAEOD-DLRCDF technique uses multi-angle joint calibration where
the spatial sparse alignment of the heterogeneous data of the camera and Radar is realized
with image falsification disregarded. Besides, the IAEOD-DLRCDF technique applies YOLOv8
object detector for radar and camera target detection individually which are then integrated
with the image plane. Moreover, the detected objects are then classified via the
bidirectional long short-term memory (BiLSTM) model. Furthermore, the Adam optimizer is
used for the optimum hyperparameter selection of the BiLSTM network which results in a
better recognition rate. The performance assessment of the IAEOD-DLRCDF method is tested
under benchmark dataset. The empirical analysis stated that the IAEOD-DLRCDF method
gains better performance over other models.
Keywords: Object detection; Deep learning; Data fusion; Radar; YOLOv8; Adam optimizer;
Machine learning
Mahya G.Z. Hashemi, Ehsan Jalilvand, Hamed Alemohammad, Pang-Ning Tan, Narendra N.
Das,
Review of synthetic aperture radar with deep learning in agricultural applications,
ISPRS Journal of Photogrammetry and Remote Sensing,
Volume 218, Part A,
2024,
Pages 20-49,
ISSN 0924-2716,
[Link]
([Link]
Abstract: Synthetic Aperture Radar (SAR) observations, valued for their consistent
acquisition schedule and not being affected by cloud cover and variations between day and
night, have become extensively utilized in a range of agricultural applications. The advent of
deep learning allows for the capture of salient features from SAR observations. This is
accomplished through discerning both spatial and temporal relationships within SAR data.
This study reviews the current state of the art in the use of SAR with deep learning for crop
classification/mapping, monitoring and yield estimation applications and the potential of
leveraging both for the detection of agricultural management practices. This review
introduces the principles of SAR and its applications in agriculture, highlighting current
limitations and challenges. It explores deep learning techniques as a solution to mitigate
these issues and enhance the capability of SAR for agricultural applications. The review
covers various aspects of SAR observables, methodologies for the fusion of optical and SAR
data, common and emerging deep learning architectures, data augmentation techniques,
validation and testing methods, and open-source reference datasets, all aimed at enhancing
the precision and utility of SAR with deep learning for agricultural applications.
Keywords: SAR; Deep learning; Crop classification; Phenology; Yield prediction; Agricultural
management practice
Guozheng Wang, Qinzhe Lv, Liyi Liu, Rong Yang, Bowen Bie, Yaojun Wu, Yinghui Quan,
An end-to-end deep learning framework for separation and parameter measurement of
composite intermittent sampling repeater jamming,
Aerospace Science and Technology,
Volume 168, Part H,
2026,
111218,
ISSN 1270-9638,
[Link]
([Link]
Abstract: To address the severe challenge posed by composite intermittent sampling
repeater jamming (ISRJ) to aerospace radar, we propose a novel end-to-end deep learning
framework for simultaneous signal separation and parameter estimation. At the core of this
framework is a custom deep separation network (CISRJ-SN), which features a unique hybrid
attention architecture. This architecture synergistically fuses one-dimensional convolution
for local feature extraction with a gated attention unit for global dependency modeling,
thereby achieving high-fidelity jammer signal separation even at a low Jammer-to-Noise
Ratio (JNR). Results from both simulations and real-world hardware-in-the-loop experiments
collectively validate the superior performance of our framework. In two-component and
multi-component scenarios, it achieves accuracy rates of 99.5 % and 94.4 % respectively,
marking a significant improvement over the next-best methods. In addition to high accuracy,
the framework demonstrates exceptional estimation precision, reducing the Mean Absolute
Error (MAE) of key parameters by over 60 %. This proves the high stability and reliability of
its estimation results, offering a promising solution for future intelligent sense-and-
countermeasure closed-loop systems.
Keywords: Composite intermittent sampling repeater jamming; Parameter estimation; Deep
learning; Radar anti-jamming; Attention mechanism
Shulin Pang, Zhanqing Li, Lin Sun, Biao Cao, Zhihui Wang, Xinyuan Xi, Xiaohang Shi, Jing Xu,
Jing Wei,
Enhancing cloud detection across multiple satellite sensors using a combined Swin
Transformer and UPerNet deep learning model,
Remote Sensing of Environment,
Volume 334,
2026,
115206,
ISSN 0034-4257,
[Link]
([Link]
Abstract: Cloud detection is crucial in many applications of satellite remote sensing data.
Traditional cloud detection methods typically operate at the pixel level, relying on
empirically tuned thresholds or, more recently, machine learning classification schemes
based on training datasets. Motivated by the success of the Transformer with its self-
attention mechanism and convolutional neural networks for enhanced feature extraction,
we propose a new encoder-decoder method that captures global and regional contexts with
multi-scale features. This new model takes advantage of two advanced deep-learning
techniques, the Swin Transformer and UPerNet (named STUPmask), demonstrating
improved cloud detection accuracy and strong adaptability to diverse imagery types,
spanning spectral bands from visible to thermal infrared and spatial resolutions from meters
to kilometers, across a wide range of surface types, including bright scenes such as ice and
desert, globally. Training and validation of the STUPmask model are conducted using data
obtained from the Landsat 8 and Sentinel-2 Manually Cloud Validation Mask datasets on a
global scale. STUPmask accurately estimates cloud amount with a marginal difference
against reference masks (0.27 % for Landsat 8 and −0.81 % for Sentinel-2). Additionally, the
model captures cloud distribution with a high overall classification accuracy (97.51 % for
Landsat 8 and 96.27 % for Sentinel-2). Notably, it excels in detecting broken, thin, and semi-
transparent clouds across diverse surfaces, including bright surfaces like urban and barren
lands, especially with acceptable accuracy over snow and ice. These encompass the majority
of challenging scenes encountered by cloud identification methods. It also adapts to cross-
sensor satellite data with varying spatial resolutions (4 m–2 km) from both Low-Earth-Orbit
(LEO) and Geostationary-Earth-Orbit (GEO) platforms (including GaoFen-2, MODIS, and
Himawari-8), with an overall accuracy of 94.21–97.11 %. The demonstrated successes in the
automatic identification of clouds with a variety of satellite imagery of different spectral
channels and spatial resolutions render the method versatile for a wide range of remote
sensing studies.
Keywords: Cloud detection; Cross-sensor; STUPmask; Swin Transformer; UPerNet
Farhana Ahmed Chowdhury, Md Kamal Hosain, Md Sakib Bin Islam, Md Shafayet Hossain,
Promit Basak, Sakib Mahmud, M. Murugappan, Muhammad E.H. Chowdhury,
ECG waveform generation from radar signals: A deep learning perspective,
Computers in Biology and Medicine,
Volume 176,
2024,
108555,
ISSN 0010-4825,
[Link]
([Link]
Abstract: Cardiovascular diagnostics relies heavily on the ECG (ECG), which reveals significant
information about heart rhythm and function. Despite their significance, traditional ECG
measures employing electrodes have limitations. As a result of extended electrode
attachments, patients may experience skin irritation or pain, and motion artifacts may
interfere with signal accuracy. Additionally, ECG monitoring usually requires highly trained
professionals and specialized equipment, which increases the treatment's complexity and
cost. In critical care scenarios, such as continuous monitoring of hospitalized patients,
wearable sensors for collecting ECG data may be difficult to use. Although there are issues
with ECG, it remains a valuable tool for diagnosing and monitoring cardiac disorders due to
its non-invasive nature and the detailed information it provides about the heart. The goal of
this study is to present an innovative method for generating continuous ECG waveforms
from non-contact radar data by using Deep Learning. The method can eliminate the need for
invasive or wearable biosensors and expensive equipment to collect ECGs. In this paper, we
propose the MultiResLinkNet, a one-dimensional convolutional neural network (1D CNN)
model for generating ECG signals from radar waveforms. With the help of a publicly
accessible radar benchmark dataset, an end-to-end DL architecture is trained and assessed.
There are six ports of raw radar data in this dataset, along with ground truth physiological
signals collected from 30 participants in five distinct scenarios: Resting, Valsalva, Apnea, Tilt-
up, and Tilt-down. By using strong temporal and spectral measurements, we assessed our
proposed framework's ability to convert ECG data from Radar signals in three distinct
scenarios, namely Resting, Valsalva, and Apnea (RVA). ECG segmentation performed better
by MultiResLinkNet than by state-of-the-art networks in both combined and individual
cases. As a result of the simulations, the resting, valsalva, and RVA scenarios showed the
highest average temporal values, respectively: 66.09523 ± 19.33, 60.13625 ± 21.92, and
61.86265 ± 21.37. In addition, it exhibited the highest spectral correlation values
(82.4388 ± 18.42 (Resting), 77.05186 ± 23.26 (Valsalva), 74.65785 ± 23.17 (Apnea), and
79.96201 ± 20.82 (RVA)), along with minimal temporal and spectral errors in almost every
case. The qualitative evaluation revealed strong similarities between generated and actual
ECG waveforms. As a result of our method of forecasting ECG patterns from remote radar
data, we can monitor high-risk patients, especially those undergoing surgery.
Keywords: ECG; Raw radar data; MultiResLinkNet; CNN; Deep learning
Yuhao Wu, Bin Li, Jun Li, Yonglou Liang, Naiqiang Zhang, Anlai Sun,
Enhancing nighttime cloud detection for moderate resolution imagers using a transformer
based deep learning network,
Remote Sensing of Environment,
Volume 332,
2026,
115067,
ISSN 0034-4257,
[Link]
([Link]
Abstract: Accurate cloud detection is essential for the quantitative applications of satellite
imager observations, but nighttime cloud detection has challenges due to limited spectral
bands, for example, physical methods using only infrared (IR) bands without using spatial
textures as input for cloud detection often result in high uncertainties, especially in some
situations such as cryosphere surface. Although numerous segmentation-style deep learning
cloud detection algorithms have proposed in previous studies, they are inadequate for
nighttime due to the difficulty in acquiring two-dimensional truth data for training and
validation. To overcome these challenges, the Transformer based Nighttime Cloud Detection
(TNCD) framework, which integrates spatial features and utilizes an advanced Transformer
architecture with relative position encoding, layer scaling, and channel attention
mechanisms, is proposed and investigated for nighttime cloud detection. The model was
trained on labels derived from CALIOP data, utilizing a dataset comprising nearly one
hundred million segments from MODIS. Independent validation indicates that TNCD
achieves robust and consistent performance across various scenarios, with an overall
accuracy (OA) of 93.26 % and over 90 % in cryosphere regions. The proposed algorithm
avoids the pattern noise appeared in the traditional physical methodology due to the
utilization of auxiliary data at coarser resolutions, it also mitigates the negative impact of
stripes in IR images for cloud detection. Moreover, TNCD shows high transferable
practicability across sensors, with over 90 % OA for MERSI. More importantly, our research
underscores the importance of water vapor absorption bands for nighttime cloud detection
over the cryosphere. TNCD's high accuracy and robustness provide unique methodology that
could be used operationally for nighttime cloud detection.
Keywords: Nighttime cloud detection; Transformer; Deep learning; MODIS; MERSI; CALIPSO
Yizhen Jia, Hui Chen, Bang Huang, WenKai Jia, Wen-Qin Wang,
Riemannian gradient deep network for joint waveform and filter optimization in MIMO radar
against chopping forwarding jamming and clutter,
Signal Processing,
Volume 239,
2026,
110257,
ISSN 0165-1684,
[Link]
([Link]
Abstract: With the rise of digital radio frequency memory technology, active deception
jamming poses a significant threat to radar systems, especially in detecting targets amid
mainlobe jamming and non-Gaussian clutter. Traditional methods like space–time matched
filtering struggle in such scenarios. This study introduces the Riemannian gradient deep
network (RGDN), a framework for joint optimization of transmit waveforms and receive
filters to improve target detection. Unlike conventional signal-to-clutter noise ratio (SCNR)
maximization, RGDN leverages information geometry to maximize the Kullback–Leibler
Divergence (KLD) between targets and clutter. By modeling non-Gaussian data with a
Gaussian mixture distribution and constructing a Riemannian manifold, the framework
achieves effective jamming suppression through receive filter term in the loss function,
minimizing jamming effects while enhancing target-clutter distinguishability. To address non-
convex optimization, Riemannian gradient descent is integrated into a deep network.
Numerical experiments show that RGDN achieves superior detection performance compared
to SCNR maximization method.
Keywords: Riemannian gradient; KL divergence; Waveform design; MIMO radar; Mainlobe
deception jamming; Deep learning
Jiahao Deng, Yiqing Qian, Feifei Cui, Yanshuang Liu, Jialong Lai,
Research on lunar regolith of the Chang'E-4 landing site: An automated analysis method
based on deep learning framework,
Icarus,
Volume 425,
2025,
116338,
ISSN 0019-1035,
[Link]
([Link]
Abstract: On January 3, 2019, the Chang'E-4 lander successfully landed within the Von
Kármán crater, located in the South P ole-Aitken Basin (SPA) on the farside of the Moon
(45.5°S, 177.6°E), marking the first soft landing on the lunar farside. The lander, equipped
with the Lunar Penetrating Radar (LPR) system, aimed to provide insights into the structure
and evolution of the Moon. Previous research often relied on manually identifying
hyperbolic features to analyze the lunar shallow subsurface properties. This inefficient
approach may lead to subjective biases, resulting in unstable outcomes. This research
constructed an automatic analysis framework by integrating the Swin Transformer with a 3D
velocity spectrum, which is then applied to analyze the properties of the Chang'E-4 LPR data.
The experimental results indicate that the framework achieved a precision of 98.9 % and a
recall of 96.7 % in hyperbolic feature identification, with an F1 of 0.9782 and AP of 94.8 %.
Additionally, it has been experimentally validated that the framework can accurately invert
hyperbolic features' two-way travel time and velocity. Finally, the framework is applied to
analyze the lunar shallow subsurface structure and properties within the landing area of the
Chang'E-4 mission.
Guanliang Liu, Wenchao Chen, Bo Chen, Bo Feng, Penghui Wang, Hongwei Liu,
Supervised contrastive deep Q-Network for imbalanced radar automatic target recognition,
Pattern Recognition,
Volume 161,
2025,
111264,
ISSN 0031-3203,
[Link]
([Link]
Abstract: In the presence of limited and extremely imbalanced data, deep learning methods
for radar automatic target recognition (RATR) often suffer from significant performance
degradation and overfitting. To tackle this issue, we propose Supervised Contrastive Deep Q-
network (SCDQ), a novel end-to-end reinforcement learning method, for multi-class
imbalanced RATR. SCDQ formulates the imbalanced recognition problem as a Markov
decision process (MDP) and optimizes the classifier through an enhanced Q-learning
paradigm. In order to augment the model’s feature extraction capabilities under the
constraint of limited samples, we tightly integrate reinforcement learning (RL) with
supervised contrastive learning, introducing an innovative feature enhancement module. To
further enhance the model’s adaptability to challenging samples, we integrate a
meticulously designed priority sampling into the proposed SCDQ framework, denoted as
SCDQ-P. Experimental results on both simulated and real datasets demonstrate the reliability
and effectiveness of the proposed method.
Keywords: Imbalanced RATR; Deep learning; Deep reinforcement learning (DRL); Supervised
contrastive learning; Priority sampling
Sridhara Murthy B, Gandla Madhu, Varaganti Manisha, Muttukuri Madhu Krishna, Padmam
Divyavalli, Reteneni Naveen,
ODC-net: Scalable and efficient object detection and classification in multi-object CCTV
environments using deep transformer YOLO,
Franklin Open,
Volume 13,
2025,
100404,
ISSN 2773-1863,
[Link]
([Link]
Abstract: The increasing demand for real-time, accurate, and scalable surveillance systems is
driven by the rapid rise in urban Closed-Circuit Television (CCTV) deployments, with global
video surveillance expected to generate over 3 billion video hours daily. However, existing
multi-object detection and classification approaches struggle with poor scalability, temporal
inconsistencies, and sub-optimal accuracy, particularly in complex CCTV environments with
dynamic object motion. To address these limitations, this work proposes Object Detection
Classification Network (ODCNet), a scalable and efficient framework for multi-object
detection, classification, and tracking in CCTV video streams. The model leverages the MS-
COCO-2017 dataset for robust training on diverse object categories, followed by a novel
Hierarchical Spatial Temporal Aggregation (HSTA) feature extraction technique that enhances
spatio-temporal consistency and contextual learning. The core detection and classification
module are powered by Deep Transformer You Only Look at Once V8 (DT- YOLOV8),
combining transformer attention mechanisms with YOLO's real-time detection capabilities.
During testing, input CCTV videos undergo preprocessing through video-to-frame
conversion, with each frame processed using HSTA and DT- YOLOV8 for precise object
detection and classification. Furthermore, to ensure reliable object tracking across video
frames, a Feature Adaptive Continual-Learning Tracker (FACLT) is integrated, enabling
consistent object association and high-quality output video generation with real-time
annotations. Extensive experiments demonstrate the superior performance of ODCNet,
achieving an impressive 99.477 % accuracy, 99.334 % precision, 99.409 % recall, and 99.727
% F1-score, establishing its effectiveness for real-world multi-object surveillance in dynamic
CCTV environments.
Keywords: Cctv videos; Object detection; Real-time video analytics; Data preprocessing;
Hierarchical video intercorrelated similarity; Deep transformer
Ziwei Zhang, Mengtao Zhu, Yunjie Li, Yan Li, Shafei Wang,
Joint recognition and parameter estimation of cognitive radar work modes with LSTM-
transformer,
Digital Signal Processing,
Volume 140,
2023,
104081,
ISSN 1051-2004,
[Link]
([Link]
Abstract: The recent developed cognitive radars can implement flexible work modes with
programmable modulation types and optimized modulating values for each mode definition
parameter. Automatic analysis of these work modes is a significant challenge for modern
electromagnetic reconnaissance receivers. In this paper, a Multi-Output Multi-Structure
(MOMS) learning-based processing framework is proposed for Joint inter-pulse automatic
Modulation Recognition and Parameter Estimation (JMRPE-MOMS). We propose a label
construction method as a feature interpretation method of the network to facilitate MOMS
learning and utilize the correlations between labels for performance gain. Moreover, an
LSTM-Transformer is designed to mine deep time-series characteristics, which can model
local and global relationships and reduce quantization loss. The proposed framework can
perform joint modulation recognition and parameter estimation (JMRPE) tasks
simultaneously with flexible output structures including scalar output and vector output
with fixed or variable sizes. Extensive simulations are performed based on the simulated
radar work modes defined with pulse repetition interval (PRI) sequences. The simulation
results validate the effectiveness and superiority of the proposed method especially under
non-ideal electromagnetic environments.
Keywords: Radar work mode; Automatic modulation recognition; Modulation parameter
estimation; Multi-output learning; Transformer
Ziqi Zhou, Baichun Wang, Zirui Huang, Xiaohui Wu, Weidong Yang, Gang Guo, Shuichangtian
Qiu, Jiakuan Yang, Aijiao Zhou,
Floc image-driven deep learning enhanced by temporal windows and transformers for
carbon emission reduction in drinking water treatment plants,
Water Research,
Volume 289, Part A,
2026,
124868,
ISSN 0043-1354,
[Link]
([Link]
Abstract: Using machine learning (ML) and deep learning (DL) algorithms for precise
coagulant dosing in drinking water treatment plants (DWTPs) helps ensure drinking water
safety and supports greenhouse gas (GHG) emission reduction. The effectiveness of these
algorithms depends heavily on the availability of long-term data. Short-term data are used in
this study to explore the potential of four traditional ML algorithms and four DL algorithms
for precise coagulant dosing. Three strategies were introduced: an innovative method for
floc morphological feature extraction, selection of temporal windows, and integration of
transformer architecture. Based on these strategies, 16 different scenarios were
constructed, resulting in 96 models for analysis. Results show that without any strategy
applied, ML models achieved 5.0% higher R and 10.5% higher R² than DL models. This is due
to their simplicity, faster convergence, and suitability for low-dimensional data. However,
with the proposed strategies, DL models significantly improved and outperformed ML
models. Given the time-lagged dependencies across DWTP treatment units, optimized DL
models N better captured complex nonlinear temporal relationships. The best-performing
model was the temporal convolutional network (TCN) with floc morphological features, 4-h
temporal window, and transformer architecture, achieving R and R2 values of 0.99. The
model was trained with only one month of data and rapidly deployed. A weekly self-
updating mechanism was integrated to ensure long-term adaptability. The model has been
operating stably in a DWTP for over six months. It has reduced coagulant dosage by 20% and
carbon dioxide equivalent (CO2-eq) emissions by an estimated 70 tons annually. This study
demonstrates the strong potential of optimized DL algorithms to improve water purification
and reduce carbon emissions.
Keywords: Floc morphological feature; Temporal window; Transformer; Deep learning;
Carbon emission reduction
Lei Xia, Shurui Zhang, Yuhang Hu, Renli Zhang, Song Li, Weixing Sheng,
A deep learning-based maneuvering target tracking with temporal convolutional networks,
Signal Processing,
Volume 239,
2026,
110322,
ISSN 0165-1684,
[Link]
([Link]
Abstract: Traditional algorithms of the target tracking rely on predefined target motion
states to modify sensor observations. However, these algorithms struggle to accurately and
promptly model the maneuvering state of a target, thereby failing to provide precise state
estimation when the target exhibits maneuvering behavior. To address this challenge, we
propose a maneuvering target tracking algorithm based on temporal convolutional networks
(TcnMTT). The TcnMTT model employs a constant velocity model-based unscented Kalman
filter to decompose the input trajectory into high maneuver state and low maneuver state.
Furthermore, the model directly maps the input observations to the true trajectory through
a set of symmetric TCN networks. Additionally, TcnMTT incorporates an instance
normalization module to project features into a specific feature space and combines a
channel attention mechanism to extract feature correlations. Simulation results demonstrate
that the proposed TcnMTT model outperforms existing methods in tracking maneuvering
targets.
Keywords: Maneuvering target tracking; Temporal convolutional network; Radar tracking;
Deep learning algorithm
Junyu Zhou, Yuting Fu, Sihan Dong, Yuemeng Liu, Han Sun, Yanmin Li, Xunbin Wei,
Multi-model deep learning on photoacoustic flow cytometry signals for real-time melanoma
circulating tumor cells detection and biological characterization,
Expert Systems with Applications,
Volume 308,
2026,
131123,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Background
Melanoma remains one of the most aggressive forms of skin cancer, with early detection
being critical for patient outcomes. This study introduces a novel photoacoustic
fingerprinting approach integrated with advanced machine learning for non-invasive
melanoma detection and circulating tumor cell (CTC) identification.
Methods
We developed a three-tiered photoacoustic fingerprinting system combining photoacoustic
flow cytometry (PAFC) with machine learning algorithms. A uniform PAFC configuration
employed a 532 nm laser for vascular localization followed by a 1064 nm laser targeting
melanin-rich melanoma cells. The study included 50 melanoma patients and healthy
controls, analyzing spectral features across multiple wavelengths. We compared self-
supervised learning architectures (PAFCMamba vs. Transformer) and developed a hybrid
CNN-Transformer model for simultaneous CTC identification, staging, and metastatic
dissemination prediction.
Results
The photoacoustic fingerprinting system achieved exceptional diagnostic discrimination
between melanoma patients and healthy controls. Random Forest achieved area under
curve (AUC) values up to 0.97. The PAFCMamba model outperformed the Transformer
architecture (accuracy 0.75 vs. 0.62, AUC 0.785 vs. 0.730). The hybrid CNN-Transformer
architecture achieved exceptional performance with AUCs up to 0.974 and precision > 94 %
in simultaneous CTC detection, staging, and metastasis prediction. High-immunogenicity
genes including MLANA, GPR89B/A, and PIGF were identified as potential immunotherapy
targets, with photoacoustic signatures serving as non-invasive surrogate biomarkers for
underlying molecular characteristics.
Conclusions
This study establishes photoacoustic fingerprinting as a clinically viable, non-invasive
approach for melanoma detection and CTC monitoring, achieving performance comparable
to conventional methods. The integration of machine learning with photoacoustic
biomarkers provides a scalable framework with interpretable features that facilitates clinical
translation.
Keywords: Photoacoustic fingerprinting; Circulating tumor cells; Deep learning; Non-invasive
diagnosis; Biomarkers; Precision medicine
Ruixing Wang, Wanying Gao, Jianfa Wu, Chunling Wei, Renjian Hao, Huida Yan,
Transformer-enhanced reinforcement learning for spacecraft evasion of asymmetric swarm
threats under complex multi-constraints,
Aerospace Science and Technology,
Volume 168, Part G,
2026,
111200,
ISSN 1270-9638,
[Link]
([Link]
Abstract: Aiming at the challenge of asymmetric threat targets characterized by ”large
numbers, high maneuverability, large fuel capacity and excellent measurement capability”
continuously approaching, a spacecraft threat evasion method based on a Transformer-
Decoder deep reinforcement learning (DRL) architecture is proposed under complex multi-
constraints such as fuel, maneuverability, time, lighting conditions, measurement capability,
and mission continuity. By introducing a multi-head attention mechanism, the method
dynamically allocates attention weights to different threat targets, enabling efficient evasion
of variable threats. Extensive simulation results demonstrate that the proposed approach
achieves superior learning efficiency, convergence speed, and generalization capability
compared with conventional DRL methods. Moreover, it maintains high success rates and
stability across swarms of different sizes and strategies, while outperforming conventional
methods in terms of maneuverability, fuel consumption, and mission continuity. These
results highlight the effectiveness of the proposed approach in enhancing the autonomous
evasion capability of spacecraft, providing a novel solution for safety control in complex
asymmetric threat environments.
Keywords: Spacecraft guidance; Intelligent decision-making; Asymmetric threat evasion;
Deep reinforcement learning; Attention mechanism
Tianwei Mou, Yang Liu, Lintao Tan, Lianhui Wu, Yaya Zhang, Chunming Li, Huan Zhou, Irene
D. Alabia,
Mapping subtle-featured oyster rafts with high-resolution imagery and deep learning
techniques,
ISPRS Journal of Photogrammetry and Remote Sensing,
Volume 231,
2026,
Pages 216-229,
ISSN 0924-2716,
[Link]
([Link]
Abstract: Raft farming is the primary method of oyster aquaculture, with both production
output and environmental impact directly linked to the density of oyster rafts. However, the
absence of accurate monitoring technologies to regulate raft density presents a significant
challenge for the sustainable development of the raft-based aquaculture industry. In this
study, we aimed to integrate deep learning techniques with high-resolution satellite imagery
to improve and support the effective management of offshore aquaculture. We applied a
Transformer-based semantic segmentation model (Mask2Former) alongside several
benchmark models (UNet, PSPNet, Segformer, and KNet) to the Jilin-1 high-resolution
satellite imagery obtained in 2024. This analysis focused on offshore oyster aquaculture in
Rushan City, China. Our results showed that Mask2Former outperformed the other models
in terms of generalizability and validation accuracy. It achieved a mean Intersection over
Union (mIoU) of 88.15 %, a mean accuracy (mAccuracy) of 92.22 %, and a mean F-score (mF-
score) of 93.39 % on the test dataset. Using the extraction maps generated by Mask2Former,
we identified over 40,000 oyster rafts in the offshore region of Rushan, covering an
aquaculture area of 20.43 km2. Compared to traditional optical imagery and extraction
methods, our approach was able to identify individual offshore aquaculture structures,
specifically oyster culture rafts, that are often overlooked. This allows for a more detailed
spatial and structural analysis. This case study underscores the significant potential of
integrating deep learning models with high-resolution satellite imagery to enhance the
sustainable management of offshore aquaculture.
Keywords: Offshore aquaculture; Deep-learning-based model; Semantic segmentation;
Aquaculture management
Yangfan Zhao, Deyi Chen, Yuxuan Wang, Baojie Nie, Yungang Zhao, Qi Li, Shilian Wang,
Dezhong Wang,
Transformer-based deep learning architecture for multivariable radioactive source term
inversion,
Journal of Environmental Radioactivity,
Volume 291,
2026,
107835,
ISSN 0265-931X,
[Link]
([Link]
Abstract: Inversion for the radioactive source term has received growing attention in the
post-Fukushima era. Under some special scenarios, the source term, including release rate,
height and position, is necessary for nuclear emergency response and consequence
assessment. Here, a transformer-based deep learning architecture was developed for
multivariable source term estimation. The CALMET-LAPMOD coupling model validated by
the Kincaid tracer experiment was employed to produce the datasets. The datasets were
systematically constructed for five representative scenarios with the following time-varying
parameters: release rate, release height, release location, coupling of release rate and
height, and coupling of all three variables. Subsequently, a Transformer model with Bayesian
optimization for adaptive hyperparameter tuning was developed. The results demonstrated
excellent performance in source term inversion, with a determination coefficient (R2) of
above 0.96 for release rate and height, and an average distance error of 1.19 km at a 95 %
confidence level for location prediction. Regarding the coupling of all three variables
scenario, the R2 for release rate and location remained above 0.92, whereas the height
achieved R2 of 0.72. Additionally, feature ablation analysis revealed that monitoring points
with high concentration values contribute significantly to inversion, providing quantitative
insights to optimize the monitoring network layout.
Keywords: Nuclear emergency; Atmospheric dispersion; Source term inversion; Transformer
architecture
Gang Xiong, Tao Zhen, Wenyu Huang, Bingxu Min, Wenxian Yu,
Fractal-domain deep learning with Transformer architecture for SAR ship classification,
ISPRS Journal of Photogrammetry and Remote Sensing,
Volume 230,
2025,
Pages 208-226,
ISSN 0924-2716,
[Link]
([Link]
Abstract: This paper extends deep learning from spatiotemporal-frequency to the fractal
domain, and to the best of our knowledge, introduces for the first time the concept of fractal
domain deep learning. Firstly, a Fractal Domain Transformer (FracFormer) model
architecture is proposed to address the challenging problem of SAR image target
classification in complex scenarios. Based on the Singularity Exponent-Domain Image
Feature Transform (SIFT), FracFormer transforms original images into the fractal-domain
feature images, utilizes fractal feature filters and combiners for iterative learning, and
ultimately achieves image classification through fractal feature mixers and classifiers.
Particularly, we derived the fractal feature filtering theorem based on SIFT and the feature
combination theorem based on SIFT, providing theoretical support for the design of the core
modules of FracFormer. On the OpenSARShip2.0 dataset, our model outperforms baseline
models, with improvements ranging from 0.37 % to 11.83 % on average. Besides, extensive
visualization analysis of the model’s fractal domain feature learning results indicates that
FracFormer accords with the two theorems, representing good interpretability. Furthermore,
FracFormer demonstrates fast convergence and strong generalization in low signal-to-noise
ratio scenarios. Specifically, at 0 dB sea clutter, it achieves a 9.96 % improvement in
classification performance over frequency domain GFNet and accelerates convergence by
approximately 36 %. The findings of this study are expected to provide new learning
paradigms and model architectures for the fields of deep learning and computer vision.
Keywords: Fractal domain deep learning; Fractal signal processing; SAR image recognition;
Fractal transformer; Learnalble fractal filtering
Yunfeng Fang, Zheng Tong, Tianqing Hei, Siqi Wang, Tao Ma,
Deep learning applications in ground-penetrating radar inversion: A review,
Measurement,
Volume 258, Part D,
2026,
119399,
ISSN 0263-2241,
[Link]
([Link]
Abstract: The complex nonlinear relationship between the subsurface medium and ground-
penetrating radar signals results in the pervasive ill-posedness and non-uniqueness of
conventional inversion methods. Deep learning, with its powerful feature extraction
capabilities and advantages in modeling complex nonlinear relationships, has unique
strengths in handling complex signals and nonlinear problems, making it especially suitable
for GPR inversion tasks. This paper reviews the latest applications of deep learning in GPR
inversion, summarizing the application strategies of deep learning from two perspectives:
data-driven and data-physics hybrid-driven. Commonly used model architectures and their
performance in signal feature extraction, multi-scale information fusion, and data
preprocessing are discussed, along with the application of various loss functions in inversion
tasks. Finally, current challenges, such as limited model generalization, model dependence
on the dataset and computational efficiency constraints, are discussed, and potential future
research directions are proposed to further advance deep learning in GPR inversion.
Keywords: Ground-penetrating radar; Deep learning; Inversion
Huayan Chen, Yongbin Liu, Yi Liang, Yu Chen, Wenbin Lin, Chaojiang Fu, Caisong Luo,
Multimodal deep learning framework for predicting the evolution of impact damage,
Advanced Engineering Informatics,
Volume 69, Part D,
2026,
104029,
ISSN 1474-0346,
[Link]
([Link]
Abstract: Accurate prediction of damage evolution in structures under impact loading is
crucial for engineering structural safety assessment. Traditional methods primarily rely on
post-hoc qualitative grading systems or simplified models. These approaches struggle to
effectively handle the complex nonlinear characteristics and multi-source heterogeneous
data inherent in impact processes. This study proposes a deep learning framework based on
multimodal fusion for predicting impact damage evolution. The framework first employs the
YOLOv8 instance segmentation model to perform precise mask segmentation on impact
process video frames, extracting time-varying geometric features of structural deformation,
which are subsequently fused with video features extracted via the ResNet18-CBAM-TSM
(RCT) architecture. Subsequently, a Bidirectional Multimodal Autoregressive Transformer
(BiMAR-Transformer) is constructed to achieve deep fusion of video features with physical
features. In addition, energy absorption rate (EAR) and support rotation angle (SRA) are used
to establish a reliable quantitative benchmark for damage evolution prediction. Impact
experimental validation conducted on castellated beams demonstrates that the proposed
framework reduces the Mean Absolute Error (MAE) in EAR prediction by 36.8% compared to
traditional GARCH models, with a Peak Prediction Error (PPE) of merely 1.81% for SRA.
Ablation experiments demonstrated that removing the instance segmentation module
resulted in a 17% reduction in the coefficient of determination (R2) compared to the
complete model, while removing the multimodal fusion module led to increases in PPE of
107.32% and 141.44% for EAR and SRA, respectively. The proposed framework establishes a
novel paradigm for predicting structural impact damage evolution, which is applicable to
various complex structures.
Keywords: Impact damage evolution; Multimodal fusion; Castellated beam; Deep learning;
Instance segmentation; Transformer
Yongjun Zhang, Xiaoshuan Zhang, Baotian Li, Wang Han, Shuran Feng,
VfiA: A vitality fusion identification approach for live fish based on wearable electrical
impedance sensing and deep learning technology,
Sensors and Actuators A: Physical,
Volume 394,
2025,
116947,
ISSN 0924-4247,
[Link]
([Link]
Abstract: This study aims to develop a novel vitality fusion identification approach (VfiA) that
integrates wearable electrical impedance sensing (WEIS) and deep learning technology for
non-invasive live fish monitoring. It builds upon the VMD-SSA-BiLSTM model, which
demonstrates superior identification performance compared to unimodal impedance
monitoring by integrating multi-impedance features and critical survival indicators (e.g.,
blood stress indexes, BSIs; respiratory rate, RR). Specifically, this model utilizes the Sparrow
Search Algorithm (SSA) to optimize a VMD-denoising-enhanced Bidirectional Long Short-
Term Memory (BiLSTM) classifier, thereby achieving higher classification accuracy.
Experimental results on groupers demonstrate that the proposed model attains superior
performance compared to SSA-Transformer, GWO-BiLSTM, PSO-BiLSTM, Transformer, and
BiLSTM within the 11–13℃ range, with an accuracy of 87.5 %, precision of 88.8 %, recall of
88.9 %, and an F1-score of 87.9 %. In conclusion, these findings validate VfiA as a robust
vitality assessment solution, highlighting its potential for optimizing temperature-zone
management in the smart live fish logistics industry.
Keywords: Impedance sensing; Deep learning; Vitality classification; Swarm intelligence
algorithm; Temperature optimization
Shuaiying Zhang, Zhen Dong, Huadong Lin, Zhendong Zhang, Jinran Wu, Sinong Quan,
Wentao An, Tong Li, Rajiv Pandey,
Integrating linear and circular polarization features for PolSAR land cover classification with
deep learning,
International Journal of Applied Earth Observation and Geoinformation,
Volume 146,
2026,
105090,
ISSN 1569-8432,
[Link]
([Link]
Abstract: Existing deep learning methods for ecological monitoring using polarimetric
synthetic aperture radar (PolSAR) imagery primarily rely on coherency (T) or covariance (C)
matrices derived from the linear polarization basis, often overlooking scattering information
inherent in alternative polarization representations. To address this limitation, this study
proposes a novel classification framework that explicitly incorporates a circular polarization
basis into the PolSAR deep learning workflow. A circular coherency matrix (Cir), analogous to
the conventional T matrix, was first derived through polarization basis transformation.
Subsequently, a multi-basis input scheme was introduced to fuse linear and circular
polarization features to enhance feature representation and information utilization. The
proposed framework was validated on two benchmark datasets using multiple deep learning
models, achieving state-of-the-art classification accuracies of 97.70% and 98.58%.Compared
with standard linear-basis approaches, the proposed scheme yielded accuracy
improvements of 2.86% over the T-matrix-based method and 2.26% over the C-matrix-based
method. In addition, the incorporation of circular polarization features significantly
enhanced physical interpretability, particularly for structurally complex targets such as
forests and buildings. Overall, the findings provide an effective technical pathway for
intelligent land cover classification and broader ecological monitoring. The source code and
datasets are available at [Link]
Implementation.
Keywords: PolSAR image classification; Deep learning; Circular polarization basis; Ecological
monitoring; Sustainable land use; Classification performance analysis
Raihan Ahamed Rifat, Fuyad Hasan Bhoyan, Md Humaion Kabir Mehedi, Md Kaviul Hossain,
Md. Jakir Hossen, M.F. Mridha,
ConMatFormer: A multi-attention and transformer integrated ConvNext based deep learning
model for enhanced diabetic foot ulcer classification,
Results in Engineering,
Volume 28,
2025,
108248,
ISSN 2590-1230,
[Link]
([Link]
Abstract: Diabetic foot ulcer (DFU) detection is a clinically significant yet challenging task due
to the scarcity and variability of publicly available datasets. Limited annotated samples
restrict the ability of conventional deep learning models to achieve robust generalization in
real-world clinical scenarios. To solve these problems, we propose ConMatFormer, a new
hybrid deep learning architecture that combines ConvNeXt blocks, multiple attention
mechanisms convolutional block attention module (CBAM) and dual attention network
(DANet), and transformer modules in a way that works together. This design facilitates the
extraction of better local features and understanding of the global context, which allows us
to model small skin patterns across different types of DFU very accurately. To address the
class imbalance, we used data augmentation methods. A ConvNeXt block was used to obtain
detailed local features in the initial stages. Subsequently, we compiled the model by adding a
transformer module to enhance long-range dependency. This enabled us to pinpoint the
DFU classes that were underrepresented or constituted minorities. Tests on the DS1
(DFUC2021) and DS2 (diabetic foot ulcer (DFU)) datasets showed that ConMatFormer
outperformed state-of-the-art (SOTA) convolutional neural network (CNN) and Vision
Transformer (ViT) models in terms of accuracy, reliability, and flexibility. The proposed
method achieved an accuracy of 0.8961 and a precision of 0.9160 in a single experiment,
which is a significant improvement over the current standards for classifying DFUs. In
addition, by 4-fold cross-validation, the proposed model achieved an accuracy of 0.9755
with a standard deviation of only 0.0031. We further applied explainable artificial
intelligence (XAI) methods, such as Grad-CAM, Grad-CAM++, and LIME, to consistently
monitor the transparency and trustworthiness of the decision-making process. These
human-readable tools enhance the comprehension of the explanations and can substantially
increase the practical use of our methodology. Our findings set a new benchmark for DFU
classification and provide a hybrid attention transformer framework for medical image
analysis.
Keywords: Diabetic foot ulcers classification; Multi-attention; Transformer; GradCam;
Explainable AI; LIME
Saidul Islam, Hanae Elmekki, Ahmed Elsebai, Jamal Bentahar, Nagat Drawel, Gaith Rjoub,
Witold Pedrycz,
A comprehensive survey on applications of transformers for deep learning tasks,
Expert Systems with Applications,
Volume 241,
2024,
122666,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Transformers are Deep Neural Networks (DNN) that utilize a self-attention
mechanism to capture contextual relationships within sequential data. Unlike traditional
neural networks and variants of Recurrent Neural Networks (RNNs), such as Long Short-Term
Memory (LSTM), Transformer models excel at managing long dependencies among input
sequence elements and facilitate parallel processing. Consequently, Transformer-based
models have garnered significant attention from researchers in the field of artificial
intelligence. This is due to their tremendous potential and impressive accomplishments,
which extend beyond Natural Language Processing (NLP) tasks to encompass various
domains, including Computer Vision (CV), audio and speech processing, healthcare, and the
Internet of Things (IoT). Although several survey papers have been published, spotlighting
the Transformer’s contributions in specific fields, architectural disparities, or performance
assessments, there remains a notable absence of a comprehensive survey paper that
encompasses its major applications across diverse domains. Therefore, this paper addresses
this gap by conducting an extensive survey of proposed Transformer models spanning from
2017 to 2022. Our survey encompasses the identification of the top five application domains
for Transformer-based models, namely: NLP, CV, multi-modality, audio and speech
processing, and signal processing. We analyze the influence of highly impactful Transformer-
based models within these domains and subsequently categorize them according to their
respective tasks, employing a novel taxonomy. Our primary objective is to illuminate the
existing potential and future prospects of Transformers for researchers who are passionate
about this area, thereby contributing to a more comprehensive understanding of this
groundbreaking technology.
Keywords: Transformer; Self-attention; Deep learning; Natural language processing (NLP);
Computer vision (CV); Multi-modality
Farhatullah, Xin Chen, Deze Zeng, Rahmat Ullah, Rab Nawaz, Jiafeng Xu, Tughrul Arslan,
A deep learning approach for non-invasive Alzheimer’s monitoring using microwave radar
data,
Neural Networks,
Volume 181,
2025,
106778,
ISSN 0893-6080,
[Link]
([Link]
Abstract: Over 50 million people globally suffer from Alzheimer’s disease (AD), emphasizing
the need for efficient, early diagnostic tools. Traditional methods like Magnetic Resonance
Imaging (MRI) and Computed Tomography (CT) scans are expensive, bulky, and slow.
Microwave-based techniques offer a cost-effective, non-invasive, and portable solution,
diverging from conventional neuroimaging practices. This article introduces a deep learning
approach for monitoring AD , using realistic numerical brain phantoms to simulate scattered
signals via the CST Studio Suite. The obtained data is preprocessed using normalization,
standardization, and outlier removal to ensure data integrity. Furthermore, we propose a
novel data augmentation technique to enrich the dataset across various AD stages. Our deep
learning approach combines Recursive Feature Elimination (RFE) with Principal Component
Analysis (PCA) and Autoencoders (AE) for optimal feature selection. Convolution Neural
Network (CNN) is combined with Gated Recurrent Unit (GRU), Bidirectional Long Short Term
Memory (Bidirectional-LSTM), and Long Short-Term Memory (LSTM) to improve
classification performance. The integration of RFE-PCA-AE significantly elevates
performance, with the CNN+GRU model achieving an 87% accuracy rate, thus outperforming
existing studies.
Keywords: Alzheimer’s disease; Classification; Deep learning; Data augmentation; Microwave
scattering; Signal processing
Caihua Hao, Xinyong Mao, Tao Ma, Songping He, Bin Li, Hongqi Liu, Fangyu Peng, Lei Zhang,
A novel deep learning method with partly explainable: Intelligent milling tool wear
prediction model based on transformer informed physics,
Advanced Engineering Informatics,
Volume 57,
2023,
102106,
ISSN 1474-0346,
[Link]
([Link]
Abstract: With the trend of lightweight in the field of intelligent electric vehicles and 3C, the
demand for high precision machining of aluminum alloy parts is growing. And tool condition
monitoring (TCM) is very important for quality control of parts, so intelligent high-accuracy
wear prediction of aluminum alloy high precision machining tools has great industrial
application value at present and in the future. This paper presents a novel TCM model (Conv-
PhyFormer) of Transformer with physics informed. The model has excellent ability to capture
short-term and long-term dependencies from nonlinear cutting time series data when there
are few training samples. The embedded hard physical constraint and soft physical
constraint in the model make the model partially interpretable. Soft physical constraint in
the form of one-dimensional causal convolution can help the proposed model better learn
the local context. Hard physical constraint in the form of the mathematical equation
representing cutting physical knowledge are embedded, thus the model does not need to
learn this knowledge from time series data from scratch. A large number of analysis results
of aluminum alloy machining experimental data show that the proposed Conv-PhyFormer
has significantly superior prediction accuracy and robustness compared with the current
three popular deep learning models for TCM. Embedded soft and hard physical constraints
can significantly reduce the training epochs of Transformer prediction model.
Keywords: Tool condition monitoring; Deep learning; Transformer; Cutting physics
knowledge; Explainability; High Precision Machining
Zhenhua Li, Jiuxi Cui, Heping Lu, Feng Zhou, Yinglong Diao, Zhenxing Li,
Prediction method for instrument transformer measurement error: Adaptive decomposition
and hybrid deep learning models,
Measurement,
Volume 253, Part D,
2025,
117592,
ISSN 0263-2241,
[Link]
([Link]
Abstract: The measurement accuracy of current transformers is crucial for power system
protection and trade fairness. The high penetration of renewable energy into the power grid
has affected the transient performance of power systems, posing significant challenges for
accurate current transformer measurement. To address this issue, this paper proposes a
prediction model for transformer measurement accuracy based on an adaptive dual-modal
decomposition strategy and a hybrid deep learning architecture. The framework integrates
an enhanced Adaptive Time-Varying Filter (A-TVF), an enhanced Adaptive Variational Mode
Decomposition (A-VMD), the Residual Error Index (REI), and the Maximum Information
Coefficient (MIC). First, A-TVF preprocesses the collected data by setting REI as the
optimization objective to adaptively adjust filter construction parameters, including the B-
spline order, bandwidth threshold, and decomposition number, and decomposes the
collected ratio error sequence to reduce the non-stationarity of the original sequence.
Subsequently, indices such as PE and Kurt are used to screen the decomposed sub-
sequences and reconstruct the complex components. Then, A-VMD is applied to further
decompose the complex components, minimizing MIC by adaptively determining the
decomposition number, penalty factor, convergence accuracy, and fidelity parameters.
Afterward, the complexity of the subcomponents obtained from the secondary
decomposition is calculated, and the entire sequence is reconstructed. Finally, a hierarchical
prediction model integrating Temporal Convolutional Networks (TCN), Bidirectional Gated
Recurrent Units (BiGRU), and a Multi-Head Attention mechanism (MHA) is employed to
predict the reconstructed components and generate the final results. Experimental results
demonstrate that the proposed adaptive dual-modal decomposition method significantly
improves prediction performance: compared with non-decomposition models, RMSE, MAE,
and SMAPE were reduced by an average of 50.12%, 46.09%, and 37.70% in global
decomposition scenarios, and by 25.92%, 23.69%, and 19.96% in rolling decomposition
scenarios, respectively. These results validate the effectiveness of the proposed method in
reducing data complexity and improving the accuracy and stability of Ratio Error predictions.
Keywords: ECT; Ratio error prediction; Adaptive dual-modal decomposition; Decomposition
and combination strategy; Hybrid deep model; Measurement accuracy
Wei Wang, Ruobing Song, Yunxiao Wu, Li Zheng, Wenyu Zhang, Zhaoxi Chen, Gang Li, Zhifei
Xu,
Deep learning-based automated diagnosis of obstructive sleep apnea and sleep stage
classification in children using millimeter-wave radar and pulse oximeter,
Sleep Health,
Volume 11, Issue 6,
2025,
Pages 859-867,
ISSN 2352-7218,
[Link]
([Link]
Abstract: Study objectives
Due to the high cost, complexity, and workload of polysomnography, a radar-based sleep
monitoring device, QSA600, has been developed as a more simplified alternative for
children. This study evaluates its agreement with polysomnography for obstructive sleep
apnea diagnosis and sleep staging.
Methods
This diagnostic accuracy study included 281 children (1-18 years) who underwent
simultaneous polysomnography and QSA600 monitoring at Beijing Children's Hospital from
September-November 2023. QSA600 recordings were automatically analyzed using a deep
learning model, while polysomnography data were manually scored.
Results
The obstructive apnea-hypopnea index (OAHI) obtained from QSA600 and polysomnography
demonstrates a high level of agreement with an intraclass correlation coefficient of 0.945
(95% CI: 0.93-0.96). Bland-Altman analysis indicated that the mean difference of obstructive
apnea-hypopnea index between QSA600 and polysomnography was −0.10 events/h (95% CI:
−11.15 to 10.96). The deep learning model evaluated through cross-validation showed good
sensitivity (81.8%, 84.3%, and 89.7%) and specificity (90.5%, 95.3%, and 97.1%) values for
diagnosing children with OAHI >1, OAHI >5, and OAHI >10. The area under the receiver
operating characteristic curve was 0.923, 0.955, and 0.988, respectively. For sleep stage
classification, the model achieved Kappa coefficients of 0.854, 0.781, and 0.734, with
corresponding overall accuracies of 95.0%, 84.8%, and 79.7% for Wake-Sleep classification,
Wake-REM-Light-Deep classification, and Wake-REM-N1-N2-N3 classification, respectively.
Conclusions
QSA600 has demonstrated high agreement with polysomnography in diagnosing obstructive
sleep apnea and performing sleep staging in children. The device is portable, low-burden,
and suitable for follow-up and long-term pediatric sleep assessment.
Keywords: Obstructive sleep apnea; Children; Deep learning; Millimeter-wave radar;
Portable sleep monitoring device; Polysomnography
Tai Dinh, Dat Tran, Zdena Dobešová, Huynh Van Hong, Daniil Lisik, Rameesh Khan,
An efficient fusion-based deep learning framework for land use and land cover image
clustering,
Engineering Applications of Artificial Intelligence,
Volume 161, Part B,
2025,
112061,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Land use and land cover (LULC) analysis is vital for understanding spatial dynamics
and informing environmental management, urban planning, and sustainable development.
Traditional approaches, such as manual surveys and conventional image clustering methods,
often face limitations in scalability and adaptability. This paper presents a novel deep
learning framework that combines the Vision Transformer (ViT) and Variational Autoencoder
(VAE) to extract complementary feature representations for LULC image clustering. The ViT
tokenizes image patches to capture high-level semantic features, while the VAE models
latent structures to integrate contextual and structural information. To further improve
clustering performance, the framework incorporates Uniform Manifold Approximation and
Projection (UMAP) for dimensionality reduction followed by k-means++ clustering, enabling
a scalable and robust solution for diverse datasets. Experiments on multiple datasets,
including the Urban Atlas LULC 2018 dataset and recent LULC maps of Japan and Vietnam,
demonstrate the framework’s superior ability to capture complex LULC patterns compared
to traditional methods. The datasets and source code will be made publicly available at
[Link] This framework has broad applications across
geospatial and remote sensing engineering, civil and environmental engineering, agricultural
planning, transportation, and urban development.
Keywords: Artificial intelligence; Land use and land cover; Urban land use; Deep image
clustering; Transformer; Vision transformer; Variational autoencoder; Uniform manifold
approximation and projection for dimension reduction (UMAP); K-means++
Yilin Bao, Xiangtian Meng, Xingnan Liu, Xue Wang, Zhengchao Qiu, Huanjun Liu, Mingchang
Wang, Abdul Mounem Mouazen,
Integrating bi-dynamic strategy and multivariate deep learning algorithms to predict high-
accuracy, long-term cropland soil organic matter,
International Soil and Water Conservation Research,
2025,
100598,
ISSN 2095-6339,
[Link]
([Link]
Abstract: Soil organic matter (SOM) is a key indicator for assessing soil health and carbon
neutrality, while environmental heterogeneity and dynamic sensitivity changes between
SOM and environmental variables can reduce prediction accuracy. This study developed a
deep learning model that accounts for spatial variability and sensitivity changes in the soil
environment, using 2284 soil samples and 64,802 Landsat TM/OLI images as inputs. A bi-
dynamic strategy is proposed, utilizing a Gaussian mixture model to dynamically partition
the study area and account for changes in environmental heterogeneity across periods. A
multivariate deep learning algorithm, T-CNN-GNN, which combines Transformer, a
convolutional neural network (CNN), and a graph neural network (GNN), was developed to
extract advanced spatio-temporal features, thereby enhancing the accuracy of long-term
SOM spatial distribution predictions. The results showed that the integration of the bi-
dynamic strategy and multivariate deep learning model achieved the highest prediction
accuracy, with a root mean square error (RMSE) of 9.49 g/kg and a coefficient of
determination (R2) of 0.77. The dynamic partitioning strategy effectively captured spatial
variations in environmental heterogeneity across periods. Over the past 40 years, the SOM
content in Northeast China decreased from 41.52 ± 0.24 g/kg to 37.94 ± 0.21 g/kg. This
study showed that SOM maps generated by the bi-dynamic strategy closely matched
measured SOM values, offering a promising method for long-term, high-accuracy soil
mapping.
Keywords: Soil organic matter; Gaussian mixed partitioning; Structural equation modelling;
Deep learning
Yanfei Peng, Jiang He, Qiangqiang Yuan, Shouxing Wang, Xinde Chu, Liangpei Zhang,
Automated glacier extraction using a Transformer based deep learning approach from multi-
sensor remote sensing imagery,
ISPRS Journal of Photogrammetry and Remote Sensing,
Volume 202,
2023,
Pages 303-313,
ISSN 0924-2716,
[Link]
([Link]
Abstract: Glaciers serve as sensitive indicators of climate change, making accurate glacier
boundary delineation crucial for understanding their response to environmental and local
factors. However, traditional semi-automatic remote sensing methods for glacier extraction
lack precision and fail to fully leverage multi-source data. In this study, we propose a
Transformer-based deep learning approach to address these limitations. Our method
employs a U-Net architecture with a Local-Global Transformer (LGT) encoder and multiple
Local-Global CNN Blocks (LGCB) in the decoder. The model design aims to integrate both
global and local information. Training data for the model were generated using Sentinel-1
Synthetic Aperture Radar (SAR) data, Sentinel-2 multispectral data, High Mountain Asia
(HMA) Digital Elevation Model (DEM), and Shuttle Radar Topography Mission(SRTM) DEM.
The ground truth was obtained for a glaciated area of 1498.06 km2 in the Qilian mountains
using classic band ratio and manual delineation based on 2 m resolution GaoFen (GF)
imagery. A series of experiments including the comparison between different models, model
modules and data combinations were conducted to evaluate the model accuracy. The best
overall accuracy achieved was 0.972. Additionally, our findings highlight the significant
contribution of Sentinel-2 data to glacier extraction.
Keywords: Glacier; Deep learning; Transformer; Multi-source data; Qilian mountain
Lv Zhou, TianLiang Chen, Fei Yang, YuanJin Pan, Ling Huang, Xiang Huang, PengDe Lai,
MSFlood-Net: A physically informed deep learning model integrating multi-source data for
flood inundation mapping,
Environmental Modelling & Software,
Volume 196,
2026,
106779,
ISSN 1364-8152,
[Link]
([Link]
Abstract: This study proposes MSFlood-Net, an enhanced U-Net-based deep learning model
for flood extent mapping. The model integrates Synthetic Aperture Radar (SAR), optical
imagery, and topographic inputs including the Digital Elevation Model (DEM) and Height
Above Nearest Drainage (HAND). Multi-scale attention and dilated convolutions enhance
feature representation in complex terrain. A multi-source dataset built upon the publicly
available GF-FloodNet dataset was constructed for training and evaluation. MSFlood-Net
achieves 97.187 % accuracy and a 96.756 % F1 score, outperforming U-Net and DeepLabV3
baselines. It shows strong robustness in delineating flood extent across rivers, reservoirs,
and urban transition areas. Physically informed inputs help reduce false positives from
shadows, clouds, and wet surfaces. MSFlood-Net provides a practical and scalable solution
for flood monitoring and supports integration into broader environmental modelling
systems.
Keywords: Flood inundation mapping; Deep learning; Multi-source remote sensing data;
Physically informed
Hua Wang, Qiangyu Zeng, Hao Wang, Jianxin He, Tiantian Yu, Guangpu Liu,
Temporal super-resolution reconstruction of weather radar echoes using a deep learning
approach,
Expert Systems with Applications,
Volume 300,
2026,
130189,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Severe convective weather events are characterised by rapid evolution and high
destructive potential, requiring weather radars to provide observations with high temporal
resolution. However, current S-band weather radar systems, constrained by their volumetric
scanning strategies, often fail to capture the rapidly changing features of these systems
promptly. To address this limitation, we propose EMAIRA-VFI, a deep learning–based
method for temporal super-resolution reconstruction of radar echoes, which enhances the
temporal resolution of radar data to meet the demands of severe convective weather
monitoring. By introducing an inter-frame attention mechanism, the proposed method
effectively fuses spatiotemporal features from sequential radar echoes, enabling accurate
modelling of dynamic weather evolution and the generation of continuous, high-temporal-
resolution radar echoes. Compared with conventional temporal interpolation methods,
EMAIRA-VFI demonstrates significant improvements in both interpolation accuracy and the
preservation of fine-scale meteorological structures. Experimental results show that the
model not only enhances the capability of S-band radars in monitoring rapidly evolving
weather events but also provides a new perspective for spatiotemporal fusion and the
intelligent application of radar data. We have open-sourced the code for this work at
[Link]
Keywords: Temporal super-resolution; Radar echo; Inter-frame attention mechanism
Kaixiang Zhang, Jiaxiang Zhang, Xinrui Han, Yilin Wang, Bo Wang, Quanhua Liu,
OSCJC: An open-set compound jamming cognition method for radar systems in high-
intensity electromagnetic warfare,
Defence Technology,
Volume 55,
2026,
Pages 436-455,
ISSN 2214-9147,
[Link]
([Link]
Abstract: In high-intensity electromagnetic warfare, radar systems are persistently subjected
to multi-jammer attacks, including potentially novel unknown jamming types that may
emerge exclusively under wartime conditions. These jamming signals severely degrade radar
detection performance. Precise recognition of these unknown and compound jamming
signals is critical to enhancing the anti-jamming capabilities and overall reliability of radar
systems. To address this challenge, this article proposes a novel open-set compound
jamming cognition (OSCJC) method. The proposed method employs a detection-
classification dual-network architecture, which not only overcomes the false alarm and
misdetection issues of traditional closed-set recognition methods when dealing with
unknown jamming but also effectively addresses the performance bottleneck of existing
open-set recognition techniques focusing on single jamming scenarios in compound
jamming environments. To achieve unknown jamming detection, we first employ a
consistency labeling strategy to train the detection network using diverse known jamming
samples. This strategy enables the network to acquire highly generalizable jamming
features, thereby accurately localizing candidate regions for individual jamming components
within compound jamming. Subsequently, we introduce contrastive learning to optimize the
classification network, significantly enhancing both intra-class clustering and inter-class
separability in the jamming feature space. This method not only improves the recognition
accuracy of the classification network for known jamming types but also enhances its
sensitivity to unknown jamming types. Simulations and experimental data are used to verify
the effectiveness of the proposed OSCJC method. Compared with the state-of-the-art open-
set recognition methods, the proposed method demonstrates superior recognition accuracy
and enhanced environmental adaptability.
Keywords: Radar compound jamming cognition; Open-set recognition; Detection-
classification dual-network; Time-frequency analysis; Contrastive learning
Ting Dai, Liye Mei, Yue Zhang, Biao Tian, Rui Guo, Teng Wang, Shan Du, Shiyou Xu,
UAVs and birds classification using robust coordinate attention synergy residual split-
attention network based on micro-Doppler signature measurement by using L-band staring
radar,
Measurement,
Volume 222,
2023,
113692,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Developing unmanned aerial vehicles (UAVs) and birds surveillance technologies to
produce accurate descriptions and achieve high classification accuracy is critical in the field
of radar automatic target recognition (RATR). This article proposes a grayscale spectrogram
image-based UAVs and birds classification method using a robust coordinate attention
synergy residual Split-Attention network (RCA-ResNeSt) under the holographic staring radar
system. Specifically, the ResNet structure with Split-Attention is used as an m-D feature
extractor. The CrossNorm and SelfNorm (CNSN) mechanism is then incorporated into the
network to advance generalization robustness. After that, to consider the spatial direction of
the m-D signature, a coordinated attention (CA) mechanism is introduced at the tail end of
the network to enable fine-grained mining of potential m-D features. Experiments are
carried out using a designed radar system. The results show the superiority of the proposed
method over existing approaches in classification accuracy and noise robustness.
Keywords: Micro-doppler (m-D) signature; Time–frequency representation (TFR); Robust
coordinate attention synergy residual split-attention network (RCA-resNeSt); Holographic
staring radar; Radar automatic target recognition (RATR)
Wei Quan, Wenjing Cheng, Yike Yang, Haiquan Zhao, Zhaoyu Chen, Yunfan Luo,
A signal fingerprint feature extraction method based on decomposition and fusion for radar
emitter individual identification,
Digital Signal Processing,
Volume 164,
2025,
105257,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Radar emitter individual identification is one of the key technologies of modern
electronic countermeasure reconnaissance and electronic intelligence. With the
advancement of radar technology and the increasingly complex electromagnetic
environment, existing methods for identifying emitter are gradually becoming unable to
meet the performance requirements of modern radar individual identification. Aiming at
improving the adaptability of feature extraction for non-cooperative radar emitter signals
and the robustness of individual identification in the complex modern electronic warfare
environment, a signal fingerprint feature extraction method based on decomposition and
fusion is proposed. It firstly integrates signal decomposition and scattering convolution
networks (SCN) to adaptively extract the multi-scale intra-pulse feature of the signal, while
removing the potential noise of the redundant component by energy proportion. And then a
deep feature fusion model based on multi-head self-attention and residual connection is
proposed to fuse the multi-scale features and the time domain features to further extract
signal fingerprint of radar emitter. Experimental results based on the real radar emitter
signals demonstrate that the identification method proposed in this paper can more
effectively extract signal fingerprint features and the identification accuracy reaches 96.45%,
which outperforms other existing identification methods.
Keywords: Radar emitter individual identification; Signal fingerprint feature; Signal
decomposition; Scattering convolution networks (SCN); Fusion
Ayesha Ibrahim, Muhammad Zakir Khan, Muhammad Imran, Hadi Larijani, Qammer H.
Abbasi, Muhammad Usman,
RadSpecFusion: Dynamic attention weighting for multi-radar human activity recognition,
Internet of Things,
Volume 33,
2025,
101682,
ISSN 2542-6605,
[Link]
([Link]
Abstract: This paper presents RadSpecFusion, a novel dynamic attention-based fusion
architecture for multi-radar human activity recognition (HAR). Our method learns activity-
specific importance weights for each radar modality (24 GHz, 77 GHz, and Xethru sensors).
Unlike existing concatenation or averaging approaches, our method dynamically adapts
radar contributions based on motion characteristics. This addresses cross-frequency
generalization challenges, where transfer learning methods achieve only 11%–34% accuracy.
Using the CI4R dataset with spectrograms from 11 activities, our approach achieves 99.21%
accuracy, representing a 15.8% improvement over existing fusion methods (83.4%). This
demonstrates that different radar frequencies capture complementary information about
human motion. Ablation studies show that while the three-radar system optimizes
performance, dual-radar combinations achieve comparable accuracy (24GHz+77GHz: 96.1%,
24GHz+Xethru: 95.8%, 77GHz+Xethru: 97.2%), enabling flexible deployment for resource-
constrained applications. The attention mechanism reveals interpretable patterns: 77 GHz
radar receives higher weights for fine movements (superior Doppler resolution), while 24
GHz dominates gross body movements (better range resolution). The system maintains
71.4% accuracy at 10 dB SNR, demonstrating environmental robustness. This research
establishes a new paradigm for multimodal radar fusion, moving from cross-frequency
transfer learning to adaptive fusion with implications for healthcare monitoring, smart
environments, and security applications.
Keywords: Human activity recognition; Multi-modal fusion; Attention mechanisms; Cross-
frequency transfer learning
Tiantian Wang, Nan Yan, Chaosan Yang, Zeliang An, Gongjing Zhang, Yuqing Xu,
Electromagnetic signal recognition using multimodal tri-branch semantic fusion network in
the UAV-assist integrated sensing and communication systems,
Digital Signal Processing,
Volume 171,
2026,
105820,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Driven by the proliferation of integrated sensing and communication (ISAC)
systems, the accurate recognition of unauthorized unmanned aerial vehicle (UAV) signals in
dynamic electromagnetic environments has emerged as a critical challenge for spectrum
security and cognitive radio applications. Conventional automatic modulation recognition
(AMR) frameworks suffer from significant performance degradation in low signal-to-noise
ratio (SNR) regimes and exhibit limited adaptability to resource-constrained edge computing
platforms. To address these limitations, we propose a novel Multimodal Tri-branch Fusion
Network (MTF-Net) architecture that synergistically integrates time-frequency analysis with
statistical feature learning. The framework systematically processes binarized time-
frequency images (B-TFIs) and higher-order cumulant vectors through three collaboratively
operating branches: (1) A primary temporal feature extractor employing dilated convolution-
residual blocks (DCRBlocks) with hierarchical dilatation factors, incorporating channel
attention mechanisms to dynamically emphasize discriminative temporal patterns; (2) Dual
auxiliary branches based on Edge-Transformer modules (ETFormers), which achieve efficient
spatial-structural learning through depthwise separable convolutions (DSC) while capturing
long-range spectral dependencies via additive attention mechanisms with linear complexity;
(3) A hierarchical fusion module implementing cross-branch feature recalibration through
learnable parameter matrices. Extensive Monte Carlo experiments demonstrate that our
MTF-Net significantly outperforms traditional methods in recognition accuracy for radar and
communication signals under low SNR conditions, establishing a new benchmark for
lightweight AMR solutions in ISAC systems.
Keywords: Multi-modal feature fusion; Unmanned aerial vehicle(UAV); Integrated sensing
and communication (ISAC); Lightweight neural network; Transformer
Yanwen Han, Xiaopeng Yan, Jiawei Wang, Sheng Zheng, Hongrui Yu, Jian Dai,
A sparse moving array imaging approach for FMCW radar with dual-aperture adaptive
azimuth ambiguity suppression and adaptive QR decomposition,
Defence Technology,
Volume 50,
2025,
Pages 254-271,
ISSN 2214-9147,
[Link]
([Link]
Abstract: Range-azimuth imaging of ground targets via frequency-modulated continuous
wave (FMCW) radar is crucial for effective target detection. However, when the pitch of the
moving array constructed during motion exceeds the physical array aperture, azimuth
ambiguity occurs, making range-azimuth imaging on a moving platform challenging. To
address this issue, we theoretically analyze azimuth ambiguity generation in sparse motion
arrays and propose a dual-aperture adaptive processing (DAAP) method for suppressing
azimuth ambiguity. This method combines spatial multiple-input multiple-output (MIMO)
arrays with sparse motion arrays to achieve high-resolution range-azimuth imaging. In
addition, an adaptive QR decomposition denoising method for sparse array signals based on
iterative low-rank matrix approximation (LRMA) and regularized QR is proposed to
preprocess sparse motion array signals. Simulations and experiments show that on a two-
transmitter-four-receiver array, the signal-to-noise ratio (SNR) of the sparse motion array
signal after noise suppression via adaptive QR decomposition can exceed 0 dB, and the
azimuth ambiguity signal ratio (AASR) can be reduced to below −20 dB.
Keywords: Frequency modulated continuous wave (FMCW); Sparse motion array; Range-
azimuth imaging; Azimuth ambiguity suppression; DAAP; Adaptive QR decomposition
Shuai Guo, Ting Chen, Penghui Wang, Jun Ding, Junkun Yan, Hongwei Liu,
Knowledge embedding fusion based on language model for enhanced radar target
recognition,
Signal Processing,
Volume 238,
2026,
110199,
ISSN 0165-1684,
[Link]
([Link]
Abstract: Traditional radar target recognition methods typically model only single echoes,
neglecting the crucial information that domain knowledge can provide for understanding
data. In this paper, we propose a knowledge embedding fusion (KEF) method for enhanced
high-resolution range profile (HRRP) recognition, which utilizes the target state descriptions
available during radar detection. KEF leverages a language model (LM) to integrate textual
knowledge with echo features for fusion recognition. It consists of three components: HRRP
feature extraction, measurement-based knowledge construction, and knowledge embedding
fusion module. First, we perform feature extraction on the HRRP to obtain echo tokens.
Next, in the knowledge construction module, the measurement statuses are standardized to
a natural language format, and the LM is utilized to extract semantic information, resulting in
text tokens. Finally, in the knowledge embedding fusion module, a cross-attention HRRP-text
fusion strategy is employed to facilitate interaction between echo tokens and textual tokens.
We also design a combination of HRRP-text matching loss and fusion classification loss to
guide model training. Experiments are conducted on a real measured dataset, and the
results indicate that KEF effectively enhances recognition performance across multiple
scenarios compared with approaches that only utilize echoes.
Keywords: High-resolution range profile (HRRP); Knowledge embedding fusion; Language
model (LM); Radar target recognition
Xiaosong Tang, Feng Yang, Xu Qiao, Jialin Liu, Haitao Zuo, Liang Gao, Jianshe Zhao, Suping
Peng,
GPR-HIDiff: A diffusion-based model for horizontal interference suppression in urban
underground detection radar profiles,
Underground Space,
Volume 26,
2026,
Pages 458-478,
ISSN 2467-9674,
[Link]
([Link]
Abstract: Automated subsurface utility detection systems in construction rely heavily on the
quality of ground-penetrating radar (GPR) profiles, which are often degraded by high-
amplitude horizontal interference. Existing low-rank decomposition methods lack the
intelligence and flexibility required for multi-site data processing and involve labor-intensive
parameter tuning, impeding their integration into intelligent construction workflows. To
address these challenges, this paper proposes a horizontal interference suppression
algorithm based on a diffusion model, termed GPR-HIDiff. The proposed model replaces
conventional sequential convolutional operators with ResBlocks throughout the encoder,
intermediate layer, and decoder of the UNet architecture, enhancing training stability.
Lightweight agent attention modules are embedded between ResBlocks at each level to
improve global information modeling capability. A spatial attention mechanism is deployed
between the encoder and decoder to achieve adaptive spatial feature optimization.
Furthermore, the forward diffusion phase adopts a cosθ schedule-based strategy to ensure a
smooth temporal variation of noise variance. A standardized dataset comprising real-world
measured samples and finite difference time domain simulation samples of urban road
models has also been constructed. The effectiveness of the hybrid dataset, the introduced
modules, the robustness analysis, and the cosθ schedule is validated through training with
single/mixed datasets, ablation studies, evaluation of metric variations before and after the
introduction of different noise levels, and comparative experiments with constant, linear,
and cosθ schedules. Experimental results demonstrate that GPR-HIDiff significantly
outperforms both traditional methods and state-of-the-art deep learning models on both
simulated and real-world test samples. It effectively suppresses horizontal artifacts,
preserves target hyperbolic contours, and avoids excessive reduction of target scattering,
showcasing its exceptional performance. This method provides a powerful algorithmic
foundation for high-resolution GPR imaging and target detection.
Keywords: Ground-penetrating radar; Horizontal interference; Diffusion model; Agent
attention module; Spatial attention; Hybrid dataset
Yalan Wang, Weiwei Liu, Yibei Wang, Xiaoniu Peng, Zefeng Liu, Changping Wang, Anle Wang,
Multi-format frequency-hopping RF signals generation based on the soliton optoelectronic
oscillator incorporating FDML mechanism,
Optics & Laser Technology,
Volume 194,
2026,
114448,
ISSN 0030-3992,
[Link]
([Link]
Abstract: Emerging optoelectronic oscillators (OEO) are exploring ways to expand their signal
generation formats while retaining excellent phase noise performance to meet the demand
for high-quality signals in information systems such as radar and electronic warfare. Here, a
soliton OEO with reconfigurable operating states is proposed and experimentally
demonstrated for generating multi-format frequency-hopping radio-frequency (RF) signals.
By introducing a composite Fourier-domain mode-locking (FDML) mechanism into the
soliton oscillation state, the proposed scheme can extend the generated signal format from
single-frequency hopping to other hopping patterns −without any structural modifications-
simply by adjusting relative loop parameters. Dual filtering mechanisms in the OEO loop are
realized based on a phase-shift fiber Bragg grating (PS-FBG) cavity and the stimulated
Brillouin scattering (SBS) effect, respectively. Experiments verify the excellent functional
reconfiguration capability of the proposed OEO architecture, generating the aforementioned
signals with adjustable time-domain or frequency-domain parameters. Furthermore, when
the constructed OEO degenerates into a conventional OEO, it can generate different types of
dual continuous-wave signals, further demonstrating its potential as a high-quality arbitrary
waveform generator. Thus, the proposed method not only expands the scope of soliton OEO
but also the application field of the OEO technology.
Keywords: Soliton optoelectronic oscillator; Frequency-hopping; Multi-format; Stimulated
Brillouin scattering; Fourier-domain mode locking
Tingpei Huang, Rongyu Gao, Haotian Wang, Jianhang Liu, Shibao Li,
mBox: 3D object detection based on millimeter-wave radar,
Measurement,
Volume 246,
2025,
116568,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Millimeter-wave radar is utilized for 3D object detection in autonomous driving
due to its advantage of not being affected by lighting conditions. Previous millimeter-wave
radar object detection algorithms have not sufficiently used the distributional and statistical
properties of sparse point clouds. This paper introduces mBox, a 3D object detection
framework using only millimeter-wave radar. To eliminate unnecessary information, we
propose a background filtering algorithm based on subtraction(BGFS) that matches and
differentiates the point cloud frame by frame. We propose a voting-based algorithm for
generating object centers(CPGV), which expands the valid data to generate more accurate
initial anchor boxes. To utilize feature information across different scales and capture the
structure of each granularity, we propose a multi-scale feature fusion network based on the
attention mechanism(GLFF-Net). We conduct experiments using the Pointillism and Astyx
datasets. The results show that the mBox outperforms the comparative methods in terms of
3D mean average precision(mAP).
Keywords: Object detection; Point cloud; Millimeter-wave radar
Ge Junkai, Sun Huaifeng, Shao Wei, Liu Dong, Yao Yuhong, Zhang Yi, Liu Rui, Liu Shangbin,
GPR-TransUNet: An improved TransUNet based on self-attention mechanism for ground
penetrating radar inversion,
Journal of Applied Geophysics,
Volume 222,
2024,
105333,
ISSN 0926-9851,
[Link]
([Link]
Abstract: Convolutional Neural Networks (CNN) are widely applied to Ground Penetrating
Radar (GPR) inversion because they have strong data-driven capabilities and are suitable for
the data structure form of GPR. For CNN, the computation increases with the distance that
the convolutional block moves from one region to another when it calculates the
relationship between two regions. For GPR data, the target reflection exists in the
surrounding traces and full time-window of the target, which leads to high degree of remote
relationship. In this paper, we propose GPR-TransUNet, a deep-learning based inversion
network which use self-attention mechanism. According to the characteristics of GPR data,
regression network and GPR-Loss mechanism were used. Both numerical and model
experiments were arranged to test the performance of the network, and the result as well as
comparative analysis demonstrate the superiority of GPR-TransUNet. Finally, we applied this
method to the field GPR data of Guangxi as an attempt.
Keywords: GPR; Inversion; Deep learning
Chaofeng Huang, Xiaowo Xu, Fan Fan, Shunjun Wei, Xiaoling Zhang, Dongmei Liu, Min Gu,
A low-SNR-adaptive temporal network with smart mask attention for radar signal
modulation recognition,
Digital Signal Processing,
Volume 168, Part D,
2026,
105640,
ISSN 1051-2004,
[Link]
([Link]
Abstract: The automatic modulation recognition of radar signals is a key technology in
electronic warfare and communication systems. However, traditional handcrafted features
often struggle to achieve high recognition accuracy under low signal-to-noise ratio (SNR)
conditions. With the rapid development of artificial intelligence technologies, deep learning-
based approaches have emerged as a promising alternative for modulation recognition. In
this article, a low-SNR-adaptive network architecture is proposed, which integrates a
bidirectional temporal convolutional network (Bi-TCN) and dual-channel smart mask
attention (DSMA) modules. The DSMA adaptively highlights informative features and
suppresses noise through complementary attention masks, enhancing robustness in low-SNR
conditions. Experimental results demonstrate that the autocorrelation domain outperforms
both time and frequency domains, with recognition accuracy improvements of 13.33 % and
14.71 %, respectively. Compared to state-of-the-art models, the proposed network achieves
63 % accuracy at -20 dB and more than 99 % accuracy at -6 dB, significantly enhancing radar
signal modulation recognition.
Keywords: Modulation recognition; Deep learning; Radar signal analysis,
Changlong Wang, Jiawei Jiang, Chong Han, Hengyi Ren, Lijuan Sun, Jian Guo,
Through-Wall Multihuman Activity Recognition Based on MIMO Radar,
Computers, Materials and Continua,
Volume 83, Issue 3,
2025,
Pages 4537-4550,
ISSN 1546-2218,
[Link]
([Link]
Abstract: Existing through-wall human activity recognition methods often rely on Doppler
information or reflective signal characteristics of the human body. However, static
individuals, lacking prominent motion features, do not generate Doppler information.
Moreover, radar signals experience significant attenuation due to absorption and scattering
effects as they penetrate walls, limiting recognition performance. To address these
challenges, this study proposes a novel through-wall human activity recognition method
based on MIMO radar. Utilizing a MIMO radar operating at 1–2 GHz, we capture activity data
of individuals through walls and process it into range-angle maps to represent activity
features. To tackle the issue of minimal variation in reflection areas caused by static
individuals, a multi-scale activity feature extraction module is designed, capable of extracting
effective features from radar signals across multiple scales. Simultaneously, a temporal
attention mechanism is employed to extract keyframe information from sequential signals,
focusing on critical moments of activity. Furthermore, this study introduces an activity
recognition network based on a Deformable Transformer, which efficiently extracts both
global and local features from radar signals, delivering precise human posture and activity
sequences. In experimental scenarios involving 24 cm-thick brick walls, the proposed
method achieves an impressive 97.1% accuracy in activity recognition classification.
Keywords: MIMO radar; human activity; Transformer; through-wall
Muhammad Fahad Munir, Abdul Basit, Wasim Khan, Ahmed Saleem, Aleem Khaliq, Nauman
Anwar Baig,
Next-Gen solutions: Deep learning-enhanced design of joint cognitive radar and
communication systems for noisy channel environments,
Computers and Electrical Engineering,
Volume 120, Part A,
2024,
109663,
ISSN 0045-7906,
[Link]
([Link]
Abstract: In recent years, the dual-function radar and communication (DFRC) paradigm has
emerged as a focal point in addressing spectrum congestion challenges. However, prevailing
research heavily relies on computationally complex likelihood-based approaches for
communication signals with an added Gaussian noise based single waveform. Note that, a
single waveform for diverse scenarios e.g., presence of a communication receiver in the
radar main lobe, side lobe, etc., may lead to a deteriorated detection performance in a DFRC
design. Therefore, in this paper, we present a cognitive DFRC architecture that utilizes a
diverse set of orthogonal waveforms at the transmitter. Specifically, based on a perception-
action cycle, a QAM-based waveform is employed for communication when both the radar
target and communication receiver are within the main lobe, while a PSK-based waveform is
used when the radar target is in the main lobe and the communication receiver is in the side
lobes. Furthermore, to enhance the feature-based estimation, the communication receiver
integrates a Convolutional Neural Network (CNN) architecture designed to autonomously
learn and extract features from received signals with different Signal-to-Noise ratio (SNR).
Next, the adaptive nature of the system enables proficient discernment of the received
signal type and its corresponding SNR value. Moreover, deep learning techniques are applied
in realistic scenarios with various channel impairments to extract features from received
signals, departing significantly from likelihood-based methods and reducing computational
complexity. The proposed methodology’s effectiveness is validated through Monte Carlo
simulations, underscoring its potential to address challenges associated with DFRC under
real-world conditions.
Keywords: CNN based DFRC; DFRC modulation classification; Channel estimation by deep
learning; DFRC spectrum efficiency optimization; DFRC cognitive architecture
Jinyang Xie, Kanghui Zhou, Lei Han, Liang Guan, Maoyu Wang, Yongguang Zheng, Hongjin
Chen, Jiaqi Mao,
Enhancing multi-task learning-based Tornado identification using spatial and temporal
information from weather radar images,
Applied Soft Computing,
Volume 184, Part B,
2025,
113834,
ISSN 1568-4946,
[Link]
([Link]
Abstract: Tornadoes, as dynamic weather phenomena, exhibit unique spatial and temporal
evolution characteristics that reflect their formation and development. Existing tornado
detection algorithms often struggle with high false alarm rates, primarily due to insufficient
capture of temporal correlations in tornado development. As an improvement, we propose a
multi-task tornado identification network with three-dimensional temporal and spatial
information (TS-MTINet). Taking continuous three-frame radar data as input, the Multi-
frame Temporal Interaction Block (MTIB) utilizes multi-head attention to model the dynamic
interaction information between the radar data, thus exploring in-depth the temporal
features during tornado development. Further, we design a Spatial-Temporal Enhancement
Module (STEM), which analyzes the difference information between continuous data to
extract local and global spatial and temporal feature variations about tornadoes. Based on
this architecture, TS-MTINet incorporates a multi-task learning framework to perform
tornado detection and number estimation tasks simultaneously, thus extracting
comprehensive information related to tornadoes. To validate the performance of the
proposed model, we construct the first Chinese tornado identification dataset with fine
radar features. The experimental results show that the proposed method shows significant
advantages in several evaluation metrics, especially in reducing false alarms. In practical case
studies, compared to the traditional TVS method, TS-MTINet achieves an increase in POD of
approximately 30% and a decrease in FAR of about 20% in several typical tornado events.
Particularly in environments with strong interference, TS-MTINet demonstrates higher
detection accuracy, reflecting greater robustness and practical value.
Keywords: Deep learning; Multi-task learning; Tornado identification; Weather radar;
Attention mechanisms
Huihui Ma, Haihong Tao, Yaxing Yue, Tiantian Zhong, Yunfei Fang, Le Wang,
Multiparameter Estimation for Bistatic EMVS-FDA-MIMO Radar with Arbitrarily Configured
Arrays,
Digital Signal Processing,
2026,
105928,
ISSN 1051-2004,
[Link]
([Link]
Abstract: This study explores the multiparameter estimation challenge within bistatic
frequency diverse array multiple-input-multiple-output (FDA-MIMO) radar system that
employs arbitrarily configured electromagnetic vector sensor (EMVS) arrays. The signal
reception model for the presented radar architecture is established. Building on this
foundation, a subspace-based algorithm is proposed to achieve accurate estimation of
spatial-polarization angles and ranges. First, rotation invariant structures in spatial domain
are formed by constructing several virtual steering matrices, from which the normalized
electromagnetic field vectors are derived. Then the two-dimensional direction-of-departure
(2D-DOD) and two-dimensional direction-of-arrival (2D-DOA) estimates are computed
through vector cross-product operation. Thereafter, polarization angles are determined
using least squares (LS) approach. Finally, by compensating the steering matrix with the
obtained 2D-DOD, the range estimation can be achieved. Furthermore, the developed
framework is evaluated for its identifiability, flexibility, computational demands, and Cramér-
Rao bound (CRB). It successfully estimates the targets’ spatial-polarization angles and
ranges, while also achieving automatic parameters pairing. Simulation results demonstrate
the validity of the developed approach.
Keywords: spatial-polarization angles estimation; range estimation; bistatic EMVS-FDA-
MIMO radar; arbitrary array
Xu Meng, Zhaogang Huang, Xin Deng, Hai Liu, Hongyuan Fang, Chao Liu, Xiaoyu Zhang, Jie
Cui,
Leakage detection and localization of buried water pipe using ground penetrating radar,
Measurement,
Volume 254,
2025,
117902,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Leakage in water distribution networks is a critical issue in urban cities worldwide.
Various non-destructive testing methods have been utilized to detect and localize water
leakages in buried pipes. Among these methods, ground penetrating radar (GPR) draws
attention to its advantages of fine resolution, high efficiency, and good portability. Previous
research has shown that water pipe leaks cause noticeable changes in GPR profiles.
However, most studies focus on comparing reflection patterns before and after water leaks,
typically through controlled experiments. The propagation mechanisms of electromagnetic
waves reflected from a leaky pipe and its surrounding media have not been investigated,
hindering direct extraction of leaked information in the GPR data. This paper proposes a co-
simulation algorithm to study electromagnetic wave propagation as water leaks appear and
evolve. Two laboratory experiments were conducted to validate the effectiveness of the co-
simulations. Results indicate that the wetting zone initially spread around the leak point and
finally presented a strawberry shape under the action of gravity. When the leaking water
reaches a certain volume, oscillating signals, composed of reflections from the pipe, the
heterogeneous wetting zone, and the creeping wave, occur in the GPR images. Based on the
findings, a wavelet-entropy-based method is proposed to localize the leak position from 3D
GPR data. Laboratory and field case results demonstrate that the proposed method is
effective for leak detection and possesses a fantastic application prospect.
Keywords: Non-destructive testing (NDT); Ground penetrating radar (GPR); Water pipe
leakage; Combined simulation; Leak localization
Pengfei Wang, Peilin Shu, MingHao Yang, Hongqiu Zhang, Jianqi Wang, Cong Wang, Hongbo
Jia,
Dual-task physiological learning for radar-based continuous blood pressure monitoring:
classification-regularized regression,
Biomedical Signal Processing and Control,
Volume 112, Part C,
2026,
108790,
ISSN 1746-8094,
[Link]
([Link]
Abstract: Continuous blood pressure (BP) monitoring is critical for hypertension
management, yet conventional non-contact radar systems face challenges such as feature
space overlap under low signal-to-noise ratio conditions and inaccurate estimation during
physiological state transitions due to the susceptibility of millimeter-level cardiovascular
vibration signals to environmental interference. To address these challenges, we propose a
dual-task learning model integrating classification-constrained regression with multi-scale
spatiotemporal feature extraction. Our framework combines: (1) A hybrid ResNet-BiGRU
backbone capturing waveform morphology and hemodynamic continuity through multi-
scale convolutions (kernels = 15/7/3) and triple-layer bidirectional gating, enabling robust 2-
second-interval predictions; (2) A physiological regularization mechanism where
classification-derived BP-range probabilities (10 mmHg bins) constrain regression outputs,
suppressing implausible fluctuations during dynamic states. Validated on 30 subjects across
resting, Valsalva, and tilt-table tests, results indicate clinically relevant accuracy (SBP:
−0.21 ± 6.74 mmHg; DBP: 0.25 ± 4.81 mmHg) at 0.5 Hz sampling rate, while demonstrating
improved dynamic-state performance versus benchmarks with DBP RMSE reductions up to
8.4 % by dual-task strategy. This work suggests dual-task learning can mitigate radar-specific
SNR constraints and physiological nonstationarity while fulfilling clinical real-time monitoring
demands (beat-to-beat resolution), contributing to practical deployment of cuffless BP
devices.
Keywords: Blood pressure; Radar; Dual-task learning; Non-contact monitoring; ResNet;
Feature combination
Junkai Ge, Huaifeng Sun, Xiaodong Li, Xushan Lu, Xuening Wang, Li Li, Kejia Hu,
Decoding the stone Buddha: Three-dimensional ground penetrating radar attribute insights
into cracks and restoration history of Sumeru throne,
Journal of Cultural Heritage,
Volume 76,
2025,
Pages 39-51,
ISSN 1296-2074,
[Link]
([Link]
Abstract: The Northern Wei dynasty stone Buddha was built in 517 AD and is currently
housed in Qingdao Museum, located in Laoshan District, Qingdao, China (36°6 ′5.58″N,
120°28′23.42″E). Over the centuries, the natural weathering process and the damage caused
by various relocations has led to internal cracks on its Sumeru throne that threatens the
stability of the Buddha. Previous restoration attempts are visible on the surface of the
throne. To guarantee the quality and effectiveness of further restoration measures, it is
essential to thoroughly investigate the cracks developments and all invisible past restoration
efforts that might interfere future restoration. An ultra-wideband stepped-frequency
continuous wave (SFCW) ground penetrating radar (GPR) system was employed to perform a
non-invasive investigation of the Buddha Sumeru throne. We used a systematic imaging
method to tackle the challenges of detecting tiny internal features within the throne.
Leveraging scattering-based velocity estimation, advanced GPR signal enhancement, Stolt
migration, and envelope attribute extraction, this approach unveils a high-resolution three-
dimensional (3D) image, offering unprecedented insights into subsurface structures. The
obtained images revealed the internal cracks, details of past restoration effort, offering
valuable insights for guiding future restoration efforts. Finally, we discussed the advantages
of GPR for investigating stone statues.
Keywords: Stone Buddha; Crack detection; Restoration history; Ground-penetrating radar
(GPR); Attribute analysis
Shih-Lin Lin,
Advanced Multi-Channel Echo Separation Techniques for High-Interference Automotive
Radars,
Computers, Materials and Continua,
Volume 85, Issue 1,
2025,
Pages 1365-1382,
ISSN 1546-2218,
[Link]
([Link]
Abstract: This paper proposes an integrated multi-stage framework to enhance frequency
modulated continuous wave (FMCW) automotive radar performance under high noise and
interference. The four-stage pipeline is applied consecutively: (i) an improved independent
component analysis (ICA) blindly separates the two-channel echoes, isolating target and
interference components; (ii) a recursive least-squares (RLS) filter compensates amplitude-
and phase-mismatches, restoring signal fidelity; (iii) variational mode decomposition (VMD)
followed by the Hilbert-Huang Transform (HHT) extracts noise-free intrinsic mode functions
(IMFs) and sharpens their time-frequency signatures; and (iv) HHT-based beat-frequency
estimation reconstructs a clean echo and delivers accurate range information. Finally, key
IMFs are reconstructed into a clean signal, and a beat-frequency estimation via HHT confirms
accurate distance results, closely aligning with theoretical predictions. On synthetic data
with an input signal-to-noise ratio (SNR) of 12.7 dB, the pipeline delivers a 7.6 dB SNR gain,
yields a mean-squared error of 0.25 m2, and achieves a range root-mean-square error
(Range-RMSE) of 0.50 m. Empirical evaluations demonstrate that this enhanced ICA and
VMD/HHT scheme effectively restores the fundamental echo signature, providing a robust
approach for advanced driver assistance systems (ADAS).
Keywords: Automotive radar; FMCW; radar noise and interference; independent component
analysis (ICA); variational mode decomposition (VMD); hilbert-huang transform (HHT)
Maged Marghany,
Chapter 3 - Quantized synthetic aperture radar signal: a comprehensive exploration,
Editor(s): Maged Marghany,
Synthetic Aperture Radar Image Processing Algorithms for Nonlinear Oceanic Turbulence
and Front Modeling,
Elsevier,
2024,
Pages 51-88,
ISBN 9780443191558,
[Link]
([Link]
Abstract: This chapter presents a novel perspective on integrating quantum mechanics to
elucidate the complexities of synthetic aperture radar (SAR). Distinguishing itself from
previous works, this chapter avoids an exhaustive exploration of radar theory and signal
processing, as these topics have been thoroughly addressed elsewhere. Instead, the focus is
on introducing a pioneering concept of quantization to comprehend the mechanisms
involved in radar imaging of sea turbulence. Subsequent chapters will delve deeper into this
concept. This chapter seizes the opportunity to clarify the quantization of radar, providing a
lucid understanding of the term. The foundational quantization of electromagnetic waves,
crucial to radar operations, is initially examined. The discussion begins by elucidating why
photons, as fundamental units, exhibit quantization—an essential aspect for understanding
quantized electromagnetic waves. This inquiry reveals why photons possess discrete energy
amounts within a specific energy spectrum, departing from a continuous range. Expanding
beyond conventional discourse, the exploration extends to the detection of quantum
microwave propagation—a novel discussion in the realm of synthetic aperture radar
publications. Moreover, the quantization of the radar equation is implemented to introduce
a fresh perspective on the entanglement between radar cross section and sea surface
turbulence—an aspect previously unexplored in radar oceanography. Noteworthy is the
chapter’s foray into the novel domain of quantum pulse-compression ranging. In the
concluding sections, a comprehensive list of sensors associated with synthetic aperture
radar satellites is compiled, detailing their physical characteristics, including bands and
frequencies. In summary, the quantization of radar signals proposes an innovative
speculation regarding the entanglement of radar cross section with turbulence spectra, such
as the Kolmogorov energy spectra—a contribution unprecedented in the domain of radar
oceanography.
Keywords: Quantum physics; physics; optics; mesoscopic physics; mathematical physics;
quantum mechanics; instrumentation; physical chemistry; space physics; superconductivity;
physics education; computational physics; computer science; statistical applications;
radiation physics; emergent computing; plasma physics; materials characterization; atomic
physics; quantum cosmology
Jiayi Cai, Zhaocheng Yang, Ping Chu, Juntao Guo, Jianhua Zhou,
Robust hand gesture detection and recognition using 4D millimeter-wave radar in a
ubiquitous scene,
Measurement,
Volume 253, Part C,
2025,
117545,
ISSN 0263-2241,
[Link]
([Link]
Abstract: In current research on HGR using radar sensors, hand gestures are typically
confined to a smaller region. However, in ubiquitous scenarios, unrestricted human body
movements and unexpected hand gesture motions usually occur, which results in a large
false alarms and recognition performance degradation. To address this issue, we propose a
robust hand gesture detection and recognition method in ubiquitous scenarios using
Frequency-Modulated Continuous Wave (FMCW) Multiple-Input Multiple-Output (MIMO)
radar. The core idea is to progressively define and classify motions in a cascaded manner,
gradually filtering out non-specific movements, reducing false positives, and enhancing the
applicability of HGR. Specifically, we first propose a suspected hand gesture motion
detection method to help identify suspicious hand gestures. Then, the velocity and position
features of the mutated signal and the stable signal are extracted. A mutated signal motion
recognition method based on a single-layer long short-term memory (LSTM) network is used
to effectively distinguish non-hand gesture motions from hand gestures. Finally, the two-
dimensional trajectory features are extracted, and cascaded with a LSTM network combined
a Gaussian probability model is developed to enhance the ability of open-set recognition.
Experimental results show that the proposed method can achieve the recognition accuracy
of 99.53% for designed hand gestures, the false alarm rate of 1.5% for unexpected hand
gestures and 0.11% for non-hand gesture motions.
Keywords: Hand gesture recognition; Non-hand gesture motion; Feature extraction;
Probability models; Ubiquitous scene
Qiangyu Zeng, Ling Li, Hao Wang, Jianxin He, Hua Wang, Yao Gao,
MCDA-UNet: A satellite data-based model for radar composite reflectivity retrieval,
Atmospheric Research,
Volume 330,
2026,
108619,
ISSN 0169-8095,
[Link]
([Link]
Abstract: The weather radar network in China exhibits an uneven spatial distribution, with
dense coverage in the eastern regions and sparse deployment in the west, resulting in
substantial detection blind spots in areas with complex terrain. This severely limits the
continuity and precision of weather monitoring and early warning in these regions. To
address this challenge, a multi-channel deep learning model, MCDA-UNet, is proposed for
radar composite reflectivity retrieval, aiming to reconstruct and enhance radar echo patterns
in regions lacking radar coverage by leveraging the extensive spatial coverage and
continuous observation capabilities of geostationary meteorological satellites. The model
employs a multi-channel input architecture to extract features from different spectral bands,
while spatial and channel attention models are incorporated to improve the representation
of key meteorological information, thereby enhancing retrieval accuracy and regional
adaptability. Comparative experiments conducted under varying precipitation intensities
demonstrate that MCDA-UNet consistently outperforms existing models across multiple
evaluation metrics, particularly in reconstructing weather radar echo structures and edge
details. These results validate the model’s capability to adapt to the full dynamic range of
weather radar reflectivity and highlight its potential for accurate precipitation retrieval in
radar blind-spot regions.
Keywords: Weather radar composite reflectivity; Satellite data retrieval; Multi-channel
structure; Full dynamic range
Zhipeng Qing, Kecheng Ge, Shunsheng Zhang, Jing Yang, Zhijin Wen, Youlei Pu,
An inverse synthetic aperture radar imaging framework based on multi-layer networks and
heat conduction attention,
Engineering Applications of Artificial Intelligence,
Volume 167, Part 1,
2026,
113708,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Accurate compensation is essential for achieving high-resolution inverse synthetic
aperture radar (ISAR) imaging. Traditional parametric methods usually rely on iterative
optimization of the objective function to compensate for target motion in radar echoes.
However, the iteration process is often computationally intensive, difficult to integrate into
deep learning frameworks, and may discard sufficiently acceptable intermediate solutions.
To address these challenges, this study proposes a deep unfolding-based translational
compensation network that combines unsupervised learning with gradient back-
propagation. A prototype network is incorporated to monitor the imaging process, enabling
early termination of iterations. Moreover, a U-shaped network architecture based on a heat
conduction attention is employed to enhance ISAR image resolution and focusing
performance. To solve the problem of offset or splitting in the imaging results caused by
residual motion errors, a learnable affine transformation is employed for automatic
centering. These modules are integrated into an echo-to-image ISAR imaging framework.
Experimental results on both simulated and real radar data demonstrate the framework’s
effectiveness and robustness.
Keywords: Inverse synthetic aperture radar imaging; Translational compensation; Heat
conduction; Deep unfolding network; Affine transformation
Cries Avian, Jenq-Shiou Leu, Hang Song, Jun-ichi Takada, Nur Achmad Sulistyo Putro,
Muhammad Izzuddin Mahali, Setya Widyawan Prakosa,
RCTrans-Net: A spatiotemporal model for fast-time human detection behind walls using
ultrawideband radar,
Computers and Electrical Engineering,
Volume 120, Part C,
2024,
109873,
ISSN 0045-7906,
[Link]
([Link]
Abstract: Ultrawideband (UWB) radar systems are becoming increasingly popular for
detecting human presence, even through walls. Recent advancements in signal processing
use deep learning techniques, which are known for their accuracy. While earlier methods
focused on spatial information using Convolutional Neural Networks (CNNs), newer research
highlights the importance of temporal information, such as how data peaks shift over time.
This study introduces RCTrans-Net, a deep-learning architecture that combines RCNet (a
Residual CNN) for spatial features with TransNet (a Transformer) for temporal features. This
fusion improves human presence classification in fast-time signal processing. Tested under
various conditions—different materials, body orientations, ranges, and radar heights—
RCTrans-Net achieved high performance with F1-scores of 0.997±0.000 for static,
0.967±0.004 for dynamic, and 0.978±0.001 for combined scenarios. The architecture
outperforms previous methods and offers real-time processing with an inference time of
about one millisecond.
Keywords: Human presence behind the wall; Residual network; Spatiotemporal'
Transformer; Ultrawideband radar system
Yuanbo Li, Wenwu Zhang, Songtao Lv, Jing Yu, Dongdong Ge, Jiawei Guo, Lin Li,
YOLOv11-CAFM model in ground penetrating radar image for pavement distress detection
and optimization study,
Construction and Building Materials,
Volume 485,
2025,
141907,
ISSN 0950-0618,
[Link]
([Link]
Abstract: Ground Penetrating Radar (GPR) is an effective technology for detecting
underground structures and has been widely utilized for monitoring road damage.
Traditional B-scan-based one-dimensional images often fail to preserve continuous spatial
information, thus inadequately reflecting the nuances of damage patterns. This paper
investigates the accurate recognition of hidden internal road damage using 3D-sliced C-scan
images. While YOLO is one of the most effective and rapid neural network models for object
detection, it still suffers from low recognition accuracy and a high rate of missed detections.
To address these issues, this study proposes an improved Convolution and Attention Fusion
Module (CAFM) fusion network model for YOLOv11, which combines the CAFM with the
C2PSA global-local feature extraction mechanism to significantly enhance the recognition
performance for complex road damage. Experimental comparisons between the YOLOv11m-
CAFM and the YOLOv11 model reveal that the combined metrics for the small (n/s) and large
(l/x) models are lower than those for the medium model (m). The YOLOv11m-CAFM
demonstrates strong performance in key metrics such as precision, recall, mAP50, and
mAP50:95, achieving values of 0.840, 0.850, 0.881, and 0.584, respectively, representing
improvements of 0 %, 4.6 %, 1.8 %, and 2.0 % over the baseline model. The confidence level
in detecting standardized targets (e.g., pipelines and well covers) exceeds 0.89. Borehole
validation confirms that the model's localization error is less than 0.15 m, and the detection
frame rate reaches 71 FPS, satisfying the requirements for rapid road assessment. This study
offers a novel method for the intelligent interpretation of GPR images, considering both
detection accuracy and real-time performance, which holds significant engineering
applications in identifying hidden road damages.
Keywords: Pavement disease detection; Ground penetrating radar; Neural network; Object
detection; Deep learning algorithm
Qi Cheng, Shiwen Zhang, Xiaoyang Chen, Hongbiao Cui, Yunfei Xu, Shasha Xia, Ke Xia, Tao
Zhou, Xu Zhou,
Inversion of reclaimed soil water content based on a combination of multi-attributes of
ground penetrating radar signals,
Journal of Applied Geophysics,
Volume 213,
2023,
105019,
ISSN 0926-9851,
[Link]
([Link]
Abstract: Rapid, accurate, and non-destructive acquisition of the distribution of reclaimed
soil moisture information can provide data for the rapid monitoring of reclaimed soil in areas
experiencing coal mining subsidence. However, the water content inversion methods based
on ground penetrating radar (GPR) are mostly single-attribute analysis methods, which are
easily affected by the soil structure. This paper proposes a multi-attribute joint analysis
method, which can reduce the influence of complex soil structure on the prediction results.
Surveys and soil sampling using GPR were performed on a subsided reclamation area in
Huaibei City, Anhui Province, China. Correlation analysis of GPR attribute information and
volumetric water content (VWC) showed that frequency peak (FP), average envelope
amplitude (AEA), energy, instantaneous amplitude area, instantaneous frequency area, and
average instantaneous frequency values were significantly related to the VWC of reclaimed
soil. The applicability of different single-attribute analysis methods under the condition of
reclaimed soil was compared. Results showed that the predictive effects of FP and AEA
attributes were better than other attributes. Generally, the single-attribute analysis method
was greatly affected by the structure of reclaimed soil, so the accuracy and reliability of this
method need to be further optimized. The prediction results of the single- and multi-
attribute joint analysis methods in the structure of reclaimed soil were compared and
analyzed. This showed that the prediction accuracy and model reliability of the multi-
attribute method are both higher than those of the single-attribute method. The multi-
attribute method can overcome the problem of insufficient accuracy of water content
detection of ground-seeking radar in soil with a complex structure. Finally, the multi-
attribute method was used to obtain the water distribution information in the reclamation
area. Results of this study provide new methods and ideas for the prediction of water
content by GPR in soils with complex structure.
Keywords: Ground penetrating radar; Volumetric water content; Attribute analysis; Land
reclamation
Haoming Feng, Huaqing Li, Wenwen Zhu, Denghao Li, Yukun Huang,
Micro-motion enhanced multi-person activity recognition with millimeter-wave radar,
Measurement,
Volume 258, Part B,
2026,
119090,
ISSN 0263-2241,
[Link]
([Link]
Abstract: As a non-contact sensing device, millimeter-wave radar exhibits unique strengths
in human activity recognition (HAR). Existing methods rely on micro-Doppler signatures for
activity classification, but they often encounter feature aliasing in multi-person activity
recognition (MPAR) scenarios. Although point cloud-based approaches can distinguish
individual targets, they primarily extract static morphological features, neglecting the micro-
motion information of human joints, which is crucial for accurate activity recognition. To
address these limitations, we proposes an innovative MPAR framework that integrates
spatial point clouds and micro-motion features. First, an improved point cloud data
association algorithm is applied to achieve multi-target point cloud feature separation,
followed by a dynamic projection mechanism to construct time–Doppler feature maps.
Then, a torso micro-motion enhancement algorithm is designed to enhance the details of
human body movements. Finally, a CNN-LSTM hybrid network architecture with a temporal-
attention is constructed for action classification. Experimental results show that the
proposed micro-motion enhancement algorithm improves recognition accuracy by 27.1%
and 2.3%, compared to two traditional time–frequency analysis methods. Furthermore,
MPAR task in occlusion scenarios achieves recognition accuracy of 93.5%. In summary,
proposed framework not only retains the inherent advantages of millimeter-wave radar but
also significantly enhances multi-person activity recognition in complex scenarios.
Keywords: Human activity recognition (HAR); Multi-person activity recognition (MPAR);
Feature separability; Temporal attention; Occlusion scenarios
Ligen Chen, Nannan Zhu, Hongbo Chen, Yonghao Dong, Yue Zhang, Nian Cai,
A causality-inspired single-source domain generalized method for low-slow-small threat
target recognition through holographic Doppler radar,
Expert Systems with Applications,
Volume 287,
2025,
128104,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Low-slow-small (LSS) target monitoring is critical for airport safety management,
particularly when LSS objects such as unmanned aerial vehicles (UAVs) and birds
unexpectedly enter airport airspace, posing significant risks to regular flights and airport
operations. Recent advances have taken advantage of deep learning for the recognition of
LSS radar targets, achieving promising classification accuracy. However, existing LSS radar
target recognition approaches often rely on statistical correlations, including unstable
spurious correlations, which can undermine the generalization performance of classification
networks, limiting their effectiveness in all-time radar recognition. To address this, we
propose a causality-inspired single-source domain generalization method for radar LSS target
recognition. Our method introduces a Causal-Symmetric Transformation (CST) module for
data augmentation, combining Non-Causal Augmentation for global perturbations and
Symmetric Transformation for local motion reversal, enhancing data diversity and reducing
bias. Additionally, we propose a Causal Mining (CM) module with a Causal Consistency loss
to extract causal features that boost generalization. A Fourier-Aware Attention (FAA) module
leverages frequency-domain information to strengthen feature representation and preserve
causal information. Extensive experiments on four real-world datasets validate the
effectiveness of our approach.
Keywords: Radar target recognition; Single-source domain generalization; Causility-inspired
model; Low-slow-small target; Holographic Doppler radar
Yun Zhou, Yinglin Zhu, Haohao Ren, Jiahao Kang, Xuegang Wang,
Refined multi-modal feature learning framework for marine target detection using radar
sensor,
Digital Signal Processing,
Volume 170,
2026,
105816,
ISSN 1051-2004,
[Link]
([Link]
Abstract: The fusion of time and time-frequency characteristics in radar echoes offers a
novel approach for marine target detection. However, echo amplitude alone cannot fully
characterize the time-domain information, as it fails to capture the temporal correlation
between sampling points. Therefore, this article introduces the Gramian Angular Summation
Field (GASF) for processing raw radar echoes to obtain the temporal information. Concretely,
to enable the detector to utilize features from diverse signal representations of the same
target echoes, we first preprocess the echoes of radar with two signal processing methods,
GASF and STFT, which aim to reflect the temporal dependence and dynamic changes of
frequency components, respectively. Subsequently, we develop a dual-stream feature
extraction network, i.e., time-frequency self-attention learning and GASF-based spatial-
temporal correlation learning, to deeply extract the discriminative features from two
modalities of the same radar echo. Then, to overcome the heterogeneity of multimodal
features during feature fusion, we propose a cross-modal feature fusion strategy to map
multi-modal features to a unified space. Finally, the fused features are fed into the detection
module. Numerous evaluation experiments on the publicly available measured IPIX dataset
demonstrate that the proposed detector is competitive with some state-of-the-art detectors
for marine target detection.
Keywords: Radar target detection; Signal processing; Gramian angular summation field;
Deep learning; Short-time Fourier transform
Yuanjia Xia, Guobing Chen, Zhen Zhang, Shuang Zhao, Zhifang Fei, Kunfeng Li, Xiaoxiao Xia,
Zichun Yang,
Design, research progress and prospects of high temperature infrared/radar compatible
stealth materials,
Optical Materials,
Volume 165,
2025,
117156,
ISSN 0925-3467,
[Link]
([Link]
Abstract: Multispectrum-compatible stealth materials, and in particular, infrared/radar
compatible materials, constitute one of the most important research areas in the stealth
technology field. Although such materials have been extensively investigated at room
temperature, those intended for the high temperature power parts of weapons and
equipment have recently gained increasing attention. This study first analyses and
summarises several typical conventional infrared/radar compatible stealth materials
intended for high temperature conditions from a structural design and mechanistic
viewpoint, then briefly summarises the research status of infrared/radar compatible stealth
metamaterials applicable under high temperature conditions, and finally offers insights into
future development directions.
Keywords: High temperature; Radar wave; Infrared; Compatible stealth
Xiaoyuan Zhang, Shaohang Jing, Jingshu Li, Yechao Bai, Feng Yan,
Cognitive radar recognition with Kolmogorov-Smirnov test and momentum gradient descent,
Digital Signal Processing,
Volume 163,
2025,
105212,
ISSN 1051-2004,
[Link]
([Link]
Abstract: The emission parameters of cognitive radars can adaptively change according to
the environment, which poses a challenge to radar electronic countermeasures (ECM). To
counter cognitive radars, it is essential to identify the cognitive characteristics. In this paper,
a method is proposed to recognize cognitive radars with power allocation function. The
signal-to-interference-plus-noise ratio (SINR) distribution of cognitive radars is derived
through feature functions, and hypothesis test is used to identify whether the target radar
has cognitive function by designing a Kolmogorov-Smirnov (K-S) detector to recognize
adaptive optimization power allocation. Subsequently, a momentum gradient descent
algorithm is used to optimize the signal of the jamming machine to reduce type II error
probability of radar recognition. K-S detector is simulated and compared with Afriat
detector, SVM and MLP detector. Results demonstrate that the K-S detector outperforms
both the Afriat and MLP detectors in identifying cognitive radars with dynamic power
allocation functionality. At the same detection probability, the K-S detector achieves a 2 dB
improvement over the MLP detector and a 4 dB improvement over the Afriat detector.
Keywords: Cognitive radar; Electronic countermeasures (ECM); Kolmogorov–Smirnov test;
Momentum gradient descent algorithm; Afriat theorem
Wenxu Zhang, Lin An, Wencheng Yang, Zhongkai Zhao, Feiran Liu,
Open set recognition of radar specific emitter based on adversarial reciprocal point learning,
Signal Processing,
Volume 238,
2026,
110137,
ISSN 0165-1684,
[Link]
([Link]
Abstract: Radar specific emitter identification (SEI) is a key technology in electromagnetic
spectrum control. Although the emergence of deep learning has promoted the development
of SEI, there are still many shortcomings in the current research results. Most of the
traditional deep learning algorithms are applicable to closed-set identification and can only
be used when the database is complete. In addition, individual differences in radar signals
are susceptible to noise interference, but traditional denoising methods are usually
independent of the feature extraction process, making it difficult to ensure that certain
individual information is not lost. Therefore, in this paper, we propose a new radar emitter
open set recognition method called adversarial reciprocal point learning with adaptive
denoising (ARPLAD). Firstly, we design a new feature extraction network for one-dimensional
signals, which combines deep residual shrinkage network with efficient attention mechanism
to autonomously denoise signals and focus on important parts of signal features. Secondly,
we train the network using adversarial reciprocal point learning combined with center loss
to extract discriminative features with compact intraclass distances and separable interclass
distances, which can efficiently discriminate unknown signals and reduce the risk of open set
identification. The experimental results show that ARPLAD exhibits excellent performance in
different conditions, providing an effective solution for SEI in open electromagnetic
environments.
Keywords: Deep learning; Radar specific emitter identification; Open set recognition;
Adversarial reciprocal point learning
Nanyu Jiang, Yuyuan Fang, Lei Zhang, Chao He, Zhenhua Wu,
IFM-PointNet++: Achieving efficient radar signal waveform recognition with instantaneous
frequency measurement,
Digital Signal Processing,
Volume 168, Part B,
2026,
105507,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Efficient and robust identification of radar signal waveforms is an essential task in
electronic reconnaissance. Current deep learning-based methods can obtain satisfying
accuracy, but they are usually with high computational burden. To address the issue, this
article develops an efficient algorithm IFM-PointNet++. This algorithm transfers the radar
signal waveform recognition into the point cloud recognition task by adopting the
instantaneous frequency measurement (IFM). By integrating IFM with PointNet++ network,
our method achieves superior efficiency and accuracy. To demonstrate this effectiveness, we
conduct comprehensive comparison experiments with YOLOv8 waveform recognition on the
time-frequency images. The results demonstrate that our proposed method significantly
accelerates signal waveform recognition while maintaining high accuracy.
Keywords: Radar signal waveform recognition; Signal recognition; PointNet++; Instantaneous
frequency measurement; Deep learning
Yunfeng Fang, Zheng Tong, Tianqing Hei, Siqi Wang, Tao Ma,
Deep learning applications in ground-penetrating radar inversion: A review,
Measurement,
Volume 258, Part D,
2026,
119399,
ISSN 0263-2241,
[Link]
([Link]
Abstract: The complex nonlinear relationship between the subsurface medium and ground-
penetrating radar signals results in the pervasive ill-posedness and non-uniqueness of
conventional inversion methods. Deep learning, with its powerful feature extraction
capabilities and advantages in modeling complex nonlinear relationships, has unique
strengths in handling complex signals and nonlinear problems, making it especially suitable
for GPR inversion tasks. This paper reviews the latest applications of deep learning in GPR
inversion, summarizing the application strategies of deep learning from two perspectives:
data-driven and data-physics hybrid-driven. Commonly used model architectures and their
performance in signal feature extraction, multi-scale information fusion, and data
preprocessing are discussed, along with the application of various loss functions in inversion
tasks. Finally, current challenges, such as limited model generalization, model dependence
on the dataset and computational efficiency constraints, are discussed, and potential future
research directions are proposed to further advance deep learning in GPR inversion.
Keywords: Ground-penetrating radar; Deep learning; Inversion
Mingyang Du, Ping Zhong, Xiaohao Cai, Daping Bi, Aiqi Jing,
Robust Bayesian attention belief network for radar work mode recognition,
Digital Signal Processing,
Volume 133,
2023,
103874,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Understanding and analyzing radar work modes play a key role in electronic
support measure system. Many classifiers, for example those based on convolutional neural
network (CNN) and recurrent neural network (RNN), are available for recognizing radar work
modes as well as emitter types from their waveform parameters. However, the performance
of these methods may suffer significantly when confronting different types of signal
degradation, e.g., measurement error, lost pulse and spurious pulse. To tackle this issue, we
in this paper develop a Bayesian attention belief network (BABNet) based on Bayesian neural
networks in which the probability distribution over weights can help to enhance the model
robustness for corrupted data. In particular, we adopt pre-trained CNN as the Bayesian
inference prior. This not only accelerates the convergence speed, but also avoids the training
process getting stuck in bad local minima. Meanwhile, instead of using RNNs which are
difficult to be implemented in parallel, the combination of padding operation and attention
module in the proposed BABNet enables CNN, as the backbone, to process sequential data
with variable length. Extensive experiments are conducted to demonstrate the recognition
capability and robustness of the BABNet in different environments.
Keywords: Radar work mode; Pulse descriptor word; Attention mechanism; Bayesian neural
network; Robustness; Recognition
Pianzhang Duan, Li Wang, Cheng Fang, Ziying Song, Ming Gao, Mo Zhou, Ying Li, Yibo Zhang,
Wei Fan, Bin Xu,
Global relationship awareness 3-dimensional object detection using 4-dimensional radar,
Engineering Applications of Artificial Intelligence,
Volume 164, Part B,
2026,
113318,
ISSN 0952-1976,
[Link]
([Link]
Abstract: 4D (4-dimensional) radar sensing technology is essential for high-precision
autonomous driving perception systems, as its superior detection capabilities at increased
distances, compared to traditional LiDAR (Light Detection and Ranging). However, due to the
sparsity of point clouds and the low resolution of millimeter-wave radar, voxel-based
methods may fail to detect distant or closely adjacent objects, leading to inadequate
detection accuracy. To mitigate the accuracy issues arising from the sparse nature of point
clouds in such scenarios, we propose a novel object detection network: GRA-Net (Global
Relation-Aware object detection Network). By leveraging a self-attention mechanism, GRA-
Net effectively learns critical features from each radar pillar, enhancing the network’s
capacity to capture relevant information about nearby objects. Furthermore, we introduce a
global perception module that integrates key features within the pillars and global features,
mitigating the impact of point cloud sparsity, particularly in distant regions. We conducted a
series of experiments to evaluate the performance of GRA-Net. On the Astyx HiRes 2019
dataset, our method achieved 33.63 mAP (mean Average Precision) and 43.93 mAP at the
moderate level; On the View-of-Delft dataset, our method achieved 47.74 mAP in the entire
annotated area and 69.25 mAP in the driving corridor area.
Keywords: 4-dimensional radar; 3-dimensional object detection; Self-attention mechanism;
Autonomous driving
Liangang Qi, Hongzhuo Chen, Qiang Guo, Shuai Huang, Mykola Kaliuzhnyi,
GLS: A hybrid deep learning model for radar emitter signal sorting,
Digital Signal Processing,
Volume 161,
2025,
105117,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Radar emitter signal sorting is a pivotal aspect of radar reconnaissance signal
processing. The increasing density of the electromagnetic environment in modern radar
pulse streams, coupled with the growing complexity and variability of operational modes
and signal forms, results in extremely limited reference data. Consequently, most existing
sorting methods fall short of meeting the performance requirements of modern electronic
warfare. To enhance sorting performance under conditions of limited samples and labeled
data, this paper proposes a radar emitter signal sorting model based on ResGCN-BiLSTM-SE
(GLS). Firstly, we propose a novel adaptive weighted adjacency matrix construction method
that aggregates multi-scale information of local and global features. Based on this, for GLS
networks, the graph convolutional network (ResGCN) is combined with the bidirectional long
short-term memory (BiLSTM) network. The GCN is employed to extract attribute features
from interleaved radar pulse sequences, while the BiLSTM is utilized to deeply capture the
temporal dependence in interleaved pulse sequences after feature extraction. Finally, an
improved squeeze-and-excitation (SE) module is applied to perform weighted fusion of
critical channel information from both spatial and temporal features. Simulation results
demonstrate that the proposed method not only achieves higher accuracy under small
sample conditions compared to existing methods, but also exhibits strong robustness in
challenging scenarios involving measurement errors, missing pulses, and spurious pulses.
Keywords: Radar emitter signal sorting (RESS); Adaptive weighted adjacency matrix; GLS
model; Features fusion
Ayesha Jabbar, Muhammad Kashif Jabbar, Asif Jabbar, Ahmed S. Almasoud, Faijan Akhtar,
Maryam Zulfiqar, Tariq Mahmood, Amjad Rehman,
Enhancing radar tracking accuracy using combined Hilbert transform and proximal gradient
methods,
Results in Engineering,
Volume 24,
2024,
103479,
ISSN 2590-1230,
[Link]
([Link]
Abstract: Accurate radar tracking is crucial in defense, navigation, and surveillance
applications, where high precision and resilience to noise are essential. Traditional radar
tracking techniques, such as Kalman Filters and Particle Filters, often struggle with
performance limitations in noisy and non-linear environments, leading to inaccuracies in
target tracking. To address these challenges, we propose a hybrid radar tracking approach
combining the Hilbert Transform with the Proximal Gradient Method within a convex
optimization framework. This combination leverages the Hilbert Transform's signal
enhancement capabilities with the Proximal Gradient Method's optimization strength,
improving accuracy and robustness under challenging conditions. Experimental results
demonstrate that the proposed method achieves a 23% reduction in Mean Squared Error
(MSE) and a 20% increase in tracking accuracy compared to conventional methods,
alongside a Signal-to-Noise Ratio (SNR) of approximately 18.3 dB, indicating superior noise
resilience. While the hybrid method offers significant improvements, it does involve
increased computational complexity and may be sensitive to initial parameter settings,
requiring careful tuning for optimal performance. Nevertheless, this method represents a
promising advancement over traditional techniques, providing a more accurate and resilient
solution for modern radar tracking applications.
Keywords: Radar tracking; Proximal gradient method; Hilbert transform; Trajectory
estimation; Convex optimization
Hua Wang, Qiangyu Zeng, Hao Wang, Jianxin He, Tiantian Yu, Guangpu Liu,
Temporal super-resolution reconstruction of weather radar echoes using a deep learning
approach,
Expert Systems with Applications,
Volume 300,
2026,
130189,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Severe convective weather events are characterised by rapid evolution and high
destructive potential, requiring weather radars to provide observations with high temporal
resolution. However, current S-band weather radar systems, constrained by their volumetric
scanning strategies, often fail to capture the rapidly changing features of these systems
promptly. To address this limitation, we propose EMAIRA-VFI, a deep learning–based
method for temporal super-resolution reconstruction of radar echoes, which enhances the
temporal resolution of radar data to meet the demands of severe convective weather
monitoring. By introducing an inter-frame attention mechanism, the proposed method
effectively fuses spatiotemporal features from sequential radar echoes, enabling accurate
modelling of dynamic weather evolution and the generation of continuous, high-temporal-
resolution radar echoes. Compared with conventional temporal interpolation methods,
EMAIRA-VFI demonstrates significant improvements in both interpolation accuracy and the
preservation of fine-scale meteorological structures. Experimental results show that the
model not only enhances the capability of S-band radars in monitoring rapidly evolving
weather events but also provides a new perspective for spatiotemporal fusion and the
intelligent application of radar data. We have open-sourced the code for this work at
[Link]
Keywords: Temporal super-resolution; Radar echo; Inter-frame attention mechanism
Yi Zhou, Yu Zhang, Changsheng Chen, Lele Li, Danya Xu, Robert C. Beardsley, Weizeng Shao,
Assessment of radar freeboard, radar penetration rate, and snow depth for potential
improvements in Arctic sea ice thickness retrieved from CryoSat-2,
Cold Regions Science and Technology,
Volume 231,
2025,
104408,
ISSN 0165-232X,
[Link]
([Link]
Abstract: The accuracy of Arctic sea ice thickness retrieved from the CryoSat-2 satellite is
significantly influenced by the sea ice surface roughness, snow backscatter, and snow depth.
In this study, four updated cases incorporating physical model-based radar freeboard, newly
estimated radar penetration rate, and well-validated satellite snow depth were constructed
to evaluate their potential improvements to the Alfred Wegener Institute's CryoSat-2 sea ice
thickness (AWI CS2). The updated cases were then compared with airborne remotely sensed
observations from the National Aeronautics and Space Administration's Operation IceBridge
(OIB) and CryoSat Validation Experiment (CryoVEx) in 2013 and 2014, as well as with ground-
based observations during the Multidisciplinary drifting Observatory for the Study of Arctic
Climate (MOSAiC) expedition from October 2019 to April 2020. The results showed that all
updated cases had the potential to improve the accuracy of sea ice thickness, maintaining
comparable correlation coefficients and significantly reducing statistical errors compared to
the AWI CS2. In the evaluation with OIB, CryoVEx, and MOSAiC, the four updated cases
reduced the root mean square error of AWI CS2 by up to 0.68 m (55 %) against OIB, 0.76 m
(53 %) against CryoVEx, and 0.47 m (76 %) against MOSAiC. The updated sea ice thicknesses
retained the main distribution patterns generated by AWI CS2, but generally showed thinner
sea ice thicknesses. From 2013 to 2018, the interannual variation trends between the
updated cases and AWI CS2 varied regionally, but both show significant decreasing trends
along the northern coasts of the Canadian Arctic Archipelago and Greenland. The updated
schemes provided new insights into the retrieval of sea ice thickness using CryoSat-2,
thereby further contributing to the quantification of the sea ice volume in the context of a
warming climate.
Keywords: Arctic; Sea ice thickness; Satellite retrieval; Assessment
Yadong Xie, Xu Yue, Junfei Zheng, Guangwei Chen, Lin Kong, Dongya Ren,
Mechanistic and spatiotemporal evolution of alkali-pumping in newly constructed bridge
pavements in subtropical environments,
Construction and Building Materials,
Volume 509,
2026,
145163,
ISSN 0950-0618,
[Link]
([Link]
Abstract: In hot and humid regions, early alkali-pumping frequently occurs in newly
constructed asphalt pavement layer on cement concrete bridge deck, with surface whitening
often observed even before the bridge is opened to traffic. This phenomenon severely
compromises the durability and service performance of the bridge deck system. To elucidate
its formation mechanism and dominant influencing factors, this study investigates a newly
built cement concrete bridge deck located in a typical subtropical climate zone, where
extensive surface whitening occurred even before the bridge was opened to traffic. A
combination of field investigation, permeability testing, ground penetrating radar (GPR),
computed tomography (CT) scanning, and X-ray analyses (XRD/XRF) was employed to
systematically explore the water migration pathways and the spatiotemporal characteristics
of alkali-pumping evolution. The results indicate that the early occurrence of alkali-pumping
is closely related to insufficient interlayer compaction, moisture accumulation in structural
depressions, and preferential infiltration through poorly drained zones such as shoulders and
joints. CT analysis demonstrated the presence of interconnected pores within the asphalt
layer, which serve as channels for upward moisture migration and calcium ion transport. XRD
and XRF tests confirmed that the alkali-pumping products are primarily composed of calcium
carbonate, originating from the free calcium components in the cement concrete decks. This
study advances the theoretical understanding of alkali-pumping in cement concrete bridge
decks under hot and humid environments and provides a scientific basis and technical
reference for improving structural design and early-stage damage prevention.
Keywords: Bridge deck pavement; Early alkali-pumping; Formation mechanism; Water
migration path; CT scanning technology
Long He, Kun Zheng, Huihua Ruan, Shuo Yang, Jinbiao Zhang, Cong Luo, Siyu Tang, Yunlei Yi,
Yugang Tian, Jianmei Cheng,
A spatiotemporal mixed-enhanced generative adversarial network for radar-based
precipitation nowcasting,
Computers & Geosciences,
Volume 200,
2025,
105919,
ISSN 0098-3004,
[Link]
([Link]
Abstract: Skillful precipitation nowcasting with high resolution and detailed information
holds promise for providing reliable alerts about severe weather events to society. Radar
echo extrapolation is an essential method for precipitation nowcasting, but traditional
methods struggle to capture rapidly changing regions. Deep learning (DL)-based methods
exhibit superior performance. However, existing DL-based methods face challenges such as
low accuracy, particularly in producing clear forecasts over longer lead times and accurately
forecasting moderate to heavy rainfall events. To address these challenges, we developed a
novel radar-based precipitation nowcasting model, STMixGAN, which can be described as a
nonlinear proximity forecasting model. This model effectively aggregates global-to-local
information and imposes constraints to represent the complex evolution of rainfall
efficiently. Consequently, STMixGAN produces realistic and spatiotemporally consistent
predictions. Using radar observations from South China, STMixGAN successfully forecasted
radar maps for the next 1 h using 24 min of input data. Two traditional methods (Persistence
and Optical flow) and five DL-based methods (ConvLSTM, Rainformer, IAM4VP, REMNet, and
GAN-argcPredNet) were employed as benchmarks to validate STMixGAN’s forecasting
capabilities. The experimental results demonstrate STMixGAN’s superior performance and
provide valuable insights for enhancing heavy rainfall forecasting.
Keywords: Precipitation nowcasting; Spatiotemporal mixed enhancement; Generative
adversarial networks; Self-attention
Xuqian Bai, Zhitao Zhang, Haorui Chen, Long Qian, Tianjin Dai, Ruiqi Li, Shuailong Fan, Sisi
Jing, Junying Chen, Maosheng Ge,
A spatiotemporal fusion algorithm based on Fourier transform is developed to generate daily
surface soil moisture with 20 m spatial resolution,
Geoderma,
Volume 463,
2025,
117548,
ISSN 0016-7061,
[Link]
([Link]
Abstract: Accurate soil moisture data with detailed spatial and temporal resolutions are
essential for hydrological modeling, precision agriculture, and climate research. Nonetheless,
the intrinsic trade-off between spatial and temporal resolution in remote sensing limits the
accessibility of soil moisture products at granular scales. This study presents a
spatiotemporal fusion algorithm utilizing Fourier transform (STFFT), integrated with Random
Forest (RF), the Water Cloud Model (WCM), and the radiative transfer model (PROSAIL) to
create a comprehensive framework for downscaling surface soil moisture (SSM). Employing
Sentinel-1 and Sentinel-2 datasets, we downscaled Soil Moisture Active and Passive (SMAP)
soil moisture products to generate daily Soil Surface Moisture (SSM) maps at a 20-meter
spatial resolution for the study area. The findings indicate that STFFT is more adept at
accommodating SSM data marked by significant heterogeneity and scale discrepancies
compared to traditional spatiotemporal fusion algorithms. Furthermore, STFFT exhibits
computational efficiency and is independent of reference image selection. The
amalgamation of RF with WCM and PROSAIL adeptly elucidates the intricate correlations
between remote sensing variables and soil moisture; the suggested framework attains
precise soil moisture mapping, evidenced by an average correlation coefficient (R) of 0.892
and a root mean square error (RMSE) of 0.034 m3/m3 across diverse land cover types.
Compared to benchmark methods that produce an average R of 0.753 and an RMSE of
0.043 m3/m3, STFFT demonstrates markedly enhanced accuracy and robustness, particularly
in heterogeneous terrains. This study introduces an improved methodology for producing
fine-scale soil moisture products characterized by enhanced spatiotemporal continuity and
reliability.
Keywords: Surface soil moisture; Downscaling; Spatiotemporal fusion algorithm; SMAP;
Sentinel-1/2
Quan Shi, Xiaoliang Xu, Jianlin Li, Huafeng Deng, Qinghai Zhang, Delin Tan, Yu He,
Spatiotemporal effect driven landslide susceptibility mapping at fine scales: a deep learning
model based on multidimensional feature fusion and source data adaptation,
Engineering Applications of Artificial Intelligence,
Volume 156, Part B,
2025,
110924,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Landslides, as frequent natural hazards, pose severe threats to human life and
property safety. This study proposes a deep learning model named Local and Global Feature
Convolutional Network with Interferometric Synthetic Aperture Radar (LGCR-Net). The
model aims to address the issues of multi-source data, heterogeneity, and interference
present in Landslide Susceptibility Mapping (LSM), while also compensating for the
timeliness limitations of LSM. Firstly, the model introduces a parallel structure for extracting
multidimensional features, achieving effective feature fusion through a Multidimensional
Feature Fusion Module(MF). Subsequently, an image enhancement module is incorporated
to process the source data, significantly enhancing the model's stability and source data
adaptability. Furthermore, by integrating the Local and Global Features Convolutional
Network (LGC-Net) model with Interferometric Synthetic Aperture Radar (InSAR) technology,
the LGCR-Net model is formed, enabling the output LSM to possess timeliness. This model
has been successfully applied to the study of the Baihetan reservoir and its surrounding
areas in China. Experimental results indicate that after incorporating the MF and image
enhancement module, the model's Precision increased by 0.0151 and 0.0213, respectively,
while Recall improved by 0.0239 and 0.0334, respectively. Compared to traditional deep
learning models, the LGC-Net model demonstrates superior practicality and predictive
reliability. With the integration of InSAR technology, the improved LSM not only addresses
the model's sensitivity shortcomings in medium and low susceptibility areas but also
maintains predictive performance in high susceptibility areas, providing a more
comprehensive and accurate representation of the potential landslide risks within the study
area.
Keywords: Landslide susceptibility mapping; Multidimensional features; Interferometric
synthetic aperture radar; Multidimensional feature fusion module; Image enhancement
module; Integrating
Jianao Cai, Dongping Ming, Feng Liu, Wenyi Zhao, Mingzhi Zhang, Xiao Ling, Mengyuan Zhu,
Lu Xu, Tingting Lu, Ningjie Liu, Yanfei Wei, Ming Huang,
An enhanced spatiotemporal prediction method on landslide displacement with LDP-
ConvFormer and MT-InSAR observations,
ISPRS Journal of Photogrammetry and Remote Sensing,
Volume 232,
2026,
Pages 594-612,
ISSN 0924-2716,
[Link]
([Link]
Abstract: Landslide Displacement Prediction (LDP) implementation for Landslide Early
Warning Systems (LEWS) using the Multi-Temporal Interferometric Synthetic Aperture Radar
(MT-InSAR) technique poses significant challenges in the Three Gorges Reservoir Area
(TGRA). On the one hand, the limited revisit frequency of satellites fails to satisfy the high-
frequency monitoring requirements of LEWS. On the other hand, traditional LDP methods
concentrate on single-point modeling. It neglects the spatial correlation between
displacement points and the landslide surface. To enhance the low-frequency MT-InSAR
observations, this paper proposes a new hybrid algorithm that integrates the Kalman Filter
(KF) and LDP-ConvFormer to achieve enhanced spatiotemporal LDP. First, multi-orbit MT-
InSAR measurements are transformed to downslope displacement. Subsequently, the multi-
orbit downslope displacements are integrated by KF to generate time series data with
enhanced temporal resolution (5/7-day intervals). The KF estimations indicate that the
integrated higher-resolution time series achieves high accuracy, with an RMSE of 0.431 cm
and an R2 of 0.974 compared to GNSS. Finally, to overcome the limitation of single-point
modeling, a novel LDP-ConvFormer is constructed for enhanced spatiotemporal LDP. The
Spatiotemporal Displacement Prediction Transformer (STDP-Former) employs Local Spatial
Multi-Head Self-Attention (LSMHSA) and Temporal Multi-Head Self-Attention (TMHSA) to
capture the displacement dependencies between different spatial locations at the same time
steps and temporal relationships across different time steps, respectively. Additionally, the
spatiotemporal feature map is decomposed into trend and periodic components, which are
modeled separately and then summed for final predictions. Experimental results
demonstrate that the constructed model can accurately establish the nonlinear relationship
between the landslide displacement and its triggering factors. The LDP-ConvFormer
outperforms benchmark methods, achieving RMSE: 46.29 mm, MAE: 26.7 mm, SSIM:
0.8187, PSNR: 35.62, R2: 0.9574, and EVar: 0.9603. Moreover, LDP-ConvFormer shows
notable superiority in LDP over medium to long periods (60-90d) in the TGRA. The enhanced
spatiotemporal LDP method provides extremely valuable reference for LEWS of translational
landslides in the TGRA.
Keywords: MT-InSAR; Deep learning; Landslide displacement prediction; Kalman filter; TGRA
Yiming Liu, Huadong Guo, Lu Zhang, Dong Liang, Qi Zhu, Zhuoran Lv, Xinyu Dou, Xiaobing Du,
A study of PM2.5 transport pathways in China from 2000 to 2021 with a novel
spatiotemporal correlation method,
Geoscience Frontiers,
Volume 16, Issue 5,
2025,
102116,
ISSN 1674-9871,
[Link]
([Link]
Abstract: In the context of urbanization, air pollution has emerged as a significant
environmental challenge. A thorough understanding of their transport pathways, especially
at a national scale, is essential for environmental protection and policy-making. However, it
remains partially elusive due to the constraints of available data and analytical methods. This
study proposed a data-driven spatiotemporal correlation analysis method employing the
Dynamic Time Warping (DTW). We represented the first comprehensive attempt to chart the
long-term and nationwide transport pathways of PM2.5 utilizing an extensive dataset
spanning from 2000 to 2021 across China, which is crucial for understanding long-term air
pollution trends. Compared with traditional chemical transport models (CTMs), this data-
driven method can generate transport pathways of PM2.5 without requiring extensive
meteorological or emission data, and suggesting fundamentally consistent spatial
distribution and trends. Our analysis reveals that China’s transport pathways are notably
pronounced in the Northwest (34% of the total pathways in China), Southwest (22%), and
North (21%) regions, with less significant pathways in the Northeast (10%) region and
isolated occurrences elsewhere. Additionally, a notable decrease in the number of China’s
PM2.5 transport pathways, similar to annual average concentrations, was observed after
2013, aligning with stricter environmental regulations. Furthermore, we have demonstrated
the feasibility of applying our method to the transport pathways of other gaseous pollutants.
The approach is effective in detecting and quantifying air pollutants’ transport pathways,
even in regions like the Northwest with limited monitoring infrastructure, which may aid in
environmental decision-making. The study will notably improve the current understanding
of air pollutants’ transport process, providing a new perspective for studying the large-scale
spatiotemporal correlations.
Keywords: Air pollutants; Spatiotemporal correlation; Big Earth Data; Transport pathways;
PM2.5; Sustainable development goals
Weicheng Liu, Xia Shi, Wenjun Yan, Shang Gou, Douglas J. Parker, Zhuxia Xu,
Spatiotemporal patterns and propagation characteristics of convective activity on the
northeast slope of Tibetan Plateau: A high-resolution radar perspective,
Atmospheric Research,
Volume 331,
2026,
108609,
ISSN 0169-8095,
[Link]
([Link]
Abstract: The northeastern slope of the Tibetan Plateau, situated in a complex terrain and
the monsoon-westerly transition zone, experiences frequent convective storms and high
disaster risk. Based on CINRAD radar and ERA5 reanalysis data from 2015 to 2019, a high-
resolution convective climatology was established, and its environmental fields were
diagnosed. Results indicate a significant topographic anchoring effect on convection, with
persistent hotspots located in the Yellow River valley-Xinglong Mountain, the eastern Qilian
Mountains, and the sharp-bend reach of the Yellow River. Moderate convection dominates
(64.7 %), while deep convection has a low frequency but high local intensity. The most active
month seasonally is July, with June and August exhibiting similar levels of activity. The
convective activity in July is most active during the season, with levels in June and August
being similar. The configuration of synoptic patterns indicates a synergistic mode of “upper-
level trough, lower-level convergence, and strong moisture transport” for convective days.
Diurnal variation is characterized by a peak in the afternoon (13:00–18:00 BJT) and is
weakest in the early morning to morning, consistent with the synergistic trigger between
solar radiation and terrain convergence. Atmospheric environment diagnostics reveal that,
compared to non-convective events, convective events have higher CAPE, stronger updrafts,
greater vertical wind shear, and more abundant water vapor in the two hours prior to
triggering. The statistical distribution of storm scales exhibits an exponential decay with a
“long-tail” pattern, with approximately 80 %–85 % of convective events having a propagation
distance of less than 20 km and a max area of less than 200 km2. The propagation direction
of convective activity exhibits inter-monthly shifts, trending eastward/southward in June,
shifting northward to east-southeastward in July, and westward/northward in August. These
findings reveal the mechanisms by which large-scale circulation and local topography jointly
influence convection, providing critical scientific support for monitoring and early warning
systems for severe convection in the Tibetan Plateau and its surrounding areas, as well as for
disaster risk prevention and control.
Keywords: Tibetan Plateau; Convective storm; Radar climatology; Propagation
characteristics; Orographic forcing
Chuangwei Xu, Jie Liu, Shiyuan Han, Xiaoqi Duan, Lei Xiang, Tong Zhang,
FourCastLSTM: A precipitation nowcasting model integrating global and local spatiotemporal
features,
Computers & Geosciences,
Volume 204,
2025,
105966,
ISSN 0098-3004,
[Link]
([Link]
Abstract: Accurate precipitation nowcasting is crucial for transportation, agriculture, urban
planning, and tourism, and it is highly beneficial in disaster prevention, resource allocation,
and service optimization. Existing precipitation nowcasting methods often integrate
convolution neural networks and recurrent neural networks or employ vision transformers
to capture spatiotemporal correlations. However, convolutional operators struggle to
capture global information, and vision transformers based global modeling may
overemphasize heavy rainfall while neglecting moderate and light precipitation. In this study,
Fourier nowCasting LSTM (FourCastLSTM) is introduced to effectively capture and fusion
spatiotemporal global and local features of precipitation, enhancing prediction accuracy for
different precipitation intensities. A Fourier nowCasting LSTM Cell (FourCastCell), which
combine the Adaptive Fourier Neural Operator (AFNO) with a simplified LSTM, is proposed
to reinforce the representation of global spatiotemporal precipitation patterns by replacing
traditional convolutional layers with AFNO. An Image Detail Enhancement module (IDE) is
adopted to strengthen local precipitation detail features by integrating difference
convolutional neural network. Finally, the adaptive feature fusion module embedded in the
IDE, can dynamically adjust the integration weights of global and local features based on the
specific spatiotemporal features of precipitation events, ensuring a balanced fusion of
features with different intensities. Experiments on synthetic datasets (MovingMNIST++) and
real-world datasets (RadarCIKM) demonstrate that the proposed FourCastLSTM outperforms
state-of-the-art approaches by 15.6 % and 9.6 % in B-MAE and B-MSE metrics, respectively.
Keywords: Precipitation nowcasting; Spatiotemporal prediction; Spatiotemporal
precipitation feature integration; Heavy rainfall
Yuchen Han, Yiyang Wang, Lei Wang, Changze Zhou, Shijin Yuan,
Extreme-oriented loss: Powering a novel framework for improving spatiotemporal met-
ocean forecasting,
Expert Systems with Applications,
Volume 309,
2026,
131115,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Forecasting spatiotemporal met-ocean series in imbalanced datasets poses a
significant research challenge that warrants substantial attention. Despite specialised
techniques for extreme event prediction, existing methods often prioritise extreme event
prediction accuracy at the expense of accuracy on normal samples. In this paper, we
innovatively proposed a pixel-wise Extreme-oriented Loss (EoL). Distinct from previous
studies that mainly concentrated on the overall prediction difficulty of frames, EoL uniquely
accentuates pixel-wise extreme characteristics and ingeniously integrates potential
imbalances in both temporal and spatial dimensions. The loss applies stronger penalties to
underestimated extreme values while attenuating penalties on their overestimation, thereby
improving the reliability of extreme-event modeling without compromising normal-event
accuracy. Furthermore, a multi-scale feature extraction module is introduced to effectively
capture features across various spatial scales. Additionally, a dependency enhancement
strategy is incorporated, making use of intermediate step prediction information as auxiliary
information to enhance the relatively long-term prediction accuracy. Extensive experimental
evaluations on real-world climate datasets from two distinct domains demonstrate that our
framework consistently outperforms representative state-of-the-art approaches, reducing
MAE and RMSE by up to 24% and yielding notable improvements in R2.
Keywords: Met-ocean forecasting; Spatiotemporal prediction; Data imbalanced; Extreme-
oriented loss
Zhifei Liu, Kang Zheng, Yongze Song, Jianing Zhang,
Daily high-resolution PM2.5 mapping using spatiotemporal CNN-transformer-KAN model,
International Journal of Applied Earth Observation and Geoinformation,
Volume 144,
2025,
104900,
ISSN 1569-8432,
[Link]
([Link]
Abstract: Daily high-resolution mapping of fine particulate matter (PM2.5) is critical for air
quality monitoring and public health. However, current methods struggle to achieve high
accuracy over large spatial and temporal scales due to limitations in modeling complex
spatiotemporal dependencies. This study proposed a novel hybrid deep learning model—
CNN-Transformer-KAN Network (CTKNet)—which utilizes Convolutional Neural Networks
(CNN) to capture spatial features, Transformers for capturing long-range dependencies, and
the Kolmogorov–Arnold Network (KAN) for nonlinear representation learning. Utilizing
spatially continuous satellite aerosol optical depth (AOD) data and other multi-source
spatiotemporal inputs, CTKNet estimated daily PM2.5 at a spatial resolution of 1 km across
China for the period 2015–2020, marking the first application of KAN in PM2.5 estimation. It
outperformed existing models, achieving a high cross-validation coefficient of determination
(R2) of 0.95 (sample-based), 0.90 (station-based), and 0.78 (time-based), and corresponding
RMSEs of 8.13, 11.03, and 17.76 µg/m3. Yearly sample-based cross-validation R2 values
ranged from 0.91 to 0.96 with RMSEs below 11.74 µg/m3, while seasonal R2 values ranged
from 0.82 to 0.89 with RMSEs below 22.14 µg/m3. Analysis reveals a significant decline in
PM2.5 nationwide, especially in eastern and central China. Seasonal peaks occur in winter,
with minima in summer, influenced by meteorology. Spatially, PM2.5 is highest in eastern
and northern regions; urban agglomerations like Beijing–Tianjin–Hebei (BTH) show severe
pollution, while Pearl River Delta (PRD) exhibits the lowest levels due to favorable
conditions. CTKNet also holds promise for other fine-scale environmental mapping tasks
using multi-source spatiotemporal data.
Keywords: PM2.5 estimation; Kolmogorov–Arnold Network; Hybrid deep learning model;
Satellite AOD; Air quality assessment
Zhiyun Yang, Hao Wu, Qi Liu, Xiaodong Liu, Yonghong Zhang, Xuefei Cao,
A self-attention integrated spatiotemporal LSTM approach to edge-radar echo extrapolation
in the Internet of Radars,
ISA Transactions,
Volume 132,
2023,
Pages 155-166,
ISSN 0019-0578,
[Link]
([Link]
Abstract: In recent years, the number of weather-related disasters significantly increases
across the world. As a typical example, short-range extreme precipitation can cause severe
flooding and other secondary disasters, which therefore requires accurate prediction of
extent and intensity of precipitation in a relatively short period of time. Based on the echo
extrapolation of networked weather radars (i.e., the Internet of Radars), different solutions
have been presented ranging from traditional optical-flow methods to recent deep neural
networks. However, these existing networks focus on local features of echo variations to
model the dynamics of holistic radar echo motion, so it often suffers from inaccurate
extrapolation of the radar echo motion trend, trajectory, and intensity. To address the
problem, this paper introduces the self-attention mechanism and an extra memory that
saves global spatiotemporal feature into the original Spatiotemporal LSTM (ST-LSTM) to form
a self-attention Integrated ST-LSTM recurrent unit (SAST-LSTM), capturing both spatial and
temporal global features of radar echo motion. And several these units are stacked to build
the radar echo extrapolation network SAST-Net. Comparative experiments show that the
proposed model has better performance on different real world radar echo datasets over
other recent methods.
Keywords: Radar echo extrapolation; Self-attention; Long short-term memory;
Spatiotemporal prediction
Chuyao Luo, Xinyue Zhao, Yuxi Sun, Xutao Li, Yunming Ye,
PredRANN: The spatiotemporal attention Convolution Recurrent Neural Network for
precipitation nowcasting,
Knowledge-Based Systems,
Volume 239,
2022,
107900,
ISSN 0950-7051,
[Link]
([Link]
Abstract: Precipitation nowcasting is an important task in the fields of transportation, traffic,
agriculture, and tourism. One of the main challenges is radar echo maps forecasting. It is
regarded as a spatiotemporal sequence prediction problem. The prevailing approaches
including the state-of-the-art methods are all based on the ConvRNN which combines the
Convolution Neural Network (CNN) and Recurrent Neural Network (RNN). However, the
feature flow delivered in multi-layer CNNs and RNN usually accompanies the information
loss. Therefore, these algorithms fail to model the long-term dependency and the heavy
rainfalls tend to be underestimated. In addition, they cannot predict the increasing intensity
trend of heavy rainfalls. In this paper, we propose a PredRANN model by embedding the
Temporal Attention Module (TAM) and Layer Attention Module (LAM) into the prediction
unit to preserve more representation from temporal and spatial dimensions respectively.
The extensive experimental results on both synthetic data sets and real world data sets
demonstrate the effectiveness and superiority of the proposed method over state-of-the-art
methods. Ablation studies also validate the developed TAM and LAM components. To
reproduce the results, we release the source code at:
[Link]
Keywords: Precipitation nowcasting; Spatiotemporal sequence prediction; Self attention
Hansheng Zeng, Yuqi Li, Ruize Niu, Chuanguang Yang, Shiping Wen,
Enhancing spatiotemporal prediction through the integration of Mamba state space models
and Diffusion Transformers,
Knowledge-Based Systems,
Volume 316,
2025,
113347,
ISSN 0950-7051,
[Link]
([Link]
Abstract: This paper presents an advanced architecture for spatiotemporal prediction MAD,
integrating Mamba modules with Diffusion Transformers for efficient spatiotemporal
modeling. The model consists of three phases: encoding, reconstruction, and prediction.
Initially, the encoder transforms raw spatiotemporal data into compact latent embeddings.
In the reconstruction phase, the Mamba module processes these embeddings through
normalization and bidirectional state space models, generating reconstructed
representations which are then decoded to restore the input data. The prediction phase
utilizes the Diffusion Transformer to model spatiotemporal features, incorporating time
embeddings and leveraging self-attention mechanisms to capture complex spatiotemporal
dependencies. Finally, the model jointly trains the reconstruction and prediction paths to
achieve high-precision spatiotemporal forecasts. Experimental results demonstrate the
model’s superior performance across various spatiotemporal prediction tasks, validating its
effectiveness and robustness. Our codes are available at
[Link]
Keywords: Deep learning; Spatio-temporal prediction; Mamba; Diffusion
Tuo Xie, Xinyao Yun, Gang Zhang, Hua Li, Kaoshe Zhang, Ruogu Wang,
Charging station cluster load prediction: Spatiotemporal multi-graph fusion technology,
Renewable and Sustainable Energy Reviews,
Volume 206,
2024,
114855,
ISSN 1364-0321,
[Link]
([Link]
Abstract: In recent years, single-station charging load prediction technology for electric
vehicles has gradually matured, but there are few prediction studies at the charging station
cluster level. Therefore, this research propose a load prediction framework for electric
vehicle charging station groups based on multi-graph fusion. First, a distance map, a traffic
network map, and a traffic density map are established to extract the topological
information of the charging station group, and the fusion operation is performed by
establishing the relationship between the influencing factors based on the mutual
correlation of the influencing factors and the multi-graph attention mechanism; Secondly, a
spatiotemporal prediction model was constructed, multi-level feature extraction was
performed, and multiple charging stations were predicted at the same time; Finally, taking
the cluster load of charging stations in an urban area as an example, a comparative
experiment was conducted to compare the model proposed in this study with the
mathematical model, the prediction performance of different variants of machine learning
models, common deep learning models and the model proposed in this study, and a
comparative test with multiple prediction horizons was conducted. The research results
show that the model proposed in this work improves the accuracy of multi-station
forecasting and provides new ideas for data-driven charging station cluster forecasting
research.
Keywords: Spatiotemporal forecasting of charging load; Attention mechanism; Graph neural
network; Maximum information coefficient; Spatiotemporal feature mining; Graph
convolutional neural network; Graph attention neural network; Multi-foresight prediction
Yong Liu, Chenyang Lu, Liang Li, Xiangchao Meng, Qiuping Jiang, Feng Shao,
Interactive feature fusion for camera-radar-based vehicle segmentation in bird’s-eye view,
Pattern Recognition,
Volume 172, Part D,
2026,
112698,
ISSN 0031-3203,
[Link]
([Link]
Abstract: Vehicle segmentation in Bird’s-Eye View (BEV) is a fundamental task for
autonomous driving, and integrating multi-modal sensory inputs, e.g., cameras and radars,
could enhance perception capability by leveraging their complementary strengths. However,
cross-modal feature fusion raises additional challenges due to the sparse and noisy
characteristics of radar data and the inherent misalignment between radar and camera
features. While existing fusion methods frequently leverage powerful attention mechanisms,
they often overlook the aforementioned heterogeneities and their impact on achieving
consistent, fine-grained alignment across modalities. We introduce the Interactively
Enhanced Camera-Radar Fusion (IECRF) framework, a novel approach that effectively bridges
cross-modal discrepancies in two stages through three new modules: Camera-Radar Feature
Aggregation (CRFA), Multi-Scale Radar Enhancer (MSRE), and Camera-Radar Feature Fusion
(CRFF). Specifically, the CRFA module explicitly models the complementary features of visual
appearance and radar geometry through two attention mechanisms, enabling fine-grained
alignment and interactive enhancement between the two modalities. The MSRE module
further refines radar representations through a modality-specific down- and up-sampling
design, amplifying salient targets while suppressing background noise in sparse radar
features. The aggregated features are then fused using the CRFF module at each stage for
lateral decoding. Extensive evaluations on the nuScenes dataset demonstrate that our IECRF
framework can operate with multiple backbones and configurations, achieving higher
vehicle segmentation accuracy even when using a lightweight EfficientNet backbone, which
is six times faster than the existing state-of-the-art approach equipped with an advanced ViT
backbone. The source code and trained models are available at
[Link]
Keywords: Vehicle segmentation; Bird’s-Eye View (BEV); Radar perception; Visual-radar
fusion; Autonomous driving
Chaofeng Huang, Xiaowo Xu, Fan Fan, Shunjun Wei, Xiaoling Zhang, Dongmei Liu, Min Gu,
A low-SNR-adaptive temporal network with smart mask attention for radar signal
modulation recognition,
Digital Signal Processing,
Volume 168, Part D,
2026,
105640,
ISSN 1051-2004,
[Link]
([Link]
Abstract: The automatic modulation recognition of radar signals is a key technology in
electronic warfare and communication systems. However, traditional handcrafted features
often struggle to achieve high recognition accuracy under low signal-to-noise ratio (SNR)
conditions. With the rapid development of artificial intelligence technologies, deep learning-
based approaches have emerged as a promising alternative for modulation recognition. In
this article, a low-SNR-adaptive network architecture is proposed, which integrates a
bidirectional temporal convolutional network (Bi-TCN) and dual-channel smart mask
attention (DSMA) modules. The DSMA adaptively highlights informative features and
suppresses noise through complementary attention masks, enhancing robustness in low-SNR
conditions. Experimental results demonstrate that the autocorrelation domain outperforms
both time and frequency domains, with recognition accuracy improvements of 13.33 % and
14.71 %, respectively. Compared to state-of-the-art models, the proposed network achieves
63 % accuracy at -20 dB and more than 99 % accuracy at -6 dB, significantly enhancing radar
signal modulation recognition.
Keywords: Modulation recognition; Deep learning; Radar signal analysis,
Xiao Cui, Baisheng Nie, Hengyi He, Peng Liu, Kaidan Bai, Haowen Zhou, Jingtao Yang,
Explainable deep learning for spatiotemporal high-temperature evolution and predictive
modeling in coal seam enhanced combustion,
Process Safety and Environmental Protection,
Volume 204,
2025,
108084,
ISSN 0957-5820,
[Link]
([Link]
Abstract: Temperature monitoring during deep coal seam combustion is essential for
optimizing in-situ heat extraction and ensuring operational safety. Accurate prediction of
temperature variations under enhanced combustion thus becomes a critical tool for
maintaining both safety and efficiency. In this study, continuous-ventilation coal combustion
experiments were performed to examine the spatiotemporal evolution of temperature, gas
emissions, and mass variation. The results showed that the high-temperature zone migrated
along the airflow direction, accompanied by pronounced spatiotemporal fluctuations in gas
concentrations. Using 13 input features—Coal Weight, Cross Section, Coal Quality Loss Rate,
CO/CO2, CO/H2, C3H8, C3H6, C2H4, C2H6, CH4, H2, CO2, and CO—predictive models for
enhanced combustion temperature were developed with multiple machine learning
methods. Eleven models were assessed, including LSTM, CNN-LSTM, CNN-LSTM with
Attention, BP, RNN, CNN, GRU, Transformer, RF, RBF, and XGBoost. Among them, the CNN-
LSTM-Attention model achieved the best performance, with an R2 of 0.987, MAE of 0.41,
RMSE of 0.59, and MAPE of 0.47—substantially outperforming the other ten models. To
enhance interpretability, SHapley Additive exPlanations (SHAP) were applied, revealing that
CO2 concentration had the strongest impact on prediction (mean SHAP value: 0.0555),
followed by the CO/CO2 ratio (0.0346) and Cross Section (0.024). Overall, this study
proposes a robust and interpretable approach for high-precision temperature prediction
during underground coal combustion, offering important guidance for thermal monitoring in
in-situ heat extraction systems.
Keywords: Enhanced coal combustion; High-temperature migration; Convolutional neural
network (CNN); Long short-term memory (LSTM); Attention mechanism; SHAP
interpretability
Hongchu Yu, Chenxi Jiang, Qinglong Fang, Tianming Wei, Lei Xu,
Deep learning driven spatiotemporal prediction of global carbon emissions from container
shipping,
Transportation Research Part D: Transport and Environment,
Volume 151,
2026,
105169,
ISSN 1361-9209,
[Link]
([Link]
Abstract: Container shipping is a significant source of global CO2 emissions, making accurate
predictions essential for meeting international environmental targets. This study proposes
ConvLSTM-CBAMNet, a deep learning model integrating channel and spatial attention
mechanisms to capture complex spatiotemporal emission trends. The model significantly
outperforms four deep learning baselines, including Transformer, ConvGRU, CNN-LSTM, and
the traditional ConvLSTM. Compared to the strongest baseline, ConvLSTM, it demonstrates
marked improvements, reducing the Root Mean Square Error (RMSE) by 19.4% to 0.0914
and the Mean Absolute Error (MAE) by 16.8% to 0.0432, while increasing the Structural
Similarity Index (SSIM) by 5.8% to 0.9035. These predictions can inform targeted
environmental policies, dynamically adjust Emission Control Areas (ECAs), optimize port
scheduling to mitigate pollution peaks, and develop environmental early-warning systems,
thereby supporting the shipping industry’s transition toward sustainability.
Keywords: Container ships; Carbon emissions; AIS data; Deep learning; Spatiotemporal
prediction
Jiaxin Qian, Jie Yang, Weidong Sun, Lingli Zhao, Lei Shi, Hongtao Shi, Chaoya Dang, Qi Dou,
Application potential and spatiotemporal uncertainty assessment of multi-layer soil moisture
estimation in different climate zones using multi-source data,
Journal of Hydrology,
Volume 645, Part B,
2024,
132229,
ISSN 0022-1694,
[Link]
([Link]
Abstract: Accurately estimating multi-layer soil moisture (SM) through remote sensing
methods presents inherent challenges and limitations. Multi-layer SM provides valuable
insights into the intricate interactions within the “soil-vegetation-atmosphere” system. This
study explored the temporal dynamics of multi-layer SM in the Shandian River Basin, China,
from 2019 to 2020. Through sensitivity analysis, we demonstrated the feasibility of using
multi-source data for estimating multi-layer SM, including dual polarization radar data,
optical vegetation descriptors, terrain factors, soil parameters, and meteorological indices.
Initially, surface soil moisture (SSM) at depths of 3 cm and 5 cm was estimated using the
modified change detection (MCD) model, which reduces the impact of vegetation.
Incorporating constraints from soil parameters during the solving process improved the
estimation accuracy of multi-layer SM. Subsequently, the water balance model, involving
precipitation and evaporation, was applied to further correct the estimation results of SSM.
Based on this, the infiltration process was considered to estimate deeper SM, including near-
surface soil moisture (NSSM) at depths of 10 cm and 20 cm, and root zone soil moisture
(RZSM) at depths of 40–50 cm. Under this framework, the estimation errors for multi-layer
SM were satisfactory (RMSE = 0.041–0.045 cm3/cm3). Finally, we explored the upper limits
of multi-layer SM estimation using multi-input and multi-output machine learning regression
(MLR) algorithms. With the incorporation of multi-source data, advanced MLR algorithms
achieved higher estimation accuracy (RMSE = 0.015–0.022 cm3/cm3) and showed potential
for cross-temporal transfer (RMSE = 0.030–0.037 cm3/cm3). Moreover, spatiotemporal
robustness revalidation of multi-layer SM was conducted across 17 observation networks
distributed cross different climatic zones in China. The results shown that the MCD model
achieved satisfactory results in estimating multi-layer SM (RMSE = 0.053–0.064 cm3/cm3),
whereas the regression models displayed higher accuracy (RMSE = 0.039–0.051 cm3/cm3).
Both the MCD and MLR models yielded similar conclusions, indicating that the estimation
accuracy of NSSM and RZSM surpassed that of SSM, primarily due to the relatively lower
variability of the former and their strong coupling with vegetation productivity. This study
also specifically discussed the influence of factors such as radar incidence angles, soil texture
types, and vegetation types on the estimation accuracy of multi-layer SM. This study
introduced a novel concept and framework for regional multi-layer and profile SM
estimation and real-time prediction through multi-source data, exhibiting high potential for
practical applications.
Keywords: Multi-layer soil moisture; Dual-polarization SAR data; Multi-source data
collaboration; Various climatic zones; Spatiotemporal estimation and prediction
Jongyun Byun, Jaehoon Cha, Jeyan Thiyagalingam, Hyeon-Joon Kim, Changhyun Jun,
Enhancing rainfall prediction accuracy through image fusion of radar and numerical weather
prediction models,
Expert Systems with Applications,
Volume 303,
2026,
130516,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Precipitation is one of the most challenging atmospheric phenomena to predict
due to the complexity involved in solving dynamic and thermodynamic atmospheric
equations. To address this challenge, extensive research has been conducted to enhance the
precision of numerical weather prediction models and radar-based extrapolation data, in
conjunction with the development of various blending techniques. However, traditional
methods have proven insufficient in capturing the diversity and nonlinearity of weather
phenomena. In response to these limitations, this study introduces a novel methodology
that leverages machine-learning based image fusion models to merge radar-based
extrapolation and numerical weather prediction rainfall datasets, thereby enhancing
prediction accuracy. An image fusion model was developed using radar-based extrapolation
data and numerical weather prediction data as input datasets, with radar observation data
utilized as target dataset. To identify the most suitable image fusion model for capturing the
complex patterns of rainfall data, two experiments were conducted: 1) Impact of model
topology, and 2) Effect of model size. A systematic analysis of the model outputs was
performed using eight evaluation metrics categorized under pixel-based metrics, feature-
based metrics, structural similarity metrics, and categorical verification metrics.
Experimental results indicated that image fusion model based on a Residual Network
(ResNet) outperformed other models in terms of model topology. Regarding model size, it
was observed that the performance did not increase proportionally with the number of
residual blocks; the most suitable performance was achieved with a specific number of
residual blocks (Case 5: 8 blocks). Additionally, the metrics compared with radar observation
data indicated that the proposed model delivered superior performance, thus offering a
high-accuracy rainfall prediction methodology.
Keywords: Image fusion; Radar; Numerical weather prediction; Deep learning; Precipitation
Cunyang Zhang, Yongmao Hou, Xiaohe Xia, Jin-Jian Chen, Yue Pan,
Multimodal feature fusion deep learning for spatiotemporal prediction of deformation and
environmental impacts in pipe-roof tunnel construction,
Advanced Engineering Informatics,
Volume 69, Part C,
2026,
104022,
ISSN 1474-0346,
[Link]
([Link]
Abstract: The pre-support tunnel construction involves complex construction conditions and
multisource data, posing challenges for efficient data transfer and accurate spatiotemporal
predictions. This study proposes an attention-based multimodal feature fusion deep learning
(AMFF-DL) framework that establishes a computational link between construction activities
and geotechnical responses. The AMFF-DL framework comprises two core components: a
multisource data preprocessing (MDP) module that systematically integrates on-site
information—including construction records, structural design parameters, and geological
surveys—into a unified, structured database, and an Attention-based Multimodal Feature
Fusion (AMFF) model that enables effective feature extraction, multimodal fusion, and
predictive learning. Applied to a real-world tunnel project in Shanghai, China, AMFF-DL
demonstrates strong predictive performance, achieving a mean absolute error (MAE) of
0.80 mm for deformation forecasts and a structural similarity index measure (SSIM) of
0.8936 for deformation cloud maps. It also accurately predicts key environmental indicators
such as pore water pressure, soil pressure, and horizontal displacement. Compared to
conventional prediction approaches, AMFF-DL performs credible data-transferring and
reliable predictions through its structured multimodal database and attention-based feature
fusion. Practically, AMFF-DL provides actionable insights into tunnel-induced impacts and
supports intelligent, data-informed risk management in complex underground construction
environments.
Keywords: Multimodal feature fusion; Deep learning; Spatiotemporal prediction;
Deformation and environmental impact; Pre-support tunnel construction
Xiaofei Zhang, Zhengping Fan, Xiaojun Tan, Qunming Liu, Yanli Shi,
Spatiotemporal adaptive attention 3D multiobject tracking for autonomous driving,
Knowledge-Based Systems,
Volume 267,
2023,
110442,
ISSN 0950-7051,
[Link]
([Link]
Abstract: Three-dimensional (3D) multiobject tracking (MOT) is an essential perception task
for autonomous vehicles (AVs). Studies have indicated that multimodal data fusion can
provide more stable and efficient perception information to AVs than a single sensor.
Therefore, this paper proposes a new spatiotemporal adaptive attention 3D (3DSTAA)
tracker, which attempts to improve the tracking performance of the end-to-end 3D MOT by
adaptively correlating spatiotemporal data. The novelty of this paper includes the following.
(1) Different from nonintelligent fusion methods, this paper uses an efficiently adaptive
spatial-guided fusion (SGFus) module for multimodal feature fusion. As a result, the 3D
structural information obtained from point cloud data can provide additional spatial
information as complementary information to the 2D texture information extracted from the
image data, collaboratively facilitating and refining the perception information
representation in the margin area. (2) This paper develops a spatiotemporal object-unique
attention (STOUA) module that calculates the relational degree of each perceived object
between two adjacent frames through attentional encoding. At the same time, an adaptive
weighting strategy is used to further study the spatiotemporal correlation of unique objects,
reducing the similarity among various objects and the differences across the same object.
Experiments tested using the KITTI tracking benchmark show that the 3DSTAA tracker is
highly competitive in both inference time and tracking performance compared with state-of-
the-art (SOTA) methods. Our corresponding code will be released on the
[Link]
Keywords: 3D multiobject tracking; Multimodal data fusion; Spatiotemporal attention
mechanism; Adaptive data association
Changlong Wang, Jiawei Jiang, Chong Han, Hengyi Ren, Lijuan Sun, Jian Guo,
Through-Wall Multihuman Activity Recognition Based on MIMO Radar,
Computers, Materials and Continua,
Volume 83, Issue 3,
2025,
Pages 4537-4550,
ISSN 1546-2218,
[Link]
([Link]
Abstract: Existing through-wall human activity recognition methods often rely on Doppler
information or reflective signal characteristics of the human body. However, static
individuals, lacking prominent motion features, do not generate Doppler information.
Moreover, radar signals experience significant attenuation due to absorption and scattering
effects as they penetrate walls, limiting recognition performance. To address these
challenges, this study proposes a novel through-wall human activity recognition method
based on MIMO radar. Utilizing a MIMO radar operating at 1–2 GHz, we capture activity data
of individuals through walls and process it into range-angle maps to represent activity
features. To tackle the issue of minimal variation in reflection areas caused by static
individuals, a multi-scale activity feature extraction module is designed, capable of extracting
effective features from radar signals across multiple scales. Simultaneously, a temporal
attention mechanism is employed to extract keyframe information from sequential signals,
focusing on critical moments of activity. Furthermore, this study introduces an activity
recognition network based on a Deformable Transformer, which efficiently extracts both
global and local features from radar signals, delivering precise human posture and activity
sequences. In experimental scenarios involving 24 cm-thick brick walls, the proposed
method achieves an impressive 97.1% accuracy in activity recognition classification.
Keywords: MIMO radar; human activity; Transformer; through-wall
Vipina Valsan, A.M. Abhishek Sai, Anu G Kumar, Aryadevi Remanidevi Devidas, Maneesha
Vinodini Ramesh, Kanakasabapathy P,
Deep learning-enabled spatiotemporal sustainable energy strategy for rural multiple
microgrids,
Results in Engineering,
Volume 28,
2025,
107569,
ISSN 2590-1230,
[Link]
([Link]
Abstract: Sustainable economic development entails effective energy management. Yet
there is a dearth of schemes that jointly capture spatial and temporal energy dynamics,
while ensuring fairness in underserved regions for geographically distant microgrids through
virtual energy sharing. This work proposes a spatiotemporal framework, for microgrid
energy management, stimulating exchange of excess renewable energy. The recommended
scheme facilitates intelligent utilization of renewable energy, with moderated costs of
generation and distribution. A simulated case study of two rural microgrid locations in India,
compliant with India's National Grid, validated the propriety of the proffered approach.
Findings showed a 30.7% reduction in dependency on the main grid, with drops in annual
energy imports from (19.1 - 13.2) MWh. Besides, multi-microgrid coordination delivered cost
savings exceeding INR 15,000, validated during monsoon and winter seasons. This study
corroborates the potential of coordinated microgrids to enhance resilience, inclusivity and
clean energy, in support of UN Sustainable Development Goals 7, 9, 11, 12, and 13.
Keywords: Deep learning; Multiple microgrids; Renewable energy; Spatiotemporal;
Sustainable energy sharing
Xiaofang Sun, Meng Wang, Junbang Wang, Guicai Li, Xuehui Hou,
Deep learning classification of winter wheat from Sentinel optical-radar image time series in
smallholder farming areas,
Advances in Space Research,
Volume 75, Issue 3,
2025,
Pages 2683-2695,
ISSN 0273-1177,
[Link]
([Link]
Abstract: As crop yield stagnation, climate change, and the rising demand for agricultural
products pose increasing challenges, mapping crop systems is becoming more and more
important. Winter wheat is one of the major cereal crops cultivated in China, ranking as the
third largest crop in terms of production and harvested area. Accurately mapping winter
wheat is necessary for implementing effective farm management practices. While many
studies have successfully produced high spatiotemporal resolution land cover maps,
relatively few map products of crop types are available in China. The growing archive of
satellite image time series provides enormous opportunities to map crops more closely. This
research presents a two-step method to map winter wheat based on Sentinel-1 and
Sentinel-2 time-series data from Shandong Province using the deep learning approaches.
The winter crops were firstly mapped using time-series optical vegetation indices employing
the deep learning methods. Then winter wheat was extracted from the winter crops mask by
coupling optical and synthetic aperture radar time-series images. The results indicated that
the precision of mapping winter wheat using Temporal Convolution Neural Networks
(TempCNN) achieved the highest precision in mapping winter wheat, with an overall
accuracy of 93.7 %, a kappa coefficient of 0.907, and an F1-score of 0.989. This was followed
sequentially by the Residual 1D convolutional neural networks (ResNet), the Multi-Layer
Perceptron (MLP), and the Lightweight Temporal Self-Attention Encoder (L-TAE). The
Temporal Attention Encoder (TAE) model demonstrated the lowest precision among the
compared models. The results agree well with independent county-level official census
winter wheat area data (R2 = 0.936). The proposed framework can also be applied in other
regions to generate maps of different crops, so future work can extend the proposed model
to other agricultural regions, where an increased number of crop types and natural
vegetation types can be included and tested.
Keywords: Sentinel-1; Sentinel-2; Classification; Winter wheat mapping; Time series; Deep
learning
Yehao Wang, Zijian Liu, Yingying Jin, Xiaoliang Wang, Lingyu Xu, Lei Wang, Jie Yu, Wenjuan
Dai, Jingxia Gao, Feng Zhang,
Interpreting spatiotemporal dynamics of Ulva prolifera blooms in the southern yellow sea
using an attention-enhanced transformer framework,
Environmental Pollution,
Volume 384,
2025,
126999,
ISSN 0269-7491,
[Link]
([Link]
Abstract: Harmful algal blooms dominated by Ulva prolifera have posed recurring ecological
and economic challenges in the southern Yellow Sea. To better understand and predict the
complex spatiotemporal dynamics of these blooms, we developed an enhanced
Transformer-based deep learning framework, incorporating multi-head self-attention
mechanisms. This model dynamically captures spatial dependencies, providing a
comprehensive understanding of bloom dynamics. Utilizing twelve key marine
environmental factors, we systematically explored all possible feature combinations to
determine the optimal predictive subset. Experimental results demonstrated superior
predictive performance of the model (MAE: 0.0213, MSE: 0.0016, R2: 0.9923) compared to
conventional deep learning models and recent spatiotemporal deep learning models.
Training dynamics revealed efficient convergence, especially with comprehensive
environmental information. Spatial attention analysis revealed that offshore regions
consistently received higher attention, indicating their critical role as informative and
generalizable environmental references. Furthermore, exhaustive feature attribution
experiments identified an optimal combination of eight environmental factors—including
temperature, salinity, current velocity, precipitation, wind direction, dissolved iron,
phosphate, and silicate—were found to significantly enhance prediction accuracy. This study
highlights the capability of attention-enhanced Transformer models for interpretable and
precise ecological forecasting, providing valuable insights for targeted mitigation and
management of U. prolifera blooms.
Keywords: Harmful algal blooms; U. prolifera; Deep learning; Spatiotemporal dynamics;
Environmental factors; Southern yellow sea
Jiabing Liu, Jianhao Sun, Haiwen Wei, Qilei Li, Junzhi Shi, Mingliang Gao,
Cloud prediction via spatiotemporal-frequency differential and attentional network,
Engineering Applications of Artificial Intelligence,
Volume 166, Part A,
2026,
113476,
ISSN 0952-1976,
[Link]
([Link]
Abstract: Cloud prediction is pivotal for meteorology, aviation safety, and renewable energy
management. A fundamental challenge in existing deep learning approaches lies in the
trade-off among prediction accuracy, computational efficiency, and long-term stability. To
bridge this gap, we introduce an end-to-end encoder–decoder architecture, termed
Spatiotemporal-Frequency Differential and Attentional Network (SFDANet). SFDANet
constructs an encoder–decoder architecture with a unique SFFE block, which integrates
spatiotemporal and frequency-domain analysis to simultaneously capture localized cloud
textures and global evolutionary dynamics. Between the encoder and decoder, an innovative
Multi-scale Differential Pyramid (MDP) module is built to selectively enhance high-frequency
details critical for rapid cloud evolution while inherently suppressing noise. To explicitly
model complex temporal dynamics, we propose a new module parallel to MDP, named
Multi-order Projection Attention (MPA). This module operates by projecting input features
into a set of parallel subspaces. Crucially, these subspaces are designed to be both linear and
non-linear. Through this architectural design, the module is capable of simultaneously
capturing predictable low-order trends and intricate high-order patterns within the data.
Comprehensive experiments on the WeatherBench dataset demonstrate that SFDANet
achieves superior accuracy and long-term stability, while it maintains remarkable efficiency
with only 0.65M parameters. The code is available at [Link]
Keywords: Cloud prediction; Spatiotemporal-frequency; Differential pyramid; Multi-order
projection
Bao-Lin Ye, Peng Wu, Lingxi Li, Weimin Wu, Bo Song, Xianchao Zhang,
Multi-intersection traffic signal control based on dynamic spatiotemporal memory enhanced
learning,
Control Engineering Practice,
Volume 165,
2025,
106606,
ISSN 0967-0661,
[Link]
([Link]
Abstract: In multi-intersection traffic signal control, spatial information contains rich traffic
state features, including intersection topology and lane associations. Effectively extracting
and integrating this spatial information is crucial for accurately characterizing the evolution
of traffic state. However, most existing methods rely on static graph structures and,
therefore, cannot dynamically model the spatiotemporal correlations of traffic states,
limiting their adaptability to real-time traffic scenarios. To address these limitations, we
propose a multi-intersection traffic signal control method based on dynamic spatiotemporal
memory enhanced learning (DSMEL). First, we develop an adaptive update mechanism for
dynamic heterogeneous graphs that analyzes correlations among heterogeneous traffic
features in real time to generate adaptive representations of spatial relationships. Second,
we introduce a dual-memory enhancement model based on spatiotemporal decoupling that
uses a multi-head attention mechanism to process spatial and temporal features separately.
Specifically, we design a temporal memory module to model periodic temporal dynamics
within the road network and a spatial memory module to track the evolving topological
relationships among traffic nodes. This enables a fine-grained capture of dynamic
spatiotemporal features within the road network. Finally, we propose an adaptive weight
learning method based on double entropy regularization that incorporates online learning of
dynamic game weight matrices and integrates entropy-constrained policy optimization with
novel reward and loss functions to enhance system stability and promote optimal
convergence in multi-agent coordination. Extensive experiments on synthetic and real-world
scenarios show that, compared with baseline methods, DSMEL reduces queue length by
27.19% to 49.89%, occupancy rate by 10.93% to 43.94%, and vehicle count by 11.39% to
43.68%. Furthermore, DSMEL demonstrated superior performance over both traditional
traffic signal control methods and reinforcement learning-based approaches in extreme
traffic scenarios, reducing the average queue length by 21.23% and the maximum queue
length by 18.20%.
Keywords: Deep reinforcement learning; Traffic signal control; Multi-agent; Dynamic graph
Pengfei Wang, Peilin Shu, MingHao Yang, Hongqiu Zhang, Jianqi Wang, Cong Wang, Hongbo
Jia,
Dual-task physiological learning for radar-based continuous blood pressure monitoring:
classification-regularized regression,
Biomedical Signal Processing and Control,
Volume 112, Part C,
2026,
108790,
ISSN 1746-8094,
[Link]
([Link]
Abstract: Continuous blood pressure (BP) monitoring is critical for hypertension
management, yet conventional non-contact radar systems face challenges such as feature
space overlap under low signal-to-noise ratio conditions and inaccurate estimation during
physiological state transitions due to the susceptibility of millimeter-level cardiovascular
vibration signals to environmental interference. To address these challenges, we propose a
dual-task learning model integrating classification-constrained regression with multi-scale
spatiotemporal feature extraction. Our framework combines: (1) A hybrid ResNet-BiGRU
backbone capturing waveform morphology and hemodynamic continuity through multi-
scale convolutions (kernels = 15/7/3) and triple-layer bidirectional gating, enabling robust 2-
second-interval predictions; (2) A physiological regularization mechanism where
classification-derived BP-range probabilities (10 mmHg bins) constrain regression outputs,
suppressing implausible fluctuations during dynamic states. Validated on 30 subjects across
resting, Valsalva, and tilt-table tests, results indicate clinically relevant accuracy (SBP:
−0.21 ± 6.74 mmHg; DBP: 0.25 ± 4.81 mmHg) at 0.5 Hz sampling rate, while demonstrating
improved dynamic-state performance versus benchmarks with DBP RMSE reductions up to
8.4 % by dual-task strategy. This work suggests dual-task learning can mitigate radar-specific
SNR constraints and physiological nonstationarity while fulfilling clinical real-time monitoring
demands (beat-to-beat resolution), contributing to practical deployment of cuffless BP
devices.
Keywords: Blood pressure; Radar; Dual-task learning; Non-contact monitoring; ResNet;
Feature combination
Hongjin Chen, Kanghui Zhou, Zhonghua Zheng, Lei Han, Yongguang Zheng,
TorViNet: A spatiotemporal deep learning network for tornado detection in user-captured
social media videos,
Expert Systems with Applications,
Volume 308,
2026,
131093,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Tornadoes are extremely destructive yet short-lived severe weather events, and
their rapid development often escapes timely confirmation by conventional meteorological
instruments. In recent years, the widespread availability of smartphones and the
proliferation of social media platforms have enabled user-captured videos to become a
valuable supplementary source for real-time severe weather monitoring. Weather radar may
indicate tornadic signatures but often cannot verify whether a tornado has touched down. In
contrast, social-media videos provide direct visual evidence of actual occurrence, making
them valuable for ground-truth validation and situational awareness. Accordingly, this study
presents TorViNet, an AI-driven spatiotemporal recognition framework designed as a
decision-support module for expert systems to detect tornadoes directly from user-captured
social media videos. To address real-world challenges such as redundant or irrelevant
frames, dynamic occlusions, and small-scale visual targets, TorViNet integrates three
domain-adapted components: a dynamic frame-level selection strategy that filters out
temporally uninformative content, a spatial-frequency attention mechanism that enhances
fine-grained vortex structures, and a contrast-aware refinement module that suppresses
background distractions. Experiments on a newly curated dataset of over 10,000 verified
tornado and non-tornado clips demonstrate that TorViNet achieves 91% accuracy and an F1-
score of 0.89, outperforming a wide range of dominant video classification models. Its
robustness under noisy, unstable, and far-visibility conditions highlights its potential for
integration into operational meteorological expert systems, providing timely situational
awareness and enhancing early warning capabilities for severe tornado events.
Keywords: Tornado detection; Social media; Meteorology systems; Deep learning; Attention
mechanism
Yong He, Zi-Long Duan, Xiang-Hong Ding, Zhao Zhang, Raud Eucaristia Mayoulou, Kao-Fei
Zhu,
Spatiotemporal prediction for groundwater heavy metal contamination using Soft-DTW-
based clustering and graph neural network framework,
Water Research,
Volume 291,
2026,
125245,
ISSN 0043-1354,
[Link]
([Link]
Abstract: Accurate prediction of groundwater heavy metal contaminant spatiotemporal
dynamics is essential for monitoring optimization and remediation decision-making at
contaminated sites. However, heterogeneous contamination distribution and complex
spatiotemporal correlations among monitoring wells pose significant prediction challenges.
In this study, a Soft Dynamic Time Warping clustering-based Graph Neural Network
(SDCGNN) was proposed for contamination zone identification, integrating multi-scale
spatiotemporal modeling. Using Soft Dynamic Time Warping (Soft-DTW) distance-based
clustering, the model partitions monitoring wells into source, plume, and attenuation zones
according to temporal contamination patterns, while a hierarchical local-global graph fusion
framework captures both zone-specific dynamics and cross-zone transport processes.
Evaluation on two-year hourly monitoring data from 25 monitoring wells at on-site
contaminated sites demonstrated that SDCGNN achieved average Mean Absolute Error
(MAE) of 0.213 mg/L and Mean Absolute Percentage Error (MAPE) of 5.51%, improving upon
baseline models by 49.9% and 61.4%, respectively. Zone-specific predictions revealed high
accuracy across source, plume, and attenuation areas despite varying concentration ranges
and spatial heterogeneity. Furthermore, spatiotemporal analysis confirmed that the model
accurately reproduced observed contamination transport patterns, including plume
migration directions and concentration gradient evolution. The proposed zone-aware
modeling approach shows promise for advancing groundwater heavy metal contamination
prediction capabilities and facilitates the optimization of site remediation strategies.
Keywords: Heavy metal contamination; Dynamic time warping; Graph clustering; Graph
neural networks
Wei Tian, Lei Yi, Xianghua Niu, Rong Fang, Lixia Zhang, Huanhuan Liu, Zhuo Xu, Shengqin
Jiang, Yonghong Zhang,
RadarNet: A parallel spatiotemporal encoder network for radar extrapolation,
Neurocomputing,
Volume 591,
2024,
127665,
ISSN 0925-2312,
[Link]
([Link]
Abstract: Radar extrapolation has been one of the most important means for nowcasting.
Most current models achieve good performance in high-frequency sequences (e.g., video,
more than 24 fps), while the temporal resolution of radar echo sequences is much lower (1
frame every 6 min) and the transforms are much more complex. The spatiotemporal
characters with some similarities would not change a lot in video sequences; however, the
radar echo sequences include more intangible changes (e.g., the echo evolution of
generation or vanish, and so on), which leads to unique distinct spatial and temporal
characters, respectively. Therefore, the singular peculiarity would be mitigated, leading to a
rapid decline in precision and sharpness during the extrapolation process. In general,
temporal feature extraction is utilized to understand the variation in pixel locations, while
spatial feature extraction is employed to capture the distribution variation of specific
regions. In this work, we propose a feature decomposition network, termed as RadarNet to
improve the extrapolation precision. The parallel independent encoders are used to enhance
multi-scale spatial feature extraction and temporal motion feature capture of radar echoes,
respectively. In addition, we design a specialized cross fusion mechanism to achieve the
inputs of the decoder which may enhance the performance of the extrapolation. The
extrapolation experiments are conducted on real radar echo datasets from Shijiazhuang and
Nanjing that demonstrate the effectiveness of our model.
Keywords: RadarNet; Radar extrapolation; Spatiotemporal prediction
Nana Chu, Kam K.H. Ng, Xinting Zhu, Ye Liu, Lishuai Li, Kai Kwong Hon,
Towards dynamic flight separation in final approach: A hybrid attention-based deep learning
framework for long-term spatiotemporal wake vortex prediction,
Transportation Research Part C: Emerging Technologies,
Volume 169,
2024,
104876,
ISSN 0968-090X,
[Link]
([Link]
Abstract: The conservative and distance-based static wake vortex-related separation may
restrict runway operational efficiency. Recent studies have demonstrated the potential of
wake separation reduction under the Re-categorisation scheme of Aircraft Weight (RECAT).
Furthermore, dynamic time-based flight separation considering vortex evolution with
respect to aircraft pairs and meteorological conditions will be the ultimate objective for
improving runway operational capacity without compromising safety. This paper presents a
hybrid deep learning framework for aircraft wake vortex recognition, evolution prediction,
and preliminary dynamic separation assessment in the final approach. Two-stage Deep
Convolutional Neural Networks (DCNNs) are utilised to identify vortex locations and strength
from wake images. Subsequently, we propose the Attention-based Temporal Convolutional
Networks (ATCNs) for future long-term vortex decay and transport forecasts based on initial
vortex information from DCNNs. 17,254 wake sequences generated by arrival flights at Hong
Kong International Airport (HKIA) are used in this study. The proposed ATCN models
outperform the specific benchmarks. Furthermore, the hybrid DCNN-ATCN model shows
great benefits in mining both spatial vortex characteristics and temporal dependencies in
vortex evolution, and achieves a computational speed of approximately 7 s per sequence.
The final vortex duration assessment demonstrates a significant potential for separation
reduction in the final approach when the crosswind speed exceeds 3 m/s. This study
provides important implications for online and fast-time wake behaviour monitoring and
state estimation. The results of vortex duration analysis conform to the RECAT-EU standards
and present an efficient strategy for developing dynamic flight separation systems.
Keywords: Flight separation; Aircraft wake turbulence; Recurrent neural network; Attention
mechanism; LiDAR
Wei Zhang, Xinyu Zhang, Junyu Dong, Xiaojiang Song, Renbo Pang,
CIDM: A comprehensive inpainting diffusion model for missing weather radar data with
knowledge guidance,
ISPRS Journal of Photogrammetry and Remote Sensing,
Volume 221,
2025,
Pages 299-309,
ISSN 0924-2716,
[Link]
([Link]
Abstract: Addressing data gaps in meteorological radar scan regions remains a significant
challenge. Existing radar data recovery methods tend to perform poorly under different
types of missing data scenarios, often due to over-smoothing. The actual scenarios
represented by radar data are complex and diverse, making it difficult to simulate missing
data. Recent developments in generative models have yielded new solutions for the problem
of missing data in complex scenarios. Here, we propose a comprehensive inpainting
diffusion model (CIDM) for weather radar data, which improves the sampling approach of
the original diffusion model. This method utilises prior knowledge from known regions to
guide the generation of missing information. The CIDM formalises domain knowledge into
generative models, treating the problem of weather radar completion as a generative task,
eliminating the need for complex data preprocessing. During the inference phase, prior
knowledge of known regions guides the process and incorporates domain knowledge
learned by the model to generate information for missing regions, thus supporting radar
data recovery in scenarios with arbitrary missing data. Experiments were conducted on
various missing data scenarios using Multi-Radar/MultiSensor System data sourced from the
National Oceanic and Atmospheric Administration, and the results were compared with
those of traditional and deep learning radar restoration methods. Compared with these
methods, the CIDM demonstrated superior recovery performance for various missing data
scenarios, particularly those with extreme amounts of missing data, in which the restoration
accuracy was improved by 5%–35%. These results indicate the significant potential of the
CIDM for quantitative applications. The proposed method showcases the capability of
generative models in creating fine-grained data for remote sensing applications.
Keywords: Weather radar data; Comprehensive inpainting; Diffusion models; Extreme
missing cases; Knowledge guidance
Qiangyu Zeng, Ling Li, Hao Wang, Jianxin He, Hua Wang, Yao Gao,
MCDA-UNet: A satellite data-based model for radar composite reflectivity retrieval,
Atmospheric Research,
Volume 330,
2026,
108619,
ISSN 0169-8095,
[Link]
([Link]
Abstract: The weather radar network in China exhibits an uneven spatial distribution, with
dense coverage in the eastern regions and sparse deployment in the west, resulting in
substantial detection blind spots in areas with complex terrain. This severely limits the
continuity and precision of weather monitoring and early warning in these regions. To
address this challenge, a multi-channel deep learning model, MCDA-UNet, is proposed for
radar composite reflectivity retrieval, aiming to reconstruct and enhance radar echo patterns
in regions lacking radar coverage by leveraging the extensive spatial coverage and
continuous observation capabilities of geostationary meteorological satellites. The model
employs a multi-channel input architecture to extract features from different spectral bands,
while spatial and channel attention models are incorporated to improve the representation
of key meteorological information, thereby enhancing retrieval accuracy and regional
adaptability. Comparative experiments conducted under varying precipitation intensities
demonstrate that MCDA-UNet consistently outperforms existing models across multiple
evaluation metrics, particularly in reconstructing weather radar echo structures and edge
details. These results validate the model’s capability to adapt to the full dynamic range of
weather radar reflectivity and highlight its potential for accurate precipitation retrieval in
radar blind-spot regions.
Keywords: Weather radar composite reflectivity; Satellite data retrieval; Multi-channel
structure; Full dynamic range
Pengfei Jia, Helmi Zulhaidi Mohd Shafri, Shengrui Yu, Zhi Zheng, Shiqing You, Abdul Rashid
Mohamed Shariff,
Radar-optical fusion of Sentinel-1/2 for high-resolution NDVI reconstruction and landscape-
driven carbon flux assessment in Kuala Selangor, Malaysia (2020–2024),
International Journal of Applied Earth Observation and Geoinformation,
Volume 145,
2025,
104966,
ISSN 1569-8432,
[Link]
([Link]
Abstract: Reliable quantification of carbon fluxes in humid tropical regions is constrained by
persistent cloud cover, heterogeneous land mosaics, and the limited resolution of existing
products. To address these challenges, this study developed a Cloud-Resilient Fusion
Network (CRFNet) that integrates Sentinel-1 SAR backscatter with cloud-screened Sentinel-2
NDVI using a CNN–BiLSTM–multi-head attention architecture. The framework reconstructed
10 m NDVI time series in Kuala Selangor, Malaysia (2020–2024), achieving annual R2 above
0.82 and RMSE below 0.12, thereby improving temporal continuity under heavy cloud–
rainfall interference. The reconstructed NDVI was used to drive a light-use-efficiency model
for net ecosystem productivity (NEP) estimation, supported by temperature-based
heterotrophic respiration. Results showed a 7.4 % decline in mean annual NEP across five
years, with degraded mangroves and sloping croplands emerging as hotspots of sink-to-
source transitions. Landscape analysis revealed strong structure–function coupling: stable
forests and mangroves were characterized by large cohesive patches with largest patch index
values above 40 % and edge density below 20 m ha-1, while croplands and degraded slopes
exhibited higher patch numbers, reduced patch dominance, and greater edge complexity,
which increased carbon source risk. By linking fine-scale NDVI reconstruction with process-
based carbon modeling and landscape metrics, this study provides a transferable workflow
for high-resolution carbon flux monitoring and a robust scientific basis for carbon budget
assessment, ecosystem management, and carbon-neutrality planning in tropical monsoon
regions.
Keywords: Cloud-resilient fusion network (CRFNet); NDVI reconstruction; Radar–optical
fusion; Net ecosystem productivity; Landscape metrics
Weidong Fang, Xibin Lin, Ji Zhang, Jiacheng Hu, Linrun Huang, Guangqian Yuan,
Spatiotemporal charging demand forecasting for EV stations via cross-attention fusion,
Applied Soft Computing,
Volume 189,
2026,
114475,
ISSN 1568-4946,
[Link]
([Link]
Abstract: With the rapid growth of electric vehicles (EVs), accurately predicting charging
demand has become crucial for intelligent transportation and energy management. Existing
deep learning methods usually neglect the integration of economic principles with complex
spatiotemporal dependencies. To this end, this paper proposes a novel framework, the Bi-
CAPNet model, which integrates a bidirectional temporal convolutional network (BiTCN), a
bidirectional gated recurrent unit (BiGRU), a crisscross attention mechanism, an economics-
informed neural network (EINN), and a sparrow search algorithm (SSA). This architecture
captures multi-scale temporal features, models sequential dependencies, integrates
spatiotemporal data, and incorporates price–demand elasticity to characterize the response
of charging demand to dynamic pricing, thereby improving economic interpretability.
Utilizing extensive real-world data collected from Shenzhen, the empirical evaluation
confirms that the proposed approach substantially outperforms representative baseline
models in terms of both predictive accuracy and economic interpretability, indicating its
strong potential for urban EV charging demand forecasting.
Keywords: Charging demand forecasting; Economics-informed neural networks;
Spatiotemporal feature fusion; Cross-attention mechanism
Cries Avian, Jenq-Shiou Leu, Hang Song, Jun-ichi Takada, Nur Achmad Sulistyo Putro,
Muhammad Izzuddin Mahali, Setya Widyawan Prakosa,
RCTrans-Net: A spatiotemporal model for fast-time human detection behind walls using
ultrawideband radar,
Computers and Electrical Engineering,
Volume 120, Part C,
2024,
109873,
ISSN 0045-7906,
[Link]
([Link]
Abstract: Ultrawideband (UWB) radar systems are becoming increasingly popular for
detecting human presence, even through walls. Recent advancements in signal processing
use deep learning techniques, which are known for their accuracy. While earlier methods
focused on spatial information using Convolutional Neural Networks (CNNs), newer research
highlights the importance of temporal information, such as how data peaks shift over time.
This study introduces RCTrans-Net, a deep-learning architecture that combines RCNet (a
Residual CNN) for spatial features with TransNet (a Transformer) for temporal features. This
fusion improves human presence classification in fast-time signal processing. Tested under
various conditions—different materials, body orientations, ranges, and radar heights—
RCTrans-Net achieved high performance with F1-scores of 0.997±0.000 for static,
0.967±0.004 for dynamic, and 0.978±0.001 for combined scenarios. The architecture
outperforms previous methods and offers real-time processing with an inference time of
about one millisecond.
Keywords: Human presence behind the wall; Residual network; Spatiotemporal'
Transformer; Ultrawideband radar system
Pengfei Yang, Feng Wu, Minyang Liu, Ting Zhong, Fan Zhou,
Beyond pillars: Advancing 3D object detection with salient voxel enhancement of liDAR-4D
radar fusion,
Pattern Recognition,
Volume 173,
2026,
112841,
ISSN 0031-3203,
[Link]
([Link]
Abstract: The fusion of LiDAR and 4D radar has emerged as a promising solution for robust
and accurate 3D object detection in complex and adverse conditions. Existing methods
typically rely on pillar-based representations, which, although computationally efficient, fail
to provide fine-grained structural details necessary for precise object localization and
recognition. In contrast, voxel-based representations offer richer spatial information but face
challenges such as background noise and data quality disparity. To address these limitations,
we propose SVEFusion, a voxel-based 3D object detection framework that integrates LiDAR
and 4D radar data using a salient voxel enhancement mechanism. Our method introduces an
adaptive feature alignment module and a novel spatial neighborhood attention module for
efficient early-stage multi-modal voxel feature integration. Furthermore, we design a salient
voxel enhancement mechanism that assigns higher weights to foreground voxels using a
multi-scale weight prediction strategy, progressively refining weight accuracy with
supervision loss. Experimental results demonstrate that SVEFusion significantly outperforms
state-of-the-art methods, establishing a new benchmark in multi-modal 3D object detection.
The source code and network weighting for reproducibility are available at
[Link]
Keywords: Object detection; Lidar; 4D Radar; Multi-modal fusion; Autonomous driving
Jinbo Fu, Hong Cao, Zhe Wang, Kuan Chang, Haitao Wang, Bo Chen, Jiuchun Sun,
WT-DANet-STAF: A spatiotemporal adaptive fusion-based denoising method for foundation
pit enclosure structure deformation data,
Measurement,
Volume 263,
2026,
120155,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Monitoring deformation in foundation pit enclosure structures is crucial, serving as
the foundation for predicting deformations, issuing early warnings for exceedances, and
implementing effective control measures. However, the collected monitoring data often
contain noise, necessitating robust denoising techniques to ensure accuracy. Existing
denoising methods for foundation pit monitoring data often struggle to effectively remove
noise while preserving key features. To address these challenges, a novel method based on
Wavelet transform (WT) and DenseNet-Attention (DANet) is proposed for denoising
deformation data of foundation pit enclosure structures, complemented by a spatiotemporal
adaptive fusion network (STAF) to optimize the results further. Initially, Wavelet transform is
applied to decompose the deformation data across spatial and temporal dimensions at
multiple scales, extracting detailed and approximation components at varying frequency
levels. Subsequently, DANet is employed to denoise and reconstruct these decomposed
components, yielding denoised data in both spatial and temporal dimensions. Finally, the
spatiotemporal adaptive fusion network integrates the denoised outputs from spatial and
temporal dimensions to generate high-quality results. The effectiveness of the proposed
method is validated through experiments using data from a road improvement project in
southern China. Comparisons with mainstream denoising techniques and evaluations via
multiple metrics demonstrate the superior performance of the proposed approach in noise
reduction.
Keywords: Foundation pit; Deformation monitoring; Wavelet transform; DenseNet-Attention
network; Spatiotemporal fusion denoising
Sasan Babaee, Mohammad Amin Khalili, Rita Chirico, Anna Sorrentino, Diego Di Martire,
Spatiotemporal characterization of the subsidence and change detection in Tehran plain
(Iran) using InSAR observations and Landsat 8 satellite imagery,
Remote Sensing Applications: Society and Environment,
Volume 36,
2024,
101290,
ISSN 2352-9385,
[Link]
([Link]
Abstract: Urban areas worldwide are increasingly facing challenges related to land
subsidence, a phenomenon exacerbated by uncontrolled groundwater extraction and urban
expansion. This research focuses on the Tehran plain, Iran's capital city, where significant
subsidence has been observed due to uncontrolled migrations influenced by various
economic and political factors. This expansion has increased demand for energy, notably
water, leading to irregular water withdrawals from underground sources and, consequently,
land subsidence. Monitoring this subsidence, particularly its effects on urban infrastructure,
has become a critical challenge. This research first reviewed the existing body of knowledge
related to subsidence measurement in the Tehran plain with an emphasis on their findings
and limitations and then used radar images to study the subsidence patterns in the Tehran
plain from 2016 to the end of 2020. Finally, the results collaborated by optical imagery
analysis to find the relationship between surface change detection and spatiotemporal
distribution of subsidence. As a result, through processing Sentinel-1A SAR images,
consistent vertical displacements (subsidence) were observed, especially in areas heavily
reliant on groundwater from wells, with some areas experiencing a rate of more than
−20 mm/year. Horizontal displacement, however, was approximately about ±8 mm/year.
Also, our results show that the subsidence rate in this plain has decreased in recent years.
Therefore, the study integrated multispectral satellite data to clarify this issue and
compensate for missing groundwater level data, specifically the Normalized-Difference
Vegetation Index (NDVI) and Normalized-Difference Moisture Index (NDMI). These datasets
were used to monitor changes in vegetation cover distribution and moisture in response to
the variations of groundwater depth over time. The results of this research can be beneficial
in adequately managing groundwater resource utilization to reduce the potential damage to
infrastructure and the environment.
Keywords: Spatiotemporal subsidence pattern; Radar interferometry; Sentinel-1A; Tehran
plain
Yongchao Zhu, Qiuling Lu, Maorong Ge, Xiaochuan Qu, Tingye Tao, Kegen Yu, Shuiping Li,
Attention enhanced ResNet for ocean surface wind speed retrieval using CYGNSS
observables,
Advances in Space Research,
2025,
,
ISSN 0273-1177,
[Link]
([Link]
Abstract: Global Navigation Satellite System Reflectometry (GNSS-R) has emerged as a
pivotal technique for ocean surface wind speed retrieval; however, establishing robust multi-
parameter retrieval models remains challenging due to the nonlinear relationships between
GNSS-R observables and geophysical variables. An Attention-enhanced Residual Network
(Att-ResNet) is proposed to address this challenge, leveraging Cyclone Global Navigation
Satellite System (CYGNSS) bistatic radar data for wind speed estimation. The CYGNSS
datasets were processed to extract multi-parameter observables, including Delay-Doppler
Maps (DDMs), normalized bistatic radar cross-section (NBRCS), and incidence angle, which
served as inputs for training wind speed retrieval models using diverse backbone
architectures (e.g., ResNet and AlexNet). Ablation experiments employing the Att-ResNet
framework were systematically conducted, with ERA5 (European Centre for Medium-Range
Weather Forecasts Reanalysis 5) and CCMP (Cross-Calibrated Multi-Platform) wind products
providing benchmark validation. Comparative analysis revealed that the Att-ResNet-retrieved
wind speeds exhibited strong spatiotemporal consistency with ERA5 and CCMP data.
Quantitative evaluations showed root mean square errors (RMSEs) of 1.379 m/s (ERA5) and
1.390 m/s (CCMP), with minimal biases (−0.069 m/s and −0.014 m/s, respectively) and
unbiased RMSEs (ubRMSEs) of 1.377 m/s and 1.390 m/s. The study demonstrates that the
Att-ResNet architecture, through its attention-driven feature selection and residual learning
mechanisms, significantly enhances spaceborne GNSS-R wind retrieval accuracy. This
artificial intelligence-driven framework establishes a new paradigm for high-resolution
spatiotemporal ocean surface wind monitoring, demonstrating the transformative potential
of deep learning in advancing GNSS-R applications.
Keywords: Residual network; GNSS-R; Wind speed; Deep learning; CYGNSS
Jiaying Li, Weidong Wang, Guangqi Chen, Zheng Han, Chongzheng Zhu, Chen Chen,
Spatiotemporal LSA modeling incorporating comprehensively the momentary effects of
rainfall and earthquake: A case study of the Liangshan Prefecture, China,
Advances in Space Research,
Volume 76, Issue 11,
2025,
Pages 6725-6740,
ISSN 0273-1177,
[Link]
([Link]
Abstract: Landslides are one of the most destructive geo-hazards, and the landslide
susceptibility assessment (LSA) can effectively reduce landslide risks and strengthen
landslide prevention. The present study explores a spatiotemporal LSA method considering
comprehensively the momentary effects of rainfall and earthquakes. Logistic regression
model, random forest model, deep belief network (DBN) model, and grey wolf optimizer
(GWO)-DBN model were used to analyze the spatial LSA, and the optimal spatial LSA
obtained using the GWO-DBN model was chosen using various evaluation metrics to analyze
the spatiotemporal LSA. Meanwhile, the historical landslide data during the year before the
study time, namely from July 5, 2020 to July 5, 2021, and the data of rainfall and earthquake
before various landslides were collected, and their effective rainfall and seismic peak ground
acceleration were calculated to construct the temporal LSA regression model. The temporal
LSA map in the study time was thus obtained and coupled with the optimal spatial LSA map
to generate the spatiotemporal LSA map. Due to dynamic changes over time of
spatiotemporal LSA, the precise landslide locations and ranges were obtained using small
baseline subset interferometric synthetic aperture radar, and the results were coupled with
spatiotemporal LSA map. There were 86.92% landslide regions with very high and high
susceptibility, and the accuracy of spatiotemporal LSA was verified, which provides a
reference for the spatiotemporal LSA verification method.
Keywords: Spatiotemporal LSA; Deep belief network model; GWO-DBN model; Temporal LSA
regression model; SBAS-InSAR
Saihan Chen, Peng Liu, Puchen Zhang, Xiaokang Ma, Ran Bao, Zixu Wang, Haixu Yang, Xiao
Ke,
Data-driven spatiotemporal fault detection in Lithium-ion batteries using isometric mapping
and modified independent component analysis,
Journal of Energy Storage,
Volume 149,
2026,
120062,
ISSN 2352-152X,
[Link]
([Link]
Abstract: Accurate detection and localization of thermal faults in lithium-ion batteries (LIBs)
are crucial for ensuring safety and preventing accidents. However, the intricate
thermodynamic behavior of large-format LIBs and battery systems, which operate as high-
dimensional distributed-parameter systems, poses significant challenges. This investigation
presents a data-driven spatiotemporal framework for detecting and localizing battery
thermal faults. Isometric mapping models the spatiotemporal dynamics of the battery's
thermal processes by decomposing high-dimensional spatiotemporal temperature data into
linear combinations of low-dimensional temporal coefficients and discrete spatial basis
functions (SBFs), with radial basis functions used to construct continuous-space SBFs.
Modified independent component analysis is further applied to temporal coefficients to
build process monitoring models and generate real-time monitoring statistics for fault
detection. Finally, spatiotemporal reconstruction yields a continuous-space contribution map
of abnormal statistics for fault localization. The proposed method is validated through
internal short circuit experiments on large-format LIBs and thermal runaway propagation
simulations of battery systems, covering 33 thermal fault scenarios under various operating
conditions. Results indicate that the method achieves high-precision fault detection and
localization using only six-dimensional temporal coefficients and corresponding SBFs, with an
average F1 score of 98.1 %, a fault detection delay of 3.4 sampling steps, and a fault
localization accuracy of 93.3 %.
Keywords: Battery thermal process; Fault detection; Fault localization; Spatial construction;
Internal short circuit; Thermal runaway
Mingyue Lu, Chuanwei Jin, Manzhu Yu, Qian Zhang, Hui Liu, Zhiyu Huang, Tongtong Dong,
MCGLN: A multimodal ConvLSTM-GAN framework for lightning nowcasting utilizing multi-
source spatiotemporal data,
Atmospheric Research,
Volume 297,
2024,
107093,
ISSN 0169-8095,
[Link]
([Link]
Abstract: Lightning phenomena can instigate a cascade of calamities, encompassing fires,
electrical infrastructure damage, and risks to human safety. Deep-learning-based lightning
nowcasting models have demonstrated significant effectiveness in disaster prevention and
mitigation. However, existing studies often neglect the impacts of surface features on
lightning activities, and conventional lightning prediction techniques based on convolutional
and recurrent networks face challenges such as the loss of feature information. Addressing
these issues, this paper presents a novel model for lightning nowcasting, the Multimodal
ConvLSTM-GAN for Lightning Nowcasting (MCGLN). This model integrates a Generative
Adversarial Network (GAN) with a Convolutional Long Short-Term Memory network
(ConvLSTM), utilizing multi-source data as inputs. It incorporates a spatiotemporal encoder-
forecaster framework within the Generator to improve the capture of multidimensional
spatiotemporal feature information, thus boosting predictive accuracy. MCGLN offers
probabilistic prediction results, allowing users to customize warning thresholds following
their specific tolerance for false and missed alarms. The performance of the MCGLN model is
evaluated through empirical analysis, utilizing real lightning datasets sourced from Zhejiang
and surrounding areas. Experimental results demonstrate that: (a) The MCGLN model
outperforms existing methods in terms of detection capability and overall performance,
showing significant improvements in the modeling process. (b) Increasing the number of
data sources improves detection capabilities, reduces the probability of false alarms, and
boosts the model performance. (c) The use of radar data enhances the recognition of high-
probability lightning occurrences, and the inclusion of surface feature data increases the
capture of terrestrial lightning genesis.
Keywords: Lightning prediction; MCGLN; Lightning nowcasting model; Ningbo
Yixuan Liu, Alim Samat, Peijun Du, Jin Chen, Jilili Abuduwaili, Kaiyue Luo, Enzhao Zhu, Dana
Shokparova,
High-resolution spatiotemporal analysis and driver attribution of floods in Kazakhstan using
SHAP and remote sensing integration,
Climate Risk Management,
Volume 51,
2026,
100783,
ISSN 2212-0963,
[Link]
([Link]
Abstract: The escalating impacts of global climate change and extreme weather have
intensified flood risks worldwide, including in arid and semi-arid regions traditionally
considered low-risk. This study examines the spatiotemporal dynamics of flood events across
Kazakhstan from 2000 to 2024 by integrating remote sensing (RS) with machine learning
(ML). Using Google Earth Engine (GEE), we address data gaps and cloud interference through
spatiotemporal fusion (STARFM), denoising, smoothing, and sample transferring techniques.
In addition, this study incorporates the Time-Disaggregated Water Frequency (TWF) method,
which enables the identification of water bodies with temporal variability, eliminates
permanent water bodies, and distinguishes flood from non-flood conditions in seasonal
water bodies, thereby enhancing the accuracy of flood reconstruction and enabling precise
delineation of flood inundation areas. Landsat and MODIS imagery are combined to produce
high-resolution flood distribution maps, while spectral similarity indicators guide the transfer
of samples from the Global Flood Database. A range of spectral, texture, environmental, and
socioeconomic features is extracted, with flood classification performed using random forest
(RF) and attribution analysis conducted via XGBoost and SHAP. Results highlight a high flood
risk in northern, southwestern, and western Kazakhstan, primarily driven by changes in
precipitation (PRE), temperature (TEM), soil moisture (SM), and land use. Floods occur most
frequently in spring — especially in March and April — due to snowmelt and extreme
precipitation. The ML models achieve over 80 % classification accuracy, demonstrating their
reliability. This work improves flood monitoring and provides essential insights for climate
adaptation and targeted flood risk management in Kazakhstan.
Keywords: Flood; Spatiotemporal; SHapley additive exPlanations (SHAP); Machine learning
(ML); Kazakhstan; Remote sensing
Haoming Feng, Huaqing Li, Wenwen Zhu, Denghao Li, Yukun Huang,
Micro-motion enhanced multi-person activity recognition with millimeter-wave radar,
Measurement,
Volume 258, Part B,
2026,
119090,
ISSN 0263-2241,
[Link]
([Link]
Abstract: As a non-contact sensing device, millimeter-wave radar exhibits unique strengths
in human activity recognition (HAR). Existing methods rely on micro-Doppler signatures for
activity classification, but they often encounter feature aliasing in multi-person activity
recognition (MPAR) scenarios. Although point cloud-based approaches can distinguish
individual targets, they primarily extract static morphological features, neglecting the micro-
motion information of human joints, which is crucial for accurate activity recognition. To
address these limitations, we proposes an innovative MPAR framework that integrates
spatial point clouds and micro-motion features. First, an improved point cloud data
association algorithm is applied to achieve multi-target point cloud feature separation,
followed by a dynamic projection mechanism to construct time–Doppler feature maps.
Then, a torso micro-motion enhancement algorithm is designed to enhance the details of
human body movements. Finally, a CNN-LSTM hybrid network architecture with a temporal-
attention is constructed for action classification. Experimental results show that the
proposed micro-motion enhancement algorithm improves recognition accuracy by 27.1%
and 2.3%, compared to two traditional time–frequency analysis methods. Furthermore,
MPAR task in occlusion scenarios achieves recognition accuracy of 93.5%. In summary,
proposed framework not only retains the inherent advantages of millimeter-wave radar but
also significantly enhances multi-person activity recognition in complex scenarios.
Keywords: Human activity recognition (HAR); Multi-person activity recognition (MPAR);
Feature separability; Temporal attention; Occlusion scenarios
Yun Zhou, Yinglin Zhu, Haohao Ren, Jiahao Kang, Xuegang Wang,
Refined multi-modal feature learning framework for marine target detection using radar
sensor,
Digital Signal Processing,
Volume 170,
2026,
105816,
ISSN 1051-2004,
[Link]
([Link]
Abstract: The fusion of time and time-frequency characteristics in radar echoes offers a
novel approach for marine target detection. However, echo amplitude alone cannot fully
characterize the time-domain information, as it fails to capture the temporal correlation
between sampling points. Therefore, this article introduces the Gramian Angular Summation
Field (GASF) for processing raw radar echoes to obtain the temporal information. Concretely,
to enable the detector to utilize features from diverse signal representations of the same
target echoes, we first preprocess the echoes of radar with two signal processing methods,
GASF and STFT, which aim to reflect the temporal dependence and dynamic changes of
frequency components, respectively. Subsequently, we develop a dual-stream feature
extraction network, i.e., time-frequency self-attention learning and GASF-based spatial-
temporal correlation learning, to deeply extract the discriminative features from two
modalities of the same radar echo. Then, to overcome the heterogeneity of multimodal
features during feature fusion, we propose a cross-modal feature fusion strategy to map
multi-modal features to a unified space. Finally, the fused features are fed into the detection
module. Numerous evaluation experiments on the publicly available measured IPIX dataset
demonstrate that the proposed detector is competitive with some state-of-the-art detectors
for marine target detection.
Keywords: Radar target detection; Signal processing; Gramian angular summation field;
Deep learning; Short-time Fourier transform
Xiangyang Luo, Ying Lu, Bibo Zhang, Yadan Yang, Jiaxin Li, Wanying Fu, Xinke Bu, Cong Li,
Identification and spatiotemporal analysis of braided rivers in the Yarlung Tsangpo basin
using an enhanced U-Net approach,
Journal of Hydrology,
Volume 666,
2026,
134796,
ISSN 0022-1694,
[Link]
([Link]
Abstract: Braided river systems, characterized by their unique ecological functions, play a
vital role in maintaining biodiversity, due to the unique braided morphology of braided river
systems, the primary prerequisite for conducting related research is the accurate
identification of braided water bodies within their catchments. However, most existing
studies rely heavily on field surveys or aerial imagery to extract information on braided river
networks. Such methods are costly, operationally complex, and insufficient for applications
requiring extensive spatial and temporal coverage. To address this challenge, this study
proposes an improved U-Net model, termed MSU-Net, which incorporates a multi-scale dual
attention gate module. By integrating spatial and channel attention mechanisms, the model
enhances the extraction of water features in complex environments. This study constructs a
monthly remote sensing dataset of the Yarlung Tsangpo River Basin from 2018 to 2023 using
Sentinel-1(SAR) and Sentinel-2 (optical) imagery. A model was trained based on manually
corrected labels and data augmentation strategies, and the spatial and temporal variations
of surface water area in the basin were analyzed based on the model outputs. The research
presents a more efficient method for identifying braided river systems and analyzes the
spatiotemporal dynamics of the braided channels in the Shannan section of the Yarlung
Tsangpo River Basin and their association with climatic factors, providing a scientific basis for
watershed water resource management.
Keywords: Braided river systems; Remote sensing; Yarlung Tsangpo–Brahmaputra River;
Water surface area; River channel variability
Xingyu Wang, Zhen Yang, Jichuan Huang, Bao Zhang, Yuhe Zhang, Deyun Zhou,
Collaborative strategy for hybrid actions of radar modes and maneuver decisions under
observation errors,
Engineering Applications of Artificial Intelligence,
Volume 160, Part A,
2025,
111774,
ISSN 0952-1976,
[Link]
([Link]
Abstract: The rapid advancement of airborne avionics has driven modern air combat to rely
heavily on information-centric operations, with radar serving as a primary tool for
information acquisition and playing a critical role in air combat. However, existing research
on air combat strategies often overlooks the impact of different radar operating modes on
maneuvering strategies, as well as the challenges posed by learning strategies under
observational disturbances. To address these gaps, this study investigates the problem of
hybrid actions decision-making for radar modes and maneuver decisions in the presence of
observational errors. Specifically, the characteristics of various radar operating modes are
analyzed and modeled, followed by an exploration of the convergence process of
reinforcement learning strategies under observational disturbances. To mitigate the
instability and volatility in strategy learning caused by observation errors, Entropy-
Decoupling-Noisy-net Proximal Policy Optimization-Advanced (EDN-PPOA) algorithm is
proposed, which significantly enhances the robustness and exploratory capability of the
model. Simulation results demonstrate that the proposed algorithm effectively achieves
coordinated tactical integration of radar modes and maneuvers in complex hybrid action
spaces, producing flexible tactical strategies that outperform expert-designed heuristics.
Furthermore, compared to the existing algorithms, the proposed method exhibits superior
stability and robustness in noisy observational environments, providing a reliable technical
foundation for intelligent decision-making in complex adversarial scenarios.
Keywords: Radar mode; Deep reinforcement learning; Hybrid actions; Maneuvering
decision-making
Pengfei Ge, Mi Chen, Roberto Tomás, Hui Liu, Kailun Fan, Xi Cheng, Xingyuan Fu, Shuang
Wang,
Insights into land deformation processes in the Yellow River Delta (China) from synthetic
aperture radar interferometry time series and machine learning,
Advances in Space Research,
2025,
,
ISSN 0273-1177,
[Link]
([Link]
Abstract: The Yellow River Delta is globally recognized as a highly dynamic region due to its
continuous transformations at the land-sea interface, and it abounds in valuable natural
resources, including oil, brine-rich groundwater and natural gas. The region experiences
impacts from tectonic activity, the natural compression and consolidation of loose
sediments, and, in particular, human economic activities. These factors lead to various types
of land subsidence, posing potential risks to local communities and economic operations.
Hence, effective monitoring and acquisition of the spatiotemporal patterns of land
subsidence in the Yellow River Delta play a crucial role in minimizing geological challenges
and financial losses. In this study, surface deformation data for the Yellow River Delta were
derived by processing 70 scenes of Sentinel-1 A/B data (32 ascending, 38 descending
acquisitions) using Interferometric Synthetic Aperture Radar (InSAR) time series technique,
with the Persistent Scatterer InSAR (PS-InSAR) and Small Baseline Subset InSAR (SBAS-InSAR)
methods applied over the period from January 2020 to December 2021. Moreover,
additional datasets, such as groundwater levels, precipitation, and areas of oil field and brine
extraction, were integrated to examine the factors affecting land subsidence and analyzed
using random forest analysis and post-interpretation methods. The findings indicate land
subsidence in the Yellow River Delta region displays an uneven distribution pattern, with
areas of severe subsidence primarily concentrated in Hekou District, Kenli District, Dongying
District and Guangrao County, characterized by the mean annual subsidence rate greater
than −120 mm/year. The spatial distribution of groundwater funnels, oil fields and brine
mining areas aligns to some extent with that of the severe land subsidence zones. The
random forest model outcomes reveal that the main contributors to land subsidence in the
Yellow River Delta are brine mining and soft soil thickness. Additionally, there exists regional
variability in the influencing factors among the various typical subsidence bowls. The post-
interpretation analysis further highlights shifts in the correlations among the various impact
factors and land subsidence.
Keywords: Land subsidence; PS-InSAR; SBAS-InSAR; Random forest; Yellow River Delta
Ji Ge, Hong Zhang, Lijun Zuo, Lu Xu, Jingling Jiang, Mingyang Song, Yinhaibin Ding, Yazhe Xie,
Fan Wu, Chao Wang, Wenjiang Huang,
Large-scale rice mapping under spatiotemporal heterogeneity using multi-temporal SAR
images and explainable deep learning,
ISPRS Journal of Photogrammetry and Remote Sensing,
Volume 220,
2025,
Pages 395-412,
ISSN 0924-2716,
[Link]
([Link]
Abstract: Timely and accurate mapping of rice cultivation distribution is crucial for ensuring
global food security and achieving SDG2. From a global perspective, rice areas display high
heterogeneity in spatial pattern and SAR time-series characteristics, posing substantial
challenges to deep learning (DL) models’ performance, efficiency, and transferability.
Moreover, due to their “black box” nature, DL often lack interpretability and credibility. To
address these challenges, this paper constructs the first SAR rice dataset with
spatiotemporal heterogeneity and proposes an explainable, lightweight model for rice area
extraction, the eXplainable Mamba UNet (XM-UNet). The dataset is based on the 2023
multi-temporal Sentinel-1 data, covering diverse rice samples from the United States, Kenya,
and Vietnam. A Temporal Feature Importance Explainer (TFI-Explainer) based on the
Selective State Space Model is designed to enhance adaptability to the temporal
heterogeneity of rice and the model’s interpretability. This explainer, coupled with the DL
model, provides interpretations of the importance of SAR temporal features and facilitates
crucial time phase screening. To overcome the spatial heterogeneity of rice, an Attention
Sandglass Layer (ASL) combining CNN and self-attention mechanisms is designed to enhance
the local spatial feature extraction capabilities. Additionally, the Parallel Visual State Space
Layer (PVSSL) utilizes 2D-Selective-Scan (SS2D) cross-scanning to capture the global spatial
features of rice multi-directionally, significantly reducing computational complexity through
parallelization. Experimental results demonstrate that the XM-UNet adapts well to the
spatiotemporal heterogeneity of rice globally, with OA and F1-score of 94.26 % and 90.73 %,
respectively. The model is extremely lightweight, with only 0.190 M parameters and 0.279
GFLOPs. Mamba’s selective scanning facilitates feature screening, and its integration with
CNN effectively balances rice’s local and global spatial characteristics. The interpretability
experiments prove that the explanations of the importance of the temporal features
provided by the model are crucial for guiding rice distribution mapping and filling a gap in
the related field. The code is available in [Link]
Keywords: Synthetic aperture radar; Rice mapping; Explainable deep learning; Feature
importance
Hu Liu, Zhenghua Zhang, Jing Yang, Jörg Benndorf, Xiaofei Wang, Jiaqi Dong, Zitao Lin,
Guoliang Chen,
GhostPointNet: A deep learning-based method for ghost point noise detection in four-
dimensional (4D) millimeter-wave radar point clouds of underground mine,
Engineering Applications of Artificial Intelligence,
Volume 161, Part C,
2025,
112380,
ISSN 0952-1976,
[Link]
([Link]
Abstract: The high dust concentration, multi-metal supports, and narrow winding tunnels in
underground mines collectively lead to frequent ghost point noise in four-dimensional (4D)
millimeter-wave radar point clouds, posing serious challenges for mining perception and
localization. To address this, we propose a deep learning algorithm, named GhostPointNet,
for 4D millimeter-wave radar ghost point detection in underground mining environments.
From an artificial intelligence perspective, this model thoroughly considers the multi-modal
features of 4D millimeter-wave radar and the environmental complexity of underground
mines. It incorporates multi-parameterized spatial information inputs in both Cartesian and
Spherical coordinates, coupled with “Double T-Net” adaptive alignment correction, while
integrating non-spatial information such as radar power and Doppler data to achieve multi-
modal representation and end-to-end discrimination between ghost points and real points.
Experimental validation shows that GhostPointNet achieves excellent performance in
underground mining scenarios with 92.45 % accuracy and 95.84 % F1-score, outperforming
traditional filtering, clustering, and machine learning algorithms. From an engineering
application perspective, GhostPointNet is specifically designed for ghost noise detection in
underground mines. Even in complex scenarios such as mine tunnel intersections and turns,
it preserves critical structural points. Its end-to-end neural network simplifies post-
processing procedures, enhances operational efficiency, and provides stable and reliable
perceptual support for subsequent tasks such as autonomous mine locomotive navigation
and three-dimensional (3D) structure reconstruction. Experimental results demonstrate that
this method surpasses baseline approaches in ghost point detection, real point preservation,
and generalization capability, providing significant support for improving underground
mining safety and efficiency.
Keywords: Deep learning; Four-dimensional (4D) millimeter-wave radar; Ghost noise;
Underground mining; Point cloud segmentation
Decai Jin, Xiufang Zhu, Ying Qu, Jianbo Qi, Hanyi Wu, Yaozhong Pan,
A robust and efficient deep optimization network for spatiotemporal data fusion,
Information Fusion,
Volume 127, Part C,
2026,
103939,
ISSN 1566-2535,
[Link]
([Link]
Abstract: Satellites strive to strike a delicate balance between temporal and spatial
resolution, thereby rendering the achievement of high resolution in both aspects
challenging. Spatiotemporal fusion algorithms have emerged as a promising solution to
tackle this challenge. However, with changes in spatiotemporal conditions, existing
spatiotemporal fusion methods, particularly those based on deep learning, face challenges
such as decreased prediction accuracy and poor reconstruction accuracy in areas of abrupt
changes. This presents significant challenges for the fusion of multi-source remote sensing
data to generate cloud-free remote sensing images on a daily scale. In this context, the study
proposes a multiscale Attention-Guided deep optimization network for Spatiotemporal Data
Fusion (AGSDF) method. The algorithm is designed to generate daily fine images using
coarse image, based on historical reference fine images. Specifically, it firstly attempts to use
a physical attention mechanism to mitigate the effects of climate change in time-series
images. Implementing a continuous spatiotemporal fusion process across multiple scales
significantly enhances the model's robustness. The performance of AGSDF was evaluated
and compared to nine methods at six sites worldwide. The experimental results indicate that
AGSDF achieved a top score in the assessment. Consequently, AGSDF holds high potential to
produce accurate remote sensing products with high temporal and spatial resolution across
extensive regions.
Keywords: Spatiotemporal fusion; Data fusion; MODIS; Landsat; Super resolution
Li Wang, Baicheng Hu, Yuan Zhao, Kunlin Song, Jianmin Ma, Hong Gao, Tao Huang, Xiaoxuan
Mao,
A hybrid spatiotemporal model combining graph attention network and gated recurrent unit
for regional composite air pollution prediction and collaborative control,
Sustainable Cities and Society,
Volume 116,
2024,
105925,
ISSN 2210-6707,
[Link]
([Link]
Abstract: Machine learning (ML) models have been extensively applied in air quality
prediction. However, many of these models often failed to unveil complex mechanisms and
regional spatiotemporal variations of composite air pollution. This brings uncertainties in
using ML models for effective composite air pollution control. The present study developed a
novel hybrid spatiotemporal model framework combining Graph Attention Network (GAT)
and Gated Recurrent Unit (GRU), namely the GAT-GRU model, to foresee composite air
pollutions with a focus on PM2.5 and O3. By extracting attention matrices for PM2.5O3
composite pollution and applying the Louvain algorithm, the framework established
effective community network divisions for coordinated control of PM2.5O3 composite
pollution. The framework was applied and tested in China's “2 + 26″ cities, a city cluster with
most heavy PM2.5 and O3 pollution and precursor emission sources. The results
demonstrate that the framework successfully captured spatiotemporal evolution of
combined PM2.5 and O3 pollution. The attention matrix is autonomously generated during
course of the model learning process with the aim to interpret the complex interactions
among “2 + 26″ cities. The framework provides a new perspective for the interpretability of
artificial intelligence models and offers a methodological support and scientific evidence for
formulating regional pollution cooperative governance strategies.
Keywords: GAT-GRU; PM2.5, O3; Attention matrix; Community network
Liangang Qi, Hongzhuo Chen, Qiang Guo, Shuai Huang, Mykola Kaliuzhnyi,
GLS: A hybrid deep learning model for radar emitter signal sorting,
Digital Signal Processing,
Volume 161,
2025,
105117,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Radar emitter signal sorting is a pivotal aspect of radar reconnaissance signal
processing. The increasing density of the electromagnetic environment in modern radar
pulse streams, coupled with the growing complexity and variability of operational modes
and signal forms, results in extremely limited reference data. Consequently, most existing
sorting methods fall short of meeting the performance requirements of modern electronic
warfare. To enhance sorting performance under conditions of limited samples and labeled
data, this paper proposes a radar emitter signal sorting model based on ResGCN-BiLSTM-SE
(GLS). Firstly, we propose a novel adaptive weighted adjacency matrix construction method
that aggregates multi-scale information of local and global features. Based on this, for GLS
networks, the graph convolutional network (ResGCN) is combined with the bidirectional long
short-term memory (BiLSTM) network. The GCN is employed to extract attribute features
from interleaved radar pulse sequences, while the BiLSTM is utilized to deeply capture the
temporal dependence in interleaved pulse sequences after feature extraction. Finally, an
improved squeeze-and-excitation (SE) module is applied to perform weighted fusion of
critical channel information from both spatial and temporal features. Simulation results
demonstrate that the proposed method not only achieves higher accuracy under small
sample conditions compared to existing methods, but also exhibits strong robustness in
challenging scenarios involving measurement errors, missing pulses, and spurious pulses.
Keywords: Radar emitter signal sorting (RESS); Adaptive weighted adjacency matrix; GLS
model; Features fusion
Laura Pedretti, Pietro Teatini, Tommaso Letterio, Guadalupe Bru, Carolina Guardiola-Albert,
Roberto Tomás, María I. Navarro-Hernández, Alessandro Bondesan, Yuri Taddia, Claudia
Meisina,
Vertical land movements assessment integrating Interferometric Synthetic Aperture Radar,
in-situ data, and engineering-geological model: The case study of the reclaimed farmland of
the Po River Delta (Italy),
Engineering Geology,
Volume 363,
2026,
108544,
ISSN 0013-7952,
[Link]
([Link]
Abstract: Low-elevation reclaimed coastlands face significant challenges from land
subsidence and sea-level rise, making long-term monitoring of ground movements crucial to
ensure infrastructure safety and preserve the natural environment. This study aims to
reconstruct the long-term historical ground deformation of the reclaimed farmland in the Po
River Delta by: i) integrating nearly 30 years of multisource, multi-temporal, and multisensor
Interferometric Synthetic Aperture Radar (InSAR) satellite data (ERS-1/2, RADARSAT-1/2,
Sentinel-1); ii) combining multisource InSAR datasets generated using different algorithms
covering distinct or overlapping time periods (Sentinel-1 PSI, P-SBAS, and IPTA); and iii)
developing a 3D engineering-geological model focused on the under-consolidated fine-
grained deposits that are more prone to subsidence. By combining multiple monitoring
techniques, this multidisciplinary approach reveals that land subsidence is primarily driven
by autocompaction of under-consolidated finegrained sediments, locally accelerated by
building construction, as evidenced by InSAR data. The highest subsidence rates occur in the
youngest reclaimed areas with thicker under-consolidated fine-grained deposits. While
integrating multisensor InSAR datasets from diverse sources to reconstruct longterm ground
deformation presents challenges, it also yields valuable insights. In this work, we
demonstrate that heterogeneous datasets can still be valuable when interpreted carefully
and that the feasibility of combining legacy and modern InSAR data for long historical
deformation reconstruction is a practical challenge in real-world data integration. Moreover,
this comprehensive approach enables updating spatial and temporal records of land
movement and identifying conditioning factors for inclusion in land movement susceptibility
and risk maps supporting land planning.
Keywords: Long-term monitoring; Land subsidence spatiotemporal evolution; Multi-
sourcetemporal-sensor InSAR; Under-consolidated fine-grained sediments;
Engineeringgeological model; Land reclamation; Po River delta
Shaopeng He, Mingjun Wang, Nicola Forgione, Andrea Pucciarelli, W.X. Tian, S.Z. Qiu, G.H.
Su,
A multi-task Transformer-Mamba-Seq framework for real-time estimation of spatiotemporal
thermal stratification in passive residual heat exchanger,
International Communications in Heat and Mass Transfer,
Volume 169, Part D,
2025,
109868,
ISSN 0735-1933,
[Link]
([Link]
Abstract: Passive Residual Heat Removal Heat Exchanger (PRHR HX) is a critical component in
Generation-III nuclear power systems. Its spatiotemporal thermal stratification
characteristics directly influence residual heat removal capacity and serve as key inputs for
multiphysics coupling analyses. However, the complexity of input conditions challenges
traditional simulation and AI approaches, particularly under abnormal and accident
scenarios. To address this, we propose a multi-task Transformer-Mamba-Seq framework that
integrates multi-head attention with a selective scan mechanism. Compared to conventional
models, it demonstrates superior performance in both 5-fold cross-validation and
K. Venkata rao,
A study on performance characteristics and multi response optimization of process
parameters to maximize performance of micro milling for Ti-6Al-4V,
Journal of Alloys and Compounds,
Volume 781,
2019,
Pages 773-782,
ISSN 0925-8388,
[Link]
([Link]
Abstract: In machining of hard metals, surface roughness, tool vibration and tool wear are
used as performance characteristics to estimate overall performance of process. This work is
aimed to maximize overall performance in micromachining of Ti-6Al-4V and investigate
effect of process parameters on performance characteristics. As per orthogonal array of L27,
twenty seven experiments are carried out on the proposed metal with cemented carbide
tools at three levels of cutting speed, feed and depth of cuts. According to user's preference
rating, graph theory and matrix approach is used to estimate weights and preference scales
for the performance characteristics. Utility concept is used to calculate overall performance
of the process using weights and preference scales. Responses surface methodology is used
to optimize process parameters to maximize overall performance of the process. In this
study, maximum performance of the process is found at cutting speed of 19.78 m/min, feed
of 75 μm/tooth and depth of cut of 50 μm. In addition to that, metal recovery in machining
due to elastic recovery is also estimated theoretically and measured practically for different
uncut chip thicknesses.
Keywords: Micro milling; Optimization; Utility concept; Taguchi; Graph theory and matrix
approach (GTMA)
Bachina Harish Babu, Sujith Bobba, T.C.H. Anil Kumar, NB. Prakash Tiruveedula, Talluri
Srinivasarao,
Optimization of dead metal zone to reduce cutting forces in micro milling of Inconel 718
using RSM,
Materials Today: Proceedings,
2023,
,
ISSN 2214-7853,
[Link]
([Link]
Abstract: This study focuses on the mechanism of DMZ (dead metal zone) creation, as well
as the impact of cutting edge geometries (sharp, chamfered, double chamfered, and blunt
edges), cutting speed, and coefficient of friction on DMZ formation while milling Inconel 718
material (FEM). A non-contact type sensor called a laser doppler vibrometer (LDV) is used to
monitor the vibration of rotating surfaces. In current research work, the LDV is used to
measure the mill cutter vibration in micro-milling of Inconel 718 in terms of acoustic optic
emission signals. A FFT (fast fourier transformer) is used for signals processing in to
frequency domain. Design of experiments as per Taguchi, experiments were performed on
the alloy at three levels of spindle speeds, depth of cuts, feed rates. Experimental results on
the amplitude o vibration of tool along X and Y directions, surface roughness were measured
and analysed using response surface methodology. Analysis obtained from the variance was
used to recognize the significant parameters which effect the vibration of tool and roughness
of surface. RSM was implemented and optimized process parameters for the minimum
vibration amplitude and surface roughness.
Keywords: Laser Doppler vibrometer; Inconel 718; Surface roughness; Tool vibration; DMZ,
Response surface methodology (RSM)
Nayeemul Islam Nayeem, Shirin Mahbuba, Sanjida Islam Disha, Md Rifat Hossain Buiyan,
Shakila Rahman, M. Abdullah-Al-Wadud, Jia Uddin,
A YOLOv11-Based Deep Learning Framework for Multi-Class Human Action Recognition,
Computers, Materials and Continua,
Volume 85, Issue 1,
2025,
Pages 1541-1557,
ISSN 1546-2218,
[Link]
([Link]
Abstract: Human activity recognition is a significant area of research in artificial intelligence
for surveillance, healthcare, sports, and human-computer interaction applications. The
article benchmarks the performance of You Only Look Once version 11-based (YOLOv11-
based) architecture for multi-class human activity recognition. The article benchmarks the
performance of You Only Look Once version 11-based (YOLOv11-based) architecture for
multi-class human activity recognition. The dataset consists of 14,186 images across 19
activity classes, from dynamic activities such as running and swimming to static activities
such as sitting and sleeping. Preprocessing included resizing all images to 512 × 512 pixels,
annotating them in YOLO’s bounding box format, and applying data augmentation methods
such as flipping, rotation, and cropping to enhance model generalization. The proposed
model was trained for 100 epochs with adaptive learning rate methods and hyperparameter
optimization for performance improvement, with a mAP@0.5 of 74.93% and a mAP@0.5-
0.95 of 64.11%, outperforming previous versions of YOLO (v10, v9, and v8) and general-
purpose architectures like ResNet50 and EfficientNet. It exhibited improved precision and
recall for all activity classes with high precision values of 0.76 for running, 0.79 for
swimming, 0.80 for sitting, and 0.81 for sleeping, and was tested for real-time deployment
with an inference time of 8.9 ms per image, being computationally light. Proposed
YOLOv11’s improvements are attributed to architectural advancements like a more complex
feature extraction process, better attention modules, and an anchor-free detection
mechanism. While YOLOv10 was extremely stable in static activity recognition, YOLOv9
performed well in dynamic environments but suffered from overfitting, and YOLOv8, while
being a decent baseline, failed to differentiate between overlapping static activities. The
experimental results determine proposed YOLOv11 to be the most appropriate model,
providing an ideal balance between accuracy, computational efficiency, and robustness for
real-world deployment. Nevertheless, there exist certain issues to be addressed, particularly
in discriminating against visually similar activities and the use of publicly available datasets.
Future research will entail the inclusion of 3D data and multimodal sensor inputs, such as
depth and motion information, for enhancing recognition accuracy and generalizability to
challenging real-world environments.
Keywords: Human activity recognition; YOLOv11; deep learning; real-time detection; anchor-
free detection; attention mechanisms; object detection; image classification; multi-class
recognition; surveillance applications
Md Shohag Mollik, Tanveer Saleh, Khairul Affendy Bin Md Nor, Mohamed Sultan Mohamed
Ali,
A machine learning-based classification model to identify the effectiveness of vibration for
μEDM,
Alexandria Engineering Journal,
Volume 61, Issue 9,
2022,
Pages 6979-6989,
ISSN 1110-0168,
[Link]
([Link]
Abstract: Micro electro-discharge machining (μEDM) uses electro-thermal energy from
repetitive sparks generated between the tool and workpiece to remove material from the
latter. However, one of the bottlenecks of μEDM is the phenomenon of short circuits due to
the physical contact between the tool and debris (formed during the erosion of the
workpiece). Adequate flushing of the debris can be achieved by applying low amplitude
high-frequency vibration to the workpiece. This study, however, shows that the application
of vibration does not yield beneficial results for the μEDM for all the parametric conditions.
This research used an off-the-shelf piezo vibrator as the high-frequency, low amplitude
vibration source to the workpiece during the μEDM process. The experiments were
conducted with and without vibration with the variation of applied discharge energy and
μEDM speed. The samples were characterized using scanning electron microscopes to gather
various data related to μEDM outputs. The results of this study revealed that vibration-
assisted μEDM becomes less effective as the discharge energy is increased (primarily by
increasing the capacitor value of the RC pulse generator). Similarly, the reduction of the
occurrence of the short circuit was profound when the low discharge energy level with low
voltage and low capacitor setting of the RC Pulse generator was used. The overall scale of
the overcut with various discharge energy and μEDM speed varied from 15.5 μm to 42 μm
for the conventional μEDM process. However, the scale above slightly reduced to 14.5 μm to
39 μm using an ultrasonic vibration device. Also, the taperness of the machined hole was
of ∼7%).
slightly reduced by applying the vibration device during the μEDM operation (overall average
Keywords: Machine learning; Micro electro discharge machining; Ultrasonic vibration; MRR;
Tool wear; EDM
K. Venkata Rao,
Power consumption optimization strategy in micro ball-end milling of D2 steel via TLBO
coupled with 3D FEM simulation,
Measurement,
Volume 132,
2019,
Pages 68-78,
ISSN 0263-2241,
[Link]
([Link]
Abstract: The present challenge in the manufacturing industry is to improve efficiency of
production activities while reducing wastage of power consumption. Past research focused
on multi response optimization of process parameters to improve performance of the
process. The present study proposed an optimization-based strategy to reduce power
consumption in micro ball end milling of D2 steel. As the power consumption is directly
proportional to cutting forces, the process parameters such as cutting speed, feed and depth
of cut were optimized to reduce cutting forces using teaching learning based optimization
(TLBO) technique coupled with 3D finite element method (FEM) simulation. During the
optimization, amplitude of cutter vibration and surface roughness were taken as constraints
as 60 µm (ISO 10816) and 2 µm (ISO 1302) respectively. Three best combinations of cutting
speed, feed and depth of cut were obtained for minimum cutting force. Among them,
combination of cutting speed of 15 m/min, feed of 112.5 µm/tooth and depth of cut of
85.25 µm has low power consumption of 67 W with tool vibration of 36.5 µm. However,
remaining two combinations were also considered to be the next best optimal cutting
conditions. Numerical simulation was carried out for the three best solutions and the cutting
forces and amplitude of cutter vibration were predicted. There was good agreement
between simulation results and experimental results that verified the acceptance of the
simulation. It was also found that the three best candidate solutions were having same the
cutting speed of 15 m/min (minimum cutting speed). Hence, the induced stresses in the
work piece were found to be with low values around 350 Mpa.
Keywords: Power consumption; Ball end milling; Simulation; Optimization; TLBO; Tool
vibration
Ayesha Ibrahim, Muhammad Zakir Khan, Muhammad Imran, Hadi Larijani, Qammer H.
Abbasi, Muhammad Usman,
RadSpecFusion: Dynamic attention weighting for multi-radar human activity recognition,
Internet of Things,
Volume 33,
2025,
101682,
ISSN 2542-6605,
[Link]
([Link]
Abstract: This paper presents RadSpecFusion, a novel dynamic attention-based fusion
architecture for multi-radar human activity recognition (HAR). Our method learns activity-
specific importance weights for each radar modality (24 GHz, 77 GHz, and Xethru sensors).
Unlike existing concatenation or averaging approaches, our method dynamically adapts
radar contributions based on motion characteristics. This addresses cross-frequency
generalization challenges, where transfer learning methods achieve only 11%–34% accuracy.
Using the CI4R dataset with spectrograms from 11 activities, our approach achieves 99.21%
accuracy, representing a 15.8% improvement over existing fusion methods (83.4%). This
demonstrates that different radar frequencies capture complementary information about
human motion. Ablation studies show that while the three-radar system optimizes
performance, dual-radar combinations achieve comparable accuracy (24GHz+77GHz: 96.1%,
24GHz+Xethru: 95.8%, 77GHz+Xethru: 97.2%), enabling flexible deployment for resource-
constrained applications. The attention mechanism reveals interpretable patterns: 77 GHz
radar receives higher weights for fine movements (superior Doppler resolution), while 24
GHz dominates gross body movements (better range resolution). The system maintains
71.4% accuracy at 10 dB SNR, demonstrating environmental robustness. This research
establishes a new paradigm for multimodal radar fusion, moving from cross-frequency
transfer learning to adaptive fusion with implications for healthcare monitoring, smart
environments, and security applications.
Keywords: Human activity recognition; Multi-modal fusion; Attention mechanisms; Cross-
frequency transfer learning
Sui Tan, Li Zhu, Jia-Huan Li, Zhi-Ruo Cui, Guan-Yuan Zhao, Jia-Yang Tu,
Non-contact vibration displacement measurement method for railway bridges based on
computer vision,
Structures,
Volume 80,
2025,
109803,
ISSN 2352-0124,
[Link]
([Link]
Abstract: To address the demand for efficient, non-contact vibration measurement of railway
bridges, this study proposes a novel structural displacement measurement method based on
Recurrent All-pairs Field Transforms (RAFT) optical flow estimation. It utilizes a pre-trained
RAFT network to infer inter-frame optical flows, extracting the average within a selected
region of interest (ROI) as the pixel displacement, and then converting it into real-world via
coordinate transformation. An indoor cable-stayed bridge model vibration test
demonstrated that all three methods exhibit increasing EMSE with increasing amplitude, but
the proposed method shows the slowest increase in error, with a relative RMSE still below
20 % at 0.297 mm (RMSE: 0.057 mm), while the LightTrack-based method (RMSE: 0.067 mm)
and the GMFlow-based method (RMSE: 0.097 mm) both have relative RMSE exceeding 20 %.
Under large amplitude vibrations (13.222–30.469 mm), all three methods exhibit
comparable accuracy, with relative RMSE below 10 %. Frequency identification results show
a relative error of less than 2.9 % for the first-order frequency across all cases. To further
validate its applicability in practical engineering, a Linear Variable Differential Transformer
(LVDT) was installed at the midspan of a railway simply-supported beam bridge, and visual
measurement devices were arranged both under and on the bridge, with vibration images
and LVDT data captured during train passing. Results show the proposed method’s midspan
deflection closely matches LVDT data, with a peak absolute error of 0.087 mm and RMSE of
0.037 mm, lower than the LightTrack - based method (0.258 mm peak error, 0.080 mm
RMSE). Ultimately, the proposed method was used to identify the deflection at the quarter
points, the relative displacement between the beam end and pier, and the relative
displacement between the track and track slab as well, which provides a reference for
similar engineering applications.
Keywords: Computer vision; Non-contact detection; Railway bridge health monitoring; Deep
learning; Vibration test
Archana Mathur, Abbas Mufaddal Dudhiyawala, Sudeepa Roy Dey, Snehanshu Saha,
Toward accurate breast cancer classification: A review of multi-modal machine learning
approaches,
Methods,
Volume 246,
2026,
Pages 48-61,
ISSN 1046-2023,
[Link]
([Link]
Abstract: The innovations in classifying breast cancer into malignant and benign categories
and further categorizing it into molecular subtypes have reshaped healthcare services,
enabling accurate diagnosis of these complex conditions. Identification of molecular
subtypes of breast cancer is one of the most important treatment challenges, as these
subtypes can have an enormous effect on the prognosis and treatment approaches. Data
integration from various modalities, such as transcriptomics, imaging, and genomics, has
been crucial in leveraging new opportunities to increase classification accuracy and improve
individualized treatment plans. These heterogeneous data sources are examined by applying
deep learning algorithms, which provide further insights into the complex patterns that
traditional approaches often overlook. In this paper, we explore the various modalities
researchers use to investigate breast cancer and the intriguing fusion techniques employed
to combine these modalities. We also review the most recent models (traditional, machine
learning, and deep learning), emphasizing their improvements over traditional classification
methods and the molecular subtype categorization of breast cancer. Furthermore, the
emphasis of this review is to examine techniques to process the entire image of the breast
tissue slide, which is challenging, particularly due to its size. We explore recent advances in
multiple instance learning tasks and the use of attention-based transformers and similar
architectures for annotating the WSI slides before using them for cancer classification. We
additionally discuss the interpretability tools—attention maps, saliency maps and model
explainability— in the context of transformers. In a nutshell, we aim to provide an in-depth
look at the revolutionary capabilities of deep learning models in precision oncology and
guide future research paths in this crucial field by synthesizing existing studies.
Keywords: Multimodality; Molecular subtype classification; Breast cancer prediction; Feature
fusion; Multiple instance learning; Whole slide imaging
Tiantian Wang, Nan Yan, Chaosan Yang, Zeliang An, Gongjing Zhang, Yuqing Xu,
Electromagnetic signal recognition using multimodal tri-branch semantic fusion network in
the UAV-assist integrated sensing and communication systems,
Digital Signal Processing,
Volume 171,
2026,
105820,
ISSN 1051-2004,
[Link]
([Link]
Abstract: Driven by the proliferation of integrated sensing and communication (ISAC)
systems, the accurate recognition of unauthorized unmanned aerial vehicle (UAV) signals in
dynamic electromagnetic environments has emerged as a critical challenge for spectrum
security and cognitive radio applications. Conventional automatic modulation recognition
(AMR) frameworks suffer from significant performance degradation in low signal-to-noise
ratio (SNR) regimes and exhibit limited adaptability to resource-constrained edge computing
platforms. To address these limitations, we propose a novel Multimodal Tri-branch Fusion
Network (MTF-Net) architecture that synergistically integrates time-frequency analysis with
statistical feature learning. The framework systematically processes binarized time-
frequency images (B-TFIs) and higher-order cumulant vectors through three collaboratively
operating branches: (1) A primary temporal feature extractor employing dilated convolution-
residual blocks (DCRBlocks) with hierarchical dilatation factors, incorporating channel
attention mechanisms to dynamically emphasize discriminative temporal patterns; (2) Dual
auxiliary branches based on Edge-Transformer modules (ETFormers), which achieve efficient
spatial-structural learning through depthwise separable convolutions (DSC) while capturing
long-range spectral dependencies via additive attention mechanisms with linear complexity;
(3) A hierarchical fusion module implementing cross-branch feature recalibration through
learnable parameter matrices. Extensive Monte Carlo experiments demonstrate that our
MTF-Net significantly outperforms traditional methods in recognition accuracy for radar and
communication signals under low SNR conditions, establishing a new benchmark for
lightweight AMR solutions in ISAC systems.
Keywords: Multi-modal feature fusion; Unmanned aerial vehicle(UAV); Integrated sensing
and communication (ISAC); Lightweight neural network; Transformer
Vladislav Semenyuk, Ildar Kurmashev, Alberto Lupidi, Dmitriy Alyoshin, Liliya Kurmasheva,
Alessandro Cantelli-Forti,
Advances in UAV detection: integrating multi-sensor systems and AI for enhanced accuracy
and efficiency,
International Journal of Critical Infrastructure Protection,
Volume 49,
2025,
100744,
ISSN 1874-5482,
[Link]
([Link]
Abstract: This review critically examines the progress in unmanned aerial vehicle (UAV)
detection and classification technologies from 2020 to the present. It highlights a range of
detection methods, including radar, radio frequency (RF), optical, and acoustic sensors, with
particular emphasis on the integration of these technologies through advanced sensor
fusion techniques. The paper explores the core technologies driving improvements in
detection accuracy, range, and reliability, with a special focus on the transformative role of
artificial intelligence and machine learning. These innovations have significantly enhanced
system performance, enabling more precise and efficient UAV detection. The review
concludes with insights into emerging trends and future developments that promise to
further refine UAV detection technologies, ensuring greater security and operational
reliability.
Keywords: UAV Detection; UAV Classification; Radar Technology; Sensor Fusion; Optical
Sensor; Acoustic Sensor
Jitao Zhang, Xingkui Mu, Qingfang Zhang, Natallia Poddubnaya, Dmitry Filippov, Jiagui Tao,
Fang Wang, Liying Jiang, Lingzhi Cao,
Structural, micro-structure, magnetic and dynamic magneto-elastic properties of samarium-
doped nickel-zinc spinel ferrites for efficient power conversion applications,
Journal of Magnetism and Magnetic Materials,
Volume 602,
2024,
172176,
ISSN 0304-8853,
[Link]
([Link]
Abstract: Development of high-quality ferrites behaved enhanced properties via inclusion of
ions in 3dn and 4fn series are desirable for efficient power conversion solid-state devices.
Nevertheless, inadequate microscopic behaviors with complex chemistry in materials
enables researchers to design and produce the macroscopic target device that remained
rudimentary and mindless. In this work, the microscopic properties in nickel-zinc spinel
ferrite series of Ni0.8Zn0.2SmxFe2-xO4 (x = 0, 0.02, 0.04, 0.06, 0.08) encompassing XRD,
SEM, EDS, VSM, FTIR and ESR were systemically characterized, and the evolutionary
mechanism of the corresponding structural, micro-structure, elemental composition and
distribution, magnetic, cations exchange, and micro-magnetic behavior was profoundly
revealed. Under microscopic examinations, the optimum composition at x = 0.02 with well-
arranged spinel structure, dense texture with expected elemental composition, favorable
soft magnetic properties and micro-magnetic behaviors is evident. Fortunately, this optimum
is in coincidence with the achievable maximum magneto-mechanical coefficient, even the
macroscopic electric properties with stronger magnetoelectric (ME) interactions and higher
power conversion efficiency (PE) in tri-layered ME samples as expected. Experimental results
show that the eventual PE reaches its maximum of 75.67 % under Ropt = 33kΩ for samples
at x = 0.02, and exhibit a 2.24 times higher PE than that of sample without samarium
substitution. These findings provide a holographic perspective to connect the microscopic
beneficial effects of materials to bulk device that are promising for efficient power
conversion solid-state electronics.
Keywords: Samarium-doped spinel ferrites; Dynamic magneto-elastic properties; Power
conversion devices
Ruige Yang, Peng Shan, Yang He, Hongming Xiao, Lin Zhang, Yuliang Zhao, Qiang Fu,
A lightweight bionic flapping wing drone recognition network based on data enhancement,
Measurement,
Volume 239,
2025,
115476,
ISSN 0263-2241,
[Link]
([Link]
Abstract: In recent times, there has been a growing focus on research into bionic drones,
which seek to mimic biological behavior and structure, thus overcoming the limitations of
conventional drones. The ability of bionic drones to blend into their surroundings presents a
significant challenge for identification. This study presents a dataset of bionic drones and
introduces the Bionic Drone Identification Network (BDRNet). The dataset was enriched
using data augmentation techniques to improve model recognition. Moreover, an
Aggregated Attention Mechanism (AAM) captures input feature correlation. Furthermore, a
Merged and Integrated Detector Head (MIDHead) and Multi-scale Lightweight Convolution
(MLWConv) have been proposed to lessen computational costs. The findings reveal that
BDRNet achieves an AP0.5 of 94.4 %, a Params of 2.6 M, and a Flops of 5.5G, surpassing
mainstream object detection models such as Faster R-CNN and YoloV5. This suggests that
BDRNet demonstrates robust recognition capabilities and holds potential for deployment in
resource-constrained embedded devices.
Keywords: Bionic flapping-wing drone Identification; Data augmentation; Lightweight;
Engineering application
J.R.J. Bennett, G.P. Škoro, John Back, S.J. Brooks, T.R. Edgecock, S.A. Gray, A.J. McFarland, K.J.
Rodgers, C.N. Booth,
Lifetime and strength tests of tantalum and tungsten under thermal shock for a Neutrino
Factory target,
Nuclear Instruments and Methods in Physics Research Section A: Accelerators,
Spectrometers, Detectors and Associated Equipment,
Volume 646, Issue 1,
2011,
Pages 1-6,
ISSN 0168-9002,
[Link]
([Link]
Abstract: A description is given of tests on tantalum and tungsten wires to evaluate their
lifetime and strength under the thermal shock that will be experienced when a solid target is
bombarded with short pulses of high energy protons in a Neutrino Factory. The results of
lifetime tests and measurements of dynamic strength characteristics at high temperatures,
stresses and strain rates using a laser Doppler vibrometer are given. The tests show that a
solid tungsten target will have a life of at least 3 years, which, with other beneficial
characteristics, make it an excellent candidate for the Neutrino Factory.
Keywords: Tantalum; Tungsten; Thermal shock; Material strength; Target lifetime; Neutrino
Factory
Pieter G.G. Muyshondt, Lukas Prochazka, Merlin Schär, Michail Chatzimichalis, Bastian
Baselt, Guy Fierens, Flurin Pfiffner,
Finite-element modelling of the 3D motion of the malleus-incus complex validated with 3D
laser Doppler vibrometry,
Hearing Research,
Volume 469,
2026,
109477,
ISSN 0378-5955,
[Link]
([Link]
Abstract: Three-dimensional (3D) motions of the middle ear (ME) are investigated with
finite-element (FE) modelling by comparison with 3D laser Doppler vibrometer (LDV)
measurements of the malleus-incus complex. 3D point velocity measurements are converted
to 3D rigid-body motion (RBM) components of the malleus and incus under acoustic
excitation of the ME from 0.2 kHz to 8 kHz. The parameters in the FE model are adjusted to
provide qualitative agreement with the 3D motion measurements for three separate model
geometries. The results show a dominant hinge-like motion for malleus and incus across the
frequency range, but with an increase of other components at high frequencies to yield a
more complex motion. Incudomallear joint flexibility increases the relative motion between
malleus and incus and is shown to contribute most to the ME transformer ratio at low and
especially high frequencies, including the phase delay across the two ossicles. The dominant
motion direction of the umbo coincides with the medial-lateral axis across the frequency
range. The malleus head, incus head and incus long process show a deviation from this
motion direction between 1.5 kHz and 5 kHz, associated with dips in the corresponding
velocity magnitude. Motion trajectories at these points follow a line below 1.5 kHz but
alternate between a line and ellipse at higher frequencies. While the tympanic membrane
influences the 3D motion of malleus and incus in a similar way, the ME suspensory ligaments
affect the motion components of the ossicles to varying degrees depending on the location
on the ossicles.
Keywords: Middle ear mechanics; Middle ear modelling; 3D laser Doppler vibrometry; Finite
element modelling
Dongfang Wang, Yufeng Yang, Jilin Lei, Baojian Wang, Qiming Ouyang, Penghao Yin,
Frontier exploration of image recognition in fuel spray diagnostics: Hybrid deep learning
models and multimodal data fusion,
Journal of the Energy Institute,
Volume 123,
2025,
102274,
ISSN 1743-9671,
[Link]
([Link]
Abstract: Owing to its high detection accuracy and real-time processing capabilities, image
recognition technology has become an indispensable tool for extracting spray morphological
characteristics and analyzing dynamic evolution processes in combustion systems. This
review systematically summarizes recent advances in image recognition technology, with a
focus on its applications in multiphase flow coupling and detailed feature extraction within
spray environments, while also providing a forward-looking discussion of current challenges
and future trends. Conventional methods are limited by weak anti-interference capability,
low feature extraction efficiency, and poor generalization. Deep learning techniques have
been increasingly adopted to enhance boundary segmentation precision and quantitative
feature parameter extraction. However, a major challenge remains in adapting these
technologies to complex environments, as most existing models struggle to balance
lightweight design with measurement accuracy—a critical barrier to real-time engineering
applications. Emerging approaches, including hybrid CNN–Transformer architectures and
novel Mamba-based models such as UltraLight_VM_UNet, have demonstrated significant
potential. The model achieves a segmentation accuracy of up to 95.43 % mIoU for complex
sprays, while reducing computational costs to just 0.05M parameters and 0.33 GFLOPs.
These advancements significantly improve robustness and generalization under noisy and
dynamic spray conditions. Future developments are expected to focus on computational
efficiency, robustness in extreme scenarios, and more effective global–local feature fusion,
thereby paving the way for real-time diagnostic applications in combustion systems.
Keywords: Spray combustion diagnostics; Image recognition technology; Image acquisition;
Parameter extraction; Deep learning
Juan José Villamarín Marrugo, Juan Manuel Naranjo Piñeros, Erwin Hernando Hernandez
Rincon,
Evidence on the Utility of Artificial Intelligence in the Interpretation of Diagnostic
Radiological Images in Low and Middle-Income Countries: A Scoping Review,
Academic Radiology,
2025,
,
ISSN 1076-6332,
[Link]
([Link]
Abstract: Rationale and Objectives
Access to diagnostic imaging in low- and middle-income countries (LMICs) is limited by
scarce equipment, geographic barriers, weak digital infrastructure, and shortages of trained
personnel. Artificial intelligence (AI) has emerged as a promising tool to mitigate these gaps
by improving diagnostic accuracy, assisting non-specialist health workers, and optimizing
workflows. This scoping review aimed to synthesize current evidence on the use of AI for
interpreting radiological diagnostic images in LMICs.
Materials and Methods
A scoping review was conducted in July 2025 following Arksey and O’Malley’s framework
and PRISMA-ScR guidelines. Searches were performed in PubMed, Scopus, and Clinical Key
for studies published between 2000 and July 2025 in English and Spanish. Eligible studies
included clinical applications of AI in radiological imaging within LMICs, reporting relevant
outcomes.
Results
From 620 records, 51 studies conducted across 33 LMICs were included. Most were
published between 2022 and 2025 and focused on ultrasound, X-ray, and computed
tomography. AI consistently improved diagnostic sensitivity, specificity, and applicability,
particularly for tuberculosis, pneumonia, obstetric care, and oncologic screening. Magnetic
resonance imaging showed promising yet mostly experimental evidence, while
mammography research remained scarce. Frequent limitations included small sample sizes,
single-center designs, reliance on public datasets, and limited multicenter validation.
Conclusion
AI demonstrates significant potential to enhance the interpretation of diagnostic radiological
images in LMICs, with consistent gains in sensitivity, specificity, and applicability across
modalities such as ultrasound, X-ray, and computed tomography. Several studies also
reported improvements in workflow efficiency and support for non-specialist providers,
underscoring AI’s dual role as a diagnostic and operational tool. Nonetheless,
methodological heterogeneity and infrastructural challenges highlight the need for
multicenter validation and context-adapted implementation strategies to ensure sustainable
integration.
Keywords: artificial intelligence; diagnostic imaging; low- and middle-income countries;
radiology; scoping review
Yu Han, Panpan Wen, Zhuoying Liu, Rui Yi, Yinuo Chen, Sheng Cao,
From Recognition to Action: Integrating Deep Learning and Robotic Control in Transthoracic
Echocardiography,
Ultrasound in Medicine & Biology,
2026,
,
ISSN 0301-5629,
[Link]
([Link]
Abstract: Population aging has driven a rise in heart failure cases, increasing the clinical
burden on cardiac diagnostics. As a first-line imaging method, transthoracic
echocardiography (TTE) faces limitations due to operator dependence, patient variability,
and workflow inefficiencies. Meanwhile, advances in artificial intelligence (AI) and robotic
ultrasound systems offer new potential pathways toward automated diagnosis. This review
examines the current landscape of AI-based image analysis and robotic-assisted
echocardiography. It presents a detailed analysis of advancements in artificial intelligence
(AI) applied to echocardiography and the evolution of robotic ultrasound systems, aiming to
introduce a discussion on semantic-to-motion mapping. By synthesizing recent progress and
outlining future directions, we can correctly recognize the current maturity level of artificial
intelligence development in the field of ultrasound examination and prepare well for the
subsequent work.
Keywords: Transthoracic echocardiography; Deep learning for medical imaging; Robotic
ultrasound scanning; Semantic-guided control; Multimodal intelligence
Taofeng Gu, Yang Liang, Yangtian Yan, Wenjun Jiang, Haiyan Yue, Gang Hu, Jize Zhang,
Towards high-fidelity urban wind profiles for the built environment: a neural field to fuse
multi-source observational data in Guangzhou, China,
Building and Environment,
Volume 288,
2026,
114009,
ISSN 0360-1323,
[Link]
([Link]
Abstract: Accurate urban wind analysis is critically hampered by sparse and heterogeneous
observational data. This work presents a solution through NF-MW (stands for Neural Field
for Multi-source Winds), a model that fuses data from Doppler LiDAR and wind profiler radar
into a continuous high-resolution wind field. By learning a direct mapping from spatio-
temporal coordinates to wind values, NF-MW can reconstruct wind speed and direction at
any arbitrary height and time. The framework uniquely handles the 360∘ periodicity of wind
direction and uses Fourier-enriched features to capture high-frequency gusts and turbulence
often missed by other models. In a Guangzhou case study, NF-MW achieved a Mean
Absolute Error of 0.55 m/s for wind speed and 8.95∘ for wind direction, demonstrating
superior accuracy over traditional methods. This approach provides the building and
environment community with a robust method to generate the realistic dynamic wind data
essential for applications ranging from pedestrian comfort assessments to urban air quality
modeling.
Keywords: Urban wind environment; Neural fields; Doppler LiDAR; Wind profiler radar; Data
fusion; Deep learning
Lixing Shi, Xueling Liang, Wenchao Chen, Yaoqiang Liu, Tong Ding, Kun Qin, Bo Chen,
Hongwei Liu,
Masked variational transformer for complex clutter modeling and target detection,
Signal Processing,
Volume 239,
2026,
110236,
ISSN 0165-1684,
[Link]
([Link]
Abstract: Weak target detection commonly encounters intense clutter interference, which
overshadows weak signals and complicates the task. Taking advantage of the powerful data
mining capability of neural networks, more and more deep learning-based methods are
applied to radar target detection. Among the approaches, those founded upon unsupervised
learning methodologies exhibit remarkable merit because they dispense with the
requirement for target samples within the training step, making them highly applicable in
practical target detecting scenarios. However, existing methods suffer from limitations in
leveraging the range-Doppler (R-D) two-dimensional correlation and finely modeling in
multiple clutter scenarios. In this paper, an unsupervised Transformer-based detector (TrDet)
is proposed to break through the boundary of modeling capability. First, with the designed
two-dimensional position embedding (2-DPE) and global query embedding (GQE)
techniques, an unsupervised training strategy for R-D spectrum based on Transformer
framework is utilized to achieve refined clutter modeling. Then, radar target detection is
formulated as an out-of-distribution (OOD) detection task to mitigate clutter interference.
Moreover, the masked variational Transformer-based detector (MVTrDet) is further
proposed to prevent target information leakage when the target is in close proximity to the
clutter in Doppler domain. Compared with several relative algorithms, our proposed
methods are better suited for radar target detection in complex clutter environments. The
experimental results derived from both measured data and simulated data verify the
effectiveness of our proposed methods.
Keywords: Radar target detection; Clutter modeling; Range-Doppler (R-D) spectrum;
Unsupervised learning; Out-of-distribution detection; Transformer
Changlong Wang, Jiawei Jiang, Chong Han, Hengyi Ren, Lijuan Sun, Jian Guo,
Through-Wall Multihuman Activity Recognition Based on MIMO Radar,
Computers, Materials and Continua,
Volume 83, Issue 3,
2025,
Pages 4537-4550,
ISSN 1546-2218,
[Link]
([Link]
Abstract: Existing through-wall human activity recognition methods often rely on Doppler
information or reflective signal characteristics of the human body. However, static
individuals, lacking prominent motion features, do not generate Doppler information.
Moreover, radar signals experience significant attenuation due to absorption and scattering
effects as they penetrate walls, limiting recognition performance. To address these
challenges, this study proposes a novel through-wall human activity recognition method
based on MIMO radar. Utilizing a MIMO radar operating at 1–2 GHz, we capture activity data
of individuals through walls and process it into range-angle maps to represent activity
features. To tackle the issue of minimal variation in reflection areas caused by static
individuals, a multi-scale activity feature extraction module is designed, capable of extracting
effective features from radar signals across multiple scales. Simultaneously, a temporal
attention mechanism is employed to extract keyframe information from sequential signals,
focusing on critical moments of activity. Furthermore, this study introduces an activity
recognition network based on a Deformable Transformer, which efficiently extracts both
global and local features from radar signals, delivering precise human posture and activity
sequences. In experimental scenarios involving 24 cm-thick brick walls, the proposed
method achieves an impressive 97.1% accuracy in activity recognition classification.
Keywords: MIMO radar; human activity; Transformer; through-wall
Hana Sebia, Thomas Guyet, Mickaël Pereira, Marco Valdebenito, Hugues Berry, Benjamin
Vidal,
Vascular segmentation of functional ultrasound images using deep learning,
Computers in Biology and Medicine,
Volume 194,
2025,
110377,
ISSN 0010-4825,
[Link]
([Link]
Abstract: Segmentation of medical images is a fundamental task with numerous
applications. While MRI, CT, and PET modalities have significantly benefited from deep
learning segmentation techniques, more recent modalities, like functional ultrasound (fUS),
have seen limited progress. fUS is a non invasive imaging method that measures changes in
cerebral blood volume (CBV) with high spatio-temporal resolution. However, distinguishing
arterioles from venules in fUS is challenging due to opposing blood flow directions within
the same pixel. Ultrasound localization microscopy (ULM) can enhance resolution by tracking
microbubble contrast agents but is invasive, and lacks dynamic CBV quantification. In this
paper, we introduce the first deep learning-based application for fUS image segmentation,
capable of differentiating signals based on vertical flow direction (upward vs. downward),
using ULM-based automatic annotation, and enabling dynamic CBV quantification. In the
cortical vasculature, this distinction in flow direction provides a proxy for differentiating
arteries from veins. We evaluate various UNet architectures on fUS images of rat brains,
achieving competitive segmentation performance, with 90% accuracy, a 71% F1 score, and
an IoU of 0.59, using only 100 temporal frames from a fUS stack. These results are
comparable to those from tubular structure segmentation in other imaging modalities.
Additionally, models trained on resting-state data generalize well to images captured during
visual stimulation, highlighting robustness. Although it does not reach the full granularity of
ULM, the proposed method provides a practical, non-invasive and cost-effective solution for
inferring flow direction—particularly valuable in scenarios where ULM is not available or
feasible. Our pipeline shows high linear correlation coefficients between signals from
predicted and actual compartments, showcasing its ability to accurately capture blood flow
dynamics.
Keywords: Functional ultrasound; Segmentation; Ultrafast ultrasound localization
microscopy; Preclinical; Medical images; Neuroscience
Pengfei Yan, Wushuang Gong, Minglei Li, Jiusi Zhang, Xiang Li, Yuchen Jiang, Hao Luo, Hang
Zhou,
TDF-Net: Trusted Dynamic Feature Fusion Network for breast cancer diagnosis using
incomplete multimodal ultrasound,
Information Fusion,
Volume 112,
2024,
102592,
ISSN 1566-2535,
[Link]
([Link]
Abstract: Ultrasound is a critical imaging technique for diagnosing breast cancer. However,
the multimodal breast ultrasound diagnostic process is time-consuming and labor-intensive,
heavily dependent on the physician’s extensive expertise. Therefore, developing a computer-
aided diagnosis system for breast cancer is essential. Existing diagnostic systems fail to
consider the varying impacts of different ultrasound modalities on diagnostic results and
struggle to address the issue of missing modalities in clinical practice. Consequently, this
paper proposes the Trusted Dynamic Feature Fusion Network (TDF-Net) for diagnosing
breast cancer using incomplete multimodal ultrasound data. Initially, this method introduces
a dual-branch feature extraction module to capture modality-specific information.
Meanwhile, a contrastive clustering loss is designed to enforce the consistency constraint,
ensuring the coherence of different modal features within the semantic space for each
sample. Additionally, an invertible neural network-based recovery method is suggested to
establish mappings between different modalities, enabling the recovery of missing
modalities. Finally, a trusted dynamic feature fusion module based on the Dirichlet
distribution is proposed to quantify each modality’s contribution to the diagnostic result by
considering uncertainty, thereby achieving the dynamic fusion of each modality’s features
across different samples. The proposed method is validated on an established multimodal
breast ultrasound dataset, demonstrating superior diagnostic performance compared to
existing methods, with an average AUC of 98.31%, a 95% confidence interval of [96.98%,
99.64%], and a p-value < 0.05. A pilot study is planned to assess the effectiveness and
usability of TDF-Net in clinical settings.
Keywords: Breast cancer; Transformer; Ultrasound; Multimodality; Missing modality
Mao Li, Sen Wang, Tao Liu, Xiaoqin Liu, Chang Liu,
Rotating box multi-objective visual tracking algorithm for vibration displacement
measurement of large-span flexible bridges,
Mechanical Systems and Signal Processing,
Volume 200,
2023,
110595,
ISSN 0888-3270,
[Link]
([Link]
Abstract: Visual displacement measurement methods for flexible structural bodies like large-
span bridges has gained wide popularity in recent years, but practical applications still have
some limitations. For instance, when acquiring images of large-span flexible bridges at a
distance, the slight angular tilt of the detection target due to irregular vibrations can cause
extremely serious misfit errors in the displacement curves returned by the vision
measurement algorithm. To improve the reliability of vibration displacement measurement
of flexible structural bodies, this paper takes the bridge subjected to external excitation in
the acquired image sequence as the object of vibration displacement measurement and
uses a designed high-precision displacement measurement algorithm for a single-stage
rotating target tracking anchor-free box to track the vibration displacement of the target in
the flexible structural body. We first extract multi-scale feature information of bridge model
image sequences using the improved YOLOv5-s backbone network and combine the
Transformer self-attention mechanism with PANet to perform a top-down and bottom-up bi-
directional fusion of target feature maps at three different scales to achieve semantic feature
fusion of shallow and deep information. Second, the improved Efficient Decoupled Head
performs the detection of rotating target centroid offset and bounding box size. Finally, the
detected results are passed into the multi-objective tracking algorithm ByteTrack, which
strengthens the spatio-temporal correlation between frames and obtains a better-fitting
vibration displacement curve. The validation and comparison of traditional visual
measurement methods and deep learning measurement methods on cable-stayed bridge
models, small arch bridges, and large span bridges show that the vibration displacement
trajectories regressed by the algorithm in this paper have the best fit with the actual
vibration displacement trajectories, which also verifies that the algorithm in this paper has
good potential for engineering applications and implementation space in the field of
condition monitoring of flexible structural bodies.
Keywords: Flexible structure; Visual vibration measurement; Tilted targets; Rotating box;
Multi-target visual tracking; Deep convolutional neural network
Haiqiao Wang, Hong Wu, Zhuoyuan Wang, Peiyan Yue, Dong Ni, Pheng-Ann Heng, Yi Wang,
A Narrative Review of Image Processing Techniques Related to Prostate Ultrasound,
Ultrasound in Medicine & Biology,
Volume 51, Issue 2,
2025,
Pages 189-209,
ISSN 0301-5629,
[Link]
([Link]
Abstract: Prostate cancer (PCa) poses a significant threat to men's health, with early
diagnosis being crucial for improving prognosis and reducing mortality rates. Transrectal
ultrasound (TRUS) plays a vital role in the diagnosis and image-guided intervention of PCa.
To facilitate physicians with more accurate and efficient computer-assisted diagnosis and
interventions, many image processing algorithms in TRUS have been proposed and achieved
state-of-the-art performance in several tasks, including prostate gland segmentation,
prostate image registration, PCa classification and detection and interventional needle
detection. The rapid development of these algorithms over the past 2 decades necessitates
a comprehensive summary. As a consequence, this survey provides a narrative review of this
field, outlining the evolution of image processing methods in the context of TRUS image
analysis and meanwhile highlighting their relevant contributions. Furthermore, this survey
discusses current challenges and suggests future research directions to possibly advance this
field further.
Keywords: Transrectal ultrasound; Prostate cancer; Medical image processing; Deep
learning; Machine learning; Medical image segmentation; Medical image registration;
Classification; Computer-assisted detection; Computer-assisted diagnosis
Table of Content,
Chinese Journal of Aeronautics,
Volume 38, Issue 8,
2025,
103685,
ISSN 1000-9361,
[Link]
([Link]
Xiang Li, PengTao Guo, Yuan Ding, Zhiwei Chen, Xu Wang, Qibao Lv,
A generalized electromechanical coupled model of standing-wave linear ultrasonic motors
and its nonlinear version,
Mechanical Systems and Signal Processing,
Volume 186,
2023,
109870,
ISSN 0888-3270,
[Link]
([Link]
Abstract: For the systematization of research on standing-wave linear ultrasonic motors
(SWLUMs), this work develops a generalized electromechanical coupled model for
characterizing SWLUMs. The proposed model focuses on dealing with modeling the
generalized two-stage energy conversion in SWLUMs. The first-stage energy conversion is
modeled by a four-terminal equivalent circuit model, which with a phase shifter and a
couple of electromechanical transformers based on the electromechanical analogy method.
The second-stage energy conversion is modeled by a physics-based, friction-driven system,
involving contact nonlinearities between the stator and the mover. Furthermore, the
effectiveness of this model and its further extension considering nonlinear vibrations of
piezoelectric transducer (stator) are exemplified and discussed by a classical SWLUM with V-
configuration stator, thereby indicating that the presented generalized model is valuable and
pragmatic in simulating and characterizing SWLUMs both in electrical and mechanical
domains.
Keywords: Linear ultrasonic motor; Piezoelectric transducer; Generalized model; Contact
nonlinearities; Nonlinear vibration
Bofeng Liang, Li-Yun Fu, Mian Lin, Tobias Müller, Wubing Deng, Tongcheng Han,
Seismic efficiency: From hydraulic fracturing-acoustic emission laboratory experiments of
shale based on energy budget,
Geoenergy Science and Engineering,
Volume 252,
2025,
213917,
ISSN 2949-8910,
[Link]
([Link]
Abstract: Fluid injection-triggered earthquakes have been documented worldwide and quite
a number of events have significant moment magnitudes (Mw ≥ 3). Seismic efficiency (η),
defined as the ratio of injection volume to net seismic moment release in hydraulic
fracturing operations, is a crucial parameter to evaluate seismic hazard. However, a
quantitative assessment of seismic and non-seismic (aseismic) energy release is a key aspect
of understanding the intrinsic properties of cracking rocks. Therefore, we develop a novel η
model based on the hydraulic-fracturing-propagation energy budget and performed
laboratory experiments on hydraulic fracturing in shale by injecting distilled water at
different rates under pseudo-triaxial stress conditions with simultaneous monitoring of
acoustic emission (AE). We estimate AE energy accurately with absolute value correction of
sensors using a laser Doppler vibrometer, and the dissipation of the potential energy using
displacement and pressure sensors. The results show that the proposed η model can
evaluate the induced seismic characteristics effectively compared with field data and the
injection rate controls the change of η to some extent. Moreover, there is a log-linear
relationship between seismic efficiency and injection efficiency (ratio of AE energy and
injection energy), which may provide an experiential method for evaluating seismicity during
the early phase of hydraulic fracturing.
Qiming Zhang, Yang Li, Zhi Zhang, Shibo Yin, Lin Ma,
Marine target detection for PPI images based on YOLO-SWFormer,
Alexandria Engineering Journal,
Volume 82,
2023,
Pages 396-403,
ISSN 1110-0168,
[Link]
([Link]
Abstract: For the task of detecting marine targets, numerous machine-learning methods
have been suggested, which can achieve comparable accuracy. Despite the successful
application of deep learning in the field of marine target detection in recent years, existing
detection methods face challenges due to the significant interference of sea clutter. As
artificial intelligence technology advances, the Swin Transformer can serve as an effective
backbone for extracting discriminative features. However, it has not yet been utilized for
target detection, and the combination of Swin Transformer and YOLO architecture has not
been applied to similar missions. In light of this, we propose a novel method, YOLO-
SWFormer, which combines the Swin Transformer and YOLO framework for target detection.
Our method can extract discriminative features from plan-position indicator (PPI) images
despite the interference of sea clutter, thereby reducing computational complexity and
enhancing target detection accuracy. Experimental results on a Sea Clutter Database
demonstrates that our method surpasses existing methods in terms of accuracy, indicating
its potential as a promising solution for marine target detection tasks.
Keywords: PPI images; Target detection; Sea clutter suppression; Swin transformer; YOLO
Zhi Liu, Dexiang Le, Tianyu Zhang, Qingrong Lai, Jiansheng Zhang, Bin Li, Yunfeng Song, Nan
Chen,
Detection of apple moldy core disease by fusing vibration and Vis/NIR spectroscopy data
with dual-input MLP-Transformer,
Journal of Food Engineering,
Volume 382,
2024,
112219,
ISSN 0260-8774,
[Link]
([Link]
Abstract: Moldy core is a highly contagious internal disease of apples, and even a small
number of diseased apples can trigger large-scale infections during the storage stage. In this
study, a combined acoustic vibration and Vis/NIR spectroscopy method for moldy core apple
identification was proposed to improve the accuracy of moldy core identification. The
vibration signals and Vis/NIR spectroscopy of apples were collected using a self-designed
micro-LDV detection device and Vis/NIR spectroscopy online detection device respectively
for constructing multiple moldy core apple classification models. The results showed that
the classification model combining vibration spectrum and Vis/NIR spectral data had
significant advantages in apple moldy core identification accuracy compared to the
classification model using a single vibration spectrum or Vis/NIR spectral data. Ultimately,
dual-input MLP-Transformer (DMLPT) demonstrated the best recognition performance with
an overall classification accuracy of 99.31% for the model, with 100%, 97.56%, 100.00% and
100% accuracy for normal, mild moldy core disease, moderate moldy core disease and
severe moldy core disease apples, respectively. This study demonstrated the excellent
performance and great potential of acoustic vibration and visible/near-infrared spectral data
fusion for fruit internal quality detection.
Keywords: Laser Doppler vibrometer; Visible near-infrared spectroscopy; Dual-input MLP-
Transformer; Non-destructive detection; Moldy core apple
Bin Zhang, Xinru Ma, Xiaoping Ma, Xiaohong Jia, Xuejun Zhang,
The impacts of rainfall on MEMS Lidar SNR and detecting ability,
Infrared Physics & Technology,
Volume 150,
2025,
106009,
ISSN 1350-4495,
[Link]
([Link]
Abstract: This study investigates the impact of rainfall on performance degradation of Lidar
through a combination of experiments and analytical modeling. A Micro-electronic-
mechanical system (MEMS) Lidar is employed for single-point detection to analyze how
rainfall affects the range-dependent signal-to-noise ratio (SNR). Consequently, we establish
an ’exponential plus linear’ decay model of the reciprocal of target distance square.
Comparative outdoor and indoor trials are conducted to assess the impacts of raindrops on
Lidar range detectability and scanning angle. The results demonstrate the raindrop-induced
attenuation has a greater impact on the Lidar point cloud than on SNR. In addition, droplets
changing laser direction causes outlier points in the point cloud. The findings underscore the
importance of rain-resilient perception strategies for reliable Lidar performance in adverse
weather conditions.
Keywords: Lidar; MEMS; SNR; Raindrop; Point cloud
Yuhan Liu, Jinlin Ye, Zecheng He, Mingyue Wang, Changjun Wang, Jie Lang, Yidong Zhou, Wei
Zhang,
Deep learning assisted non-invasive lymph node burden evaluation and CDK4/6i
administration in luminal breast cancer,
iScience,
Volume 28, Issue 7,
2025,
112849,
ISSN 2589-0042,
[Link]
([Link]
Abstract: Summary
Precise lymph node evaluation is fundamental to optimize CDK4/6 inhibitor therapy in
luminal breast cancer, particularly given contemporary trends toward axillary surgery de-
escalation that may compromise traditional lymph node staging for recurrence risk
evaluation. The lymph node prediction network (LNPN) was developed as a multi-modal
model incorporating both clinicopathological parameters and ultrasonographic
characteristics for lymph node burden differentiation. In a multicenter cohort of 411
patients, LNPN demonstrated robust performance, achieving an AUC of 0.92 for binary
lymph node burden classification (N0 vs. N+) and 0.82 for ternary lymph node burden
classification (N0/N1–3/N ≥ 4). Notably, among patients undergoing sentinel lymph node
biopsy (SLNB) with confirmed 1–2 metastatic lymph nodes, LNPN predicted high-burden
metastases (N ≥ 4) with an AUC of 0.77. LNPN provided a non-invasive method to assess
lymph node metastasis and recurrence risk, potentially reducing unnecessary axillary lymph
node dissection (ALND), and facilitating decision-making regarding the intervention of
CDK4/6i in luminal breast cancer patients.
Keywords: Cancer; Machine learning
Roy W. Martin, David A. Gilbert, Fred E. Silverstein, Michele Deltenre, Guido Tytgat,
Rhealond K. Gange, John Myers,
An endoscopic Doppler probe for assessing intestinal vasculature,
Ultrasound in Medicine & Biology,
Volume 11, Issue 1,
1985,
Pages 61-69,
ISSN 0301-5629,
[Link]
([Link]
Abstract: Flexible fiberoptic endoscopes permit the physician to inspect the mucosal surface
of the upper gastrointestinal tract and colon. However, this visual inspection provides little
information about the underlying vascular supply to the intestinal wall. We tested the
hypothesis that a Doppler probe could be constructed small enough to pass through the
biopsy channel of a fiber endoscope and be used with it while performing endoscopy. The
purpose would be to determine the location of patent arteries or veins, determine the
magnitude and waveforms of the velocity in them, and estimate their contribution or
potential contribution to intestinal bleeding. For this purpose, a miniature catheter probe
(1.8 mm O.D. and 2 m in length) and an electronic range limited pulsed Doppler unit were
developed. This probe and unit were studied in a series of 13 dogs to determine efficacy of
detecting arterial and enous flow and to test the safety of the device. The duodenum was
surgically exposed and opened in the region of the common bile duct (CBD). Arteries and
veins surrounding the CBD were studied. Particular attention was directed to arterial
structures which clinically pose a risk of bleeding when performing endoscopic papillotomy,
a therapeutic technique in which the papilla of Vater is cut to release bile duct stones. The
results of the study revealed that the probe could indeed detect arterial and venous
structures accurately. There was no evidence that the probe produced any injury to the
common bile duct or pancreas by histological or serum amylase studies and the device was
determined safe and suitable for clinical evaluation.
Keywords: Ultrasound; Intestinal blood flow; Papillotomy; Varices; Blood velocity; Doppler;
Endoscope; Catheter probe
Zhaochun Ding, Xiang Li, Jiang Wu, Jinshuo Liu, Lipeng Wang, Yu Tian, Yanhu Zhang, Xuewen
Rong, Yibin Li,
External-pipe-climbing piezoelectric actuator with high climbing/towing capability and
untethered movement,
International Journal of Mechanical Sciences,
Volume 306,
2025,
110840,
ISSN 0020-7403,
[Link]
([Link]
Abstract: To accomplish high climbing/towing capability and untethered movement, a
miniature external-pipe-climbing piezoelectric actuator (MEPCPA) is developed by
integrating a pair of wing-shaped transducers driven by piezoelectric stack plates and an
onboard circuit. Here, the transducers provide the climbing and clamping functions with the
driving feet and the spring, respectively; these interestingly imitate the propelling and
hugging functions of the sloth’s lower and upper limbs. The micro controller, boost module,
and transistors arranged in the H-bridge shape form the minimum system of a lightweight
onboard circuit. To verify our proposal, first, by constructing a vibration model, the
transducer was designed to enhance the driving force without excessively increasing the
weight. Meanwhile, the friction coefficient was modified by considering the surface
roughness to predict the climbing/towing performance. Then, a prototype whose
mechanical part had the size of 52 × 35 × 72 mm3 and the weight of 20.5 g was fabricated
for performance assessment. In a tethered manner, the MEPCPA climbed up the glass tube
vertically, towed the maximal weight of 120 g (equal to 5.9 times the mechanical part’s
weight), and yielded the maximal speed of 103.8 mm/s. Installed with a 12-V 300-mAh
battery, the MEPCPA successfully climbed up the tube having the tilting angle of 45° with the
ground and it produced the maximal towing weight, the maximal climbing speed, and the
minimal stepwise displacement of 20 g, 18 mm/s, and 0.36 μm, respectively, at the tilting
angle of 30° To the best of our knowledge, this study is an initial report regarding
piezoelectric actuators climbable in an untethered manner, and provides fundamental
technique for designing miniature piezoelectric actuators potentially applicable to narrow
environments particularly lacking external power source.
Keywords: Piezoelectric driving; Miniature actuator; Bioinspired actuation; Climbing
piezoelectric actuator; Towing capability; Untethered movement
Liu Zhi, Chen Nan, Le Dexiang, Lai Qingrong, Li Bin, Wu Jian, Song Yunfeng, Liu Yande,
Acoustic vibration multi-domain images vision transformer (AVMDI-ViT) to the detection of
moldy apple core: Using a novel device based on micro-LDV and resonance speaker,
Postharvest Biology and Technology,
Volume 211,
2024,
112838,
ISSN 0925-5214,
[Link]
([Link]
Abstract: Moldy-core is a common internal disease in apples, and apples infected with this
disease cannot be directly identified according to their external characteristics. In this study,
a novel acoustic vibration device based on micro-LDV, resonance speaker and microphone
was employed to detect moldy-core in apples, the acoustic vibration signals of healthy
apples and apples with different degrees of moldy-core were converted into acoustic
vibration multi-domain images (AVMDI), which consisted of time-frequency images
generated through continuous wavelet transform (CWT), as well as time-domain and
frequency-domain images generated through Gramian Angular Field (GAF). Subsequently
the combination of AVMDI and Vision Transformer (ViT) was applied to the identification of
internal defects in fruits. The outcomes evince that the classification efficacy of the model,
amalgamating sound and vibration signals, surpasses that of models reliant on solitary
sound or vibration signals. The AVMDI-ViT model achieved an overall classification accuracy
of 97.96 %. Specifically, it achieved 100 % accuracy in identifying healthy apples, 94.74 %
accuracy in identifying mild moldy-core apples (≤ 7 %), 97.50 % accuracy in identifying
moderate moldy-core apples (> 7 % and ≤ 15 %), and 100 % accuracy in identifying severe
moldy-core apples (> 15 %). The proposed method demonstrates a high level of accuracy in
the identification of moldy-core apples, while also offering advantages in terms of simplicity,
speed, cost-effectiveness and non-contact detection.
Keywords: Moldy core; Apple; Micro-LDV; Resonance speaker; Acoustic vibration multi-
domain images; Vision Transformer
Chenglin Yao, Jianfeng Ren, Ruibin Bai, Heshan Du, Jiang Liu, Xudong Jiang,
Progressively-orthogonally-mapped EfficientNet for action recognition on time-range-
Doppler signature,
Expert Systems with Applications,
Volume 255, Part D,
2024,
124824,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Although 2D radar signal representations, such as spectrograms and range-Doppler
maps have been widely used for target recognition, 3D time-range-Doppler (TRD) has been
less studied, partially because of the difficulties in extracting features from the TRD
representation, i.e., shallow 3D neural networks have limited discriminant power, but
repeatedly applying 3D convolutions will lead to an oversized 3D network. A hybrid 3D–2D
network architecture, Progressively-Orthogonally-Mapped EfficientNet (POMEN), is
proposed to address these challenges. More specifically, the proposed POMEN utilizes 3D
convolutions in the earlier stages to capture the information embedded in the sparse 3D
TRD representation, and to avoid the oversized feature map caused by excessively applying
3D convolutions, we propose to progressively map the 3D features into three sets of 2D
features corresponding to the range-time signature, range-Doppler map and time-Doppler
signature (spectrogram), respectively. Subsequently, 2D EfficientNet blocks were designed to
extract discriminant information from the three sets of 2D feature maps. This hybrid 3D–2D
network design effectively extracts features from the 3D TRD representation, thereby
avoiding oversized features from full-sized 3D networks and the information loss of 2D
networks on 2D representations. Finally, a homogeneous gated fusion network was designed
to fuse the three sets of 2D features. The proposed method was evaluated on the UGRS,
MIMOGR, and mmWRWD datasets. The experimental results for all datasets demonstrate
that the proposed POMEN significantly and consistently outperforms the state-of-the-art
models in both 2D and 3D representations.
Keywords: Radar activity recognition; Progressively-orthogonally-mapped EfficientNet;
Homogeneous gated fusion; Time-range-Doppler representation
Yongchao Zhu, Qiuling Lu, Maorong Ge, Xiaochuan Qu, Tingye Tao, Kegen Yu, Shuiping Li,
Attention enhanced ResNet for ocean surface wind speed retrieval using CYGNSS
observables,
Advances in Space Research,
2025,
,
ISSN 0273-1177,
[Link]
([Link]
Abstract: Global Navigation Satellite System Reflectometry (GNSS-R) has emerged as a
pivotal technique for ocean surface wind speed retrieval; however, establishing robust multi-
parameter retrieval models remains challenging due to the nonlinear relationships between
GNSS-R observables and geophysical variables. An Attention-enhanced Residual Network
(Att-ResNet) is proposed to address this challenge, leveraging Cyclone Global Navigation
Satellite System (CYGNSS) bistatic radar data for wind speed estimation. The CYGNSS
datasets were processed to extract multi-parameter observables, including Delay-Doppler
Maps (DDMs), normalized bistatic radar cross-section (NBRCS), and incidence angle, which
served as inputs for training wind speed retrieval models using diverse backbone
architectures (e.g., ResNet and AlexNet). Ablation experiments employing the Att-ResNet
framework were systematically conducted, with ERA5 (European Centre for Medium-Range
Weather Forecasts Reanalysis 5) and CCMP (Cross-Calibrated Multi-Platform) wind products
providing benchmark validation. Comparative analysis revealed that the Att-ResNet-retrieved
wind speeds exhibited strong spatiotemporal consistency with ERA5 and CCMP data.
Quantitative evaluations showed root mean square errors (RMSEs) of 1.379 m/s (ERA5) and
1.390 m/s (CCMP), with minimal biases (−0.069 m/s and −0.014 m/s, respectively) and
unbiased RMSEs (ubRMSEs) of 1.377 m/s and 1.390 m/s. The study demonstrates that the
Att-ResNet architecture, through its attention-driven feature selection and residual learning
mechanisms, significantly enhances spaceborne GNSS-R wind retrieval accuracy. This
artificial intelligence-driven framework establishes a new paradigm for high-resolution
spatiotemporal ocean surface wind monitoring, demonstrating the transformative potential
of deep learning in advancing GNSS-R applications.
Keywords: Residual network; GNSS-R; Wind speed; Deep learning; CYGNSS
Fnu Neha, Deepshikha Bhati, Deepak Kumar Shukla, Sonavi Makarand Dalvi, Nikolaos
Mantzou, Safa Shubbar,
An analytics-driven review of U-Net for medical image segmentation,
Healthcare Analytics,
Volume 8,
2025,
100416,
ISSN 2772-4425,
[Link]
([Link]
Abstract: Medical imaging (MI) plays a vital role in healthcare by providing detailed insights
into anatomical structures and pathological conditions, supporting accurate diagnosis and
treatment planning. Noninvasive modalities, such as X-ray, magnetic resonance imaging
(MRI), computed tomography (CT), and ultrasound (US), produce high-resolution images of
internal organs and tissues. The effective interpretation of these images relies on the precise
segmentation of the regions of interest (ROI), including organs and lesions. Traditional
methods based on manual feature extraction are time-consuming, inconsistent, and not
scalable. This review explores recent advances in artificial intelligence (AI)-driven
segmentation, focusing on Convolutional Neural Network (CNN) architectures, particularly
the U-Net family and its variants—U-Net++, and U-Net 3+. These models enable automated,
pixel-wise classification across modalities and have improved segmentation accuracy and
efficiency. The review outlines the evolution of U-Net architectures, their clinical integration,
and offers a modality-wise comparison. It also addresses challenges such as data
heterogeneity, limited generalizability, and model interpretability, proposing solutions
including attention mechanisms and Transformer-based designs. Emphasizing clinical
applicability, this work bridges the gap between algorithmic development and real-world
implementation.
Keywords: Medical image analysis; Deep learning models; Image segmentation; Healthcare
imaging; Pattern recognition; Artificial intelligence
Lingwei Xu, Haiyang Sun, Kai Wang, Gaofeng Nie, Zhe Chen, T. Aaron Gulliver,
An intelligent wireless sensing algorithm for complex cross-domain scenarios based on DB-
FA-YoLov6,
Expert Systems with Applications,
Volume 296, Part A,
2026,
128912,
ISSN 0957-4174,
[Link]
([Link]
Abstract: Wireless sensing technology can identify human motion via feature information
from WiFi signals. The popularity of smartphones, wearable devices, and other smart
devices has increased the use of wireless sensing in fields such as smart homes, smart
healthcare, human–computer interaction, and autonomous vehicles. However, the mobile
communication environment is complex and dynamic which makes wireless sensing
challenging. The issues include low model sensing accuracy, poor scene generalization
ability, and high environmental dependence. Therefore, this paper proposes a cross-domain
intelligent wireless sensing algorithm based on a double branch frequency attention
mechanism Yolov6 network called DB-FA-YoLov6. This integrates a Yolov6 neural network,
frequency attention module, and residual module to provide efficient extraction of signal
features and enhance model generalization. The goal is to reduce the effect of the
environment on sensing tasks and improve robustness, portability, and cross-domain
accuracy. The DB-FA-YOLOv6 model integrates two types of residual modules, BasicBlock and
Bottleneck. It replaces the large modules in the Yolov6 network model with lightweight
structures, which can decrease the number of parameters, improve the efficiency of model
training and testing, and reduce the complexity. Compared with current sensing algorithms
such as Vision Transformer Network for Multiple Vision Tasks (ViT-MVT), Environment
Independent (EI), and Joint Adversarial Domain Adaptation (JADA), the proposed DB-FA-
YOLOv6 algorithm has better sensing accuracy, sensing efficiency, and cross-domain
performance. For the in-domain scenario, the proposed algorithm achieves improvements of
10.0 % in sensing accuracy and 10.1 % in sensing efficiency. The sensing accuracy of the
proposed algorithm in cross-domain scenarios, namely location and orientation, is improved
by 10.5 % and 9.7 %, and the sensing efficiency is improved by 7.0 % and 52.1 %,
respectively.
Keywords: Intelligent wireless sensing; Cross-domain sensing; Attention mechanism; Double-
branch Yolov6 neural network
Adrito Das, Danyal Z. Khan, Dimitrios Psychogyios, Yitong Zhang, John G. Hanrahan,
Francisco Vasconcelos, You Pang, Zhen Chen, Jinlin Wu, Xiaoyang Zou, Guoyan Zheng, Abdul
Qayyum, Moona Mazher, Imran Razzak, Tianbin Li, Jin Ye, Junjun He, Szymon Płotka, Joanna
Kaleta, Amine Yamlahi, Antoine Jund, Patrick Godau, Satoshi Kondo, Satoshi Kasai, Kousuke
Hirasawa, Dominik Rivoir, Stefanie Speidel, Alejandra Pérez, Santiago Rodriguez, Pablo
Arbeláez, Danail Stoyanov, Hani J. Marcus, Sophia Bano,
PitVis-2023 challenge: Workflow recognition in videos of endoscopic pituitary surgery,
Medical Image Analysis,
Volume 106,
2025,
103716,
ISSN 1361-8415,
[Link]
([Link]
Abstract: The field of computer vision applied to videos of minimally invasive surgery is ever-
growing. Workflow recognition pertains to the automated recognition of various aspects of a
surgery, including: which surgical steps are performed; and which surgical instruments are
used. This information can later be used to assist clinicians when learning the surgery or
during live surgery. The Pituitary Vision (PitVis) 2023 Challenge tasks the community to step
and instrument recognition in videos of endoscopic pituitary surgery. This is a particularly
challenging task when compared to other minimally invasive surgeries due to: the smaller
working space, which limits and distorts vision; and higher frequency of instrument and step
switching, which requires more precise model predictions. Participants were provided with
25-videos, with results presented at the MICCAI-2023 conference as part of the Endoscopic
Vision 2023 Challenge in Vancouver, Canada, on 08-Oct-2023. There were 18-submissions
from 9-teams across 6-countries, using a variety of deep learning models. The top
performing model for step recognition utilised a transformer based architecture, uniquely
using an autoregressive decoder with a positional encoding input. The top performing model
for instrument recognition utilised a spatial encoder followed by a temporal encoder, which
uniquely used a 2-layer temporal architecture. In both cases, these models outperformed
purely spatial based models, illustrating the importance of sequential and temporal
information. This PitVis-2023 therefore demonstrates state-of-the-art computer vision
models in minimally invasive surgery are transferable to a new dataset. Benchmark results
are provided in the paper, and the dataset is publicly available at:
[Link]
Keywords: Endoscopic vision; Instrument recognition; Step recognition; Surgical AI; Surgical
vision; Workflow analysis
Haoming Feng, Huaqing Li, Wenwen Zhu, Denghao Li, Yukun Huang,
Micro-motion enhanced multi-person activity recognition with millimeter-wave radar,
Measurement,
Volume 258, Part B,
2026,
119090,
ISSN 0263-2241,
[Link]
([Link]
Abstract: As a non-contact sensing device, millimeter-wave radar exhibits unique strengths
in human activity recognition (HAR). Existing methods rely on micro-Doppler signatures for
activity classification, but they often encounter feature aliasing in multi-person activity
recognition (MPAR) scenarios. Although point cloud-based approaches can distinguish
individual targets, they primarily extract static morphological features, neglecting the micro-
motion information of human joints, which is crucial for accurate activity recognition. To
address these limitations, we proposes an innovative MPAR framework that integrates
spatial point clouds and micro-motion features. First, an improved point cloud data
association algorithm is applied to achieve multi-target point cloud feature separation,
followed by a dynamic projection mechanism to construct time–Doppler feature maps.
Then, a torso micro-motion enhancement algorithm is designed to enhance the details of
human body movements. Finally, a CNN-LSTM hybrid network architecture with a temporal-
attention is constructed for action classification. Experimental results show that the
proposed micro-motion enhancement algorithm improves recognition accuracy by 27.1%
and 2.3%, compared to two traditional time–frequency analysis methods. Furthermore,
MPAR task in occlusion scenarios achieves recognition accuracy of 93.5%. In summary,
proposed framework not only retains the inherent advantages of millimeter-wave radar but
also significantly enhances multi-person activity recognition in complex scenarios.
Keywords: Human activity recognition (HAR); Multi-person activity recognition (MPAR);
Feature separability; Temporal attention; Occlusion scenarios
Vera Lucia Da Silveira Nantes Button,
Chapter 5 - Displacement, Velocity, and Acceleration Transducers,
Editor(s): Vera Lucia Da Silveira Nantes Button,
Principles of Measurement and Transduction of Biomedical Variables,
Academic Press,
2015,
Pages 155-219,
ISBN 9780128007747,
[Link]
([Link]
Abstract: This chapter presents the main methods for measuring displacement, velocity, and
acceleration, commonly used in biomedical determination of other quantities, such as
pressure, flow, and force. This chapter will discuss resistive transducers with particular
emphasis on strain gauges, both metallic and semiconductor, capacitive, piezoelectric and
inductive transducers, with special emphasis on the functioning of linear variable differential
transformer. The functioning principle of transducers for velocity and acceleration
measurements, tachometers, and accelerometers, respectively, will be presented and their
biomedical applications will be exemplified.
Keywords: Resistive transducer; strain gage; capacitive transducer; LVDT; accelerometer
M.E. Pleydell,
Laser Doppler vibration measurement using a polarisation-based device,
Measurement,
Volume 6, Issue 1,
1988,
Pages 10-18,
ISSN 0263-2241,
[Link]
([Link]
Abstract: A non-contacting laser-based displacement and vibration measuring device is
presented. It is based on the coherent detection of the Doppler shift introduced into light
scattered by an optically rough target object. Geometric manipulation of the beams within
the system compensates for the speckle in the scattered light. The sense of the target
motion is derived using a passive technique based on controlled polarisation of the
reference beam. In its simplest mode of operation it has a resolution of half of the laser
wavelength, but this may be increased by a factor of four. Continuous measurements of
displacement of up to 10 cm have been made with an experimental system, with less than
1% error; vibrations of smaller amplitude have been measured with greater accuracy. In the
existing configuration the maximum target velocity in the direction of the measuring beam is
limited by the 100 kHz analogue signal processing roll-off to 32 mm/s; however, this may
easily be extended by using more sophisticated electronics.
Keywords: Laser doppler; vibration; speckle; polarisation
Federico Carlos Gallardo, Jorge Luis Bustamante, Clara Martin, Cristian Marcelo Orellana,
Mauricio Rojas Caviglia, Guillermo Garcia Oriola, Agustin Ignacio Diaz, Pablo Augusto Rubino,
Vicent Quilis Quesada,
Novel Simulation Model with Pulsatile Flow System for Microvascular Training, Research, and
Improving Patient Surgical Outcomes,
World Neurosurgery,
Volume 143,
2020,
Pages 11-16,
ISSN 1878-8750,
[Link]
([Link]
Abstract: Background
Simulation allows surgical trainees to acquire surgical skills in a safe environment. With the
aim of reducing the use of animal experimentation, different alternative nonliving models
have been pursued. However, one of the main disadvantages of these nonliving models has
been the absence of arterial flow, pulsation, and the ability to integrate both during a
procedure on a blood vessel. In the present report, we have introduced a microvascular
surgery simulation training model that uses a fiscally responsible and replicable pulsatile
flow system.
Methods
We connected 30 human placentas to a pulsatile flow system and used them to simulate
aneurysm clipping and vascular anastomosis.
Results
The presence of the pulsatile flow system allowed for the simulation of a hydrodynamic
mechanism similar to that found in real life. In the aneurysm simulation, the arterial flow
could be evaluated before and after clipping the aneurysm using a Doppler ultrasound
system. When practicing anastomosis, the use of the pulsatile flow system allowed us to
assess the vascular flow through the anastomosis, with verification using the Doppler
ultrasound system. Leaks were manifested as “blood” pulsatile ejections and were more
frequent at the beginning of the surgical practice, showing a learning curve.
Conclusions
We have provided a step-by-step guide for the assembly of a replicable and inexpensive
pulsatile flow system and its use in placentas for the simulation of, and training in,
performing different types of anastomoses and intracranial aneurysms surgery.
Keywords: Anastomosis; Aneurysm; Microsurgery; Neurosurgery; Placenta; Training
simulation
Peng Zhu, Kai Chen, Cong Xu, Shuangfei Zhao, Ruiqi Shen, Yinghua Ye,
Development of a monolithic micro chip exploding foil initiator based on low temperature
co-fired ceramic,
Sensors and Actuators A: Physical,
Volume 276,
2018,
Pages 278-283,
ISSN 0924-4247,
[Link]
([Link]
Abstract: The performances of exploding foil initiator has been improved in terms of high
safety and high reliability over the other electrical initiators. This work develops a monolithic
micro-chip exploding foil initiator (McEFI) based on low temperature co-fired ceramic. This
McEFI has a really monolithic construction without any adhesive or bonding parts, which
endows it with the inherent benefits of large volume/low-cost production and high-
reproducibility. Using this method, it can eliminate complicated fabrication processes such as
precise machining, aligning and bonding that are inevitable in conventional manufacturing
process of EFIs. Its primary characteristics for electrical burst, flyer acceleration and
detonating capability are presented and discussed. Results show that McEFI could reliably
detonate hexanitrostilbene at 2.5 kV/0.22 μF, and its performance could be improved by
optimizing design parameters.
Keywords: Exploding foil initiator; Low temperature co-fired ceramic; Monolithic
construction; Firing characteristics
Pengfei Xue, Peng Xiong, Heng Hu, Tao Wang, Mingyu Li, Qingxuan Zeng,
Integration of the exploding foil initiator with capacitor discharge unit and its performance
characterization,
Measurement,
Volume 242, Part C,
2025,
116069,
ISSN 0263-2241,
[Link]
([Link]
Abstract: This study presents a novel integrated exploding foil initiator system (EFIs),
consisting of a planar switch, an EFI chip, and a capacitor. Additionally, we propose an
evaluation method for integrated EFIs that combines numerical simulations with
experimental testing, enabling a thorough assessment of their performance. Under test
conditions of 900 V/0.22 μF, the EFI demonstrated an inductance of 10.2 nH and a resistance
of 107.5 mΩ, measured via the sampling resistance method. The operational behavior of the
EFIs was analyzed through a two-dimensional metal electrical explosion model and a one-
dimensional flyer propulsion model, with the results compared to experimental data. Results
indicated that under 900–1200 V/0.22 μF, the deviations between measured and calculated
values for electrical explosion performance and flyer velocity were both within 5 %,
validating the accuracy of the evaluation approach. Firing tests further confirmed that the
EFIs successfully detonated HNS-IV at 1000 V/0.22 μF.
Keywords: Exploding foil initiator system; Flexible Printed Circuit; Micro-Electro-Mechanical-
Systems; Calculation model
Yanfang Yu, Dadian Wang, Huibo Meng, Jinyu Guo, Zhiying Han,
Spatiotemporal evolution analysis of bubble swarms in gas-liquid static mixers based on an
improved CNN,
International Journal of Heat and Mass Transfer,
Volume 255, Part 1,
2026,
127731,
ISSN 0017-9310,
[Link]
([Link]
Abstract: Static mixers (SM) are significant in multiphase mixing due to their high efficiency,
energy saving and reliability. Komax static mixer (Komax) effectively promotes bubble
breakup and liquid turbulence that leads to the formation of numerous small and medium-
sized bubbles. Three-dimensional features of bubbles with high deformation, overlap and
offset angle under turbulent conditions in SM are difficult to be extracted. In this paper, a
lightweight multidimensional backbone neural network incorporating a deformable
attention transformer is proposed for parallelled and targeted extracting the
multidimensional information of bubbles. The deviation between convolutional neural
network-predicted void fractions (Vf) and experimentally measured values was below 10%.
Compared with the empty pipe, the flow pattern could maintain bubble flow even if the Vf is
up to 0.7 in Komax. The bubble spatiotemporal characteristics and radial Sauter mean
diameter (d32) distribution demonstrated the local slug flow occurred when average velocity
differences between large and small bubble swarms exceeded 0.37 m/s. This phenomenon is
effectively restrained when the bubble velocity uniformity factor remains below 0.1 in
Komax. Chaotic characteristics analysis combined with longitudinal bubble distribution
revealed the bubble flow was stable when the superficial liquid velocity (UL) = 0.0283–
0.0424 m/s and the slip velocity (US) = 0.15–0.25 m/s. Additionally, the comprehensive
breakup efficiency (H) of Komax is 7.9%‒11.8% higher than that of Quatro static mixers at US
= 0.193‒0.259 m/s. Finally, the generalization of the improved model for predicting flow
patterns and extracting bubble information was excellent with the precision exceeding 0.96.
Keywords: Komax static mixer; Neural network; Bubble dynamics; Deformable attention
transformer; Chaos analysis
Yinshen Wang, Zhengxuan Hu, Ping Zhang, Zhihua Fan, Wenming Li, Xuejun An, Xiaochun Ye,
A real-time edge SAR imaging acceleration architecture utilizing multi-level dataflow
parallelism,
Journal of Systems Architecture,
Volume 170,
2026,
103635,
ISSN 1383-7621,
[Link]
([Link]
Abstract: Synthetic Aperture Radar (SAR), a key radar signal processing technology, is widely
deployed on edge devices due to its high resolution, long-range detection, and all-weather,
all-day imaging. However, achieving real-time SAR imaging on resource-constrained edge
platforms is challenging because SAR algorithms involve complex workflows and diverse
operators. Prior works focused on DSP, FPGA, and GPU platforms have struggled to balance
performance and power efficiency. Furthermore, frequent kernel switching necessitates
repeated reconfigurations and memory accesses, increasing latency. To address these
challenges, we propose a SAR-tailored dataflow model enabling multi-level dataflow
parallelism across various operators. First, we introduce a reconfigurable architecture that
integrates customized processing elements optimized for SAR. Second, we propose a multi-
level dataflow model exploiting parallelism at the task, instruction, and node levels.
Additionally, we present an instruction switching mechanism and a preprocessing method
for matrix transposition to reduce kernel switching overhead. Experimental results
demonstrate that for an 8K × 8K image, our approach achieves a processing time of 0.66 s,
with a 37.1× performance improvement over a CPU (i5-9500) and a 1.42× improvement over
a GPU (NVIDIA Orin). Evaluations of SAR operators across diverse scales indicate a 1.45×
performance gain over state-of-the-art reconfigurable architectures featuring dataflow
modeling.
Keywords: Synthetic aperture radar (SAR); Reconfigurable architecture; Dataflow model;
Multi-level parallelism; Real-time imaging
Abdulrhman H. Al-Jebrni, Saba Ghazanfar Ali, Bin Sheng, Huating Li, Xiao Lin, Ping Li,
Younhyun Jung, Jinman Kim, Li Xu, Lixin Jiang, Jing Du,
SynTaskNet: A synergistic multi-task network for joint segmentation and classification of
small anatomical structures in ultrasound imaging,
Computer Vision and Image Understanding,
Volume 263,
2026,
104616,
ISSN 1077-3142,
[Link]
([Link]
Abstract: Segmenting small, low-contrast anatomical structures and classifying their
pathological status in ultrasound (US) images remain challenging tasks in computer vision,
especially under the noise and ambiguity inherent in real-world clinical data. Papillary
thyroid microcarcinoma (PTMC), characterized by nodules ≤1.0 cm, exemplifies these
challenges where both precise segmentation and accurate lymph node metastasis (LNM)
prediction are essential for informed clinical decisions. We propose SynTaskNet, a synergistic
multi-task learning (MTL) architecture that jointly performs PTMC nodule segmentation and
LNM classification from US images. Built upon a DenseNet201 backbone, SynTaskNet
incorporates several specialized modules: a Coordinated Depth-wise Convolution (CDC) layer
for enhancing spatial features, an Adaptive Context Block (ACB) for embedding contextual
dependencies, and a Multi-scale Contextual Boundary Attention (MCBA) module to improve
boundary localization in low-contrast regions. To strengthen task interaction, we introduce a
Selective Enhancement Fusion (SEF) mechanism that hierarchically integrates features across
three semantic levels, enabling effective information exchange between segmentation and
classification branches. On top of this, we formulate a synergistic learning scheme wherein
an Auxiliary Segmentation Map (ASM) generated by the segmentation decoder is injected
into SEF’s third class-specific fusion path to guide LNM classification. In parallel, the
predicted LNM label is concatenated with the third-path SEF output to refine the Final
Segmentation Map (FSM), enabling bidirectional task reinforcement. Extensive evaluations
on a dedicated PTMC US dataset demonstrate that SynTaskNet achieves state-of-the-art
performance, with a Dice score of 93.0% for segmentation and a classification accuracy of
94.2% for LNM prediction, validating its clinical relevance and technical efficacy.
Keywords: Medical image classification; Medical image segmentation; Multi-task learning;
Synergistic learning; Small anatomical structures
Heng Zhao, Zhili Long, Shuyuan Ye, Jianzhong Ju, Yuxiang Li,
Ultrasonic tool shank with multiple vibration mode for micro-nano drilling: Design,
optimization and experiment,
Sensors and Actuators A: Physical,
Volume 366,
2024,
114987,
ISSN 0924-4247,
[Link]
([Link]
Abstract: The machining effects of hole drilling is limited by merely changing the amplitude.
Adjusting ultrasonic vibration frequency to expand the matching range of machining
parameters optimization is necessary to improve hole drilling effects. Therefore, the design,
optimization, and experiment evaluation for a novel integrated, fined, low-cost dual-
frequency ultrasonic tool shank is proposed, which can work in three vibration modes for
micro-nano precision drilling. The principle of the dual-frequency ultrasonic transducer is
introduced. The electro-mechanical equivalent circuit model with step horn is applied to
study the dual-frequency characteristics. The transducer displacement equation model
based on the wave equation theory is deduced to find the shared vibration node of the first
and third frequency. Sensitivity analysis by equivalent circuit model and Finite element
method (FEM) simulation of the key parameter is both conducted to optimize ultrasonic
transducer geometry dimension. The vibration node of mounting flange is specified in an
identical position for the dual vibration modes. A prototype of ultrasonic tool shank is
fabricated and evaluated in experiments. It demonstrates that the dual working frequency of
the ultrasonic transducer is 34.9 kHz and 104.5 kHz in the first and third mode resonant
vibration, which is benefit to the low and high frequency drilling. Three working modes of
the low, high, and coupling frequency vibration is measured by a self-developed ultrasonic
generator. When 40 V voltage is excited, the vibration amplitudes of the low and high
frequency are 17.8 µm, and 9.3 µm, respectively, which can meet with the drilling
requirement. Moreover, the coupling vibration with the low and high mode are excited
simultaneously, and the motion trajectory of the coupling vibration is observed in a “M”
shape, which is a novel trajectory for the precision micro-nano drilling. Micro drilling
experiments show that the 35 kHz ultrasonic vibration drilling can achieve better machining
effect than conventional drilling (CD). The 105 kHz ultrasonic drilling outperforms the CD,
35 kHz and coupling ultrasonic drilling at cutting force reduction. The coupling ultrasonic
vibration drilling achieves the lowest chipping diameter. It is to be noted that the 105 kHz-
8 µm produce the fined smoothly surface. The 35 kHz, 105 kHz and coupling ultrasonic
drilling can produce better surface micro morphology with less quantity and size of craters
than that of CD.
Keywords: Ultrasonic tool shank; Dual-frequency; Multiple vibration mode; Ultrasonic
transducer
Dingcheng Ji, Jing Lin, Fei Gao, Jiadong Hua, Wenhao Li,
A deep learning-based spatial gradient reconstruction method for efficient damage
identification in composite with high-sparsity Lamb wavefield,
Mechanical Systems and Signal Processing,
Volume 224,
2025,
112018,
ISSN 0888-3270,
[Link]
([Link]
Abstract: The structural integrity and safety of carbon fiber reinforced plastics (CFRP) are
vulnerable to delamination, which is often imperceptible to the naked eye. Although the
Scanning Laser Doppler Vibrometer (SLDV) has shown promise in damage quantification of
CFRP, its time-consuming measurement process limits its application in engineering
scenarios. To address this, we introduce a novel damage index, the spatial gradient, which
captures the interaction between delamination and the wavefield. We have also developed a
neural network capable of reconstructing the spatial gradient directly from high-sparsity
Lamb wavefield data obtained at an extremely low spatial sampling rate, thereby
significantly reducing measurement time. To enhance the network’s capability to detect
wavefield anomalies, we employ the cross-attention technique, allowing for the direct
injection of shallow features representing local wavefield distortions caused by damage into
the decoder. Additionally, we integrate multiple reconstruction layers to guide the wavefield
reconstruction process, ensuring meaningful information is captured at each stage. Our
method achieves substantial improvements in reconstruction accuracy, increasing from 70 %
to 92 % in single-damage scenario and from 14 % to 72 % in multi-damage scenario
compared to the previous state-of-the-art techniques. By using the reconstructed spatial
gradient field for damage imaging through spatial covariance analysis, our approach
demonstrates its feasibility and generalizability across various damage locations. This
suggests its potential as a reliable solution for fast and accurate damage characterization,
reducing the measurement burden and enhancing practical applicability.
Keywords: High-sparsity wavefield reconstruction; Lamb waves; Deep learning; Spatial
gradient imaging
Binglei Yue, Aili Jiang, Chun Yang, Junwei Lei, Heng Liu, Yin Zhang,
Deep Learning-Enhanced Human Sensing with Channel State Information: A Survey,
Computers, Materials and Continua,
Volume 86, Issue 1,
2025,
Pages 1-28,
ISSN 1546-2218,
[Link]
([Link]
Abstract: With the growing advancement of wireless communication technologies, WiFi-
based human sensing has gained increasing attention as a non-intrusive and device-free
solution. Among the available signal types, Channel State Information (CSI) offers fine-
grained temporal, frequency, and spatial insights into multipath propagation, making it a
crucial data source for human-centric sensing. Recently, the integration of deep learning has
significantly improved the robustness and automation of feature extraction from CSI in
complex environments. This paper provides a comprehensive review of deep learning-
enhanced human sensing based on CSI. We first outline mainstream CSI acquisition tools
and their hardware specifications, then provide a detailed discussion of preprocessing
methods such as denoising, time–frequency transformation, data segmentation, and
augmentation. Subsequently, we categorize deep learning approaches according to sensing
tasks—namely detection, localization, and recognition—and highlight representative models
across application scenarios. Finally, we examine key challenges including domain
generalization, multi-user interference, and limited data availability, and we propose future
research directions involving lightweight model deployment, multimodal data fusion, and
semantic-level sensing.
Keywords: Channel State Information (CSI); human sensing; human activity recognition;
deep learning
Chen Nan, Liu Zhi, Le Dexiang, Lai Qingrong, Jiang Bingnian, Li Bin, Wu Jian, Song Yunfeng,
Liu Yande,
Prediction of yellow flesh peach firmness using a novel device and data augmentation
acoustic vibration multi-domain images array Swin Transformer (DA-AVMDIA-SwinT),
Computers and Electronics in Agriculture,
Volume 235,
2025,
110402,
ISSN 0168-1699,
[Link]
([Link]
Abstract: Firmness is an important indicator closely related to the ripeness of yellow flesh
peaches. Non-destructive detection of yellow flesh peach firmness is beneficial for managing
yellow flesh peaches during storage and for quality grading. In this study, a novel acoustic
vibration device based on micro-LDV, a microphone, and a resonance loudspeaker was used
to predict the firmness of yellow flesh peaches. The acoustic vibration response signals of
yellow flesh peaches with different firmness were transformed into acoustic vibration multi-
domain images arrays (AVMDIA), and Swin Transformer (SwinT) model was employed to
extract the acoustic vibration response features of the fruits. Firstly, the prediction
performance of the models built with acoustic multi-domain images array (AMDIA),
vibration multi-domain images array (VMDIA) and AVMDIA as inputs were compared.
Subsequently, the effects of using data augmentation (DA) and not using DA on model
performance were compared. Finally, the prediction performance of the SwinT model, the
Vision Transformer (ViT) model, and the Resnet50 model (CNN-based) for yellow flesh peach
firmness was compared. The results showed that the SwinT model based on data
augmentation acoustic vibration multidomain data array (DA-AVMDIA-SwinT) gave the best
prediction of firmness for yellow flesh peaches, with RP2 = 0.951, RMSEP = 0.515 N/mm,
RPDP = 4.524. In this paper, a method is reported for converting fruit acoustic vibration
response signals into two-dimensional images for fruit firmness prediction is reported for
the first time, which can accurately predict fruit firmness in a simple, fast and cost-effective
manner.
Keywords: Firmness prediction; Yellow flesh peach; Acoustic vibration multi-domain images
array; Data augmentation; Swin transformer
Yuxiang Li, MARIIA KIREEVA, Xicheng Liu, Zhili Long, Shuyuan Ye, Jianzhong Ju,
Design and performance analysis of bidirectional vibration ultrasonic transducer for wire
bonding,
Applied Acoustics,
Volume 238,
2025,
110791,
ISSN 0003-682X,
[Link]
([Link]
Abstract: Ultrasonic frequency, amplitude and vibration mode are the key factors affecting
the stability and reliability of ultrasonic wire bonding. Conventional wire bonding is realized
by ultrasonic transducer (UT) with single frequency and longitudinal vibration. To optimize
the bonding process and achieve high-performance bonding joints, we propose an UT that
utilizes two longitudinal vibration and one bending vibration mode, which realized by the full
PZT and regional polarization PZT. By employing a step-type flexible structure, the flanges for
three vibration modes are optimized at the same node position. To drive the UT, the
amplifier module circuit with 39.98 times amplification and 130 kHz bandwidth is designed.
Under the 9 kg·cm torque for assembling UT, the longitudinal vibration frequencies of the UT
are 71.2 kHz and 122.9 kHz, and the bending vibration frequency is 43.1 kHz. The self-
developed amplifier module is designed to drive the UT for both single and coupled
vibration, and the test results show that the amplitude of the bending vibration is 7.49 µm at
a driving voltage of 20 V, and the amplitudes of the 1st longitudinal vibration and the 2nd
longitudinal vibration are 5.72 µm and 4.44 µm, respectively, which satisfy the amplitude
requirements for wire bonding. Furthermore, the coupled vibration trajectory of the UT
forms a parallelogram, and the vibration area can be adjusted by changing the amplitude.
Experimental results have shown that the bending mode is excited by regionally polarized
PZT, achieving sufficient vibration in both x/y directions while ensuring a small and
lightweight structure. This provides a potential application solution for the new bonding
method of wire bonding.
Keywords: Wire bonding; Ultrasonic transducer; Bidirectional vibration; Regional polarization
PZT; Amplifier module
Van Ngoc Dang, Ngoc Chau Hoang, Quoc Cuong Nguyen, Minh Thuy Le,
Advancing robust human activity recognition via informative mmWave radar characteristics
and a lightweight spatio-spectro-temporal network,
Measurement,
Volume 256, Part A,
2025,
118056,
ISSN 0263-2241,
[Link]
([Link]
Abstract: Human activity recognition (HAR) is increasingly important in aiding our daily life,
with millimeter-wave (mmWave) radar sensors emerging as a promising noninvasive solution
thanks to their excellent spatial and velocity resolution. Although existing radar-based
systems have shown strong performance, they primarily focus on micro-Doppler signatures
while neglecting angle information, which can hinder practical deployment in real-world
scenarios. Moreover, current state-of-the-art recognition models using mmWave radar often
require substantial computational resources, making integration into resource-constrained
devices challenging. This work proposes an efficient radar-based HAR system that leverages
angle and spectro-temporal information from micro-Doppler signatures. Our system utilizes
a multi-channel micro-Doppler representation corresponding to the number of virtual
antenna receivers as input. Then, a lightweight dilated convolutional network, namely SST-
DCN, extracts spatial-aware multi-scale spectro-temporal information through time-
frequency dilated convolutions. Experimental results on our real-world dataset demonstrate
the superiority of our approach compared to conventional features and other state-of-the-
art radar-based HAR systems.
Keywords: Human activity recognition; Millimeter-wave radar; Deep learning; Lightweight
network; Dilated convolution
Yu Dou, Yongjian Li, Shuaichao Yue, Yang Li, Jianguo Zhu,
Measurement of alternating and rotational magnetostrictions of Non-oriented silicon steel
sheets,
Journal of Magnetism and Magnetic Materials,
Volume 571,
2023,
170566,
ISSN 0304-8853,
[Link]
([Link]
Abstract: Extensive numerical analyses and experimental tests have shown that while major
regions of magnetic cores in electromagnetic devices, such as transformers and rotating
electric machines, are dominated by the alternating magnetic fields, the rotating magnetic
fields exist in the corner joints of multi-phase transformers and the yoke area behind the
teeth of rotating electric machines. These magnetic fields cause core losses due to hysteresis
and eddy currents, mechanical vibration and acoustic noises due to magnetostriction. Many
studies have been reported in the literature on alternating and rotational core losses and
alternating magnetostriction, but not so much on rotational magnetostriction, especially in
non-oriented (NO) silicon steel sheets. The vibration and noise caused by rotational
magnetostriction are much higher than those caused by alternating magnetostriction. This
paper reports the development of a magnetostriction measurement system consisting of a
2D symmetrical single sheet tester (SST) and resistance strain gauges. The selection
considerations of resistance strain gauges are summarized. The alternating and rotational
magnetostriction characteristics of a NO silicon steel, B35A300 (0.35 mm), are measured.
Because the rotating magnetic field in the stator is not always ideally circular but more often
elliptical, the magnetostriction under elliptical magnetizations with different axis ratios is
also measured. The magnetostriction anisotropy of the NO steel sheet is observed and
analyzed. The elongation and contraction under alternating, circular and elliptical rotating
magnetizations with different axis ratios are discussed. The results can provide data support
for calculating and designing multi-phase transformers and rotating electric machines under
different types of magnetizations. The research can provide theoretical and practical
guidance for mitigating mechanical vibration and acoustic noise caused by magnetostriction.
Keywords: Rotational magnetostriction; Non-oriented silicon steel sheets; Strain gauge
Jitao Zhang, Han Qiao, Qingfang Zhang, Bingfeng Ge, D.A. Filippov, Jie Wu, Fang Wang, Jiagui
Tao, Jing Chen, Liying Jiang, Lingzhi Cao,
Compact magnetoelectric power splitter with high isolation using ferrite/piezoelectric
transformer composite,
Journal of Magnetism and Magnetic Materials,
Volume 574,
2023,
170691,
ISSN 0304-8853,
[Link]
([Link]
Abstract: A compact, passive magnetoelectric (ME) power splitter, with a
ferrite/piezoelectric transformer bilayer composite with a coil wound around it, as well as a
structure-constructing strategy, was presented and developed in this research. The
impedance differences in a Rosen-type PT induced by integrated transverse/longitudinal
polarizations facilitated the device realization with higher isolation rather than elementary
LCR lumped elements. Furthermore, the measurements of the material properties and
electrical resonance behaviors for the presented ME splitter were implemented, and the
power-dividing capabilities were verified by measuring the output and corresponding power
for each port of the device. The experimental results demonstrated that the power for Port I
in the ME power splitter reached its maximum values of 154.26 nW at R = 6 kΩ and 475.95
nW at R = 5.5 kΩ. Correspondingly, the power for Port II reached its maximum values of
12.73 nW at R = 30 kΩ and 27.85 nW at R = 12 kΩ. Consequently, a power division ratio of
3:1 was obtained under optimum load resistance and EMR conditions with constant input
power. These results provided an innovative approach to constructing novel functional
power electronics with solid-state materials while suggesting the promising applications for
some special scenarios such as powering and controlling multi-channel LCD strings.
Keywords: Power splitter; Magnetoelectric composite; Piezoelectric transformer
Dapeng Zhang, Yifan Xie, Yining Zhang, Zhengjie Liang, Yutao Tian,
Experimental Advances in Airfoil Dynamic Stall and Transition Phenomena,
Fluid Dynamics and Materials Processing,
Volume 21, Issue 4,
2025,
Pages 697-739,
ISSN 1555-256X,
[Link]
([Link]
Abstract: Airfoil structures play a crucial role across numerous scientific and technological
disciplines, with the transition to turbulence and stall onset remaining key challenges in
aerodynamic research. While experimental techniques often surpass numerical simulations
in accuracy, they still present notable limitations. This paper begins by elucidating the
fundamental principles of transition, dynamic stall, and airfoil behavior. It then provides a
systematic review of six major experimental methodologies and examines the emerging role
of artificial intelligence in this domain. By identifying key challenges and limitations, the
study proposes strategic advancements to address these issues, offering a foundational
framework to guide future research in airfoil structures and related fields.
Keywords: Airfoil; dynamic stall; transition; experimental methodologies; artificial
intelligence
Jiale Ren, Hengyi Li, Aihui Wang, Kenshi Saho, Lin Meng,
Radar-based gait analysis by Transformer-liked network for dementia diagnosis,
Biomedical Signal Processing and Control,
Volume 91,
2024,
105986,
ISSN 1746-8094,
[Link]
([Link]
Abstract: Providing reliable diagnostic evidence to doctors while minimizing financial and
physical burdens on patients is a prominent focus of current research. Gait features show
potential as a clinical marker for dementia diagnosis. Radar is capable of efficient,
contactless collecting human motion. This paper proposes a novel radar-based gait analysis
strategy with a Transformer-liked network for dementia diagnosis. The gait data are collected
by Micro-Doppler radar. Then Welch’s power spectral density estimation is adopted to
obtain the frequency features while unifying and compressing data size. The network is
designed to explore the relationship between dementia and gait features. In the network,
1D convolution is crucial in extracting local features and encoding features into deeper
dimensions. The attention-based module, inspired by the encoder of the Transformer,
possesses an edge in capturing long-sequence dependencies. The dual-stage gating
mechanism enhances the discriminative power of the learned representations by fine-tuning
the weights of extracted features. To validate the effectiveness of the proposed strategy,
comparative experiments are performed with prevailing networks in both time and
frequency domains. Experimental results demonstrate the superiority of the frequency
domain processing method over the time domain processing method and fusion time-
frequency processing method. Notably, the proposed model outperforms others, achieving
the highest accuracy of 94.93% in frequency domain processing-based experiments — 5.91%
and 4.25% higher than the highest accuracies in time domain processing-based experiments
and fusion time-frequency processing-based experiments respectively. The overall findings
illustrate that our proposal can provide a reliable reference for dementia diagnosis
effectively.
Keywords: Dementia diagnosis; Gait analysis; Power spectral density estimation; Attention
mechanism; Convolutional neural network; Gating mechanism
Michael A. Rothfuss, Jignesh V. Unadkat, Michael L. Gimbel, Marlin H. Mickle, Ervin Sejdić,
Totally Implantable Wireless Ultrasonic Doppler Blood Flowmeters: Toward Accurate
Miniaturized Chronic Monitors,
Ultrasound in Medicine & Biology,
Volume 43, Issue 3,
2017,
Pages 561-578,
ISSN 0301-5629,
[Link]
([Link]
Abstract: Totally implantable wireless ultrasonic blood flowmeters provide direct-access
chronic vessel monitoring in hard-to-reach places without using wired bedside monitors or
imaging equipment. Although wireless implantable Doppler devices are accurate for most
applications, device size and implant lifetime remain vastly underdeveloped. We review past
and current approaches to miniaturization and implant lifetime extension for wireless
implantable Doppler devices and propose approaches to reduce device size and maximize
implant lifetime for the next generation of devices. Additionally, we review current and past
approaches to accurate blood flow measurements. This review points toward relying on
increased levels of monolithic customization and integration to reduce size. Meanwhile,
recommendations to maximize implant lifetime should include alternative sources of power,
such as transcutaneous wireless power, that stand to extend lifetime indefinitely. Coupling
together the results will pave the way for ultra-miniaturized totally implantable wireless
blood flow monitors for truly chronic implantation.
Keywords: Battery-less; Blood flow monitor; Flowmeter; Free flap; Wireless power;
Transcutaneous wireless power
Jens Schwarz, Brian Hutsel, Thomas Awe, Bruno Bauer, Jacob Banasek, Eric Breden, Joe
Chen, Michael Cuneo, Katherine Chandler, Karen DeZetter, Mark Gilmore, Matthew Gomez,
Hannah Hasson, Maren Hatch, Nathan Hines, Trevor Hutchinson, Deanna Jaramillo, Christine
Kalogeras Loney, Ian Kern, Derek Lamppa, Diego Lucero, Larry Lucero, Keith LeChien, Mike
Mazarakis, Thomas Mulville, Robert Obregon, John Porter, Pablo Reyes, Alex Sarracino,
Daniel Scoglietti, Gabriel Shipley, Trevor Smith, Brian Stoltzfus, William Stygar, Adam Steiner,
David Yager-Elorriaga, Kevin Yates,
Mykonos: A pulsed power driver for science and innovation,
High Energy Density Physics,
Volume 53,
2024,
101144,
ISSN 1574-1818,
[Link]
([Link]
Abstract: Sandia National Laboratories has been operating the Mykonos linear transformer
driver (LTD) in a five-cavity configuration since 2014. The machine operates at 1MA output
current, 500kV output voltage, with a 10–90% current rise time of 85ns, which enables small
scale physics and engineering pulsed power experiments. Mykonos provides hands-on
pulsed power experimental training for students and staff alongside senior Sandia scientists
in an environment that is more accessible than the Z Facility. Over the years, we have fielded
and accumulated a wide variety of optical, x-ray and electrical diagnostics and we are
preparing to open this facility to outside users. Here, we are presenting the pulsed power
and diagnostic capability of Mykonos as well as some recent experiments that have been
performed on the facility. The goal of this publication is to attract researchers across the
pulsed power and high energy density (HED) community to collaborate with Sandia on
exciting, innovative science and to train the next generation of researchers for the National
Nuclear Security Agency (NNSA) and the nation. As such, we have established a Mykonos
Academic Access Program (MAAP) as part of ZNetUS to enable academic utilization of the
Mykonos Pulsed Power Facility.
Keywords: Pulsed power; ZNetUS; Linear transformer driver
Derek Ka-Hei Lai, Li-Wen Zha, Tommy Yau-Nam Leung, Andy Yiu-Chau Tam, Bryan Pak-Hei So,
Hyo-Jung Lim, Daphne Sze Ki Cheung, Duo Wai-Chi Wong, James Chung-Wai Cheung,
Dual ultra-wideband (UWB) radar-based sleep posture recognition system: Towards
ubiquitous sleep monitoring,
Engineered Regeneration,
Volume 4, Issue 1,
2023,
Pages 36-43,
ISSN 2666-1381,
[Link]
([Link]
Abstract: Sleep posture monitoring is an essential assessment for obstructive sleep apnea
(OSA) patients. The objective of this study is to develop a machine learning-based sleep
posture recognition system using a dual ultra-wideband radar system. We collected
radiofrequency data from two radars positioned over and at the side of the bed for 16
patients performing four sleep postures (supine, left and right lateral, and prone). We
proposed and evaluated deep learning approaches that streamlined feature extraction and
classification, and the traditional machine learning approaches that involved different
combinations of feature extractors and classifiers. Our results showed that the dual radar
system performed better than either single radar. Predetermined statistical features with
random forest classifier yielded the best accuracy (0.887), which could be further improved
via an ablation study (0.938). Deep learning approach using transformer yielded accuracy of
0.713.
Keywords: Obstructive sleep apnea; Deep learning; Sleep monitoring; Feature extraction;
Ablation study
Zhongrui Bai, Fanglin Geng, Hao Zhang, Xianxiang Chen, Lidong Du, Peng Wang, Pang Wu,
Gang Cheng, Zhen Fang, Yirong Wu,
Non-contact blood pressure estimation using FMCW radar: A two-stream approach focused
on central arterial activity,
Biomedical Signal Processing and Control,
Volume 106,
2025,
107718,
ISSN 1746-8094,
[Link]
([Link]
Abstract: This paper proposes a radar-based two-stream blood pressure (BP) estimation
framework (R2S-BP), focusing on central arterial activity. It separately analyzes central-
arterial pulse transit time (caPTT) and pulse wave morphology using multi-location Doppler
Cardiogram (DCG) data from millimeter wave FMCW radar. Specifically, phase information at
harmonic heart rate frequencies is used to compute time delay arrays, representing caPTT-
related features. Additionally, k-Shape clustering is employed to select optimal DCGs from
the neck and chest regions that contain BP-related morphological features. These features
are processed through a two-stream neural network combining BiLSTM, ResNet, and multi-
head attention modules. Subject-independent 9-fold cross-validation results show that the
standard deviations of the errors for systolic and diastolic BP are 7.33 and 5.36 mmHg,
respectively. The intra-subject correlation coefficient for both systolic and diastolic BP
averages 0.82. Comparative and ablation studies demonstrate the superiority of the two-
stream approach and the critical importance of its components. This approach integrates
physiologically guided manual feature construction with a deep learning model, fully
leveraging the capabilities of FMCW radar data.
Keywords: Non-contact blood pressure estimation; Two-stream neural network; Doppler
Cardiogram; Central-artery pulse transit time
Marco Antonacci, Emanuele Riva, Attilio Frangi, Alberto Corigliano, Valentina Zega,
Planar GRIN lenses: Numerical modeling and experimental validation,
Journal of Sound and Vibration,
Volume 537,
2022,
117217,
ISSN 0022-460X,
[Link]
([Link]
Abstract: Phononic Crystal (PnC) Gradient Index (GRIN) lenses have been intensively
investigated in recent years for their promising applications in energy harvesting. Here we
propose and verify, both numerically and experimentally, three designs of PnC GRIN lenses
with amplification factors at the focal points of 4.33x, 7.35x and 7.49x. A design procedure
based on the combination of two mechanisms of refraction index gradient formation is
employed to boost the efficiency of the lens. Moreover, thanks to the planarity of the design
and to the single-phase constitutive material, the proposed lenses are fully compatible with
microfabrication processes and can be therefore employed as micro energy harvesters after
proper miniaturization.
Keywords: GRIN lenses; Phononic crystals; Energy harvesting; Numerical modeling;
Experiments