Conference
Conference
ABSTRACT 1 INTRODUCTION
Early and accurate detection of atrial fibrillation (AF), especially ECG is a non-invasive tool most commonly used for the detection of
paroxysmal AF (PAF), is critical for preventing stroke and cardio- cardiovascular diseases, based on a time series signal that provides
vascular complications. Traditional diagnostic tools often fail to information about morphology, amplitudes, heart rate, etc[15][11].
capture short-term, intermittent rhythms in limited ECG segments. One electrophysiologic disturbance that ECG signals observe within
In this work, we propose SimSANet, a deep learning model com- the atria is atrial fibrillation. Atrial fibrillation is the most frequent
bining Cross-Scale Attention (CSA) and SimAM, a parameter-free type of arrhythmia observed in clinical practice today[10][6]. Still,
attention mechanism, for robust multiclass classification of AF, usually the paroxysmal atrial fibrillation goes unnoticed, as in PAF,
non-AF (NAF), and PAF. the episodes are shorter and intermittent[7][25].
We preprocess 10-second lead-II ECG signals from the CPSC These episodes come and go irregularly and frequently, making
2021 Challenge using a second-order Butterworth bandpass filter it difficult to detect them through conventional methods[7]. Even
(0.5–45 Hz), convert them into 2D Constant-Q Transform (CQT) though atrial fibrillation in itself isn’t life-threatening, it signifi-
spectrograms, and use MobileNetV2 as a lightweight backbone cantly increases the risk of serious conditions like stroke and heart
for feature extraction. SimSANet enhances discriminative learning failure[38][6]. It’s early detection and proper treatment are crucial
through spatial and energy-based attention modules. To further to prevent it, as it serves as a precursor to prevent disease progres-
improve class separability, especially between NAF and PAF, we fur- sion through early AF surgery or drug intervention[23]. However,
ther added a hybrid loss function combining softmax cross-entropy AF detection poses a challenge because of its unpredictability and ir-
with triplet loss using hard mining. regularity. This is the reason why resting ECG is insufficient for the
Our method is evaluated on a balanced dataset of 96,000 CQT diagnosis of PAF[7][36]. To overcome this challenge, dynamic ECG
images (40k AF, 40k NAF, 16k PAF) and achieves a test accuracy is used, such as Holter monitors and wearable ECG devices[31].
of 98.48%, with a macro-averaged F1-score of 0.98. The t-SNE Before the AI era, atrial fibrillation diagnosis was done via con-
plots confirm the improved separation of PAF embeddings, and ventional methods using dynamic ECG recordings from Holter
the confusion matrix analysis reveals reduced misclassifications, monitors and wearable monitors, by analysing them manually, car-
particularly in clinically ambiguous cases. diologists inspected the recordings to identify the characteristics of
atrial fibrillation, like irregular R-R intervals, absent P waves, and
CCS CONCEPTS fibrillatory waves[6][7]. Persistent atrial fibrillation diagnosis was
• Computing methodologies → Neural networks; Supervised relatively easier as it could be observed even in a short-duration
learning by classification; Metric learning; • Applied computing ECG recording, since it could be seen throughout[7].
→ Health informatics; Life and medical sciences. But for the paroxysmal atrial fibrillation (PAF), an extended du-
ration ECG was required[36][7]. For PAF, to capture the short tran-
sient episodes, extended duration Holter monitors were used for
KEYWORDS
24-48 hours[25][31]. These inspections were time-consuming, repet-
Atrial Fibrillation, Deep Learning, Triplet loss itive, and also limited by low sensitivity for infrequent events[17].
ACM Reference Format: In recent years, we have seen advancements in machine learning
Urvi Patel and Soha Ghodeswar. 2024. A Deep Learning Pipeline for ECG- and deep learning for the detection of AF[16][26][3][1].
Based Classification of Atrial Fibrillation Using Time-Frequency Represen-
tations. In Proceedings of Indian Conference on Computer Vision Graphics 1.1 RELATED WORKS
and Image Processing (ICVGIP 2024). ACM, New York, NY, USA, 8 pages.
[Link] With the advent of deep learning and AI, multiple models have
been proposed for the detection of atrial fibrillation to automate
the diagnosis. The primary focus of earlier studies has been binary
Permission to make digital or hard copies of all or part of this work for personal or
classroom use is granted without fee provided that copies are not made or distributed classification of raw ECG signals to detect whether the subject has
for profit or commercial advantage and that copies bear this notice and the full citation atrial fibrillation or not.
on the first page. Copyrights for components of this work owned by others than the For instance, Ben-Moshe et al [2] developed and introduced
author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or
republish, to post on servers or to redistribute to lists, requires prior specific permission RawECGNet, a recurrent convolution-based deep learning model
and/or a fee. Request permissions from permissions@[Link]. that uses 30-second raw ECG segments to detect AF. This model
ICVGIP 2024, December 13–15, 2024, Bengaluru, India combines a 1D ResNet with BiGRU layers and generalizes across
© 2024 Copyright held by the owner/author(s). Publication rights licensed to ACM.
ACM ISBN 979-8-4007-1075-9/24/12 three datasets (UVAF, RBDB, and SHDB). Further, this model also
[Link] includes the estimation of AF burden. Although this model lacked
ICVGIP 2024, December 13–15, 2024, Bengaluru, India First Author et al.
interpretation using tools like Grad-CAM or attention visualization, STAGE Total Recordings AF PAF NON-AF
it was also limited to binary classification as AF and non-AF.
Stage-1 730 12 5 42
In a different approach, Li et al.[26] proposed a Self-Complementary
Stage-2 706 37 18 14
Convolutional Neural Network (SCCNN) based on a 2D Z-
Total 1436 47 23 53
shaped representation of ECG signals and attention mechanisms
to classify between AF and non-AF. Despite this, the model did Table 1: Summary of the Dataset of Dynamic ECG Recordings
not address the detection of PAF, and didn’t address the concerns
related to complex computations with respect to clinical use.
Similarly, Rahman et al.[33] introduced a dual-branch CNN This results in a total of 1,436 ECG segments (730 + 706) across
model that fuses raw ECG and discrete wavelet transformed (DWT) all classes. Both training sets were merged before preprocessing to
features to enhance temporal and frequency domain representa- create a more diverse and robust dataset for training. In this work,
tions. Although this model shows strong performance, it still does we are using only lead II ECG recordings from the dataset.
not address the detection of the subtypes of atrial fibrillation; more-
over, the interpretability wasn’t explored either. 2.1 Preprocessing and Image Generation
In yet another approach to address class imbalance and rare
When working with ECG, it is important to preprocess the sig-
event detection, Asadi et al.[1] introduced PxAF-Net, where they
nals in order to remove unwanted noise and to enhance the sig-
developed a framework that combines generative adversarial net-
nal for further improvement of the accuracy of cardiovascular
works (GANs) and neural architecture search (NAS) to improve
diagnosis[15][4].
detection of PAF. Despite the excellent performance, its reliance on
The ECG recordings in our dataset, recorded from 3-lead wear-
synthetic data and the model complexity raises concerns about the
ables as well as a 12-lead Holter system, consisted of much un-
real world deployment of the model.
necessary and unwanted noise and artifacts, which needed to be
Jayaraman et al.[22] proposed Q-Deep, which is more aligned
cleaned before conversion into an image to feed into our deep
with our approach and is based on a six-layer 2D CNN trained
learning model. This step was important, as a cleaned and prepro-
on RGB CQT spectrogram images, which were also derived from
cessed ECG would give us a higher accuracy for the detection of
10-second segments. This model performs multiclass classification
paroxysmal atrial fibrillation[4][9].
among AF, non-AF, and PAF but lacks attention-based interpretabil-
The ECG recordings were filtered using the bandpass filtering
ity as well as embedding strategies like triplet learning, which is
technique. For this purpose, we used a second-order Butterworth
important for feature refinement and class separability for condi-
filter in the range 0.5 to 45 Hz.
tions like AF and PAF that are morphologically similar.
The lower cutoff of the bandpass filter was kept at 0.5 Hz to re-
In this study, we propose an attention-based deep learning model,
move any baseline interference caused by respiration or movement
which is based on Soumyajit et al[13] SimSANet. We used this ar-
of the subject. The higher cutoff was kept at 45 Hz to remove any
chitecture to classify among three classes, AF, PAF, and NAF, using
possible powerline interference and other noise, such as muscle
short 10-second segments of ECG recordings from Lead II. Un-
artifacts[12][32]. A second-order Butterworth filter was chosen
like prior work, our method employs a MobileNetV2 backbone for
as it worked the best to preserve the morphology of ECG compo-
computational efficiency, enhanced by SimAM (Simple Attention
nents, as it has a smooth frequency response and minimal phase
Module) and Cross-Scale Attention (CSA) to focus on the most
distortion[14][24].
relevant spatial-frequency regions of the CQT spectrogram. Addi-
On performing the power spectral analysis, it was found that
tionally, we incorporate a hybrid loss function combining triplet loss
there was no noticeable interference at 50 or 100 Hz; hence, the use
with categorical cross-entropy, enabling the model to learn highly
of a notch filter was avoided. Smoothening was also avoided as the
discriminative embeddings for closely related rhythm classes, and
Savitzky-Golay filter distorted the frequency domain, which was
finally use softmax for the classification of AF, PAF, and non-AF.
noticed in the generated CQT images[29].
After filtering, the preprocessed signals were divided into 10-
2 DATASET DESCRIPTION second segments. Since ECG signals are 1D data, these segments
were then converted into CQT [Link] CQT input images
The dataset used to train our deep learning model for classifying
have been shown effective in detecting AF rhythm as demonstrated
atrial fibrillation (AF) versus normal rhythms is taken from the
in [1][22][26] and works well with CNN architectures. The time
Paroxysmal Atrial Fibrillation Events Detection from Dy-
domain ECG may not expose all patterns clearly. CQT does so by
namic ECG Recordings: The 4th China Physiological Signal
recognizing both temporal and spectral patterns, which helps the
Challenge 2021[28].
model to learn useful patterns, the morphology, and frequency
The ECG signals were acquired from either 12-lead Holter de-
variations[3].
vices or 3-lead wearable monitors. Challenge data include variable-
length ECG recordings extracted from leads I and II of long-term
dynamic ECGs, sampled at 200 Hz. To reduce ambiguity during 3 MODEL ARCHITECTURE
labeling, each AF episode contains a minimum of five consecutive In this work, we present a deep learning model for detecting parox-
heartbeats. We basically have two training sets, training set 1 and ysmal atrial fibrillation. Our main goal is to take 2D CQT im-
training set 2. We merged it together to have a greater number of ages derived from ECG signals and correctly identify whether
segments. they belong to AF, PAF, or a normal rhythm [1][26]. To do this,
A Deep Learning Pipeline for ECG-Based Classification of Atrial Fibrillation Using Time-Frequency Representations ICVGIP 2024, December 13–15, 2024, Bengaluru, India
SimAM
SimAM is a simple and lightweight parameter-free attention
module that is more effective than other complex and traditional
attention modules like SE or CBAM. It computes attention based on
local spatial variance, which allows the network to highlight more
informative features along with fewer computations and less pa-
rameter overhead. This helps us retain the lightweight and efficient
Figure 1: ECG signal before and after preprocessing. nature of MobileNetV2 without increasing the model complexity.
In contrast, popular modules like CBAM[40] or SENet[21] help
a model highlight the more discriminative features in the data, but
we break the model into three clear stages. First, we use a pre-
these modules have a larger number of learnable parameters, which
trained MobileNetV2 as the backbone to pull out low- and mid-
increases the complexity of the model.
level features[35]. Then, we boost these features with a channel-
Hence, SimAM provides an efficient way for feature enhance-
separated sequential attention (CSA) block and a SimAM attention
ment while keeping the architecture lightweight.
layer so that the model can focus on the most informative patterns at
SA Module
different scales[13][40]. Finally, during training, we add batch-hard
Our model is based upon the architecture proposed in The SA
mining to fine-tune the embedding space, helping the model sepa-
module is made of multiple SMDC blocks. The SMDC blocks are
rate classes more effectively and improve its final predictions[8][18].
the main building blocks of the CSA attention module[13]. In our
We’ll get more into details for the same in the upcoming sections.
model, there are 3 types of SMDC blocks applied: generic, row-wise,
In Fig. 1, you can see the overall architecture of the model.
and column-wise.
3.1 FEATURE EXTRACTION
Generic SMDC Block – The generic SMDC blocks apply depth-
In the first stage of our model, we aim to extract fundamental wise convolutions first vertically, then horizontally, one after the
features from the input CQT images. Each input is a 2D CQT spec- other, in a widespread manner. It does so with the help of sequen-
trogram image of size 224×224 pixels with 3 color channels (RGB). tially increasing kernel sizes from 1 x 1 to 5 x 5. This helps the
We are using MobileNetV2, pretrained on ImageNet, as the fea- model to capture spatial information at multiple scales, ranging
ture extractor in our first stage, as it is a lightweight convolutional from local to global.
architecture that relies on depthwise-separable convolutions and Row-wise and column-wise SMDC block
inverted residual blocks to capture both simple and more complex Further, to capture more directional features at different scales,
visual patterns without using excessive parameters[35][20]. we introduce specialized row and column SMDC blocks.
Next, an average pooling layer (AP) with a pool size of 2 x 2 is The row-wise SMDC applies depth-wise convolutions
applied, which enhances the feature extraction. Further, a Separable horizontally, with increasing kernel sizes – specifically 1 x 1, 2 x 1,
convolutional layer with 512 filters and a kernel size of 5x5 is added 3 x 1, 4 x 1, 5 x 1, scanning the image one row at a time.
to make feature maps more focused for attention modules, applied Similarly, the column-wise SMDC applies depth-wise convolu-
in the upcoming stage. tions across columns, i.e., vertically, using kernel sizes 1 x 1, 1 x 2, 1
x 3, 1 x 4, and 1 x 5, scanning the image one column at a time, which
XMobileNetV2 = MobileNetV2(XProcessedTensor ) (1) helps to gradually expand the receptive field column-wise. By com-
XAP = Average Pooling(XMobileNetV2 ) (2) bining these directional approaches, the SMDC structure extracts
features along both spatial axes, focusing on both column-wise and
XSepConv = SepConv(XAP ) (3)
row-wise variations, thereby expanding the overall receptive field.
Inside the SA module, the input goes through the first generic
3.2 CSA Module
SMDC, producing XGSMDC1i. Then, we apply GAP (global average
In the second stage, we use a channel-separated attention module pooling) and GMP (global max pooling). These pooled features are
that divides the features into different groups of channels, which concatenated and reshaped to form Xreshape.
further enhances the feature extraction process for more complex Next, this reshaped feature map goes into another generic SMDC
features. Each segment of the input tensor’s channels then focuses block, and its output is concatenated with the previous SMDC
on a different part of the information. Each segment is then pro- output. This result is passed through both the row-wise and column-
cessed through a sequential attention block. This block consists of wise SMDC blocks, giving us XRGSMDC1i and XCGSMDC1i. These
ICVGIP 2024, December 13–15, 2024, Bengaluru, India First Author et al.
outputs are also concatenated and then multiplied element-wise to AF detection[1][26]. We chose different pre-trained models like
form XMul1i, which is sent into the third generic SMDC. MobileNetV2, DenseNet, ResNet, and XceptionNet as a feature ex-
This pattern continues across all SMDC layers (a total of 5 blocks). tractor. The main aim is to find a lightweight model, robust, and
For each head in the CSA module, this entire process is repeated, balanced between both accuracy and computation cost, we fixed
and finally, the outputs from all heads are concatenated to get Ycsa. same hyperparameters across all the experiments, which are:
In parallel, SimAM is applied to the same input tensor to boost
spatial awareness even further. At the end, we combine the outputs Dropout: 0.3 Epochs: 50 Optimizer: Adam optimizer
from the CSA module and SimAM to form the final enhanced feature Learning rate: 1e-4 Batch size: 32 Rescale: 1./255
tensor, XFused, which is then used for classification. Scheduler:ReduceLROnPlateau (factor=0.5,patience=5,min𝑙 𝑟 =
1𝑒 − 6)
3.3 Classification and Batch-Hard Mining Stage
In the final stage of the model, we aim to improve the class predic- Pretrained Model Accuracy Precision F1-Score Recall
tions as well as the feature separability. First, the fused output is MobileNetV2 0.9871 0.97 0.97 0.97
passed through a global average pooling (GAP) layer to flatten the DenseNet121 0.9867 0.98 0.98 0.98
output. Subsequently, the output is passed into a fully connected ResNet50 0.9852 0.98 0.98 0.98
softmax classifier that directly predicts one of the three rhythm xceptionNet 0.9826 0.98 0.98 0.98
classes: AF, PAF, or Normal.
Table 2: Performance comparison of pretrained models
At the same time, we also apply batch-hard triplet mining[18]
to the embeddings. Triplet mining looks at the embeddings across
a batch of samples and constructs "hard" triplets — those where
a sample of one class is most similar to a different class or most
different from its class — then penalizes these cases. Specifically,
for each anchor embedding, hard positives and hard negatives are
identified[8].
We then encourage the model to decrease the distance between
the anchor and its hard positive, and increase the distance to the
hard negative by at least a fixed margin. Here, triplet loss helps to
refine features so that class separability is increased[19]. Figure 3: MobileNetV2
4 EXPERIMENTS AND RESULTS After analysing each model’s accuracy curve, loss curve, F1-
The first step towards training our deep learning model to classify score, precision, and accuracy, MobileNetV2 stands out as a balance
atrial fibrillation(AF), paroxysmal atrial fibrillation(PAF), and Non- between performance and efficiency. Even after achieving great
atrial fibrillation(NAF) is to choose a baseline model for feature accuracy with a pre-trained model, to strengthen the model architec-
extraction. Therefore, we selected pretrained model as they are ture, we will add attention blocks so that it gives consistent results
trained on large-scale image datasets and thus can extract features with real-world edge cases, like there might be class imbalance or
from the 2D CQT images more meaningfully,which is effective in minute morphological differences between PAF VS NAF. We first
Figure 4: DenseNet121
Figure 5: ResNet50
Figure 6: XceptionNet
combined MobileNetV2 with a channel attention block, aiming to Figure 9: t-SNE plot of MobileNetV2+channel Attention
enhance class-wise feature weighing through global pooling.
Figure 10: Confusion Matrix of MobileNetV2+CBAM atten- real-life edge cases, we added a hybrid loss combining categorical
tion cross-entropy with triplet loss using batch hard mining[18][8].
The categorical cross-entropy tries to minimize the prediction
error, but does not maximize The interclass distance can still lead
to class overlapping in boundary regions. So with triplet loss, it will
bring the same class embedding closer and push the different class
embedding further, which will help to improve generalization on
noisy or unseen data. This is the confusion matrix of hard mining
hus, the proposed SimSANet with hard-mining triplet loss proves to [5] Jason Brownlee. 2018. How to randomly split data into train and test sets in
be our most robust and reliable model for real-world deployment. Python. [Link]
machine-learning-algorithms/.
[6] A. John Camm, Gregory Y. H. Lip, Raffaele De Caterina, Irene Savelieva, Dan Atar,
Stefan H. Hohnloser, Gerhard Hindricks, and Paulus Kirchhof. 2012. 2012 focused
5 CONCLUSION AND FUTURE SCOPE update of the ESC Guidelines for the management of atrial fibrillation. European
Heart Journal 33, 21 (2012), 2719–2747. [Link]
In this study, we aimed to develop a deep learning model for classi- [7] Efstratios I. Charitos, Paul D. Ziegler, Ulrich Stierle, David R. Robinson, Hans-
fying atrial fibrillation, paroxysmal atrial fibrillation, and non-atrial Hinrich Sievers, and Thomas Hanke. 2014. A comprehensive review of the
fibrillation. To train our model, we derived a dataset from the CPSC mechanisms and detection of paroxysmal atrial fibrillation. Journal of Electrocar-
diology 47, 6 (2014), 771–778. [Link]
2021 challenge[28], which has lead 1 and lead 2 ECG recordings. [8] Davide Chicco. 2021. Siamese neural networks: An overview. Artificial Intelligence
We carefully preprocessed each ECG signal of lead-2 and divided it Review 55 (2021), 2009–2031. [Link]
[9] Vít Chudáček, Josef Spilka, George Georgoulas, and Lenka Lhotská. 2014. Pre-
into 10-second RGB CQT as CQT has strong performance in detec- processing methods for time series data: ECG signal classification. Biomedical
tion of AF[22][1][26]. images. Then we got around 96k images of Signal Processing and Control 10 (2014), 39–48.
NAF, 40k images of AF, and 16k images of PAF. To remove bias, we [10] Sumeet S Chugh, Rasmus Havmoeller, Kumar Narayanan, David Singh, Michiel
Rienstra, Emelia J Benjamin, Richard F Gillum, Young-Hoon Kim, John H McAn-
performed downsampling and used 40k RGB CQT images of the AF ulty, Zhi-Jie Zheng, et al. 2014. Worldwide epidemiology of atrial fibrillation:
and NAF classes and 16k images of the PAF class. We also resized a Global Burden of Disease 2010 Study. Circulation 129, 8 (2014), 837–847.
each image into 224*224, which aligns with image-based CNN ar- [Link]
[11] Gari D. Clifford, Francisco Azuaje, and Patrick E. McSharry. 2006. ECG statistics,
chitectures like MobileNet [35] and to train the model, we made a noise, artifacts, and missing data. Artech House, 55–99.
split of 70:25:5 of train:validation:test [5]. To train our model, we [12] Ivan K Daskalov, Ivan A Dotsinsky, and Ivaylo I Christov. 2000. Effective ECG
signal preprocessing and QRS detection algorithms in real time. Physiological
used SimSANet, which uses CSA and SimAM with MobileNetV2 as Measurement 21, 3 (2000), 287. [Link]
a backbone, which outperformed CBAM[39] and channel attention [13] Sayantan Gayen, Souvik Maity, P.K. Singh, R. Sarkar, and Bijaya Ketan Sarkar.
models [21]in terms of precision, accuracy, and F1-score, and the 2025. SimSANet: a simple sequential attention-aided deep neural network for
vehicle make and model recognition. Neural Computing and Applications 37
t-SNE plot showed strong class separation,then we further added (2025), 319–339. [Link]
batch hard mining With triplet loss to enhance PAF boundary clas- [14] Sobia Gilani and Zahid Malik. 2020. Denoising ECG signals using adaptive
sification and removed misclassification at edge cases and achieved filters and Butterworth filters. Biomedical Signal Processing and Control 59 (2020),
101917. [Link]
a low test loss of 0.0374 with 98.48% accuracy. [18][8]. [15] Ary L Goldberger, Luis AN Amaral, Leon Glass, Jeffrey M Hausdorff, Plamen Ch
In future work, We aim to check its feasibility in more real-life edge Ivanov, Roger G Mark, Joseph E Mietus, George B Moody, Chung-Kang Peng, and
H Eugene Stanley. 2000. PhysioBank, PhysioToolkit, and PhysioNet: components
cases by testing this across different datasets and diverse patient de- of a new research resource for complex physiologic signals. Circulation 101, 23
mographics (e.g., MIT-BIH Arrhythmia Database [30], PTB-XL[37], (2000), e215–e220. [Link]
Chapman University ECG dataset[27]),this cross-dataset evaluation [16] Awni Y. Hannun, Pranav Rajpurkar, Mohammad Haghpanahi, Geoffrey H. Tison,
Christopher Bourn, Mintu P. Turakhia, and Andrew Y. Ng. 2019. Cardiologist-
will help assess the model’s robustness across variations in patient level arrhythmia detection and classification in ambulatory electrocardiograms
populations, recording settings, and sampling frequencies. using a deep neural network. Nature Medicine 25, 1 (2019), 65–69. [Link]
We will also implement AF burden estimation that measures time org/10.1038/s41591-018-0268-3
[17] Junichiro Hayano, Akira Yamada, Yutaka Sakakibara, and Takanari Fujinami.
spent in AF per patient, which will be beneficial to estimate the 2005. Assessment of detection sensitivity of atrial fibrillation in 24-h Holter ECG
intensity of the risk, guide early intervention, and also can provide recordings using a new algorithm. Journal of Electrocardiology 38, 1 (2005), 10–14.
[Link]
personal treatment recommendation for high risk patients. We will [18] Alexander Hermans, Lucas Beyer, and Bastian Leibe. 2017. In Defense of the
extend our model to utilize multi-lead ECG recordings [34], While Triplet Loss for Person Re-Identification. In Proceedings of the IEEE Conference on
this study focused on lead II due to its prominence in rhythm anal- Computer Vision and Pattern Recognition (CVPR) Workshops.
[19] Elad Hoffer and Nir Ailon. 2015. Deep metric learning using triplet network. In
ysis, integrating signals from multiple leads (e.g., V1, V5, aVR) may International Workshop on Similarity-Based Pattern Recognition. Springer, 84–92.
improve the detection of complex arrhythmias, particularly in noisy [20] Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun
or low-amplitude segments. Exploring lead-specific attention or Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. 2017. MobileNets:
Efficient convolutional neural networks for mobile vision applications. In arXiv
lead fusion strategies could further enhance performance in mixed preprint arXiv:1704.04861.
clinical data. [21] Jie Hu, Li Shen, and Gang Sun. 2018. Squeeze-and-Excitation Networks. In
Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
(CVPR). 7132–7141.
[22] Shanthini Jayaraman, Hoda Asadi, Somayeh Goodarzy, Derek Abbott, and Math-
ACKNOWLEDGMENTS ias Baumert. 2024. Deep learning of short single-lead ECG segments for persistent
atrial fibrillation detection using time-frequency representations. Artificial Intel-
REFERENCES ligence in Medicine (2024). In press.
[1] Hoda Asadi, Shanthini Jayaraman, Somayeh Goodarzy, Derek Abbott, and Math- [23] Paulus Kirchhof, Stefano Benussi, Dipak Kotecha, Anders Ahlsson, Dan Atar,
ias Baumert. 2023. Q-Deep: A deep learning framework for the detection of persis- Barbara Casadei, Manuel Castella, Hans-Christoph Diener, Hein Heidbuchel,
tent atrial fibrillation from short single-lead ECG segments. Artificial Intelligence Jeroen Hendriks, et al. 2016. Early rhythm-control therapy in patients with atrial
in Medicine 141 (2023), 102466. [Link] fibrillation. New England Journal of Medicine 375, 6 (2016), 581–592. https:
[2] Doron Ben-Moshe, Ran Gelbart, Saul Greenstein, Eyal Sela, and Shai Shoham. //[Link]/10.1056/NEJMoa1602007
2023. RawECGNet: Deep Learning Generalization for Atrial Fibrillation Detection [24] R Kohli and A Arora. 2021. Baseline wander removal in ECG signals using
from the Raw ECG. Frontiers in Physiology 14 (2023), 1160309. [Link] Butterworth IIR filter. Journal of Medical Engineering & Technology 45, 5 (2021),
3389/fphys.2023.1160309 356–362. [Link]
[3] Soumyajit Bhattacharya, Arpan Choudhury, Aradhya Shukla, Rajan Saini, Ab- [25] David E. Krummen, Jason D. Bayer, Jennifer Ho, Jerry Hoang, Andreas Schricker,
hishek Dutta, Samarjit Bhattacharya, and Kaushik Roy. 2024. Cardiac rhythm clas- Shih-Ann Lin, and Sanjiv R. Narayan. 2012. Mechanisms of paroxysmal atrial
sification from imbalanced ECG signals using deep learning-based ensemble mod- fibrillation and implications for catheter ablation. Journal of Cardiovascular
els with transfer learning. Soft Computing (2024). [Link] Electrophysiology 23, 6 (2012), 625–631. [Link]
024-10480-z 02396.x
[4] R Bousseljot, D Kreiseler, and A Schnabel. 2009. ECG signals from the PTB
database. Physikalisch-Technische Bundesanstalt (PTB), Germany (2009).
ICVGIP 2024, December 13–15, 2024, Bengaluru, India First Author et al.
[26] Yang Li, Yuan Liu, Chao Tian, Xiang Liu, Chengyu Liu, Yujian Zhang, Mingkai Transformed ECG Features. In Proceedings of the 44th Annual International Con-
Jiang, and Mingkai Xu. 2023. Diagnosis of atrial fibrillation using self- ference of the IEEE Engineering in Medicine and Biology Society (EMBC). IEEE,
complementary attentional convolutional neural network. Scientific Reports 1085–1088. [Link]
13, 1 (2023), 9172. [Link] [34] Pranav Rajpurkar, Awni Y Hannun, Mohammad Haghpanahi, Christopher Bourn,
[27] Feifei Liu, Chengyu Liu, Mingkai Xu, Yujian Zhang, and Mingkai Jiang. 2021. An and Andrew Y Ng. 2019. Deep learning for ECG interpretation: comparison with
open-access database for evaluating the algorithm performance of ECG rhythm cardiologists. Nature Medicine 25, 6 (2019), 856–860. [Link]
and morphology abnormality detection. Physiological Measurement 42, 10 (2021), s41591-019-0350-9
105003. [Link] [35] Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-
[28] Feifei Liu, Xin Zhu, Chengyu Liu, Yujian Zhang, Mingkai Xu, Yuan Liu, Guox- Chieh Chen. 2018. MobileNetV2: Inverted Residuals and Linear Bottlenecks. In
ian Wu, et al. 2021. The China Physiological Signal Challenge 2021: Clas- Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
sification of 12-lead ECG. In PhysioNet/Computing in Cardiology Challenge. (CVPR). 4510–4520.
[Link] [36] Atul Verma, Allan C. Skanes, Eric N. Prystowsky, and Andrea Natale. 2014.
[29] S. Machhale, R. Suryavanshi, and M. Kokare. 2014. Comparison of Savitzky–Golay Approaches to the management of atrial fibrillation: update on rhythm control
and Butterworth filters for ECG signal preprocessing. In International Conference strategies. Canadian Journal of Cardiology 30, 7 (2014), S42–S50. [Link]
on Communication and Signal Processing (ICCSP). IEEE, 173–177. [Link] 10.1016/[Link].2014.04.012
10.1109/ICCSP.2014.6949861 [37] Philipp Wagner, Nils Strodthoff, Ralf-Dieter Bousseljot, David Kreiseler, Felix I.
[30] George B. Moody and Roger G. Mark. 2001. The impact of the MIT-BIH Arrhyth- Lunze, Wojciech Samek, and Tobias Schaeffter. 2020. PTB-XL, a large publicly
mia Database. IEEE Engineering in Medicine and Biology Magazine 20, 3 (2001), available electrocardiography dataset. Scientific Data 7, 1 (2020), 1–15. https:
45–50. [Link] //[Link]/10.1038/s41597-020-0495-6
[31] Hieu T. Nguyen, George F. Van Hare, Matthew Rudokas, Eric Bowman, and [38] Philip A. Wolf, Robert D. Abbott, and William B. Kannel. 1991. Atrial fibrillation
Phuoc V. Nguyen. 2021. Wearable devices for cardiac rhythm monitoring: ac- as an independent risk factor for stroke: the Framingham Study. Stroke 22, 8
curacy and usability. Expert Review of Medical Devices 18, 4 (2021), 295–305. (1991), 983–988. [Link]
[Link] [39] Sanghyun Woo, Jongchan Park, Joon-Young Lee, and In So Kweon. 2018. CBAM:
[32] Jiapu Pan and Willis J Tompkins. 1985. A real-time QRS detection algorithm. Convolutional Block Attention Module. In Proceedings of the European Conference
IEEE Transactions on Biomedical Engineering 3 (1985), 230–236. [Link] on Computer Vision (ECCV). 3–19.
10.1109/TBME.1985.325532 [40] Shen Yang, Qi Tan, Zhihao Zheng, Jiawei Xu, Deng-Ping Liu, and Xinchao Zhou.
[33] A. Rahman, K.M. Hasan, R. Amin, and M.A. Moni. 2022. A Deep Learning Scheme 2021. SimAM: A simple, parameter-free attention module for convolutional neural
for Detecting Atrial Fibrillation Based on Fusion of Raw and Discrete Wavelet networks. In Proceedings of the IEEE/CVF International Conference on Computer
Vision (ICCV). 11818–11827.