Multimodel Approach
Multimodel Approach
1. Introduction
Fingerprints, part of the dermatoglyphics field, are a complex and unique pattern of
curving line structures called friction ridges. The number, shape (loops, whorls, and arches),
and location of each ridge make every person unique, and they do not vary with growth or
age. The fingerprint image consists of dark lines called ridges and white lines called valleys.
In this work, we consider the International Fingerprint Liveness Detection Competition 2015
(LivDet 2015) dataset.
The third sub-block in the heartprint branch is 2D CNN architecture converts the STFT
input image of the size (26,37)26,37 to a (3,224,224)3,224,224 3D feature image.
The first fusion strategy is based on the concatenation of the extracted features of the
fingerprint image and the heartprint signal. The architecture of the two-branch neural network
is illustrated in Figure 2, which contains the fingerprint branch for the feature extraction and
the heartprint branch for height-relevant feature learning. The feature fusion module consists
of a sequence of deep learning layers. The first layer applies the concatenation of the feature
vector, arriving from the fingerprint feature extraction module and the feature vector received
from the heartprint feature extraction module. The combined feature passes through a fully
connected layer followed by a batch normalization layer, a Swish activation layer, a dropout
layer, and a second fully connected layer. Finally, a binary classifier decides in which
category, artifact or bona-fide, the elaborated fingerprint-heartprint feature belongs to.
Multimodal biometric systems seek to increase performance that may not be possible by
using a single biometric indicator by providing multiple shreds of evidence of the same
identity. An optimal fusion of multiple modalities is a fundamental request for the
development of a reliable solution. In attempting to improve the performance of the detection
system, the outputs of the fingerprint and heartprint branches are further processed using a
fusion module. This fusion module is performed by intercalating the heartprint SIFT image as
additional bands to the fingerprint image. We call this a channel-wise fusion approach and
illustrate it in Figure
A wide variety of deep learning strategies have been used to build biometric
identification systems. Usually, these methods depend on CNN to extract features from input
data. Inspired by the biological systems of humans, the attention mechanism has
revolutionized the natural language processing and computer vision systems [1,2]. The
attention mechanism has reasonably become one of the most fundamental concepts in the
deep learning field. The feature extraction module uses a state-of-the-art data-efficient ViT
variant (Deit) to extract discriminative features for image classification tasks. Deit has the
same architecture as ViT.
The input image to the ViT is split into 𝑁N patches of a fixed
size 𝐷D, 𝑋∈ℝ𝑁×𝐷X∈RN×D, the patches are flattened and fed to a linear projection to
create lower-dimensional linear embeddings, a positional embedding is added with the class
of the embedded image, and the sequence is fed to the transformer encoder. The transformer
encoder uses a Multi-Head Self Attention layer (MSA) as an attention mechanism in between
all the input vectors. The input to the attention block has three linear input layers (receive the
queries 𝑄=𝑋𝑊𝑄Q=XWQ, keys 𝐾=𝑋𝑊𝐾K=XWK, and
values 𝑉=𝑋𝑊𝑉V=XWV with 𝑊𝑄WQ,𝐾WK,𝑊𝑉∈ℝ𝐷×𝑑WV∈RD×d are the
parameters of the linear transformations), followed by a scaled dot-product attention function
to give the output matrix:
𝐴𝑡𝑡𝑒𝑛𝑡𝑖𝑜(𝑄,𝐾,𝑉)=𝑆𝑜𝑓𝑡𝑚𝑎𝑥(𝑄𝐾𝑇/𝑑−−√)𝑉AttentionQ,K,V=SoftmaxQKT/dV
(2)
where the term 𝑑−−√d provides proper normalization. The attention function is
repeated ℎh times to produce a multi-head self-attention ( ℎh heads), a concatenation
operation joins the ℎh outputs of the different heads, and a final MLP head performs the
classification task.
Introduced by Facebook AI, the structure of the Deit model, built based on the ViT
model [34], showed enhancement over previous ViT models. ViT does not generalize well
when trained on a small amount of data and needs to be pre-trained with a huge amount
(hundreds of millions) of images. Deit architecture is proposed using a ViT architecture with
a teacher-student strategy and a distillation token. The distillation token, which allows the
deep model to learn from the teacher’s output, interacts with the class token and patch tokens
through the self-attention layers to provide the hard label predicted by the teacher.
3. Results
To evaluate the proposed method, we first built our own multimodal dataset. Then, we
split the dataset into training and testing sets, by splitting based on the subjects.
The International Fingerprint Liveness Detection Competition 2015 (LivDet 2015) and
the real heartprint dataset, called Heartprint2022, are used to evaluate and validate the
performance of the proposed deep learning architecture. The LivDet 2015 dataset is provided
by Orrù et al. and can be downloaded from [6], whereas the Heartprint2022 dataset is a new
dataset collected in our lab, which is the Advanced Lab for Intelligent Systems Research
(ALISR), and can be downloaded from here [12].
The LivDet 2015 dataset has approximately 19,000 fingerprint images captured using
four different optical fingerprint sensors: GreenBit, Biometrika, Digital Persona, and
CrossMatch [6]. It aims to develop both software-based and hardware-based fingerprint
liveness detection methodologies [6]. LivDet 2015 contains a training set dataset and a testing
dataset [7]. Each set contains bona-fide (live) and artefact (fake) fingerprint images acquired
via different fingerprint scanners, as illustrated in Table 1. To mimic real scenarios, the
image capturing process includes normal mode, with dry and wet fingers, and with high and
low pressure.
Table 1. Device and image characteristics of the LivDet 2015 dataset.
1000 ×
Biometrika HiScan-PRO 1000 1000 1000 1500
1000
Digital
[Link].U 5160 252 × 324 1000 1000 1000 1500
Persona
L Scan
Crossmatch 640 × 480 1500 1500 1500 1448
Guardian
LivDet 2015 datasets contain spoof fingerprint images collected using artificial fingers.
Artificial fingers are fabricated using plasticine-like material to create a negative impression
or a mold of the real finger (cooperative method), the mold is then filled to produce the
artificial finger using gelatin, PlayDoh, or silicone. A latent fingerprint left on a surface is
another way to make artificial fingerprints (non-cooperative method); a transparency sheet is
obtained from a processed latent fingerprint and used to create the mold. Figure 4 shows
samples from the LivDet 2015 dataset.
Figure 4. Sample images of LivDet2015 dataset captured using CrossMatch, Digital Persona,
GreenBit, and Biometrica sensors. Live samples are in the green box and fake samples made
of different materials are in the red box.
LivDet 2015 dataset contains spoof images made using diverse materials, such as
Ecoflex, gelatin, latex, wood glue, liquid Ecoflex, and RTV (a two-component silicone
rubber), as shown in Table 2. The testing set includes some spoof images of materials which
were not included in the training set.
Table 2. Materials used for fabricating spoof images in the LivDet 2015 dataset. The
unknown materials that do not exist in the training part are in bold.
Green Bit
Biometrika Ecoflex, gelatin, latex, Ecoflex, gelatin, latex, wood glue, Liquid
Digital wood glue Ecoflex, RTV
Persona
The heartprint dataset is collected using the ReadMyHeart ECG device, by DailyCare
BioMedical [35]. The ReadMyHeart handheld ECG is simple to use without skin electrodes,
leads, wires, or conductive gels. The measurements are taken by placing the thumbs on the
conductive plates as shown in Figure 5. The heartprint needs only 30 s of measuring time
during which 15 s are digitalized and exported to the computer via a USB port. The heartbeat
activities of 164 persons are captured during two sessions to build an Heartprint2022 dataset
of 656 ECG records. During
Fingerprint
the preprocessing
Images
step, the Heartbeats authors used a
four-order Bona- band-pass
Artefact
Butterworth Fide filter with cut-
off frequencies # samples per of 0.25 and 40
Hz to remove 10 12 10 the different
subject
types of noise that can affect
Total number of
the heartprint, 700 840 700 such as the
samples
power-line interface,
baseline Fingerprint wanders, and
patient- Images electrode
motion Heartbeats artifacts.
Bona-
Artefact
Fide
# samples per
10 12 10
subject
Total number of
700 840 700
samples
In the first experiment, we trained different deep models on the multimodal dataset. The
average detection accuracy of the different models using a single modality (no fusion) and
multimodality (with fusion) biometric traits is reported in Table 4. In this experiment, we
achieved the fusion process between the contribution of the fingerprint and the heartprint
information part via concatenation at the feature level. We employed different pre-trained
models as the backbone networks to extract the feature of each modality. We trained the
different models on the fingerprint images and heartprint signals using the Adam optimizer
with a scalable learning rate, a batch size of 32, and 30 training epochs.
Table 4. Average accuracy of the proposed fusion by concatenation architecture.
Table 4. Average accuracy of the proposed fusion by concatenation architecture.
Average Accuracy %
Deit_tiny_patch16_224_f
Fingerprint 97.4 97.4
e
(No fusion)
Resnet18 98.0 98.0
Deit_tiny_patch16_224_f
95.0 98.8
e
Resnet18 98.3 99
We note from the reported results in Table 4 that the Resnet50 architecture outperforms
the other CNN models and achieves the highest accuracy of 98.7%. Resnet18 and
mobilenetv2_110d architectures perform with a high accuracy of 98.3%, which is not far off
the highest performance. Deit_tiny_patch16_224_fe and vit_tiny_patch16_224 architectures
achieve the lowest accuracies of 95% and 97.1%, respectively.
In the second experiment, we applied the fusion between the fingerprint and the
heartprint signal at the data level. We combined the two modalities by inserting the heartprint
image as a new channel in the fingerprint image. We trained different pre-trained models on
the combined fingerprint-heartprint to extract significant features from the combined images.
As shown in Table 4, combining fingerprints with heartprint data provides better
performance in detecting fake fingerprints than single modality biometric traits. The different
deep models, namely, Deit_tiny, Resnet18, Resnet18d, Resnet50, MobileNetv2_100, and
MobileNetv2_110d, vit_tiny_patch16_224, perform with accuracies of 98.8%, 99%, 97.6%,
98.3%, 98.3%, 98.6%, and 98.1%, respectively. We can clearly see that the multimodal
biometric system outperforms a biometric system with a single biometric indicator (i.e.,
without using a fusion). This performance is achieved thanks to converting the heartprint to a
2D image and the benefit of the 2D convolution’s power in the deep learning models. The
Resnet50 deep model achieves the highest performance (an accuracy of 99%), surpassing the
other models.
[Link] Analysis of the Number of Training Subjects
Generally, it is common knowledge that a small training dataset produces weak
approximation [37]. To assess the impact of the training set size on the system performance,
we trained and evaluated the different models on a dataset with different sizes (between 20%
and 80%) and reported the achieved performance of each model in Table 5.
Table 5. Average accuracy in terms of percentage of subjects used in the training set.
Channel-wise concatenation approach is used.
Table 5. Average accuracy in terms of percentage of subjects used in the training set.
Channel-wise concatenation approach is used.
The reported results in Table 5 reveal that increasing the size of the training dataset in
the learning process improves the classification performance during the testing phase. We can
observe this behavior from the models’ performances. Deit-tiny performs well when trained
on a dataset with a size more than 50%. With a small size of training samples (20%), the
Resnet18 and Resnet50 models achieve good accuracies and maintain their performances for
all the training sample sizes. As shown in Table 5, Resnet18 reaches an accuracy of 99.3%
and outperforms all the other deep learning models when trained on 80% of the dataset. It is
well known that training a deep learning model involves large amounts of labeled training
samples. Training a model with insufficient amounts of labeled data degrades the testing
accuracy. Despite training with small amounts of training samples, the deep models perform
well with good accuracies (95.05%, 96.5%, 97%, and 90.5%) for Deit_tiny, Resnet18,
Resnet50, and Mobilenetv2_100, respectively, when trained using only 20% of the dataset.
[Link] of the heartprint feature
During this experiment, we assess the effect of the number of heartbeats used in the
STFT heartprint image on the model performance accuracy. We repeated the experiment with
STFT heartprint images built using a different number of heartbeats (heartprint of length
ranged between 5 heartbeats and 20 heartbeats), the model performance is reported in Table
6.
Table 6. Average accuracy of the proposed fusion via concatenation architecture with respect
to the number of heartbeats.
Table 6. Average accuracy of the proposed fusion via concatenation architecture with respect
to the number of heartbeats.
5 98.70
7 98.05
10 96.83
13 97.40
15 98.05
18 98.38
20 96.75
Table 6 shows that increasing the number of heartbeats in the STFT transformation of
the heartprint to an image does not significantly affect the accuracy. The highest accuracy
(98.7%) is obtained when adopting five heartbeats of the heartprint signal to construct the
STFT image.
4. Conclusions
References
1. Oloyede, M.O.; Hancke, G.P. Unimodal and Multimodal Biometric Sensing Systems: A
Review. IEEE Access 2016, 4, 7532–7555. [Google Scholar] [CrossRef]
2. Mordini, E.; Tzovaras, D. (Eds.) Second Generation Biometrics: The Ethical, Legal and
Social Context; The International Library of Ethics, Law and Technology; Springer:
Dordrecht, The Netherlands, 2012; Volume 11, ISBN 978-94-007-3891-1. [Google Scholar]
3. González-Soler, L.J.; Gomez-Barrero, M.; Chang, L.; Suárez, A.P.; Busch, C. On the Impact
of Different Fabrication Materials on Fingerprint Presentation Attack Detection. In
Proceedings of the 2019 International Conference on Biometrics (ICB), Crete, Greece, 4–7
June 2019. [Google Scholar]
4. ISO/IEC 30107-1:2016; Information Technology—Biometric Presentation Attack Detection
—Part 1: Framework. ISO: Geneva, Switzerland, 2016.
5. Chugh, T.; Jain, A.K. Fingerprint Spoof Generalization. arXiv 2019, arXiv:1912.02710.
[Google Scholar]
6. Orrù, G.; Casula, R.; Tuveri, P.; Bazzoni, C.; Dessalvi, G.; Micheletto, M.; Ghiani, L.;
Marcialis, G.L. LivDet in Action-Fingerprint Liveness Detection Competition 2019. In
Proceedings of the 2019 International Conference on Biometrics (ICB), Crete, Greece, 4–7
June 2019. [Google Scholar]
7. Ghiani, L.; Yambay, D.A.; Mura, V.; Marcialis, G.L.; Roli, F.; Schuckers, S.A. Review of the
Fingerprint Liveness Detection (LivDet) Competition Series: 2009 to 2015. Image Vis.
Comput. 2017, 58, 110–128. [Google Scholar] [CrossRef]
8. Husseis, A.; Liu-Jimenez, J.; Goicoechea-Telleria, I.; Sanchez-Reillo, R. A Survey in
Presentation Attack and Presentation Attack Detection. In Proceedings of the 2019
International Carnahan Conference on Security Technology (ICCST), Chennai, India, 1–3
October 2019; pp. 1–13. [Google Scholar] [CrossRef]
9. Micheletto, M.; Orrù, G.; Casula, R.; Yambay, D.; Marcialis, G.L.; Schuckers, S. Review of
the Fingerprint Liveness Detection (LivDet) Competition Series: From 2009 to 2021.
In Handbook of Biometric Anti-Spoofing: Presentation Attack Detection and Vulnerability
Assessment; Marcel, S., Fierrez, J., Evans, N., Eds.; Springer: Singapore, 2023; pp. 57–76.
[Google Scholar] [CrossRef]
10. Javier, G.; Fernando, A.F.; Julian, F.; Javier, O.G. A high performance fingerprint liveness
detection method based on quality related features. Future Gener. Comput. Syst. 2012, 28,
311–321. [Google Scholar] [CrossRef]
11. Coli, P.; Marcialis, G.L.; Roli, F. Vitality Detection from Fingerprint Images: A Critical
Survey. In Proceedings of the Advances in Biometrics; Lee, S.-W., Li, S.Z., Eds.; Springer:
Berlin/Heidelberg, Germany, 2007; pp. 722–731. [Google Scholar]
12. Islam, M.S.; Alhichri, H.; Bazi, Y.; Ammour, N.; Alajlan, N.; Jomaa, R.M. Heartprint: A
Dataset of Multisession ECG Signal with Long Interval Captured from Fingers for Biometric
Recognition. Data 2022, 7, 141. [Google Scholar] [CrossRef]
13. Odinaka, I.; Lai, P.-H.; Kaplan, A.D.; O’Sullivan, J.A.; Sirevaag, E.J.; Rohrbaugh, J.W. ECG
Biometric Recognition: A Comparative Analysis. IEEE Trans. Inf. Forensics Secur. 2012, 7,
1812–1824. [Google Scholar] [CrossRef]
14. Zhang, Q.; Zhou, D.; Zeng, X. HeartID: A Multiresolution Convolutional Neural Network for
ECG-Based Biometric Human Identification in Smart Health Applications. IEEE
Access 2017, 5, 11805–11816. [Google Scholar] [CrossRef]
15. Abo-Zahhad, M.; Ahmed, S.M.; Abbas, S.N. Biometric Authentication Based on PCG and
ECG Signals: Present Status and Future Directions. SIViP 2014, 8, 739–751. [Google
Scholar] [CrossRef]
16. Li, M.; Narayanan, S. Robust ECG Biometrics by Fusing Temporal and Cepstral Information.
In Proceedings of the 2010 20th International Conference on Pattern Recognition, Los
Alamitos, CA, USA, 23–26 August 2010; pp. 1326–1329. [Google Scholar]
17. Labati, R.D.; Sassi, R.; Scotti, F. ECG Biometric Recognition: Permanence Analysis of QRS
Signals for 24 h Continuous Authentication. In Proceedings of the 2013 IEEE International
Workshop on Information Forensics and Security (WIFS), Guangzhou, China, 18–21
November 2013; pp. 31–36. [Google Scholar]
18. Ribeiro Pinto, J.; Cardoso, J.S.; Lourenço, A. Evolution, Current Challenges, and Future
Possibilities in ECG Biometrics. IEEE Access 2018, 6, 34746–34776. [Google Scholar]
[CrossRef]
19. Raju, A.S.; Udayashankara, V. Biometric Person Authentication: A Review. In Proceedings
of the 2014 International Conference on Contemporary Computing and Informatics (IC3I),
Mysore, India, 27–29 November 2014; pp. 575–580. [Google Scholar]
20. NS, G.R.S.; Maheswari, N.; Samraj, A.; Vijayakumar, M.V. An Efficient Score Level
Multimodal Biometric System Using ECG and Fingerprint. J. Telecommun. Electron.
Comput. Eng. (JTEC) 2018, 10, 31–36. [Google Scholar]
21. Regouid, M.; Touahria, M.; Benouis, M.; Costen, N. Multimodal Biometric System for ECG,
Ear and Iris Recognition Based on Local Descriptors. Multimed. Tools Appl. 2019, 78,
22509–22535. [Google Scholar] [CrossRef]
22. El Rahman, S.A. Multimodal Biometric Systems Based on Different Fusion Levels of ECG
and Fingerprint Using Different Classifiers. Soft Comput. 2020, 24, 12599–12632. [Google
Scholar] [CrossRef]
23. Agrafioti, F.; Gao, J.; Hatzinakos, D.; Agrafioti, F.; Gao, J.; Hatzinakos, D. Heart
Biometrics: Theory, Methods and Applications; IntechOpen: London, UK, 2011; ISBN 978-
953-307-618-8. [Google Scholar]
24. Islam, M.S.; Alajlan, N. Biometric Template Extraction from a Heartbeat Signal Captured
from Fingers. Multimed. Tools Appl. 2017, 76, 12709–12733. [Google Scholar] [CrossRef]
25. Zhao, C.X.; Wysocki, T.; Agrafioti, F.; Hatzinakos, D. Securing Handheld Devices and
Fingerprint Readers with ECG Biometrics. In Proceedings of the 2012 IEEE Fifth
International Conference on Biometrics: Theory, Applications and Systems (BTAS),
Arlington, VA, USA, 23 September 2012; IEEE: New York, NY, USA, 2012; pp. 150–155.
[Google Scholar]
26. Alajlan, N.; Islam, M.S.; Ammour, N. Fusion of Fingerprint and Heartbeat Biometrics Using
Fuzzy Adaptive Genetic Algorithm. In Proceedings of the World Congress on Internet
Security (WorldCIS-2013), London, UK, 9 December 2013; IEEE: New York, NY, USA,
2013; pp. 76–81. [Google Scholar]
27. Hammad, M.; Wang, K. Parallel Score Fusion of ECG and Fingerprint for Human
Authentication Based on Convolution Neural Network. Comput. Secur. 2019, 81, 107–122.
[Google Scholar] [CrossRef]
28. Komeili, M.; Armanfard, N.; Hatzinakos, D. Liveness Detection and Automatic Template
Updating Using Fusion of ECG and Fingerprint. IEEE Trans. Inf. Forensics Secur. 2018, 13,
1810–1822. [Google Scholar] [CrossRef]
29. Jomaa, R.M.; Islam, M.S.; Mathkour, H. Enhancing the Information Content of Fingerprint
Biometrics with Heartbeat Signal. In Proceedings of the 2015 World Symposium on
Computer Networks and Information Security (WSCNIS), Hammamet, Tunisia, 19–21
September 2015; IEEE: New York, NY, USA, 2015; pp. 1–5. [Google Scholar]
30. Jomaa, R.M.; Islam, M.S.; Mathkour, H. Improved Sequential Fusion of Heart-Signal and
Fingerprint for Anti-Spoofing. In Proceedings of the 2018 IEEE 4th International Conference
on Identity, Security, and Behavior Analysis (ISBA), Singapore, 11–12 January 2018; IEEE:
New York, NY, USA, 2018; pp. 1–7. [Google Scholar]
31. Jomaa, R.M.; Islam, M.S.; Mathkour, H.; Al-Ahmadi, S. A Multilayer System to Boost the
Robustness of Fingerprint Authentication against Presentation Attacks by Fusion with Heart-
Signal. J. King Saud Univ. Comput. Inf. Sci. 2022, 34, 5132–5143. [Google Scholar]
[CrossRef]
32. Hammad, M.; Liu, Y.; Wang, K. Multimodal Biometric Authentication Systems Using
Convolution Neural Network Based on Different Level Fusion of ECG and
Fingerprint. IEEE Access 2019, 7, 26527–26542. [Google Scholar] [CrossRef]
33. Jomaa, R.M.; Mathkour, H.; Bazi, Y.; Islam, M.S. End-to-End Deep Learning Fusion of
Fingerprint and Electrocardiogram Signals for Presentation Attack
Detection. Sensors 2020, 20, 2085. [Google Scholar] [CrossRef] [PubMed]
34. Touvron, H.; Cord, M.; Douze, M.; Massa, F.; Sablayrolles, A.; Jegou, H. Training Data-
Efficient Image Transformers & Distillation through Attention. In Proceedings of the 38th
International Conference on Machine Learning, Virtual, 18 July 2021; Volume 139, pp.
10347–10357. [Google Scholar]
35. ReadMyHeart—Handheld ECG Recording Device (Id:976240) Product Details. Available
online: [Link]
--976239_976240.html (accessed on 2 May 2023).
36. Islam, M.S.; Alajlan, N. Augmented-Hilbert Transform for Detecting Peaks of a Finger-ECG
Signal. In Proceedings of the 2014 IEEE Conference on Biomedical Engineering and
Sciences (IECBES), Kuala Lumpur, Malaysia, 8 December 2014; pp. 864–867. [Google
Scholar]
37. Gütter, J.; Kruspe, A.; Zhu, X.X.; Niebling, J. Impact of Training Set Size on the Ability of
Deep Neural Networks to Deal with Omission Noise. Front. Remote Sens. 2022, 3, 932431.
Available
online: [Link] (accessed on 16
August 2023). [CrossRef]