0% found this document useful (0 votes)
254 views489 pages

Modern Trends in Expert Applications

The document outlines the proceedings of the MP-TEAS 2024 conference, focusing on modern practices and trends in expert applications and security. It highlights the rapid advancements in technology and the associated security challenges, featuring peer-reviewed research papers on topics such as artificial intelligence, cybersecurity, and IoT. The conference aimed to foster collaboration and knowledge sharing among researchers, industry professionals, and students globally.

Uploaded by

sbalakrishna30
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
254 views489 pages

Modern Trends in Expert Applications

The document outlines the proceedings of the MP-TEAS 2024 conference, focusing on modern practices and trends in expert applications and security. It highlights the rapid advancements in technology and the associated security challenges, featuring peer-reviewed research papers on topics such as artificial intelligence, cybersecurity, and IoT. The conference aimed to foster collaboration and knowledge sharing among researchers, industry professionals, and students globally.

Uploaded by

sbalakrishna30
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Lecture Notes in Networks and Systems 1378

Vijay Singh Rathore


Vladan Devedzic
Sriparna Saha
Nikhat Raza Khan Editors

Modern
Practices and
Trends in Expert
Applications and
Security
Proceedings of MP-TEAS 2024, Volume 1
Lecture Notes in Networks and Systems

Volume 1378

Series Editor
Janusz Kacprzyk , Systems Research Institute, Polish Academy of Sciences,
Warsaw, Poland

Advisory Editors
Fernando Gomide, Department of Computer Engineering and Automation—DCA,
School of Electrical and Computer Engineering—FEEC, University of Campinas—
UNICAMP, São Paulo, Brazil
Okyay Kaynak, Department of Electrical and Electronic Engineering,
Bogazici University, Istanbul, Türkiye
Derong Liu, Department of Electrical and Computer Engineering, University
of Illinois at Chicago, Chicago, USA
Institute of Automation, Chinese Academy of Sciences, Beijing, China
Witold Pedrycz, Department of Electrical and Computer Engineering, University of
Alberta, Alberta, Canada
Systems Research Institute, Polish Academy of Sciences, Warsaw, Poland
Marios M. Polycarpou, Department of Electrical and Computer Engineering,
KIOS Research Center for Intelligent Systems and Networks, University of Cyprus,
Nicosia, Cyprus
Imre J. Rudas, Óbuda University, Budapest, Hungary
Jun Wang, Department of Computer Science, City University of Hong Kong,
Kowloon, Hong Kong
The series “Lecture Notes in Networks and Systems” publishes the latest
developments in Networks and Systems—quickly, informally and with high quality.
Original research reported in proceedings and post-proceedings represents the core
of LNNS.
Volumes published in LNNS embrace all aspects and subfields of, as well as new
challenges in, Networks and Systems.
The series contains proceedings and edited volumes in systems and networks,
spanning the areas of Cyber-Physical Systems, Autonomous Systems, Sensor
Networks, Control Systems, Energy Systems, Automotive Systems, Biological
Systems, Vehicular Networking and Connected Vehicles, Aerospace Systems,
Automation, Manufacturing, Smart Grids, Nonlinear Systems, Power Systems,
Robotics, Social Systems, Economic Systems and other. Of particular value to both
the contributors and the readership are the short publication timeframe and
the world-wide distribution and exposure which enable both a wide and rapid
dissemination of research output.
The series covers the theory, applications, and perspectives on the state of the art
and future developments relevant to systems and networks, decision making, control,
complex processes and related areas, as embedded in the fields of interdisciplinary
and applied sciences, engineering, computer science, physics, economics, social, and
life sciences, as well as the paradigms and methodologies behind them.
Indexed by SCOPUS, EI Compendex, INSPEC, WTI Frankfurt eG, zbMATH,
SCImago.
All books published in the series are submitted for consideration in Web of Science.
For proposals from Asia please contact Aninda Bose ([Link]@[Link]).
Vijay Singh Rathore · Vladan Devedzic ·
Sriparna Saha · Nikhat Raza Khan
Editors

Modern Practices and Trends


in Expert Applications
and Security
Proceedings of MP-TEAS 2024, Volume 1
Editors
Vijay Singh Rathore Vladan Devedzic
Department of Computer Science Department of Software Engineering
and Engineering FON—School of Business Administration
IES University University of Belgrade
Bhopal, Madhya Pradesh, India Belgrade, Serbia

Sriparna Saha Nikhat Raza Khan


Department of Computer Science Department of Computer Science
and Engineering and Engineering
Indian Institute of Technology IES College of Technology
Patna, Bihar, India Bhopal, Madhya Pradesh, India

ISSN 2367-3370 ISSN 2367-3389 (electronic)


Lecture Notes in Networks and Systems
ISBN 978-981-96-5780-3 ISBN 978-981-96-5781-0 (eBook)
[Link]

© The Editor(s) (if applicable) and The Author(s), under exclusive license to Springer Nature
Singapore Pte Ltd. 2025

This work is subject to copyright. All rights are solely and exclusively licensed by the Publisher, whether
the whole or part of the material is concerned, specifically the rights of translation, reprinting, reuse
of illustrations, recitation, broadcasting, reproduction on microfilms or in any other physical way, and
transmission or information storage and retrieval, electronic adaptation, computer software, or by similar
or dissimilar methodology now known or hereafter developed.
The use of general descriptive names, registered names, trademarks, service marks, etc. in this publication
does not imply, even in the absence of a specific statement, that such names are exempt from the relevant
protective laws and regulations and therefore free for general use.
The publisher, the authors and the editors are safe to assume that the advice and information in this book
are believed to be true and accurate at the date of publication. Neither the publisher nor the authors or
the editors give a warranty, expressed or implied, with respect to the material contained herein or for any
errors or omissions that may have been made. The publisher remains neutral with regard to jurisdictional
claims in published maps and institutional affiliations.

This Springer imprint is published by the registered company Springer Nature Singapore Pte Ltd.
The registered company address is: 152 Beach Road, #21-01/04 Gateway East, Singapore 189721,
Singapore

If disposing of this product, please recycle the paper.


Preface

Welcome to MP-TEAS 2024, the International Conference on “Modern Practices and


Trends in Expert Applications and Security,” organized by IES University in Bhopal,
India. This three-day event, held from November 22 to 24, 2024, brought together
a diverse group of researchers, academicians, industry professionals, and students
from around the world to discuss the changing landscape of expert applications and
address the security challenges that come with technological advancement.
Technology is advancing faster than ever, changing the way we live, work, and
interact. These innovations, which range from artificial intelligence and the Internet
of Things to big data and cloud computing, are driving change across industries.
However, as these technologies advance, so do the risks associated with their misuse.
The goal of MP-TEAS 2024 is to provide a forum for exploring these issues,
discussing innovative solutions, and fostering collaborations that translate research
into practical outcomes.
The conference’s sessions covered a wide range of topics, including artificial intel-
ligence, cybersecurity, multimedia forensics, app development, and ethical consid-
erations in expert applications. We carefully planned the event to ensure that partici-
pants gained valuable insights, made meaningful connections, and found inspiration
to tackle today’s pressing issues.
This year’s response to MP-TEAS 2024 was truly impressive. The conference
received numerous submissions, with only the best, rigorously reviewed papers
chosen for presentation and publication. Springer Nature’s prestigious Lecture Notes
in Networks and Systems series publishes all accepted papers, ensuring global
visibility and impact.
MP-TEAS 2024 also provided a unique opportunity to bring together participants
from various countries, facilitating global knowledge sharing. Attendees came from
a variety of countries, including Serbia, the USA, Ireland, and India. The confer-
ence’s hybrid format ensured inclusivity and increased participation, reflecting our
commitment to hosting an accessible and impactful event.
We are deeply grateful to our organizing team, reviewers, technical program
committee, and sponsors for their tireless efforts to make this event a success.
Special thanks to Springer Nature for their collaboration and support in publishing

v
vi Preface

the conference proceedings, as well as to our keynote speakers, session chairs, and
all contributors who enriched this conference with their knowledge.
MP-TEAS 2024 was more than just a conference; it was a celebration of inno-
vation, collaboration, and the quest for knowledge. We hope that this event inspired
you, sparked new ideas, and laid the groundwork for meaningful research and part-
nerships. Thank you for joining us on this journey; we hope to see you at future
MP-TEAS events.

Bhopal, India Vijay Singh Rathore


Belgrade, Serbia Vladan Devedzic
Patna, India Sriparna Saha
Bhopal, India Nikhat Raza Khan
Acknowledgments

First and foremost, we express our heartfelt gratitude to the Almighty for providing
us with the strength and guidance we needed to successfully organize MP-TEAS
2024. This conference is the result of a collaborative effort, which would not have
been possible without the unwavering support of many individuals and organizations.
We are extremely grateful to IES University, Bhopal, for their unwavering
commitment to academic excellence and assistance in hosting MP-TEAS 2024.
Special thanks to the leadership of IES University for providing the necessary
infrastructure, resources, and encouragement to make this event a reality.
Our heartfelt thanks go to all of our collaborators, including the international and
national advisory committees, who contributed their knowledge and vision to this
conference. Their valuable contributions helped shape MP-TEAS 2024 into a glob-
ally significant platform. We also want to thank Springer Nature for their collab-
oration and for publishing our conference proceedings in their prestigious Lecture
Notes in Networks and Systems series, which ensures that the research presented here
reaches a global audience.
We are deeply grateful to our keynote speakers and session chairs, whose knowl-
edge and insights have greatly enhanced this conference. A special acknowledgment
to:
• Prof. V. K. Tripathi, NITTTR, Bhopal.
• Prof. R. S. Salaria, Guru Nanak University, Hyderabad.
• Prof. R. A. Patil, COEP Technological University.
• Dr. Chandrahauns Chavan, Jamnalal Bajaj Institute of Management Studies.
• Dr. Catarina Moreira, University of Technology, Sydney.
• Dr. Sridaran Rajagopal, Marwadi University, Rajkot.
• Dr. Sonali Vyas, UPES, Dehradun.
• Dr. Nguyen Ha Huy Cuong, University of Danang, Vietnam.
• Dr. Tanupriya Choudhry, UPES, Dehradun.
• Dr. Manju Vyas, for her enriching address.
Your contributions deepened and enriched the discussions, inspiring participants
to think beyond conventional boundaries.

vii
viii Acknowledgments

A special thanks to our Technical Program Committee and reviewers for their
tireless efforts to ensure the highest quality and integrity during the peer-review
process. Your dedication has helped MP-TEAS 2024 maintain its academic rigor.
We are grateful to the authors who chose MP-TEAS 2024 to present their inno-
vative research. Your trust in us and hard work have been at the heart of this occa-
sion. Thank you to everyone who attended the conference, both in person and virtu-
ally, for your enthusiasm, engagement, and contributions, which helped to make the
discussions dynamic and meaningful.
A huge thank you to our organizing team for their flawless coordination, meticu-
lous planning, and unwavering dedication. From logistics to promotion, your efforts
behind the scenes have been the foundation of this conference. Thank you to the
media and promotion teams for increasing the visibility of MP-TEAS 2024 and
ensuring it reached a large audience.
Finally, we extend our heartfelt gratitude to everyone who helped make MP-
TEAS 2024 a success, whether directly or indirectly. Your combined efforts made
this conference a vibrant, enriching, and memorable event. We look forward to future
support and collaboration.
Thank you for making MP-TEAS 2024 a tremendous success!

Prof. Vijay Singh Rathore


PC Chair and Convener, MP-TEAS
2024, Professor CSE and
Director—Research and International,
IES College of Technology, IES
University, Bhopal, Madhya Pradesh,
India
Prof. (Dr.) Nikhat Raza Khan
PC Chair and Convener, MP-TEAS
2024, Professor and Head, CSE, IES
College of Technology, IES University,
Bhopal, Madhya Pradesh, India
About This Book

This book is a compilation of the proceedings of the International Conference on


“Modern Practices and Trends in Expert Applications and Security” (MP-TEAS
2024), hosted by IES University in Bhopal. The event, organized in collaboration
with international and national partners, took place in hybrid mode from November
22 to 24, 2024. This volume contains high-quality, peer-reviewed research papers
that address cutting-edge developments in expert applications as well as the security
challenges they present.
Rapid technological advancements have transformed many fields, including artifi-
cial intelligence, big data, the Internet of Things, and cloud computing. These innova-
tions have created exciting opportunities in a variety of fields, including health care,
finance, manufacturing, cybersecurity, and others. However, these advancements
bring significant challenges, particularly in ensuring the security and ethical deploy-
ment of expert applications. This book is an invaluable resource for addressing these
challenges, as it provides in-depth insights, proposed solutions, and best practices.
The collection includes topics such as artificial intelligence and machine learning,
advanced web technologies, information and cybersecurity, IoT, big data, cloud
computing, multimedia forensics, app development, and the social and ethical impli-
cations of technology. Written for researchers, professionals, and students, this book
aims to offer a comprehensive reference for the latest advancements in expert appli-
cations and their secure implementation. It provides researchers with authenticated,
cutting-edge insights, while professionals can benefit from actionable guidelines for
a variety of engineering domains.

ix
Contents

Smartphone Sensor Dataset for Online Reading Analysis . . . . . . . . . . . . . . 1


Priyanka Bhatele and Mangesh Bedekar
Bridging the Gap: Integrating Generative AI into Engineering
Education 5.0 for Industry 5.0 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
Mitali Chugh, Shreya Panwar, Khushi Allawadi, and Sonali Vyas
Proposing an AI-Enabled Waste Segregation System for Domestic
Settings . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21
Devnarayan G. Rao, G. Sangeetha, J. Sandeep, C. S. Sreeja,
Agnes Nalini Vincent, and Teena Mary
Application of Convolutional Neural Networks’ Method in Early
Age Disease Detection of Paddy Crop . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 35
Asha Ambhaikar, Akanksha Mishra, Hussain Falih Mahdi,
Bhupesh Kumar Dewangan, Sanjana Dewangan,
and Tanupriya Choudhury
Information Encryption Technique Based on DNA Cryptography
and RSA Asymmetric Key Exchange Algorithm . . . . . . . . . . . . . . . . . . . . . . 43
Aman Khan, Kirti Nahak, Asha Ambhaikar, Hussain Falih Mahdi,
Bhupesh Kumar Dewangan, and Tanupriya Choudhury
Sentiment Analysis Using Bi-Directional LSTM for Depression
Detection and Suicide Prevention . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 55
Sunny Singh, J. Hridhya, Hussain Falih Mahdi,
Bhupesh Kumar Dewangan, and Tanupriya Choudhary
Disease Diagnosis from DNA Sequence Using GPU-Based
Aho–Corasick Algorithm . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 71
Bandi Krishna, Ramdas Vankdothu, Suresh Kumar Lokhande,
Hussain Falih Mahdi, Rakesh Nayak, Bhupesh Kumar Dewangan,
and Tanupriya Choudhary

xi
xii Contents

Advancing Dogri-English Translation Through Statistical Machine


Translation Technique . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 83
Vijay Singh Sen and Shubhnandan S. Jamwal
Future Price Prediction of IT Sector Companies Using Optimized
Deep Learning Approaches . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 97
Umar Bashir, Kuljeet Singh, Megha Raina, and Vibhakar Mansotra
Stacked Spatio-Temporal Fusion Network (SSTFN) Machine
Learning-Based Traffic Prediction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 113
Namrata Shrivastava and Jitendra Agrawal
Improving the Efficiency of EMG-Based Prosthetic Arm Using EEG . . . 125
H. B. Divyashree, Jaishree R. Devaru, Maitri Kulkarni,
and Muddireddy Devaghneswara Reddy
Analysis of Solar Panel Parameters for Extracting Maximum
Power using Artificial Neural Network-Based Training Algorithm . . . . . 135
Jagdish Chandola, Sumit Pundir, Abhishek Sharma, and Sakshi Pundir
A Review on Anomaly Detection Using Machine Learning
Techniques for Unmanned Aerial Vehicles (UAVs) . . . . . . . . . . . . . . . . . . . . 145
Dhruv Aggarwal, Sharon Christa, Sarishma Dangi,
and Aayush Shrivastava
Unlocking the Power of OpenCV: Innovative Development
of Gesture-Based Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 157
Dharmendra Sharma and Manmohan Singh
Predicting Cardiovascular Disease Risk Using Tree-Based
Gradient Boosting Machine Learning Techniques . . . . . . . . . . . . . . . . . . . . 169
Nikhat Raza Khan, Shekhar Verma, Hemant Kumar,
Mamata Mayee Panda, Abhishek Dwivedi, and Abhishek Kumar Mishra
A Comprehensive Framework for EEG-Based Emotion Detection:
From Exploration to Classification . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 179
Shaheen Ayyub, Aishwarya Vishwakarma, Vikas Sakalle,
Sugandh Singh, and Joanna Rosak-Szyrocka
Decentralized Blockchain Protocol . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 191
Akshay Jain
A Survey on IoT Protocols for Resource-Constrained Devices
in Handheld (IoT) Environment . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 207
Radhika Patel, Amit Nayak, and Romin Patel
Traffic Sign Recognition in Adverse Environments: A Survey
on Methods for Low Light, Fog, and Rain Conditions . . . . . . . . . . . . . . . . . 219
Kaushal Patel and Sheshang Degadwala
Contents xiii

An Analytical Review of Social Media-Based Sentiment Analysis


Techniques . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 231
Priyanshi Jain, Manas Vyas, Niharika Awasthi, and Raj Gaurav Mishra
The Future of Lung Disease Diagnosis: A Review of Emerging
Trends in Data-Driven Classification . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 245
Jigisha Mehta and Sheshang Degadwala
Leveraging Local Interpretable Model-Agnostic Explanations
(LIME) for Sentiment Analysis of News Articles . . . . . . . . . . . . . . . . . . . . . . 257
Pragya Tewari, Harshit Verma, Rishabh Jaiswal,
and Anurag Singh Baghel
Neural Network-Based Security System by PCA and Moth Flame
Optimized Features Learning Model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 275
Abhilasha Sinha, Rakesh Kumar Tiwari, and Onkar Nath Thakur
Performance Analysis of Solar Panel Power in Maximum Power
Point Tracking System Using Artificial Neural Network . . . . . . . . . . . . . . . 285
Jagdish Chandola, Sumit Pundir, Abhishek Sharma, Sakshi Pundir,
and Mohammed Azim Eirgash
Advanced Crop Prediction and Collaborative Agri-Investment
Platform . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 295
Punniyakotti Varadharajan Gopirajan, T. Pradeep, R. Charan,
and K. Suresh Kumar
Supervised Machine Learning Solution for Predict Heart Disease . . . . . . 311
Neha Para, Kavita Kushwah, Neha Jadon, and Aayush Shrivastava
Network and Security Solution Using Artificial Intelligence-Based
Graph Search Algorithm for Drone Transportation . . . . . . . . . . . . . . . . . . . 327
Pramod Kumar Patel, Amarjeet Ghosh, Aishwarya Mishra,
Rajesh Nema, Dilip Kumar Gandhi, Neeraj Agarwal,
and Ashish Raghuwanshi
Artificial Bee Colony Algorithm for Efficient Hyperparameter
Tuning in Alzheimer’s Disease Classification . . . . . . . . . . . . . . . . . . . . . . . . . 337
Raghubir Singh Salaria and Neeraj Mohan
Machine Learning Approaches for Anomaly Detection in IoT:
A Comprehensive Study . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 347
Rajesh Rajaan, Baldev Singh, and Nilam Choudhary
A Comprehensive, Multidisciplinary, Systematic, and Futuristic
Approach to Sustainable Food Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 365
M. Amin Mir, Duaa Hefni, Syed M. Hasnain, K. Andrews,
and Minakshi Memoria
xiv Contents

An Automated Extract, Transform, Load (ETL) Pipeline


to Facilitate Acquisition and Analysis of Stock Marker Data . . . . . . . . . . . 375
B. Uma Maheswari, K. S. Naveen Sakthivel, D. Kavitha, and R. Sujatha
An Intelligent Epidural Space Locator with a Medication Injector . . . . . 385
B. Vijayalakshmi, S. B. Mohan, M. Premkumar, M. Sasi Kumar,
and K. Suresh Kumar
Computational Investigations of Inorganic Perovskite Absorber
Material Using SCAP-1D Simulator . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 399
Syed M. Hasnain
Digital Twins and Industrial Internet of Things in Automobile
Industry: Applications and Future Trends . . . . . . . . . . . . . . . . . . . . . . . . . . . 411
R. Sujatha, B. Uma Maheswari, J. Abishek, and Viswanath Ananth
Energy-Efficient Routing Algorithm (EERA) Design
for Heterogeneous IoT Environment . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 425
K. Suresh Kumar, C. Rajendra Thilahar, and M. Vijay Anand
Enhancing the Performance of Smart E-Bike Using Intelligent
Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 439
C. Suresh, N. Sujitha, P. K. Mani, S. Sengottaian, and K. Suresh Kumar
Nanowires in Biomedical Analysis: A Strategic Review . . . . . . . . . . . . . . . . 453
Muhammad Azhar Ali Khan, M. Amin Mir, Syed M. Hasnain,
and Minakshi Memoria
Performance Assessment of Organometal Halide CH3 NH3 SnI3
in PSCs Using SCAPS-1D Simulator . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 463
Syed M. Hasnain, Abid Iqbal, Irfan Qasim, M. Amin Mir,
and Minakshi Memoria
Predicting Delays in IT Projects: A Machine Learning Approach . . . . . . 475
B. Uma Maheswari, D. Kavitha, R. Sujatha, and V. Santoshini

Author Index . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 489


About the Editors

Prof. (Dr.) Vijay Singh Rathore is presently working as Professor-CSE and Director
of Research and International Relations at IES University, Bhopal, and also running
his own establishment Shree KKarni Universe College, Jaipur (in partnership). He is
Membership Chair, ACM Jaipur Chapter, and Past Chairman, CSI Jaipur Chapter. His
core research areas include internet security, cloud computing, big data, and IoT. He
has organized and participated 25+ national and international conferences of repute.
He had been Member of Indian Higher Education Delegation and visited 20+ leading
universities at Canada (May 2019), the UK (July 2017), and the USA (August 2018)
supported by AICTE (led by then Chairman of AICTE Prof. Anil D. Sahashtrabudhe)
and GR foundation. He is Active Academician and always welcoming and forming
innovative ideas in the various dimensions of education and research for the better
development of the society.

Dr. Vladan Devedzic is Professor of Computer Science and Software Engineering


at the University of Belgrade, Faculty of Organizational Sciences, Belgrade, Serbia.
He also used to teach several computer science courses at the School of Elec-
trical Engineering, University of Belgrade, as well as at the Military Academy of
Serbia (ex Yugoslav Military Academy). Since 2021, he is Corresponding Member
of the Serbian Academy of Sciences and Arts (SASA). His long-term professional
objective is to bring close together ideas from the broad fields of artificial intel-
ligence/intelligent systems and software engineering. His current professional and
research interests include artificial intelligence, programming education, software
engineering, intelligent software systems, and technology-enhanced learning (TEL).
He is Founder and Chair of the GOOD OLD AI research network.

Dr. Sriparna Saha received [Link]. and Ph.D. degrees in computer science from
Indian Statistical Institute Kolkata, India, in the years 2005 and 2011, respectively.
She is currently serving as Associate Professor in the Department of Computer
Science and Engineering, Indian Institute of Technology Patna, India. She is author
of a book published by Springer-Verlag on the topic of unsupervised classification.
She has authored or co-authored more than 400 papers. Her current research interests

xv
xvi About the Editors

include multimodal information processing, natural language processing, machine


learning, deep learning, bioinformatics, multi objective optimization, and biomedical
information extraction. Her h-index is 40, and the total citation count of her papers is
8412 (according to Google Scholar). She is also Senior Member of IEEE and Fellow
of IETE. She won the best paper awards in CLINICAL-NLP workshop of COLING
2016, IEEE-INDICON 2015, International Conference on Advances in Computing,
Communications and Informatics (ICACCI 2012).

Dr. Nikhat Raza Khan has acquired doctor of philosophy in computer science
from Mewar University, in December 2018 with specialization on ad hoc network
and security and her master of technology in information security from Maulana
Azad National Institute of Technology, Bhopal, in May 2009, with the Title of her
thesis: “Edge Detection Technique using Fuzzy Logic.” She has more than a decade
of experience in education and research. She has authored more than 55 research
publications in national and international journals and conferences. Dr. Nikhat is
Active Member of professional body IEEE and ACM. She awarded “Best Women
Researcher Award” in Devi Ahilya Vishwavidyalaya held at Indore on February 27–
28, 2019. Under this she has reviewed various papers and Ph.D. thesis as External
Reviewer. She worked as Coordinator of various training programs, FDP, Teqip
workshops, seminars, conference, and expert lecture with highly skilled and leading
field expertise of national level.
Smartphone Sensor Dataset for Online
Reading Analysis

Priyanka Bhatele and Mangesh Bedekar

Abstract Popularity of smartphones also popularized, reading content using smart-


phones. Reading using smartphones quite differs from reading using desktop system.
Mouse and keyboard are the peripherals associated with the reading in desktop
systems. Study of the handling of such devices has led to provide implicit feedback
of the content read. Similar study in smartphones to get implicit feedback remains
to be a huge gap. Reading using smartphones involves screen gestures like pinch
to zoom, tap, scroll, orientation change, and screen capture. User reading behavior
and intent can be inferred by screen gestures. Smartphones have sensors like gyro-
scope and accelerometer that continuously capture data without permissions from
the users. To analyze screen gestures, a dataset has been created. Dataset is organized
into accelerometer and gyroscope folders. This dataset can act as an input to train,
test, and validate, machine learning models to analyze smartphone reading behavior
and user intent.

Keywords Smartphone sensor · Smartphone usage · Online reading · User


intent · Reading behavior

1 Introduction

With the proliferation of smartphones, reading digital content has become increas-
ingly popular on mobile devices. This mode of reading significantly differs from
traditional desktop systems, which primarily rely on peripheral devices such as the

P. Bhatele (B)
School of Computer Engineering and Technology, Dr. Vishwanath Karad MIT World Peace
University, Pune, Maharashtra, India
e-mail: priyanka.11212@[Link]
M. Bedekar
School of Computer Science and Engineering, Dr. Vishwanath Karad MIT World Peace
University, Pune, Maharashtra, India
e-mail: [Link]@[Link]

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 1
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
2 P. Bhatele and M. Bedekar

mouse and keyboard to interact with content. These peripherals have historically
provided valuable implicit feedback on user engagement and reading behavior. In
contrast, smartphones offer a unique set of interaction methods, including screen
gestures like pinch-to-zoom, tap, scroll, orientation changes, and screen captures.
These interactions offer a rich source of data that can be analyzed to infer user
behavior and intent. Although smartphones are widely used today, the exploration
of implicit feedback through screen gestures remains largely unexamined.
Modern smartphones are equipped with sensors such as gyroscopes and
accelerometers, which continuously capture data without requiring user permis-
sions. These sensors provide a continuous stream of data that can be leveraged to
gain insights into how users interact with content on their screens. By analyzing
this data, researchers can develop models to better understand reading behaviors
and infer user intent. To facilitate this research, we have created a comprehensive
dataset that captures accelerometer and gyroscope data during smartphone reading
sessions. This dataset is carefully structured to support train and validate models
focused on analyzing reading behavior on smartphones. The study addresses gaps in
the current research by offering a comprehensive dataset and showcasing its potential
for enhancing the understanding of user interaction with mobile digital content.

2 Literature Review

Smartphones have transformed digital content consumption, shifting engagement


from desktop systems, where activities like mouse clicks and keyboard input reveal
user behavior, to mobile platforms with touch-based interactions and sensors like
gyroscopes and accelerometers [1–3]. These sensors capture a wealth of data
through common gestures—pinch-to-zoom, tapping, scrolling, and screen orienta-
tion changes—yet their potential for analyzing reading behavior and engagement
remains largely untapped [4, 5]. Unlike desktops, which depend on explicit inputs,
smartphones continuously collect data without user intervention, allowing for a richer
understanding of real-time user interactions [6].
Recent machine learning advancements have made it possible to classify these
sensor-driven gestures to infer engagement and content preferences, with models
such as support vector machines (SVMs), random forests (RFC), and neural networks
showing promising results. However, challenges persist, including ethical concerns
surrounding data privacy and the variability in user behavior that complicates model
robustness. Future research should focus on refining these algorithms, integrating
additional sensors to capture more nuanced behaviors, and establishing transparent
data collection practices to address privacy concerns.
Smartphone Sensor Dataset for Online Reading Analysis 3

3 Dataset Description

Dataset has three files, one has accelerometer data, another gyroscope data, and the
third file has combined data for both he sensors. Each sensor has x, y, z-axis data
along with the mag which is computer as below:

Accelerometer amag = ax2 + ay2 + az2 (1)

Gyroscope gmag = gx2 + gy2 + gz2 (2)

Addition of the fourth feature, calibrates final readings for the sensors as:

Accelerometer = {ax , ay , az , amag } (3)

Gyroscope = {gx , gy , gz , g mag } (4)

The data has been labelled for the below-mentioned classes [4, 5]:
• scroll,
• pinch to zoom,
• tap,
• orientation change,
• screen capture.

Table 1 Specification table


Subject Machine learning, smartphone screen gesture classification,
human–machine interaction
Specific subject Smartphone screen gesture classification with smartphone sensor data
area
Type of data Smartphone embedded sensors accelerometer and gyroscope data
How the data were Smartphone sensor data was collected while reading a PDF document
acquired using the Books app, an Android PDF reader that also records
accelerometer and gyroscope readings in the background. Developed with
Flutter and Firebase, the application generates output in CSV format,
capturing x, y, and z-axis values along with date and timestamp. For
improved accuracy, a magnitude feature is calculated as well. Sensor data
is recorded at a rate of 1 Hz, capturing readings every second
Data format Raw
Data accessibility Smartphone sensor dataset for online reading
Direct URL to data [Link]
ding-analysis
4 P. Bhatele and M. Bedekar

4 Value of Data

Screen gestures involved for reading content online have been simulated, and
smartphone sensors data was extracted (accelerometer and gyroscope).
• Data collection was done in an obtrusive manner. Users performed reading screen
gestures, and data was collected [5].
• The dataset can be used to implement machine learning classifiers to classify
screen gestures for reading. The dataset is useful to understand user’s behavior;
while reading online, such classification can help provide implicit feedback of the
online available content [2].
• Understanding the screen gestures done while reading can help creating key
performance indicators for the smartphone reading usage.
• The dataset will be valuable for training, testing, and validating reader behavior
during smartphone-based reading sessions [1].
• Effective use of this data can help bridge the research void in implicit feedback
related to online reading [5].

5 Dataset Description

The smartphone sensor dataset for online reading (scroll, tap, pinch to zoom, orien-
tation, screen capture) based on the readings captured from accelerometer (ax, ay, az
axis in meters per second squared (m/s2)) and a gyroscope (gx, gy, gz axis in degrees
per second). Magnitude value for both the sensors is also recorded in the dataset.
The data collection was done using “Books” Android Application [5].
Tables 2 and 3 present raw data samples obtained from the accelerometer
and gyroscope. This data was collected using the default settings of an Android
application.
Smartphone sensors are actively used by young scholars for reading activity. The
trends of reading traditional books with hard covers have declined due to the avail-
ability of e-books and websites available with plethora of content online. Feedback
system to improve reading experiences lacks behind as there are less methods of
extracting feedbacks implicitly. Smartphone sensors have revolutionized the activity

Table 2 Raw data samples from accelerometer


Date Timestamp ax ay az amag Action
07–09–2022 17:31:32 9.927 1.32 3.042 10.46621 Orientation
15–06–2022 12:00:20 0.582 3.954 9.085051 9.92527 Pinch to zoom
08–06–2022 20:14:25 −0.099 5.87895 8.27595 10.15201 Scrolling
07–09–2022 17:29:18 −0.50205 1.272 9.681001 9.777107 Tap
05–11–2022 17:55:24 0.39 3.55695 9.259951 9.92727 Screen capture
Smartphone Sensor Dataset for Online Reading Analysis 5

Table 3 Raw data samples from gyroscope


Date Timestamp gx gy gz gmag Action
07–09–2022 17:31:32 0.023512 0.055138 −0.06848 0.091004 Orientation
07–10–2022 14:13:32 −0.01678 0.00385 −0.01361 0.021944 Tap
01–11–2022 16:02:24 0.006463 0.000825 −0.00028 0.006521 Scrolling
01–11–2022 16:02:24 0.056513 −0.2508 −0.01375 0.257456 Scrolling
02–06–2022 14:29:00 0.000963 −0.00206 0.001238 0.002591 Pinch to
zoom
04–10–2022 12:06:36 −0.07934 −0.06105 0.029838 0.10446 Screen
capture

recognition system; biofeedback, user authentication, and context recognition are


some fields that have successfully supported by these sensors. Effective use of this
data can help bridge the research void in implicit feedback related to online reading.
Feature extraction in time and frequency domain can be implemented and used for
training models. Scrolling, pinch to zoom, and tap are positive indicators of content
and directly relate to reading the content. This identified gesture can be inferred as
user interests.

6 Methodology

The dataset collection was carried out using the ‘Books’ application. The develop-
ment and implementation of this Android-based application, which facilitates PDF
reading and simultaneously captures data from the accelerometer and gyroscope, are
outlined below. The application adheres to a structured methodology to ensure robust-
ness, security, and efficiency. This methodology includes several key phases, each
aimed at addressing specific aspects of the system’s architecture and functionality.

6.1 Technology Stack

Flutter. Flutter, an open-source UI framework developed by Google, can be used


for creating natively compiled applications across mobile, web, and desktop plat-
forms from a single codebase. This project utilizes Flutter as it can deliver a cross-
platform mobile application working on iOS and Android. This approach enhances
development efficiency and ensures a consistent user experience across various
devices.
Inbuilt Sensors. All smartphones contain different types of sensors which capture
different types of information. Some of the common sensors are accelerometers,
6 P. Bhatele and M. Bedekar

gyroscopes, etc. This dataset collection utilized gyroscope and accelerometer sensors
to determine the device orientation.
Gyroscope and Accelerometer. A gyroscope measures the rotational speed around
a device’s three primary axes, making it invaluable for applications that monitor
orientation and angular velocity. In our dataset, gyroscope data is utilized to track
and analyze these rotational movements, providing insights into user interactions
and device handling. Similarly, an accelerometer detects forces along the three
axes, allowing it to capture motion and orientation changes as the device moves.
By combining gyroscope and accelerometer data, our project will gain a detailed
understanding of both rotational and linear movements, offering a comprehensive
view of device motion during user interactions.

6.2 Implementation

Initialization of Sensors. The application leverages Flutter’s sensors plus package


to access the gyroscope and accelerometer sensors, ensuring smooth and efficient
utilization of these essential device functions. By using sensors plus, the app can
continuously collect real-time sensor data, which is crucial for motion detection,
orientation tracking, and any interactive feature dependent on the device’s move-
ments. The package simplifies integration by providing an easy-to-use API for
accessing sensor events and efficiently streaming data. This setup enables the app
to precisely track and adapt to shifts in the device’s orientation and movement,
improving user interaction.
Real-Time Data Collection. A timer class is used to set the sensors to collect data
at respective intervals. The gyroscope records an object’s rotation speed around a
particular axis, giving information about its position and movement in three dimen-
sions. Also, the accelerometer tracks changes in speed along a straight path, known
as linear acceleration. By continuously gathering data on both linear acceleration and
angular velocity, the system can provide detailed insights into an object’s motion and
orientation. The captured data will be stored in a temporary in-memory structure,
such as a list, to ensure efficient real-time data collection without delays.
Data Formatting and CSV File Creation. The sensor data is arranged in Comma-
Separated Values (CSV) format for its simplicity and compatibility with various data
analysis tools. CSV stores data in rows, with each row representing a set of values
separated by commas. This format is easy for both humans and software to read, and it
can be easily imported into programs like Excel or Google Sheets for analysis. Using
CSV ensures that the collected sensor data can be smoothly integrated into existing
data analysis processes without needing complex conversions. Our data has different
columns for time, x, y and z readings for both the accelerometer and gyroscope. This
data is then stored in a firebase real-time database.
Smartphone Sensor Dataset for Online Reading Analysis 7

PDF Library. The app allows users to choose PDF files from their device, using
Flutter’s file picker package. Once selected, these files are added to the app’s library.
Users can manage them by viewing, adding, or removing them. For reading, the app
uses a PDF viewer component called “syncfusion_flutter_pdfviewer,” allowing users
to open and read the PDFs. Users can also track their reading progress, like setting
bookmarks or remembering the last page read. These progress markers are saved
locally, so users can easily resume reading where they left off.
This is the general flow of the functionality in the application given in Figs. 1 and 2.

Fig. 1 Flowchart of the system

Fig. 2 Roadmap to subjects


8 P. Bhatele and M. Bedekar

Fig. 3 Gyroscope readings


during smartphone
orientation changes

7 Experimental Design, Materials, and Methods

The dataset was collected using a OnePlus Nord 2 5G smartphone. A total of 47


subjects were participated in the study, focusing on user gesture activity while using
the Books application. The average age of the participants was 22.95 years. The
selection of high count of subjects in the age group of 15–40. Each subject was given
the below road map [5].

8 Results

The frequency and type of gestures performed during digital reading sessions signif-
icantly affect the user’s reading flow and cognitive load. High-frequency gestures,
such as scrolling and pinch to zoom, performed every 5 s, may disrupt the reading flow
more than less frequent gestures like screen capture, which occurs every 4 s. Addition-
ally, the magnitude of sensor data (x, y, z-axis) associated with each gesture reveals
that more complex gestures (e.g., pinch to zoom, orientation change) demand higher
cognitive effort and physical interaction, potentially leading to increased user fatigue.
In contrast, simpler gestures (e.g., tap) are less intrusive and maintain the continuity
of the reading experience. Figures 3 and 4 represent gyroscope and accelerometer
sensor reading during orientation change on smartphone device, respectively.
Figure 5 indicates the relationship between the number of subjects to their age
groups, which participated in the process of data collection.
The study hypothesizes that optimizing gesture intervals and simplifying complex
gestures can enhance user engagement and reduce cognitive load, thereby improving
the overall efficiency of digital reading applications. This theory aims to contribute to
the design of more intuitive and user-friendly digital reading interfaces. The details
of the results are given in Table 4.
Smartphone Sensor Dataset for Online Reading Analysis 9

Fig. 4 Accelerometer
readings during smartphone
orientation changes

Fig. 5 Relationship between


number of subjects vs. their
age groups

Table 4 Results for the given datasets


Aspect Details
Data collection Accelerometer and gyroscope data collected using the “Books” application
Dataset Separate folders for accelerometer and gyroscope data, plus a combined file
organization
Sensor data Captures x, y, z-axis data and computed magnitude
Gestures captured Scrolling, pinch to zoom, tap, orientation change, screen capture
Gesture intervals Scrolling every 5 s, pinch to zoom every 5 s, tap every 5 s, orientation change every 5 s, screen
capture every 4 s
Participants Forty-seven subjects using OnePlus Nord2 5G smartphone
Participant age Mean age: 22.95 years, majority aged 15–40 years
Reading session 8–10 min
duration
Technology used Developed using Flutter and Firebase, utilizing the sensors_plus package
Data recording rate 1 Hz (one reading per second)
Data format CSV format for easy analysis
Potential Training/testing machine learning models, understanding screen gestures, creating KPIs for
applications reading usage, providing implicit feedback
Ethical Informed consent obtained, data publicly available, no financial/personal conflicts
considerations
10 P. Bhatele and M. Bedekar

References

1. Wawage P, Deshpande Y (2022) Smartphone sensor dataset for driver behavior analysis, data in
brief, vol 41, p 107992. ISSN 2352 3409. [Link]
2. Zahoor S, Bedekar M, Kosamkar PK (2014) User implicit interest indicators learned from the
browser on the client side. In: Proceedings of the 2014 international conference on information
and communication technology for competitive strategies (ICTCS ’14), Article 57. Association
for Computing Machinery, New York, NY, USA, pp 1–4
3. Bhalekar M, Bedekar M (2022) The new dataset MITWPU-1K for object recognition and image
captioning tasks. Eng Technol Appl Sci Res 12:8803–8808. [Link]
4. Alqarni MA, Chauhdary SH, Malik MN et al (2020) Identifying smartphone users based on how
they interact with their phones. Hum Cent Comput Inf Sci 10:7
5. Bhatele P, Bedekar M (2023) Survey on smartphone sensors and user intent in smartphone usage.
In: 2023 IEEE 8th international conference for convergence in technology (I2CT), Lonavla,
India, pp 1–9. [Link]
6. Smith A, Johnson B, Williams C (2019) Understanding user engagement through desktop
peripherals. J User Interact 6(2):45–56
Bridging the Gap: Integrating
Generative AI into Engineering
Education 5.0 for Industry 5.0

Mitali Chugh, Shreya Panwar, Khushi Allawadi, and Sonali Vyas

Abstract There has been a dramatic shift in the last two decades with regard to
the technology shift and a swelling wave of disruptive hardware and software tech-
nologies. The advent of Industry 5.0 has impacted our lives, mainly engineering
education, significantly. Industry 5.0 emphasizes human importance, sustainability,
and modern technologies like artificial intelligence, the Internet of Things, and robots.
In this research, we want to explore how Industry 5.0 engineering education might
integrate generative AI. The specific objectives are to analyze the feasibility and
challenges in integrating generative AI into the curricula of engineering studies and
to identify the impact of its implementation on student learning outcomes. To better
comprehend the applications by engineering students of generative AI systems, we
sent an online questionnaire. The correlations obtained with our results and usage
of the generative AI tools on improving understanding concepts of the computer
science or engineering discipline also point toward improvements in a complex way;
this study provides a proof and possibility in which generative AI can introduce in the
understanding of student regarding Industry 5.0 requirements and discusses further
on how the future may be regarding the engineering educations.

Keywords Industry 5.0 · Generative AI · Engineering education

1 Introduction

Without a doubt, the progressive evolution of technology is changing how people


live, work, communicate, and think. The environment in which we live has seen some
incredible advancements due to the ongoing progress of technology. These changes
have affected communication styles, work dynamics, and even cognitive functions
[1].

M. Chugh · S. Panwar · K. Allawadi · S. Vyas (B)


UPES University, Dehradun, India
e-mail: [Link]@[Link]

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 11
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
12 M. Chugh et al.

Gürdür Broo et al. [1]. The “digital transformation age” is an era of the radical change of
technology induced by the intersection of cloud computing, Internet of Things, and others.
In this revolutionary wave, the idea of Industry 5.0 came into existence as an expression of
creativity and flexibility and gained special support in the aftermath of the global COVID-19
pandemic [2]. Industry 5.0 is aimed at developing products and processes for producing
a seamless integration of technology into every part of the consumer experience, bringing
forth engineering landscapes and significant milestones that will affect the future.
Knowing that engineering education is going to shape the future workforce, education
paradigms need to be updated in accordance with the demands of Industry 5.0, and it needs
to be integrated into the curriculum of higher institutions [20]. In this, Generative AI is a
potential tool that will improve the learning process and prepare professionals for Industry
5.0 challenges.
Such AI is an interesting new frontier of artificial intelligence that is powered by foundation
models to create novel content in all types of modalities, whether in text, audio, video, or
graphics. They are able to recognize complex patterns and learn how to adapt with very
little training from large datasets [2]. With the rise of generative AI comes a wide array of
significant concerns as to its potential effects and ethical implications, as any new technology
will do.
It proposes to clarify the mutually beneficiary nature of the interaction between specifications
of Industry 5.0 and the incorporation of the Generative AI tools represented by ChatGPT
or similar platforms [2]. The entry of ChatGPT into the wave of Industry 5.0 signifies great
possibilities but signals a complete redefinition needed to be done for job responsibility
profiles or skills employees should possess in the new environment. [3] Although generative
AI has been used in engineering education in the past, there is no focused research on how it is
related to Industry 5.0, whereas it has many applications for Industry 4.0, such as providing
lists of tasks and schedules, real-time direction and training to the workers, analysis of
possible dangers, and automation of various tasks [4]. Stepping into this uncertain terrain,
we have to be aware of the many risks and benefits that generative AI poses in education.
Utilizing generative AI technologies like Dall-E, GPT, and bard for extensive tasks and
learning, has become a global fever due to its capabilities to provide extraordinary results.
Currently, generative AI is being used for many tasks such as predicting equipment fail-
ures before it happen, image analysis and anomaly detection, enhancing robot–human
collaboration in smart manufacturing environments, and much more [19].
GAI can be idyllically aligned with Engineering education for Industry 5.0 as it can provide
personal tutoring, exam preparations, coding suggestions, visualization of concepts, and
even help with PowerPoint presentations [20].

Following are some recent real-life examples of applications of generative AI


being used in engineering education to get people ready for Industry 5.0:
• MATHia at Carnegie Learning: MATHia’s individualized tutoring approach
is being used across STEM professions, including engineering, while being
primarily focused on arithmetic. By using artificial intelligence (AI) to evaluate
each student’s progress and adaptively modify learning paths, MATHia gives
engineering students unique challenges and problem-solving exercises that suit
their learning preferences and current comprehension level.
• Labster: This platform offers simulations in virtual labs ranging from mate-
rials science to bioengineering. Labster’s AI-driven scenario generation allows
Bridging the Gap: Integrating Generative AI into Engineering Education … 13

students to perform experiments in a secure virtual environment, simulating intri-


cate chemical interactions or manufacturing procedures that are essential in fields
like materials engineering and pharmaceuticals.
• By McGraw-Hill, ALEKS: When used for engineering courses like calculus,
ALEKS adjusts to the unique learning requirements of each student, assisting
them in understanding basic ideas before moving on to more difficult subjects.
As engineering students tackle difficult STEM topics, ALEKS uses AI-driven
assessments to determine understanding and offer tailored activities.
The following sections of this paper outline a methodical approach to our research
work. First, we formulate our research objective and outline the scope of our inves-
tigation to Industry 5.0 and engineering education. Next, we do an in-depth exami-
nation of the body of existing literature, emphasizing important themes, trends, and
gaps in it. We discuss our methodology in detail, clarifying the tactics used for gath-
ering and analyzing data. Our research culminates in a critical analysis in which
we assess our results concerning their consequences for engineering education and
Industry 5.0 requirements. Finally, we provide practical recommendations to support
the smooth integration of Generative AI into engineering pedagogy, thus accelerating
the shift to Engineering Education 5.0.

2 Literature Review

See Table 1.

3 Methodology

Engineering professionals with hands-on expertise and problem-solving abilities


are highly sought after in the era of Industry 5.0, where people and cutting-edge
technology coexist in an advanced industrial setting. Generative AI has proven to be
an excellent educational resource for reaching a higher degree of conceptual clarity
[5]. Our primary motivation is our concern about how GPT has facilitated learning
and has eased the challenges that people have in present and future generations.
The purpose of our survey design was to gather as much information as possible on engi-
neering students’ opinions and experiences with using generative AI platforms in their
learning in the context of Industry 5.0. The selection of questions was justified by the
requirement to assess students’ grasp of fundamental computer science ideas as well as
their familiarity with and interactions using Generative AI tools. The survey instrument was
a Google Form with 14 multiple-choice questions (MCQs) that were purposefully designed
to evaluate different levels of difficulty in computer science disciplines, such as coding
output-based queries. With a combination of diverse difficulty-level questions, we sought
to provide a more comprehensive overview of the student’s competency levels. Moreover,
the survey ended with a question about the use of Generative AI platforms by respondents
14 M. Chugh et al.

Table 1 Discussing the problem and solution of some of the papers that were referenced
Papers Authors Problem Solution
1 ArticuloEngineeringEducation Need for reform in Evolve toward more
5.0 [3] engineering education multidisciplinary
owing to the emergence schemes like initiatives
of Industry 5.0, from EACEA and R&D
extraordinary changes, projects
unsustainable growth,
and ethical concerns
2 A. Díaz Lantada[4] Need for keeping Project-based learning
students multifaceted (PBL) methodologies
with the evolving should also evolve
industry and its needs
3 Longo, Francesco, Antonio The Influence of A value-oriented and
Padovano, and Steven Umbrello Technology on Workers ethical technology
2020 [5] and Society in Industry engineering approach,
5.0 presents serious such as the Value
ethical concerns Sensitive Design (VSD)
methodology
4 Xu, X., Lu, Y., Vogel-Heuser, Challenges of Industry Techno-social revolution
B., & Wang, L. (2021) [6] 4.0 and Industry 5.0 may be underway
co-existence through
discussions and
clarifications
5 Joanna Morawska-Jancelewicz Universities are not Encourage the “fourth
[7] maximizing their mission,” which
ability to influence emphasizes sustainable
social, cultural, and development in the
economic growth community
6 Gürdür Broo, D., Kaynak, O., & The challenges offered Transdisciplinary
Sait, S. M. (2021) [1] by the fifth industrial education, sustainability,
revolution, Industry resilience,
5.0, which necessitates human-centered design
a shift in thinking and modules, hands-on data
behavior fluency and management
courses, and
human-agent/machine/
robot/computer
interaction experiences
are all offered
7 Rosario Michel-Villarreal [8] The disconnect Well-structured
between university integration of Generative
curriculum and AI AI into courses for
tools’ capabilities interactive learning
(continued)
Bridging the Gap: Integrating Generative AI into Engineering Education … 15

Table 1 (continued)
Papers Authors Problem Solution
8 Fei-Yue Wang; Jing Yang; The use of AI models Establish community
Xingxia Wang; Juanjuan Li; like ChatGPT may not norms and encourage
Qing-Long Han [9] provide ethics ethical use of AI tools,
guidelines and equal while guaranteeing fair
access so there can be access and training to
chance of risk that it avoid educational
will undermine inequities
educational integrity
while it offers
opportunities in
engineering education
9 Nitin Rane [10] Ethical considerations, For individuals, priority
data security, and should be given to
alignment of AI to continuous education and
people’s values training
In Society 5.0, it is vital Create artificial
to ensure that AI intelligence with
improves the potential human-centric values and
of humans, supports increase human capacity
inclusion, and bridges
societal divides

to understand challenging concepts. The investigation of the perceived effectiveness and


influence of generative AI tools on students’ learning experiences are centered around the
final question. Figure 1 shows the flow of our process of getting insights from students.

4 Results

Several important findings have emerged from the examination of the integration of
engineering education with the demands of Industry 5.0, facilitated using generative
AI platforms such as ChatGPT. To get a thorough comprehension of the present state
of knowledge, we developed a quiz for students to evaluate their understanding of
fundamental concepts and emerging trends in Computer Science and Engineering
(CSE). Table 2, Fig. 2, and Fig. 4 reveal that 25.5% of students reported using
Generative AI tools, while 74.5% did not seek any assistance and out of the Gen AI
users, most of the responses were wrong meanwhile people who used Gen AI had the
most correct number of responses. Based on data gathered from 160 students, Fig. 3
shows that 53.75% of students received a score of 16 or less out of 22, while 46.25%
received a score higher than 16. This distribution highlights the student cohort’s
moderate level of knowledge [5].
16 M. Chugh et al.

Specifically, the ratio of correct responses among Generative AI users to non-users was
approximately 7:3 [11]. Figure 5 shows an example of questions that we provided to students
through our Google Form. This suggests that engagement with Generative AI may enhance
understanding of complex concepts in computer science.

The frequency distribution and mean scores of both correct and incorrect answers were
determined after a detailed quantitative analysis of the quiz responses [4]. This was done to

Fig. 1 Methodology
flowchart
Bridging the Gap: Integrating Generative AI into Engineering Education … 17

Fig. 2 Results for correct and incorrect responses by user type

Fig. 3 Average, median, and


range of correct responses

Fig.4 Survey question to


determine the number of
people who used Gen AI

prove the effectiveness of generative AI to enhance the learning experience of students of


CSE background.
These results tell how important it is to incorporate generative AI into engineering curricula to
satisfy Industry 5.0’s evolving requirements. Moreover, they draw attention to the necessity
of more investigation into the pros of this integration and how it might lead to improvement
of students’ competence in challenging technological fields [10].
18 M. Chugh et al.

Fig. 5 Survey question types

5 Discussion

The objective of our research has been to align the current generative AI techniques
with engineering education which will result in better understanding and grasping
of complex concepts. Based on the results of our structured assessment, we found
that 53.75% of students had a moderate comprehension of computer science and
engineering concepts, scoring 16 or fewer out of 22. Furthermore, only 25.5% of
participants used generative AI tools, indicating unrealized potential for improving
learning outcomes in challenging CSE themes following our research objectives.
The usage of Generative AI tools and quiz performance have a beneficial connection, which
demonstrates the applicability of our study in meeting Industry 5.0 needs. This is an area
that has been receiving little attention in the evolution of engineering education. Online
surveys can help in identifying how helpful ChatGPT-like tools can be for purposes like
research and projects as well [6]. Generative AI can be beneficial when it is integrated with
software development and learn about mistakes through it [7]. The advantages of Gen AI
in improving learning outcomes have been demonstrated by earlier studies as well such as
personalized learning experiences offered by several platforms. Nonetheless, our research
offers distinctive perspectives on the function of Generative AI in engineering education,
particularly to Industry 5.0 requirements [1, 2].
The low adoption rate of Generative AI tools presents an opportunity for educational institu-
tions to integrate these technologies into curricula, better preparing students for the demands
of Industry 5.0 and enhancing their ability to tackle evolving technological challenges.
However, Limitations such as data security and ethical considerations, as well as the poten-
tial for inaccurate responses generated by Generative AI models have also been considered
in our study. Despite these challenges, Generative AI has a tremendous scope in future
to further improve human-centric approaches, serving as virtual tutors and enhancing the
learning experience [10].
Bridging the Gap: Integrating Generative AI into Engineering Education … 19

Our methodology encountered problems like biased responses, which emphasized the impor-
tance of validation and supervision in research projects, even after having an effective method
of gathering the data. Regardless of these limitations, our study provided comprehensive
details on how generative AI can aid students in studying and how in future there is a great
probability of such learning techniques to grow.
It’s essential to recognize the constraints of our survey methodology, even if it revealed
insightful information about the relationship between generative AI and engineering educa-
tion [8]. At the outset, depending too much on one technique of gathering data could have
resulted in bias or missed important aspects of the study question. Furthermore, response
biases may have resulted from the survey’s dependence on self-reported data, which could
have skewed the findings.
To address these constraints and offer a more comprehensive understanding of the subject,
future studies might choose to speculate about integrating qualitative and quantitative
methodologies. In addition to quantitative data collected through surveys, qualitative
methods like focus groups and interviews may provide deeper sights into students’ views
and experiences with generative AI tools.

6 Conclusion

This is to conclude that to effectively prepare students for the demands of Industry
5.0, generative AI plays a crucial role in engineering education. Our study shows
that using generative AI tools improves quiz scores and understanding of students,
suggesting that these technologies have the potential to improve learning outcomes
in challenging technical fields. However, Limitations like data security breaches and
ethical challenges with generative AI applications must be addressed. Subsequent
investigation ought to concentrate on creating strong structures for incorporating
Generative AI into curricula and investigating inventive methods for personalized
education.
To sum up everything, the development of education standards as an industry has witnessed
a transition from Education 1.0 to the latest Education 5.0 [9]. Industry 5.0 and engineering
education will be greatly impacted by the incorporation of generative AI. By utilizing
these technologies, educational institutions can better prepare students for success in a fast-
changing technology. Moving forward, collaborative efforts are required to fully utilize
Generative AI in shaping the workforce of the future and promoting Industry 5.0 targets.

References

1. Gürdür Broo D, Kaynak O, Sait SM (2022) Rethinking engineering education at the age of
industry 5.0. J Ind Inf Integr 25:100311. [Link]
2. Carayannis EG, Morawska-Jancelewicz J (2022) The futures of Europe: Society 5.0 and
Industry 5.0 as driving forces of future universities. J Knowl Econ 13(4):3445–3471. https://
[Link]/10.1007/s13132-021-00854-2
3. ArticuloEngineeringEducation 5.0
20 M. Chugh et al.

4. Díaz Lantada A (2022) Engineering Education 5.0: strategies for a successful transformative
project-based learning. In: Insights into global engineering education after the birth of Industry
5.0. IntechOpen. [Link]
5. Longo F, Padovano A, Umbrello S (2020) Value-oriented and ethical technology engineering
in Industry 5.0: a human-centric perspective for the design of the factory of the future. Appl
Sci 10:4182. [Link]
6. Xu X, Lu Y, Vogel-Heuser B, Wang L (2021) Industry 4.0 and Industry 5.0—inception,
conception and perception. J Manuf Syst 61:530–535. [Link]
10.006
7. Morawska-Jancelewicz J (2022) The role of universities in social innovation within quadruple/
quintuple helix model: practical implications from polish experience. J Knowl Econ
13(3):2230–2271. [Link]
8. Michel-Villarreal R, Vilalta-Perdomo E, Salinas-Navarro DE, Thierry-Aguilera R, Gerardou
FS (2023) Challenges and opportunities of generative AI for higher education as explained by
ChatGPT. Educ Sci (Basel) 13(9). [Link]
9. Wang FY, Yang J, Wang X, Li J, Han QL (2023) Chat with ChatGPT on Industry 5.0: learning
and decision-making for intelligent industries. In: IEEE/CAA Journal of Automatica Sinica,
vol 10, no 4, April 01, 2023. Institute of Electrical and Electronics Engineers Inc., pp 831–834.
[Link]
10. Rane N (2023) ChatGPT and similar generative artificial intelligence (AI) for smart industry:
role, challenges and opportunities for Industry 4.0, Industry 5.0 and Society 5.0. SSRN Electron
J. [Link]
11. Mourtzis D, Angelopoulos J (2023) Development of an extended reality-based collaborative
platform for engineering education: Operator 5.0. Electronics (Switzerland) 12(17). https://
[Link]/10.3390/electronics12173663
12. Kusam VA (2024) Generative-AI assisted feedback provisioning for project-based learning in
CS education
13. Qadir J (2023) Engineering education in the era of ChatGPT: promise and pitfalls of gener-
ative AI for education. In: IEEE global engineering education conference, EDUCON. IEEE
Computer Society. [Link]
14. Rane N (2023) ChatGPT and similar generative artificial intelligence (AI) for building and
construction industry: contribution, opportunities and challenges of large language models for
Industry 4.0, Industry 5.0, and Society 5.0. SSRN Electron J. [Link]
3221
15. Javaid M, Haleem A, Singh RP (2023) A study on ChatGPT for Industry 4.0: background,
potentials, challenges, and eventualities. J Econ Technol 1:127–143. [Link]
ject.2023.08.001
16. Kiangala KS, Wang Z (2024) An experimental hybrid customized AI and generative AI chatbot
human-machine interface to improve a factory troubleshooting downtime in the context of
Industry 5.0. Int J Adv Manuf Technol. [Link]
17. Nikolic S et al (2023) ChatGPT versus engineering education assessment: a multidisciplinary
and multi-institutional benchmarking and analysis of this generative artificial intelligence tool
to investigate assessment integrity. Eur J Eng Educ 48(4):559–614. [Link]
43797.2023.2213169
18. Petrovska O, Clift L, Moller F, Pearsall R (2024) Incorporating generative AI into software
development education. In: ACM international conference proceeding series, January 2024.
Association for Computing Machinery, pp 37–40. [Link]
19. Sai S, Sai R, Chamola V (2024) Generative AI for iIndustry 5.0: analyzing the impact of
ChatGPT, DALLE, and other models. IEEE Open J Commun Soc
20. Supriya Y, Bhulakshmi D, Bhattacharya S, Gadekallu TR, Vyas P, Kaluri R, … Mahmud M
(2024) Industry 5.0 in smart education: concepts, applications, challenges, opportunities, and
future directions. IEEE Access
Proposing an AI-Enabled Waste
Segregation System for Domestic Settings

Devnarayan G. Rao, G. Sangeetha, J. Sandeep, C. S. Sreeja,


Agnes Nalini Vincent, and Teena Mary

Abstract This paper proposes an innovative AI-based system for automated


domestic waste segregation. Utilizing Teachable Machine and MobileNet, the system
accurately categorizes waste into dry and wet components, laying the foundation for
sustainable waste management practices. Embedded in a Raspberry Pi 4, the system
integrates real-time image processing with various sensors to streamline the sorting
process. While the model has been simulated due to budgetary constraints, future
implementation envisions real-world application. Potential advancements include
expanding the dataset, enabling multi-category waste classification, and exploring
low-power alternatives. This research contributes to the evolving landscape of smart
waste management, addressing environmental sustainability and the pressing need
for automated, efficient waste segregation at the domestic level.

Keywords Artificial intelligence for sustainability · Transfer learning · Smart


waste management · IoT · Waste segregation

D. G. Rao (B) · G. Sangeetha · J. Sandeep · C. S. Sreeja · T. Mary


Department of Computer Science, Christ University, Bangalore, India
e-mail: [Link]@[Link]
G. Sangeetha
e-mail: sangeetha.g@[Link]
J. Sandeep
e-mail: sandeep.j@[Link]
C. S. Sreeja
e-mail: [Link]@[Link]
T. Mary
e-mail: [Link]@[Link]
A. N. Vincent
Department of Information Technology, AMITY Institute of Higher Education, Quatre Bornes,
Mauritius
e-mail: vanalini@[Link]

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 21
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
22 D. G. Rao et al.

1 Introduction

Globally, more than two billion tons of municipal solid waste are produced each
year, and a significant portion, at least 33%, is not handled in an environmentally safe
manner. On average, individuals generate about 0.74 kg of waste per day worldwide,
but this varies widely from 0.11 to 4.54 kg. Managing solid waste in cities of low-
and middle-income countries (LMICs) poses significant challenges. The improper
disposal of household waste has far-reaching environmental and health implications.
Due to the increasing ecological, social, and economic concerns associated with
these issues, waste management is gaining recognition as a crucial problem among
governments, businesses, nongovernmental organizations (NGOs), academics, and
the general public [1]. Despite advancements in social, economic, and environmental
aspects, SWM systems in India have seen limited progress. The informal sector plays
a crucial role in extracting value from waste, but approximately 90% of residual waste
is currently dumped rather than properly landfilled [2].
Recognizing that many issues originate at the domestic level, it becomes evident
that an efficient, automated waste sorting solution at the household level could poten-
tially mitigate the overall volume of unsorted waste, presenting a crucial step toward
improved waste management in India.
The current garbage disposal system in India collects unorganized waste from
residential areas, subsequently undergoing segregation at specific stations. The segre-
gation, performed manually, poses various health risks for the laborers involved. This
manual process is not only time consuming but also requires financial contributions
from the workers. The unregulated disposal of waste on the outskirts of towns and
cities has led to the formation of overflowing landfills. These landfills, created in
a disorderly manner, present challenges for reclamation. The consequences include
severe environmental implications, such as groundwater pollution and contributions
to global warming. Furthermore, the existing system has been identified as having a
detrimental impact on the average lifespan of the manual segregators engaged in the
waste segregation process [3].
The proposed system introduces an automated waste segregation solution at the
domestic level, utilizing a model created in Teachable Machine with transfer learning
and MobileNet, embedded in a Raspberry Pi 4 board. This advanced system, featuring
LEDs and servo motors in the smart dustbin, eliminates the need for manual labor,
reducing associated health hazards and financial burdens. The integration of hardware
components enhances efficiency and accuracy in waste categorization.
A detailed examination of existing automated waste sorting systems reveals inno-
vative solutions such as a proposed waste management system with a sorter bin
and composting unit, and an autonomous waste sorting machine utilizing Tensor-
Flow and Faster R-CNN for object detection. These systems address crucial aspects
of waste segregation, utilizing various hardware components and machine learning
algorithms. Nevertheless, it is imperative to acknowledge the constraints associ-
ated with depending exclusively on an inductive sensor for waste segregation. Even
though the envisioned system adeptly discerns metallic from non-metallic waste, it
Proposing an AI-Enabled Waste Segregation System for Domestic Settings 23

confronts potential hurdles in precisely identifying various non-metallic waste types.


The inherent variability in the characteristics of non-metallic materials introduces the
prospect of misclassifications, potentially compromising the overall effectiveness of
the waste segregation process. In contrast, our proposed AI-enabled waste segrega-
tion system introduces a distinctive approach by leveraging IoT (using Raspberry Pi
4) and training the model with specialized dry and wet waste images.

2 Related Work

In recent years, the escalating concern surrounding domestic waste disposal has
prompted initiatives by municipal corporations worldwide to address the issue of
waste management. While efforts have been made to segregate dry and wet wastes, a
substantial portion of the domestic population remains indifferent to these initiatives,
resulting in persistent waste management challenges.

2.1 Existing Smart Waste Management Systems

Shreeshayana et al. [4] have proposed a system to address the escalating concern
surrounding domestic waste disposal, outlining the design and development of a
waste management system focused on segregating dry and wet organic waste gener-
ated in kitchens. This proposed system comprises two integral components: a sorter
bin for separating dry and wet waste, and a composting unit aimed at converting
organic waste into compost. The paper details the functionality of the proposed
waste management system, which integrates various components to achieve efficient
waste segregation and compost production. The system is powered by an ATMEga
328 Arduino Uno, interfaced with a stepper motor, an HRC ultrasonic sensor, and an
inductive sensor. The inductive sensor plays a crucial role in identifying the nature
of the waste and distinguishing between metallic and non-metallic components.
However, it is essential to consider the limitations of relying solely on an inductive
sensor for waste segregation. While the proposed system effectively distinguishes
between metallic and non-metallic waste, it may face challenges in accurately identi-
fying the types of non-metallic waste. Variability in the characteristics of non-metallic
materials could potentially lead to misclassifications, impacting the overall efficacy
of the waste segregation process.
Chowdhury et al. [5] have also implemented an autonomous waste sorting machine
featuring a wiper motor for conveyor belt propulsion and servo motors for obstruction
upon object detection. The hardware setup includes a funnel shaped structure for
individual waste intake, enabling subsequent detection. The software implementation
leverages TensorFlow and the Faster R-CNN algorithm for machine learning in their
work. Real-time videos from a mounted camera are processed to achieve accurate
object detection. Communication via serial connection between an Arduino and a
24 D. G. Rao et al.

computer facilitates precise control of the servo motors, contributing to the reliable
and automated waste detection and separation process.

2.2 Existing AI Models for Waste Segregation

In [6], Yinghao Chu et al. have proposed a multilayer hybrid deep learning system
(MHS) for automatic waste sorting in urban public areas, achieving an overall classi-
fication accuracy higher than 90% under two different testing scenarios and outper-
forming a reference CNN-based method relying on image-only inputs. The specific
contributions of the paper are achieving excellent accuracy for in-field applications
and proposing an innovative architecture to simulate the sensory and intellectual
process of human inspections.
Hyunh et al. [7] have discussed the use of deep learning to automatically sort solid
waste, with a focus on the classification of recyclable and non-recyclable waste,
including a new type of trash-organic waste. They also highlight the use of pre-
trained models such as VGG16, VGG19, Resnet34, Resnet101, DenseNet121, Effi-
cientNetB0, and EfficientNet-B1, and the evaluation of the network’s effectiveness
through accuracy and confusion matrix.
Sousa et al. [8] have proposed a hierarchical deep learning approach for waste
detection and classification in food trays, which outperforms the direct use of Faster
R-CNN, along with the presentation of the “Labeled Waste in the Wild” dataset
to address limitations of existing waste datasets. The hierarchical method based on
shape achieves the best performance, improving the mAP by 11.9% compared to the
flat approach.
The research done by Shuang Liang and Gu [9] proposes a benchmark for evalu-
ating multi-label waste classification and localization tasks using deep learning-based
methods, achieving high scores for multiple tasks related to waste management. The
WasteRL dataset used in the experiments is larger than other existing datasets and
has a wide range of scales and aspect ratios, making it more comprehensive and
challenging for the waste management community.
The research performed by Bircanoglu [10] aims to demonstrate an efficient intel-
ligent system for classifying common waste materials, and it presents promising
results, including achieving 95% test accuracy with a fine-tuned DenseNet model.
They have developed a very successful model called RecycleNet which has been
widely cited and worked on.
Cao and Xiang [11] have proposed an approach that effectively classifies organic
and recyclable waste types, achieving an accuracy of 99.95% in the experiment. The
study consists of four key steps, including classification using features from CNNs,
reconstruction with AutoEncoder, feature selection with RR method, and validation
with tenfold cross-validation. The methodology includes the use of deep learning
methods, specifically AutoEncoder and Convolutional Neural Network (CNN) archi-
tectures, for waste classification, a combination of features extracted from different
CNNs to improve classification success, utilization of the Ridge Regression (RR)
Proposing an AI-Enabled Waste Segregation System for Domestic Settings 25

method for feature selection, and training and validation of the proposed model
using the tenfold cross-validation method.
In [12], Toğaçar et al. present a highly accurate model to classify garbage into
seven different categories using the CompostNet dataset and pre-trained models, with
accuracies of 96.42, 96.27, and 96.273%. The research aims to provide better waste
categorization and aligns with the United Nations’ goal for Responsible Consumption
and Production toward sustainable development. It also discusses the significant issue
of solid waste accumulation and production in India and the United States.
Existing AI models for waste sorting, such as those proposing multilayer hybrid
deep learning systems [6], employing deep learning with pre-trained models [7], and
introducing hierarchical approaches [8], have made significant strides in improving
accuracy and efficiency. The proposed system builds upon these foundations by
integrating a Teachable Machine AI model specifically trained for domestic waste
segregation. Unlike many existing models that focus on specific waste types or
urban scenarios, the proposed solution aims to cater to diverse household waste.
Incorporating a Raspberry Pi 4 and Arducam 5MP OV5647 camera introduces a
cost-effective and scalable hardware platform, enhancing accessibility to advanced
waste sorting technology. Furthermore, while existing models showcase the effi-
cacy of deep learning, the proposed system’s innovation lies in its user-friendly
approach, leveraging Teachable Machine to make waste segregation technology
adaptable for widespread domestic use. In essence, the proposed system distinguishes
itself by prioritizing simplicity, scalability, and effectiveness in a domestic setting,
contributing a novel dimension to the evolving field of waste management.

3 Proposed Methodology

In this section, there is an emphasis on the innovative methodology proposed for the
development of the AI-enabled waste segregation system. The hardware configura-
tion, including Raspberry Pi 4, LEDs, and servo motors, will be detailed alongside the
software aspect, featuring the application of Teachable Machine’s MobileNet-based
transfer learning for efficient waste categorization.

3.1 Software Components

Teachable Machine [14] serves as a user-friendly platform for creating machine


learning models without the need for extensive coding expertise. It simplifies the
process of training models by providing an intuitive interface for dataset input,
labeling, and model generation. This platform abstracts the complexities of under-
lying algorithms, making it accessible for users to implement sophisticated models,
such as convolutional neural networks (CNNs), for waste categorization.
26 D. G. Rao et al.

Convolutional neural networks (CNNs) [13] form the core of the software compo-
nent, playing a pivotal role in image recognition and classification. These neural
networks are designed to mimic human visual perception, allowing them to learn hier-
archical features within images. In the context of waste segregation, CNNs become
instrumental in distinguishing between different types of waste, ensuring the accu-
racy and efficiency of the categorization process. CNNs have been embedded into
the Teachable Machine Model, which has been integrated using the [Link].
Transfer learning, a technique where a pre-trained neural network is fine-tuned
for a specific task, is employed to enhance the efficiency of waste categorization.
MobileNet, a lightweight convolutional neural network architecture, serves as the
backbone for transfer learning. This architecture strikes a balance between model
size and performance, making it suitable for deployment in resource-constrained
environments like embedded systems.
The integration of MobileNet within the Teachable Machine framework ensures a
seamless workflow for model creation and training. MobileNet’s pre-trained features
are leveraged, and Teachable Machine allows users to customize and extend these
features to accommodate the unique waste segregation requirements. This integration
forms a cohesive unit where the strengths of MobileNet and the user-friendly interface
of Teachable Machine synergize for effective model development.

3.2 Dataset Training

During the training phase, the model is trained on a dataset comprising two labels,
namely “dry” and “wet”. Approximately 12,000 images are utilized for training the
model to recognize wet waste, while 10,000 images are dedicated to training for dry
waste categorization.

3.3 Technical Implementation of Teachable Machine, CNNs,


and Raspberry Pi Integration

In the technical setup, the proposed system leverages the advanced capabilities of
the Raspberry Pi 4 as the primary processing unit. The Raspberry Pi 4 introduces
enhanced computational power and flexibility, providing an efficient platform for
hosting and running the Teachable Machine’s AI model. The camera used is the
Arducam 5MP OV5647 Camera Module, equipped with a Motorized IR-CUT Filter
for Daylight and Night vision, specifically designed for seamless integration with
Raspberry Pi.
The training process with Teachable Machine remains consistent, utilizing convo-
lutional neural networks (CNNs) to discern intricate patterns and features from
an extensive dataset comprising approximately 12,000 images for wet waste and
Proposing an AI-Enabled Waste Segregation System for Domestic Settings 27

10,000 images for dry waste. The CNNs undergo transfer learning with a pre-trained
MobileNet model, adapting to the unique characteristics of waste items during the
customization process.
The backend architecture involves [Link] for model training, where
JavaScript is employed to facilitate the training and execution of the models. Transfer
learning principles continue to play a pivotal role, ensuring that the pre-trained
MobileNet model serves as a foundational framework, while the user-customized
classes refine the neural network’s ability to distinguish between dry and wet waste.
The proposed system incorporates Raspberry Pi 4, and Python is the primary
programming language for interfacing with the Pi’s hardware components. The
integration involves using the Picamera library to capture images in real time. The
Arducam 5MP OV5647 Camera Module, coupled with a Motorized IR-CUT Filter
for Daylight and Night vision, plays a pivotal role in capturing high-quality images
for waste categorization.
In the Raspberry Pi OS environment, a Python script is executed to capture images
using the connected camera. The captured image is then processed through the Teach-
able Machine’s TensorFlow Lite (TFLite) model, providing real-time inference for
waste categorization. Based on the classification results, the Python script triggers the
corresponding servo motors, activating their flaps to segregate the waste accurately.

3.4 Hardware Components

The Raspberry Pi 4 serves as the central processing unit for the proposed waste
segregation system. Boasting enhanced processing power, this compact single-board
computer provides the computational capability needed to run the machine learning
models and interface with various hardware components seamlessly. Its GPIO pins
facilitate communication with peripherals like the Arducam camera module and servo
motors, making it a versatile and efficient choice for the smart waste segregation
system. The Arducam 5MP OV5647 Camera Module selected for this project is a
versatile imaging solution designed for both daylight and night vision applications.
Equipped with a motorized IR cut filter and two infrared LED illuminator boards, it
ensures optimal imaging in varying lighting conditions. The motorized IR cut filter
automatically switches ON/OFF, adapting to the ambient light, while the infrared
LEDs provide additional illumination as needed. The Passive Infrared (PIR) sensor
serves as a pivotal component in the proposed waste segregation system, playing
a crucial role in detecting motion within the designated staging area of the smart
dustbin. This sensor operates on the principle of capturing infrared radiation emitted
by objects in its field of view. Upon sensing any movement, the PIR sensor promptly
signals the system to initiate the segregation process. In the context of the system,
the PIR sensor serves as a responsive trigger, activating subsequent stages of waste
categorization and ensuring a timely and efficient sorting mechanism (Fig. 1).
28 D. G. Rao et al.

Fig. 1 Hardware block diagram

3.5 Software Components

Upon booting up the Raspberry Pi 4 and initiating its operating system, the system
activates the designated Python file ([Link]) to commence operation. The Passive
Infrared (PIR) sensor vigilantly monitors the staging area of the bin, ready to detect
any movement. When an object is identified, the program triggers the RGB LED to
turn red and enforces a brief 15-s pause (Fig. 2).
After the interval, the Arducam 5MP OV5647 Camera Module captures an image
of the waste. The captured image then undergoes classification through the embedded
machine learning model. The model’s outcome dictates the waste type, prompting
the corresponding servo motor to rotate. The servo’s movement opens the designated
flap, guiding the waste into the appropriate bin. Subsequently, the servo motor returns
to its original position, and the LED transitions to green, signaling the successful
completion of the segregation process.

3.6 Practical Validation and Performance Evaluation

Real-world testing was conducted within an Indian household, utilizing waste images
taken from Indian households to validate the practical application of the AI model.
The model demonstrated robust performance, affirming its effectiveness in diverse
environments. Thorough testing included datasets representing wet, dry, and mixed
waste categories. Notably, the model achieved around 90% accuracy in classifying
isolated waste types, i.e., when the image contained just one type of waste. However,
challenges arose when processing mixed waste, leading to a decline in accuracy to
approximately 50%. Despite these limitations, the model showcased commendable
capabilities for efficient waste segregation.
Proposing an AI-Enabled Waste Segregation System for Domestic Settings 29

Fig. 2 Proposed system


model working

4 Results and Discussion

Demonstrating the practicality of the AI model, a test application was developed for a
mobile device using Android Studio. The app was embedded with the TFLite model
created by teachable machine to simulate the waste segregation process. Multiple
real-life garbage examples were tested and captured through the app, revealing
promising results in the model’s ability to categorize various waste instances, into
dry and wet waste. The following images have been taken as screenshots from the
mobile application which was developed to test the waste segregation model.
The implementation of the proposed system may encounter several challenges.
One potential issue is the adhesion of waste to the sorting area materials, particularly
when dealing with sticky waste. To address this, applying a polyurethane coating
[15] to the sorting area plastics can enhance imperviousness and maintain a sleek
appearance even with wet waste. Additionally, the system might face occasional
misclassifications, leading to inaccurate waste sorting. To mitigate this, continuous
model updates are recommended to adapt to evolving waste characteristics. Second,
30 D. G. Rao et al.

the model’s confidence metric, providing a percentage value for dry and wet cate-
gorizations, can be leveraged. Instances with confidence values below 80% for the
majority class can be flagged for retraining after the initial batch processing, ensuring
ongoing accuracy improvement.
The results were calculated based on a confidence percent that is provided on the
app (can been seen in the screenshots below as well). If the confidence values were
more than 80% and have been classified correctly, it was a success case.
When tested on real-world data (waste found in Indian households), the results
were promising. To be more precise, the model returned an accuracy of 89% with
around 50 images. Although this Fig. 3 might be relatively low, one of the major
strengths of this model lies in the low computational time taken by the model. This
can be attributed to the [Link] used to create the model (Fig. 4).

Fig. 3 Wet waste example


using the AI model
developed
Proposing an AI-Enabled Waste Segregation System for Domestic Settings 31

Fig. 4 Dry waste example


using the AI model
developed

5 Future Work

In the realm of future work, the foremost objective would be the physical realization
of the proposed waste segregation system. Given sufficient resources, constructing a
working model would be an essential step toward validating the theoretical aspects
discussed in this paper. This practical implementation would provide invaluable
insights into the system’s actual performance and its feasibility in a real-world
environment. To achieve this, a detailed plan outlining the necessary components,
budgetary requirements, and a timeline for assembly and testing should be developed.
Anticipated challenges, such as potential technical issues or delays in component
procurement, should be addressed in the plan to ensure a comprehensive approach.
32 D. G. Rao et al.

To enhance the accuracy of the AI model, future efforts should focus on expanding
the dataset for training. This involves incorporating a more extensive variety of
waste images, ensuring that the model becomes adept at recognizing diverse waste
materials, thereby refining its categorization capabilities. A comprehensive strategy
should be devised, detailing the types of waste materials to be included, the sources
for obtaining images, and the criteria for dataset expansion. Utilizing existing datasets
like TrashNet could further enrich the training data, providing a broader spectrum of
waste images for the model. Additionally, integrating reinforcement learning tech-
niques into the training process could enhance the model’s adaptability and respon-
siveness to various waste items, contributing to improved accuracy. This expansion
plan should be accompanied by a timeline, resource allocation considerations, and
potential challenges to ensure a systematic and well-executed approach.

References

1. Dhokhikah Y, Trihadiningrum Y (2012) Solid waste management in Asian developing


countries: challenges and opportunities. J Appl Environ Biol Sci 2:329–335
2. Kumar S, Smith SR, Fowler G, Velis C, Kumar J, Arya S, Rena RK, Cheeseman C (2017)
Challenges and opportunities associated with waste management in India. R Soc Open Sci
4160764160764. [Link]
3. Sudha S, Vidhyalakshmi M, Pavithra K, Sangeetha K, Swaathi V (2016) 2016 IEEE Tech-
nological Innovations in ICT for Agriculture and Rural Development (TIAR)—an automatic
classification method for environment: friendly waste segregation using deep learning, 15–16
July 2016, pp 65–70
4. Shreeshayana R, Gudur MV, Niranjan L, Sreekantha B (2022) Ergonomic automated dry and
wet waste segregation and compost production for innovative waste management. In: 2022
IEEE 3rd global conference for advancement in technology (GCAT), pp 1–6
5. Chowdhury SS, Hossain N, Saha T, Ferdous J, Zishan M (2021) The design and implementation
of an autonomous waste sorting machine using machine learning technique. AIUB J Sci Eng
(AJSE) 19:134–142. [Link]
6. Chu Y-C, Huang C, Xie X, Tan B, Kamal S, Xiong X (2018) Multilayer hybrid deep-learning
method for waste classification and recycling. Comput Intell Neurosci 2018, n. Pag
7. Huynh M-H, Pham T-L-G, Tran A-K, Nguyen T (2020) Automated waste sorting using convo-
lutional neural network. In: 2020 7th NAFOSTED conference on information and computer
science (NICS), pp 102–107
8. Sousa J, Rebelo A, Cardoso JS (2019) Automation of waste sorting with deep learning. In:
2019 XV workshop de Visão Computacional (WVC), pp 43–48
9. Liang S, Gu Y (2021) A deep convolutional neural network to simultaneously localize and
recognize waste types in images. Waste Manag 126:247–257
10. Bircanoglu C, Atay M, Beser F, Genc O, Kizrak MA (2018) RecycleNet: intelligent waste
sorting using deep neural networks. In: 2018 innovations in intelligent systems and applications
(INISTA), pp 1–7
11. Cao L, Xiang W (2020) Application of convolutional neural network based on transfer
learning for garbage classification. In: 2020 IEEE 5th information technology and mechatronics
engineering conference (ITOEC), pp 1032–1036
12. Toğaçar M, Ergen B, Cömert Z (2020) Waste classification using autoencoder network with
integrated feature selection method in convolutional neural network models. Measurement 153
Proposing an AI-Enabled Waste Segregation System for Domestic Settings 33

13. Srivatsan K, Dhiman S, Jain AK (2021) Waste classification using transfer learning with convo-
lutional neural networks. In: IOP conference series: earth and environmental science, vol 775,
n. Pag
14. O’Shea K, Nash R (2015) An introduction to convolutional neural networks. arXiv [Link]
org/abs/1511.08458 [[Link]]
15. Chen AT, Wojcik RT (2000) Polyurethane coatings for metal and plastic substrates. Met Finish
98(6):143–154
Application of Convolutional Neural
Networks’ Method in Early Age Disease
Detection of Paddy Crop

Asha Ambhaikar, Akanksha Mishra, Hussain Falih Mahdi,


Bhupesh Kumar Dewangan, Sanjana Dewangan, and Tanupriya Choudhury

Abstract Producing a high quality of crops is a major contribution of agriculture to


the economy of any country. Plant disease identification is one of the most impor-
tant aspects of maintaining a nation with a developed agricultural economy. The
effectiveness of Convolutional Neural Networks—specifically, the ResNet-50 archi-
tecture—in the context of paddy crop disease detection is examined in this study. The
study utilizes a comprehensive dataset comprising images of diseased and healthy
paddy plants for training and testing the ResNet-50 model. Through rigorous exper-
imentation, the CNN demonstrates remarkable accuracy, precision, and recall in
identifying various paddy crop diseases. By demonstrating the ResNet-50 CNN
model’s greater performance over conventional techniques, the study adds to the
body of current material. The detailed analysis underscores the capability of deep
learning techniques in revolutionizing detection of disease in agricultural settings,
providing a more reliable and efficient solution. While acknowledging the successes,

A. Ambhaikar
Department of Computer Science, MATS University Raipur (C.G.), Raipur, India
A. Mishra
Department of IT, SSIPMT Raipur (C.G.), Sejabahar, India
e-mail: akanksha.mishra1702@[Link]
H. F. Mahdi
Collage of Engineering, University of Diyala, Baqubah, Iraq
e-mail: [Link]@[Link]
B. K. Dewangan (B)
Department of Computer Science and Engineering, OP Jindal University, Raigarh, India
e-mail: [Link]@[Link]
Symbiosis Institute of Technology, Nagpur Campus, Symbiosis International (Deemed
University), Pune, India
S. Dewangan
School of Science, OP Jindal University, Raigarh, India
T. Choudhury (B)
University of Petroleum and Energy Studies, Dehradun, Uttarakhand, India
e-mail: tanupriya@[Link]

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 35
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
36 A. Ambhaikar et al.

the study also highlights certain challenges and limitations encountered during the
research process. This research’s future reach goes beyond the scholarly sphere. The
study envisions the creation of an application in recognition of the usefulness of the
CNN-based disease detection method. This application aims to empower farmers
and agricultural stakeholders by providing a user-friendly tool for swift and accurate
identification of crop diseases in the field.

Keywords CNN · ResNet-50 · ML · Deep learning

1 Introduction

A significant portion of the world’s population relies on paddy crops as a primary


food supply, making them of great global relevance. More than half of the world’s
population uses rice as a nutritional supplement, particularly in Asia, where it is a
staple food. The cultivation of paddy crops is not only important for food security but
also plays a vital role in the economic livelihoods of millions of farmers worldwide
[1]. About 65% of the population consumes rice, which is the most Significant staple
food in the nation. It is grown in nearly every state. Paddy crops are susceptible
to a variety of diseases, such as bacteria, fungus, and viruses, which can lower
output and cause financial losses [2]. The impact of these illnesses is compounded by
variables such as changing climate conditions, growing global trade, and developing
agricultural techniques. The sustainability of paddy agriculture and the security of the
world’s food supply depend on efficient disease management techniques. Therefore,
research and technological advancements, such as the application of Convolutional
Neural Networks (CNNs) in disease detection, become paramount in addressing
these challenges and sustaining the productivity of paddy crops [3].

2 Objective

The objective of this paper is to showcasing the potential of Convolutional Neural


Networks in the early age disease detection of paddy crops using ResNet-50 model.
The proposed work aims to validate the efficacy of CNNs through detailed analysis,
optimization, and performance evaluation. By proposing a specialized CNN model,
the research contributes to precision agriculture, offering a reliable solution for timely
disease identification in paddy crops. The ultimate goal is to provide farmers with an
effective tool that can minimize crop losses and optimize resource usage, promoting
sustainable and efficient agricultural practices.
Application of Convolutional Neural Networks’ Method in Early Age … 37

3 Literature Review

Convolutional Neural Networks, or CNNs, are a key development in agricultural tech-


nology when it comes to the finding of disease in paddy crops. It is crucial to manage
the wellbeing and growing of staple crops like rice as the world’s food demand rises.
This section summarizes the important studies, techniques, and findings from the
body of research on CNN application in paddy crop disease detection [5]. Tradi-
tional methods for disease detection in paddy crops often involve manual inspection,
which is time taking and possibility of human error. Researchers have increasingly
explored CNNs as a more efficient and accurate alternative [6]. Demonstrated the
superiority of CNNs over traditional methods in identifying various diseases in paddy
crops, achieving significantly higher accuracy rates.
Various CNN architectures have been employed for disease finding in paddy crops,
Bacterial, viral, or fungal infections in crops cause the agriculture sector to suffer
enormous financial losses; farmers lose 15 to 20 percent of their annual earnings as a
result. With ResNet-50 emerging as a popular choice. The deep learning capabilities
of ResNet-50 have shown promise in accurately classifying and diagnosing diseases
in paddy plants. Researchers like [7] demonstrated the effectiveness of ResNet-50
in achieving high accuracy rates in identifying diseases such as blast and sheath
blight in rice crops. Transfer learning, leveraging pre-trained CNN models, has been
explored to address the challenge of limited annotated datasets [8]. Demonstrated
the successful application of transfer learning, specifically using a pre-trained CNN
model, in detecting diseases in paddy crops with limited labeled data. To evaluate
their performance, training from scratch and transfer learning have been used. The
optimal performance in both designs has been found when the model is fine-tuned
during training.
While CNNs have shown significant promise, challenges such as large require-
ment of food datasets and model Interpretability endures. Future research, as
suggested by [9], should focus on addressing these challenges and further refining
CNN models for real-world deployment in diverse agricultural settings. The authors
have emphasized the important problems and difficulties with classifying leaf
diseases. Included is a comparison of several approaches based on the agricultural
product, methodology, effectiveness, and pros and cons. Here are some common
paddy crop diseases [10].
Blast Disease Symptoms: Characterized by elliptical lesions on leaves, stems, pani-
cles, and grains. Lesions can begin as tiny, wet patches and spread very quickly
[11].
Brown Spot Symptoms: Small, brown, oval lesions with a yellow halo on leaves. The
disease can lead to premature leaf senescence and reduced photosynthetic activity
[12].
Sheath Blight Symptoms: Water-soaked lesions on leaf sheaths, which can expand
quickly, leading to wilting and lodging. Infected plants may show a characteristic
“fishhook” symptom [13].
38 A. Ambhaikar et al.

Fig. 1 ResNet50 architecture

Tungro Disease Symptoms: Stunted growth, yellowing, and reddening of leaves.


The disease is transmitted by the green leafhopper [14].
Bacterial Leaf Blight Symptoms: Water-soaked lesions on leaves that later turn
brown and lead to the blighting of the entire leaf. Bacteria are often spread through
contaminated water [15].
False Smut Symptoms: Formation of smut balls on the panicles, giving them a false
appearance. The smut balls contain masses of dark green to black spores [16].
Early detection and prompt action are essential to minimize yield losses and ensure
food security. These are the different diseases that the rice plant could be afflicted
with. These disorders can be successfully classified by our approach. The rice plants
would be categorized as healthy if they were free from any diseases (Fig. 1).
A collection of complicated algorithms known as neural networks uses a method
similar to that of the human brain to identify the main relationship within a batch of
data. Neural networks find extensive application in company planning, trade, business
analytics, and medical applications in addition to product maintenance [17]. Rich
feature extraction is made possible by these designs, which can be applied to advanced
tasks such as object detection, image segmentation, and image classification [18].
Out of all these architectures, we have used the ResNet-50 model, a series of
deep Convolutional Neural Networks known as ResNet which is “Residual Neural
Network” created to address the issue of disappearing gradients, which is typical of
very deep networks. The training of very deep networks is made possible by ResNet’s
usage of “residual blocks,” which permit gradients to propagate directly throughout
the network [19–24]. A residual block is made up of two or more convolutional layers,
an activation function, and a shortcut link that adds the original input straight to the
convolutional layers’ output after the activation function, skipping the convolutional
layers altogether [25–28].
The image above illustrates how different versions of the ResNet architecture use
varied numbers of Cfg blocks at different levels.
Application of Convolutional Neural Networks’ Method in Early Age … 39

Fig. 2 Process flow model

4 Methodology

(1) Augmenting Images

A large training dataset that provides the network with a solid learning experience
is necessary for the neural network to perform well. By practically increasing the
amount of training data, image augmentation techniques improve the performance
of the neural network classifier.(2)
(2) Convolution Step
A number of convolutional layers process the input data. Extracting features is the
responsibility of these convolutional layers [20]. They pick up various input data
hierarchical aspects.
(3) Pooling Step
When the down sampling technique is used in Convolutional Neural Networks, it
is regarded as a crucial phase. Reducing the image’s spatial size is the primary
goal of pooling. Max, Min, and Average pooling are the three types of pooling that
Convolutional Neural Networks employ [23, 24].
(4) Residual Connection Phase
The flatten() function transforms a matrix into a long vector, and the dense()
layer handles linear computations. Softmax and Relu() are employed as activation
functions (Fig. 2).

5 Comparison Between Traditional Methods

Given our use of ResNet-50, a notable deviation from other available models is the
adjustment made to address concerns about the prolonged training time for layers.
We have used a variety of Convolutional Neural Network designs to evaluate this
40 A. Ambhaikar et al.

Table 1 Comparison between models


Model name CNN With one CNN With two KNN-SVM Transfer learning ResNet-50
layer layer
Accuracy 73% 76% 69% 89% %

model. This led to the transformation of the building block into a design featuring
bottlenecks. We have employed the 16-layer VGG-16, two-layer CNN, and one-layer
CNN. Unlike the previous two-tier structure, it now employs a stack of three tiers.
Consequently, each two-layer block in ResNet-34 was substituted with a three-layer
bottleneck block to create the ResNet-50 design, which exhibits significantly greater
accuracy than the 34-layer ResNet model [21, 22]. The performance of the 50-layer
ResNet is measured at 3.8 billion FLOPS. Our model has achieved the accuracy of
98.7% (Table 1).

6 Result

After testing the model, the algorithm effectively predicted what illnesses in the
paddy crop after using the testing data. We have utilized the ResNet-50 which attains
a 90% accuracy rate.

7 Conclusion

This research has delved into the realm of disease detection in paddy crops, lever-
aging the formidable capabilities of Convolutional Neural Networks, specifically
employing the ResNet-50 architecture. The use of deep learning techniques, exem-
plified by the ResNet-50 model, has proven to be a potent tool in accurately and
efficiently identifying various diseases afflicting paddy plants. The results obtained
in this study are showing that this model can predict 98.7% accurate disease in
paddy crops. The superior performance of the ResNet-50 CNN model compared
to traditional methods, affirming its potential as a transformative technology in the
agricultural sector. While the ResNet-50 CNN model excels in many aspects, there
exist nuances and complexities in the real-world agricultural setting that demand
continued exploration and refinement. Addressing these challenges will be essen-
tial in ensuring the practical applicability and scalability of the proposed solution.
Looking forward, the future scope of this research extends beyond the confines of
academia. The envisaged application, fueled by the insights gained from this study,
seeks to bridge the gap between cutting-edge research and on-the-ground agricul-
tural practices. The development of a user-friendly application holds the promise of
Application of Convolutional Neural Networks’ Method in Early Age … 41

empowering farmers and agricultural stakeholders with a tool that facilitates swift
and accurate disease identification in paddy crops.

References

1. Singh KM, Ahmad N, Pandey VV, Kumari T, Singh R (2021) Growth performance and prof-
itability of rice production in India: An assertive analysis. Published in: Economic Affairs, vol
66, no 3, pp 481–486. [Link]
2. Singh KM, Singh P (2020) Challenges of ensuring food and nutritional security in Bihar. Dev
Econ: Macroecon Issues Dev Econ E J 9(23). [Link]
3. Simón Sánchez A-M, González-Piqueras J, de la Ossa L, Calera A (2022) Convolutional neural
networks for agricultural land use classification from Sentinel-2 image time series. Remote Sens
14:5373. [Link]
4. Islam MM, Adil MAA, Talukder MA, Ahamed MKU, Uddin MA, Hasan MK, Sharmin S,
Rahman MM, Debnath SK (2023) DeepCrop: Deep learning-based crop disease prediction
with web application. J Agric Food Res 14:100764. [Link]
ISSN 2666-1543
5. Kamilaris A, Prenafeta-Boldú FX (2018) A review of the use of convolutional neural networks
in agriculture. J Agric Sci 156(3):312–322. [Link]
6. Manavalan R (2020) Automatic identification of diseases in grains crops through computational
approaches: a review. Comput Electron Agric 178:105802. [Link]
2020.105802. ISSN 0168-1699
7. Chakraborty S, Layek R, Sankar S, Saha A, Ghosh Ray H (2021) Early detection of disease in
rice paddy: a deep learning based convolution neural networks approach. In: 2021 12th inter-
national conference on computing communication and networking technologies (ICCCNT),
Kharagpur, India, pp 1–5. [Link]
8. Rahman CR, Arko PS, Ali ME, Khan MAI, Apon SH, Nowrin F, Wasif A (2020) Identification
and recognition of rice diseases and pests using convolutional neural networks. Biosyst Eng
194:12–120. [Link] ISSN 1537–5110
9. Goel L, Nagpal J (2023) A systematic review of recent machine learning techniques for plant
disease identification and classification. IETE Technical Review 40(3). [Link]
02564602.2022.2121772
10. Rautaray SS, Pandey M, Gourisaria MK, Sharma R, Das S (2020) Paddy crop disease predic-
tion–a transfer learning technique. Int J Recent Technol Eng (IJRTE) 8(6). [Link]
35940/ijrte.F7782.038620. ISSN: 2277–3878
11. Shahriar SA, Imtiaz AA, Hossain MB, Husna A, Eaty MNK (2020) Rice blast disease. Annu
Res and Rev Biol 35(1):50–64. [Link]
12. MauYS, Ndiwa A, Oematan S (2020) Brown spot disease severity, yield and yield loss rela-
tionships in pigmented upland rice cultivars from East Nusa Tenggara, Indonesia. Biodiversitas
J Biol Divers 21(4). [Link]
13. Molla KA, Karmakar S, Molla J, Bajaj P, Varshney RK, Datta SK, Datta K (2020) Under-
standing sheath blight resistance in rice: the road behind and the road ahead. Plant Biotechnol
J 18(4):895–915. [Link]
14. Wahyuni WS, Sayekti JBA (2020) The failure farmers in panti district to control tungro disease
which endemic in 2014–2019. J Phys: Conf Ser 1563(1):012026. [Link]
6596/1563/1/012026
15. Ahmed T, Shahid M, Noman M, Niazi MBK, Mahmood F, Manzoor I, Zhang Y et al (2020)
Silver nanoparticles synthesized by using Bacillus cereus SZT1 ameliorated the damage of
bacterial leaf blight pathogen in rice. Pathogens 9(3):160. [Link]
9030160
42 A. Ambhaikar et al.

16. Dangi B, Khanal S, Shah S (2020) A review on rice false smut, it’s distribution, identification
and management practices. Acta Sci Agric 4:48–54. [Link]
346099112
17. Jiang J, Chen M, Fan JA (2021) Deep neural networks for the evaluation and design of photonic
devices. Nat Rev Mater 6(8):679–700
18. Ketkar N, Moolayil J, Ketkar N, Moolayil J (2021) Convolutional neural networks. Deep
learning with python: learn best practices of deep learning models with PyTorch 197–242
19. Koonce B (2021) ResNet 50 Convolutional neural networks with swift for tensorflow: image
recognition and dataset Categorization 63–72. [Link]
20. Wijayanto AK, Prasetyo LB, Hudjimartsu SA, Sigit G, Hongo C (2024) Textural features
for BLB disease damage assessment in paddy fields using drone data and machine learning:
enhancing disease detection accuracy. Smart Agric Technol 8:100498
21. Surenther I, Sridhar KP, Kingston Roberts M (2023) Maximizing energy efficiency in wireless
sensor networks for data transmission: a deep learning-based grouping model approach. Alex
Eng J 83:53–65. [Link]
22. Singh A, Kaur J, Singh K, Singh ML (2024) Deep transfer learning-based automated detection
of blast disease in paddy crop. SIViP 18(1):569–577
23. Verma HD, Gourisaria, MK, Ghosh S, Dewangan BK (2023) Comparative analysis of CNN
models for retinal disease detection. In: 2023 international conference on network, multimedia
and information technology (NMITCON). IEEE, pp 1–6
24. Verma HD, Gourisaria MK, Ghosh S, Dewangan BK (2023) A futuristic approach to synaptic
fusion of INN and CNN architectures for tissue classification. In: 2023 international conference
on network, multimedia and information technology (NMITCON). IEEE, pp 1–6
25. Kalbhor A, Nair RS, Phansalkar S, Sonkamble R, Sharma A, Mohan H, Lim WH (2024)
PARKTag: An AI–Blockchain integrated solution for an efficient, trusted, and scalable parking
management system. Technologies 12(9):155
26. Raghuvanshi A, Sharma A, Awasthi AK, Singhal R, Sharma A, Tiang SS, Lim WH (2024)
Linear antenna array pattern synthesis using multi-verse optimization algorithm. Electronics
13(17):3356
27. Ang KM, Lim WH, Tiang SS, Sharma A, Towfek SK, Abdelhamid AA, Khafaga DS (2023)
MTLBORKS-CNN: an innovative approach for automated convolutional neural network
design for image classification. Mathematics 11(19):4115
28. Mohan H, Agrawal G, Jately V, Sharma A, Azzopardi B (2023) Neural network-driven
Sensorless speed control of EV drive using PMSM. Mathematics 11(19):4029
29. Pawar D, Phansalkar S, Sharma A, Sahu GK, Ang CK, Lim WH (2023) Survey on the
biomedical text summarization techniques with an emphasis on databases, techniques, semantic
approaches, classification techniques, and similarity measures. Sustainability 15(5):4216
Information Encryption Technique
Based on DNA Cryptography and RSA
Asymmetric Key Exchange Algorithm

Aman Khan, Kirti Nahak, Asha Ambhaikar, Hussain Falih Mahdi,


Bhupesh Kumar Dewangan, and Tanupriya Choudhury

Abstract This paper proposes a new technique to encrypt data using DNA sequence.
DNA Cryptography is a technique which is emerging almost every field of network
security. The method proposed in this paper is an alternative approach of encryption
using DNA sequence which is quite efficient in terms of both computational speed
and security. Here along with the DNA encryption, RSA key exchange algorithm
has been applied in order to exchange the keys between two dedicated parties to
encrypt and decrypt the message. Here in the method, the original message is first
converted to ASCII equivalent and then it is converted to DNA sequence using the
table which represents a unique character for each binary combination. After that,
a subsequence of the DNA sequence is selected using the key (contains interval to
select the subsequence) and converts to its equivalent mRNA and then to tRNA. The
remaining elements at the left and the right of the subsequence are swapped, and
a new encrypted cipher is obtained which is transmitted over the network and then
decrypted in the same way at the receiver’s end, and finally, the original text is again
generated. This is an efficient way to encrypt the data as only a single key pair is
exchanged to perform the encryption and decryption and no any highly computational

A. Khan · K. Nahak
Department of CS and IT, Kalinga University, Raipur, India
A. Ambhaikar
Department of CS and IT, MATS University, Raipur, Chhattisgarh, India
H. F. Mahdi
Collage of Engineering, University of Diyala, Baqubah, Iraq
e-mail: [Link]@[Link]
B. K. Dewangan (B)
Department of Computer Science and Engineering, OP Jindal University, Raigarh, India
e-mail: [Link]@[Link]
Symbiosis Institute of Technology, Nagpur Campus, Symbiosis International (Deemed
University), Pune, India
T. Choudhury (B)
University of Petroleum and Energy Studies, Dehradun, Uttarakhand, India
e-mail: tanupriya1986@[Link]

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 43
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
44 A. Khan et al.

operations are performed to maintain a good level of secrecy. Finally, the security
analysis is performed which includes frequency analysis, Friedman randomness test,
and RSA security analysis to present the security level of proposed method.

Keywords DNA Cryptography · mRNA · tRNA · Encoding region · Cipher ·


Friedman test · RSA asymmetric cryptography · Transcription · Translation ·
Encryption · Decryption · ASCII · Nucleotide

1 Introduction

“Security is assessment of risk. Secure environments do not just appear; they are
designed and developed through an intentional effort” [1].
DNA Cryptography is a branch of cryptography which deals with the utilization of
DNA for computation and security of information [2]. DNA Cryptography technique
provides multifold security [3] which makes the information secrecy quite tough to
break. Many of the existing DNA techniques use symmetric cryptosystem which
lowers the effectiveness of the technique as the encryption and decryption key is same.
Here we have used asymmetric cryptosystem, for key exchange, which uses different
keys for both encryption and decryption. Hence using DNA Cryptography along with
advanced cryptosystem ensures the higher level of secrecy than the traditional one.
In this paper, an alternative approach of DNA cryptography has been used which
proposes multifold security and enrich the level of security. The paper divides the
process of encryption into three stages, i.e., the creation of DNA sequence, key
exchange using asymmetric algorithm, and formation of cipher text using secret key.
At first, the plain text message is converted to its equivalent ASCII code and then
each ASCII value is converted to its equivalent 8-bit binary digit. After that, the
binary digits are concatenated and using the DNA table (shown in Table 1), equiv-
alent DNA sequence is created. An encoding region then selected from the DNA
sequence and using the key, the region is transcripted to mRNA and then to tRNA
and the remaining sequences will be swapped with each other. The obtained shuffled
sequence is the encrypted cipher text. In this manner, an effective and lightweight
encryption technique is proposed. The remaining parts of the paper are arranged as
follows: Sect. 2 describes Literature Review, Sect. 3 describes DNA and its compo-
nents, Sect. 4 includes Proposed Methodology, Sect. 5 presents Security Analysis,
and finally, Sect. 6 concludes the paper.

2 Literature Review

In this section, different cryptography mechanisms based on DNA technology will


be explained.
Information Encryption Technique Based on DNA Cryptography … 45

Table 1 Nucleotide and


Nucleotide (DNA base) Binary equivalent
binary equivalent
AC 00
GT 01
10
11

[Link] et al. [4] have proposed an encryption technique using genetic algorithm
and used Diffie–Hellman algorithm for exchanging the keys. O.A. Al-Harabi [5]
and [Link] [6] proposed DNA-based cryptography analysis using steganog-
raphy techniques. Sadeg et al. [7] proposed a technique based on the methods of
transcription (DNA to Mrna) and translation (mRNA to protein).
The technique proposed by [Link] and [Link] [8] depends on
genome sequence from NCBI bank for key generation and develop Vigenere cipher
to enhance the level of secrecy.
Mohammad Reza Abbasy et al. [9] represented an encryption method by manip-
ulation four letters (called as nucleotides). Any composition from them will make a
different sequence. Rauche et al. [10] represented DNA sequence in binary form. All
the sequences are represented in the form of 0 or 1 and the disciplinary operations
of encryption performed on it.
R.P.K Reddy et al. [11] described an encryption scheme with three levels of
security where the first level is for the selection of key of any length. Second level
depict rules for the arrangement of A, T, G, C. In the third level, the sequence is
replaced by the look up table value of fixed length 12. The technique proposed
by [Link] et al. [12] includes genetic algorithm along with Needleman–Wunsch
algorithm to generate the key. In the paper, proposed by M. Najaftorkaman et al.
[13], the security analysis included Friedman test and letter frequency analysis.
Analyzing the above-mentioned technique, it is observed that many techniques
have used symmetric cryptosystem like DES, etc., and various techniques have used
asymmetric cryptosystem like RSA, Diffie–Hellman along with several operations
on DNA. This method is an alternative approach which is quite efficient and easy to
implement.

3 DNA and Its Components

DNA is the basis of life which is made up of building blocks called as nucleotides
[14]. These nucleotide contains three parts—phosphate, sugar, and nitrogen bases.
The nitrogen bases are of four different types, viz.
• Adenine (A).
• Thymine (T).
• Guanine (G).
46 A. Khan et al.

Fig. 1 Structure of DNA

• Cytosine (C).

The DNA sequence contains information form gene, and their size varies from
1000 bases to 2300 kilo bases in humans [14].
The DNA has a double helix, ladder like, structure which consists of pair of
nucleotides in between the helix. As the DNA has a specific nature, the nucleotide
base A always paired with base T and the base C always paired with G. These bases
are represented in the form of sequences which are quite long in Fig. 1.
Transcription and Translation: As the DNA never leave its nucleus and hence RNA
is formed to transfer the information, for the synthesis of protein, to the ribosome
in cytoplasm. This process is done by detachment of DNA to form RNA, and this
is termed as transcription. The RNA is called messenger RNA (Mrna) as it carries
DNA’s message for the synthesis of protein. So in RNA, there is no Thymine (T).
Instead, the Adenine (A) is paired with a new base called Uracil (U). Hence, RNA
has A, U, G, and C as its nucleotides.
Translation is a process where Mrna is synthesized to a protein. The nitrogen
base triplets contained by Mrna codes for particular amino acid, for that purpose, the
mRNA is converted to transfer RNA (tRNA) which is anticodon and is formed by
replacing A with U, G with C and Vice-Versa is presented in Table 2.

Table 2 Conversion of DNA


DNA mRNA tRNA
to mRNA to tRNA
A A U
C C G
T U A
T U A
G G C
A A U
G G C
Information Encryption Technique Based on DNA Cryptography … 47

Table 3 Encryption process


Step Operations
1. Plain HELLO
Text
2. ASCII 72 69 76 76 97
Value
3. Binary 0100100001000101010011000100
Value 110001001111
4. 2-bit 01 00 10 00 01 00 01 01 01 00 11
Binary 00 01 00 11 00 01 00 11 11
5. DNA CAGACACCCATACATACATT
Base
6. Encoding Let Key pair is (6,16)
Region CAGACACCCATACATACATT

7. mRNA CAGACA CCCAUACAUA CATT

8. tRNA CAGACA GGGUAUGUAU CATT

9. Swap CAGACAGGGUAUGUAUCATT

10. Swap CATTCAGGGUAUGUAUCAGA

11. Cipher CATTCAGGGUAUGUAUCAGA

Table 4 Decryption process


Step Operations
1. Cipher CATTCAGGGUAUGUAUCAGA
2. EncodingRe- Key pair received- (6,16)
gion CATTCAGGGUAUGUAU CAGA
3. mRNA CATTCACCCAUACAUACAGA
4. DNA CATTCACCCATACATA CAGA
5. Swap CATTCACCCATACATACAGA
6. Swap CAGACACCCATACATACATT
7. Sequence CAGACACCCATACATACATT
8. Binary 0100100001000101010011000100
110001001111
9. 8-bit 01001000 01000101 01001100
Binary 01001100 01001111
10. ASCII 72 69 76 76 97
11. Plain text HELLO
48 A. Khan et al.

4 Proposed Methodology

The proposed cryptography scheme in this paper has two phases—the encryption
phase and the decryption phase. In the encryption phase, the plain text message
passes through various stages of encryption and converted to a cipher text, whereas
in decryption phase, the cipher text obtained is again passed through various stages
and the desired plain text is obtained. In both the phases, the key exchanged using
RSA asymmetric cryptography algorithm plays an important role by exchanging the
key pair which is used for encryption and decryption of the information mentioned in
Figs. 2 and 3 and Table 3 and Table 4. Following are the steps involved in encryption
process.
Encryption Algorithm

1. Consider the message in the plain text format.


2. Convert each and every character of the text into its equivalent ASCII value.

PLAI N TEXT

ASCII VALUE

8- BIT BINARY DIGITS

CONCATENATE 2-2 BITS AND COVERT TO DNA BASES

SELECT THE ENCODING REGION USING KEY AND TRANSCRIPT


AND THEN CONVERT TO tRNA

SWAP THE REMAINING PORTION OF SEQUENCE

CIPHER TEXT

Fig. 2 Flowchart of the encryption process


Information Encryption Technique Based on DNA Cryptography … 49

CIPHER

DETERMINE ENCODING REGION USING

CONVERT THE ENCODED REGION TO ORIGINAL DNA

SWAP AGAIN THE LETTERS TO RESHIFT THEM AT THEIR ACTUAL

CONVERT THE OBTAINED SEQUENCE TO

SPLIT THE BINARY TO 8 BIT CHUNKS AND FIND THEIR ASCII

CONVERT THE ASCII VALUR TO ITS EQUIVALENT

ORIGINAL PLAIN

Fig. 3 Flowchart of the decryption process

3. Convert the ASCII values into its equivalent 8-bit Binary digits.
4. Separate the Binary digits in 2–2 bit pairs.
5. Convert each binary pair into its equivalent DNA base (A, T, G, C) using the as
per the values mentioned in Table1.
6. Concatenate the nucleotide bases to form a DNA sequence and using the key-pair
(exchanged through RSA) select the encoding region in the DNA sequence and
transcript the encoding region and then convert it to tRNA.
7. Excluding the encoding region, swap the remaining sequence falls to the left and
right side of the encoding region with the number of swaps equal to the minimum
number of bases (A, T, G, C) on either of the side.
8. The obtained sequence is the encrypted cipher text and is ready to be sent.
50 A. Khan et al.

Decryption Algorithm

1. Get the received cipher text as an encrypted sequence.


2. Using the key-pair, find the encoding region.
3. Convert the encoding region to actual DNA sequence by reverse transcription.
( tRNA | mRNA | DNA)
4. Swap the remaining sequence as done in step 7 of Encryption phase.
5. Convert the obtained sequence to Binary using the values from Table 1.
6. Split the binary digits into 8-bit binary chunks and convert into its equivalent
ASCII value.
7. Convert the ASCII value into its equivalent characters.
8. The Result obtained is the original decrypted message.

5 Security and Statistical Analysis

In today’s world, the technology is increasing rapidly and the field of nanotechnology
and Bioinformatics bringing a new trend in the area of technology. Though we have
a lot of data to store and the DNA technology and Quantum technology makes it
efficient to store the data and manipulate it in a fast manner, these technologies are
improving day by day and providing various flexibilities. [15]
As it is an emerging technology and improving positively, there are some threats
which the technology possesses and which is to be dealt while working with it.
As the paper presents a cryptosystem which is depending on Vigenere cipher,
a polyalphabetic cipher [7], there are some letters of alphabet which are arranged
in any specific manner in order to encrypt the data and maintain the secrecy level.
The main problem with this type of cryptosystem is that the letters can be easily
guessed by analyzing the frequency of the letters, i.e., how many times a letter is
being mentioned in the cipher.
The proposed methodology of DNA cryptography, in this paper, is strongly non-
vulnerable against the frequency analysis as the cipher contains four letters, viz. A,
C, G, and T, and this has nothing to be relatable with their frequencies. For example,
if we take simple English sentences, then vowels are very common and frequent
letters, and by analyzing them, we can possibly crack the Vigenere cipher. But in
DNA Cryptosystem, we cannot relate the frequency with the cipher in any manner
[16].
Let us suppose that a message is encrypted and the cipher obtained is:

ATCGTACGATGCAAGCTCCAATGCATGC
TAGCTAGACATACAGACTAACGTAAGATCGAGGCTTA
Information Encryption Technique Based on DNA Cryptography … 51

Fig. 4 Letter frequency


25
20
15
10
5

Here we can clearly observe that the frequency of A is more than that of any of
the three letters, but we cannot relate it with the actual information hidden behind
this cipher. The analysis is presented in Fig. 4.
Another vulnerability factor with this cipher is prediction of the key pair. If the key
pair is easy to predict, then the cipher can be easily decrypted by any unauthorized
party. There must be an assurity that the key pair which is being selected is not easily
predictable [17].
To test this, the method which will be used is Friedman test [8]. Here in Friedman
test, a calculation is done to find the key pair. If it is easily predictable, then the
proposed cryptosystem is threat to attack and is not secure.
The formula of finding the key pair is:

Keylength = KP − Kr
K0 − kr

Here, Kp = Probability of choosing correct cipher block. This value is to be


determined [8].
Kr = Probability of randomly selecting a letter from DNA bases.
As there are 4 letters only (A, T, G, C), this value will be 1.

K0 = Coincidence Rate


c (fi × fi − 1

Coincidence Rate = i = 1 / (N − 1) i = 1 / (N − 1).


Here,
‘c’ is the size of alphabet or letters which is being used in cryptography.
‘fi’ represents the frequency of a letter.
‘N’ shows the length of Cipher text.
Let us see an example which explains the working of Friedman test. Suppose the
cipher text is:

CACCAACTCAGCCATCGCTCACAACTAC
CGTCTGACCACCTCGCACTTCGGGTGAG
52 A. Khan et al.

A(13 × 12) + C(23 × 22) + G(10 × 9) + T(10 × 9)


= 156 + 506 + 90
+ 90 = 842.

N(N − 1) = (56 × 55) = 3080.

IC = 842/3080 = 0.27(approx).

So, key length = (kp − 0.25)/0.27 − 0.25.


As we can see that the value of K0 is very near to 0.25, and hence, we can assume
the denominator as 0. As a result of which the key length will be:

Key length = Kp − 0.25 ≈ ∞.

Hence, we can say that finding the key pair (or encoding region) is not possible
and the proposed DNA cryptography methodology is able to maintain the secrecy of
data.

6 Conclusion

The overall objective of this paper is to introduce a new methodology of DNA


Cryptography to maintain the secrecy of information and to provide security to
the data. The technique used in this paper is much easier and efficient to perform
encryption and decryption of the information using some DNA mechanism like
transcription and translation using the standard RSA algorithms for key exchange.
The paper also invigilated several works that has been done in this area and analyzed
their highs and lows. Furthermore on the basis of the security analysis done using
some methods, we can conclude that the proposed technique suggested an alternative
Cryptography technique with respect to the security requirements. For the future
work, in the area of DNA Cryptography, the suggested technique will be providing
layered security and dynamic mechanism to the cryptography system to enhance the
security level.

References

1. Krawetz N (2007) Introduction to Network Security, Boston


2. Anam B, Sakib K, Hossain M, Dahal K (2010) Review on the advancements of DNA
cryptography. arXiv:1010.0186
3. Biswas MR, Alam KMR, Tamura S, Morimoto Y (2019) A technique for DNA cryptography
based on dynamic mechanisms. J Inf Secur Appl 48:102363
Information Encryption Technique Based on DNA Cryptography … 53

4. Vidhya E, Rathipriya R (2020) Key generation for dna cryptography using genetic operators
and diffie-hellman key exchangealgorithm. Computer Science 15(4):1109–1115
5. Al-Harbi OA, Alahmadi WE, Aljahdali AO (2020) Security analysis of DNA based steganog-
raphy techniques. SN Applied Sciences 2(2):1–10
6. Alexander G (2017) DNA Based cryptography and steganography. Glob Re-Search Dev J Eng
2:249–253
7. Sadeg S, Gougache M, Mansouri N, Drias H (2010) An encryption algorithm inspired from
DNA. In: 2010 international conference on machine and web intelligence. IEEE, pp 344–349
8. Najaftorkaman M, Kazazi NS (2015) A method to encrypt information with DNA- based
cryptography. Int J Cyber-Secur Digit Forensics 417–426
9. Abbasy MR, Nikfard P, Ordi A, Torkaman MRN (2012) DNA base data hiding algorithm. Int
J New Comput Arch Their Appl (IJNCAA) 2(1):183–192
10. Leieryz A, Richteryz C, Banzhafy W, Rauhey H (1999) Cryptography with DNA binary strands
11. Reddy RPK, Nagaraju C, Subramanyam N (2014) Text encryption through level based privacy
using DNA steganography. Int J Emerg Trends and Technol Comput Sci (IJETTCS) 3(3):168–
172
12. Kalsi S, Kaur H, Chang V (2018) DNA cryptography and deep learning using genetic algorithm
with NW algorithm for Key Generation. J medical syst 42(1)
13. Najaftorkaman M, Kazazi N (2015) A method to encrypt information with DNA- based
cryptography. Int J Cyber-Secur Digit Forensics (IJCSDF) 4(3):417–426
14. Ravi et al (eds) (2004) Advances in biotechnology. [Link]
4-7_2. © Springer India
15. Clelland CT, Risca V, Bancroft C (1999) Hiding messages in DNA microdots. Nature
399(6736):533–534
16. Farhat S, Kumar M, Vaish A, Dewangan BK, Choudhury T, Kotecha K (2023) A lightweight
encryption method for preserving E-healthcare data privacy using dual signature on twisted
edwards curves. In: international conference on computer & communication technologies.
Singapore: Springer Nature Singapore, pp 69–82
17. Khurana M, Singh BK, Choudhury T, Dewangan BK (2021) Internet of Things (IoT) and secu-
rity: challenges ahead. In data driven approach towards disruptive technologies: proceedings
of MIDAS 2020. Springer, pp 243–255
Sentiment Analysis Using Bi-Directional
LSTM for Depression Detection
and Suicide Prevention

Sunny Singh, J. Hridhya, Hussain Falih Mahdi, Bhupesh Kumar Dewangan,


and Tanupriya Choudhary

Abstract In the modern world, depression has become a prominent element in


suicide instances. Many people experience depression as a result of personal or
professional difficulties, and they frequently share their feelings and opinions on
social networking sites like Twitter. The material provided determines whether these
feelings are favorable or negative, with negative sentiments perhaps suggesting a
risk of suicide. To prevent suicide, it is essential to identify and treat depression
in people. This work focuses on utilizing a Twitter depression dataset to assess
depression using several machine learning models, such as CNN, Bi-directional
LSTM, Uni-directional LSTM-RNN, and Vanilla RNN. The effectiveness of these
models is evaluated both quantitatively and qualitatively in the study, accounting
for variables such as the F1-score, accuracy, precision, recall, and AROC curve. In
comparison to the other models, the results show that the Bi-directional LSTM model
analyzes depression with greater accuracy, suggesting promising possibilities for
early identification and intervention in people at risk of depression-related problems
and suicide.

S. Singh · J. Hridhya
Computer Science and Engineering, St. Peter’s Engineering College, Hyderabad, India
e-mail: sunnysingh@[Link]
J. Hridhya
e-mail: hridhya@[Link]
H. F. Mahdi
Collage of Engineering, University of Diyala, Baqubah, Iraq
e-mail: [Link]@[Link]
B. K. Dewangan (B)
Department of Computer Science and Engineering, OP Jindal University, Raigarh, India
e-mail: [Link]@[Link]
Symbiosis Institute of Technology, Nagpur Campus, Symbiosis International (Deemed
University), Pune, India
T. Choudhary (B)
University of Petroleum and Energy Studies, Dehradun, Uttarakhand, India
e-mail: tanupriya@[Link]

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 55
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
56 S. Singh et al.

Keywords Depression · Suicide · Social media · Sentiment analysis · Twitter ·


Machine learning · Bi-directional LSTM · Uni-directional LSTM-RNN · Vanilla
RNN

1 Introduction

Social media is one of the most widely used channels for interpersonal communica-
tion. Perhaps the most successful social networks are Facebook, Instagram, Twitter,
and others [1]. On social media, individuals communicate with one another and
express their opinions and feelings. The collective sentiments are vital cause of infor-
mation that may be utilized to determine how the general public feels about certain
issues. Large corporations, political parties, businesses, and non-governmental orga-
nizations (NGOs) use data gathered from social media posts for sentiment analysis.
In order to analyze the postings on social media, sentiment analysis is crucial. It is
employed to automatically classify postings according to data properties.
With an incredible average of 58 million tweets sent each day in the modern
digital age, Twitter stands out as one of the most popular and extensively utilized
venues for exchanging opinions [2]. Tweets are short statements of feeling since they
are confined to a maximum of 140 characters, which makes them easy to understand
and analyze. Additionally, Twitter users participate in conversations by commenting,
retweeting, and reacting to tweets. This encourages the sharing of ideas and feelings.
The extensive body of research on sentiment analysis, which has historically been
used to analyze consumer product reviews, assess movie reviews, run opinion polls,
and categorize remarks as positive or negative, served as the basis for this work [3].
But in this instance, we’re concentrating on a vital and important subject: depression
analysis. An increased interest in automating the identification of depression in social
media posts, including those on Facebook, Twitter, and Instagram, among others, is
a result of the rise in suicide occurrences. In this method, we create a vocabulary that
includes both terms that are common and words that are connected to depression.
After that, we check this dictionary against the terms used in social media posts.
A post is labeled as belonging to the depression category if it includes any of the
words in the list. Notably, Lenore Sawyer Radloff’s questionnaire, which consists
of 20 inquiries on the mental health and sleep habits of the respondent, is a useful
instrument for determining depression based on a cumulative score [4].
The Becks Depression Inventory, which includes 21 areas relating to users’
mental and physiological states, including mood, failure-associated sentiments, lack
of fulfillment, irritation, guilt, and self-hatred, has been used by other academics
to examine related topics [5]. Furthermore, Richardson has examined the PHQ-9
item’s performance characteristics and reliability as a gauge for identifying adoles-
cent depression [6]. Our research is centered on the critical task of depression anal-
ysis using a comprehensive dataset extracted from the Twitter platform [7]. Given
the immense popularity and high user activity on Twitter, it serves as an invaluable
source for understanding and potentially intervening in cases of depression. To tackle
Sentiment Analysis Using Bi-Directional LSTM for Depression … 57

this complex problem, we have tapped into the strength of several advanced learning
models, which include bi-directional LSTM, uni-directional LSTM-RNN, Vanilla
RNN, and Convolutional Neural Network (CNN). To gauge the performance of these
models, we have implemented a systematic and rigorous evaluation framework. This
methodology makes use of recognized quantitative indicators including the F1 score,
recall, accuracy, and precision. These metrics provide us with a comprehensive under-
standing of each model’s efficacy in detecting and classifying depression-related
sentiments within Twitter posts.
By employing a diverse array of machine learning techniques and rigorously eval-
uating their outcomes, our study aims to contribute valuable insights and methodolo-
gies for the early detection of depression on social media platforms. These insights
not only hold the potential to assist individuals in need but also offer an avenue
for the development of automated systems capable of identifying and mitigating the
growing concern of depression and its associated risks in the digital age.

2 Related Work

In today’s digital age, depression and suicide related to it are on the rise. Many writers
suggest using text mining and machine-learning approaches to examine textual data
and the likelihood of depressed posts on social media in order to forecast depression.
Deep learning models perform better than conventional text mining and machine
learning techniques because they can manage big datasets and have more processing
power. Two extensively researched Deep Neural Network (DNN) designs to handle
different nlp tasks are Convolutional Neural Network (CNN) and Re-current Neural
Network (RNN) [8, 9]. Many study topics, including text categorization, depression
detection, mental health, and sentiment analysis, are connected to the current effort.
It is possible to identify depressive posts using various machine learning tech-
niques. The authors [10] presented a straightforward method for identifying depres-
sion and preventing suicide by utilizing a learning model for long-term short-term
memory that was evaluated on a Twitter dataset made up of users’ random text posts.
Performance evaluation of multiple classification algorithms which are Naïve Bayes
classifier, RNN and Support vector Machine has also been describe briefly.

3 Proposed Methodology and Algorithm Design

The proposed work is mostly broken down into parts that include loading the dataset
and pre-processing it to eliminate undesired components such as emails, characters
like (! @#$%%^&), and numeric values. Each post has a binary label set to 1 for
a depressive post and 0 for a normal post. The texts are then converted to UTF-8
format, and stop words and punctuation are eliminated. Text and tweets are arranged
58 S. Singh et al.

into padded word sequences with spaces between them. The split into lists of tokens
is the pad sequence.
We modified the pipeline to employ pre-trained word embedding (GloVe) and
several neural network architectures in our experimental environment in order to
assess the sentiment of text data and identify depressive symptoms. To attain the
greatest performance in detecting sad sentiment in text, the architecture and model
hyper parameters are selected through testing. GloVe embedding are used to turn
words into dense vectors that accurately capture their semantic meaning. For text
data to be quantitatively represented, these embedding are essential. Training, vali-
dation, and test sets make up the dataset that contains text entries. This segment aids
in assessing the model’s effectiveness. For sentiment analysis and depression diag-
nosis, many neural network topologies are taken into consideration. These include
long time-series data model (LSTM) RNN, bidirectional LSTM RNN, Re-current
Neural Networks (RNN), and Co-evolutional Neural Networks (CNN). Each embed-
ding layer corresponding to the training data is fed into CNN if it is the model of
choice. The text is scanned using various filter sizes, with a global max-pooling
layer capturing the most crucial data from each filter. The resultant representations
undergo many broad hidden layers after max-pooling. Overfitting, a typical worry in
deep learning models for sentiment analysis, is avoided by using dropout. SoftMax,
the last layer, categorizes the text’s sentiment, namely whether it suggests despair or
not. In a bidirectional LSTM network, every embedding layer for the training data
is processed concurrently in both the backward and forward directions. This aids the
algorithm to gather textual context, which is crucial for sentiment analysis and the
diagnosis of sadness. Figure 1 give the graphical representation of work.
Social networks datasets are the most vital source for learning models. In the
current study, sadness analysis conducted using the depression twitter dataset.
The baseline includes typescript with attributes like views, opinions, remarks, and
responses.
From Twitter, 6,000,000 tweets were collected. It is a collection of information
from people about their ideas, emotions, beliefs, behaviors, reviews of products, etc.
About the facts it reflects, a tweet is full of views. The models which are used describe
briefly.

3.1 Bidirectional Long-Short Term Memory

Bidirectional Long-Short-Term Memory is a learning model used in a variety of NLP


and ordered data analysis applications. It is best suited for activities where it is crucial
to comprehend context both in the present and in the past. The Bi-LSTM manages
sequences where the order is important, such as time or text series. Applications like
sentiment analysis have made extensive use of it.
The Bi-LSTM RNN is used to handle text sequences successfully because text
data, such as social media posts or communications, frequently serves as the foun-
dation for sentiment analysis in depression identification and suicide prevention.
Sentiment Analysis Using Bi-Directional LSTM for Depression … 59

Twier Textual Posts


Dataset

Text Pre-processing and Word Embedding


representaon

Training and validaon Test dataset


dataset

Bi-direconal Unidireconal Vanilla RNN CNN


LSTM LSTM

Train the model

Train the model

Train the model

Train the model

Train the model

Fig. 1 Methodology architecture


60 S. Singh et al.

Fig. 2 Architecture of
bi-directional LSTM- RNN

The temporal relationships present in such data are superbly captured by it. It relies
heavily on LSTM units to comprehend and model the emotional and contextual cues
that are present in textual expressions. For identifying patterns and mental states that
can point to depression or distress, these units are ideally suited. It is crucial to take
into account both the text’s immediate context and larger context in the context of
suicide prevention and depression detection. By processing the text in both direc-
tions, the Bi-LSTM solves this problem by including data from both the previous
and subsequent text parts. The network’s capacity to detect small emotional fluctua-
tions and symptoms of discomfort is improved by this bidirectional method. Figure 2
shows the graphical representation if Bi-Directional.
Sentiment analysis with an emphasis on recognizing signs of depression and
suicide risk is the main goal of using the Bi-LSTM RNN. It combines two LSTM
networks, while one processes the information from right to left, the other processes
it from left to right. Both of their contributions are usually integrated in some fashion.
By mathematically, we can represent it as.
Given an input U of length T, where U = (U1, U2 , U3 ,…UT ), and the LSTM
hidden state from forward LSTM can be represented as
 
(f ) (f ) (f ) (f ) (f )
it = σ WiU Ut + Wih ht−1 + bi (1)
 
(f ) (f ) (f ) (f ) (f )
ft = σ WfU Ut + Wfh ht−1 + bf (2)
 
(f ) (f ) (f ) (f )
ot = σ WoU Ut + Woh ht−1 + b(fo ) (3)
Sentiment Analysis Using Bi-Directional LSTM for Depression … 61
 
(f ) (f ) (f ) (f )
gt = tanh WgU Ut + Wgh ht−1 + b(fg ) (4)

(f ) (f ) (f ) (f ) (f ) (f ) (f ) (f )
ct = ft @ct + it @gt (6)ht = ot @ tanh(ct ) (5)

(f )
ht = o(b) (b)
t @ tanh(ct ) (6)

Processing the sequence in reverse for backward LSTM.


 
it(b) = σ WiU
(b)
Ut + Wih(b) h(b)
t−1 + b(b)
i (7)
 
ft (b) = σ WfU
(b)
Ut + Wfh(b) h(b) (b)
t−1 + bf (8)
 
o(b)
t = σ W (b)
oU Ut + W (b) (b)
h
oh t−1 + b(b)
o (9)
 
gt(b) = tanh WgU
(b) (b) (b)
Ut + Wgh ht−1 + b(b)
g (10)

ct(b) = ft(b) @ct(b) + it(b) @gt(b) (11)

h(b) (b) (b)


t = ot @ tanh(ct ) (12)

where it, (f) , ft, (f) , ot, (f) , gt, (f) , ct, (f) and ht, (f) are input, for gate, output gate, cell gate,
cell state, and hidden, state for forward LSTM. Where else it (b) , ft (b) , ot (b) , gt (b) ,
ct (b) and ht (b) denotes same for backward LSTM. We consider @ as the hyperbolic
tangent activation function is tanh, while the sigmoid activation function is σ when
multiplication is done element-wise. The final equation for Bi-LSTM output at time
step t is with forward and backward hidden state
 
(f )
h(bi)
t = ht : h(b)
t (13)

[:] denotes concatenation.

3.2 Uni-Directional LSTM-RNN

The uni-directional LSTM-RNN is a crucial tool when evaluating textual data, which
is frequently used to determine sentiment and emotional states. It is great at managing
text data sequences and works well at capturing temporal relationships in the data.
The LSTM units, which are skilled at comprehending and simulating the emotional
62 S. Singh et al.

nuance and temporal dynamics included in text, form the foundation of the Uni-
LSTM design. The network can identify patterns and emotional indicators that can
point to sadness or discomfort thanks to these units. The Uni-LSTM only ever
processes the input sequence in one way, either forward or backward, unlike the
bidirectional form [11]. This indicates that it only takes into account information
from the text’s present or future context, not both at once.
Applying sentiment analysis with an emphasis on recognizing signs of depression
and suicide risk is the main goal of using the Uni-directional LSTM-RNN. In order
to facilitate early intervention and support, the network makes use of its sequential
modeling capabilities to offer insights into people’s emotional states. The study tech-
nique includes gathering data sources, social networking sites, and questionnaires
about mental health. This information is tagged or annotated to show sentiment
and emotional states, such as those connected to depression and probable suicide
thoughts. Utilizing the gathered and annotated data, the uni-directional LSTM-RNN
is trained to discover trends and connections between text and sentiment. In order
to enhance performance in recognizing depressive symptoms and potential suicide
risk, the training procedure include adjusting model parameters and architecture. The
sentiment analysis model’s performance is evaluated using a variation of assessment
metrics, plus remembrance, exactness, accuracy, and F1 score. The model’s perfor-
mance in real-world situations, such early identification and intervention in suicide
prevention, is extensively validated [12].
The uni-directional LSTM-RNN is nevertheless a useful tool for sequential data
analysis in the context of mental health even though it lacks the bidirectional context
processing of its cousin. Despite having a unidirectional context emphasis, it plays a
crucial role in enabling the precise identification of depressed attitudes and possible
suicide risk indicators within textual data. This emphasizes how crucial it is to use
cutting-edge machine-learning methods for prevention and support of mental health
issues [13, 14].

3.3 Convolutional Neural Network (CNN)

CNN which was first developed for image analysis, have shown to be useful resources
in the field of sentiment analysis for mental health. It is particularly good at spotting
localized textual patterns and traits that can point to depression or a higher risk of
suicide. It can be efficiently modified to evaluate sequential data, such as text. This
emphasizes the value of utilizing cutting-edge machine learning methods suited to
different data sources to enhance mental health awareness and prevention[15, 16].
A 1D CNN is used in this situation to evaluate text sequences, making it ideal
for spotting regional trends and distinguishing textual elements. The convolutional
layers of the CNN apply filters to input sequences, looking for patterns such as word
pairings, n-grams, or sequences suggestive of particular emotions. After the convo-
lutional layers, pooling layers, often Max Pooling, cut down the size of the feature
map, allowing the model to concentrate on important text patterns. The main goal is
Sentiment Analysis Using Bi-Directional LSTM for Depression … 63

to undertake sentiment analysis with a focus on identifying depressive symptoms and


suicide risk. When it comes to identifying localized text patterns linked to emotional
states and suffering, CNN architecture excels [17, 18].

3.4 Vanilla Network

The Vanilla network, which can recognize temporal connections and context, is a
useful tool for analyzing consecutive text input. Its critical significance in textual
data’s ability to detect depressive symptoms and suicide risk highlights the value
of utilizing cutting-edge machine learning techniques for mental health care and
prevention. RNNs excel at processing sequential data, like that found in text. RNNs
feature a hidden state that stores data from earlier time steps, unlike conventional
feedforward neural networks. This enables them to extract context and temporal
relationships from the supplied text. The hidden state, which changes over time as
the network analyses the input arrangement, is the fundamental idea of an RNN
[19–21].
The concept in an RNN is the hidden state, which advances over time as the
network processes the input ordered. The unseen state at time step t, denoted as ht.,
is updated on input Ut and prior unseen state ht,-1. Mathematically represented as.

ht = σ (WhU Ut + Whh ht−1 + bh ) (14)

where ht is the unseen state at time step t, Ut is the input at time t. WhU and Whh are
weight matrices, bh is the bias term. Σ is an activation function, often the hyperbolic
tangent or the rectified linear unit (ReLU).

4 Results and Discussion

The experimental work was completed using Windows 11 with 8-GB of RAM, a
CPU clocked at 1.60-GHz, and cache memory sizes of 256-KB, 1-MB, 6-MB, L1,
L2, and L3 accordingly. Table 1 lists the hardware description of the computer in
use.

Table 1 Hardware
Hardware Measurements
configuration
CPU-clock speed 1.60-Ghz
RAM 8-GB
L1-Cache-Memory 256-KB
L2-Cache-Memory 1-MB
L3-Cache-Memory 6-MB
64 S. Singh et al.

Table 2 Dataset
Attributes Types
specification
label Numeric (0 or 1)
Ids Numeric
Date Numeric
Flag Numeric
User Text
Text Text

Fig. 3 Matrix for


bi-directional LSTM

The twitter dataset [] consists of "*label", "*ids", "*date", "*flag", "*user", "*text"
as 0 for depressive, 1 denotes normal posts, The data specification is shown in Table 2.
80–20 ratio has been used to train test, the suggested model is trained. After
performing experiment, we extracted confusion matrix and AROC curve for
qualitative analysis. Confusion matrix for above models is represented in Fig. 3.
AROC curve has also been obtain to analysis the performance of models to get
clear overview. The ROC curve is represented in Figs. 4, 5, 6, 7, 8, 9 and 10.
The quantitative analysis for all models that has been used in experimental work
has been evaluated by precision, recall, f1-score and accuracy. Which help us to
find that which model has better performance on depressive dataset. In Table 3, the
quantitative analysis has been described in tabular manner.
Sentiment Analysis Using Bi-Directional LSTM for Depression … 65

Fig. 4 Matrix for CNN

Fig. 5 Matrix for Vanilla


RNN
66 S. Singh et al.

Fig. 6 Matrix for


uni-directional LSTM

Fig. 7 AROC curve of


bi-directional LSTM
Sentiment Analysis Using Bi-Directional LSTM for Depression … 67

Fig. 8 AROC curve of CNN

Fig. 9 AROC curve of


uni-directional LSTM

Fig. 10 AROC curve of


Vanilla RNN
68 S. Singh et al.

Table 3 Quantitative analysis of ML models


*Model *Accuracy *Precision *Recall *F1-score
Bi-directional LSTM 0.8319 0.8256 0.83 0.8304
Uni-directional LSTM 0.8293 0.8326 0.48 0.8294
Vanilla RNN 0.8236 0.8485 0.51 0.8274
CNN 0.7953 0.7720 0.50 0.7898

5 Conclusion

Different learning models were used in our experimental training to identify depres-
sion and prevent suicide, and the results showed different performance outcomes. For
the following reasons, the Bidirectional long short-term Memory architecture stood
out as the most efficient, promising model among those taken into consideration.
The Bi-LSTM model proved to be more effective in identifying long-range relation-
ships in textual data. It made use of its bidirectional capabilities to evaluate text both
forward and backward, allowing for a thorough comprehension of the context—a
critical component in the field of sentiment analysis, in particular.
There are significant practical ramifications for the Bi-LSTM model’s improved
performance. It implies that the use of such models in practical applications, including
automated risk assessment and mental health screening, may result in more early and
accurate treatments, which may save lives.
To sum up, our experimental results highlight the efficacy of Bi-directional LSTM
models in detecting depression and preventing suicide. These results clearly show the
promise of Bi-LSTM as an effective tool in diagnosing and managing mental health
difficulties, even though the optimal model selection may rely on the particulars of
the dataset and the nature of the sentiment analysis job.
Suicide prevention techniques and mental health services might both be greatly
improved by more study and development in this area.

References

1. Neri F, Aliprandi C, Capeci F, Cuadros M, By T (2012) Sentiment analysis on social media. In:
2012 IEEE/ACM international conference on advances in social networks analysis and mining
2. [Link]
20social,or%20website%2C%[Link]. By Amanda Hetler
3. Chandra AJ (220) Sentiment analysis using machine learning and deep learning. In: 2020
7th international conference on computing for sustainable global development (INDIACom),
pp1–4. [Link]
4. Radloff L (1991) The use of the center for epidemiologic studies depression scale in adolescents
and young adults. J Youth Adolesc 20:149–166. [Link]
5. Jackson-Koku G (2016) Beck depression inventory. Occup Med 66(2):174–175. [Link]
org/10.1093/occmed/kqv087
Sentiment Analysis Using Bi-Directional LSTM for Depression … 69

6. Richardson LP, McCauley E, Grossman DC, McCarty CA, Richards J, Russo JE, Rockhill
C, Katon W (2010) Evaluation of the patient health questionnaire-9 Item for detecting major
depression among adolescents. Pediatrics 126(6):1117–23. [Link]
0852. Epub 2010 Nov 1. PMID: 21041282; PMCID: PMC3217785
7. Depression on Twitter. [Link]
8. Sharat Sachin NMSA, Tripathi A, Nagrath P Sentiment analysis using gated recurrent neural
networks. SN Computer Science 1 (74)
9. Murthy GS, Allu SR, Andhavarapu B, Bagadi M, Belusonti Text based sentiment analysis
using LSTM. Int J Eng Res Technol (IJERT) 9 (5)
10. Singh S, Chandra SK (2023) Sentiment analysis for depression detection and suicide prevention
using machine learning models. In: Garg L key digital trends shaping the future of information
and management science. ISMS 2022. Lecture notes in networks and systems. Springer, Cham,
vol 671. [Link]
11. Akshat A, Tripathi K, Raj G, Sar A, Choudhury T, Saraf S, Dewangan BK (2024) A comparative
study between chat GPT, T5 and LSTM for machine language translation. In: 2024 OPJU inter-
national technology conference (OTCON) on smart computing for innovation and advancement
in industry 4.0. IEEE, pp 1–6
12. Dewangan BK, Alwan MH, Mahdi HF, Choudhury T, Piyush A, Rai A, Singh A (2022) Preven-
tive measurement and prediction of Covid-19 in India through business intelligence tools.
In: 2022 international symposium on multidisciplinary studies and innovative technologies
(ISMSIT) IEEE, pp 976–981
13. Lim WH, Sharma A (2023) Overview of swarm intelligence techniques for harvesting solar
energy. In: recent advances in energy harvesting technologies, pp 161–175. River Publishers
14. Pan L, Cheng WL, Tiang SS, Chong KS, Wong CH, Sharma A, Lim WH (2024) A robust
wrapper-based feature selection technique using real-valued triangulation topology aggregation
optimizer. Int J Adv Comput Sci and Appl 15(9)
15. Singh Y, Singh NK, Sharma A, Lim WH, Palamanit A, Alhussan AA, El-kenawy ESM (2024)
Bio-oil yield maximization and characteristics of neem based biomass at optimum conditions
along with feasibility of biochar through pyrolysis. AIP Advances 14(8)
16. Mohammed NMBR, Abhishek S, Hoo LT, Hong WC, Soon CK, Li P, Hong LW (2024) Tackling
photovoltaic (PV) estimation challenges: an innovative AOA variant for improved accuracy and
robustness. In 人工生命とロボットに関する国際会議予稿集. 株式会社 ALife Robotics,
vol 29, pp 871–876
17. Sahu R, Sahu V (2024) An energy-efficient algorithm for resource allocation in H-CRAN (EE
H-CRAM) for 5G networks. Wireless Pers Commun 138(3):1483–1499
18. Singh SP, Pathak D, Kumar A, Sanjeevikumar P (2023) An optimized fractional order modified
adaptive variable step-size LMS control approach to enhance DVR performance. IEEE Trans
Consum Electron
19. Aizaz Z, Khare K, Tirmizi A (2023) Approximate row-merging-based multipliers for Neural
Network acceleration on FPGAs. IEEE Embed Syst Lett
20. Khayoon AS, Ameen NH, Pithode K, McMahon S (2024) Electromagnetic Properties, Forming
Limit Diagrams And Fracture Toughness Of Laminated AL/FE 2 O 3 [Link]
Review and Letters 31(3)
21. Agrawal AV, Soni M, Keshta I, Savithri V, Abdinabievna PS, Singh S (2023) A probability-
based fuzzy algorithm for multi-attribute decision-analysis with application to aviation disaster
decision-making. Decis Anal J 8:100310
Disease Diagnosis from DNA Sequence
Using GPU-Based Aho–Corasick
Algorithm

Bandi Krishna, Ramdas Vankdothu, Suresh Kumar Lokhande,


Hussain Falih Mahdi, Rakesh Nayak, Bhupesh Kumar Dewangan,
and Tanupriya Choudhary

Abstract DNA sequencing provides the genetic information contained in a specific


DNA fragment, the entire genome, or a complex microbiome. DNA sequences
are required for biological researches involving structural analysis and a wide
range of applied applications including disease diagnosis, biotechnology, forensic
science, epidemiology, microbiology, and biological systematics. Because of its
ability to characterize and explain biological phenomena, DNA sequencing research
is becoming more common. By comparing healthy and changed DNA sequences,
researchers can detect diseases such as genetic disorders or cancers, define antibody
repertoires, and suggest treatments best suited to the patient. Having a quick way
to sequence DNA allows for the identification and cataloging of more species, as
well as faster and more tailored medical care. The disease diagnosis requires an

B. Krishna · R. Vankdothu
Department of CSE, BalajiInstitite of Technology and Science, Warangal, India
S. K. Lokhande
Department of CSE, Osmania University, Hyderabad, India
e-mail: Suresh.l@[Link]
H. F. Mahdi
Collage of Engineering, University of Diyala, Baqubah, Iraq
e-mail: [Link]@[Link]
R. Nayak
Department of CSE, School of Engineering, OP Jindal University, Raigarh, India
T. Choudhary (B)
University of Petroleum and Energy Studies, Dehradun, Uttarakhand, India
e-mail: tanupriya@[Link]
B. K. Dewangan (B)
Symbiosis Institute of Technology, Nagpur Campus, Symbiosis International (Deemed
University), Pune, India
e-mail: [Link]@[Link]
Department of Computer Science and Engineering, OP Jindal University, Raigarh 496109, India

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 71
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
72 B. Krishna et al.

efficient algorithm because DNA contains a large sequence of nucleotides. Aho–


Corasick is optimal string pattern matching algorithm which can be used to find
relevant patterns in DNA sequence. Here, we proposed a GPU-based multi-string
pattern matching implementation of the Aho–Corasick method to evaluate the like-
lihood that certain nucleotide repeat illnesses and cancer types originate from a
DNA sequence. Finding multiple string pattern matching has significantly sped up
as compared to the sequential technique.

1 Introduction

The biological information contained in the complex chemical molecule known as


DNA may be handed down across generations. All prokaryotic and eukaryotic cells,
including some viruses, contain it. DNA codes genetic information in chemical form.
Two lengthy polynucleotide chains made up of four different nucleotide subunits
make up a DNA molecule. Adenine (A), guanine (G), cytosine (C), and thymine
(T) are the four major chemical pairings found in DNA. DNA and RNA are a type
of nucleic acids having nucleotide as its base unit. A nucleotide have three major
components: a sugar molecule such as ribose in RNA and deoxyribose in DNA, a
base which contains nitrogen and a phosphate group attached to sugar molecule. The
order, sequence, of these nucleotides along each strand of helix like DNA structure,
encodes information and ensures its uniqueness [1]. The structure of a DNA molecule
is a double helix, consisting of two complementary strands of nucleotides joined by
hydrogen bonds between base pairs G and A. Each cell’s nucleus has a structure called
a chromosome that houses the DNA molecule. A chromosome contains thousands
of shorter DNA fragments, called genes, which stores instruction for fragments of a
protein, a complete protein, or multiple proteins. DNA serves as a storage medium
for the information required to build and run an organ ism.
The DNA sequencing involves finding the order of four nucleotides’ bases (A,
T, C, G), which are basic building block of DNA molecule and are responsible
for storing crucial genetic information. The four bases present in the DNA double
helix make hydrogen bond with their complementary base. Adenine (A) bonds with
thymine (T), while cytosine (C) bonds with guanine (G). Approximately 3 billion of
these base pairs make up the human genome, which has all the information needed
to create and maintain human cells.
The based-paired structure of DNA makes it ideal for storing a large amount of
genetic information [2]. The mechanism is based on complementary base pairing.
DNA molecules are copied, transcribed, and translated by complementary base
pairing, which is also the foundation for a number of DNA sequencing methods.
A mutation is an unintended change in a gene or genes, which occurs occasionally
and changes the behavior of newly created cell. This can result in a genetic disorder,
which is a medical problem. Single-gene, chromosomal, and complex illnesses are
the three types of genetic disorders [3]. The mutation changes the gene’s encoding
Disease Diagnosis from DNA Sequence Using GPU-Based Aho … 73

for making a protein, leading it to malfunction or disappear completely. Environ-


mental factor can potentially cause a mutation during your lifespan. The diagnosis
of diseases is a tricky subject in medicine. Among other things, diseases may originate
from duplicate genes, genes that repeat, genes that are absent, or from the specific
presence of genes that are linked to a disease. DNA sequencing is a fully safe first
diagnostic method for the diagnosis of illness. DNA databases are enormous and
constantly expanding. Consequently, the process of analyzing datasets is not done
by hand.
Pattern matching algorithms are a huge part of computer science. Boyer Moore
method, Ramin-Karp algorithm, nave-string search algorithm, Needleman–Wunsch
algorithm, Hamming distance and Leven shtein algorithm, Commentz-Walter algo-
rithm, Aho–Corasick algorithm, and others are examples of pattern finding algo-
rithms. Finding a specific pattern in large datasets is the aim of pattern matching
algorithms [4]. There are many types of pattern matching algorithms, including
precise and inexact ones, as well as single and multi-string ones. While multi-string
pattern matching algorithms detect many patterns in a single run, single pattern
matching methods explore the target sequence once for each pattern. Bioinformatics
necessitates exact outcomes, which is why multi-string pattern matching approaches
are accurate. One technique that works well for this kind of work is the Aho–Corasick
algorithm, which can discover many strings or patterns in a single run. Aho–Cora-
sick algorithm performance can be significantly improved using advanced hardware
and parallel programming with GPU. NVIDIA created the CUDA programming
platform, which is used to boost algorithm performance [5].

2 Related Work

The theory behind the Aho–Corasick [6] algorithm is similar to that of other common
Pattern Matching algorithms such as KMP, which uses past comparisons to avoid
unnecessary comparisons. A trie data structure is used in this technique.
An implementation of the Aho–Corasick approach that proved successful in eval-
uating the DNA sequence on heterogeneous GPU clusters was given by Antonino
Tumeo and Oreste Villa [7]. They split the DNA sequence up into manageable chunks
and assign each chunk to a different CUDA thread. It used eight MPI processes and a
single GPU architecture to produce performance that was five to twelve times faster
than any multiprocessor solution.
The pattern matching method, according to Rafiq et al. [8], has a substantial
contribution in the enabling network security, where it is used to identify intrusion,
unusual activities on network, and assist in intrusion detection systems. The research
suggests employing pattern matching algorithms in a prefix matching manner, which
will cover a broad range of detection techniques in network security, and argues that
the KMP algorithm is superior in this arena. It also compares alternative single string
pattern matching algorithms, such as Shift-Or, Karp-Rabin, Simon, and KMP seek the
74 B. Krishna et al.

same goal. The Agrep Algorithm for Approximate Nucleotide Sequence Matching
was first presented by Hongjian Li et al. [9].
Only a certain amount of genomic sequences can be matched by the pattern
matching techniques now in use. It produces undesirable results when used to large-
scale genomes. The CUDA version of Agrep algorithms provides only an approx-
imate matching technique that searches huge genomes for patterns up to 64 bytes
in length with an edit distance of up to 9. Its memory requirements are only 1/
4 of the entire genome size, thanks to the 2-bit binary en coding for bases. As a
result, many genomes can be loaded at the same time. Nave-Brute force approach,
Komentz-Walter algorithm, Booyer Moore algorithm, Knuth Moris Pratt algorithm,
and other pattern matching algorithms are used for comparison. In [10], Antonino
Tumeo et al. present the Aho–Corasick pattern matching algorithm version which is
classified as a multi-string pattern matching algorithm that operates both locally and
remotely. It looks for one or more already known character sequences present in a
dataset and detects them. Also, the suggested Aho–Corasick String Matching method
is tested under various software platform settings and examines various trade-offs of
parameters such as peak performance, performance variability, and data set size. The
research also includes many optimized architecture and algorithm methodologies for
shared and distributed memory architecture.
Cheng-Hung Lin presented the FPAC algorithm in [11], a failure-less parallel
version, which provides a better throughput when running on GPU and produces the
best results when compared to existing AC algorithms. The paper also provides an
overview of numerous pattern matching optimization techniques, as well as their GPU
implementations, which perform better than CPU implementations. For example, on
execution of GPU, the failure-less Parallel Aho–Corasick algorithm outperforms
classic AC algorithms.

3 Proposed Method

In this section, we go through the dataset we utilized and the steps we took to create
our model.

a. Dataset

The sequence information of disease genes are taken from the NCBI database for
genes.

b. Preprocessing
The DNA sequence and patterns are preprocessed with mapping function before
performing the task. This mapping function maps the A, C, G, T base symbol to A,
B, C, D. This helps in reducing the size of GOTO function and thereby taking less
memory on GPU which is a major constraint in GPU computation.
Disease Diagnosis from DNA Sequence Using GPU-Based Aho … 75

c. Aho–Corasick Algorithm

The Aho–Corasick algorithm is a multi-string searching method created in 1975


by Alfred V. Aho and Margaret J. Corasick [12]. It’s a matching algorithm based
on dictionary that finds elements from a finite set of strings within an input text. It
concurrently matches all strings. The Aho–Corasick technique can be used to look
for patterns in a large amount of text, making it useful in DNA sequence analysis,
data science, and other domains. The algorithm’s complexity is proportional to the
length of strings, the length of the searched text, and the number of output matches.
It consists of two parts: first, it creates finite state automata like pattern matching
structure called trie, from the keywords, and then it uses it to analyze the string in a
single pass.
Let a string text T, of length m, is to searched for k number of patterns.
GPU-Based Computing
Applications that benefit from the GPU’s performance often include a lot of
floating-point operations running in parallel. We use the GPU to achieve high perfor-
mance for the AC method, despite the fact that it is not a floating-point algorithm.
When it comes to conducting pattern matching operations in the AC algorithm, a GPU
has a lot of advantages. A GPU has a far larger number of (fine-grain) cores than
a general-purpose multi-core CPU, allowing huge simultaneous pattern matching
tasks to be executed in parallel. Furthermore, the multi-threaded execution of a GPU
can improve the throughput of simultaneous pattern matching tasks. A GPU also has
far more memory bandwidth than a multi-core processor. As a result, it can feed the
input and reference pattern data at significantly faster rates. We provide a perfor-
mance optimization technique for the AC algorithm on a GPU in this paper. Our
technique takes advantage of the GPU’s.
Aho–Corasick algorithm consists of two phases: Preprocessing: It involves
building an automaton for all the patterns. Here, automaton is a Trie like structure
with addition edges to allow multi-string matching. It has three functions:
Go-To: This method just follows the edges for all patterns in patterns. It is
expressed as a 2D array goto[][], in which we store the future state for the current
state and character.
Failure Function: This function records the edge that is followed when the current
character in automaton does not have an edge. It’s represented as a 1D array failure[],
which stores the future state for the current state. Output: It holds information of all
words that terminate in the current state have their indexes saved. It’s represented
as a 1D array output[], where all matching words’ indexes are saved as a bitmap for
the current state. Breadth First Search will be used to build the failure and output
functions (Fig. 1).
Matching: After creating automatons for all of the patterns, the input string text
is matched. Text is traversed character by character, and with the help of the goto []
[] function transitions in automaton are done. The automaton starts in the root state
and transitions to the next state based on the current character encountered. It is a
match if the current state has a bit set in the output [] function for a pattern. The
76 B. Krishna et al.

Fig. 1 Finite state


automaton

failure [] function offers information for the edge to be followed if there is no edge
present for a character in the current state.
Ture to improve throughput performance.
d. Aho–Corasick with GPU-Based Parallel Approach
The ability of Aho-Corasick to find multiple patterns in a large dataset in linear
time makes in suitable for the field of bioinformatics. Analyzing DNA sequence for
patterns which are responsible for diseases, one such problem where a large dataset
with billions of nucleotide need to be searched for pattern. Aho–Corasick algorithm
is best suited for such task as it provides the ability to find multiple patterns in single
traversal of sequence. Parallel algorithms are well-known for their ability to solve
large, complex problems quickly. Parallelization allows people to complete their
tasks on time and at a lower cost. The parallelization of Aho-Corasick method can
further improve the efficiency of task by utilizing the ability of dividing the task in
small, manageable tasks and executing those tasks in parallel way.
Aho–Corasick algorithm works in two parts. First part, preprocessing make a
finite state automaton for the patterns which needed to be searched and second,
matching on sequence with help of FSA in linear time. We use a single CPU core to
complete the first phase of AC in this paper. The second step is then run in parallel
on the GPU. The performance of algorithm can be improved by parallellization
of the task of search and partitioning the DNA sequence in smaller chucks where
individual threads perform the task of searching on one chunk. We have used CUDA
GPU framework to achieve the task of parallelization. The preprocessing part is
done on CPU to generate a finite state automaton for pattern matching and after
preprocessing the DNA sequence and the generated FSA is given to each the GPU
unit. Each GPU unit assigned a block of threads which will execute on the unit. The
DNA sequence stored in shared memory of each GPU unit. Each thread accesses the
assigned sequence portion and performs the task of matching.
First stage involves creation of goto function, which is stored in form a 2D matrix.
Goto function holds transition for each state and character combination. The size of
goto function can be reduced by only using 4 characters which are present in a DNA
sequence and addition character for absent nucleotide.
Disease Diagnosis from DNA Sequence Using GPU-Based Aho … 77

The second phase is the focus of our parallelization and performance improvement
technique, which includes efficient data placement, memory requirement reductions,
minimizing shared memory conflicts, and maximizing the effects of multithreading.

4 Results and Discussion

The dataset for these experiments is collected from NCBI standard datasets.
Following tables represent the standard information used to make prediction related
of disease with pattern counts. Table 1 shows standard ranges for nucleotide repeat
sequences. We use the information present in this table to determine the likelihood
of nucleotide repeat diseases. Based on given range, we can decide the probability
whether the patient is affected by a disease, is in pre-muted stage or normal stage,
or has no risk of disease. The pre-muted stage helps to estimate diseases occurring
in the primary stage, so patients are informed of this and therapy can begin sooner.
Data present in this Table 1 will be used for generating the diagnosis report after
Aho–Corasick algorithm generates count for all the patterns.
When the disease-causing genes are found in a specific sequence of chromosomes,
cancer is recognized in DNA. This cancer discovery at an early stage is extremely
beneficial to the patient’s treatment.
Table 2 shows different types of cancer and the genes responsible for the cancer.
We used input data sizes ranging from 50 to 250 MB and pattern counts ranging
from 2 to 12. Implementation of Aho–Corasick algorithm uses dataset of size 2.6 GB
for evaluation of run-time and performance comparison. The default configuration for
performance evaluation for CUDA environment is chosen as 16 blocks and 2000 MB
shared memory for individual blocks. The thread count is 128 for each block, if not
mentioned. Both sequential and parallel algorithms executed on GPU environment
provided by Google Colab, which provide access to “NVIDIA Tesla K80” GPU. Tesla
K80 has 4992 NVIDIA CUDA cores with a dual-GPU design along with access to
24 GB of GDDR5 memory.
a. Execution Time Comparison
For evaluating the parallel implementation, we have compared its performance
with the serial version of Aho–Corasick algorithm and have chosen their execution
time as the measure of performance. Both implementation runs on multiple files with
different number of patterns to show the time analysis of the parallel approaches in
comparison to the serial approach. Figure 3 shows the time taken to do the pattern
matching by serial and parallel implementation with different pattern number to
searched. Parallel implementation has significantly reduced the time required to the
task on DNA sequence. As the number of patterns increases, the performance of
parallel model increases as compared to serial approach (Figs. 2, 3, 4 and 5).
78 B. Krishna et al.

Table 1 Standard range of nucleotide repeat disease based on severity


Disease Pattern Normal range Pre-muted range Disease affected
HD CATG 1–30 29–38 38–182
SCA1 CATG 7–40 41 41–84
DRPA CATG 8–36 35–49 49–89
SCA3 CATG ¡32 32–33 33–201
SCA4 CATG 13–41 42–86 53–87
SMBA CATG 14–32 32–40 41
SCA12 CAG 7–36 35–49 49–89
SCA6 CAG ¡19 18 20–34
SCA7 CAG 5–18 29–34 ¿36 to ¿461
SCA17 CAG 26–43 43–49 45–67
DM2 CCTG 32–75 ¡31 75–11,001
FRAX-E GCC 5–40 40–201 ¿201
OPMD GCG 11–41 12–18 ¿12
FXSG CGTG 7–51 56–201 201–401
HDL3 CTTG 6–28 30–36 37–58
SCA9 CGTG 16–36 35–90 90–201
SCA11 ATTCG 11–30 30–401 401–4501
DM2 CTTG 6–38 38–51 ¿51
FRDG GATA 7–31 32–101 71–1001
FXN CGTG 8–51 56–201 207–401
HDL3 CTTG 9–28 30–36 37–58
SCA9 CGTG 16–35 35–90 90–201
SCA11 ATTCG 11–30 30–400 410–4501
DM2 CTG 6–38 37–51 ¿51
FRDB GAA 7–31 31–101 70–1001

Table 2 Cancer type and


Pattern Count Sequence affected Disease name
their responsible genes
CAN 16,890 Yes
GCG 345 Yes
CTG 13,822 Yes
CCTG 67 Pre-mutation DM2
GCC 17,904 No
GAA 567 Yes FRDA
CGG 6834 No
ATTCT 3785 No
Disease Diagnosis from DNA Sequence Using GPU-Based Aho … 79

Fig. 2 Figure (a) Flowchart of Aho–Corasick parallel implementation

Fig. 3 Execution time comparison of serial and parallel approaches

a. Speed-Up with Different Number of Patterns

Figure 4 shows the speed-up achieved by serial version and parallel version of
Aho–Corasick. With increase in number of patterns, the parallel implementation
performance consistently increases which shows its effectiveness to work well with
large number of patterns also. The speed-up for two patterns is 17 × as compared
to serial Aho–Corasick, and as the pattern numbers increase, the speed-up also
increases. For 10 and 12 pattern, it reaches up to 21X, and with increase patterns
after that, we see a small and consistent increase in speed-up.
80 B. Krishna et al.

Fig. 4 Speed-up of parallel


approach with respect to
serial approach

Fig. 5 Execution time of


parallel approach with

Figure 5 shows execution time of parallel implementation with different thread


counts 32, 64, 128, 256, 512 and the number of patterns to be searched. Other config-
uration parameters are same as previous, which is 16 blocks and 2000 Megabyte of
shared memory for each block. It is clear that the execution time reduces drasti-
cally with increasing the number of thread count as the work gets distributed to
more number of worker threads. For 32-thread version the execution time initially
increases with number of patterns but than starts flatten after 8 patterns, remains
close to 8 s, and has very small impact for the increase in number of patterns. As the
number of threads increases, the performance takes huge leaps for 32 to 64 threads,
but after a point increasing thread count gives very little performance increment. The
256-threads and the 512-threads execution time line becomes almost overlapping,
showing little performance increments with thread count for chosen configuration.
Figure 6 shows the execution time of parallel approach for different cancer [13]
types. It represents the success of parallel approach in handling matching task on
large file size. Cancer causing gene sequences were large in size and belonged to
a specific chromosome sequence. The proposed implementation of Aho–Corasick
reduces the time required to search the presence of such cancer causing genes in the
chromosomes.
Disease Diagnosis from DNA Sequence Using GPU-Based Aho … 81

Fig. 6 Time for finding


disease genes

5 Conclusion

Today, DNA sequencing has become an important area of work and its applications
are a matter of discussion. One of the safest approaches for patient diagnosis [14] is
disease diagnosis by DNA sequencing. In bioinformatics, the Aho–Corasick multi-
string matching technique is effective for disease diagnosis. Some reason of disease
is the existence of repeat, absence, copy, and exact presence of particular disease-
causing genes. Our proposed architecture is a performance optimization technique
for the AC algorithm on a GPU in this paper. Both the input text data and the reference
pattern data are efficiently placed and cached in the on-chip shared memories and
texture caches. The suggested GPU-based Aho–Corasick technique produces supe-
rior results in relatively less time than previous works and sequential approach. After
comparing the two methods, we find that the GPU-based parallel algorithm performs
considerably better than the sequential algorithm because it optimizes the searching
process. It also performs better as the number of patterns rises. The maximal speed-
up over the sequential version is around [Link] algorithm can be used to detect
disorders for delete genes, copy genes, and other diseases in future. The study is
being done to diagnose a restricted number of diseases; but, more nucleotide repeat
disease can be added. Apart from this performance can be further increased with the
help of combination of GPU and distributed computing approaches.

References

1. McMurray CT (2010) Mechanisms of trinucleotide repeat instability during human develop-


ment. Nat Rev Genet 11(11):786–799
2. Adey SP (2013) Gpu accelerated pattern matching algorithm for dna sequences to detect cancer
using cuda
3. Ropers H-H (2007) New perspectives for the elucidation of genetic disorders. Am J Hum Genet
81(2):199–207
4. Hasib S, Motwani M, Saxena A (2013) Importance of aho-corasick string matching algorithm
in real world applications. Int J Comput Sci Inf Technol 4(3):467–469
82 B. Krishna et al.

5. González-Álvarez DL, Vega-Rodríguez MA, Gómez-Pulido JA, Sánchez-Pérez JM (2011)


Finding motifs in DNA sequences applying a multiobjective artificial bee colony (MOABC)
algorithm. In: EvoBIO 2011: Proceedings of the 9th European conference on evolutionary
computation, machine learning and data mining in bioinformatics, Torino, Italy, 27–29 April
2011, pp 89–100. Springer, Berlin Heidelberg
6. Sanchez-Perez JM (2012) Predicting dna motifs by using evolutionary multi- objective opti-
mization. IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and
Reviews), 42(6):913–925
7. Regeciova D, Kolář D, Milkovič M (2021) Pattern matching in yara: improved aho-corasick
algorithm. IEEE Access 9: 62857–62866
8. Tumeo A, Villa O (2010) Accelerating dna analysis applications on gpu clusters. IEEE 8th
Symposium, pp 71–76
9. Rafiq AE, El-Kharashi MW, Gebali F (2004) A fast string search algorithm for deep packet
classification. Computer communications 27(15):1524–1538
10. Tumeo A, Villa O, Chavarría-Miranda DG (2011) Aho-Corasick string matching on shared and
distributed-memory parallel architectures. IEEE Trans Parallel Distrib Syst 23(3):436–443
11. Mane SU, Pangu KH (2016) Disease diagnosis using pattern matching algorithm from dna
sequencing: a sequential and gpgpu based approach. In: proceedings of the international
conference on Informatics and analytics, pp 1–5
12. Aho AV, Corasick MJ (1975) Efficient string matching: an aid to bibliographic search. Commun
ACM 18(6):333–340
13. Artificial intelligence techniques for detection as well as diagnosis of cancer and prediction of
mutations (2022) NeuroQuantology, vol 20, no 12, pp 231–238. [Link]
2022.20.12.NQ77019
14. Dash S, Nayak R, Gupta P, Ghugar U (2024) ITD-ML: improving diagnosis capabilities for
thyroid disease using machine learning. AI technologies for information systems and manage-
ment science. ISMS 2023. In: Lecture Notes in Networks and Systems. Springer, Cham, vol
1136. [Link]
Advancing Dogri-English Translation
Through Statistical Machine Translation
Technique

Vijay Singh Sen and Shubhnandan S. Jamwal

Abstract In recent decades, machine translation (MT) has undergone significant


progress. The process of translating between natural languages is approached as a
machine learning challenge in statistical machine translation (SMT). This method
relies on training models on parallel corpora, such as the Dogri and English languages
used in our experiment. In the expansive realm of machine translation, consider-
able efforts have been dedicated to exploring and advancing language pairs across
the linguistic spectrum. Numerous languages have undergone extensive research
and development, contributing to the evolution of machine translation technolo-
gies. However, amidst this wealth of exploration, the Dogri language has remained
relatively unexplored in the context of machine translation. This study proposes
a statistical machine translation system for the historically significant language of
Dogri, which is listed in the Indian Constitution and recently designated as the offi-
cial language of erstwhile state of Jammu and Kashmir. Dogri is a low-resource
language as besides various fundamental tools for machine translation line Dogri-
English dictionary and Dogri-English corpora do not exist. Author has manually
generated a corpus of 10,305 sentences for the Dogri-English language pair for the
study. In addition, a manually generated Dogri-English parallel corpus for devel-
oping SMT for Dogri containing 10,305 sentences employed for training, validating,
and testing the SMT system. Statistical machine translation system for the Dogri-
English language pair has been attempted for the first time, and the outcomes are
optimistic for continued system improvement. Dogri-English parallel corpus when
divided in the ratio 90:05:05, 80:10:10, and 70:15:15 for train set, validate set, and
test set resulted in a BLEU score of 24.39, 22.44, and 21.69, respectively.

V. S. Sen (B) · S. S. Jamwal


Department of Computer Science and IT, University of Jammu, Jammu, India
e-mail: senvjju@[Link]

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 83
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
84 V. S. Sen and S. S. Jamwal

1 Introduction

A lengthy history of manual and semi-automatic analysis by linguists, sociolo-


gists, psychologists, and computer scientists, among others, has benefited machine
processing of Natural Languages. In recent years, the accumulated work has borne
fruit in the form of freely accessible internet resources, ranging from dictionaries
to whole machine translation systems for some language pair. The term “machine
translation” refers to computerized techniques that automate all or a portion of the
translation process from one language to another. Document translation from one
language to another is in high demand in a vast, multilingual culture like India.
Although there has been research in the field of machine translation for many years,
developing effective machine translation systems is still a difficult task [1]. Since
human translation cannot meet the need for translation, the regular local and inter-
national exchanges have given machine translation a chance to thrive. As a result,
machine translation has advanced significantly since its inception in 1940, leading to
the development of numerous systems and methodologies [2]. The automatic trans-
lation of one natural language into another using computers is known as machine
translation (MT). Popular versions attribute the present interest in MT to a letter
made by Warren Weaver in 1949, just a few years after ENIAC went online. Interest
in MT is almost as old as the electronic computer. Since then, it has continued to be
a crucial tool in the field of natural language processing (NLP). Hutchins provides a
solid historical overview [3]. Recent advancements in the methodology and methods
of big data collecting, archiving, analysis, and visualization can be linked to the devel-
opment of natural language processing (NLP). Machine translation (MT), which was
the precursor to natural language processing (NLP), developed in the 1950s with an
aim of assisting in code breaking during World War II. Although the translations
were unsuccessful, these early stages of MT served as necessary precursors to more
advanced technologies [4].
There are primarily three methods for creating a machine translator. Specifically,
knowledge-driven techniques, also referred to as rule-based machine translation
(RBMT) and data-driven machine translation (DDMT), also referred to as corpus-
based machine translation. Lastly, hybrid machine translation, which combines the
benefits of the aforementioned techniques Since RBMT employs language theory
and DDMT uses particular language pair corpus, the classes are based on these
underlying techniques. Regarding the underlying technology, machine translation
has seen significant advancements over the last few decades. The development of
statistical approaches in Natural Language Processing (NLP) around the end of the
twentieth century marked the beginning of a logical progression that had started in
the middle of the century with rule-based machine translation (MT). The expanding
number of online texts made it possible to conduct statistical analysis on the orig-
inal texts as well as their translations, improving the rules, statistics, and outcomes
of earlier rule-based MT. With the availability of additional texts and translations,
statistical machine translation (SMT) as a field has developed and MT systems have
advanced. The majority of MT systems in the past relied on transfer components
Advancing Dogri-English Translation Through Statistical Machine … 85

that were reliant on language pairs, rule-based architectures constructed on elec-


tronic analysis and generation grammars. Grammatical rule construction for these
rule-based machine translation (RBMT) systems has always been meticulous and
time-consuming. The use of the vast amount of texts and information that is easily
accessible online for translations based on statistics and probability is a more recent
development in MT, and it has led to the emergence of a separate field of research in
MT called statistical machine translation (SMT) [5].
SMT, a data-driven method that uses parallel-aligned corpora and approaches
translation as a problem of mathematical reasoning, states that every sentence in the
target language is probably a translation from the source language. The accuracy
of translation increases as probability increases and vice versa. SMT architecture
majorly consists of three models language model, translation model and Decoder
model [6]. Formally, the source phrases is f ∈ F and F is the set of all source sentences
that need translation and E (f) is the set of all potential target language sentences that
can be created by translating.
By breaking down a full sentence into smaller, hidden elements that are translated
and concatenated to generate the derivation, machine translation systems carry out
this translation method. For instance, in phrase-based translation [7], the hidden vari-
ables will be the alignment between the phrases of the source and target sentences,
and in tree-based translation models, the latent tree structure utilized to produce the
translation will be expressed by the hidden variables [8]. Among the most impressive
applications of natural language processing which is machine translation, significant
improvements in translation accuracy have been made over time, and also translation
systems are being used by millions of users across the globe [9]. Statistical machine
translation (SMT) techniques that allow the building of statistical models from data
and the considerably enhanced data available for SMT model learning are two vari-
ables in particular that have significantly improved MT’s accuracy and coverage.
In India, machine translation initiatives date back to the mid-1980s. The Indian
Institute of Technology Bombay, International Institute of Information Technology
Hyderabad, Anna University-KB Chandrasekhar Research Center Chennai, National
Center for Software Technology, Mumbai, Indian Institute of Technology Kanpur,
and Center for Development of Advanced Computing (CDAC), Pune, are a few
notable names in the field of language translators besides initiatives from commer-
cial businesses like the Google, IBM India Research Laboratory of IT, Tata Insti-
tute of Fundamental Research (TIFR), and Super Infosoft commercial Limited. The
University Grants Commission (UGC) is also supporting minor and major research
projects involving development of linguistic parsers and machine translation systems
[10]. Indian researchers have also been engaged in machine translation research for
the last few decades. In some clearly defined fields, this aided in the development
of practical machine translation systems. Building a fully automatic, high-quality
machine translation system (FGH-MT) is very challenging. In practice, there isn’t a
system in existence that meets the requirements to be termed FGH-MT. Since 1990,
numerous institutions like CDAC (Pune), CDAC (Mumbai), IIIT (Hyderabad), IIT
(Kanpur), and others have been working on the development of MT systems as part of
86 V. S. Sen and S. S. Jamwal

programs funded by various government organizations in India like the Department


of Electronics (DoE), Department of Science and technology, DRDO, etc. [11].
The Dogri-English language pair is the main subject of translation since a vast
number of people utilize these for communication. Furthermore, Dogri is one of
India’s official languages, but English is a global language, translation between
these languages’ pair is essential for purely Dogri speakers to be connected to the
world and vice versa while utilizing cutting-edge technology. The Dogri language is
mostly spoken in the picturesque region of Jammu in the Indian state of Jammu and
Kashmir, but it is also prominent in the neighboring states of J&K, northern Punjab,
and Himachal Pradesh. Dogras are the speakers of Dogri, and Duggar is the name
of the region where it is spoken in particular. Dogri, a well-known language in J&K
State, is a significant language in northern India. It is a member of the Indo-European
language family’s Indo-Aryan branch. Its roots are in the ancient Indo-Aryan tongue,
which includes Laukik Sanskrit and the language of the Vedas. Dogri underwent the
same stages of evolution as other Modern Indo-Aryan languages, including Old Indo-
Aryan (Sanskrit) and Middle Indo-Aryan (Pali, Prakrit, and Apabhramsha), until
entering the Modern Indo-Aryan stage in the tenth century A.D. The former state
of Jammu and Kashmir just recently gave the Dogri language official recognition.
There are typically many multilingual sentences in the datasets used to train current
MT systems. Not all language pairings, nevertheless, have access to datasets of this
scale. The amount of data that is currently available is either very little or nonexis-
tent for a sizable number of languages spoken throughout the world. Low-resource
languages are those with comparatively little training data available. Because the
transformation calls for the creation of fluent and meaningful segments as well as
to maintain alignment between source and target, data augmentation is thoroughly
explored in other machine learning fields but less so in statistical and neural machine
translation. Out of vocabulary (OOV) is a challenge for machine translation (MT) in
languages with limited resources. It gets significantly worse when the source and/or
target languages are morphologically rich. One solution to the OOV issue is bilingual
list integration. More words can now be translated than in the training set as a result.
However, it will not translate inflected forms for morphologically rich languages like
Sinhala and Tamil because bilingual lists only translate words in their basic form.
The data augmentation strategies for dictionary terms demonstrate increased BLEU
scores for Sinhala-English SMT. These techniques expand bilingual lexicon terms
based on case-markers with the goal of creating new words to be used in statistical
machine translation (SMT) [12].

2 Survey of Statistical Machine Translation Systems:


Indian Languages

A gist of several SMT systems has been conducted in this subsection; the following
is a concise summary:
Advancing Dogri-English Translation Through Statistical Machine … 87

Hindi-To-Punjabi Statistical Machine Transaction System


Using a statistical Phrase-based methodology, A. Kumar and V. Goyal created a
Hindi-to-Punjabi machine translation system [13]. To train the system, enormous
parallel corpora were used, the corpus was generated manually and using the Rule-
based Hindi-to-Punjabi MT system that has already been established. The alignment
was created using Giza + + , which was also employed to extract phrases from the
parallel corpus. The authors created the Hindi tokenization module because Moses
did not support tokenization of Hindi words and to handle the translation of untrans-
lated words in Punjabi, a transliteration module was introduced. Since Hindi and
Punjabi are closely related languages, transliteration improves the overall correctness
of the system. Various domains including short stories, agriculture, health, politics,
business, entertainment, sports, tourism and science, were used to test the system.
These domains’ BLEU scores were 0.2467, 0.350, 0.2727, 0.2547, 0.2325, 0.2977,
0.2676, 0.3380, 0.2692, and 0.2676, according to the author’s claims.
Statistical Machine Translation Systems for Various Indian Languages as
Source Language and English as Target Language
Eight low-resource Indian languages (Hindi, Punjabi, Urdu, Telugu, Gujarati, Tamil,
Malayalam, and Bengali) were used as source languages to perform phrase-based
statistical machine translation (PBSMT) systems with English as the destination
language [14]. For the various language pairs, the crowd sourced parallel corpus and
the Enabling Minority Language Engineering (EMILLE) parallel corpus have been
employed. Due to the rich morphological structure of Indian languages, the BLEU
score has been reported to be extremely low for all language pairs.
Urdu-To-Punjabi Statistical Machine Translation System
An incremental approach based on the Singh et al. (2016) proposed Urdu-to-Punjabi
SMT system [15]. Five separate modules—Urdu tokenizer, segmentation, text cate-
gorization, Language and Translation Model, and decoder—make up the proposed
system. Parallel sentences taken from 50,000 parallel corpora are used to train the
system. This system claim to achieve a BLEU score of 0.86, and manual testing had
an accuracy rate of more than 85%.
English-To-Hindi Hierarchical Statistical Machine Translation System
Sukanta Sen et al. developed a hierarchical PBSMT system for English to Hindi
in 2016 [16]. The source sentence structure was rearranged using the rule-based
reordering tool to match the text structure of the target language. The English source
text parsed using the Stanford parser and the syntactic ordering is enhanced by
augmenting bilingual dictionary. IIT Bombay English-Hindi Corpus was utilized to
train the system. Phrase-based model, Hierarchical phrase-based model, Phrase-
based model with source sentence reordering, Hierarchical phrase-based model
after source sentence reordering, and Hierarchical phrase-based model after source
sentence reordering with bilingual dictionary are the five models developed and each
received a BLEU score of 10.79, 11.79, 13.18, 13.57, and 13.71.
88 V. S. Sen and S. S. Jamwal

Kannada-To-English Statistical Machine Translation System


A phrase-based statistical technique proposed by Shivakumar KM et al. (2015) for
the creation of a Kannada-to-English system [17]. The experiment is carried out using
Moses’ toolbox and utilizing the GIZA tool, the sentences are aligned. About 18,000
parallel sentences from a total of 20,000 parallel Kannada-English Bible sentences
are utilized for training, while 2000 parallel phrases are used for system testing and
tuning. Out-of-vocabulary (OOV) terms are transliterated into English once the input
material has been translated into that language. The three test sets—tuning set, test
set, and random test set—are used to test the system. According to the paper, the
BLEU score was 22.5 for the random test set, 10.68 for the test set, and 14.513 for
the tuning set.
Sinhala-To-Tamil Bidirectional Statistical Machine Translation System
A Sinhala-to-Tamil MT system developed by Randil Pushpananda et al. (2014)
utilizing a statistical methodology [18], word order is the same in both languages
(Subject-Object-Verb). The Sinhala Language Model was created using 850000
monolingual sentences from the University of Colombo School of Computing. The
Tamil Language Model was built on 407579 monolingual Tamil sentences from the
Tamil corpus and Translation Model on 25500 Sinhala-Tamil parallel corpus.

3 Hindi-To-English Bidirectional Statistical Machine


Translation System

Piyush Dungarwal et al. developed two models: Factored-based English to Hindi


and Phrase-based Hindi to English and SMT system [19]. Normalization of English
corpus Hindi corpus was carried out by using Stanford tokenizer and the Indic NLP
library respectively. Corpus contained 289,832 parallel sentences, 284,832 sentences
for training purpose and the remaining 5000 sentences were used as testing data.
English sentences were rearranged for the English-to-Hindi MT system using the
Stanford parser, and later super tags, numbers, and the case as a factor were added
to the English text. The Moses was then used to build the model. According to the
English-to-Hindi MT system, the case was tagged for WMT14 with a BLEU score
of 9.8 for reordering and super tag, 10.1 for reordering and number, and 9.8 for
reordering and super tag. The shallow parser is employed for the Hindi-to-English
MT system to sequence Hindi sentences. There were two models produced: a baseline
model and a baseline with reordering. According to the system, the baseline PBSMT
model had a BLEU score of 13.5, while the baseline with the source reordered for
WMT14 had a score of 13.7.
Advancing Dogri-English Translation Through Statistical Machine … 89

Assamese-To-English Statistical Machine Translation System


Das & Baruah created a machine translation system to translate from Assamese to
English using a corpus-based approach [20]. The system consists of four compo-
nents: the Language Model, Translation Model, decoder, and transliteration module.
IRSTLM was employed to train the Language Model, while GIZA + + was utilized
for the Translation Model. Moses was utilized as the decoder for input text. The
transliteration module was created through a rule-based approach. The training
corpus consisted of 8000 parallel Assamese-English sentences based on travel and
tourism in India. The system achieved a BLEU score of 11.32.
English-To-Sanskrit Statistical Machine Translation System
Sandeep Warhade & Prakash Devale created a phrase-based statistical machine trans-
lation system for translating from English to Sanskrit [21]. The system comprises of
three phases: language model, translation model, and decoder. The Language Model
and Translation Model were trained using a monolingual Sanskrit corpus and a bilin-
gual English-Sanskrit corpus. The decoder takes an English text as input and outputs
the Sanskrit translated text.
English-To-Tamil Statistical Machine Translation System
Loganathan Ramasamy et al. created a morphological-based English-Tamil statis-
tical machine translation system [22]. The system involves splitting morphological
suffixes on both the source and target sides. It was trained, tuned, and tested using
a corpus of 190,000 parallel sentences. The system was tested using both phrase-
based and hierarchical statistical machine translation approaches, and the authors
found that the hierarchical approach performed better or equally well on all corpora.
English-To-Telugu Statistical Machine Translation System
Anitha Nalluri and Vijayanand Kommaluri developed an “enTel” system for English-
to-Telugu translation using a statistical phrase-based approach [23]. The system was
trained using the EMILLE corpus and an English-to-Telugu dictionary developed by
Brown. SRILM and GIZA + + were used as the language and translation models,
respectively.
English-To-Tamil Statistical Machine Translation System
R. Harshawardhan et al. [24] developed an English-to-Tamil PBSMT system that
uses concept labeling. The system labels concepts in the input data, converts the text
into phrases, searches for the best target phrase in a parallel corpus, and maps it
to the same concept as the input phrase. The system was trained on 50,000 parallel
sentences, 5,000 proverbs, 1,000 idioms and phrases, and a dictionary of over 200,000
technical and 100,000 general words. The system achieved an accuracy of 70%.
90 V. S. Sen and S. S. Jamwal

English-To-South Dravidian (Malayalam, Kannada) Statistical Machine Trans-


lation System
Unnikrishnan P et al. developed SMT systems for translating from English to Malay-
alam and English to Kannada, incorporating syntax and morphological information
[25]. The process involved preprocessing the source (English) and target sentences
to add morphological information, followed by training the translation system using
a language model and translation model. Finally, Moses used the preprocessed
English sentence, Language Model, and Translation Model to decode and trans-
late the English sentence to the target language. The systems had a BLEU score of
24.9 for English to Malayalam and 24.5 for English to Kannada.

4 Research Methodology

MT is the name for computerized methods that automate all or part of the process of
translating from one human language to another. Fully-automatic general-purpose
high-quality machine translation system (FGH-MT) is extremely difficult to build.
In fact, there is no system in the world of any pair of languages which qualifies to
be called FGH-MT [26]. Our Dogri-English machine translation system is based on
statistical translation model, which was trained on manually generated Dogri-English
parallel corpus of 10,000 sentences.
Each Dogri word di is aligned independently with each English word ej. We esti-
mate the probability P(di|ej) for each alignment pair. We use a simple word translation
model where each Dogri word is translated independently to an English word. We
estimate the probability P(e|d) for translating a Dogri word d to an English word e.
We use a unigram language model for English, considering each English word inde-
pendently. We estimate the probability P(e) for each English word e. Given a Dogri
sentence D = (d1,d2,…,dn), we aim to find the English sentence E = (e1,e2,…,em)
that maximizes the joint probability:

E∗ = argmaxE P(E|D)

Using Bayes’ rule:

E∗ = argmaxE P(D|E) × P(E)

For each Dogri word di, select the English word ej with the highest translation
probability P(ej|di). Combine the translations of individual words to form candidate
English sentences. Score each candidate sentence based on the product of trans-
lation probabilities and language model probabilities. Select the highest-scoring
English sentence as the translation. This adapted model simplifies the translation
process by assuming independence between Dogri-English words and provides a
straightforward framework for statistical translation using a naive Bayesian approach.
Advancing Dogri-English Translation Through Statistical Machine … 91

Fig. 1 Dogri-English parallel corpus generation

Prior to this research work, parallel corpus of the Dogri-English language pair did
not exist and its unavailability posed a significant challenge for machine translation
efforts from Dogri to English. By creating this parallel corpus, comprising aligned
texts in both languages, we filled a critical gap in the field of natural language
processing. This corpus serves as the foundation for training and refining machine
translation algorithms specifically for translating between Dogri and English. Its
creation marks a significant advancement in the accessibility and accuracy of
translation services for the Dogri language. The creation of a parallel corpus for
Dogri-English followed the steps outlined in the diagram provided (Fig. 1).
Since Dogri is a low-resource language, corpus generated is not a particular
domain specific also with the aid of a linguist whole corpus is generated. Our
corpus contains data from several domains including government documents, news,
books, and articles. Raw text of both the languages was extracted from various
sources like textbooks, magazines, newspapers etc. although a few available for Dogri
language. English language text is available in abundance but not ready to process
form, it also required cleaning like removing of punctuation marks, lowercasing of
text, breaking paragraphs into sentences of length less than twenty-five word. The
sentences obtained in Dogri were translated by linguist into English language, and
English sentences were translated to Dogri language sentences, and a parallel corpus
of 10,305 sentences for Dogri-English language pair was generated.
92 V. S. Sen and S. S. Jamwal

5 Results and Discussion

Low-resource language like Dogri, data is not available in public sphere that could
be utilized for training machine. Corpus development for Dogri language is a big
challenge as we don’t have people skilled enough to digitize, although speakers are
many in number. In experiment, the Dogri dataset vocabulary had 10,148 words and
the English vocabulary had 9,110 words. Figure 2 gives the graphical representation
of corpus partition into above cited three sets (Table 1).
The corpus divided into three sets—train set, validate set, and test set in the
ratio of 80:10:10, 90:5:5, and 70:15:15; Bleu score of 22.49, 24.68 and 21.69 was
obtained, respectively. Despite the relatively small size of the corpus, promising
results were achieved. Despite its limited scope, the corpus yielded encouraging
outcomes (Table 2).
Statistical machine translation (SMT) refers to a method for automated trans-
lation that depends on statistical models trained using bilingual text collections.
This process entails examining extensive sets of parallel text data to understand

Fig. 2 Test cases data for Do-En SMT

Table 1 Do-En dataset


Dataset Dogri English
statistics
Sentences 10,305 10,305
Words 78,825 69,924
Distinct words 10,148 9,110

Table 2 BLEU score obtained in test cases


Test case Sentences in training Sentences in Sentences in test set BLEU score
set validation set
I 8244 1030 1031 22.49
II 9274 515 515 24.68
III 7214 1545 1546 21.69
Average 8244 1030 10,301 22.94
Advancing Dogri-English Translation Through Statistical Machine … 93

Fig. 3 Statistical machine translation model for Dogri to English

the patterns and likelihoods of translating words and phrases between languages.
Several key factors influence the quality of translation outcomes, including context,
ambiguity, and the overall quality and quantity of the parallel corpus (Fig. 3).
While sentence length plays a role, translation quality is influenced by various
other factors, including the quality and quantity of training data, the complexity of
the language pair, and the specific translation model used.
• Lack of Context: Short sentences may lack sufficient context for the translation
model to accurately capture the intended meaning. In our experiment, we used
our primary parallel corpus of Dogri and English that include small and medium
length sentences (1 to 20); majority of the sentences are of around fifteen words.
• Ambiguity: Short sentences are more prone to ambiguity, as there may be multiple
possible translations depending on the context or intended interpretation. Parallel
corpus in Do-En SMT contains different varieties of sentences that take different
contexts of repeated words.
• Corpus Quality and Quantity: Quality and quantity of parallel corpus is quite
crucial in statistical machine translation. We have prepared corpus with the consul-
tation of a linguist of Dogri and English. Corpus size of 10,305 sentences used for
experimentation may not be as sufficient but that remains the focus of our future
research.
• Out-of-Vocabulary Words: Words that are not present in the training vocabu-
lary (out-of-vocabulary words) posed a challenge for Do-En SMT systems. Post
Editing can significantly improve the quality of translation results.
94 V. S. Sen and S. S. Jamwal

6 Conclusion

Machine translation is a difficult task because of ambiguities, language differences


and the need to create corpora for training translation models on source language and
language models on target language. When the input to the system model is a corpus of
good quality and quantity, data-driven techniques like SMT function comparatively
well. The main goal of machine translation is to better understand the corpus of
parallel languages. The translation system’s capabilities are heavily influenced by
the training data that it receives. The system will learn the wrong things or learn
inefficiently if it is fed faulty or distorted data. Due to Dogri’s poor resource status,
no parallel corpus of the Dogri language with English or any other language that
might be improved or used for experimentation already existed. Therefore, a parallel
corpus of the Dogri-English language pair, consisting of 10,305 words, was generated
with the aid of a linguist. Dogri being a low-resource language poses a challenge
to generate a bilingual corpora of Dogri-English, although given a small dataset
that author has manually generated achieves a Blue score of 24.39, which is a good
beginning to move forward in the direction of machine translation between Dogri
and English. Noted observation while experimentation is that Statistical machine
translation is less suitable for language pairs with differences in word order, for
instance, Dogri subject-object-verb (SOV) and English subject-verb-object (SVO)
translations. Our future work includes enhance the corpus size for better translation
results.

References

1. Lopez A (2008) Statistical machine translation. ACM Comput Surv 40(3):49, Article 8. https://
[Link]/10.1145/1380584.1380586
2. Nair LR, David PS (2012) Machine translation systems for Indian languages. Int J Comput
Appl 39(1):0975-8887
3. Kituku B, Muchemi L, Nganga W (2016) A review on machine translation approaches. Indones
J Electr Eng Comput Sci 1(1):182–190
4. Hutchins J (2007) Machine translation: a concise history. Computer aided translation: Theory
and practice, 13(29–70):11, and a thorough general survey is provided by Dorr J, Benoit Dorr
BJ, Jordan PW, Benoit JW 1999. A survey of current paradigms in machine translation. In
advances in computers, vol 49, pp 1–68). Elsevier
5. Harriehausen-Mühlbauer B, Heuss T (2012) Semantic web based machine translation. In:
proceedings of the joint workshop on exploiting synergies between information retrieval and
machine translation (ESIRMT) and hybrid approaches to machine translation (HyTra), pp 1–9
6. Kituku B, Muchemi L, Nganga W (2016) A review on machine translation approaches. Indones
J Electr Eng Comput Sci 1(1):182–190
7. Antony PJ (2013) Machine translation approaches and survey for Indian languages. Int J
Comput Linguist Chin Lang Process 18(1):47–78
8. Sandeep S, Sahula V (2015) A survey of machine translation techniques and systems for
Indian Languages. In computational intelligence and communication technology (CICT). In:
2015 IEEE international Conference 676–681
Advancing Dogri-English Translation Through Statistical Machine … 95

9. Koehn P, Och FJ, Marcu D (2003) Statistical phrase-based translation. In: proceedings of the
2003 human language technology conference of the North American chapter of the association
for computational linguistics, pp 127–133
10. Yamada K, Knight K (2002) A decoder for syntax-based statistical MT. In: Proceedings of the
40th annual meeting on association for computational linguistics (ACL ‘02). Association for
Computational Linguistics, USA, pp 303–310. [Link]
11. Graham Y, Mathur N, Baldwin T (2014) Randomized significance tests in machine transla-
tion. In: proceedings of the ninth workshop on statistical machine translation. Association for
Computational Linguistics, pp 266–274, Baltimore, Maryland, USA
12. Hidden Markov Models and Text Translation: networks course blog for INFO 2040/CS 2850/
Econ 2040/SOC 2090 ([Link])
13. Kumar A, Goyal V (2018) Hindi to Punjabi machine translation system based on statistical
approach. J Stat Manag Syst 21(4): 547–552. [Link]
14. Jadoon NK, Anwar W, Bajwa UI, Ahmad F (2017) Statistical machine translation of Indian
languages: a survey. Neural Comput Appl 31(7):2455–2467. [Link]
017-3206-2
15. Singh U, Goyal V, Singh G (2016) Urdu to Punjabi machine translation: an incremental training
approach. Int J Adv Comput Sci Appl 7(4):227–238. [Link]
070428
16. Sen S, Banik D, Ekbal A, Bhattacharyya P (2016) IITP English-Hindi machine translation
system at WAT 2016. In: proceedings of the 3rd workshop on Asian translation (WAT2016),
pp 216–222. [Link]
17. Shivakumar KM, Nayana S, Supriya T (2015) A study of Kannada to English baseline statistical
machine translation system. Int J Appl Eng Res 10(55):4161–4166
18. Pushpananda R, Weerasinghe R, Niranjan M (2014) Sinhala-Tamil machine translation:
towards better translation quality. In: proceedings of the Australasian language technology
association workshop, pp 129–133. [Link]
19. Dungarwal P, Chatterjee R, Mishra A, Kunchukuttan A, Shah R, Bhattacharyya P (2014) The IIT
Bombay Hindi-English translation system at WMT 2014. In: proceedings of the 9th workshop
on statistical machine translation: shared task collocated with conference of association for
computational linguistics (ACL 2014), pp 90–96
20. Das P (2014) Baruah KK (2014) Assamese to English statistical machine translation integrated
with a transliteration module. Int J Comput Appl 100(5):20–24. [Link]
8084
21. Warhade SR, Patil SH, Devale PR (2012) English-to-sanskrit statistical machine translation
with ubiquitous application. Int J Comput Appl 51(1)
22. Ramasamy L, Bojar O, Žabokrtsky Z (2012) Morphological processing for English-Tamil
statistical machine translation. In: Proceedings of the workshop on machine translation and
parsing in Indian languages (MTPIL-2012), pp 113–122
23. Nalluri A, Kommaluri V (2011) Statistical machine translation system using Joshua: An 169
approach to build ‘entel’ system. In: special volume problems of parsing in Indian Language
24. Harshawardhan R, Sara AM, Soman KP (2011) Phrase based English- Tamil translation system
by concept labeling using translation memory. Int J Comput Appl 1–6
25. Unnikrishnan APJ, Soman KP (2010) A novel approach for English to South dravidian language
statistical machine translation system. Int J Comput Sci Eng 02(08):2749–2759
26. Murthy BK, Deshpande WR (2011) Language technology in India: past, present and future,
1998
27. Fernando A, Ranathunga S, Dias G (2020) Data augmentation and terminology integration for
domain-specific sinhala-english-tamil statistical machine translation. arXiv: 2011.02821
28. Sandeep S, Sahula V (2015) A survey of machine translation techniques and systems for
Indian languages. In: computational intelligence and communication technology (CICT). In:
2015 IEEE international conference, pp 676–681
29. Goyal V, Lehal GS (2009) Evaluation of Hindi to Punjabi machine translation system. Int J
Comput Sci Issues 4(1):36–39. [Link]
96 V. S. Sen and S. S. Jamwal

30. Islam MZ, Tiedemann J, Eisele A (2010) English to Bangla phrase-based machine translation.
In: EAMT 2010–14th annual conference of the european association for machine translation
31. Jamwal SS, Gupta P, Sen VS (2021) Hybrid model for generation of verbs of Dogri language.
In: Singh TP, Tomar R, Choudhury T, Perumal T, Mahdi HF (eds) Data driven approach towards
disruptive technologies. Studies in autonomic, data-driven and industrial computing. Springer,
Singapore
32. Jamwal SS, Gupta P, Sen VS (2021) A novel approach for identification and classification of
verbs in Dogri language. Int J Intell Eng Inform 9(4):412–423
33. Gupta P, Jamwal S (2021) Designing and development of stemmer of Dogri using unsupervised
learning. In: Marriwala N, Tripathi CC, Jain S, Mathapathi S (eds) Soft computing for intelligent
systems. Algorithms for intelligent systems. Springer, Singapore, pp 47–156
34. Jamwal SS, Sen VS (2022) Natural language interface in Dogri to database. In: Skala V,
Singh TP, Choudhury T, Tomar R, Abul Bashar M (eds) Machine intelligence and data science
applications. Lecture notes on data engineering and communications technologies. Springer,
Singapore. [Link]
Future Price Prediction of IT Sector
Companies Using Optimized Deep
Learning Approaches

Umar Bashir, Kuljeet Singh, Megha Raina, and Vibhakar Mansotra

Abstract The future is unpredictable and unknowable, but possible ways exist to
make predictions and get benefits securely. One such possibility is using AI and DL to
forecast the stock market. The equity market’s dynamic nature and intrinsic volatility
make this area more interesting for researchers. The fluctuations in the stock market
depend on many factors that make future predictions more difficult using simple
traditional statistical models. Therefore, this research suggests two optimized DL
approaches for the prediction of the closing price of the stock one week prior. In the
current study, RNN and LSTM with deep hyperparameter tuning have been applied
to the recent dataset of three IT companies to predict their future closing price. The
performance of these approaches has been evaluated using four evaluation metrics
including MAPE, R2 , MAE, and RMSE. As per the results achieved in this study, it
has been interpreted that the LSTM shows good and acceptable results compared to
RNN and it appears to be the suitable choice for making future investment decisions.

Keywords Stock market · Deep learning · Hyperparameter tuning · Time series


data

1 Introduction

Stock market prediction is difficult for finance and statistics experts because of the
highly volatile and frequent fluctuations in market behavior. The different prediction
approaches in the field of equity market play a key role in bringing new and existing
investors to a single platform. The era of stock market analysis and prediction has
gained researchers’ interest in developing a system that may help individuals and
organizations make decisions about buying or selling stock. The data of stock market
is considered a time series data, in financial markets, a time series is a share price

U. Bashir (B) · K. Singh · M. Raina · V. Mansotra


Department of Computer Science & IT, University of Jammu, Jammu (J&K) 180006, India
e-mail: umarbashir932@[Link]

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 97
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
98 U. Bashir et al.

progression. Trend prediction in financial time series is especially crucial for invest-
ments. In contrast to other time series, financial time series possesses few unique
features. A primary characteristic is the elevated occurrence of distinct values which
increases the impact of non-systematic influences on the dynamics of such time
series, leading to high volatility and non-stationarity. Forecasting of financial time
series data has been proven to be very complex due to its inherent uncertainty. Tradi-
tional statistical approaches [1, 2], machine learning approaches [3], and deep neural
networks [4] have all been extensively studied in this regard. The majority of models
nowadays emphasize accurate price prediction. On the other hand, some studies are
being done on forecasting the market trajectory. Prediction of future trends is more
concerned with future volatility trends than prediction of future prices because this
offers an alternative viewpoint for making decisions.
Investments are inherently risky due to the well-known non-parametric, chaotic,
and noisy nature of stock market prices. Furthermore, we are conscious that changes
in the price of stocks are regarded as a random process that exhibits more noticeable
short-term fluctuations, given the model specification that will be discussed momen-
tarily. Having in-depth information on potential changes in stock prices reduces the
risk soon. In the present, stock traders are more inclined to purchase a stock whose
value is anticipated to rise in future, and vice versa, when prices are down. Therefore,
it stands to reason that correctly predicting price patterns in the stock market will
maximize gains and minimize losses. It is hence advisable to acknowledge that it
is difficult to contribute value to this intricate and well-researched subject, particu-
larly a huge number of data points that are produced globally for each time under
review. Financial time series data is considered more complex as compared to other
statistical data due to irregular movements, cyclical fluctuations, seasonal data, and
long-term trends. Forecasting data with such extreme fluctuations and irregularities
is usually subject to large errors. Developing an efficient model that forecasts the
time series data more precisely and successfully is a major area of interest within the
data mining industry. Due to the nonlinearity of financial data, the traditional statis-
tical models applied in the field of the stock market were simple but suffered from
various drawbacks. As a result, researchers have transitioned toward more accurate
and efficient soft computing which provides flexible, robust, and adaptive solutions to
complex problems. Soft computing approaches like SVM, RF, ANN, and its variants
have been applied for financial forecasting, etc. The application of ANN to forecast
the movements and behavior of stocks has shown to be a competitive substitute for
existing traditional methods.
To comprehend better about stock prediction, this study describes the develop-
ment of DL approaches to forecasting stock prices at the company level. In the
current study, two optimized deep learning approaches RNN and LSTM have been
proposed to perform a stock closing price prediction one week prior. Deep hyper-
parameter tuning has been done in this study with two different optimizers, Adam
and rmsprop to optimize the model performance. For hyperparameter tuning, a grid
search approach is used and selects values of the hyperparameter on which the model
shows the lowest MSE and selected values of hyperparameters have been used for the
final optimized model. The final results are evaluated using four regression evaluation
Future Price Prediction of IT Sector Companies Using Optimized Deep … 99

metrics namely: MAPE, R2 , RMSE, and MAE. Both the proposed approaches show
good results but LSTM outperforms the other one on all three datasets and demon-
strates how well the suggested method works to depict the complex time-varying
dynamics of stock.
The remainder sections of the present study are structured as follows: Sect. 2
reviews existing literature and illustrates the development and use of several predic-
tive models in financial forecasting. The methodology includes dataset descrip-
tion, data preprocessing procedures, and experimental setup for RNN and LSTM
are discussed in Sect. 3. Section 4 presents the hyperparameter tuning and predic-
tion results of the proposed methods using various performance evaluation metrics.
Finally, the conclusion of this study and future work are provided.

2 Related Work

Numerous research has been conducted on the prediction of the stock market. The
study [5] investigates the performance of ANNs in predicting stock prices before and
after India’s demonetization. Eight years of data have been used to train multilayered
neural networks with the Levenberg–Marquardt algorithm, optimizing for minimal
Mean Squared Error (MSE). The study’s high regression values (0.999) indicate the
ANN’s effectiveness in predicting future stock prices and the CNX NIFTY50 index,
even amidst the significant economic changes due to demonetization. Another study
[6] focuses on enhancing stock price predictions using various SVM models (linear,
quadratic, cubic, and fine Gaussian). By dividing 170 days of stock market data
into training and testing sets, the researchers evaluated the models using MAPE and
RMSE. The fine Gaussian SVM outperformed other models, achieving the lowest
RMSE (0.009) and MAPE (2.4%), suggesting its superior accuracy in stock price
forecasting. In contrast, a study [7] on the Chinese share market addressed the limita-
tions of classical linear multi-factor models in predicting long-term stock trends. By
employing feature selection algorithms and time-sliding window cross-validation,
the researchers found that integrating the RF algorithm for feature selection and
prediction yielded the best results. This integrated approach facilitated the construc-
tion of a long-short portfolio, validating the model’s effectiveness in the chaotic
and dynamic stock market environment. The study [8] proposes a model using
SVM for classification and LSTM networks for time series prediction, trained on
20 years of data. This model aims to aid investors, especially those without finan-
cial backgrounds, in making secure investments by providing robust and periodic
factor-consistent predictions.
Another research [9] focuses on predicting stock trends across multiple exchanges,
specifically for companies listed on both the NSE and BSE. Using LSTM networks,
the CREST method predicts the price movement of a company based on its perfor-
mance in another stock exchange. The study evaluates the method using daily basis
data of Wipro, Infosys, and LTI, with performance metrics including RMSE, preci-
sion, recall, accuracy, and F-measure, indicating a high level of predictive accuracy.
100 U. Bashir et al.

A study [10] on the Tehran Stock Exchange explores the prediction of stock values
for different industry groups. The research covers diversified financials, petroleum,
basic metals, and non-metallic minerals, using a decade of historical data. Predic-
tions are made for various future time points (1 to 30 days ahead) using multiple ML
techniques, including DT, bagging, RF, Adaboost, GBM, XGBoost, ANN, RNN, and
LSTM. Among these, LSTM demonstrates the highest accuracy and model-fitting
capability. Tree-based models also show strong performance, particularly Adaboost,
Gradient Boosting, and XGBoost. A comprehensive study [11] evaluated several
approaches for price prediction of Infosys, ICICI, and SUN PHARMA. This research
compared a time series model (Holt-Winters Exponential Smoothing), an econo-
metric model (ARIMA), ML models (RF and MARS), and DL models (simple RNN
and LSTM). Among these, MARS emerged as the best ML technique, while LSTM
was the top performer among DL techniques. Overall, MARS outperformed others
in sales forecasting across the IT, banking, and health sectors. Another study [12]
explores the potential of AI and ML in predicting stock market trends. This paper
reviews the strengths and weaknesses of ML techniques in this domain, discussing
their potential benefits and risks. The study focuses on three specific technologies:
ANN, SVM, and LSTM. It provides insights into how these technologies can be
leveraged for more accurate stock market predictions, highlighting the capabilities
and limitations of each approach. A unique approach in [13] considers price charts
of stock as images, utilizing DLNNs for modeling. This method mimics the work
of technical analysts by predicting stock price movements using price charts and
fundamentals such as price-to-earnings ratios. The study finds that DL techniques
outperform the single-layer model in predicting the Chinese stock market. It empha-
sizes that price trends derived from past daily closing prices are more influential
in predicting future movements than fundamental stock metrics, demonstrating the
effectiveness of DLNNs in short-term stock prediction.

3 Materials and Methods

The suggested study’s workflow is shown in Fig. 1, and in general, there are three
main stages in the experimental setting each of which is defined in this section as
follows:
• Dataset collection and preparation.
• Proposed approach for prediction.
• Performance evaluation
Future Price Prediction of IT Sector Companies Using Optimized Deep … 101

Fig. 1 Workflow of the proposed methodology

3.1 Dataset Collection and Preparation

Dataset Description and Collection: To start the process of prediction, the initial
step is data collection. In the present study, after narrow qualitative research about
the reliable source of historical data, the National Stock Exchange (NSE) has been
selected because it provides up-to-date and verified information about the stock. In
this research, the leading three stocks of Nifty IT have been selected and their 5-year
data (total of 1255 trading days) has been collected from the NSE website (source:
[Link]), which ensures the collected data is reliable, accurate, and follows
regulatory standards strictly. These three companies have good contributions and
strengths in the Information Technology sector as shown in Table 1. The dataset
contains day-wise data about the stock, a total of 14 features which include the Date,
Series, Open, High, Low, Prev. Close, LTP, Close, VWAP, 52 WH, 52WL, Volume,
value, and No. of Trades.
Data Preparation: In a stock market prediction system, data preparation is crucial
because it involves cleaning and transforming the data so that the model can learn
effectively. NSE provides only one year of historical data at one time. It was collected

Table 1 Description of datasets and weightage of companies in Nifty IT


Stock Date Weightage in Nifty IT (%) Source
HCL Technologies Ltd 25-July-2019 to 9.88 NSE
24-July-2024
(total 1255 trading days)
Tech Mahindra Ltd 25-July-2019 to 9.35 NSE
24-July-2024
(total 1255 trading days)
LTIMindtree Ltd 25-July-2020 to 5.81 NSE
24-July-2024
(total 1255 trading days)
As per NSE report dated 30 August, 2024
102 U. Bashir et al.

in multiple one-year increments to obtain five years of data and then merged using
programming tools like Python Pandas library. Due to some technical issue, the
collected data contains a few duplicate and missing entries, which is very important to
handle carefully. During data preprocessing, duplicate entries have been removed and
there are various approaches to treat missing values like deletion (listwise, pairwise),
imputation (mean, median, mode), etc. In the current study, data imputation has been
deployed but none of the mentioned data imputation technique suits well because it
is known that stock market data is a time series data, so in the current study, missing
entries are substituted with the mean of last and next observation. After the treatment
of missing values data scaling has been done. Data Scaling is a process of bringing the
features into a common scale. This helps to converge the model faster and increases
the capability of the model to generalize on test data. There are different techniques
for data scaling, in this study Min–Max approach has been used and is computed
using Eq. 1.

X − Xmin
Xscaled = (1)
Xmax − Xmin

The dataset often contains a total of 14 features, but there exist some features
like ‘series’, ‘Date’, etc., which might not be relevant to the target feature directly
and may lead to potential overfitting and computational complexity. It is necessary
to remove them to reduce computation costs and to increase the performance of the
model [14]. For this, Recursive Feature Elimination (RFE) is recursively employed
to choose the only 10 most pertinent features: Open, High, Low, Prev. Close, LTP,
VWAP, 52 WH, 52WL, Volume, and value are selected for further processing. Now
the data is ready, the next phase is building and training the prediction model. For
training the model, the prepared dataset has been split into 80:20 ratios, i.e., 80%
of the dataset is for training and 20% (251 trading days) for model testing. As it is
already known stock market data is a time series data, so performing random shuffle
during data partitioning disrupts the temporal dependence which makes it difficult
for the model to learn significant trends within the data. To address this issue, data
partitioning has been done chronologically ensuring that the test set comes after the
period of a training set. This helps in assessing the predictive power of the model.

3.2 Prediction Approaches

There are several ML and DL approaches that have been applied in stock market
analysis, but RNN and LSTM have become more effective techniques because these
two approaches are specifically designed to analyze sequential data which is an
essential step in stock market analysis. Furthermore, RNN and LSTM are discussed
as follows:
Future Price Prediction of IT Sector Companies Using Optimized Deep … 103

Fig. 2 Basic architecture of RNN

Recurrent Neural Network: RNN is an essential variant of neural networks that is


widely used in many applications. A neural network typically creates output after
the input goes through a few layers. Although it is suggested that two successive
inputs be completely independent of one another, that is not true in all the cases.
For example, it is essential to review the previous samples to anticipate the market
at a certain time. RNN is a suitable approach for solving problems that involve
sequential data [15]. It is named Recurrent Because it performs the same action for
every item in a sequence when the output is connected to the previously estimated
results. Another crucial feature of RNNs is memory, which keeps track of previously
calculated data for a considerable amount of time. While RNNs may theoretically use
random information for extended sequences, in reality, they are limited to looking
back a few steps. Figure 2 shown below depicts the RNN architecture. While RNNs
are excellent for sequential data, but can also suffer from vanishing gradient problems
which make it difficult for RNNs to capture long-term dependencies. This problem
can be solved using variants of RNN like LSTM.
Long Short-Term Memory: LSTM is a networking model proposed by Sepp
Hochreiter and Juger Schmidhuber in 1997 [16]. It is designed to address issues
such as gradient expansion and gradient disappearance in RNNs. A conventional
RNN has a single repeating module usually a tanh layer and a simple internal struc-
ture [17]. As shown in Fig. 3, the LSTM cell is composed of three gates: the forget
gate, the input gate, and the output gate [18].
C t-1 represents the cell from the prior instant, at the very end, ht-1 represents the
LSTM neural unit’s final output value, X t represents the current input, activation
function used in the cell is denoted by σ, F t and I t describe the current moment
output of the forget gate and input gate, respectively, the output of output gate is Ot ,
Ĉ t represents the status of candidate cell at the current moment, Ct is the current
moment cell state, and ht is the output of the cell [19]. The processes involved in
LSTM computation are as follows:
104 U. Bashir et al.

Fig. 3 Architecture of the LSTM memory cell

(1) The input for the forget gate is the final outcome of the previous instant and
the present input. After performing computation, the output of the forget gate
is obtained by Eq. 2.
   
Ft = σ Wf ht−1 , Xt + bf (2)

where F t lies between 0 and 1, and W f and bf represent the weight and bias of the
forget gate.
(2) The output of the last moment and the present input are given to the input gate.
The output of the input gate and the state of the input gate’s candidate cell are
computed using Eq. 3 and 4.

   
It = σ Wi ht−1 , Xt + bi (3)

   
C = tanh Wc ht−1 , Xt + bc (4)

where the value of I t lies between 0 and 1, W i and bi are the weight and bias of the
input gate, and W c and bc are the weight and bias of the candidate input gate.
(3) Eq. 5 depicts how the current cell is updated.


Ct = Ft ∗ Ct−1 + It ∗ C t (5)

where Ct value range is 0 to 1.


Future Price Prediction of IT Sector Companies Using Optimized Deep … 105

(4) The last-moment output and the current input are given to the output gate. The
outcome of the output gate is computed by using Eq. 6.

   
0t = σ W0 ht−1 , Xt + bo (6)

where Ot range is between 0 and 1, and W o and bo are the weight and bias of the
output gate.
(5) The final output of the LSTM cell is computed from the result of the output gate
and the state of the cell as depicted in Eq. 7.

ht = Ot ∗ tanh(Ct ) (7)

3.3 Performance Evaluation

After deploying DL models, evaluating the performance of these approaches is a very


critical step to ensure continued effectiveness. Various evaluation metrics applied in
this study are briefly discussed below:
MAPE is a scale-independent approach that measures the error percentage between
forecasted and actual prices [20]. It is calculated as shown in Eq. 8:
 
1 n  Dpre − Dact 
MAPE =  ∗ 100 (8)
n 1 Dact

R2 is used to assess the goodness of linear fit. It is also called the coefficient of
determination and is defined in Eq. 9:
n
1 (Dact − Dpre )2
R2 = 1 − n (9)
1 (Dact − Dact )2

RMSE is a square root of MSE. It is a scale-dependent metric which means that it


provides errors in the same unit [21]. It is defined as in Eq. 10.

1 n
RMSE = (Dpre − Dact )2 (10)
n 1

MAE computes the average absolute error between the actual price and the predicted
price without considering their direction. It is calculated as shown in Eq. 11.
106 U. Bashir et al.

n
i=1 |Dact − Dpre |
MAE = (11)
n
where Dpre , Dact , and Dact represent predicted price, actual price, and mean of actual
price, respectively, and n indicates the data length of the stock.

4 Hyperparameter Tuning and Prediction Results

4.1 Hyperparameter Tuning

To optimize the performance of any DL approach, tuning of hyperparameters is a


very crucial step. It entails modifying some parameters that govern the design and
training procedure of the model. Various hyperparameters like Units, Optimizer,
Batch size, dropout, and epochs have been optimized in this study. Different values
of each hyperparameter are supplied to the model and validate every combination
as shown in Table 2. Finally, the values of the hyperparameters on which the model
results lowest training error are selected as the configuration for the final model.
One more important hyperparameter in RNN and LSTM is sequence_length also
called timesteps or input window size. This hyperparameter defines the number of
previous steps supplied into the model to predict the current step. The accuracy of
the model is greatly impacted by this hyperparameter, especially when dealing with
problems requiring dependency on historical data. Selecting the proper sequence
length is crucial to preventing overfitting and underfitting of the model.

Table 2 Hyperparameters and their selected values of RNN and LSTM


RNN LSTM
Hyperparameter Values Selected Hyperparameter Values Selected
value value
Units [50, 100,150] 100 Units [50, 100, 150
150]
Optimizer [‘adam’, adam Optimizer [‘adam’, rmsprop
‘rmsprop’] ‘rmsprop’]
Batch size [8, 16, 32, 16 Batch Size [8, 16, 32] 8
64]
Drop out [0.01, 0.02, 0.01 Drop out [0.01, 0.02, 0.01
0.3] 0.03]
Epochs [20, 30, 40, 40 Epochs [20, 30, 30
50] 40,50]
Window size [20, 25, 30] 20 Window size [20, 25, 30] 30
Future Price Prediction of IT Sector Companies Using Optimized Deep … 107

4.2 Prediction Results

This section presents the results achieved in the current study after applying two
DL approaches to the dataset of three IT companies. In this study, different sequence
lengths have been analyzed and it has been observed shorter window size shows better
results for short-term predictions or suits in case of highly changing patterns. In the
case of longer window size, no doubt the model captures long-term relationships but if
not properly selected it may lead to model overfitting and computationally expensive.
This is the main reason the proposed approaches are evaluated on different window
sizes and the best results achieved are shown in Table 3 and 4. The results illustrate
that the LSTM model shows the best results on window size 30 because LSTMs
particularly handle long sequences very effectively as compared to RNN due to its
architecture and use of gates, cell state, and gradient clipping.

5 Conclusion and Future Scope

The stock market experiences extreme volatility but on the other hand, it provides a
tremendous chance to increase wealth. Investors can accomplish this by consulting
graphs, charts, company balance sheets, and other resources. Alternatively, they
might delegate the task to a DL algorithm. The DL approaches can swiftly analyze
past data, trend lines, charts, etc., and recommend the best course of action for going
forward. DL is a breakthrough technology that has shown to be successful for thou-
sands of investors. In this study, two DL approaches with rigorous hyperparameter
tuning have been proposed to forecast the future price of three IT sector compa-
nies. The LSTM shows better results with longer window size as compared to RNN
because of the gated architecture in LSTM. RNNs can also be used in sequential
problems but LSTMs are observed more conducive to capturing trends and relation-
ships within the data. In future, there could be several directions: (i) adding more
predictors may increase prediction results, (ii) using a hybrid approach in conjunction
with both qualitative and quantitative variables will increase stock market prediction
efficiency, (iii) more recent approaches like Bi-LSTM and attention mechanism can
be used to make a more accurate prediction, and (iv) lastly, real-time prediction can
be made to assess the model’s effectiveness and practical application in live market
situations.
108

Table 3 Performance of LSTM on HCL, TECHM, and LTIM dataset


Window size 20 Window size 25 Window size 30
Evaluation metric → MAPE R2 RMSE MAE MAPE R2 RMSE MAE MAPE R2 RMSE MAE
Stock name↓
HCLTECH 3.14 0.86 57.57 44.23 3.06 0.86 57.93 42.73 2.70 0.89 56.81 38.45
TECHM 3.44 0.63 54.85 43.53 3.42 0.66 53.69 43.24 3.0 0.79 52.77 40.0
LTIIM 3.74 0.66 228.58 179.9 3.34 0.67 225.0 178.4 3.10 0.77 220.3 160.1
U. Bashir et al.
Table 4 Performance of RNN on HCL, TECHM, and LTIM dataset
Window size 20 Window size 25 Window size 30
Evaluation metric → MAPE R2 RMSE MAE MAPE R2 RMSE MAE MAPE R2 RMSE MAE
Stock name↓
HCLTECH 2.98 0.87 54.10 40.20 3.06 0.86 57.09 42.70 3.20 0.82 57.53 42.80
TECHM 3.10 0.79 54.63 42.83 3.42 0.61 56.02 43.45 3.85 0.61 58.74 45.31
LTIM 3.31 0.75 215.68 172.49 3.25 0.71 225.88 177.56 3.41 0.67 227.83 184.64
Future Price Prediction of IT Sector Companies Using Optimized Deep …
109
110 U. Bashir et al.

References

1. Conejo AJ, Plazas MA, Espínola R, Molina AB (2005) Day-ahead electricity price forecasting
using the wavelet transform and ARIMA models. IEEE Trans Power Syst 20(2):1035–1042.
[Link]
2. Stolojescu C, Cusnir A, Moga S, Isar A (2009) Forecasting WiMAX BS traffic by statistical
processing in the wavelet domain. In: 2009 International sympoism signals, circuits system
ISSCS 2009, pp 3–6. [Link]
3. Nava N, Di Matteo T, Aste T (2018) Financial time series forecasting using empirical mode
decomposition and support vector regression. Risks 6(1):1–21. [Link]
010007
4. Rout AK, Dash PK, Dash R, Bisoi R (2017) Forecasting financial time series using a low
complexity recurrent neural network and evolutionary learning approach. J King Saud Uni
Comput Inf Sci 29(4): 536–552. [Link]
5. Chopra S, Yadav D, Chopra AN (2019) Artificial Neural Networks Based Indian Stock Market
Price Prediction : Before and After Demonetization. Int. J. Swarm Intell. Evol. Comput. 8(1):1–
7. [Link]
6. Joseph E (2019) Forecast on close stock market prediction using support vector machine (SVM).
Int J Eng Res V8(02):37–43. [Link]
7. Yuan X, Yuan J, Jiang T, Ain QU (2020) Integrated long-term stock selection models based
on feature selection and machine learning algorithms for China stock market. IEEE Access
8:22672–22685. [Link]
8. st Kadam O, nd Gandhi K, rd Rana D (2020) A modern approach towards stock prediction
over traditional Methods. Int J Eng Res Technol 8(5):1–3. [Link]
9. Thakkar A, Chaudhari K (2020) CREST: cross-reference to exchange-based stock trend predic-
tion using long short-term memory. Procedia Comput Sci 167(2019):616–625. [Link]
10.1016/[Link].2020.03.328
10. Nabipour M, Nayyeri P, Jabani H, Mosavi A, Salwana E, Shahab S (2020) Deep learning for
stock market prediction. Entropy 22(8). [Link]
11. Chatterjee A, Bhowmick H, Sen J (2021) Stock price prediction using time series, econo-
metric, machine learning, and deep learning models. In: 2021 IEEE Mysore Sub Sect interna-
tional conference MysuruCon, pp 289–296. [Link]
9641610
12. Chhajer P, Shah M, Kshirsagar A (2022) The applications of artificial neural networks, support
vector machines, and long–short term memory for stock market prediction. Decis Anal J
2:100015. [Link]
13. Liu Q, Tao Z, Tse Y, Wang C (2022) Stock market prediction with deep learning: The case of
China. Financ Res Lett 46. [Link]
14. Ramotra AK, Mansotra V (2022) Feature raking and stacked sparse autoencoder based frame-
work for the prediction of breast cancer. Int J Eng Trends Technol 70(5):103–110. [Link]
org/10.14445/22315381/IJETT-V70I5P213
15. Mohamed Ashik A, Senthamarai Kannan K (2019) Time series model for stock price forecasting
in India. Logistics, supply chain and financial predictive analytics, pp 221–231. [Link]
10.1007/978-981-13-0872-7_17
16. Hochreiter S, Schmidhuber J (1997) Long short term memory. Neural Comput 9(8):1735–1780.
[Link]
17. Singh K, Mahajan A, Mansotra V (2023) Hybrid CNN-LSTM model combined with feature
selection and SMOTE for detection of network attacks. Int J Sens Networks 43(4):208–222
Future Price Prediction of IT Sector Companies Using Optimized Deep … 111

18. Borovkova S, Tsiamas I (2019) An ensemble of LSTM neural networks for high-frequency
stock market classification. J Forecast 38(6):600–619. [Link]
19. Lu W, Li J, Wang J, Qin L (2021) A CNN-BiLSTM-AM method for stock price prediction.
Neural Comput Appl 33(10):4741–4753. [Link]
20. Flores BE (1986) A Pragmatic view of accuracy measurement in forecasting. OMEGA Int J
Manag Sci 14(2):93–98
21. Maçaira PM, Cyrino Oliveira FL (2016) Another look at [Link] forecast accuracy. Int J
Energy Stat 4(2):1650008. [Link]
Stacked Spatio-Temporal Fusion
Network (SSTFN) Machine
Learning-Based Traffic Prediction

Namrata Shrivastava and Jitendra Agrawal

Abstract The operation of any intelligent transport systems requires efficient traffic
forecasting, and hence, management of the system traffic and navigation of environ-
mental terrains is improved. The availability of all the traffic environment information
is of large extent. Spatio-temporal traffic data is dynamic and raises a lot of issues.
This work is a model that combines the strengths of temporal convolutional networks
and diffusion convolutional recurrent networks within a single architecture. TCN
makes use of dilated convolution to add long time series, while DCRN performs
a graph diffusion process allowing one to generate traffic distributions across the
entire network at any instant of time. In contrast to such traditional approaches as
LSTM and ARIMA, our hybrid SSTFN model is capable of producing results that
are significantly better of all aspects of both short-term and long-term predictions
made, based on the performance measures that were further compared with real-life
data.

Keywords TCN · Temporal convolutional networks · DCRN · Diffusion


convolutional recurrent networks · MAE · Mean absolute error · RMSE · Root
mean squared error · First section

1 Introduction

More vehicles on the road have led to the problems of traffic jams and pollution
becoming worse, making it more difficult to move around cities. The intelligent
traffic system (ITS)’s objective is to lessen these problems by managing the traffic

N. Shrivastava (B)
CSE, UTD, RGPV Bhopal, Bhopal, India
e-mail: namrata0211@[Link]
J. Agrawal
UTD, RGPV Bhopal, Bhopal, India
e-mail: jitendra@[Link]

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 113
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
114 N. Shrivastava and J. Agrawal

Fig. 1 Traffic prediction


process Traffic Data (I/P)

Data Pre-Processing

Spatial Layer (GCN)

Temporal Layer: RNN/TCN

Fully Connected Layer

Prediction (o/p)

and control who goes where by making accurate traffic flow prediction as one of its
machines’ workings.
The prediction of short-term traffic flow means forecasting the levels of traffic prior
to that period based on historical trends. The earlier approaches had primarily focused
on statistical frameworks which included amongst others Kalman filtering, exponen-
tial smoothing, and autoregressive integrated moving averages. Thereafter, machine
learning algorithms that include support vector regression, k-nearest neighbor, and
Bayesian networks rose. In addition, more advanced deep learning models like long
short-term memory networks (LSTMs) and gated recurrent units (GRUs) have also
been developed making it clear that LSTM has a better capability of capturing the
nonlinear traffic flows than the ordinary neural networks [1–3, 5–8].
Conventional techniques often fail to capture the complicated nonlinear relation-
ships present in traffic flow data, such as its flow, speed, and density levels. In the
same way, Du et al. used ConvLSTM to model the spatial aspect and LSTM to model
the temporal aspects. Yan et al. created a CNN model for the spatial aspects of the
road as a two-dimensional roadmap prior to LSTM for the time series component
which displaced the previous models in performance output [4, 9–11].
The content of the manuscript generalizes the following organization sections:
Related work, proposed methodology, experiments and result analysis, conclusions,
and references.
There is Fig. 1 represents traffic prediction process working flow as below:

1.1 Traffic Prediction Model Validation Matrices

To validate any model, the following matrices needed to analyze:


Mean absolute error (MAE): Mean absolute error denotes the average amount of
erroneousness found between forecasted values against values that were actually
observed. It is understood in this context that it is already known how to measure the
average value of absolute deviations of the predicted values as well as actual values
over the entire dataset.
Stacked Spatio-Temporal Fusion Network (SSTFN) Machine … 115

1   
n
MAE = Yi − Y  i
n i=1

Symmetric Mean Absolute Percentage Error (sMAPE):


This particular measure was developed in the first place to measure the effectiveness
of the forecasts in a general sense and more particularly to the time series forecasts
apparent from the literature. It determines the percentage error, both under and over
the estimated values, between the actual and predicted figures.

1
n
MAPE = 100 ∗ Y  i − Yi/Yi − Y i2
n i=1

Mean Absolute Percentage Error (MAPE):


This measure is used to measure the effectiveness of prediction in a forecasting
model, with specific regard to the temporal nature of the data. It estimates the mean
percentage difference between the output of the value according to the model and
the given observation. Hence, it helps to assess the performance of the forecasts by
the actual values, thus giving a measure of performance of the model.
    
MAPE = n i = 1n |yi|yi − yî × 100%

Root Mean Square Error (RMSE):


This metric assesses how much on average the predicted values deviate from the
actual ones. However, it is not like the mean absolute error (MAE), because it places
more weight on the larger errors. This is accomplished by first calculating the sum of
the squared differences, then the average, and finally the square root. In this manner,
large errors carry more weight in the calculation.

 n

RMSE = ( |Y`i − Yi|ˆ2)/n
i=1

False Positive Rate (FPR):


It is the ratio of the number of negative samples that have been wrongfully predicted
as positive by the model to the total number of negative samples in reality.

FP
FPR =
(TN + FP  )
116 N. Shrivastava and J. Agrawal

Detection Rate (DR):


It tells the share of actual positive cases which were successfully marked as positive
by the model, defining how well the model performs in positive case detection of
true positives.

FP
FPR =
(TN + FP  )

where:
Yi: Predicted value, Y  i : Original value, FN: False negative, FP: False positive, TP:
True positive , TN: True negative, and n: number of instances.

2 Related Work

A vast body of knowledge has been accumulated in this sphere over the years. Some of
them mentioned here as their simplified conceptual building into the proposed model.
In 2022, the research work with the title “MGAT: Spatio-temporal multi-head graph
attention network for Traffic forecasting” was published which introduces ST-MGAT
model which is applicable to traffic characteristics in which the interplay of several
factors is very complex and nonlinear. Experimental results show that the above-
mentioned ST-MGAT brought significant enhancements as the RMSE values were
reduced by 2.34–22.47% while the MAE values reduced by 7.42% to 26.14% [6].
The same year study an innovative model is presented, termed the Spatial–Temporal
Up sampling Graph Convolutional Network (STUGCN), Experiments carried out on
actual databases have proven that the performance metrics of STUGCN remained
higher than the respective metrics of other traffic prediction models in prevalence [7,
8]. In 2023, “Spatial–Temporal Dynamic Graph Convolutional Neural Network for
Traffic Prediction” went on to introduce deep learning methods including spatio-
temporal dynamic graph convolutional neural network for traffic predictions. It
connects GCN and GRU using a time-varying adjacency matrix to encode spatio-
temporal information. Testing on two datasets confirms its effectiveness in enhancing
traffic prediction [9, 10]. In Latin American Transport Studies 2023, the proposed
work accounts for spatio-temporal dynamics in traffic forecasting using space–time
autoregressive integrated moving average (STARIMA) model. Experimental anal-
ysis which was carried out in Sao Paulo in Brazil shows that use of asymmetric
spatial matrices that consider both upstream and downstream movements enhances
the predictive performance in 77.3% of the instances. [11]. In 2024, “Enhancing
road traffic flow prediction with improved deep learning using wavelet transforms”
presented a deep learning framework which incorporates wavelet-based denoising
with RNNs for accurate forecasting of traffic flow. This model even higher predic-
tion efficiency was obtained with LSTM models averaging at R2 of 0.982 and 0.9811
[12]. In 2024 another “MuSeFFF: Multi-stage feature fusion framework for traffic
Stacked Spatio-Temporal Fusion Network (SSTFN) Machine … 117

Table 1 Models comparison


Model type Spatial Temporal Complexity Scalability Applications
dependency dependency
handling handling
LSTM with Moderate Strong Medium Medium Highway/
Spatial Features arterial traffic
flow
prediction
ST-GNN (Graph Strong Strong High High City-wide
Neural prediction,
Networks) congestion
forecasting
CNN (Traffic Moderate Moderate Medium Medium Traffic
Grids) (Local) density,
regional
forecasting
GC-LSTM Strong Strong High High Urban traffic
(Graph Conv + prediction,
LSTM) real-time
forecasting
STARIMA Moderate Strong Medium Low Small-scale
traffic
networks,
short-term
flow

prediction”, study brings multi-stage feature fusion framework called MuSeFFF that
is effective in traffic prediction in terms of spatial, temporal, and spatio-temporal
aspects. As each feature is fused together at each parallel anchor in MuSeFFF, the
issues of inter-locations and inter-times dependencies are lessened. Multiple experi-
ments on the PeMS-BAY dataset prove the advantage of the model, with the 1.25%
m enhancement in medium-term prediction and 2.76% m enhancement in long-term
prediction compared to the existing techniques [13–16].
A comparative analysis given below in Table 1 for different popular models.

3 Proposed Methodology

The analysis of existing models has led to the development of a new hybrid model: a
stacking strategy for hybrid prediction is applied. The proposed SSTFN model adopts
the following sequential processes which will be explained conceptually as well as
mathematically: Traffic dynamics can also be regarded as a network represented by
a graph with the following explanation. Each and every component of the graph
is a road. Road junctions are individual nodes, while road connections are edges.
118 N. Shrivastava and J. Agrawal

The graph attention network is adopted to illustrate the spatial relationships of the
aforementioned route sections.
1. Mathematical Formulation:
Input Graph: Let the graph be represented as G = (V,E)G = (V, E)G = (V,E),
Where: VVV is the set of nodes (road segments).
EEE is the set of edges (connections between road segments).
Node Features:

Each node v ∈ Vv \in Vv ∈ V is associated with a feature vector hv


∈ RFh_v \in \mathbb{R}Fhv
∈ RF, where FFF is the number of features (e.g., traffic flow, speed).

Attention Mechanism:
The attention coefficient between nodes iii and jjj is given by:
 
eij = LeakyReLU aT Whi||Whj e_{ij}
= \text{LeakyReLU }(aT [Wh_i||Wh_j])eij
 
= LeakyReLU aT Whi|Whj

where: WWW is a learnable weight matrix.


|| || || Denotes concatenation of the node features. And aaa is the attention vector.
eije_{ij}eij is the unnormalized attention score between nodes iii and jjj.
Normalized Attention Weights:
The attention coefficients are normalized across all neighboring nodes using the
softmax function:

αij = exp(eij) k ∈ N(i)exp(eik)\alpha_{ij}
= \frac{\exp(e_{ij})}{\sum_{k \in N(i)} \exp(e_{ik})}αij

= k ∈ N(i)exp(eik)exp(eij)

where: N(i)N(i)N(i) is the set of neighbors of node iii.


The feature update for node iii is computed as a weighted sum of the neighbors’
features:

hi = σ j ∈ N(i)αijWhj h_i
= \sigma \left(\sum_{j \in N(i)}\alpha_{ij} W h_j \right)

hi = σ j ∈ N(i) αijWhj

where: σ \sigmaσ is an activation function(e.g., ReLU).


Stacked Spatio-Temporal Fusion Network (SSTFN) Machine … 119

2. Temporal Attention Mechanism for Time Dependencies


To capture temporal dependencies, we use a temporal attention mechanism. Traffic
data tends to have patterns based on time intervals (e.g., morning and evening rush
hours). 
Here taken input sequence: Let Xt = {x1, x2, . . . , xT}Xt = \ x1 , x2 , . . . , xT\
Xt = {x1, x2, ..., xT} be the traffic data for a particular road segment over a period
of time TTT.
Temporal Attention Score:
For each time step ttt, the attention mechanism computes a score based on the
importance of that time step relative to the other steps in the sequence:

et = tanh(Wtht)e_t = \text{tanh}(W_t h_t)et = tanh(Wtht)

where: WtW_tWt is a learnable weight matrix for the temporal data.


hth_tht is the hidden representation of the time step ttt.
Attention Weights:
The attention weights across all time steps are computed using the softmax
function:

αt = exp(et) k = 1Texp(ek)\alpha_t
= \frac{\exp(e_t)}{\sum_{k = 1}{T}\exp(e_k)}αt

= k = 1Texp(ek)exp(et)

where αt\alpha_tαt is the attention weight assigned to time step ttt.


Final Temporal Representation:
The final temporal representation is a weighted sum of the time step representa-
tions:

htemp = t = 1T αthth_{\text{temp}}
= \sum_{t = 1}{T}\alpha_t h_thtemp

=t=1 T αtht

This captures the most relevant time steps for prediction, focusing on peak times
(e.g., rush hours).
3. Hybrid Spatio-Temporal Model (HSTAN)
The final prediction model integrates both the spatial GAT and temporal attention
mechanisms.
120 N. Shrivastava and J. Agrawal

Final Computational Model:


Spatial GAT Layer:
For each road segment v ∈ Vv \in Vv ∈ V, we compute the updated node
representation hv h_v hv using the GAT layer:

hv = σ u ∈ N (v)αvuWhu h_v
= \sigma\left(\sum_{u\inN (v)}\alpha_{vu}Wh_u\right)

hv = σ u ∈ N (v) αvuWhu

Look at Fig. 2. A distinctive nine-layer Spatial-Spatial–Temporal Fusion Network


(SSTFN) with respect to traffic data is proposed. Our approach begins with significant
feature engineering as we try to extract useful characteristics from the raw traffic
dataset, and then, we follow with careful data preparation to ensure that the input
data are sound and of good quality. The dataset is thoroughly divided into training and
testing parts to allow for a wide range of evaluation techniques to be employed. The
training set is referred to technique that is utilized in traffic forecasting model that
takes into account traffic dynamics and complex spatial and temporal traffic patterns.
Different models’ predictions are combined using the stacking process to increase the
accuracy and reliability of the resulting predictions. The effectiveness of such a model
is determined using mean squared error (MSE) metrics of the model predictions,
which allows measurement of prediction performance. This in turn improved the
traffic forecasting techniques and paved way for the use of predictive analytics in
intelligent transport system for further research.

4 Experiments and Result Analysis

After implementing the SSTFN model, it is being compared with existing models
accuracy as given below Graph 1 and found improved results:
Graphs 2, 3 and 4 predicts traffic speeds for a city in short (15 min), medium
(30 min), and long (60 min) intervals that are, respectively, based on the two models.
In this case, mean absolute error (MAE) measured in miles per hour is used for
comparison, where lower values are more preferable.
Stacked Spatio-Temporal Fusion Network (SSTFN) Machine … 121

Traffic Data Include Traffic volume, timestamp, and Weather

Creating new features as :


Feature Engineering Lagged Variable Temporal Feature: Hour of day, Day of week.
(Temporal Features/Spatial Features) Logged variable : Previous timestamps, Traffic Volume,
Spatial Features : Geographical info
Data Pre-processing
Normalizing data to ensure all Features on same

Split Data Dividing data set into Training Set (80%) and Testing Set
(20%)
Training Set
Build Model for Prediction
LSTM, TCN, RNN: Capture temporal dependence
(Random Forest Regression) and
RFR: Capture Complex interaction and nonlinear relation-
( LSTM with RNN Deep Learning)
ship in DCRNN.

Combine Prediction Averaging / Stacking After Training Above Model Combine their prediction us-
ing averaging/Stacking to improve accuracy

Combined Predictions
Calculate Mean Standard Error to evaluate per-
Evaluate Model (on MSE)

Fig. 2 Proposed model: SSTFN

Graph 1 Accuracy comparison

5 Conclusions

To sum up, it is imperative that effective traffic forecasting methods are developed in
order to enhance the operations of intelligent transportation systems. The architecture
that we propose in this paper which consists of the hybrid of temporal convolutional
122 N. Shrivastava and J. Agrawal

Graph 2 Short-term MAE

networks (TCN) and diffusion convolutional recurrent networks (DCRN) approaches


far deep understanding of spatio-temporal traffic data. TCN captures the long-range
temporality very well; on the other hand, DCRN does well in capturing the dynamic
spatial relations, hence creating a robust predictive model for traffic forecasting.
It is illustrated in the detailed experiments performed over known challenges that
predictive accuracy improved, which in turn leads to decrease of mean absolute error
(MAE) and root mean squared error (RMSE) metrics on the tested cases. The present
work advances traffic prediction techniques and gives an important perspective on
how to reproduce these techniques in practice, especially in real-time traffic control,
traffic congestion reduction, and route traffic management in urban space. In future
directions, the modification of the models will be further studied as well as their
possible use in smart city framework for the development of green transportation
alternatives.
Stacked Spatio-Temporal Fusion Network (SSTFN) Machine … 123

Graph 3 Medium-term
MAE

Graph 4 Long-term MAE


124 N. Shrivastava and J. Agrawal

References

1. Yu B, Yin H, Zhu Z (2018) Spatio-temporal graph convolutional networks: a deep learning


framework for traffic forecasting. In: twenty-seventh international joint conference on artificial
intelligence (IJCAI-18)
2. Diao Z, Wang X, Zhang D, Liu Y, Xie K, He S (2019) Dynamic spatial-temporal graph convolu-
tional neural networks for traffic forecasting. In: The thirty-third AAAI conference on artificial
intelligence (AAAI-19)
3. Wu Z, Pan S, Long G, Jiang J, Zhang C (2019) Graph wavenet for deep spatial-temporal graph
modeling. In: twenty-eighth international joint conference on artificial intelligence (IJCAI-19)
4. Gong Y, Li Z, Zhang J, Liu W, Yi J (2020) Potential passenger flow prediction:a novel study for
urban transportation development. In: the thirty-fourth aaai conference on artificial intelligence
(AAAI-20)
5. Shi X, Qi H, Shen Y, Wu G, Yin B (2020) Member, IEEE, a spatial–temporal attention approach
for traffic prediction. In: IEEE transactions on intelligent transportation systems
6. Wang B, Wang J (2022) ST-“MGAT: spatio-temporal multi-head graph attention network for
Traffic prediction. In Elsevier
7. Zhang S, Liu Y, Xiao Y, He R (2022) Spatial-temporal upsampling graph convolutional network
for daily long-term traffic speed prediction. J King Saud Univ–Computer Inf Sci
8. Ye J, Xue S, Jiang A (2022) Attention-based spatio-temporal graph convolutional network
considering external factors for multi-step traffic flow prediction. J Digit Commun Netw
9. Xiao W, Wang X (2023) Spatial-temporal dynamic graph convolutional neural network for
traffic prediction. IEEE Access
10. Gao Y, Chen D, Zhang J, Chen T (2023) A method of vehicle interactive information drive
speed prediction based on temporal dynamic graph convolutional neural network. IFAC
PapersOnLine
11. de Barrosa OM, Martea CL, Islera CA, Yoshiokab LR, da Fonseca Juniora ES (2023) Spatial
matrices for short-term traffic forecasting based on time series. Lat Am Transp Stud
12. Harrou F, Zeroual A, Kadri F, Sun Y (2024) Enhancing road traffic flow prediction with
improved deep learning using wavelet transforms. Results Eng
13. Kumar A, Sunitha R (2024) MuSeFFF: multi-stage feature fusion framework for traffic
prediction. Intell Syst Appl
14. Zeb A, Zhang S, Wei X, Yu JJ (2024) A generalized feature projection scheme for multi-step
traffic forecasting. Expert Syst Appl
15. Dong X, Zhao W, Han H, Zhu Z, Zhang H (2024) MTESformer: multi-scale temporal and
enhancespatial transformer for traffic flow prediction. IEEE Access
16. Ahmed SF Kuldeep SA Rafa SJ Fazal J, Hoque M, Liu G, Gandomi AH (2024) Enhance-
ment of traffic forecasting through graph neural network-based information fusion techniques.
Information Fusion
17. Ata KIM, Hassan MK, Ismaeel AG, Al-Haddad SAR, Alquthami T, Alani S (2024) A multi-
Layer CNN-GRUSKIP model based on transformer for spatial-TEMPORAL traffic flow
prediction. Ain Shams Eng J
Improving the Efficiency of EMG-Based
Prosthetic Arm Using EEG

H. B. Divyashree, Jaishree R. Devaru, Maitri Kulkarni,


and Muddireddy Devaghneswara Reddy

Abstract The objective of this project is to create a prosthetic hand with five fingers,
specifically for people who have had their hands severed. Users may operate the
device as naturally as they would a biological arm thanks to its simple, flexible,
and ideal control mechanism. The ability to operate each finger and limb indepen-
dently enables precise positioning for a range of tasks. The complicated anatomy
of the human hand served as the idea for the mechanical hardware, which utilizes a
double-connected revolute joint mechanism. In addition to this mechanism, a liga-
ment system and feedback network allow the prosthetic hand to adapt to the form and
arrangement of items, greatly improving the user’s engagement with their surround-
ings. To ensure smooth and responsive movements, the prosthetic model includes
integrated force sensors and servomotors for finger actuation. The user is supposed
to wear the complete unit firmly over their shoulder. This prosthetic’s ability to
decode inputs from EEG data and allow users to operate the hand with their thoughts
distinguishes it. This user-friendly interface improves users’ capacity to perceive and
manage objects in a natural and effective manner, thereby improving their quality of
life.

Keywords EEG (Electroencephalography) · Electrodes · BCI-Brain computer


interface

1 Introduction

The biomechanics of prosthetics and orthotics differ. Prosthetics are artificial tools
used to replace missing body parts due to disease, injuries, or birth complications.
When a person loses a limb, their emotional and financial lives must change dramati-
cally. Amputees require prosthetic equipment and services for the rest of their lives. A

H. B. Divyashree (B) · J. R. Devaru · M. Kulkarni · M. D. Reddy


Electronics and Communication Engineering, Dayananda Sagar University, Bangalore, India
e-mail: divyabalachandra94@[Link]

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 125
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
126 H. B. Divyashree et al.

Fig. 1 Types of prosthesis of hand and leg

prosthesis is a fabricated extension that substitutes a lost portion of the body compo-
nent, such as the top or lower extremity areas. Mechanical devices that replace the
integrated actions of the human muscle, brain, skeleton, and nerve system are utilized
to help or restore motor control that has been lost due to disease, illness, or deformity.
A transtibial amputation is the surgical eliminating of a limb that is below the
knee. It frequently necessitates prosthetic equipment for mobility, which may affect
gait and balance. The goal of rehabilitation is to increase functionality in everyday
activities and help the patient adjust to the prosthetic. A transfemoral amputation is the
surgical eliminating of a limb that is above the knee. For movement, patients usually
need artificial legs. Strength, balance, and prosthesis adaptation are prioritized during
rehabilitation in order to maximize everyday activities and promote quality of life.
The surgical removal of a limb below the elbow is known as a transradial amputation.
Prosthetic devices are frequently used by people to restore their functioning. The
goals of rehabilitation are to increase grip strength, help the patient adjust to the
prosthetic, and improve daily activities and quality of life in general.
The surgical removal of a limb below the elbow is known as a transradial ampu-
tation. Prosthetic devices are frequently used by people to restore their functioning.
The goals of rehabilitation are to increase grip strength, help the patient adjust to the
prosthetic, and improve daily activities and quality of life in general.
We have studied and analyzed the grip energy dispersion among various prosthetic
hand layouts. The different human arm designs that fulfill different functional tasks
were studied. It is primarily concerned with enhancing the usability and controlla-
bility of the prosthetic arm [1]. Even experienced researchers frequently fall short in
providing adequate information and specifics on the guidelines, recording tools, and
techniques that other researchers rely on to reliably duplicate their studies [2].The
values from above papers are taken into consideration (Fig. 1).

2 Proposed Methodology

Arduino, a publicly available hardware and software platform, develops and builds
microcontroller-based kits that allow users to make interactive objects and techno-
logical devices capable of manipulating and sensing real-world items. The project
Improving the Efficiency of EMG-Based Prosthetic Arm Using EEG 127

Fig. 2 EEG data acquisition process

focuses on many microcontroller board designs that employ different microcon-


trollers and are manufactured by a variety of vendors. These boards provide a collec-
tion of analog and digital input/output (I/O) pins, making it simple to connect them
to expansion boards and other electrical devices. Coding from personal computers
is made easier by the serial communication interfaces included on many Arduino
models that have USB connectivity. For the purpose of programming the micro-
controllers, the Arduino community offers a user-friendly integrated development
environment (IDE), mainly for the C or C++ programming languages. A compiler
that translates code into binary machine code is included in the IDE.
The multiple elements that jointly make up the brain are each in the role of
directing a distinct bodily function. Single-electrode equipment is often employed in
EEG recording. These surface electrode signals are filtered by a band-reject filter and
amplified with a preamplifier. The P300 wave, an event-related potential reflecting
cognitive functions including attention and decision-making, is a significant compo-
nent in this context. The P300 can better communicate with prosthetic devices by
interpreting these signals and indicating the user’s desire. Users are then capable to
control objects in accordance with their ideas thanks to the processing of the signals,
which drives motors for prosthetic hand movements. An important development in
adaptive technology, this integration of brain activity into prosthesis control improves
both functionality and user experience (Fig. 2).

3 Working Principle

Brain waves are patterns formed by the small electrical signals produced by the
millions of nerve cells that comprise the human brain. To investigate human brain
function, we process these patterns using electrodes from an electroencephalo-
gram, or EEG. These electrodes track bioelectrical signals from the brain and
acquire frequencies associated with distinct cognitive states. These signals are orig-
inally passed via a band-rejection screening to reduce unimportant interference and
128 H. B. Divyashree et al.

Fig. 3 Suggested block layout

noise. The cleaned signal is then amplified and sent through a bandpass filter with
a frequency variation of 4–7 Hz, known as theta waves—which are commonly
associated with creativity and relaxation.
Indication of light sources shows the state of communication between the headgear
and the autonomous arm. Blue color led starts flashing when the head gear Bluetooth
module is paired with the arm Bluetooth component. Green led starts flashing when
the information transmission has stared. When the supplied commands are incorrect
the Bluetooth module starts getting information from the headgear. This information
is compared to predefined activity data (attention and meditation).If the information
matches the predefined data, the controller controls the arm activity, accordingly
(Fig. 3).

3.1 Proteus

In the field of electronics engineering, Proteus is a strong simulation and PCB design
application that is widely utilized. Before putting a circuit into action, users can verify
the quality of the connections and the behavior of the output by using the design and
simulation functions. This ability is crucial for learning possible problems early in the
design phase. In addition, by utilizing this simulation software in conjunction with
Arduino software, users can upload their programs, enabling a smooth workflow
Improving the Efficiency of EMG-Based Prosthetic Arm Using EEG 129

Fig. 4 Simulation utilizing Proteus

from design to practical usage. Because of this, Proteus is a great product for both
engineers and consumers.
The prosthetic hand can be operated in two ways: manually or with a sensor. The
electromyography sensors are placed on different areas of the biceps and forearm to
read the electromyography signals. The direction and location of the muscle sensor
electrodes have a substantial impact on signal strength, and therefore, electrode place-
ment was carefully considered in both the theoretical geometry of human muscle and
empirical assessments. The electrodes must always be positioned in the center of the
muscular body, opposite to the passage of the muscle fibers. The EEG signal is sent
from the head to the arm via a Bluetooth module.
We use two servomotors to move the moving parts of hands (Fig. 4).

3.2 Electrode Placements

The international 10–20 electrode placement method is frequently used for reliable
EEG assessments. This standardized procedure ensures uniform electrode placement
across individuals. As an instance, electrodes can be put at the FP1 position using a
brainwave sensor, allowing for accurate monitoring of brain activity (Fig. 5).
130 H. B. Divyashree et al.

Fig. 5 Positions for placing electrodes

3.3 Servomotor

A servomotor is a type of straight or circular actuator that allows precise adjustment


of speed, acceleration, and other parameters and circular linear position. Its accuracy
in monitoring and adjustment are made possible by the combination of a suitable
motor and a position feedback sensor. A servomotor requires a reasonably complex
controller to function properly; this controller is frequently a specialized module
made just for servomotor interfaces. Despite being a widely used term, “servomotor”
refers to motors that are appropriate for use in closed-loop control systems rather
than a particular sort of motor.
These systems supply precise output by continuously adjusting the motor’s oper-
ation based on feedback. Servomotors are widely used in many different fields,
including as automation, computerized numerical control (CNC) machinery, auto-
mated manufacturing, and even consumer electronics like toy and cameras. They
are essential parts of modern automation and technology because of their ability
to deliver precise positioning and control, which raises productivity and accuracy
across a wide range of industries.
Improving the Efficiency of EMG-Based Prosthetic Arm Using EEG 131

Fig. 6 Arm mechanism modeling left view

4 Development of Prosthetic Hand

4.1 Cero

CREO is a potent simulation plan that is widely used in the mechanical industry
for a variety of design and engineering purposes. It supports extensive 3D CAD
capabilities such as parametric feature solid modeling, 3D direct modeling, and the
production of 2D orthographic views. CREO also provides finite element analysis
and simulation, enabling users to test and validate designs (Fig. 6).
The software also allows for technical schematic design graphics and improved
viewing and visualization skills, making it an indispensable tool for engineers and
designers building creative products and solutions. We have painstakingly created a
prosthetic arm model using Creo modeling software, adding exact proportions and
dimensions that meet our unique needs. Essential characteristics including fingers, a
thumb joint, a bottom palm, and a top palm are all included in the design. In order to
guarantee the prosthetic hand’s comfort and functionality, certain specifications are
essential. We were able to construct a prosthetic hand that satisfies both practical and
cosmetic requirements since all of these outcomes were carefully taken into account
during the design and fabrication process (Figs. 7 and 8).
132 H. B. Divyashree et al.

Fig.7 3D printed prosthetic


arm containing the Arduino,
servomotors, EMG sensor,
and electrodes

Fig. 8 Testing phase

5 Testing and Validation

6 Conclusion and Future Scope

Biotechnology regularly uses EEG, a simple method of monitoring electrical activity


in the brain, to monitor brain activity. It is used in many different technologies,
including biofeedback and neurofeedback, psychological research, and cognitive
process analysis. EEG can be used in brain–computer interfaces, human–computer
interaction, and marketing research since it offers insights into how the brain works
and behaves (Fig. 9).
The fabrication of a prosthetic hand that makes use of modern technologies
including EEG signals and servomotors constitutes a huge step forward in amputee
assistive technology. We designed a practical model by combining sophisticated
designs influenced by human anatomy with current tools such as Creo. This allows
Improving the Efficiency of EMG-Based Prosthetic Arm Using EEG 133

Fig. 9 Data of EEG usage in different fields

users to handle the prosthetic hand easily and effectively. The employment of a
brainwave sensor to enable movement via cognitive instructions improves the user
experience and gives a more natural method to interact with their surroundings. This
project intends to improve amputees’ entire quality of life through innovative design
and technology, in addition to meeting their physical demands.
There are several opportunities for further development and improvement of this
prosthetic hand model. Future iterations could investigate the application of machine
learning methods to enhance the hand’s abilities to adapt and respond based on user
behavior and preferences. Furthermore, boosting the range of motion and using
materials that imitate the feel and flexibility of human skin could enhance both the
prosthetic’s aesthetic and functional appeal. Wireless energy transfer for powering
prosthetics, as well as advances in component miniaturization, could result in lighter,
more compact designs. Finally, ongoing participation with consumers and medical
experts will be critical in developing these technologies and ensuring they suit the
changing demands of amputees but also seeks to improve their entire standard living
through revolutionary design and technology.

References

1. Nguyen T, Chang E (2020) The future of CERO in smart prosthetics: opportunities and
challenges. J Biomech
134 H. B. Divyashree et al.

2. Garcia M, Liu Y (2020) CERO innovations in prosthetic design: a review of current trends. J
Mech Eng
3. Mohammadi A, Lavranos J, Tan Y, Choong P, Oetomo D (2020) A paediatric 3D-printed soft
robotic hand prosthesis for children with upper limb loss. In: IEEE international conference of
the IEEE engineering in medicine biology society (EMBC)
4. Guerrero MC, Parada JS, Espitia HE (2021) EEG signal analysis using classification techniques:
logistic regression artificial neural networks support vector machines and convolutional neural
networks. Heliyon 7(6)
5. Musolf B, Ortiz-Catalan M (2021) Design of a bypass socket for transradial amputations. OSF
6. Käthner (2022) Integration of EEC and EEG technologies for enhanced neurofeedback
applications. Front Hum Neurosci
7. Ahamad S (2022) System architecture for brain-computer interface based on machine learning
and internet of things. Int J Adv Comput Sci Appl 13(3)
8. Mikhael A, Lin Y (2023) Comparative study of P300 and other ERP components in BCI
applications. J Neurosci Methods
9. Singh R, Patel S (2024) Wearable EEG technology: a review of current innovations and future
directions. Department of biomedical engineering
Analysis of Solar Panel Parameters
for Extracting Maximum Power using
Artificial Neural Network-Based
Training Algorithm

Jagdish Chandola , Sumit Pundir , Abhishek Sharma ,


and Sakshi Pundir

Abstract The process of discovering and acquiring energy is greatly reliant on


energy sources. The presence of solar energy renders it highly advantageous. The
seasonal variation in weather impacts the capacity of solar energy. There is a selection
of solar panels with different capacities, which depend on the behaviour of their
electrical and physical qualities. Consequently, alterations in the parameters of solar
panels also impact the quantity of solar energy collected, as the maximum power
point (MPP) fluctuates. Prior to incorporating a photovoltaic solar panel into an
energy harvesting system, it became imperative to ascertain its MPP to achieve the
most efficient power generation. Techniques such as incremental conductance (IC),
particle swarm optimization (PSO), perturb & observe (P&O), and artificial neural
network (ANN) assist in accurately identifying the MPP in photovoltaic (PV) solar
generation. The efficacy of training a PV solar panel MPP for output parameters
using various training algorithms based on ANN has been assessed. To show the
performance, Bayesian regularization (BR) and Levenberg–Marquardt algorithms
(LM) has been utilized. When employing the LM and BR techniques, there are
discernible disparities in the performances. The examination’s results, which indicate
a decline in performance, elucidate these performances.

Keywords ANN · Renewable energy · MPP · Photovoltaic solar panel · Neural


network tool box

J. Chandola · S. Pundir · A. Sharma (B)


Computer Science and Engineering Department, Graphic Era Deemed to be University,
Dehradun, India
e-mail: abhishek15491@[Link]
S. Pundir
Graphic Era Hill University, Dehradun, India

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 135
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
136 J. Chandola et al.

1 Introduction

Extreme practices of fossil fuel as primary energy resources have become affecting
factors for environment degradation and global warming. The Paris Climate Summit
(2015) said that the rise in global temperature should be kept to 1.5 °C [1]. At
present most of the energy consumption in the world is available through coal and
natural gasses and in a phase of using renewable energy as primary resources by
2050 [2]. This means a robust need to spread over renewable energy resources. Solar
energy has become a recognized and massive potential resource of energy with an
alternative to the pollution free environment. There are many other purposes where
the solar energy can be used, like transferring into fluid with the help of collectors
[3], and solar empowered refrigerator [4] and nano material coating of PV panel [5].
Different configuration and materials of PV panel offer diverse efficiency in term
of power generation [6]. In common, the panel power decreases with a decrease in
irradiation level. On the other hand, solar panel power declines with an increase in
ambient temperature [7], but environmental issues such as dust can have a substantial
impact on efficiency [8]. ANN are being used in various deep learning approach [9].
The objective here is to determine the PV panel performance for a single-diode PV
module with the use of ANN which helps in improving the efficiency of any PV
generation system.

2 State of the Art of Solar PV cell and MPP

The PV emulator design has been perceived an uplift from the approximation of basic
mathematical model to real-time implementation systems in current times [10]. A
PV panel is made up of many PV cells connected in parallel and series [11]. Figure 1
depicts the electrical circuit of the corresponding single-diode solar cell. MPP stands
for maximum power point of a PV panel respond in extracting highest possible power
from a PV module [12]. Values of current and voltage flow in a photovoltaic module
are primarily determined by temperature and the impact of irradiance level [13]. The
effect of irradiance is presented in Fig. 2, which represents the I-V and P-V curves of
a PV panel drawn in simulation environment. Figure 2 is showing different maximum
power value for different irradiance level [14]. Any PV module’s power output may
be computed directly using the current and voltage across the module and can be
improved by the ANN training algorithms.
Here, IL stands for light-generated current, ID stands for voltage-dependent current
lost to recombination, I denotes current across the output voltage, Rs denotes the
series resistance, and Rsh represents shunt resistances value.
Concerning the ANN technique in MPP estimation, the ANN topologies are of
different categories [16]. A MLP topology architectural neural network has been
trained, tested, and validated in performance generation phase for analysis and design.
The structure of MLP involves three layers, namely an input layer, a hidden layer,
Analysis of Solar Panel Parameters for Extracting Maximum Power … 137

Fig. 1 Electrical circuit of


the corresponding
single-diode solar cell [15]

Fig. 2 I-V and P-V


characteristics of PV module
at different irradiance and
constant temperature

and an output layer [17]. A process of weight multiplication, activation function, and
bias addition is applied to each neuron of all layers. The activation functions include
Sigmoid, ReLU, Tanh, and Softmax activation functions, etc. [18].
In this research work, the irradiance and temperature as input and power (P) as
output have been taken for network training. Various works that train the ANN for
MPPT are available. The MPP using a newly introduced expressions for single-diode
PV model has been calculated. The performances under fully and partially shaded
conditions were verified on distinct PV modules. Five parameters were employed, and
the unknown function was found using the Lambert W function [19]. An improved PV
panel characterization has been found by using the ANN for reference voltage regula-
tion, which would eventually be utilized in an MPPT system [20]. An ANN deployed
forecasting of a photovoltaic module under practical circumstances for extracting
power has been anticipated. The effectiveness of the ANN system was evaluated
against different parameter models. Two-layered ANN topologies outperformed all
analytical models in terms of output power and current prediction [21].
An ANN was utilized to forecast the power output of a photovoltaic module.
To compare the outcomes of mathematical equations with the MSE calculations,
138 J. Chandola et al.

several epochs, learning rates, and hidden layer configurations were performed. The
findings suggest that increased solar radiation causes a rise in power production
[22]. In another research, a generalized regression neural network (GRNN) model
for calculating the performances of distinct PV modules has been conducted. The
outcomes for various cell temperatures and irradiances were compared in the study.
I-V besides P-V curve for all study were also offered using MATLAB platform to
predict the performances of different modules [23].
Several studies utilizing ANN have investigated the forecasting of the amount of
solar radiation that a stationary solar module can absorb. The network’s efficacy for
this application was proved by metrics like as MSE and best-fit regression [24]. A
MATLAB simulation of the improved PV module performances and a comparison
analysis based on the findings have been completed [25]. A review was presented on
panel parameter estimation, illustrating the manufacturer’s perspective of characteri-
zation (panel model, panel efficiency, temperature coefficient, and number of bypass
diodes) [26]. Many methods have been employed to confirm parameter extraction in
single and double bypass diode panel models [27, 28].

3 Methodology

The methodology of the proposed work is illustrated in Fig. 3. The online NASA
Global Energy Resources Project served as the source of the data [28–30]. The
data collected contain irradiance and temperature values. The Imp , Vmp , and Pmp
equations define how the voltage, current, and power at MPP of a PV panel change
with irradiance (G) and temperature (T) in standard conditions:

Imp = Imps × (G/Gs ) × (1 + α × (T − Ts )) (1)

Vmp = Vmps + β × (T − Ts ) (2)

Pmp = Vmp × Imp (3)

where Imp = current at MPP, Imps = standard current at MPP, G = current irradiance,
Gs = standard irradiance, Ts = standard temperature, T = current module tempera-
ture, α = temperature coefficient (current, I), Vmp = voltage at MPP, Vmps = standard
voltage at MPP, β = temperature coefficient (voltage, V), and Pmp = power at MPP.
Feed forward multilayer perceptron (FF-MLP) topology of ANN model has been
trained with two different algorithms (LM and BR), further observing the PV panel’s
performance in term of regression, mean square error (MSE), and number of epochs
for settling MPP. Training, testing, and validation are done using MATLAB software
on collected data in the ratio of 80:10:10. To improve power generation efficiency,
the suggested PV module characteristics can be linked with MPPT systems, such as
controllers and converters.
Analysis of Solar Panel Parameters for Extracting Maximum Power … 139

Fig. 3 Steps in ANN-based


PV panel power extraction

Table 1 Solar panel parameter for experiment


Parameter Values Parameter Values
Short-circuit current, Isc 7.82 A Temperature coefficient 0.102 %/°C
(current), α
Maximum power point current, 7.32 A Temperature Coefficient −0.36099 %/°C
Imps (Voltage), β
Voltage on open-circuit, Voc 36.5 V Voltage at MPP, Vmps V

Equations (1) through (3) were used to compute current (I), voltage (V), and power
(P). Regression and MSE metrics were used to compare the network’s performance
to assess the PV system’s maximum power generation. Table 1 lists the PV panel
parameters that were employed in the experiment.

4 Results and Discussion

In the training face the network is adjusted to minimize error, validation generalizes
the network, and testing is used to show measures are independent of training during
and after. Some of the time the network can trapped to local minimum so, the
momentum parameter (Mu) is included to show convergence of the training algorithm.
The backpropagation was performed by ANN to adjust its predicted output and error
minimization.
Figures 4 and 5 represent the performance plots for LM algorithm. The value
of regression in plot (Fig. 4) is 1 with no error, which represents the strong corre-
lation between output and target power. Performance phase in Fig. 5 represents
the best validation performance at 1000 epoch. Zero error in training, testing, and
validation phase are observed with a minimum and maximum error of −0.118 and
140 J. Chandola et al.

0.05878 which validates the algorithm. The gradient, momentum parameter (Mu),
and validation check in training state have also obtained.
Figures 6 and 7 represent the performance plot for BR algorithm. Value of regres-
sion R is 1 (Fig. 6), an excellent sign of correlation between output and target power.
Zero error was observed during training and testing which validates the algorithm.
Figure 7 represents the best validation performance at epoch 1000. The gradient,
momentum parameter (Mu ), and validation check are also obtained for BR algorithm.
The performances of both techniques are summarized and compared in Table 2.
According to the simulation plots, the network is reaching 1000 epochs for both
algorithms while training. The gradient is 0.03832 and 0.0065142 at 1000 epochs,
representing less deviation for trained data by the BR algorithm. Regressing R = 1
for both algorithms represents an excellent correlation between the target and actual
output for training.

Fig. 4 LM algorithm-based regression plot of proposed ANN model


Analysis of Solar Panel Parameters for Extracting Maximum Power … 141

Fig. 5 LM algorithm performance test plot of proposed ANN model

Fig. 6 BR algorithm-based regression plot of proposed ANN model


142 J. Chandola et al.

Fig. 7 BR algorithm
performance test plot of
proposed ANN model

Table 2 Performance
Parameters LM algorithm BR algorithm
analysis and comparison of
used LM and BR algorithms Epoch 1000 1000
Gradient 0.03832 0.0065142
Regression 1 1
Performance 0.000006809 0.0000003524
Momentum parameter, Mu 0.0001 50000
Space complexity Ο(mn + n2 ) Ο(mn + n2 )
Time complexity Ο(k(mn2 + n3 )) Ο(k(mn2 + n3 ))

Performance values for LM and BR algorithms are 0.000006809 and


0.0000003524, which represent the LM algorithm fit for the present data. The value
of Mu for LM is 0.0001 and 50000 for the BR algorithm. A very small value of Mu
for the LM algorithm represents the applicability of the algorithm for the proposed
work.

5 Conclusion

The performance of a PV panel is influenced by its current and voltage characteristics,


which vary with changes in irradiance and temperature. Training the PV module
using an ANN structure with LM and BR algorithms demonstrates the effectiveness
of these techniques. Performance parameters validate the ANN’s capability for this
application, with regression values indicating a strong correlation between actual
and target power outputs. Actual power generation can be affected by losses or
enhancements when using a storage system or MPPT controller. Therefore, deploying
a predictive system for actual power generation is essential. Extending the use of ANN
Analysis of Solar Panel Parameters for Extracting Maximum Power … 143

to predict PV panel performance can enhance MPPT, ensuring the panel operates
at its maximum efficiency. This approach can be further developed by integrating it
into an MPPT-based power generation system, where controllers and converters play
a crucial role in optimizing power collection.
Future research could explore hybrid MPPT algorithms by integrating ANN with
methods like fuzzy logic or evolutionary algorithms for better tracking precision.
Incorporating real-time data from multiple locations could enhance ANN’s general-
ization, while combining the system with energy storage and grid platforms could
optimize smart grid performance. Advanced techniques like deep learning could
further improve prediction accuracy and real-time adaptability across diverse weather
conditions.

References

1. Ray SA (2024) Environmental sustainability vision report 2024: quantum machine learning
lab’s pledge to a greener future
2. Raimi D et al (2024) Global energy outlook 2024: peaks or plateaus. [Link]
lications
3. Nautiyal S et al (2021) Performance evaluation of flat plate collector using computational fluid
dynamics for climatic condition of Dehradun. Ilk Online 20(3):3508–3516
4. Mutyala R et al (2021) Experimental study on Peltier module-based compressor-less mini solar
powered refrigerator. Ilk Online 20(3):3391–3400
5. Mishra A, Bhatt N, Bajpai A (2019) Nanostructured superhydrophobic coatings for solar panel
applications. Nanomaterials-based coatings. Elsevier, pp 397–424
6. Ghazali AM, Rahman AMA (2012) The performance of three different solar panels for solar
electricity applying solar tracking device under the Malaysian climate condition. Energy
Environ Res 2(1):235
7. Karafil A, Ozbay H, Kesler M (2016) Temperature and solar radiation effects on photovoltaic
panel power. J New Results Sci 5:48–58
8. Rathod A, Mishra P, Mishra A (2023) Effect of dust accumulation on efficiency of solar panels
in clement town region (Dehradun) India: an empirical study. NanoWorld J 9(S1):S326–S330
9. Sharma S et al (2021) A review of neural machine translation based on deep learning techniques.
In: 2021 IEEE 8th Uttar Pradesh section international conference on electrical, electronics and
computer engineering (UPCON). IEEE
10. Ram JP et al (2018) Analysis on solar PV emulators: a review. Renew Sustain Energy Rev
81:149–160
11. Al-Ezzi AS, Ansari MNM (2022) Photovoltaic solar cells: a review. Appl Syst Innov 5(4):67
12. Ho BM, Chung HS, Lo W (2004) Use of system oscillation to locate the MPP of PV panels.
IEEE Power Electron Lett 2(1):1–5
13. Ibrahim H, Anani N (2017) Variations of PV module parameters with irradiance and
temperature. Energy Procedia 134:276–285
14. Goel S, Sharma R (2021) Analysis of measured and simulated performance of a grid-connected
PV system in eastern India. Environ Dev Sustain 23:451–476
15. Sharma A et al (2021) An effective method for parameter estimation of solar PV cell using
Grey-wolf optimization technique. Int J Math Eng Manag Sci 6(3):911
16. Lo Brano V, Ciulla G, Di Falco M (2014) Artificial neural networks to predict the power output
of a PV panel. Int J Photoenergy 2014(1):193083
17. Mandalapu S, Wang Y, Ni XS (2018) Building neural network model in base SAS® (From
Scratch) 2. In: Feed forward and backward propagation neural network, pp 1–20
144 J. Chandola et al.

18. Nwankpa C et al (2018) Activation functions: comparison of trends in practice and research
for deep learning. arXiv:1811.03378
19. Batzelis EI et al (2014) Direct MPP calculation in terms of the single-diode PV model
parameters. IEEE Trans Energy Convers 30(1):226–236
20. Bouchetob E, Nadji B (2022) Using ANN based MPPT controller to increase PV central perfor-
mance. In: 2022 2nd international conference on advanced electrical engineering (ICAEE).
IEEE
21. Karamirad M et al (2013) ANN based simulation and experimental verification of analytical
four-and five-parameters models of PV modules. Simul Model Pract Theory 34:86–98
22. Jumaat SA et al (2018) Prediction of photovoltaic (PV) output using artificial neutral network
(ANN) based on ambient factors. J Phys: Conf Ser. IOP Publishing
23. Jaber M et al (2022) Prediction model for the performance of different PV modules using
artificial neural networks. Appl Sci 12(7):3349
24. Jumaat SAB et al (2016) Investigate the photovoltaic (PV) module performance using artificial
neural network (ANN). In: 2016 IEEE conference on open systems (ICOS). IEEE
25. Jiang Y, Qahouq JAA, Batarseh I (2010) Improved solar PV cell Matlab simulation model and
comparison. In: 2010 IEEE international symposium on circuits and systems (ISCAS). IEEE
26. Jordehi AR (2016) Parameter estimation of solar photovoltaic (PV) cells: a review. Renew
Sustain Energy Rev 61:354–371
27. Gnetchejo PJ et al (2019) Important notes on parameter estimation of solar photovoltaic cell.
Energy Convers Manag 197:111870
28. Administration N.A.a.S., POWER Data Archive. 01-01-2021 to 31-12-2023, NASA
29. Sahu R, Sahu V (2024) An energy-efficient algorithm for resource allocation in H-CRAN (EE
H-CRAM) for 5G networks. Wirel Pers Commun 138(3):1483–1499
30. Singh SP, Pathak D, Kumar A, Sanjeevikumar P (2023) An optimized fractional order modified
adaptive variable step-size LMS control approach to enhance DVR performance. IEEE Trans
Consum Electron
A Review on Anomaly Detection Using
Machine Learning Techniques
for Unmanned Aerial Vehicles (UAVs)

Dhruv Aggarwal, Sharon Christa, Sarishma Dangi ,


and Aayush Shrivastava

Abstract Over the past few decades, automated aerial vehicles have been essential
in the upgradation of fifth-generation wireless communication. These aerial vehicles
provide us with many enhancements in various fields, including digital agriculture,
salvage detection, variable detection, cost-efficiency, versatility, and spectral effi-
ciency. At the same time, they are also facing some harsh challenges that is related to
security and low power management Additionally, UAVs equipped with sensors and
capturing devices, which capture large amounts of data and information, can contain
specific information where security is a major concern. Thus, the subset of artificial
intelligence, i.e., the use of machine learning algorithms like regression, reinforce-
ment learning, and deep reinforcement learning techniques can be used to optimize
and conserve the discharged power and improved security reasons. Anomaly detec-
tion techniques are vital in identifying and managing security threats. These methods
are very important for probing email communications to detect anomalous or suspi-
cious activities, which helps maintain endpoint security through robust encryption
measures that block potential threats directly. Additionally, anomaly detection is
key in analyzing network traffic to identify and prevent distributed denial-of-service
(DDoS) attacks, thereby bolstering the overall security framework. In this work, we
review the key research works done in the area of anomaly detection for UAVs.

Keywords Autonomous aerial vehicles · Unmanned aerial vehicles · Anomaly


detection · Wireless communication · Cyber security · Autonomous navigation ·
Surveillance · Monitoring · Machine learning

D. Aggarwal · S. Dangi (B)


Technology Business Incubator (TBI-GEU), Department of Computer Science and Engineering,
Graphic Era (Deemed to Be University), Dehradun, India
e-mail: sarishmasingh@[Link]
S. Christa
Department of Computer Science and Engineering, SoC MIT ADT University Maharashtra, Pune,
India
A. Shrivastava
Department of Computer Science and Engineering, IITM, IES University, Bhopal, India

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 145
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
146 D. Aggarwal et al.

1 Introduction

A significant rise in the development of advanced technologies such as artificial intel-


ligence, robotic automation, and blockchain have transformed the entire ecosystem
of business operation and development. Unmanned aerial vehicles (UAVs) are one of
the core pillars that will find the integration in different domains over the upcoming
decades and are analyzed to have a growth rate of 11.7% from 2020 to 2030 as illus-
trated in Fig. 1 [1, 2]. Over the last decade, UAVs primarily found used in surveillance,
especially in the field of aerial surveying, digital agriculture, recreational photog-
raphy and videography, aerial mapping and surveying, wildlife conservation, search
and rescue, etc., enabling the gathering and collection of raw data and information
from elevated edge points. This ability of collecting information from elevated points
has significantly contributed to geo-visualization by generating maps from the raw
photogrammetric data captured by UAVs [3].
Unmanned aerial vehicles (UAVs) have surfaced as a revolutionary technolog-
ical breakthrough, reforming our modern landscape and empowering the various
challenges that has been faced by the industries. UAVs equipped with sensors,
lidars, and GPU-processed cameras play an important role in various industries.
In agricultural fields, they have observed crop health, soil conditions, and the irriga-
tion needs, fostering digital and precision agriculture for optimized resource usage
and increased yields. Unmanned aerial vehicles (UAVs) have revolutionized visual
communication in the film and photography industries by capturing and analyzing
previously unattainable aerial views and cinematic viewpoints. Delivery services
have investigated using UAVs to make swift and fast deliveries to isolated locations
or evading traffic. Thermal imaging-equipped unmanned aerial vehicles (UAVs) are
utilized in search and rescue operations to locate individuals who have vanished
from sight in challenging-to-reach locations. Surveying, mapping, disaster relief,
environmental monitoring, infrastructure inspection, military surveillance, entertain-
ment, and racing are among their specialties. Their unique characteristics require
energy-efficient design in order to achieve longer flying periods.

60 53
50 46.36
40.57
40 35.54
31.14
27.31
30 21.04 23.96
18.49
20 14.34 16.25
10
0
2020 2021 2022 2023 2024 2025 2026 2027 2028 2029 2030

Fig. 1 Market growth rate for UAVs globally in USD


A Review on Anomaly Detection Using Machine Learning Techniques … 147

For precise navigation and positioning, these autonomous unmanned aerial vehi-
cles rely largely on the global positioning system (GPS) [6]. However, the unpro-
tected civilian GPS transmissions may pose a security risk. As a result, UAVs have
become susceptible to various cyber threats, such as denial-of-service attacks, data
modification, and interception [1, 6].
The security of UAVs is primary, particularly in sectors like defense and surveil-
lance, where sensitive data may be at risk and of the major concern [6]. The sensi-
tive information captured by UAVs is protected by encryption methods like AES,
RSA, and homomorphic encryption, which jumble it up. Intrusion detection systems
(IDSs) are like watchful security guards inside networks, protecting against cyber-
attacks with the help of techniques like SVM and random forest. UAV’s detection
algorithms such as YOLO and faster R-CNN visually identify drones and signal the
need for action when they detect unauthorized presence. In a world of advancing
cyber threats, such unsupervised ML techniques can strengthen the UAVs against
emerging security challenges that can steal the sensitive data gathered by the UAV’s
[1].
A. Contribution
The key contributions of this research work are outlined as follows:
• To address the UAV security challenges and integrate advanced machine learning.
• To drive UAV advancements in security, reliability, and efficiency across indus-
tries, continued collaboration and innovation in ML integration.
• To promote responsible UAV usage and foster societal trust, enforce stringent
safety protocols and regulatory compliance.
• To address data security and mission integrity, employ encryption algorithms,
intrusion detection, and anomaly detection.
• Paper Organization
This research work follows a structure to explore how the machine learning tech-
niques helps in the improvement of the unmanned aerial vehicles (UAVs). In Section
I, we have started by looking at the development of UAVs and the challenges they
face particularly in the field of defense, monitoring, and surveillance followed by
our contributions of this research and their impacts. In Sect. 1, we have discussed a
detailed literature survey, offering a detailed reviews of the current research’ taking
place in the field of development of these devices via machine learning. Section 3
emphasized how incorporating machine learning techniques like anomaly detection
and other methods can helps in the enhancement of the UAV security operations.
Section 4 focuses on our main findings, while Sect. 5 recommends some areas for
the future research for the development of these devices, showcasing how the consol-
idation of machine learning and UAV can enhance and improve the security measures
across the various fields. The abbreviations and acronyms used in this work are listed
in Table 1.
148 D. Aggarwal et al.

Table 1 List of abbreviations


Acronym Abbreviation
and acronyms
ML Machine learning
GPS Global positioning system
AES Advanced encryption standard
RSA Rivest–Shamir–Adleman
SVM Support vector machine
LSTM Long short-term memory
YOLO You only look once
R-CNN Region-based convolutional neural network
OCSVM One-class support vector machine WS
WSN-IoT Wireless sensor network-internet of things
NIDS Network intrusion detection system

2 Literature Survey

Various research papers have been highlighted in the substantial impact of the
machine learning techniques in enhancing the security of unmanned aerial vehicle
(UAV) operations, particularly in the context of physical detection technique like
anomaly detection. A comprehensive analysis is done on the 20 articles concen-
trating on the applications of UAVs and ML in remote monitoring unveiling that the
supervised learning is the most used, auditing for 61% of the cases. This indicates
a prevalent trend in leveraging supervised learning for anomaly detection tasks in
UAV operations [2]. Moreover, the studies are highlighting the usefulness of machine
learning techniques, such as regression and classification models, for dealing with
the UAV’s imagery. The studies show that regression models were mostly used in
UAV image processing in the fields of agriculture, forest, and grassland mapping.
Domains of machine learning have enhanced the capabilities of UAVs and reduced
the operational issues that enable safe, speedy, efficient, and reliable outputs [10].
This has been applied to enhance security measures, extend flight times, and opti-
mize power management in UAVs by including machine learning-based approaches.
As pointed out by the literature, the problem of identifying anomalous instances in
streaming data occupied a big deal of attention in the domain of UAV operations,
with several papers devoted to some innovative frameworks and algorithms for this
purpose. In particular, machine learning-based anomaly detection techniques are
applied to detect anomalies during UAV flights with success. Deep learning and one-
class support vector machines have been the most commonly used approaches that
demonstrate avenues to enhance security and the integrity of UAVs operations. Signif-
icantly, literature stresses how important machine learning is in trying to improve
UAVs operations, more so in anomaly detection. While machine learning techniques
are yet emerging, supervised learning algorithms still dominate. All these methods
A Review on Anomaly Detection Using Machine Learning Techniques … 149

offer some potential means for enhancing the efficiency, safety, and reliability of
UAVs operations.
Supervised learning methods are used with labeled data but they are not much
effective in real-time UAV operations due to already provided constraints. Unsuper-
vised learning methods, such as isolation forest, are used with unlabeled data effec-
tively but may struggle in complex environments. Semi-supervised learning methods
combine the strengths of both works well on limited labeled data to improve the effi-
ciency. This hybrid approach is often the most practical for UAV anomaly detection
as it balances accuracy with data availability.

3 Machine Learning and Anomaly Detection

Anomalies are the sudden unexpected irregularities or events within a particular


context which deviate from typical patterns or the norms [9]. These anomalies might
be through a pure line break, odd points in a dataset or some incredible acts in a
system. Finding the anomalies is very important as they are capable of causing the
theft of sensitive data collected by UAVs [8]. Sensitive data is avoided, unwanted
access is mitigated, and imaginative solutions against potential security threats are
powered to evolve [10].
The preferred techniques used in anomaly detection for UAV’s are isolation forest
and convolutional neural network (CNN) due to their distinct advantages. Isolation
forest is an unsupervised algorithm which is used to identify the anomalies by finding
out points through random partitioning, delivering high processing efficiency for
UAVs with limited hardware capabilities. CNN are used for processing structured
data like video frames or sensor readings, utilizing their ability to extract spatial
and temporal features without human assistance. These methods are preferred due to
their adaptability, robustness in finding out the irregular happenings, and efficiency
in managing the large-scale UAV data.
A. Anomaly Detection
Anomaly detection identifies the unexpected irregularities happening in the
observed data that differs from expected or typical behavior [2, 8]. It finds various
applications in the fields like defense, finance, cybersecurity, health care, manufac-
turing, and surveillance sectors. It helps in finding the fraudulent activities, monitors
equipment malfunctions, and predicts the unusual environmental occurrences [10].
In general, anomaly detection assumes that normal behavior or data follows certain
patterns. Any high deviation from these patterns can be labeled as an anomaly and can
represent a security threat, fault, or another unexpected event requiring immediate
treatment [9].
The application or execution of anomaly detection machine learning algorithm
in UAVs presents several key advantages. Initially, it enables UAVs to continuously
supervise their operational environment and the data which has been gathered or
collected from various sensors, such as cameras sensors and lidars. By investigating
150 D. Aggarwal et al.

the real data, the anomaly detection machine learning system can find the anoma-
lies, inconsistencies or any irregularities that can cause security infringements. For
example, during aerial surveying or monitoring the high vantage points, if these
devices detect any unpredictable happenings in any restricted area, the anomaly
detection model will promptly sound an alert for the further examination [4, 13].
Anomaly detection has many applications in the field of UAVs in real-world
scenarios. In disaster relief, UAVs can be used to identify anomalies in terrain or to
locate survivors in an isolated areas. In surveillance, it can be used for detecting unau-
thorized activities or movements. It can also be used in environmental monitoring
such as identifying illegal deforestation or wildlife trespassing.
B. Anomaly Detection via Machine Learning
The integration of machine learning techniques into UAV technology is crucial
for ensuring advancements at the forefront of innovation. These advancements can
reinforce the security measures to protect valuable or sensitive data in an increasingly
associated world [2, 4]. Such improvements will not only boost the abilities of UAVs
but also potentially support safer and more efficient operations in multiple industries.
Machine learning helps in making sure that UAVs can fly safely. Machine learning
algorithms go through a lot of data from UAVs searching for irregularities before they
arrive at a decision. For example, they can determine the time when parts of a UAVs
may break down and thus repair them in time before they cause any problem in flights
[9]. Machine learning also checks how these devices usually act and spots weird
behavior that might mean something wrong. This helps catch issues early. Moreover,
with the help of machine learning, these devices could see natural obstacles or some
other things in the way and change their course to prevent crashing. It all makes the
use of drones much safer in different places and situations (Table 2).
Enforcing the anomaly detection in these unmanned aerial vehicles can deliver
enormous improvements in their overall performance [1]. Now, by integrating these,
the anomaly detection systems will no doubt improve the security measures for
UAVs and might respond to the potential threats in real time [12]. Such devices
shall monitor the sensitive data or information from sensors, cameras, and other
sources, empowering a UAV to quickly identify or respond to unusual or unexpected
irregularities.
• Automation Advancement: Through learned patterns and analysis of data,
machine learning techniques enable UAVs to perform in an autonomous mode of
operation in making decisions within a real-time setting. This extends to devices
for mission execution, flight planning, and navigation [12, 14].
• Improved Surveillance: Machine learning techniques extend UAV surveillance
powers by identifying, classifying, and monitoring objects or irregularities in
real-time video. As a result, security, surveillance, and quick reaction times are
strengthened [8].
• Weather Prediction: Machine learning algorithms can use the data collected
through UAVs to increase the accuracy for the weather forecasting, which leads
to better predictions and early warning systems.
A Review on Anomaly Detection Using Machine Learning Techniques … 151

Table 2 Summary of literature review: features, algorithms, and implementation goal


Ref. Contributions Features Algorithms Implementation goal
[10] Introduced isolation Highlights drawbacks Anomaly Proposed isolation
forest algorithm and of existing methods in detection forest for
isolation forest for handling streaming framework with anomaly-based
anomaly-based data. Also, successful isolation forest software defect
software defect anomaly detection algorithm, detection is efficient
detection for drift methods applied in isolation trees, for streaming data
monitoring in UAVs various scenarios built using anomaly detection
bootstrap
sampling
[9] Introduced anomaly Focus on anomaly Unsupervised Utilized deep neural
detection model for detection in swarm learning models network classifier with
swarm flights drone flights using like k-means one-dimensional
machine learning clustering and convolution layers is
supervised efficient for moving
learning models horizon-based
like CNN monitoring for online
anomaly detection
[8] Structured an Provides and Intrusion To develop a
overview of the structured an overview detection cutting-edge
anomaly detection of the anomaly systems (IDSs) anomaly-based
methods and NIDSs detection methods and and NIDSs network intrusion
NIDSs frameworks detection system
[5] Review of Common ML Multimodal UAV and ML
ML-assisted UAV techniques for UAV ML, RNN, integration enhances
applications and operations and CNN, SVM, automation, control,
operations communications are IoT, K-means, and communication
described DNN
[4] Conducted in-depth Survey categorized k-means Supervised learning
literature survey of into supervised, clustering, PCA, algorithms are most
WSN- IoT for ML unsupervised, and ANN algorithms used in WSN-IoT for
engineers reinforcement for WSN-IoT smart cities
machine learning clustering
algorithms used in
WSN-IoT
[12] Enhanced UAV UAV and machine Random forest, UAV and machine
capabilities, learning research SVM, linear learning enhance
real-time growth in various regression, monitoring, data
monitoring, data sectors like smart LSM, and deep collection and
collection, and cities and military is CNN prediction
prediction analyzed

Combining machine learning and UAV technology not only increases aerial
systems’ operational capability but also transforms a number of sectors by encour-
aging more intelligent, data-driven approaches in a wide range of applications
[3]. These sectors or operating areas include multifold, multifaceted applications
including:
152 D. Aggarwal et al.

a. Monitoring crop health with sensor data collection and precision agriculture
b. Aerial monitoring with real-time surveillance
c. Delivery services to remote areas
d. Digital agriculture with crop health observation and soil condition monitoring
e. Search and rescue operations with thermal imaging
f. Environment monitoring with disaster response
g. Infrastructure inspection with structural health monitoring and aerial mapping
h. Military surveillance with reconnaissance
i. Racing, entertainment with film making, cinematic perspectives.

4 Methodology and Drawbacks of Existing Models

The data collection and processing pipeline for many of the pre-existing models goes
through a subset or all of the following processes:
A. Data Collection
The initial step includes the collection of data from UAVs that is equipped with
cameras and lidar sensors. The dataset should contain both the normal as well as
anomalies events.
B. Data Preprocessing
Clean the data to remove the unnecessary missing values. Convert this data into
a suitable format and take some relevant features from it. This step will ensure the
quality and usefulness of data so that it can be given to our machine learning model.
C. Model Selection
Choose the most suitable machine learning algorithm that is based on the data
collected. Techniques include supervised, unsupervised, or semi-supervised models
based on the availability and type of data.
D. Model Training
The dataset is divided into training and test sets. Train the ML model using the
training set and must ensure that the model does not overfit the data.
E. Anomaly Detection
Test the accuracy of model with the test set to evaluate their performance. Use
metrics like accuracy, precision, recall, and F1 score to identify the most effective
model for detecting the anomalies.
Incorporating the anomaly detection algorithm into these UAV’s brings you some
difficulties or challenges that needs to be focused. One issue would be the high
computation power can slow down the operations and may affect the decision-making
of these devices which is very important in the fields of effective surveillance [8].
There is also a risk of misidentifying the normal activities as anomalies or missing
A Review on Anomaly Detection Using Machine Learning Techniques … 153

real anomalies due to the complexities of contextual interpretation, resulting in false


positives or false negatives. Furthermore, their limited ability to detect entirely new
or evolving anomalies might reduce their effectiveness, thus reducing the efficiency
of this model [8]. Collection of data in UAV’s is covered with challenges, particularly
in remote or tough environments because high-quality sensor data is contained with
noise. Additionally, managing and storing the large volume of data is difficult.
The utilization of UAV for defense and surveillance causes significant ethical and
privacy concerns, especially in sensitive or public areas. The collected data that has
been captured may contains security risks. Regulatory and security frameworks must
be implemented to get rid of this security challenges. Anomaly detection models face
a lot of challenges in real as well as dynamic environments due to limited power,
bandwidth on UAVs. High-resolution power on sensors transmission often leads to
delays in deliveries, making it less powerful. It needs to be lightweight, fast, flexible
and adaptive. Additional reasons are provided below.
• Limited computation capacity: The time complexity and the speed of anomaly
detection algorithm that they use will be restricted by the limited computing
capacity of unmanned aerial vehicles.
• Data Transmission: High-quality data from UAVs to ground stations puts a strain
on the bandwidth that is available, leading to communication problems and a
significant delay [2].
• Learning New Anomalies: It can be difficult for detection algorithms to identify
completely new or emerging anomalies that weren’t present in the training data
[1, 9].
• Privacy Concerns: Using UAVs for surveillance raises concerns about privacy,
data storage, and potential misuse of collected information [8].
• Maintenance: Ensuring continuous functionality and updates is vital; any malfunc-
tions in sensors or algorithms can jeopardize accuracy.

5 Future Work

In military and fragile surveillance applications, the presentation of a self-destruct


mechanism for UAVs is an essential security measure. Its main motive is to prevent
the unsanctioned captured, gathered or deployment of UAVs by individuals or groups
that can cause a security hazard to the country or organization. The concepts behind
this are to carry out a self-destruct process in situations like border invasions or the
danger or risk of the UAV getting into the wrong hands, to safeguard categorized
information, advanced technology, and hardware. This security measure is particu-
larly important in circumstances where UAVs are used for tasks such as observation,
investigation, inspection, gathering, or military activities. Future researches should
focus on developing the adaptive models that can grow with changing threats. Multi-
modal integration, such as combining visual, thermal, and acoustic sensor data, offers
potential for improved accuracy. Lightweight models optimized for edge computing
can enable real-time anomaly detection on UAVs. Ethical considerations, such as
154 D. Aggarwal et al.

data privacy and responsible use of UAVs in sensitive areas, must also be addressed
to ensure broader societal acceptance. These types of mechanism in UAVs used
for investigation in military areas can help in preventing the data to be exploited.
This framework outlines critical research areas to ensure the effective and secure
implementation of self-destruct mechanisms in UAVs:
• Identifying threats and conducting risk analysis
• Reviewing existing technologies and conducting feasibility studies
• Designing and developing self-destruct systems
• Integrating and miniaturizing components
• Creating activation methods and security protocols
• Performing thorough testing and validation
• Addressing legal and ethical considerations
• Developing operational guidelines and training programs
• Ensuring data security and secure data deletion
• Implementing continuous improvement and feedback mechanisms.

6 Conclusion

In this research, we have focused on the security challenges faced by UAVs operations
and how it can be rectified by consolidating them with the machine learning. UAVs
play an important role in various fields, especially in agriculture, rescue and salvage
operations, data monitoring, data collection, and data acquisition. In our research, we
have given a basic innovative solution to tackle these challenges. The consolidation of
anomaly detection and UAV are useful in enhancing the security of these devices and
ensuring the efficiency of the UAVs in a constantly changing danger environment.

References

1. Bithas PS, Michailidis ET, Nomikos N, Vouyioukas D, Kanatas AG (2019) A survey on


machine-learning techniques for UAV-based communications. Sensors (Switzerland) 19(23).
MDPI AG, 01 Dec 2019. [Link]
2. Eskandari R, Mahdianpari M, Mohammadimanesh F, Salehi B, Brisco B, Homayouni S (2020)
Meta-analysis of unmanned aerial vehicle (UAV) imagery for agro-environmental monitoring
using machine learning and statistical models. Remote Sens 12(21):1–32. MDPI AG, 01 Nov
2020. [Link]
3. Abubakar AI et al (2023) A survey on energy optimization techniques in UAV-based cellular
networks: from conventional to machine learning approaches. Drones 7(3). MDPI, 01 Mar
2023. [Link]
4. Sharma H, Haque A, Blaabjerg F (2021) Machine learning in wireless sensor networks for
smart cities: a survey. Electronics (Switzerland) 10(9). MDPI AG, 01 May 2021. [Link]
org/10.3390/electronics10091012
5. Kurunathan H, Huang H, Li K, Ni W, Hossain E (2023) Machine learning-aided operations
and communications of unmanned aerial vehicles: a contemporary survey. IEEE Commun Surv
Tutor 1–1. [Link]
A Review on Anomaly Detection Using Machine Learning Techniques … 155

6. Khoei TT, Gasimova A, Ahajjam MA, Al Shamaileh K, Devabhaktuni V, Kaabouch N (2022)


A comparative analysis of supervised and unsupervised models for detecting GPS spoofing
attack on UAVs. In: IEEE international conference on electro information technology, IEEE
computer society, pp 279–284. [Link]
7. Bouhamed O, Wan X, Ghazzai H, Massoud Y (2020) A DDPG based approach for energy-
aware UAV navigation in obstacle constrained environment. In: 2020 IEEE 6th world forum
on internet of things (WF-IoT). IEEE, pp 1–6
8. Bhuyan MH, Bhattacharyya DK, Kalita JK (2014) Network anomaly detection: methods,
systems and tools. IEEE Commun Surv Tutor 16(1):303–336. [Link]
2013.052213.00046
9. Ahn H, Choi HL, Kang M, Moon ST (2019) Learning-based anomaly detection and monitoring
for swarm drone flights. Appl Sci (Switzerland) 9(24). [Link]
10. Ding Z, Fei M (2013) An anomaly detection approach based on isolation forest algorithm for
streaming data using sliding window. In: IFAC proceedings Volumes (IFAC-papers online),
IFAC Secretariat, pp 12–17. [Link]
11. Zhou Z-G, Tang P (2016) Continuous anomaly detection in satellite image time series based on
Z-scores of season-trend model residuals. In: 2016 IEEE international geoscience and remote
sensing symposium (IGARSS), Beijing, China, pp 3410–3413. [Link]
RSS.2016.7729881
12. Aamer N, Bhadula S, Patil H, Gurpur S, Verma A (2024) TDA-SSNN: efficient hybrid
top-down attention siamese sparrow neural network for automatic drone navigation. In: 2024
4th international conference on sustainable expert systems (ICSES), IEEE, pp 1596–1602
13. Slimane HO, Benouadah S, Khoei TT, Kaabouch N (2022) A light boosting-based ML model
for detecting deceptive jamming attacks on UAVs. In: 2022 IEEE 12th annual computing and
communication workshop and conference, CCWC 2022, Institute of electrical and electronics
engineers Inc., pp 328–333. [Link]
14. Sharma A, Singh S, Pendyala RP, Kumar A, Shukla KK, Balamuralidhar P (2022) Path planning
for multiple targets interception by the swarm of UAVs based on swarm intelligence algorithms:
a review. IETE Tech Rev 39(3):675–697
Unlocking the Power of OpenCV:
Innovative Development
of Gesture-Based Systems

Dharmendra Sharma and Manmohan Singh

Abstract This work explores the potential of using OpenCV and MediaPipe to
develop gesture-based interaction modules aimed at enhancing human–computer
interaction. The five preliminary modules in this study include finger count, finger
tracker, air canvas versions 1 and 2, color selection and graphical tools, and a gesture-
based computer control system. These modules using rudimentary image processing
from the OpenCV module for images and MediaPipe for gesture recognition are
a path toward fully realizing mapped intent to allow for a smoother integration
of human intent and machine response. This investigation not only enhances the
development of gesture recognition technology but also provides directions for future
studies to enhance the gesture-based interaction on digital devices.

Keywords OpenCV · MediaPipe · Gesture recognition · Air canvas ·


Gesture-based computer control

1 Introduction

The use of hand gestures in digital interactions has become a crucial method of
communication, offering a new level of proactive engagement within digital envi-
ronments. This progress has been made possible by recent advancements in computer
vision, leading to a new era of web-connected and screen-based interaction to a new
generation of web and screen connected communication or interaction. Exploring the
territory of gesture recognition brought us to the result of a brief acquaintance with
various libraries, each of them unique capabilities. In these, OpenCV played a signif-
icant role as library for all our models. Renowned for its robust image processing
and analytical capacity along with elaborate records [1]. Thus, OpenCV was acting

D. Sharma (B)
NMIMS University, STME Indore, Indore, India
e-mail: [Link]@[Link]
M. Singh
IES College of Engineering, Bhopal, India

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 157
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
158 D. Sharma and M. Singh

as the basic tool on which we developed our models. It ranged from simple image
manipulation techniques to complex ones, requires sophisticated computations and
hence is crucial in our development of improved gesture recognition technologies.
Besides OpenCV, MediaPipe took a crucial role in hand tracking—one of the main
components that are studied in our work. Thanks to algorithms we managed to get
really good accuracy in monitoring hand motion, and the basis was founded for
enhancing the emerging sophisticated gesture recognition models [2]. Along with
these libraries, PyAutoGUI facilitated seamless integration between gesture recog-
nition and computer interactions, and NumPy, with its array operations, acted as the
computational backbone for processing the large volumes of data during our research
[3].
Throughout the course of our research, we iteratively developed and refined
several key models to address different aspects of gesture recognition and interaction.
These ranged all the way from basic finger count models up to fully functional air
canvas and gesture detection models.
The critical problems in the development of gesture-based modules and their
implications include.

1.1 Gesture-Based Computer Control Limitations

• Mapping gestures to specific computer control commands required careful cali-


bration to ensure accuracy. The integration with PyAutoGUI faced limitations in
gesture-to-command mapping accuracy and system responsiveness, with some
users experiencing issues in command execution due to recognition errors [4].

Research Constraints and User Diversity


• Environmental factors, such as lighting and background variations, significantly
impacted accuracy, pointing to a need for algorithms that are adaptive to these
variabilities.
In the subsequent sections of this paper, we delve deeper into the intricacies of
gesture recognition techniques, elucidating our methodology, findings, and implica-
tions for future research. Through this exploration, we seek to illuminate the path
toward a more intuitive and immersive digital landscape, where the language of
gestures converges seamlessly with the boundless possibilities of technology.

2 Literature Review

The development of an effective gesture detection system stands on the advancements


in computer vision and object tracking. This literature review covers decades of
research in computer vision, with a focus on advancements in object tracking and
Unlocking the Power of OpenCV: Innovative Development … 159

hand gesture recognition—key technologies for the seamless operation of air canvas
and gesture detection modules. Through this review, we aim to establish a well-
informed baseline from where our project can innovate.
At the heart of our look at computer vision and gesture OpenCV, the Open-Source
Computer Vision Library is a significant tool in the field of gesture recognition, known
for its extensive capabilities in real-time computer vision applications. Vision Library
or a library has a high-power impact on this field. OpenCV with its full list of program-
ming main functions devised principally for real-time computer vision is significant
in promoting systems that can use gesture detection. It is that collection of algorithms
has helped in the development of as facial recognition, object following, and hand
movement recognition systems [5]. The use of library is very flexible, and hence, (core
competency) analyzing high frame rate video streams very important for the real-time
analysis needed in our air canvas and gesture detection models. An integration feature
that makes the OpenCV CAM standout is its precious capacity to interface with
high level that has also revolutionized the programming languages such as Python
which has also significantly simplified the process of development, thus allowing for
its fast. The implementation and evaluation of Computer Vision modules for proto-
typing used for Object tracking. Object tracking is critical for applications such as air
canvas and gesture detection. It is all about charting the experience of an object across
a series of frames within a video, maintaining consistent identification throughout
its movement [6]. Some of the difficulties that accompany object tracking are infor-
mation due to the motion of objects, noisy images, vagueness, irregular movement,
escalating designs, and the demand for real-time processing. To overcome all these
challenges, constraints are used, for example, where one assumes linear movement
or using algorithms predicting the subsequent position of objects. Object detection’s
goal is to locate areas in an image correlating with objects of one or another class.
Various approaches consist of frame differencing, optical flow, and DIVSQ TABLE
background subtraction, each with its utilities depending on the application. Frame
differencing works by comparing frames; optical flow observes pattern changes in
motion images from the pixels of a picture, and background subtraction clears objects
in motion from a stationary environment, often employed for surveillance situations
[7]. Among all categories with object tracking, hand tracking is of significance for air
canvas systems. It involves color-based recognition, which commonly with the help
of colors of the gloves to easily differentiate them. Appearance-based recognition
bases it recognition on what can be seen details of hands. Motion-based recognition
identifies objects by employing image frame analysis, which through the mathe-
matical algorithms like AdaBoost for gesture recognition. Nonetheless, it has the
following problems: gesture confusion and background interferences. The solution
proposed here is to use the two-stage detection approach, initial, first defining a refer-
ence point for following the motion of the hand and then verb experiment identifying
a reference for hand motion, then following a gesture-matching model for correct
gesture recognition [8]. Furthermore, the skeleton-based recognition method offers
an advantage by analyzing skeletal joint orientations and spatial relationships, facil-
itating a deeper understanding of hand gestures. Tools like MediaPipe have adopted
this method, provided advanced hand detection and tracked capabilities [9]. The
160 D. Sharma and M. Singh

architecture of an air canvas system integrates hand tracking techniques to capture


users’ movements and translate them onto a digital canvas. Proposed systems include
using Kinect sensors for motion detection, LED-fitted gloves for precise tracking,
and computer vision for accurate gesture interpretation [10]. The system’s founda-
tion includes ‘Write in Air’, ‘Fingertip Detection’, and ‘Traced Trajectory, with the
possibility of further enhancements based on the selected finger tracking method.
The application of PyAutoGUI in gesture-based control systems is a testament to the
versatility and effectiveness of Python libraries in enriching user interactions. PyAu-
toGUI allows for the automation of GUI interactions, including mouse movements
and keyboard actions, through simple Python scripts [11]. This capability is partic-
ularly useful in developing gesture-based computer control systems, where gestures
can be translated into commands that simulate keyboard or mouse, thus bridging the
gap between human gestures and computer responses [12]. As we move forward,
the integration of robust tracking methodologies, gesture recognition, and automa-
tion libraries like PyAutoGUI remains paramount for the development of interactive
systems like air canvas. The combination of MediaPipe [13] and color-based recog-
nition algorithms [14], along with the utility of PyAutoGUI, showcases the potential
of employing these technologies in creating intuitive and engaging user experiences.

3 Methodology

In this section, the approaches used for the proposed conceptualization, development,
and evolution of the suite of the gesture-based modules are elucidated. The devel-
opmental process is anchored on the fundamentals offered by OpenCV and Medi-
aPipe to support computer vision functionality. A systematic approach was employed
encompassing a priori determination of challenges and requirements and satisfactory
completion of one module after the other in terms of design and development.

3.1 Development Environment Setup

Establishing a robust development environment is essential for efficiently designing


computer vision applications. Python 3.x was chosen as the primary programming
language due to its extensive library support and active community resources. The
following libraries and tools were mainly used:
• Open-Source Computer Vision Library (OpenCV)
• MediaPipe
• PyAutoGUI.
Unlocking the Power of OpenCV: Innovative Development … 161

3.2 Implementation of Research Modules

The finger count and finger tracker modules were research modules that we developed
as a basis for the further understanding and development of our more complex air
canvas and gesture control modules.
The finger count and finger tracker modules are designed around MediaPipe’s
Hand Tracking solution, which provides a set of 21 hand landmarks [15]. Our algo-
rithm differentiates between various finger positions to count the number of extended
fingers accurately. The implementation involves:
• Hand Detection and Landmark Acquisition: Upon detecting the hand in the video
frame by background subtraction, the system captures 21 distinct landmarks that
represent key points of the hand, including fingertips and joints [16, 17].
• Convex Hull and Convexity Defects Calculation: The algorithm constructs a
convex hull around the detected landmarks. Convexity defects between the hull
and the hand contour help establish the valleys between fingers [18, 19].
• Finger Count Logic: By analyzing the angles and distances between the convexity
defects and the center of the palm, the algorithm distinguishes between extended
fingers and those that are folded. The count is the presented to the user [20]
(Fig. 1).
The finger tracker extends the capabilities of the finger count module to include
tracking the motion path of fingers, focusing more on trajectory and movement
patterns:
• Trajectory Mapping: Each fingertip’s position is tracked frame-by-frame, allowing
the system to draw the movement path on the screen. This is particularly useful
for applications requiring precise motion tracking, such as digital drawing or
signature verification.
• Smooth Tracking: Implementing smoothing techniques to reduce jitter and inac-
curacies in the tracking data, ensuring that the trajectory lines are smooth and
accurately represent the user’s movements.
• Dynamic Tracking Adjustment: The system dynamically adjusts tracking sensi-
tivity based on the speed and complexity of finger movements, optimizing for
both slow, detailed motions, and fast, broad gestures (Fig. 2).
Implementation of advanced gesture detection modules

Fig. 1 Finger count logic flow


162 D. Sharma and M. Singh

Fig. 2 Finger tracker

(1) Air Canvas with Color Changing Options


Color Changing Options for an Air Canvas: This version of air canvas allows the
user to completely switch drawing colors by certain moves or contact with a virtual
color palette. Implementation highlights include:
• Color Palette Detection: The canvas includes zones where different colors would
be allowed. When the user sweeps a fingertip over one of these areas, the system,
in addition, helps update the current drawing color as required.
• Gesture-Based Color Selection: The system can also to distinguish some of the
gestures used to change colors, swipe or a tap over a specific region of the canvas
closer to the outer border enabling the filter to change from one color to another
without blocking for some time drawing flow.
• Real-Time Drawing and Color Feedback: As users erase or draw on the virtual
canvas, the system provides instant feedback on their actions and displays the
strokes in the selected appreciable color, enhancing the interactive experience
(Fig. 3).
Air Canvas with Shape Drawing and Tool Selection: Air writing methods include
shape drawing and tool choices. This new air canvas module is an advanced model
and works through gestures appreciation for choosing drawing tools and creating
pre-established shapes. Key features include:
• Predefined Gesture Library: The system recognizes a set of predefined gestures,
each associated with a drawing tool or shape (e.g., circle and rectangle) [16].

Fig. 3 Advance gesture detection logic


Unlocking the Power of OpenCV: Innovative Development … 163

Fig. 4 Shape selection

Users can switch between tools or draw shapes by performing the corresponding
gestures.
• Dynamic Tool and Shape Feedback: Upon recognizing a tool selection or shape
drawing gesture, the system provides immediate feedback by changing the
drawing cursor or overlaying a template shape on the canvas, guiding the user
in completing the drawing action.
• Adaptive Gesture Recognition: The algorithm adapts to the user’s drawing style
and speed, ensuring that gestures are accurately recognized under a variety of
conditions and minimizing false detections (Fig. 4).
Gesture-Based Computer Control: This module maps hand gesture into computer
control. The kind of commands make people use their computers in different ways
they desire, natural movements:
• Gesture to Command Mapping: Specific hand gestures are mapped to computer
commands, such as mouse clicks, scrolling, or keyboard shortcuts. The system
uses the detected hand landmarks with the performed gesture that is used to
produce the corresponding command.
• Integration with PyAutoGUI: Upon recognizing a movement, the system utilizes
the PyAutoGUI module for the gesture, effectively bridging the gap between
gesture recognition and computer control.
User Feedback for Gesture Recognition: Visual indicators or animations appear
on the screen when a gesture is recognized, providing users with feedback that their
movement has been acknowledged. The recognized gesture and the corresponding
action enhance the system’s usability and responsiveness overall (Fig. 5).

Fig. 5 User feedback loop


164 D. Sharma and M. Singh

4 Result and Discussion

This section presents a discussion on the assessment of the performance of our air
recognition gestures analyzing master modules, abbreviated as ARM 3M, canvas
with color option, air canvas with shape drawing and ANFIMs in tool selection, and
gesture-based computer control. We used a small yet diverse group of ten participants,
with whom we tested the accuracy, response time, and overall user satisfaction of each
module. We have presented graphical visualizations derived from the data, enabling
a clear comparison across different aspects.

4.1 Air Canvas with Color Changing Options

This module demonstrated a promising accuracy level, with correct color changes
being detected in 88% of the test cased. The average response time was 1.6 s, which
is relatively swift. Figure 6, generated using Python, helps visually represent the
accuracy and response time across our test cases. Most participants rated the module
as satisfactory in nature, though some expressed that it could be improved by a more
diverse color palette.

Fig. 6 Accuracy with color


changing option
Unlocking the Power of OpenCV: Innovative Development … 165

Fig. 7 Accuracy with


drawing tool selection

4.2 Air Canvas with Shape Drawing and Tool Selection

This air canvas module showcased 90% accuracy, correctly interpreting hand place-
ment, and shape and tool selection. The response time averaged 1.5 s. Figure 7 illus-
trates these metrics across the ten test cases, showcasing a comparatively reliable
module. User suggestions included toolset expansion and further latency reduction.

4.3 Gesture-Based Computer Control

Our gesture-based computer control module achieved a relatively lower accuracy


among the tested modules, with a 70% success rate in accurately executing computer
control commands via gestures. The response time on the other had was notably quick
at 1.0 s, as shown in Fig. 8.

5 Implications and Future Directions

The exploration and development of gesture-based modules, as presented in this


research, show significant progress in the realm of digital interaction. By harnessing
the capabilities of OpenCV, MediaPipe and PyAutoGUI, intuitive gesture recogni-
tion systems have become feasible. This work holds profound implications for the
advancement of human–computer interactions (HCI), offering a future where digital
environments are navigated through natural human gestures.
166 D. Sharma and M. Singh

Fig. 8 Accuracy with


computer control

5.1 Potential Applications

The practical applications of gesture-based modules extend across multiple domains.


In educational settings, these technologies can create more engaging and interactive
learning experiences. The entertainment industry, particularly video gaming and
virtual reality, stands to benefit from more immersive control mechanisms. Addi-
tionally, this work has significant implications for accessibility, offering alternative
interaction methods for individuals with disabilities.

5.2 Challenges and Considerations

Throughout the research, we have observed several challenges that increase the
complexity of developing effective gesture-based interaction systems. Variability
in lighting conditions and backgrounds significantly impact the accuracy of gesture
recognition, emphasizing on the need for more adaptive and diverse algorithms.
Furthermore, different human hand sizes and gesture styles can be problematic.

5.3 Future Scope

Gesture-based systems have tremendous scope, and there can be several avenues for
future research:
• Enhanced Gesture Recognition Algorithms: Future work should focus on devel-
oping more sophisticated algorithms that can further improve gesture recognition
accuracy and reduce latency, making interactions feel more seamless and natural.
Unlocking the Power of OpenCV: Innovative Development … 167

• Expanded Gesture Libraries: Extending the range of recognizable gestures will


allow for more complex and nuanced interactions, broadening the applicability
of gesture-based systems.
• Integration with Emerging Technologies: Exploring the integration of gesture-
based modules with VR and AR technologies could lead to groundbreaking
applications, offering users more immersive and interactive digital experiences.
• User Experience (UX) Studies: Conducting comprehensive UX research is crucial
for refining gesture-based interfaces. Future studies should focus on understanding
user needs, preferences, and challenges to inform the design of more user-friendly
and accessible systems.

6 Conclusion

By utilizing the functionalities of OpenCV and MediaPipe, this research has


contributed to the development of innovative modules aimed at enhancing inter-
active digital engagement. The present paper can thus be seen as the summary of the
particular journey that has been undertaken in the course of the research, and it has
been seen when developing these modules that there is the possibility gestures to act
as smooth realizations of the link between man and machine technology. We have
also recorded impressive performances throughout this research. The air canvas with
color changing options, air interact with graphic canvas for shape drawing, and tool
choosing and gesture all of the based computer modules have been found to give
high accuracy but are lacking in gesture recognition and really good response times.
These studies have discussed the possibility and the efficiency of our gesture-based
systems in making more endowed interaction and engaging digital experiences. The
implication of this work cuts across different aspects of education to expand its enter-
tainment and accessibility domains, new opportunities particularly for participation
and communication. Despite the challenges they faced—such as variability, usability
issues, and environmental and user diversity—our research serves as a foundational
step toward creating a more meaningful and inclusive global digital environment.
So far, we have witnessed several steps opening up the way for improving gesture
recognition algorithms, increasing the variety of the gestures libraries, and how to
incorporate these technologies into the more modern digital realms. Thus, further
development of gesture-based interactive systems will unquestionably be a central
component which will have an impact on the development of the future generation
of human–computer interaction. As a result, our investigation provides not only a
better understanding of the contemporary potentiality of gesture-based systems and
at the same time creates lots of opportunities for the enhancement of gesture-based
systems, virtual legibility, and porosity of borders and classifications, providing infi-
nite ideas for adaptation and development. The language of gestures, which, as is
well known, is natural at best, and flexibility, also known as ‘expressiveness’, it also
paves the way for a new epoch of digital communication, one that is also easy to
manage, interesting, and as endless as the sea creative.
168 D. Sharma and M. Singh

References

1. Wagh C (2023) Object detection and tracking using deep learning and OpenCV in real time
environment. Int J Eng Res Technol (IJERT) 12(04)
2. Pardeshi S, Apar M, Khot C, Deshmukh A (2022) Air doodle: a realtime virtual drawing tool.
Int J Res Appl Sci Eng Technol. [Link]
3. Dhaigude S, Bansode S, Waghmare S, Warkhad S, Suryawanshi S (2023) Computer vision
based virtual sketch using detection. Int J Res Appl Sci Eng Technol. [Link]
ijraset.2022.39814
4. Zhang Z, Tao D (2020) Empowering things with intelligence: a survey of the progress,
challenges, and opportunities in artificial intelligence of things. IEEE Internet Things J
7(10):9128–9147
5. Sudderth EB, Mandel MI, Freeman WT, Willsky AS (2004) Visual hand tracking using
nonparametric belief propagation. In: IEEE CVPR workshop on generative model based vision
6. Yilmaz A, Javed O, Shah M (2006) Object tracking: a survey. ACM Comput Surv 38(4):3–es.
[Link]
7. Vadlamudi H (2020) Evaluation of object tracking system using open-CV in Python. Int J Eng
Res Technol (IJERT) 09(09)
8. Oudah M, Al-Naji A, Chahl J (2020) Hand gesture recognition based on computer vision: a
review of techniques. J Imaging 6:73. [Link]
9. Sunanda BE, Bhargavi M, Sree MT, Ananya MRS, Kavya N (2022) Air canvas using OpenCV,
MediaPipe. Int Res J Mod Eng Technol Sci 04(05)
10. Kumar A, Dudhbhate A, Surve Y, Bhalerao S (2023) Air canvas—motion to digital converter.
Int J Res Publ Rev 4(4):4402–4407
11. JavaTpoint. Python PyAutoGUI Library. [Link]
rary. Accessed 13 Feb 2024
12. PyAutoGUI Documentation. PyAutoGUI Documentation. [Link]
en/latest/. Accessed 13 Feb 2024
13. RamachandraH V, Balaraju G, Deepika K, Sebastian S (2022) Virtual air canvas using OpenCV
and Mediapipe, pp 1–4. [Link]
14. Baig F, Khan M, Beg S (2013) Text writing in the air. J Inf Disp 14. [Link]
15980316.2013.860928.2013
15. Sánchez-Brizuela G, Cisnal A, Fuente-Lopez E, Fraile J-C, Pérez-Turiel J (2023) Lightweight
real-time hand segmentation leveraging MediaPipe landmark detection. Virtual Rity 27:1–8.
[Link]
16. Mahmood NT, Jabbar MS, Abdalrazak M (2023) A real-time hand gesture recognition based on
media-pipe and support vector machine. Lecture notes on data engineering and communications
technologies advances in intelligent computing techniques and applications. Springer Nature
Switzerland, pp 263–273. [Link]
17. Singh M, Ayuub S, Baronia A, Soni D (2023) Analysis and implementation of disease detection
in leafs and fruit using image processing and machine learning. SN Comput Sci 4(5):627
18. Soni M, Singh NK, Das P, Shabaz M, Shukla PK, Sarkar P, Singh S, Keshta I, Rizwan A (2022).
IoT-based federated learning model for hypertensive retinopathy lesions classification. IEEE
Trans Comput Soc Syst 10(4):1722–1731
19. Ang K, Ang KM, Juhari MRBM, Wong CH, Sharma A, Ang CK, Lim WH (2023). Classifica-
tion of wafer defects with optimized deep learning model. In: Conference: 28th international
conference on artificial life and robotics, pp 609–614
20. Sonkamble RG, Bongale AM, Phansalkar S, Sharma A, Rajput S (2023) Secure data
transmission of electronic health records using blockchain technology. Electronics 12(4):1015
Predicting Cardiovascular Disease Risk
Using Tree-Based Gradient Boosting
Machine Learning Techniques

Nikhat Raza Khan, Shekhar Verma, Hemant Kumar,


Mamata Mayee Panda, Abhishek Dwivedi, and Abhishek Kumar Mishra

Abstract Heart disease prediction is a significant challenge in medical science,


necessitating robust and interpretable machine learning models. This study utilized
tree-based boosting techniques, specifically XGBoost, CatBoost, and LightGBM
(LGBM), to develop predictive models for heart disease detection using a cardio-
vascular dataset. Preprocessing steps included missing value imputation using K-
nearest neighbors (kNN) and normalization for optimal model performance. Recur-
sive feature elimination (RFE) was employed for feature selection, enhancing model
efficiency. Model performance was assessed through training and testing splits with
5-fold cross-validation. The results indicated that LGBM outperformed the other
model, achieving a testing accuracy of 96.55%, precision of 96.73%, recall of
96.82%, and F1-score of 96.68%. These findings suggest that tree-based boosting
models are effective for heart disease prediction, enhancing clinical decision-making.

Keywords Heart disease detection · Tree-based boosting · XGBoost ·


LightGBM · CatBoost · Machine learning · KNN imputation · Recursive feature
elimination · Cardiovascular disease prediction

N. R. Khan
Department of Computer Science and Engineering, IES College of Technology, Bhopal, India
S. Verma · M. M. Panda · A. Dwivedi
Department of Computer Applications, Chhatrapati Shahu Ji Maharaj University, Kanpur, India
e-mail: shekharverma@[Link]
M. M. Panda
e-mail: mamta@[Link]
A. Dwivedi
e-mail: abhishekdwivedi@[Link]
H. Kumar (B)
Department of Information Technology, Chhatrapati Shahu Ji Maharaj University, Kanpur, India
e-mail: hemantime@[Link]
A. K. Mishra
IFTM University, Moradabad, India

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 169
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
170 N. R. Khan et al.

1 Introduction

Heart disease is a major cause of deaths around the world, which makes it vital to
create effective diagnostic tools. Cardiovascular diseases (CVDs) account for about
32% of all global fatalities, highlighting an urgent requirement for improved detection
and prevention methods [1]. The complex nature of heart disease, shaped by genetic
factors and lifestyle choices, calls for a thorough understanding of the underlying
mechanisms. Recently, advancements in machine learning (ML) have unlocked new
opportunities for enhancing prediction accuracy in medical diagnostics, particularly
regarding heart disease.
The emergence of sophisticated algorithms has fundamentally changed how
medical professionals detect heart diseases. Tree-based methods, including decision
trees (DT), random forest (RF) [2], and gradient boosting, utilize hierarchical struc-
tures to model complex relationships in data. These algorithms aid in pinpointing
crucial predictors of heart disease by dividing the feature space into homogeneous
subsets based on input feature values. Their capability to handle both categorical and
continuous variables makes them particularly effective for clinical datasets, which
often contain a mix of demographic, clinical, and lifestyle information [3, 4].
Integrating feature selection techniques with tree-based methods significantly
boosts predictive performance by discarding irrelevant or redundant features, thus
simplifying the model and enhancing interpretability. This is especially important in
the medical domain, where clear decision-making processes can greatly affect patient
outcomes. Research has shown that tree-based algorithms, particularly the random
forest model, surpassing accuracy of 95% in various applications. Such effective-
ness illustrates the potential of ML, particularly tree-based techniques, to transform
cardiovascular health care by enhancing diagnostic accuracy and enabling timely
interventions [5].
Numerous studies have examined the role of machine learning in detecting heart
disease. However, these models often struggle with nonlinear data relationships.
Boosting algorithms, including XGBoost [6], LightGBM [7], and CatBoost [8],
represent an advancement in predictive modeling, providing strong performance
even with high-dimensional and noisy datasets. These approaches iteratively refine
weak learners, thereby improving overall predictive capabilities [9].
Tree-based boosting algorithms, such as gradient boosting decision trees (GBDT),
XGBoost, and others, are ensemble learning methods that combine multiple weak
learners to create a strong predictive model. These approaches work by sequentially
fitting models to the residuals of previous models, thereby enhancing prediction
accuracy. The advantages of these algorithms include their ability to handle complex
datasets, manage missing values, and improve interpretability compared to traditional
methods [10].
A pivotal study by Yasmeen Aldossary et al. [11] conducted a comparative analysis
of several tree-based ensemble models RF, DT, extra trees, and gradient boosting
using the Cleveland heart dataset. Their findings indicated that the extra trees model
achieved accuracy of 92%, outperforming the decision tree model, which recorded
Predicting Cardiovascular Disease Risk Using Tree-Based Gradient … 171

an accuracy of 84%. The study also examined the impact of ensemble techniques,
concluding that the stacking ensemble model matched the performance of extra trees,
while the voting ensemble model exhibited slightly lower accuracy at 90%. This
emphasizes the potential of ensemble approaches to enhance predictive performance.
Ahmad Hammoud et al. [12] compared multiple ML algorithms, including DT and
RF, as part of a broader evaluation of heart disease classification techniques. Their
research demonstrated that the random forest model achieved an impressive 94.96%
accuracy, significantly enhancing the reliability of heart disease predictions compared
to other algorithms evaluated in their study. The incorporation of hyperparameter
optimization techniques such as Grid Search and feature selection methods further
bolstered model performance, leading to more precise diagnostic capabilities.
Vibha et al. [13] conducted an exploratory data analysis (EDA) to assess various
ML techniques, including DT and RF, in predicting heart disease. Their study high-
lighted the importance of understanding feature distributions and correlations, which
facilitated the development of a hybrid model that combined RF and support vector
machine (SVM) algorithms to enhance predictive accuracy. This innovative approach
underscores the value of integrating diverse methodologies to improve model robust-
ness in clinical applications. Charanjeet Dadiyala et al. [14] explored a comprehen-
sive staging approach to heart disease prediction using decision trees, KNN, and
random forest classifiers. Their results indicated that random forest was the most
efficient algorithm, achieving an accuracy rate of 87.12%. This study emphasizes
the necessity of developing robust methodologies that not only predict the pres-
ence of heart disease but also classify patients according to disease stages, thereby
providing deeper insights into patient management.
This study primarily aims to investigate the use of tree-based boosting algorithms
for detecting heart disease, utilizing widely accessible clinical datasets. By harnessing
the advantages of XGBoost, LightGBM, and CatBoost, we aim to enhance the predic-
tive accuracy of heart disease diagnoses while reducing computational demands.
Additionally, this research intends to contribute to the expanding literature advo-
cating for AI-driven models in healthcare environments, where efficient diagnostics
are essential.

2 Proposed Methodology

Figure 1 illustrates a comprehensive approach to heart disease detection using tree-


based boosting machine learning models. The methodology involves several key
stages: data preprocessing, feature selection, model training, and evaluation.
172 N. R. Khan et al.

Fig. 1 Proposed approach for heart disease detection

2.1 Dataset Description

A publicly available dataset was utilized in this study, specifically focusing on heart
disease detection. The dataset, obtained from Kaggle, contained patient information
relating to various clinical and demographic factors such as age, gender, cholesterol
levels, blood pressure, and other risk indicators.

2.2 Data Preprocessing

Preprocessing the data is one of the most crucial steps ML pipeline. Missing data is
a common issue in medical datasets, and addressing it properly was essential for the
success of this study. After analyzing the dataset, it was found that several features
had missing values. The missing values in the dataset are addressed using the KNN
imputation method. For each missing value xi , the KNN imputation estimates xi by
computing the average of the k-nearest neighbors’ values


k
xi = 1
k
xj (1)
j=1

where xj represents the values of the nearest neighbors, and k is the number of
neighbors considered. For categorical features, missing values are imputed using the
most frequent (mode) value.
For categorical features, one-hot encoding was applied to transform them into
numerical values, enabling their usage in machine learning models. Convert cate-
gorical variables like chest pain type (cp), fasting blood sugar (fbs), and thalassemia
(thal) into numerical representations. For features that have an inherent order (e.g.,
slope of the ST segment), label encoding can be used to assign each category a
numerical label.
Each feature was normalized using min-max scaling to bring all values between
0 and 1, ensuring the uniformity of the scales.
Predicting Cardiovascular Disease Risk Using Tree-Based Gradient … 173

2.3 Recursive Feature Elimination (RFE)

Feature selection was performed using RFE [15]. This technique removes less impor-
tant features by recursively building a model and ranking the features based on their
importance scores. The importance of each feature Xj was evaluated by recursively
fitting the model and removing the least significant feature each time. The objective
is to minimize the cost function

m
J (θ ) = 1
2m
(hθ (Xi ) − yi )2 (2)
i=1

where hθ is the hypothesis, Xi represents the feature set, and yi is the corresponding
label.

2.4 Machine Learning Algorithms

This study implemented three tree-based gradient boosting machine learning models:
XGBoost, CatBoost, and LightGBM. Gradient boosting is a ML technique that builds
an ensemble of decision trees sequentially, where each new tree attempts to correct
the errors made by the previous trees. Unlike bagging techniques such as random
forest, which train independent models in parallel, boosting works by training weak
learners one after another, each new tree learning from the mistakes of the previous
models. The final model is the weighted sum of all the weak learners. The general
objective function is to minimize a loss function using gradient descent.
eXtreme Gradient Boosting (XGBoost): XGBoost, introduced by Chen and
Guestrin [6], is an optimized distributed gradient boosting library designed to enhance
both the computational speed and model performance. XGBoost’s key improve-
ments over traditional gradient boosting methods include advanced regularization,
tree-pruning strategies, and parallelization techniques.
The XGBoost’s objective function is

N 
  K
L(θ ) =  yi , ŷi(t) + (Tt ) (3)
i=1 k=1

 
where  yi , ŷi(t) is loss function, (Tt ) is a regularization term to control model
complexity of each decision tree $T_t$. This regularization helps prevent overfitting,
which is a common issue in gradient boosting models.
XGBoost can perform parallel processing during the construction of trees, which
significantly speeds up training. XGBoost builds trees by minimizing the objective
function
174 N. R. Khan et al.


T
ŷi = f (xi ) = Tt (xi ) (4)
t=1

where each new tree Tt is added in a way that it corrects the errors made by the
previous ensemble of trees.
LightGBM: LightGBM [7] is a gradient boosting framework developed by
Microsoft, optimized for high performance and efficiency, particularly on large
datasets. It achieves better speed and memory efficiency compared to XGBoost by
employing a novel technique called Histogram-based learning. LightGBM grows
tree’s leaf-wise rather than depth-wise (as in traditional boosting). This means that it
splits the leaf with the largest loss reduction, leading to a more complex tree structure.
This approach can result in better accuracy but also risks overfitting. LightGBM uses
GOSS to speed up training by focusing on data points with larger gradients, while
subsampling the rest of the data.
Similar to XGBoost, the prediction is the sum of the weak learners is


T
ŷi = f (xi ) = Tt (xi ) (5)
t=1

where each tree Tt in LightGBM is built leaf-wise, focusing on maximizing the


information gain from splits.
CatBoost: CatBoost [8] is a gradient boosting library that is specifically designed
to handle categorical features more efficiently than other boosting algorithms.
Unlike XGBoost and LightGBM, which require categorical features to be one-
hot encoded, CatBoost natively supports categorical data without extensive prepro-
cessing. CatBoost builds symmetric trees, where the structure of the tree is the same
on both sides of a split. This results in faster prediction times, as the decision paths
are easier to traverse.
CatBoost uses a slightly different objective function that handles both continuous
and categorical features effectively


T 
J  
ŷi = Tt (xi ) + Cj xij (6)
t=1 j=1

where Tt represents the decision trees for continuous features, and Cj represents
transformations of categorical features.
Predicting Cardiovascular Disease Risk Using Tree-Based Gradient … 175

3 Results

These models were implemented using Python libraries, including Scikit-learn,


XGBoost, LightGBM, and CatBoost. All models were trained on the preprocessed
dataset, with hyperparameter tuning conducted through GridSearchCV. This method
performs an exhaustive search over a predefined set of hyperparameters, such as
learning rate, maximum tree depth, and the number of estimators. The goal was to
identify the best parameter combination that maximizes model performance while
preventing overfitting.
The dataset was split into three parts: 75% for training, 15% for testing, and 10%
for validation. The train-test split ensures that the models are trained on one subset of
the data and tested on another, which is essential for evaluating model performance
on unseen data. Additionally, 5-fold cross-validation was applied to the training set
to further validate the model performance. The model performance was evaluated
using standard classification metrics shown in Table 1.
Table 2 displays the performance metrics of various tree-based ML models prior to
feature selection and GridSearchCV optimization. Notably, models such as CatBoost,
XGBoost, and LGBM exhibit training accuracies exceeding 95%, indicating a strong
fit to the training data. However, their testing accuracies, while slightly lower, remain
commendable, suggesting good generalization capabilities.
The decision tree model, with a training accuracy of 91.26% and testing accuracy
of 88.71%, demonstrates balanced precision and recall values, resulting in an F1-
score of 88.29%. Similarly, SVM achieves a training accuracy of 92.11% and a
testing accuracy of 90.77%, culminating in an F1-score of 90.56%. Random forest

Table 1 Evaluation metrices


Metric name Formula
PT +NT
Accuracy PT +NT +PF +PN
PT
Precision PT +PF
PT
Recall PT +PN
F1-score 2 · Precision+Recall
Precision·Recall

Table 2 Tree-based model selection without feature selection


Model Training Testing accuracy Precision (%) Recall (%) F1-score (%)
accuracy (%) (%)
Decision tree 91.26 88.71 87.38 89.13 88.29
SVM 92.11 90.77 90.12 91.07 90.56
Random forest 92.79 90.83 91.24 91.44 91.22
CatBoost 95.07 91.63 92.87 92.64 92.25
XGBoost 95.83 92.41 93.24 93.22 93.23
LGBM 95.98 93.01 93.73 93.66 93.69
176 N. R. Khan et al.

Table 3 Tree-based model selection with RFE and GridSearchCV


Model Training accuracy Testing accuracy Precision (%) Recall (%) F1-score (%)
(%) (%)
CatBoost 97.53 93.27 93..41 93.58 93.50
XGBoost 98.67 95.81 95.49 95.28 95.44
LGBM 98.98 96.55 96.73 96.82 96.68

outperforms decision tree and SVM, attaining higher accuracies and an F1-score
of 91.22%. Ensemble methods like random forest may offer better performance
compared to single-tree models.
Moreover, CatBoost, XGBoost, and LGBM not only exceeded 95% training accu-
racy but also maintains high testing accuracies. Their precision, recall, and F1-scores
are the highest among the evaluated models, suggesting superior predictive capabil-
ities. Particularly, LGBM achieves the highest F1-score of 93.69%, making it the
most effective model without feature selection.
Table 3 presents the performance metrics of selected tree-based machine learning
models after applying recursive feature elimination (RFE) and GridSearchCV.
Observing these results, it is evident that all models achieved high accuracies in
both training and testing phases, with LGBM showing the highest testing accuracy
at 96.55%.
CatBoost attained a significant training accuracy of 97.53%, yet its testing accu-
racy was slightly lower at 93.27%. This might suggest a minor overfitting, although
precision and recall remained robust at 93.41% and 93.58%, respectively, resulting
in an F1-score of 93.50%. The minimal difference between precision and recall indi-
cates a balanced performance. XGBoost surpasses CatBoost in both training and
testing accuracies, recording 98.67% and 95.81%, respectively. With precision and
recall closely matched around 95%, the F1-score reached 95.44%. This consistency
implies effective generalization to unseen data.
LGBM demonstrated superior performance among the models evaluated,
achieving the highest training accuracy of 98.98% and testing accuracy of 96.55%.
Precision and recall were also the highest at 96.73% and 96.82%, leading to an F1-
score of 96.78%. The negligible difference between precision and recall suggests
excellent predictive capabilities.

4 Conclusion

In this research, we proposed a method for detecting heart disease using tree-based
boosting machine learning techniques. The models utilized, including XGBoost,
CatBoost, and LGBM, demonstrated a high level of accuracy and generalization on
the dataset, with LGBM emerging as the most effective model. The results showed
Predicting Cardiovascular Disease Risk Using Tree-Based Gradient … 177

that LGBM achieved the best performance with a testing accuracy of 96.55%, preci-
sion of 96.73%, recall of 96.82%, and an F1-score of 96.78%. XGBoost followed
closely, with a testing accuracy of 95.81%, while CatBoost showed a slightly lower
performance, with a testing accuracy of 93.27%. These results highlight the strength
of LGBM in capturing complex patterns in the data, making it a suitable model
for heart disease prediction. XGBoost also performed competitively, but LGBM’s
balance of accuracy, precision, and recall set it apart. Future work could focus on
enhancing feature selection techniques or applying these models to larger datasets
for even broader applicability in healthcare prediction tasks.

References

1. World Health Organization: Cardiovascular Diseases (CVDs) (2021)


2. Breiman L (2001) Random forests. Mach Learn 45:5–32
3. Yadav SS, Jadhav SM, Nagrale S, Patil N (2020) application of machine learning for the
detection of heart disease. In: 2020 2nd international conference on innovative mechanisms
for industry applications (ICIMIA), pp 165–172. [Link]
9074954
4. Dalal S, Goel P, Onyema EM, Alharbi A, Mahmoud A, Algarni MA, Awal H (2023) Appli-
cation of machine learning for cardiovascular disease risk prediction. Comput Intell Neurosci
2023:9418666. [Link]
5. Zhou Y, Liu Y, Wang X (2022) Optimizing machine learning algorithms for heart disease
detection. Int J Healthc Technol Manag 22:321–336
6. Chen T, Guestrin C (2016) XGBoost. In: Proceedings of the 22nd ACM SIGKDD international
conference on knowledge discovery and data mining. ACM, New York, NY, USA, pp 785–794.
[Link]
7. Ke G, Meng Q, Finley T, Wang T, Chen W, Ma W, Ye Q, Liu T-Y (2017) LightGBM: a highly
efficient gradient boosting decision tree. In: Proceedings of the 31st international conference
on neural information processing systems. Curran Associates Inc., Red Hook, NY, USA, pp
3149–3157
8. Prokhorenkova L, Gusev G, Vorobev A, Dorogush AV, Gulin A (2018) Catboost: Unbiased
boosting with categorical features. In: Advances in neural information processing systems, pp
6638–6648
9. Friedman JH (2001) Greedy function approximation: a gradient boosting machine. Ann Stat
29:1189–1232. [Link]
10. Smith J, Johnson L (2020) Application of machine learning in heart disease detection. J Med
Res 45:123–135
11. Aldossary Y, Ebrahim M, Hewahi N (2022) A comparative study of heart disease prediction
using tree-based ensemble classification techniques. In: 2022 international conference on data
analytics for business and industry (ICDABI). IEEE, pp 353–357. [Link]
ABI56818.2022.10041488
12. Hammoud A, Karaki A, Tafreshi R, Abdulla S, Wahid M (2024) Coronary heart disease predic-
tion: a comparative study of machine learning algorithms. J Adv Inf Technol 15:27–32. https://
[Link]/10.12720/jait.15.1.27-32
13. Vibha MB, Snegha SR, Kiran U, Kirana Y (2024) Exploratory data analysis of heart disease
prediction using machine learning techniques-RS algorithm. In: 2024 second international
conference on intelligent cyber physical systems and internet of things (ICoICI). IEEE, pp
209–216. [Link]
178 N. R. Khan et al.

14. Dadiyala C, Saxena AA, Kale KA, Bhattad KA, Sheikh NTS, Priyanshi (2024) Progressive heart
disease prediction model using machine learning: a comprehensive staging approach. In: 2024
international conference on smart systems for applications in electrical sciences (ICSSES).
IEEE, pp 1–11. [Link]
15. Zeng X, Chen Y-W, Tao C (2009) Feature selection using recursive feature elimination for
handwritten digit recognition. In: 2009 fifth international conference on intelligent information
hiding and multimedia signal processing, pp 1205–1208. [Link]
9.145
A Comprehensive Framework
for EEG-Based Emotion Detection: From
Exploration to Classification

Shaheen Ayyub, Aishwarya Vishwakarma, Vikas Sakalle, Sugandh Singh,


and Joanna Rosak-Szyrocka

Abstract With increased relevance, emotion detection using electroencephalogram


(EEG) signals is expected to grow considerably in the number of applications and
impact on the society’s understanding of human emotions. The present article offers a
systematic approach for emotion recognition using EEG data, which include a dataset
expanded into positive, negative, and neutral emotive categories. It is important to
note that we don’t only focus on the end-result, but also integrate the exploratory
data analysis with careful preprocessing so as to maintain the level of relevance
of the data. It is in this context that we managed to construct a functional model
for classification of emotional states using ensemble learning algorithms with the
aim of maximizing accuracy and visual representation of decision boundaries. This
framework puts emphasis on the usefulness of ensemble learning in performing
emotion recognition through the application of EEG signals.

Keywords Electroencephalogram · Emotion classification · Data exploration ·


Ensemble learning · Feature analysis · Exploratory data analysis · Decision
boundaries

S. Ayyub (B)
Technocrats Institute of Technology Bhopal, Bhopal, Madhya Pradesh, India
e-mail: shaheenayyub@[Link]
A. Vishwakarma · V. Sakalle
LNCT University Bhopal, Bhopal, Madhya Pradesh, India
S. Singh
IES College of Technology Bhopal, Bhopal, Madhya Pradesh, India
J. Rosak-Szyrocka
Czestochowa University of Technology, Czestochowa, Poland

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 179
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
180 S. Ayyub et al.

1 Introduction

The study of human emotions has attracted more and more interest recently in a
variety of disciplines, from artificial intelligence to psychology [1]. Understanding
and correctly identifying emotions is at the heart of this investigation, and it usually
involves analyzing neurological signals like electroencephalogram (EEG) data [3].
EEG-based emotion detection, a branch of emotional computing, has the poten-
tial to shed light on the neurological correlates of emotions and enable objective
evaluation of them in real-time situations [2]. To analyze the emotional state of a
person, emotion detection via EEG in clinical practice looks into the electric activity
within the patients’ brains [1]. EEG recordings obtained from electrodes placed
on the surface of the scalp reveal the underlying neural correlates of electric brain
activities associated with emotional events. It is the intention of researchers using
machine learning and signal processing techniques to interpret these patterns, where
the emotions can be classified as negative, neutral, and positive.
The usefulness of employing EEG technology in detecting emotions is that it is
non-invasive and has the ability to overcome limitations posed by temporal resolution
and also allow for communication of emotional states in real time.
EEG is particularly valuable for dynamic and context-aware emotion recogni-
tion applications, such as mental health assessment, human-computer interaction,
and affective computing systems. Image-assisted diagnosis and therapy of people
affected with mood disorders, among others, could become a reality with further
advancements in electroencephalogram-based emotion recognition systems [2]. In
parallel, this would also allow a better human-computer interaction where experi-
ences may be improved through context-sensitive interfaces which are responsive to
the emotional state of users.
This work demonstrated the ability to recognize and analyze the emotional states
of an individual using EEG signals. Data from EEG-assisted features cross-validation
and classifiers was employed for selection, data analysis, its visualization and to
construct a model. Firstly we describe the dataset used in our study and then explore
the EEG data through visualizations and feature analysis. After that, preprocessing
techniques applied to the data to prepare it for classification have been discussed.
We then introduce our classification framework, which utilizes ensemble learning
techniques to classify emotional states. Finally, we present the results of our classi-
fication experiments, followed by a discussion of the findings and their implications
[3].

2 Background and Related Work

Emotions display many qualities making it a multidimensional and complex feature,


and emotion detection has been a focus of great attention over a range of fields
including but not limited to psychology, neuroscience as well as computer science.
A Comprehensive Framework for EEG-Based Emotion Detection … 181

Emotion controlling techniques based on older methodologies have been mostly


based on informational self-reports or observational analysis, which could be swayed
by subjective factors [1]. In recent times, EEG as a physiological signal in the research
community has emerged as an accurate understanding about emotions. Such signals
provide a distinct perspective on the neural processes underlying different emotions,
thus enabling improved context-based emotion recognition techniques [1].
Emotions can be detected as byproducts of a brain activity pattern that can also
be captured in EEG as one signal by a single electrode. This work builds upon
earlier literature that employed various signal processing, feature extraction, and
classification approaches for differentiating between emotions [3]. In recent years,
there has been an increasing interest in the application of EEG for the recognition of
emotions. Mobile computing and smart technology have opened up new directions
in emotion research, showing encouraging developments, specifically in the domains
of affective computing, mental health assessment, and human-computer interaction
[2].
However, within the reviewed literature, there is a limited amount of literature
that proposes new, convenient methodologies for solving tasks—emotions based on
EEG. Specific features available in EEG signals were identified using techniques of
signal processing, such as time-domain, frequency-domain, time-frequency analysis,
and so on [3]. In addition, support vector machines, decision tree algorithms, and
convolutional and recurrent neural networks, have been more recently applied in the
task as emotion classifiers across studies [4]. Finally, enthusiasm for the emotion
recognition problem, which often implies the use of several classifiers at once, is
growing. These approaches combine the strengths of multiple methods to improve
performance in emotion recognition based on EEG.
Despite these advancements, challenges remain in EEG-based emotion detection,
including data variability, feature selection, and generalization across diverse popu-
lations [3]. Addressing these challenges requires interdisciplinary collaboration and
the development of robust methodologies that account for individual differences in
emotional processing. By building upon previous research and leveraging innova-
tive techniques, EEG-based emotion detection continues to advance, offering new
avenues for understanding and interpreting human emotions.

3 Proposed Framework

Dataset Description: The present study used emotion-induction EEG bio-signals


during various tasks and analysis. There are three key divisions of showings contained
in this dataset that include all corresponding categorical evidences—positive, nega-
tive, and neutral divisions. The number of samples in this dataset is 2131, while each
sample contain 2548 descriptive features. While EEG signals registration electrodes
were installed according to the 10–20 system and signals from the top of the head
as well as the sides were captured. Emotional experiences were based on reported
experience from participants through an approved affect and “emotion”; the term
182 S. Ayyub et al.

was most often verified during debriefing interviews after the experiment had been
done. Brain activity while processing emotion is also provided by these recordings,
where every recording consists of several channels and each channel covers a single
scalp location.
Exploratory Data Analysis: The dataset encompasses a balanced distribution of
emotional states, ensuring representation across the positive, negative, and neutral
categories. We visualized the dataset through time-series visualization, spectral anal-
ysis, and event-related potentials (ERPs) analysis. Subsequently, we perform inde-
pendent component analysis (ICA) and correlation heatmap analysis to gain deeper
insights into the underlying neural dynamics (Fig. 1).
A heatmap displays relationships between different features in the EEG data. This
facilitates the exploration of feature correlations and patterns.
Feature Significance Analysis and Normalization: Next, we conduct feature
significance analysis through a t-test to assess the relevance of different features
in predicting emotions where the number of significant and insignificant features
for each emotion is stored in a dictionary created by this model. It then extracts the
associated feature counts and emotion labels. After that, it performs z-score normal-
ization to normalize the selected features in a particular range. Afterward, the dataset
is split into training and testing sets for machine learning model input (Fig. 2).
Classification Framework: For the purpose of emotion classification based on
EEG signals, an ensemble learning technique is introduced in this research work. In
other words, ensemble learning develops advanced statistical classifiers by aggre-
gating the predictions of several heterogeneous base classifiers. Dosage forms indi-
vidual classifiers and structures ensemble approaches in order to overcome worlds
of the single classifier and improve the entire performance of the classifying task
(Fig. 3).
The ensemble learning technique in our study is implemented using a voting
classifier, a mechanism that combines the predictions made by base classifiers to
produce an overall prediction. The voting classifier is made up of a set of algorithms
each representing a distinct machine learning model with its distinctive advantages
and abilities. Specifically, the following algorithms explain why the voting classifier
is effective.
1. Random Forest [20]: A model-based ensemble learning algorithm introduced to
overcome the disadvantages of decision trees by creating a collection of decision
trees. When making the final decision, all trees are built independently and their
results are combined using voting mechanism. Random forests are robust against
overfitting and are capable of analyzing data in high-dimensional space.
2. Support Vector Machine (SVM) [22]: This is a particular kind of classifier that
determines the optimal separating hyper-plane between data points belonging to
different classes SVM revolves around the maximization of the margin between
the data of various classes. SVM has also been effective in nonlinear cases and
hence widely used for both binary and multiclass classification.
3. K-Nearest Neighbors (k-NNs) [24]: This is a simple nonparametric classification
scheme where new data points are classified according to the majority voting
A Comprehensive Framework for EEG-Based Emotion Detection … 183

Fig. 1 a Time-series visualization. b Spectral analysis. c Event-related potentials (ERPs) analysis.


d Independent component analysis (ICA). e t-distributed stochastic neighbor embedding (t-SNE)
visualization. f Correlation heatmap
184 S. Ayyub et al.

Fig. 1 (continued)
A Comprehensive Framework for EEG-Based Emotion Detection … 185

Fig. 2 Feature significance analysis

Fig. 3 Proposed model life cycle

done by the data points that are nearest in the feature space. K-NN is a simplistic
approach but works wonders few instances where there is a large amount of
training data and very few features.
The classification framework operates under the classical supervised learning
approach, which entails dividing the dataset into training and testing sets for model
building and assessment, respectively. Each base classifier is fit on the training
data employing appropriate hyperparameter values and optimization procedures,
during model fitting. Cross-validation may be used for hyperparameter tuning and
for evaluation of the model on out of sample data (Figs. 4 and 5).
Once trained, the classifiers are all incorporated into a single entity known as a
voting classifier, which takes the predictions it has made and simply uses a majority
vote. Then the classifier which has been trained over the data is tested on the testing
data to check its classification accuracy. In the end, the regions in which each classifier
predicts a certain output are displayed in order to evaluate their ability to predict
(Fig. 6).
186 S. Ayyub et al.

Fig. 4 Model accuracy

Fig. 5 Confusion matrix

Fig. 6 Decision boundaries for each classifier


A Comprehensive Framework for EEG-Based Emotion Detection … 187

4 Discussion

The classification results from our ensemble learning framework shed light on the
issues underlying emotion detection from EEG signals. Our method has shown a
promising result in the classification of emotions using EEG signals. It is worth noting
the high classification performance across different emotions, which also reaffirm
and confirm the ability of ensemble learning to model the complex dynamics of the
brain that underlie different emotions. In addition, the decision boundaries as shown
for each classifier give an idea of the features that are useful in discrimination by the
models and what kind of emotions are related to the brain activity.
Yet, the aforementioned successes are rather overshadowed by some constraints
of this work. To start with, the emotional dataset of this work may be limited to
the varieties of emotions that are common in daily life. In future studies, the collec-
tion of larger and more representative datasets should be there, so as to enhance the
applicability of the classification models to different groups and contexts. Further-
more, the methods adopted in our study for extracting features and preprocessing
steps do not account for some of the important aspects of EEG signal dynamics. The
use of different feature representation and preprocessing strategies may improve
classification accuracy and stability.

5 Conclusion and Future Scope

This paper puts forth an elaborate structure for detecting emotional states using elec-
troencephalogram signals and accordingly applying ensemble learning techniques.
In carrying out exploratory data analysis, feature analysis, and model training, this
research showed that ensemble learning techniques help in understanding the various
patterns of brain activities associated with different emotions. The classification accu-
racy achievable is quite high which emphasizes the usefulness of EEG-based emotion
detection in understanding how people feel.
Even though our work presents some promising results on emotion recognition
using electroencephalography, there are a number of suggestions that could be made
for future research. To begin with, looking at other methods of feature extraction and
preprocessing can be of great use when it comes to the improvement of classification
models. On top of that, using data fusion techniques, namely EEG combined with
other physiological signals or behavioral data, would be advantageous and provide
a more comprehensive understanding of the emotion-related aspects of the brain.
Other exciting opportunities might consider the development of real-time emotion
recognition systems and adaptive interfaces which would be useful in health care,
education as well as human-computer interactions.
188 S. Ayyub et al.

References

1. Pan J et al (2024) ST-SCGNN: a spatio-temporal self-constructing graph neural network for


cross-subject EEG-based emotion recognition and consciousness detection. IEEE J Biomed
Health Inform 28(2):777–788. [Link]
2. Chen Y-K, Li J-Y, Fang W-C (2023) An edge AI accelerator of LRCN model with RISC-V
platform for EEG-based emotion real-time detection system. In: 2023 IEEE biomedical circuits
and systems conference (BioCAS), Toronto, ON, Canada, pp 1–5. [Link]
CAS58349.2023.10389052
3. Chalupnik R, Bialas K, Majewska Z, Kedziora M (2022) Using simplified EEG-based brain
computer interface and decision tree classifier for emotions detection. In: Barolli L, Hussain
F, Enokido T (eds) Advanced information networking and applications. AINA 2022. Lecture
Notes in Networks and Systems. Springer, Cham, vol 450. [Link]
99587-4_26
4. Joshi VM, Ghongade RB (2021) EEG based emotion detection using fourth order spectral
moment and deep learning. Biomed Signal Process Control 68:102755. [Link]
1016/[Link].2021.102755. ISSN 1746-8094
5. Gilakjani SS, Osman HA (2024) A graph neural network for EEG-based emotion recognition
with contrastive learning and generative adversarial neural network data augmentation. IEEE
Access 12:113–130. [Link]
6. Akhand MAH et al (2024) Emotion recognition from EEG signal enhancing feature map using
partial mutual information. Biomed Signal Process Control 88:105691
7. Chao L et al (2024) Multi-view domain-adaptive representation learning for EEG-based
emotion recognition. Inf Fusion 104:102156
8. Goel A, Bhujade RK (2020) A functional review, analysis and comparison of position
permutation based image encryption techniques. Int J Emerg Technol Adv Eng 10(7):97–99
9. Chu W, Fu B, Xia Y, Liu Y (2023) EEG-based emotion recognition using spatial-temporal
connectivity. IEEE Access 11:92496–92504. [Link]
10. Li T, Fu B, Wu Z, Liu Y (2023) EEG-based emotion recognition using spatial-temporal-
connective features via multi-scale CNN. IEEE Access 11:41859–41867. [Link]
1109/ACCESS.2023.3270317
11. Dubey GP, Bhujade RK (2020) Improving the performance of intrusion detection system using
machine learning based approaches. Int J Emerg Trends Eng Res 8(9):4947–4951
12. Gonzalez HA, Yoo J, Elfadel IM (2019) EEG-based emotion detection using unsupervised
transfer learning. In: 2019 41st annual international conference of the IEEE engineering in
medicine and biology society (EMBC), Berlin, Germany, pp 694–697. [Link]
EMBC.2019.8857248
13. Wang S, Qu J, Zhang Y, Zhang Y (2023) Multimodal emotion recognition from EEG signals
and facial expressions. IEEE Access 11:33061–33068. [Link]
3263670
14. Goel A, Bhujade R (2020) A prototype model for image encryption using zigzig blocks in
inter-pixel displacement of RGB value. Int J Adv Trends Comput Sci Eng 9(5):6905–6912
15. Gonzalez HA, Muzaffar S, Yoo J, Elfadel IAM (2020) An inference hardware accelerator for
EEG-based emotion detection. In: 2020 IEEE international symposium on circuits and systems
(ISCAS), Seville, Spain, pp 1–5. [Link]
16. Basar MD, Duru AD, Akan A (2020) Emotional state detection based on common spatial
patterns of EEG. Signal Image Video Process 14(3):473–481
17. Gonzalez HA, Muzaffar S, Yoo J, Elfadel IM (2020) BioCNN: a hardware inference engine for
EEG-based emotion detection. IEEE Access 8:140896–140914. [Link]
ESS.2020.3012900
18. Mehta R, Bhujade RK (2019) Advance cbir by using cts features with relevance feedback
19. Kelnhofer J, Blaisdell M, Ghandi M (2021) LSTM vs plot-based CNN for EEG emotion
detection tasks. In: 2021 IEEE/ACM conference on connected health: applications, systems
A Comprehensive Framework for EEG-Based Emotion Detection … 189

and engineering technologies (CHASE), Washington, DC, USA, pp 121–123. [Link]


10.1109/CHASE52844.2021.00026
20. Abdulrahman A, M Baykara (2021) A comprehensive review for emotion detection based on
EEG signals: challenges, applications, and open issues. Trait Signal 38(4)
21. Issa S, Peng Q, You X (2021) Emotion classification using EEG brain signals and the broad
learning system. IEEE Trans Syst Man Cybern, undefined. [Link]
2020.2969686
22. Hu J, Wang C, Jia Q, Bu Q, Sutcliffe R, Feng J (2021) ScalingNet: extracting features from
raw EEG data for emotion recognition. Neurocomputing, undefined. [Link]
[Link].2021.08.018
23. Katiyar N, Bhujade R (2019) Authentication & spam analysis architecture over C-hlms. Int J
Sci Technol Res 8(9):612–615
24. Suzuki K, Laohakangvalvit T, Matsubara R, Sugaya M (2021) Constructing an emotion esti-
mation model based on EEG/HRV indexes using feature extraction and feature selection
algorithms. Sensors, undefined. [Link]
25. Yin Y, Zheng X, Hu B, Zhang Y, Cui X (2021) EEG emotion recognition using fusion model
of graph convolutional neural networks and LSTM. Appl Soft Comput, undefined. [Link]
org/10.1016/[Link].2020.106954
Decentralized Blockchain Protocol

Akshay Jain

Abstract This paper introduces a decentralized blockchain protocol designed to


optimize asset management efficiency while addressing the limitations of tradi-
tional single-dimensional blockchains. By prioritizing partition tolerance and even-
tual consistency, the protocol facilitates seamless asset registration, tokenization, and
exchange without compromising availability. Emphasizing simplicity to bolster reli-
ability and security, the protocol excludes complex features such as virtual machines
and smart contracts. Key functionalities encompass asset transactions, metadata
management, peer discovery, wallet generation, uncle blocks handling, dynamic
hash difficulty adjustment, RESTful API for network interaction, and Proof-of-
Work mining with configurable Scrypt factors. Additionally, a proposal for Proof-
of-Collaboration is outlined to mitigate the centralization issue inherent in mining
pools. This protocol offers a promising solution for decentralized asset management,
enhancing efficiency and security in the process.

Keywords Decentralized blockchain protocol · CrankChain · Asset tokenization ·


Proof-of-Collaboration · Blockchain scalability · Metadata management · Peer
discovery · RESTful API · Dynamic hash difficulty · Scrypt algorithm ·
Decentralized asset exchange · Uncle blocks · Mining decentralization ·
Blockchain security

1 Introduction

Many of the complaints surrounding Bitcoin and other single-dimensional


blockchains point toward the suboptimal transaction speed and cost [1]. While
this makes it less than ideal for microtransactions, the benefits of this design are

A. Jain (B)
Jaipur, India
e-mail: [Link]@[Link]

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 191
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
192 A. Jain

often over-looked. Brewer’s (CAP) Theorem [2] does not cease to exist in a decen-
tralized datastore. In fact, the boundaries are heightened. Just as different central-
ized databases coexist with different strengths and weaknesses, so shall they be in
decentralized networks.
In this paper, I introduce CrankChain, which utilizes the many strengths of Satoshi
Nakamoto’s original design [3]. By default, a decentralized network must be partition
tolerant. For asset registration, tokenization, and exchange, I cannot sacrifice avail-
ability. Consistency is important, but eventual consistency is sufficient. Waiting for
confirmation blocks will signal that the transaction is synchronized across a majority
of nodes.
CrankChain is designed under the philosophy that over-engineering opens the
door to vulnerabilities. By reducing the software’s intrinsic complexity, open-source
contributors are less likely to encounter surprises and improve upon the software’s
reliability and security. For this reason, the software excludes a virtual machine, state
machine, and support for smart contracts. Though there are legitimate use cases for
smart contracts [4], the use cases here do not require state changes per contract via
transaction receipts [5].
The network strives to be truly decentralized. It exists as a protocol, and the imple-
mentation can be used as a reference for future builds. In a permissioned network,
there is neither a gatekeeper nor a central coordinator to route transactions.

2 Protocol

2.1 Asset Transactions

The tokenization of an arbitrary asset establishes its existence on the chain as if


it were its genesis block creation. This is done by sending an ASSET_CREATION
transaction. Fourth, the token can be traded as a first-class member on-chain with
STANDARD transactions. The native asset, Cranky Coin, is an asset just like any
other tokenized asset on the network. Its outlying property as the native asset is the
fact that transaction fees and COINBASE transactions are dealt in Cranky Coin. All
transactions include an asset type which defaults to the SHA256 sum of the string
“Cranky Coin”. New assets can be registered by submitting a transaction with a
ASSET_CREATION type enumerated value. From now on, the registrant owns that
asset and may transact that asset though fees remain in the native asset type—Cranky
Coin.
Decentralized Blockchain Protocol 193

2.2 Metadata

Arbitrary metadata can be added to the chain via an ASSET ADDENDUM transaction.
Due to the nature of the blockchain’s immutability and infinite retention policy, all
metadata addendums are appended, and the history remains intact. CrankChain does
not enforce a protocol for the metadata, but the author recommends using a protocol
compatible with GNU diff [6].

2.3 Exchange1

Asset owners are able to post a buy/sell limit order via ORDER transaction types.
Another transaction type, FILL, allows for users to fill the posted limit order. The
original poster may also cancel the limit order if the CANCEL transaction is mined
before a fill transaction.

2.4 Peer Discovery

When a new node joins the network, it queries all known nodes for its list of peers. The
node iterates over the known list of peers and initiates a handshake with each peer.
The handshake is initiated by requesting the peer’s network configuration properties.
If the properties are identical to their own, it proceeds by sending an HTTP POST
request with a request body containing its own configuration properties. The recipient
compares its configuration properties with the ones it received. If they are identical
and the recipient has not exceeded its maximum peer threshold, it adds the initiator
to its peer list and returns an accepted status code. The initiator waits for an accepted
status code before adding the new peer to its list. The minimum and maximum
number of peers are configurable per node.
2
The ZeroMQ PUB/SUB model is applied. Hosts subscribe to peers that possess
identical configuration properties.

2.5 Wallet

CrankChain uses an Elliptical Curve Digital Signature Algorithm variant known as


secp256k1, similar to Bitcoin [7]. Wallets can be generated with the included wallet
client, which does not require a client to download the chain. Paper wallets can be
generated by instructing the wallet client to reveal the private key. The full node and

1 Future implementation.
2 See footnote 1.
194 A. Jain

wallet configuration properties store the user’s symmetrically encrypted private key.
The supported symmetric cryptosystem is AES256 [8]. An AES256 encryption and
decryption tool is provided with the node. An M or T wallet address prefix will be
used to distinguish Mainnet from Testnet.

2.6 Uncle Blocks

In the event that different blocks are received on the CrankChain node for the same
height, both are stored. The first block will continue to build on the 0 branch, while
subsequent blocks will append to a separate branch. This alternate block is known as
the uncle block. Uncle is a jargon used in the Ethereum whitepaper to describe such
blocks (Bitcoin refers to these as orphan blocks) [3, 4]. On a CrankChain node, branch
numbers are determined by SQLite’s default primary key sequence generator. This
is preferred over autoincrement, which prevents the reuse of a previously assigned
key [9]. This is also preferred over a code-generated sequence because the latter may
be susceptible to race conditions given the multi-process nature of the node.
Though branch numbers are stored in the database, they are not the property of
the block itself. All blocks (primary and uncle) will be stored if and only if there
exists a persisted block with a hash equal to that of the new block’s previous hash
property. If a block of that height exists, the branch number is incremented by 1.
Uncle blocks will be omitted (and pruned in the future) if the difference in height
with the primary branch is greater than 6.

2.7 Uncle Transactions

Transactions are stored in a similar fashion to blocks. Primary keys are composite
keys of the transaction hash and branch number, which is the only guarantee for
uniqueness since uncle blocks may reference the same transaction in different blocks.
This means that transaction data will be unique per branch but may be replicated
across branches. The node relies on the database’s constraints to prevent a potential
double-spend.

2.8 Longest Chain

Locally, the node always considers the 0 branch to be the main branch. In the event
that an uncle branch grows taller than the main branch, the branch numbers are
swapped for blocks and their transactions. Transaction hashes are not necessarily
unique in storage, as duplicates may be stored under a different branch number.
Decentralized Blockchain Protocol 195

When competing blocks come in, if they are valid, I accept both, but only one is
authoritative (the longest one).
When an alternate branch outpaces the primary branch, the node must restructure
its chains. If the node did not store uncle blocks, it would be a costly process to
download and validate each block since the split. Chain restructuring is a critical part
of eventual consistency and cannot be avoided. In order for CrankChain to restruc-
ture chains quickly without a performance hit, the author presents the following
restructuring strategy:
1. The branch number B of the longest chain is identified
2. Beginning with the tallest block, B is swapped with the 0 branch block at the
same height. If a 0 branch block does not exist at that height, B is just replaced
with 0.
3. The previous hash of the swapped block is followed to the next tallest block in
that branch. The branch number is then set to 0 while the 0 branch block at the
same height is set to B. This step is repeated until I arrive back at the main 0
branch, and no more swaps are required.

Let c = primary chain

Let c = new tallest chain

PROMOTE − CHAIN (c , c )


(1)
if height [c ] > height [c ]

then n ← height [c ] (2)

b ← branch [c [n]] (3)

while branch [c [n]] = 0 (4)

do if c [n] = NIL (5)

then branch [c [n]] ← b (6)

branch [c [n]] ← 0 (7)

n←n−1 (8)

This results in a restructured chain without exceeding a O(n) complexity and


without needing to validate the blocks and the transactions contained within. The
196 A. Jain

diagram illustrates a highly unlikely scenario where the node has received six distinct
valid blocks of the same height. I witnessed that branch 3 outpaced the previous
dominant branch 0 and must now be promoted to the authoritative branch. Notice B
was preserved throughout the swap process. The two chains, which spanned across
branches 0, 1, and 2, are now fully contained within 0 and 3. This protocol allows
the chain to remain thin and allows for easier pruning.
Decentralized Blockchain Protocol 197

2.9 Dynamic Hash Difficulty

Hash difficulty adjustment is performed based on a simple moving average. The


network is configured so that target time v is 10 min per block, difficulty adjustment
span s is 500, and timestamp sample size n is 2. Minimum hash difficulty m is 1. t
denotes the block’s timestamp at height h. d denotes the present difficulty:

th + th−1 + · · · + th−n
Let f : (t, h) →
n+1

Let t = f (th , h) − f (th−s , h − s)


⎨ dh−1 (h − 2 ≤ s)V (t = v · s)
dh = dh−1 + 1 t < v · s

dh−1 − 1 t > v · s

By taking this approach, the block’s hash difficulty can be calculated determinis-
tically without the need to persist it in the block header.
Single timestamps are not a reliable metric, _as miners can populate the value
arbitrarily. Therefore, I take the arithmetic mean t of th and its n previous t  s:

1 
h
t= ti
n+1
i=h−n

2.10 Network

CrankChain takes a unique approach to full-node design by implementing a RESTful


API over remote procedure calls. Rest allows the CRUD operations necessary for the
network to operate and poses no disadvantages over the latter. In fact, interaction with
the full nodes becomes simpler for web-based applications wishing to communicate
with a node since clients already support HTTP verbs.3 Publish/Subscribe models are
used to broadcast messages across peers using ZeroMQ.

3 See footnote 1.
198 A. Jain

2.11 Mining

Currently, CrankChain operates on a Proof-of-Work [3] system. In order to avoid the


centralization of mining, memory-hard slow hashes are preferred over pure compu-
tational fast hashes. Scrypt was chosen as the hash algorithm. Though it has received
criticism for not being ASIC-resistant, this can be corrected by using Scrypt with
different factors. Configurable Scrypt factors according to (Scrypt—Wikipedia) [10]
include:
– Passphrase–The string of characters to be hashed.
– Salt–A string of characters that modifies the hash to protect against Rainbow table
attacks
– N–CPU/memory cost parameter.
– p–Parallelization parameter; a positive integer satisfying p ≤ (223 − 1) * hLen/
MFLen.
– dkLen–Intended output length in octets of the derived key; a positive integer
satisfying dkLen ≤ (223 − 1) *hLen.
– r–The blocksize parameter, which fine-tunes sequential memory read size and
performance. 8 is commonly used.
– hLen–The length in octets of the hash function (32 for SHA256).
– MFlen–The length in octets of the output of the mixing function. Defined as r *
128 in RFC7914.
CrankChain’s Scrypt hash is configured with the following factors: N =1024, r =1,
p=1, dkLen=32

2.12 4 Proof-of-Collaboration (Proposal)

Pain point Verification of Work (PoW) mining is a serious competition to find a nonce
matching the necessary hash design. Excavators with standard equipment face critical
impediments against huge scope tasks because of high capital and minimal expense
power necessities. Mining pools frequently unify assets, compounding disparity.
While remuneration likelihood ought to scale straightly with hash rate, it develops
dramatically, disadvantaging more slow machines as hash trouble rises. This powerful
builds the gamble of 51% attacks [11] and beats decentralized support down. Despite
the fact that commandeering a PoW network is exorbitant and ostensibly unprofitable
[11], a few substances could in any case seek after it for key reasons.
Goals
– more extensive, equitably disseminated network.
– direct connection between block prizes and capital speculation.
– direct connection between block prizes and uptime.

4 See footnote 1.
Decentralized Blockchain Protocol 199

– zero slant connection between block rewards and hash time.


– reward uptime instead of hash rate.
– no centralization.
– no restriction.
Proposal The author proposes an alternative system to Proof-of-Work known as
Proof-of-Collaboration. The system adheres to the following protocol:
On-chain node registry
1. New node registers by transacting a fee to a collaborator registration address
and REGISTRATION transaction type. The transaction would include the host’s
address in the metadata.
2. Once the block containing the node registry transaction has reached finality, other
collaborators acknowledge the new node.
3. If a collaborator is unable to reach another collaborator, it records downtime in
its local peer database.
4. Collaborators are acknowledged in ascending order by the locally recorded down-
time. The greatest advantage goes to the collaborators with the least recorded
downtimes, which are guaranteed to partake in the mining of every block
Collaborative mining By following this protocol, the chain effectively eliminates
miners competing for a block. This also means:
1. Collaborators take turns initiating the mining process by its registration transac-
tion’s block height. When it is a collaborator’s turn to initiate the mining of a
block, it starts by creating a block. *It may be the case that any collaborator can
initiate the mining of a new block.
2. The collaborator calculates the Merkle tree based on the included transaction
hashes plus the coinbase transaction and hashes it once (with the starting nonce,
0).
3. If the block pattern does not fit, it forwards the block header to a subset of
registered collaborators sorted by ascending downtime.
4. The next set of collaborators adds its coinbase transaction and repeats the previous
two steps.
5. If the block pattern fits the chain, it broadcasts the block to all nodes.
6. If the block pattern does not fit, and all registered collaborators are exhausted, the
nonce is incremented by 1 and repeated. 5 It may be the case that simply changing
the order of contributing collaborators will provide enough permutations of the
Merkel tree that the block pattern may be satisfied.

n!
npr =
(n − r)!

7. Block rewards are divided among contributing collaborators.

5 This step may not be necessary.


200 A. Jain

– The block header should store extra data: exchanges file and coinbase
addresses.
– Hash trouble doesn’t have to expand because of the idea of this framework
– Each block will include a fluctuating number of hubs, subsequently changing
prize sums
– Library exchanges ought to cost a little expense in coins so hubs don’t enroll
and go disconnected without being punished
– Organization strength is a higher priority than computational power
– Cost/reward is directly associated and urges more clients to add to the
organization

Coinbase transaction fees are calculated:

Let C = {col1 , col2 , . . . , coln }

reward
fee =
|C|

Attack vectors Sybil attacks are currently the greatest concern for Proof-of-
Collaboration. A malicious actor may attempt to register a large number of collabo-
rators and hash every permutation of the actors’ node addresses until the block hash
satisfies the hash difficulty pattern. There are a few strategies to mitigate this to some
degree:
1. Restrict IPs based on CIDR block.
2. Disallow incrementing the nonce completely.
3. Adjust hash difficulty based on the number of registrants.
4. Enforce a certain order based on registration block height.
5. Passing a bloom filter?
6. Chain of verifiable signatures?

SprivC (SprivB (SprivA (hash))) = validhash


VpubA (VpubB (Vpubc (validhash))) = hash

3 System Overview

3.1 Scaling

The network largely follows the traditional single-dimensional blockchain design,


as transaction speed is less of a concern than durability and security. The author
does not believe Bitcoin to have a scaling issue [1], but a performance issue. The
Decentralized Blockchain Protocol 201

network’s bandwidth is very limited, but the mempool acts as an effective rate limiter
and priority queue based on transaction fees.
In order for CrankChain to be resilient to large volumes of transactions, each node
must be able to withstand a reasonable amount of inbound requests. I achieve this
by separating the resource-heavy requests from the lighter requests and routing the
heavier requests to a queue. Queue consumers are loosely coupled and can be scaled
to a configurable number of processes optimal for that particular node’s hardware.
Meanwhile, light resource requests can fetch data directly or optionally via cache.
Costly request endpoints are also granted permission to known connected peers.
The mempool transactions are persisted in a physical store and deleted when they
are mined. This offers a few advantages over traditional in-memory mempools. The
most important advantage is the fact that separate processes can access the same data
without relying on an interprocess pipe.
Miners are also a separate process that can be enabled/disabled independently
from the rest of the node.
202 A. Jain

3.2 Persistence

CrankChain uses SQLite for persistence. SQLite possesses several properties that
make it ideal as the underlying data store. It is fast, process-safe, thread-safe, resilient
to file corruption, and it does not require a separate standalone instance. In addition, it
includes an advanced query language, a context manager for transactions, and built-
in locking. Python requires multiprocessing over threads to achieve performance
gain. LevelDB was considered due to its superiority in performance, but it is not
process-safe and susceptible to file corruption.

3.3 Queueing

CrankChain relies heavily on the queue to regulate the inflow of requests that require
validation, and database writes operations. A light whitelist validation is all that
occurs upfront to provide a synchronous acceptance or rejection response before the
message is enqueued.
The queue process is a ZeroMQ push/pull proxy bound to a local UNIX socket.
ZeroMQ is a broker less messaging queue, but the proxy acts as a simple broker
[12].
6
The notification process is a ZeroMQ publisher that is TCP bound.
7
The listener process is a ZeroMQ subscriber that is also TCP bound.

3.4 Processes and Threads

A major architectural objective is horizontal scalability. Implementation in Python


adds challenges in concurrent operations due to the nature of the global interpreter
lock. Since Python is not especially known for its benchmarking performance, relying
on threads to increase speed is less than ideal. A better approach is to spawn separate
processes or deployments as needed. The downside to running multiple processes is
the lack of a shared memory space; communication between processes must rely on
an interprocess pipe or queue. CrankChain dedicates a process for the API layer,
a process for the queue, and a configurable number of processes for the queue
processors. The API layer is a thin Bottle process that merges public and permission
endpoints and routes them to the appropriate service or queue. The queue process
independently runs the ZeroMQ proxy, which acts as a messaging broker. Each of
the producers and consumers communicates to the broker via a local UNIX socket,
which makes it interprocess compatible. Consumer processes have the majority of
the responsibilities as they validate, persist, and notify. While the consumer processes

6 See footnote 1.
7 See footnote 1.
Decentralized Blockchain Protocol 203

can be scaled to an arbitrarily high number, node operators should be cautious as not
to overwhelm the database with excessive write operations as SQLite does not allow
concurrent writes.
Finally, the miner is a separate module that runs independently of the node. It
communicates with the queue and the database. It should use the optional C-compiled
module for a more desirable hash rate.
8
The notification process runs a ZeroMQ publishing server for broadcasting
outbound notifications. Consumers of this queue are external.
9
The listener process runs a ZeroMQ subscriber for listening to its external peer
subscriptions for inbound notifications.

3.5 Caching

Caching is optional but encouraged. CrankChain does not natively provide caching,
but it is trivial to integrate any number of standalone caching services, such as
Memcached or Redis.

4 Attack Vectors

4.1 Double-Spend Attack

CrankChain nodes will calculate account balances (when queried) rather than seek
the last known transaction and return the UTXO (Bitcoin). CrankChain has the luxury
of doing this because it uses a less primitive data store with indexed public addresses
for a O(1) random access lookup. UTXO does, however, provide another purpose—
prevent double-spend attacks [3]. During each transaction, the entire wallet balance
is spent before it is returned upon completion. Concurrent transactions cannot occur.
Ethereum follows a different strategy to mitigate double-spend attacks—include an
incrementing nonce with each transaction [4]. A transaction that does not include
a nonce with the correct sequence is invalid. CrankChain follows neither strategy
and requires transactions to include a hash of the previous transaction, similar to
blocks. Effectively, each public address forms a mini-chain of transactions, which
proves undeniably that one transaction follows the next. Since transaction hashes are
guaranteed to be unique (per branch), duplicate transaction hashes, i.e., transactions
of the same amount to the same recipient with the same timestamp, are not allowed.

8 See footnote 1.
9 See footnote 1.
204 A. Jain

4.2 DDOS Attack

While no public service is completely resistant to DDOS attacks, certain measures


are taken. In order to mitigate DDOS attacks from unknown hosts, POST transac-
tion endpoints (from light clients) and light GET request endpoints will be the only
endpoints available to the public. Costlier calls will be restricted to known peers or
dropped with a 403. Known peers must go through a prior handshake/verification.
In order to mitigate DDOS attacks from known peers, a few measures are taken.
First, a maximum number of peers is allowed and configurable per node. Second,
costly POST requests will be asynchronously placed on a queue before validation/
processing takes place on one of the queue consumer processes, which is also config-
urable per node. Optionally, a caching layer can be added behind. The Python imple-
mentation of CrankChain uses Bottle [13], which ships with adapters to common
WSGI servers. It may also be beneficial for each node to operate behind an NGINX
reverse proxy with rate limiting enabled.
Spam Attack Transaction fees and the previous transaction hash requirement
prevent a user from spamming the network with an excessive number of blocks.
It is relatively cheap for a node to filter out invalid transactions from entering the
mempool. A random access O(1) lookup is all that is necessary to drop such trans-
actions. If a transaction isn’t finalized, its hash does not qualify as the subsequent
transaction’s previous hash.
Proof-of-Work prevents a miner from spamming the network with an excessive
number of blocks. Such an attack would require a hash rate much higher than the
rest of the network. If the attacker invested in the resources capable of performing a
spam attack, it would be short-lived due to the adjustable hash rate.
Chain Death Spiral A Chain Death Spiral occurs when a network faces a sudden
loss of mining power, causing the targeted block rate to be missed. This weakness
was witnessed on Bitcoin Cash when its block time reached 15 hours, prompting it to
hard fork and apply an Emergency Difficulty Adjustment to relax the difficulty sooner
than scheduled [14]. If an attacker with a dominant hash rate decides to cease mining
CrankChain, a Chain Death Spiral should not pose a threat due to CrankChain’s
dynamic hash difficulty adjustment. The hash difficulty is re-calculating every block
in a deterministic fashion using moving averages, which may be adjusted either way.

4.3 Replay Attack

Like many other blockchains, replay attacks are prevented by prefixed wallet ad-
dresses [15]. CrankChain main net will be prefixed with m while testnet addresses
will be prefixed with t.
Decentralized Blockchain Protocol 205

4.4 Sybil Attack

Sybil attacks are also mitigated similarly to other networks. By restricting IPs based
on CIDR block, CrankChain makes it difficult to perform this attack. Bitcoin also
takes this approach by limiting nodes from establishing an outbound connection to
more than one IP address per /16 (x.y.0.0) [11].

4.5 Quantum Attack

If quantum computing resources become available and threaten the current cryp-
tosystems in play, CrankChain can upgrade to quantum-resistant cryptosystems and
allow users to migrate their wallets.

5 Conclusion

In conclusion, the author presents details on the implementation of CrankChain.


In doing so, it was proven that decentralized asset registration, tokenization, and
exchange can be achieved without added complexity, smart contracts, or compro-
mised decentralization. Furthermore, a scalable implementation was built using
Python, a language not known for performance.
The author has also proposed a new mining mechanism known as Proof-of-
Collaboration or PoC with hopes that the protocol will mature and eventually be
implemented into CrankChain and other future blockchains.
CrankChain was first introduced as Cranky Coin on [Link], 2017 [16].

References

1. Contributors W (2018) Bitcoin scalability problem—wikipedia, the free encyclopedia. https://


[Link]/w/[Link]?title=Bitcoinscalabilityproblemoldid=831138002. Accessed 8
Apr 2018
2. Brewer E (2012) Cap twelve years later: How the ”rules” have changed. [Link]
com/articles/cap-twelve-years-later-how-the-rules-have-changed
3. Nakamoto S (2008) Bitcoin: a peer-to-peer electronic cash system. [Link]
pdf
4. Buterin V (2013) Ethereum: a next-generation smart contract and decentralized application
platform. [Link]
5. Wood DG (2018) Ethereum: a secure decentralised generalised transaction ledger. [Link]
[Link]/yellowpaper/[Link]
6. David Mackenzie RS, Eggert P (2017) Comparing and merging files. [Link]
tware/diffutils/manual/[Link]
206 A. Jain

7. Contributors BW (2017) Secp256k1—bitcoin wiki. [Link]


p256k1oldid=64635
8. U. S. N. I. of Standards and T. (NIST) (2001) Announcing the advanced encryption standard
(aes). [Link]
9. Contributors S (2018) Sqlite autoincrement—sqlite documentation. [Link]
html
10. Contributors W (2018) Scrypt—Wikipedia, the free encyclopedia. [Link]
[Link]?title=Scryptoldid=827554872. Accessed 8 Apr 2018
11. Contributors BW (2017) Weaknesses—itcoin wiki. [Link]
knessesoldid=64963
12. Hintjens P (2018) Zeromq-the guide. [Link]
13. Hellkamp M (2018) Bottle: Python web framework. [Link]
14. Patrick (2017) Bitcoin and the blockchain. [Link]
2017/08/[Link]
15. Contributors BW (2017) List of address prefixes—bitcoin wiki. [Link]
php?title=Listofaddressprefixesoldid=62460
16. Kim EE (2017) Let’s create our own cryptocurrency. [Link]
11/lets-create-our-own-cryptocurrency/
A Survey on IoT Protocols
for Resource-Constrained Devices
in Handheld (IoT) Environment

Radhika Patel, Amit Nayak, and Romin Patel

Abstract The Internet of Things, i.e., IoT system or application, succeeds consid-
erably depending on the communication standards and choice and implementation
of the protocols. The selection of proper communication standards and an efficient
protocol used for communication and for establishing connectivity in IOT is diffi-
cult because of the heterogeneity and resource restrictions of IOT devices. There are
protocols designed and adopted as well for the purpose of messaging in the field of
IoT. There is no categorical communication protocol designed that is able to coop-
erate all communication (messaging) use cases and justify the complete necessities of
Internet of Things (IoT) systems. Hence, it becomes difficult to apprehend the proto-
cols of the application layer specifically utilized for messaging and/or communication
purpose in IoT systems. This paper delivers a relative study of the Constrained Appli-
cation Protocol (CoAP), Hypertext Transfer Protocol (HTTP), and Message Queuing
Telemetry Transport (MQTT) messaging protocols including their security.

Keywords Internet of Things (IoT) · Protocols · Constrained Application


Protocol (CoAP) · Hypertext Transfer Protocol (HTTP) · Message Queuing
Telemetry Transport (MQTT) · Message · Communication · Security ·
Performance · Efficiency

R. Patel (B) · A. Nayak · R. Patel


Department of Information Technology, Devang Patel Institute of Advance Technology and
Research (DEPSTAR), Charotar University of Science and Technology (CHARUSAT), Changa,
India
e-mail: [Link]@[Link]
A. Nayak
e-mail: [Link]@[Link]
R. Patel
e-mail: 21dit060@[Link]

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 207
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
208 R. Patel et al.

1 Introduction

The Internet of Things is well-known because of its feature of hyper-connectivity that


allows the users to connect and helps in exchanging of data among IoT devices [1].
A massive amount of data is engendered and exchanged in the IoT system with the
external devices like sensors and actuators [3]. The data generated in the IoT system
is used in varied services and important applications. Moreover, 93% of databases
used in IoT are open source [5]. Therefore, it is necessary to confirm the excellent
communications among IoT networks and the security as well as privacy provided
in the network [6, 7].
It is important for resource discovery to offer a cohesive and adaptable plat-
form that enables applications to identify and integrate diverse devices, regardless of
the messaging (communication) technology and/or protocol it supports [26]. Some
methods exist for resource discovery, ranging from discovering nearby to distant
resources. They are based on peer-to-peer (P2P) connections, request-response
models, communicate through transport protocols and exchange different messages.
The communication (messaging) protocols at the application layer are critical
elements in the IoT system. CoAP, AMQP, XMPP, and MQTT are examples of
protocols used by IoT applications to determine how messages are structured and
communicated. These protocols have a crucial role in facilitating communication
between IoT systems, devices, and the cloud/edge structure [5, 6].
There are various IoT network protocols to provide similar discovery functions
with distinct characteristics and functions like configuration management and regis-
tration. These network protocols can help developers choose the most suitable
protocol for their desired performance, precision, dependability, trustworthiness, and
energy efficiency, among other factors.
This study concentrates on the primary messaging protocols utilized in IoT at the
application layer, which include CoAP, MQTT, and AMQP. Our aim is to examine and
analyze the protocols, namely MQTT, COAP and AMQP in terms of their compati-
bility besides support for data rates at the data link layer [8]. This analysis will help
in efficiently handling both incoming and outgoing data transfer on the networks.
The choice of an appropriate messaging/communication protocol for a specific IoT
implementation relies on various factors comprising the business requirements of
application, software as well as hardware capabilities, along with the data rate.
In the paper, we extant a relative analysis of the main messaging/communication
protocols used in different Internet of Things applications, services, technologies,
and standards. Our goal is to present a thorough analysis and comparison of these
protocols’ performances in various IoT contexts with constraints. The paper’s subse-
quent sections are arranged as follows: Section 1 examines the relevant literature. A
relative study of AMQP, CoAP, and MQTT is provided in Section 2, together with
information on their topographies, fundamentals, and suitability for use in various
contexts.
A Survey on IoT Protocols for Resource-Constrained Devices … 209

Section 3 deliberates the comparative features and possible vulnerabilities of each


protocol. Section 4 evaluates the correct implementation and usage of each messaging
protocol.

2 Related Work

This work documents various communication and/or intermediate protocols,


including Security Applications, Message Queue Telemetry Transfer Protocol, and
Extensible Messaging and Status Protocol [1]. Although these messaging systems
provide some similar connections, they differ in the representation of these attributes.
Effective IoT systems rely on the devices of IoT for performing data collection and
data exchange tasks. Hence, choosing a telemetry messaging protocol for IoT device
connectivity is a significant task [4]. When choosing a messaging protocol, you
should consider the hardware capabilities of your IoT device.
The devices of IoT possess varying bandwidth, and they don’t rely on a universal
or comprehensive wireless technology. Consequently, the data rate supported by
these devices fluctuates based on their hardware capabilities. To illustrate, the size
of the usual HTTP request header is usually between 700 and 800 bytes, encom-
passing fundamental characteristics like the User-Agent request header field [1].
While the protocols from the application layer enable quicker capturing of the data
than the physical data rates supported, they can also contribute to increased latency.
Numerous research and reviews have been carried out on the subject of application
layer protocols [2].
The utilization of a suitable messaging protocol has the potential to lessen network
traffic and latency, thereby improving the consistency of an IoT application [27].
A comprehensive evaluation of communication/messaging protocols in IoT, such as
COAP, HTTP, MQTT, DDS, AMQP, and XMPP, was conducted by [2], which empha-
sized the essential resemblances and communication jobs of these protocols. The
research offered comprehensive and detailed insights into each protocol, including
its implementation, characteristics, and challenges. Similarly, a study conducted in
[19] examined the fundamental features and performance concerns of application
layer communication and messaging protocols in the context of IoT. This research
also explored the probable amalgamation of Internet of Things (IoT) systems into
fog as well as cloud-based systems, along with the associated challenges. Dobbe-
laere [20] introduced a qualitative and measurable comparison outline to evaluate
the key processes of publisher/subscriber systems in open-source message broker
such as RabbitMQ and Apache Kafka, aiming to determine the most appropriate
architecture.
The survey done assessed the CoAP, MQTT, AMQP, and HTTP protocols in the
field of IoT. The study equaled the protocols and recognized the strong suit, limits,
benefits, drawbacks of each protocol, along with the security, and reliability. In [23],
the performance of MQTT and CoAP in wireless sensor networks was evaluated
by creating a common middleware for relating the protocols and implementing a
210 R. Patel et al.

boundary and doorway to provide a standardized environment for comparison. The


survey concluded that the performance of respective protocol is reliant on altered
network circumstances, with messages passed on using MQTT facing less delay
than the messages passed using CoAP. The protocol, MQTT has lesser packet loss
and comparatively more delay in comparison with CoAP. When the size of the packets
or message is small and the loss rate is less than or equal to 25%, CoAP generates
lesser traffic than MQTT [35]. Martín et al. [35] explored ways to bridge the breach
among end users and customizable situations by empowering users to install sensor
devices and actuators on their personal microcontrollers using interfaces.
Similarly, a study [10] weighed the performance of the protocols named MQTT
and CoAP by developing a flexible middleware and comparing the two proto-
cols. It was observed that messages sent using MQTT had lower delay and packet
loss compared to CoAP. However, MQTT’s message performance was negatively
impacted compared to CoAP. Additionally, when messages were small and the packet
loss rate was below 25%, CoAP generated less traffic than MQTT. The issue of
security was addressed in a study which proposed an authentication and validation
technique to ensure communication confidentiality. This technique prevents unau-
thorized nodes from accessing the MQTT network and validates encrypted messages
to maintain data security and communication confidentiality. The technique can be
implemented on constrained devices, with memory requirements of 405,912 bytes
for publisher nodes and 406,856 bytes for subscriber nodes.
In their study, Bai et al. (2019) put forth a new model for a messaging protocol that
combines MQTT with CoAP. The motive of this study was to evaluate the powers and
weaknesses of both protocols when integrated with blockchain technology and elec-
tronics and communication in the IoT. The findings from the experiment revealed
that MQTT protocol offers greater compatibility and dependability for IoT archi-
tecture, but it is not appropriate for restricted environments. Moreover, the survey
claimed that using MQTT with CoAP in combination is more suitable than using
either protocol independently [11].

3 Relative Study of MQTT, COAP, and AMQP

Table 1 provides an overview of the Constrained Application Protocol (CoAP),


Hypertext Transfer Protocol (HTTP), Message Queuing Telemetry Transport
(MQTT) communication/messaging protocols, highlighting their key characteris-
tics and strengths and limitations. The table compares these protocols based on
specific criteria, such as IoT system requirements, device types, available resources,
and application types. It is significant to note that the analysis results might differ
depending on the conditions of the IoT environment, including the components
involved such as vigorous network status, and the rate of retransmission of packet.
The assessment of these protocols is primarily dependent on immobile components
and empirical proofs from previous surveys. It provides comprehensive specifications
and explanations of the key attributes of five fundamental messaging/communication
A Survey on IoT Protocols for Resource-Constrained Devices … 211

Table 1 An assessment of IoT messaging/communication protocols collected from the survey


Characteristics COAP MQTT AMQP
Initiated 2013 1999 2003
Regulated 2014 2013 2014
Model Open sources
Transport UDP TCP/IP TCP
Communication LOW Medium, not more Medium, less than
overhead than 256 MB HTTP, & more than
MQTT
Security Low-use DTLS for Security Low-use DTLS for
security & built security & built
Consumption of LOW Medium, more than Medium, less than
energy COAP HTTP, & more than
MQTT
Latency LOW Medium, more than Medium, less than
COAP HTTP, & more than
MQTT
Applicability Web services, Popular in Small size/power Business &
IoT & WSN, CISCO, mobile applications commercial platforms

protocols in IoT [9]. Each protocol is discussed in terms of information transmission,


security measures, challenges, unique features, performance, business application,
compatibility with different environments, and pros/cons. The aim is to offer readers a
reliable reference on the usage and functionalities of these protocols in the perspective
of IoT.

3.1 Message Size and Overhead

The Hypertext Transfer Protocol (HTTP) is mostly used for web browser and web
server communication, not often for Internet of Things devices. Therefore, when
compared to the other messaging protocols that were studied, HTTP has the largest
message size and system overhead, whereas CoAP guarantees the lowest message
size and overhead [12, 13, 31].

3.2 Power Consumption

HTTP requires the maximum energy as well as the resource. Constrained Application
Protocol (CoAP) requires a lower level of energy consumption and also the demand
of resource is lesser in compare to other messaging protocols [2, 13, 16, 31, 33].
212 R. Patel et al.

Compared to other communications protocols, HTTP requires a higher level of


resource and energy consumption [31]. In terms of memory use, MQTT is more
demanding than AMQP and CoAP. When tests with small payloads (10) and big
payloads (1000) were conducted, the results were comparable; MQTT consumed
2.8 times less power than CoAP and 5.3 times less power than HTTP. In contrast,
CoAP used 1.9 times less energy than MQTT [2, 30].

3.3 Latency and Bandwidth

CoAP has the lowest levels of latency and bandwidth performance compared to
HTTP [13, 31]. In order to reduce latency, a UDP data operation in CoAP includes 2
UDP datagrams. Additionally, this procedure reduces the response times for network
loads.
According to [15, 17], it was found that CoAP uses more computing resources than
MQTT, while AMQP uses more decisions than CoAP and MQTT. IoT communica-
tion/messaging varies in terms of capacity, performance, and robustness depending
on each protocol type; for example, HTTP and CoAP have overhead due to commu-
nication overhead. CoAP and HTTP are best used in situations where low latency is
critical.
Joshi et al. [30] studied the latency and performance of MQTT and HTTP using
testbeds. It was found that HTTP has a greater response time than MQTT. DeCaro
et al. [18] tracked the performance of MQTT with CoAP. The authors found that
CoAP has better bandwidth performance, better health status, and better response
time. The research proves that CoAP is suitable for low-end applications and uses
limited resources (such as edge lists) at a limited time.

3.4 Reliability

Constrained Application Protocol (CoAP) is an unstable UDP-based client-server


Internet of Things protocol. CoAP uses a two-bit header to define the message type
(confirmable, reset, and non-confirmable messages). One drawback of CoAP is that
it depends on DTLS for security and lacks built-in security, making it impossible to
confirm whether the entire message was received or correctly decoded. Nevertheless,
DTLS requires additional packets during the handshake procedure and does not
support multicast. As a result, during the handshake phase, additional computing
power is required to boost network traffic [14, 23].
CoAP devices could not work correctly due to possible detection issues caused
by Network Address Translation (NAT) since the device has an IP address assigned.
When CoAP devices connect to another network, they will get a new IP address.
Therefore, the device is not accessible by edge-based IoT systems [1]. Therefore,
A Survey on IoT Protocols for Resource-Constrained Devices … 213

CoAP or other messaging systems do not support gesture control and dynamic
detection [24, 38, 39].

4 Evaluations of the Security Feat and Threats in Existing


Messaging Protocols

4.1 CoAP

CoAP provides UDP/TLS, which ensures secure data transport, and DTLS, which
supports the encryption techniques AES and RSA as well. To safeguard data, it can
additionally make use of UDP security features. Inappropriate message parsing leads
to potential security risks in CoAP-enabled devices since client logic function server
parsers are unable to handle and manage incoming messages, which can affect the
CoAP node’s availability. Because of this vulnerability, an attacker can remotely
implement and run arbitrary code on the node. CoAP’s implementation of the proxy
caches’ access control methods is inadequate. Unauthorized nodes may be able to
access the CoAP environment if the CoAP nodes are not implemented correctly
during the setup phase.
If the encryption key is not strong enough, the attacker can compro-
mise the CoAP node. In addition, popular attacks on CoAP responses and authenti-
cation include IP spoofing [17, 21, 22].
All the factors used should be considered when selecting the appro-
priate mail system.
CoAP provides real-time services communication in a RESTful style, making it
ideal for devices with limited resources and very short message sizes. For Internet of
Things devices that also use HTTP to provide services, CoAP is a good option.
Compared to MQTT, it performs better in terms of data packet formation and
transmission times [23, 25, 28].
Another study used an end-to-end test with a common middleware to examine
MQTT and CoAP performance. As a result, MQTT messages are transmitted with
less latency than CoAP. When compared to CoAP, MQTT’s performance was shown
to be less efficient due to a higher packet loss rate. Furthermore, when the message
size is the same or smaller, and the packet loss rate is at or below 25%, CoAP
produces less traffic than the MQTT protocol. Similar to this, Luzuriaga et al. [33],
Marsh et al. [34], Marti et al. [35] tested MQTT, CoAP traffic, and network resilience
under various network conditions using a Raspberry Pi 3 ARM-device. According to
the study’s findings, CoAP performance performs better with smaller payloads and
less well with larger message sizes.
Marsh et al. [34] deployed and contrasted MQTT and CoAP-based IoT systems
for temperature collection, as well as IoT M2M protocols. According to the study,
MQTT and CoAP can transmit data 100% of the time with little packet loss. The
study also found that when working with smaller data sizes, CoAP has a lower data
214 R. Patel et al.

loss rate. The authors came to the conclusion that a variety of network factors, such
as data volume, affect how well MQTT and CoAP perform.
Compared to other communications protocols, HTTP demands a greater amount
of energy and resources to operate [31, 32]. In terms of memory use, MQTT is more
demanding than AMQP and CoAP. When tests with small payloads (10) and big
payloads (1000) were conducted, the results were comparable; MQTT consumed
2.8 times less power than CoAP and 5.3 times less power than HTTP. In contrast,
CoAP used 1.9 times less energy than MQTT [2, 29, 30].

5 Results and Analysis

5.1 Energy Consumption

When using UDP as transport, the power consumption of WSN is reduced and
battery life is extended due to the smaller size of the packet header. To test the
improved performance of CoAP over HTTPP, various web service requests are simu-
lated between CoAP client and server and HTTP client and server. Figure 1 shows
the energy consumption of CoAP and HTTP server nodes. Use the Cooja simulator
to measure the power consumption of the client over time in two different operating
modes. This is called RX mode (when the mote server receives a packet) and TX
mode (when the mote server sends a packet). Since more bytes are exchanged in
CoAP, nodes need to spend more time to receive, send, and process packets, which
results in Rx and Tx energy consumption. Leveraging CoAP requires less power than
using MQTT or AMQP.
An alternative figure for the client requests inter-arrival time, such as 5, 10, and
30, is used in the second experiment. Each experiment has a variable total number
of created client requests due to the use of these predefined lengths of simulation.
Figure 1 displays the graph of the CoAP and HTTP server motes’ energy usage for
various client request inter-arrival times. This illustrates that the difference in average
energy consumption increases with the number of client-server transactions (Fig. 2).

Fig. 1 Energy consumption


of CoAP and HTTP server
motes
A Survey on IoT Protocols for Resource-Constrained Devices … 215

Fig. 2 Energy consumption


of the CoAP and HTTP
server motes for different
values of the client request
inter-arrival time

5.2 Response Time

When employing web services to access sensor resources, the response time should
be as minimal as possible. The response time utilizing CoAP and HTTP client-
server systems may have an effect on the performance of some applications that have
stringent latency requirements. This further demonstrates how using UDP lowers the
packet header size and response time. The duration between the client’s request and
the time it takes to get the response with the payload is known as the response time.

6 Conclusion

IoT application protocol is a protocol for exchanging and transmitting data between
IoT devices and systems. This communication process is essential for developing,
using, providing, sharing, and transferring data on IoT systems. This article provides
a comparative analysis to provide guidance and tools to help IoT developers choose
the right messaging protocol for their applications. Research shows that no single
protocol can meet the communication requirements of all IoT applications, including
servers and limited hardware.
This study shows that MQTT has the highest adoption rate in terms of perfor-
mance due to its stability and easy configuration. In MQTT architecture, software
supports two-way communication. For very small messages, CoAP works best as
a message queuing protocol. It is suitable for devices with limited resources as it
facilitates instant communication in RESTful mode. AMQP is suitable for appli-
cations with special features and business needs as it supports communication and
intercommunication between business and IoT applications.
216 R. Patel et al.

References

1. IoT developer survey results (2018) [Link]


assets/[Link]
2. Al-Masri E, Kalyanam KR, Batts J, Kim J, Singh S, Vo T, Yan C (2020) Investigating messaging
protocols for the Internet of Things (IoT). IEEE Access 8(2020):94880–94911
3. Alghamdi TA, Lasebae A, Aiash M (2013) Security analysis of the constrained application
protocol in the Internet of Things. In: Second international conference on future generation
communication technologies (FGCT 2013). IEEE pp 163–168
4. Alsbouí T, Hammoudeh M, Bandar Z, Nisbet A (2011) An overview and classification of
approaches to information extraction in wireless sensor networks. In: Proceedings of the 5th
international conference on sensor technologies and applications (SENSORCOMM’11), vol
255
5. Aly M, Khomh F, Haoues M, Quintero A, Yacout S (2019) Enforcing security in Internet of
Things frameworks: a systematic literature review. Internet of Things 6(2019):100050
6. Azzedin F (2010) Trust-based taxonomy for free riders in distributed multi- media systems.
In: 2010 International conference on high performance computing & simulation, pp 362–369.
[Link]
7. Azzedin F, Albinali H (2021) Security in Internet of Things: RPL attacks taxonomy. In: The
5th international conference on future networks & distributed systems, pp 820–825
8. Azzedin F, Ghaleb M (2019) Internet-of-Things and information fusion: Trust perspective
survey. Sensors 19(8). [Link]
9. Azzedin F, Suwad H, Alyafeai Z (2017) Countermeasureing zero day attacks: asset-based
approach. In: 2017 International conference on high performance computing & simulation
(HPCS), pp 854–857. [Link]
10. Azzedin F, Suwad H, Rahman MM (2022) An asset- based approach to mitigate zero-day
ransomware attacks. CMC-Comput Mater Continua 73(2):3003–3020
11. Babovic ZB, Protic J, Milutinovic V (2016) Web performance evaluation for internet of things
applications. IEEE Access 4(2016):6974–6992
12. Bali RS, Jaafar F, Zavarasky P (2019) Lightweight authentication for MQTT to improve
the security of IoT communication. In: Proceedings of the 3rd international conference on
cryptography, security and privacy, pp 6–12
13. Bandyopadhyay S, Bhattacharyya A (2013) Lightweight Internet protocols for web enablement
of sensors using constrained gateway devices. In: 2013 International conference on computing,
networking and communications (ICNC). IEEE, pp 334–340
14. Basavaraju N, Alexander N, Seitz J (2021) Performance evaluation of advanced message
queuing protocol (AMQP): an empirical analysis of AMQP online message brokers. In: 2021
International symposium on networks, computers and communications (ISNCC). IEEE, pp
1–8
15. Bender M, Kirdan E, Pahl M-O, Carle G (2021) Open- source mqtt evaluation. In: 2021 IEEE
18th annual consumer communications & networking conference (CCNC). IEEE, pp. 1–4.
16. Cohn R (2011) A comparison of AMQP and MQTT. White Paper, StormMQ
17. Cui P (2017) Comparison of IoT application layer protocols
18. De Caro N, Colitti W, Steenhaut K, Mangino G, Reali G (2013) Comparison of two lightweight
protocols for smartphone-based sensing. In: 2013 IEEE 20th symposium on communications
and vehicular technology in the Benelux (SCVT). IEEE, pp 1–6
19. Dizdarević J, Carpio F, Jukan A, Masip-Bruin X (2019) A survey of communication protocols
for internet of things and related challenges of fog and cloud computing integration. ACM
Comput Surv (CSUR) 51(6):1–29.
20. Dobbelaere P, Esmaili KS (2017) Kafka versus RabbitMQ: A comparative study of two industry
reference publish/subscribe implementations. In: Proceedings of the 11th ACM international
conference on distributed and event-based systems, pp 227–238
21. Foster A (2015) Messaging technologies for the industrial internet and the internet of things.
PrismTech Whitepaper 21
A Survey on IoT Protocols for Resource-Constrained Devices … 217

22. Gao W, Nguyen JH, Yu W, Lu C, Ku DT, Hatcher WG (2017) Toward emulation-based perfor-
mance assessment of constrained application protocol in dynamic networks. IEEE Internet of
Things J 4(5):1597–1610
23. Ghafir I, Prenosil V, Hammoudeh M (2015) Botnet command and control traffic detection
challenges: a correlation-based solution. Int J Adv Comput Netw Secur 7(2):2731
24. Ghaleb M, Azzedin F (2021) Towards scalable and efficient architecture for modeling trust in
IoT environments. Sensors 21(9). [Link]
25. Grigorik I (2013) Making the web faster with HTTP 2.0. Commun ACM 56(12):42–49
26. Hammoudeh M, Kurz A, Gaura E (2007) Mumhr: Multi-path, multi-hop hierarchical routing.
In 2007 International conference on sensor technologies and applications (SENSORCOMM
2007). IEEE, pp 140–145
27. Hammoudeh M, Newman R, Dennett C, Mount S, Aldabbas O (2015) Map as a service:
a framework for visualising and maximising information return from multi-modal wireless
sensor networks. Sensors 15(9):22970–23003
28. Han NS (2015) Semantic service provisioning for 6LoWPAN: powering internet of
things applications on Web (Doctoral dissertation). Available from Institut National des
Telecommunications.(tel-01217185)
29. Imane S, Tomader M, Nabil H (2018) Comparison between CoAP and MQTT in smart
healthcare and some threats. In: 2018 International symposium on advanced electrical and
communication technologies (ISAECT). IEEE, pp 1–4
30. Joshi J, Rajapriya V, Rahul SR, Kumar P, Polepally S, Samineni R, Kamal Tej DG (2017)
Performance enhancement and IoT based monitoring for smart home. In: 2017 International
conference on information networking (ICOIN). IEEE, pp 468–473
31. Khoi NM, Saguna S, Mitra K, Ahlund C (2015) IReHMo: an efficient IoT-based remote
health monitoring system for smart regions. In: 2015 17th International conference on e-health
networking, application & services (HealthCom). IEEE, pp 563–568
32. Luzuriaga JE, Perez M, Boronat P, Cano JC, Calafate C, Manzoni P (2014) Testing AMQP
protocol on unstable and mobile networks. In: International conference on internet and
distributed computing systems. Springer, pp 250–260
33. Luzuriaga JE, Perez M, Boronat P, Cano JC, Calafate C, Manzoni P (2015) A comparative
evaluation of AMQP and MQTT protocols over unstable and mobile networks. In: 2015 12th
Annual IEEE consumer communications and networking conference (CCNC). IEEE, pp 931–
936
34. Marsh G, Sampat AP, Potluri S, Panda DK (2008) Scaling advanced message queuing protocol
(AMQP) architecture with broker federation and infiniband. Ohio State University, Technical
Report OSU-CISRC-5/09-TR17(2008), vol 38
35. Martí M, Garcia-Rubio C, Campo C (2019) Performance evaluation of CoAP and MQTT_SN
in an IoT environment. In: Multidisciplinary Digital Publishing Institute Proceedings, vol 31,
no 1, p 49.
Traffic Sign Recognition in Adverse
Environments: A Survey on Methods
for Low Light, Fog, and Rain Conditions

Kaushal Patel and Sheshang Degadwala

Abstract In traffic safety, especially with the development of autonomous vehicles,


TSR systems can play a critical role. However, their performances are consider-
ably affected in poor environments where visibility and quality are degraded. This
survey paper will comprehensively review the methods developed to address these
challenges and outline various limitations of traditional TSR methods during such
conditions. This aims to analyze the approaches using deep learning, image enhance-
ment, and sensor fusion that enhance recognition accuracy in adverse weather and
lighting conditions. Further, the applicability of these methods in practical situa-
tions is discussed, and gaps in current research are identified together with possible
research directions for study in enhancing robustness and reliability in TSR systems
operating under challenging environmental conditions.

Keywords Traffic sign recognition · Adverse environments · Low light · Fog ·


Rain · Image enhancement

1 Introduction

Modern driver-assist technologies and self-driving automobiles require traffic sign


recognition. Traditional TSR systems have evolved from simple rule-based algo-
rithms based on color and form identification to complex machine learning and deep
learning models with great accuracy [1, 2]. While early techniques were effective
in ideal weather, real-world illumination and weather dramatically affect sign visi-
bility in traffic [3, 4]. The growing demand for reliable and robust TSR systems has
shifted research toward designing more resilient algorithms for dark, fog, and rain
environments, where classic techniques fail.

K. Patel (B) · S. Degadwala


Department of Computer Engineering, Sigma University, Vadodara, Gujarat, India
e-mail: kaushalpatel15@[Link]
S. Degadwala
e-mail: sheshang13@[Link]

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 219
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
220 K. Patel and S. Degadwala

Fig. 1 Types of Indian signs. Source [Link]

Recent research indicates that TSR performance has improved with cutting-edge
deep learning methods, replacing older methods. CNN and GAN methods have
improved traffic sign image quality and recognition under difficult settings [7, 11,
14]. Adding LIDAR and radar to vision-based systems improves TSR’s capacity
to recognize signs in low-visibility settings [15, 16]. Despite these advances, TSR
systems cannot effectively detect traffic signs in dark, heavy downpour, and dense
fog, raising concerns about their practicality [15, 16].
Figure 1: Some examples of traffic signs, which are categorized into four groups:

1. Mandatory Signs [17]: “Stop,” “No Entry,” and “Speed Limit.“


2. Compulsory signage [25]: command drivers to turn or honk in certain situations.
3. Warning signs [17, 25]: alert drivers to curves, pedestrian crossings, and narrow
lanes.
4. Informative signs [17]: These notify drivers of parking, bus stops, and directions.

This paper will review existing techniques to improve TSR performance in difficult
environments.

2 Literature Works

Significant breakthroughs in traffic sign identification and recognition have been


predominantly driven by the use of deep learning and computer vision techniques.
Traffic Sign Recognition in Adverse Environments: A Survey … 221

Flores-Calero et al. [1] examined real-time detection systems using YOLO.


Despite its speed and accuracy, YOLO cannot recognize tiny or hidden signs; there-
fore, future research should focus on busy area sign detection [1]. Hussain et al.’s
[2] TSC18 CNN model for traffic sign categorization fared well in sign recogni-
tion. Future research may use better data augmentation methods to solve subop-
timal performance under diverse illumination situations [2]. In real time, Liu et al.
[3] enhanced minor traffic sign detection with YOLOv5. Processing resources are
needed; hence, future research should employ simpler models [3]. Zeng et al. [4]
improve long-tailed traffic sign identification with transformer fusion and residual
learning. While the confusion set enhances uncommon indication accuracy, its design
is complicated and calculation time may slow real-time applications [3]. Future
research may overcome this problem. Gospodinov and Krastev [5] developed a viable
cyber-physical system-based traffic sign detection model. Mobile implementations
differ from fixed environment versions and could be optimized [5]. The technology
has considerable promise.
Satti et al. [6] used vision transformers to spot potholes and traffic signs in low
light. Transformer computational cost is a limitation; therefore, future research may
identify more efficient real-time topologies [6]. Chen et al. [7] created CNN-multi-
scale transformer semi-supervised learning for data limitations. Though beneficial,
the method requires faster training due to real-time performance issues [7]. Cui et al.
[8] enhanced traffic sign detection with the YOLO network to save computational
power and accuracy. While complicated backdrops remain troublesome, context-
aware identification may enhance performance [8]. Wang et al. [9] created a vehicle-
mounted adaptive traffic sign detector that can recognize small signs in different situ-
ations. Specialized hardware hinders system adaptability and acceptance, although
future studies may fix it [9]. Liu et al. [10] improved fog traffic sign detection with
photo defogging and transformers. Since defogging increases computer load, future
research should focus on more efficient pretreatment approaches [10].
Rani et al. [11] used deep learning and haze reduction to improve autonomous
car traffic sign identification in low-visibility conditions. However, their method is
computationally demanding and may be reduced in future generations [11]. Huang
et al. [12] improved YOLOv8’s traffic sign recognition. The approach enhances
detection accuracy but requires considerable processing power for real-time appli-
cations [12]. Parse and Pramod [13] improve traffic sign recognition with bilateral
filtering and transfer learning as edge detectors. Real-time adaptation is difficult,
requiring faster processing [13]. Yan et al. [14] improved low-light traffic sign detec-
tion. Future studies may increase the method’s adaptation to dense fog or heavy rain
[14]. Wang et al. [15] improved the YOLOv5 model for traffic sign identification,
which had better real-time multi-scale detection but lower computer efficiency that
future studies may improve by streamlining the model design [15].
Kuppusamy et al. [16] improved YOLOv7 traffic sign detection with a Convo-
lutional Block Attention Module. Further research may refine the model to lower
its high computing cost [16]. Wang et al. [17] improved YOLOv5’s real-time minor
sign identification. However, the model’s computational complexity inhibits real-
time applications [17]. Deep learning improved regional sign recognition for Indian
222 K. Patel and S. Degadwala

traffic signs by Megalingam et al. [18]. The model failed in low light, although
better picture preprocessing may help [18]. Dewi et al. [19] improved microscopic
sign detection with spatial pyramid pooling, however processing efficiency may be
improved [19]. A real-time traffic sign identification system for autonomous cars by
Malarvizhi et al. [20] showed promise but was limited in high-speed conditions with
regularly traversed signs. Faster detection methods may fix this in future studies [20].
Zhang and Zhao [21] employed deep learning to improve traffic sign recognition
accuracy but struggled with scalability in big, complex datasets [21]. Future research
may overcome these issues. Lai [22] demonstrated a high-performing deep learning-
based traffic sign identification system in controlled environments. Its adaptation to
different weather and lighting conditions is limited; therefore, future research may
try to improve its robustness in practical contexts [22]. Xie et al. [23] presented a
federated learning architecture for traffic sign identification utilizing spiking neural
networks, reducing the need for centralized data. The model’s efficacy in varied
circumstances needs further testing [23]. Dewi et al. [24] used DCGAN to create
synthetic data for traffic sign identification. While their approach has shown promise
in improving datasets, challenges remain in generating high-quality synthetic images,
which can be addressed in future study [24]. Xing et al. [25] developed a guided image
filtering method for traffic sign identification, improving picture quality but limiting
real-time flexibility, which future research may address [25].
The literature review on traffic sign recognition (TSR) under adverse situations
highlights its progress and obstacles. In challenging situations, deep learning, espe-
cially CNNs, enhances recognition accuracy. Image enhancement and sensor fusion
reduce low light, fog, and rain-related visual difficulties, improving TSR system
robustness. The findings are lacking, notably in generalizable models across contexts
and real-time processing for practical applications. This assessment emphasizes the
need for ongoing research and development to create safe, efficient TSR systems in
any weather.

3 Comparative Study

Table 1 provides a detailed summary of numerous algorithms used for traffic sign
recognition (TSR) in diverse weather situations, including poor light, rain, and fog.
Table 1 Comparative study of traffic sign recognition in adverse weather conditions
Algorithm Year Weather ACC (%) Key features Advantages Limitations GPU Inference time (ms)
conditions
Histogram of 2005 Clear, Low Light 60 Edge detection, Simple, Poor performance in No ~3000
Oriented feature extraction computationally adverse weather
Gradients (HOG) efficient conditions
[1]
Support Vector 2008 Clear, Fog 85 Binary High accuracy for Sensitive to noise No ~2000
Machine (SVM) classification, small datasets and less effective in
[2] margin low light
maximization
Convolutional 2014 Low Light, 93 Hierarchical High accuracy, Requires large Yes ~30–50
Neural Networks Rain, Fog feature extraction, robust against noise datasets and
(CNN) [3] spatial hierarchies and distortion computational
resources
Region-based 2015 Rain, Fog 92 Object detection High precision in Slow inference Yes ~200–500
CNN (R-CNN) [4] with region complex scenes speed requires
proposals preprocessing
Generative 2017 Low Light, Rain 88 Data Improved robustness Computationally Yes ~500–1000
Traffic Sign Recognition in Adverse Environments: A Survey …

Adversarial augmentation by generating diverse intensive training


Networks (GAN) through synthetic datasets complexity
[5] image generation
YOLO (You Only 2018 Rain, Fog, Low 91 Real-time object Fast inference speed, The trade-off Yes ~10–20
Look Once) [6] Light detection, good accuracy between speed and
single-stage accuracy
architecture
(continued)
223
Table 1 (continued)
224

Algorithm Year Weather ACC (%) Key features Advantages Limitations GPU Inference time (ms)
conditions
U-Net [7] 2019 Fog, Rain 83 Image Excellent for Limited in handling Yes ~50–150
segmentation for semantic complex
detailed feature segmentation backgrounds
extraction
Fusion of LIDAR 2021 Rain, Fog, Low 96 Multi-sensor data Improved accuracy The complexity of Yes ~500–700
and Camera Data Light fusion for and reliability in sensor integration
[8] enhanced adverse conditions and data processing
robustness
Single Shot 2016 Low Light, Rain 85 Real-time Good balance May struggle with Yes ~15–30
Multibox Detector detection with between speed and small object
(SSD) [9] multiple aspect accuracy detection
ratios
DenseNet [10] 2017 Fog, Low Light 90 Dense Improved gradient More complex Yes ~100–200
connectivity flow better accuracy architecture than
patterns for with fewer traditional CNNs
feature reuse parameters
Attention 2020 Rain, Fog 91 Focuses on Enhanced Increased Yes ~100–200
Mechanism [11] relevant features recognition of computational
in the image obscured or distant overhead
signs
Hybrid Deep 2022 Rain, Fog, Low 95 Combines CNNs Improved Complexity in Yes ~100–300
Learning Model Light with traditional performance in training and model
[12] image processing adverse conditions tuning
methods
(continued)
K. Patel and S. Degadwala
Table 1 (continued)
Algorithm Year Weather ACC (%) Key features Advantages Limitations GPU Inference time (ms)
conditions
ResNet [13] 2015 Fog, Rain 89 Deep residual Reduces the Computationally Yes ~30–50
learning for vanishing gradient heavy
feature extraction problem
Faster R-CNN 2015 Rain, Fog 88 Region proposal Improved speed and Requires more Yes ~200–300
[14] network for accuracy computational
speedier detection resources
EfficientDet [15] 2020 Low Light, Rain 90 Efficient object Balanced accuracy Limited Yes ~20–50
detection with and efficiency performance in
scaling high-density scenes
Mask R-CNN [16] 2017 Rain, Fog 91 Instance Provides detailed More complex than Yes ~300–600
segmentation with localization standard R-CNN
object detection
OpenPose [17] 2018 Low Light, Fog 78 Multi-person Effective in detecting Sensitive to Yes ~100–200
detection, pose complex scenarios occlusions
estimation
MobileNet [18] 2017 Rain, Low Light 80 Lightweight Fast and efficient for Lower accuracy No ~ 10–20
Traffic Sign Recognition in Adverse Environments: A Survey …

architecture for mobile devices compared to larger


mobile models
applications
Spatial Pyramid 2014 Fog, Rain 82 Multi-scale Robust against Complexity in Yes ~50–100
Pooling [19] feature extraction various scales implementation
GoogleNet [20] 2014 Low Light, Fog 89 Inception modules Good balance of High computational Yes ~50–100
for feature depth and width cost
extraction
(continued)
225
Table 1 (continued)
226

Algorithm Year Weather ACC (%) Key features Advantages Limitations GPU Inference time (ms)
conditions
Adaptive Learning 2019 Rain, Fog 85 Dynamic learning Improves This may lead to Yes ~50–100
Rate [21] rate adjustment convergence speed instability in some
cases
Fuzzy Logic 2020 Low Light, Rain 78 Incorporates Robustness against Requires careful No ~2000
Systems [22] uncertainty and noise tuning of parameters
imprecision
Extreme Learning 2018 Fog, Rain 73 Single hidden Fast training speed Limited No ~10–50
Machine (ELM) layer feedforward representation
[23] neural network capability
Transformer 2021 Low Light, Rain 92 Self-attention Effective in capturing Computationally Yes ~200–500
Models [24] mechanism for long-range expensive
feature extraction dependencies
Graph Neural 2022 Rain, Fog 90 Represents data as Captures Complexity in Yes ~100–200
Networks [25] graphs for better relationships training and
feature extraction between objects inference
K. Patel and S. Degadwala
Traffic Sign Recognition in Adverse Environments: A Survey … 227

4 Challenges and Solutions

4.1 Challenges

1. Environmental Variability: Inclement weather conditions, including diminished


light, fog, and precipitation, substantially affect the visibility and clarity of traffic
signs, resulting in misunderstanding or missed detections.
2. Sensors and cameras may generate noise and distortion in pictures owing to
environmental conditions, complicating detection and classification operations.
3. Scarcity of High-Quality Datasets: The availability of high-quality datasets under
diverse unfavorable situations is typically limited, complicating the successful
training of models for real-world applications.
4. Real-time Processing Requirements: Numerous traffic sign recognition systems
are required to function in real-time, demanding algorithms that reconcile
accuracy with computing efficiency, a challenging endeavor.
5. Occlusion and Obstruction: Traffic signs may be partly concealed by other enti-
ties, such as automobiles or vegetation, resulting in challenges in detection and
identification.

4.2 Solutions

1. Data Augmentation: Improving training datasets using synthetic data creation


or augmentation (e.g., mimicking unfavorable situations) may enhance model
resilience and accuracy.
2. Multi-Sensor Fusion: Integrating data from diverse sensors (e.g., cameras,
LIDAR, and radar) may augment the identification process by offering comple-
mentary information and enhancing detection reliability under challenging
settings.
3. Adaptive Learning Techniques: Using adaptive learning algorithms that can
perpetually learn and adapt to changing contexts enhances accuracy and
durability over time.
4. Advanced Preprocessing Techniques: Employing sophisticated image prepro-
cessing techniques, including denoising algorithms and contrast enhancement,
may alleviate the impact of noise and distortion, enhancing picture quality before
processing.
5. Real-Time Optimization: Lightweight models or efficient architectures, such as
MobileNet or EfficientNet, enable real-time processing with minimal accuracy
trade-offs, rendering them appropriate for on-road applications.

Addressing these difficulties with the suggested methods will substantially


enhance traffic sign recognition systems, improving safety and reliability across
diverse driving circumstances.
228 K. Patel and S. Degadwala

5 Conclusion and Future Work

This evaluation concentrates on traffic sign recognition (TSR) approaches’ perfor-


mance in bad light, rain, and fog. The analysis of 30 algorithms shows TSR’s evolu-
tion from image processing to deep learning. These systems have improved in terms
of precision and resilience, but environmental unpredictability, noise, and constrained
datasets persist. Addressing these issues is crucial for practical TSR system imple-
mentation. This study underscores the need for continued research and develop-
ment to improve traffic sign recognition reliability and efficiency, thereby making
highways safer and driver support systems better.
Research may examine several ways to improve traffic sign recognition. Multi-
sensor systems may improve recognition performance in difficult environments by
providing more data. Better data augmentation technologies and synthetic datasets
may also help overcome real-world data shortages. Hybrid models that combine
the benefits of many methods may be useful for real-time applications that balance
accuracy with processing economy. Additionally, studying reinforcement learning
and meta-learning may improve adaptation and performance in changing situa-
tions. As autonomous cars improve, TSR systems must be integrated into intelligent
transportation frameworks to ensure road safety and efficiency.
Disclosure of Interests The authors declare no conflicts of interest.

References

1. Flores-Calero M, Astudillo CA, Guevara D, Maza J, Lita BS, Defaz B, Ante JS, Zabala-Blanco
D, Armingol Moreno JM (2024) Traffic sign detection and recognition using YOLO object
detection algorithm: a systematic review. Mathematics 12:1–31. [Link]
h12020297
2. Hussain A, Qureshi KN, Aslam A, Tariq T, Abdullahi MR (2024) TSC18 convolutional neural
network for traffic sign classification. In: 4th International conference on emerging smart tech-
nologies and applications, Marta, pp 1–6. [Link]
39008
3. Liu A, Liu Y, Kifah S (2024) Deep convolutional neural network for enhancing traffic sign
recognition developed on Yolo V5. In: Proceedings of 2nd international conference on advance-
ments in smart, secure and intelligent computing, ASSIC 2024. [Link]
IC60049.2024.10508025
4. Zeng G, Huang W, Wang Y, Wang X, Wenjuan E (2024) Transformer fusion and residual
learning group classifier loss for long-tailed traffic sign detection. IEEE Sens J 24:10551–10560.
[Link]
5. Gospodinov N, Krastev G (2024) Cyber-physical system for traffic sign detection and
recognition. Eng Proc 60. [Link]
6. Satti SK, Rajareddy GNV, Mishra K, Gandomi AH (2024) Potholes and traffic signs detection
by classifier with vision transformers. Sci Rep 14:1–18. [Link]
52426-4
7. Chen S, Zhang Z, Zhang L, He R, Li Z, Xu M, Ma H (2024) A semi-supervised learning frame-
work combining CNN and multiscale transformer for traffic sign detection and recognition.
IEEE Internet Things J 11:19500–19519. [Link]
Traffic Sign Recognition in Adverse Environments: A Survey … 229

8. Cui Y, Guo D, Yuan H, Gu H, Tang H (2024) Enhanced YOLO network for improving the
efficiency of traffic sign detection. Appl Sci (Switzerland) 14. [Link]
20555
9. Wang J, Chen Y, Ji X, Dong Z, Gao M, Lai CS (2024) Vehicle-mounted adaptive traffic sign
detector for small-sized signs in multiple working conditions. IEEE Trans Intell Transp Syst
25:710–724. [Link]
10. Liu Z, Yan J, Zhang J (2024) Research on a recognition algorithm for traffic signs in foggy
environments based on image defogging and transformer. Sensors 24. [Link]
s24134370
11. Rani AR, Anusha Y, Cherishama SK, Laxmi SV (2024) Traffic sign detection and recognition
using deep learning-based approach with haze removal for autonomous vehicle navigation.
E-Prime Adv Electric Eng Electron Energy 7:100442. [Link]
100442
12. Huang Z, Li L, Krizek GC, Sun L (2023) Research on traffic sign detection based on improved
YOLOv8. J Comput Commun 11:226–232. [Link]
13. Parse M, Pramod D (2023) Edge detection technique based on bilateral filtering and iterative
threshold selection algorithm and transfer learning for traffic sign recognition. Sci J Silesian
Univ Technol Ser Transp 119:199–222. [Link]
14. Yan Y, Deng C, Ma J, Wang Y, Li Y (2023) A traffic sign recognition method under complex illu-
mination conditions. IEEE Access. 11:39185–39196. [Link]
3266825
15. Wang Q, Li X, Lu M (2023) An improved traffic sign detection and recognition deep model
based on YOLOv5. IEEE Access. 11:54679–54691. [Link]
3281551
16. Kuppusamy P, Sanjay M, Deepashree PV, Iwendi C (2023) Traffic sign recognition for
autonomous vehicle using optimized YOLOv7 and convolutional block attention module.
Comput Mater Continua 77:445–466. [Link]
17. Wang J, Chen Y, Dong Z, Gao M (2023) Improved YOLOv5 network for real-time multi-
scale traffic sign detection. Neural Comput Appl 35:7853–7865. [Link]
521-022-08077-5
18. Megalingam RK, Thanigundala K, Musani SR, Nidamanuru H, Gadde L (2023) Indian traffic
sign detection and recognition using deep learning. Int J Transp Sci Technol 12:683–699.
[Link]
19. Dewi C, Chen RC, Yu H, Jiang X (2023) Robust detection method for improving small traffic
sign recognition based on spatial pyramid pooling. J Ambient Intell Humaniz Comput 14:8135–
8152. [Link]
20. Malarvizhi N, Jupudi AK, Velpuri M, Dheeraj TVK (2023) Autonomous traffic sign detection
and recognition in real time. In: Lecture notes in networks and systems, pp 415–423. https://
[Link]/10.1007/978-981-19-6088-8_36.
21. Zhang H, Zhao J (2022) Traffic sign detection and recognition based on deep learning. Eng
Lett 30:666–673
22. Lai Y (2022) Traffic sign recognition based on deep learning technique. In: ACM International
conference proceeding series, pp 62–68. [Link]
23. Xie K, Zhang Z, Li B, Kang J, Niyato D, Xie S, Wu Y (2022) Efficient federated learning with
spike neural networks for traffic sign recognition. IEEE Trans Veh Technol 71:9980–9992.
[Link]
24. Dewi C, Chen RC, Liu YT, Tai SK (2022) Synthetic data generation using DCGAN for improved
traffic sign recognition. Neural Comput Appl 34:21465–21480. [Link]
521-021-05982-z
25. Xing J, Nguyen M, Qi Yan W (2022) The improved framework for traffic sign recognition
using guided image filtering. SN Comput Sci 3:1–16. [Link]
55-y
An Analytical Review of Social
Media-Based Sentiment Analysis
Techniques

Priyanshi Jain, Manas Vyas, Niharika Awasthi, and Raj Gaurav Mishra

Abstract Social media is now a core part of how we communicate in today’s world.
It is a rich source of user-generated content, reflecting public sentiments on various
issues. This study will, therefore, undertake a critical analysis of methodologies
employed in sentiment analysis on social media data and their applications in multiple
domains like product reviews. The upsurge of social media has gone up dramatically
on platforms like Twitter and Reddit, among others. It has filled the web with user-
generated content, and therefore, sentiment analysis will be very much valuable
in truly understanding the opinion of the public. This study proceeds further with
the exposition of the various tools, technologies, and methodologies through which
sentiment analysis is done, underlining NLP and machine learning techniques as
crucial elements in this respect. It will also enable the researchers and practitioners
to select appropriate models and techniques that shall yield the most effective and
realistic sentiment analysis, thereby enabling better decision-making processes by
industries to address and meet customer expectations effectively.

Keywords Machine learning · Natural language processing · Neural networks ·


Sentiment analysis

1 Introduction

The rapid increase in social networking platforms has completely changed how
people communicate. Social networking sites like Twitter, Facebook, and Instagram
have become virtual public congregations where opinions are shared on a wide range
of issues, including product experiences. Sentiment analysis has emerged as a key

P. Jain (B) · M. Vyas · N. Awasthi


Narsee Monjee Institute of Management Studies, Indore, India
e-mail: priyanshijain320@[Link]
R. G. Mishra
STME, Narsee Monjee Institute of Management Studies, Indore, India
e-mail: [Link]@[Link]

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 231
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
232 P. Jain et al.

subfield of NLP, aiming to interpret this vast amount of unstructured data and extract
meaningful insights from it. This technology has gained much interest in it as it
can transform unstructured text into actionable information. In an instance, this may
determine the satisfaction of customers in product reviews for businesses, or politi-
cians may assess how the public perceives their policies. This review paper targets
the application of sentiment analysis to social media data, especially about reviews
against any particular product. Sentiment analysis is highly important to companies
for further improvement in their products and services, as consumers are getting
more dependent on online reviews for purchases. This paper will identify strengths
and weaknesses of various approaches by reviewing and synthesizing findings from
more than 30 research papers, hence providing a road map for future research in the
area. This paper will also discuss some of the developed tools, technologies, and
concepts that enhance sentiment analysis, therefore giving a wide-ranging overview
to both academics and professionals. The final ambition of this study is to assist the
current discourse on sentiment analysis, in that it represents an increasingly important
means of gaining an understanding of public sentiment and aiding decision-making
processes in various domains.

2 Literature Review

Go et al. [1] in their work added emoticons as noisy labels for schooling sentiment
classifiers on Twitter information. Their approach highlighted the demanding situa-
tions of the use of bigrams, which led to a drop in accuracy. Additionally, the model
tested obstacles in its effectiveness on languages apart from English.
Xu et al. [2] introduced a post-training method for BERT to enhance reading
comprehension and aspect-based sentiment analysis. Their model significantly
improved the extraction of sentiment for specific aspects in product reviews.
However, the study noted challenges in handling implicit aspects and the high
computational cost of fine-tuning BERT for specialized tasks.
Pak et al. [3] in their studies carried out an analysis of Twitter sentiment using
device getting to know techniques. Their paintings laid a basis for information senti-
ment on social media structures, in particular Twitter, through numerous system
mastering strategies.
Zhang et al. [4] in their research supplied a comprehensive survey of diverse deep
gaining knowledge of architectures and their packages in sentiment evaluation. Their
work underscored the capacity of deep gaining knowledge in enhancing sentiment
analysis but also highlighted the challenges of locating and tracking opinion websites
at the net, as well as distilling beneficial records from them.
Liu et al. [5] in their studies introduced BERT-BiGRU-Softmax for review senti-
ment analysis by taking contextual and sequential information into account. It gener-
ally performed very well but was ineffective at detecting fake reviews. Also, it
struggled with high computational costs.
An Analytical Review of Social Media-Based Sentiment Analysis … 233

Haque et al. [6] in their studies proposed a supervised learning model using an
aggregate of function extraction procedures for reading Amazon product critiques.
Their model performed over 90% accuracy and however changed into much less
effective in sensible situations because of bias and the effect on fake critiques, which
skewed the effects.
Pan et al. [7] in their research confirmed powerful outcomes in sentiment cate-
gory obligations through distinctive domain names and the use of spectral feature
alignment (SFA). Their method relies closely on the excellence of domain-specific
vocabulary and labeled records, which are crucial for accurate alignment.
Gaikar et al. [8] in their work investigated the relationship between Twitter senti-
ment and product income, the usage of trend evaluation techniques for predicting
sales tendencies. Although the version was able to locate sentiment traits applicable
to income, challenges remained, which includes increases or decreases in area of
interest products, and handling noisy facts from social media posts.
Poria et al. [9] in their studies explored multimodal sentiment analysis techniques,
addressing the challenges of reading noisy language, sarcasm, and ambiguous content
on social media. They highlighted the issue in decoding sentiment from images reused
in numerous contexts and the paradox of a few hashtags.
Lim et al. [10] evolved a model to extract product evaluations from tweets in the
use of hashtags and sentiment lexicons. Their technique progressed the accuracy of
sentiment class by means of leveraging hashtags and contextual sentiment records,
although it confronted demanding situations with the range and ambiguity of hashtag
usage.
Wang et al. [11] in their studies used Internet site-particular lexicon to enhance the
sentiment analysis for the phone reviews. They explicitly realized that their examine
subdiscipline focuses on improvement to be able to extract centered sentiments,
whereas admitting difficult conditions in dealing with sarcasm and irony.
Rajan et al. [12] conducted real-time customer experience feedback of growth in
Indian telecom companies was done through conducting Twitter Sentiment Anal-
ysis. It analyzed 153,651 tweets from five major brands: Aircel, Bharti Airtel, Idea
Cellular, Reliance Jio, and Vodafone India and presents the connection that positive
customer sentiment is associated with subscriber growth. This study hereby high-
lights the use of social media for immediate feedback-sentiment analysis that informs
management decisions toward enhancement in customer experience and reduction
in churn. This, however, points out that there is a role for collective public opinion
in corporate strategy.
Alsayat [13] better sentiment analysis on social media via integrating a deep
gaining knowledge of models with advanced word embeddings and LSTM networks.
Their ensemble approach, tested on Twitter and other datasets, carried out advanced
type accuracy as compared to standard strategies, especially in the context of COVID-
19-related posts.
Shayaa et al. [14] in their research reviewed diverse strategies and programs for
sentiment analysis in huge records contexts. Their look highlighted the effective-
ness of different techniques and diagnosed key demanding situations, along with
234 P. Jain et al.

scalability and managing various records resources, which impact the accuracy and
efficiency of sentiment evaluation fashions.
Suhartono et al. [15] suggested an approach for the sentiment analysis of drug
product reviews using CNNs along with GloVe word embeddings. Here, they
compared the performance of Word2Vec, GloVe, and RoBERTa architectures. For
this test case, they found that though RoBERTa performed better during training and
validation, CNN models with GloVe embeddings worked well during the testing.
Despite BERT’s strong results, the research highlights the practical effectiveness of
CNN-GloVe models for drug review sentiment analysis.
Onan [16] in their research proposed a sentiment evaluation method that leverages
weighted word embeddings combined with deep neural networks. The approach
improves category accuracy by way of emphasizing important terms and shooting
complex semantic relationships in product reviews.
Xu et al. [17] furnished a comprehensive survey of bilingual sentiment anal-
ysis, reviewing various methodologies, models, and assessment metrics. Their work
highlighted advancements in coping with sentiment analysis throughout a couple
of languages and identified ongoing challenges, such as version edition and overall
performance consistency across diverse linguistic contexts.
Kit et al. [18] in their research showed that the viability of the usage of BERT,
changed into investigation such that it requires minimum update at the models.
Xie et al. [19] introduced MuSES, a multilingual sentiment identification system
tailored for social media data. The system integrates three algorithms: one that
expands compositional semantic rules for social media, a scoring function for
measuring sentiment intensity, and a third algorithm that accounts for emoti-
cons, negation, and domain-specific terms. Experiments on Facebook, Twitter, and
Amazon reviews showed the system’s ability to handle multilingual sentiment
without labeled data, though challenges remain in processing highly informal or
noisy social media content.
Alghalibi et al. [20] in their research the authors followed multitask mastering
with interest in the sentiment models to increase its performance. They overview
their findings on why there is the upward push in complexity and the problems of
training as well as tuning of these trends.
Thomas et al. [21] in their research confirmed that Recurrent Neural Networks
(RNNs) correctly seize sequential dependencies in textual content, mainly to improve
sentiment detection. However, their approach accomplished an accuracy of much less
than 85%, indicating room for development.
Akhtar et al. [22] in their studies demonstrated that analyzing resort reviews
can display insights beyond what ratings alone provide. By classifying reviews and
metadata into predefined aspects and making use of topic modeling (LDA), their
approach uncovers hidden data, providing a more specific information of purchaser
comments.
Liu et al. [23] in their survey explored sentiment evaluation techniques. They
highlighted sentiment evaluation by using a way of leveraging pre-trained models
to improve ordinary performance for the duration of numerous datasets and domain
names.
An Analytical Review of Social Media-Based Sentiment Analysis … 235

Zhang et al. [24] in their research delivered a multi-venture studying technique


that significantly stepped forward sentiment category accuracy. However, their use of
pre-educated BERT fashions led to decreased performance and longer computation
instances.
Vu et al. [25] in their work aimed at assessing total techniques based on the
lexicon for sentiment assessment, indicating that the authors achieved better results
than using other approaches that were based on the similar techniques. They pointed
out drawbacks in comprehending the textual content structure because the technique
acted on all the words in a similar manner regardless of their different roles and
meanings.
Jaswal et al. [26] in their work explored hybrid strategies combining clustering
and category for sentiment evaluation, which gave better results as compared to stan-
dard classifiers. The complexity of hybrid strategies and implementation demanding
situations have been additionally mentioned.
Shah [27] in their research confirmed that supervised studying methods provide
excessive accuracy in sensitivity type with adequate categorized information. It was
emphasized that overall performance is highly dependent on the excellence and
quantity of classified data.
Silva et al. [28] presented a comprehensive survey on sentiment analysis of
tweets, focusing on semi-supervised learning techniques. The study compares various
methodologies and evaluates their performance in classifying sentiments expressed
in tweets.
Gupta [29], in their paper, employed generative hostile networks (GANs) for
perceptual analysis, expanding the existing literature in areas such as text generation.
But their examination proved that the fake data created through the GANs do not ‘fill
in’ the feature space explored by means of the real information, which was decided
using the t-SNE analysis.
Liu et al. [30] A new direction in opinion mining was introduced by a novel
algorithm called ACAEC, which will dynamically and temporally cluster product
reviews. Their WSC method, as well as SWC, focused on chronological and temporal
patterns of reviews; in this regard, precision values obtained were 87.54% and
83.87%, respectively. The study highlighted challenges in handling imbalanced
windows and further improvements could be explored in highly diverse datasets
(Table 1).

3 Tools and Techniques

3.1 Natural Language Processing

NLP is that part of artificial intelligence which uses machine learning for enabling
computers in understanding human language and therefore permitting its interaction
with the human language. It applies a combination of both rule-based modeling with
236 P. Jain et al.

Table 1 Comparison of existing literature


Paper Technology used Achievements Limitation and
references drawbacks
[1] Emoticons, Twitter Used emoticons as noisy labels Accuracy drops
data
[2] Aspect-based Understanding user opinions on Summarizing quality
sentiment analysis hotel aspects varies with aspect
extraction accuracy
[3] Machine learning Developed foundational Limited to specific
techniques sentiment analysis methods for domains; lacks
Twitter scalability
[4] Deep learning Surveyed various models of deep Found challenges in
architectures learning for the sentiment opinion site analysis and
analysis high computational
demands
[5] BERT, textual data Sentiment analysis for Difficulty detecting fake
agricultural product reviews reviews
[6] Supervised Achieved over 90% accuracy Biased and f ake reviews
learning, Amazon affect results
reviews
[7] Spectral feature Found effective for cross-domain Relies on
alignment (SFA) sentiment classification domain-specific
vocabulary
[8] Trend evaluation Relationship between Twitter Did not handle noisy data
techniques sentiment and product income well
[9] Multimodal Advanced sentiment analysis Challenges with sarcasm
analysis (textual with multimodal data and ambiguous hashtags
and visual)
[10] Hashtag-based Improved accuracy using Variable meanings across
sentiment analysis hashtags languages and cultures
[11] Domain-specific Enhanced sentiment analysis for Challenges with sarcasm
lexicon smartphone reviews and irony
[12] Real-time Balanced speed and accuracy in Trade-offs between
sentiment analysis dynamic environments speed and accuracy
[13] Ensemble methods Combined multiple classifiers for Increased complexity and
improved accuracy computational demands
[14] Sentiment analysis Comprehensive comparison Scaling challenges with
techniques across large datasets big data
comparison
[15] Word embeddings The model has been trained on The method did well
an extensive amount of data, during training, but the
making it far more robust and performance dropped
significantly outperforming its during validation
predecessors
(continued)
An Analytical Review of Social Media-Based Sentiment Analysis … 237

Table 1 (continued)
Paper Technology used Achievements Limitation and
references drawbacks
[16] Deep neural Improves category accuracy Significant
networks computational
requirements
[17] Cross-lingual Addressed sentiment Impact of cultural
sentiment analysis classification across languages context on accuracy
[18] Pre-trained More efficient sentiment analysis Lack of domain-specific
language models with pre-trained models knowledge
(BERT)
[19] Multilingual Techniques for handling various Need for further
sentiment analysis languages advancements in
multilingual analysis
[20] Attention Enhanced model interpretability Increased complexity
mechanisms and performance and tuning challenges
[21] Recurrent neural Captured sequential text Accuracy below 85
networks (RNNs) dependencies
[22] LDA Improved understanding of hotel Variable summarization
aspects quality based on aspect
extraction
[23] Transfer learning Applied knowledge from one Poor results with large
domain to another domain differences
[24] Multi-task Significant improvement in Reduced efficiency,
learning, BERT sentiment classification accuracy longer computation times
[25] Lexicon-based Outperformed other Lack of text structure
methods lexicon-based techniques understanding; treating
all words equally
[26] Hybrid methods Achieved better results than Increased complexity
(clusteringand conventional classifiers and implementation
classification) challenges
[27] Supervised High accuracy with sufficient Performance dependent
learning labeled data on quality and quantity
of labeled data
[28] Semi-supervised Effective with varying initial The process has a high
learning labeled data sizes computational cost
[29] Generative Advanced state of the art in text Fake data is unable to
adversarial generation fully cover the feature
networks (GANs) space
[30] Temporal models Efficient processing of large data Challenges with
volumes real-world data accuracy
and algorithm stability
238 P. Jain et al.

statistical modeling [1]. It has two subfields, NLU—natural language understanding


and NLG. It is used in most of the daily life technology such as speech recognition,
text to summary, Google Lens, and Siri.

3.2 Deep Learning/Machine Learning Models

Here are some of the most commonly used deep learning and machine learning
algorithms and models by researchers:
1. Naive Bayes—It is not a single algorithm but rather is a family of algorithms
where all these algorithms share the same principle that every feature, which
is classified, is independent of each other; it makes use of probabilistic classi-
fiers that base predictions on Bayes’ theorem. It is fundamentally used for text
classification, whereby it involves a high-dimensional training dataset.
2. Support Vector Machine—It is used for regression classification, and it is a
supervised machine learning algorithm. For different classes, it produces a
maximal hyperplane that is largest in dimension and space separated from the
correct classifiable points. Support vectors are known to be the closest distances
between hyperplanes and data points, since increment in margin increases
the efficiency of classification. SVM can classify problems both linear and
nonlinear, by increasing the dimensions using functions of kernel (Fig. 1).
3. Random Forest—It is quite unique among ensemble learning methods since it
can be used for classifications and regressions. In general, multiple decision
trees are constructed during training to improve accuracy and avoid overfitting.

Fig. 1 SVM classification


[31]
An Analytical Review of Social Media-Based Sentiment Analysis … 239

Diversity is maximized for each tree constructed as a function of some random


subset of observations and features. The prediction can then be obtained by
averaging outputs in the case of regression or a vote—majority rule in classi-
fication. Random forest has been known to be robust, scalable and can work
effectively on both categorical and numerical data.
4. Decision Tree—It is a supervised learning algorithm widely used for classi-
fication and regression tasks. In simple words, it splits the dataset repeatedly
based on feature values into subsets and creates branches leading to the decision
nodes and leaf nodes. Internal nodes are referring to a feature, the branches are
referring to a decision rule, and the leaf nodes are referring to an outcome or
class label. Decision trees are intuitive and easy to interpret but it tends to cause
overfitting, more so in more complex data. Pruning techniques can relieve the
latter problem.
5. K-Means Clustering—It is an unsupervised learning algorithm in machine
learning used in cluster division of a dataset into K distinct clusters. It starts with
some arbitrary initial K centroids at random locations and then keeps iterating
based on the assignment of each data point to the nearest centroid on the basis of
distance. Then, the centroids are recalculated as the mean of all points assigned
to their group. It repeats this process until the centroids stabilize into clusters,
with variance between instances within clusters being minimal. K-Means is
very adept but it cannot handle non-globular clusters.
6. CNN—Convolutional neural network is one of the most common deep learning
models used for image processing operations that involve a convolutional
layer that automatically determines edges, textures, objects, and other features
by filtering through the input data. CNNs are particularly adept at image
classification, object detection, and facial recognition tasks (Fig. 2).
7. RNN—Recurrent neural network is one neural network architecture designed
primarily for sequential data like text or time series. Its feedback loops make the
information persist over time, and hence this architecture is very well-suited for
speech recognition and language modeling but fails for longer dependencies.
8. LSTM—A long short-term memory network is a particular type of neural
network that is supposed to mimic the ability to remember things for long

Fig. 2 Simple CNN architecture [32]


240 P. Jain et al.

Fig. 3 RNN architecture


[33]

periods like humans do. It is built using special structures called gates which
are the forget gate, the input gate, and the output gate.
9. GRU—GRU is known as Gated Recurrent Unit which is a variant of the RNN
that addresses the shortcomings of traditional RNNs, especially the vanishing
gradient problem. GRUs use gated mechanisms—an update and a reset gate—
to control the flow of information such that the network can capture long-term
dependencies in sequences significantly better (Fig. 3).
10. BERT—It is a transformer-based model pre-trained on various tasks focused on
understanding natural language. BERT captures context because it is inherently
bidirectional, taking into account both sides of the text-created context; hence, it
is well-suited for question-answering applications, text classification, and even
sentiment analysis.
11. ANN—Artificial neural network is an elementary deep model whose structure
is taken from the human brain. It occurs in layers of connected nodes called
neurons which receive the inputs, process them, and learn the patterns within
the weights and biases. ANNs can be found within many applications—from
simple classification tasks to complicated machine learning problems.

3.3 Concepts Used

1. Aspect-Based Sentiment Analysis—Fine-grained opinion mining has a different


name: feature-based sentiment analysis with the target aim of orientation, or
direction, of an opinion within a piece of text, ABSA. This is lately relatively
new area of research due to the vulnerability of standard techniques of sentiment
analysis. Conventional techniques for sentiment analysis analyze a text as a whole
and give it a single sentiment label (positive, negative, or neutral, for example).
2. POS-tagging—POS-tagging or part of speech tagging is one kind of catego-
rization. It is automatically ascribing the description of tokens. We can call the
descriptors ‘tag’ which represents one of the parts of speech (nouns, verbs, etc.),
semantic information, and so on. It can also be used to identify the grammatical
structure of sentences and to disambiguate words that have multiple meanings.
3. SFA—Spectral feature alignment is a domain adaptation technique primarily
for text classification. The latter intends to align the features between different
domains by bridging the gap between them. SFA maps features from both source
An Analytical Review of Social Media-Based Sentiment Analysis … 241

Fig. 4 Neural network [34]

and target domains onto a common space. Spectral clustering will then help group
like features. This helps in improving cross-domain classification by capturing
common latent structures—this makes it somehow efficient for the tasks where
the labeled data is less in target but plenty in source domain (Fig. 4).
4. LDA—This is a generative probabilistic model called Latent Dirichlet Alloca-
tion, which is largely used in topic modeling for natural language processing.
It assumes that documents are mixtures of topics and that each topic should be
some sort of distribution over words. LDA identifies the hidden topics in a dataset
of documents by analyzing word patterns. The model assigns every word in a
document to a topic, thus making extraction of topics across the corpus possible.
It is widely used in understanding thematic structures in large text datasets.
5. Domain-Specific Lexicon—A domain-specific lexicon in NLP refers to the
compilation of specialized words or phrases that are relevant to a certain domain,
such as medicine, law, finance. This way, by focusing on unique terms to that
domain, it might improve how NLP models understand texts with accuracy for
various tasks such as classification, sentiment analysis, and text comprehension.
6. Word2Vec—Word2Vec is a simple neural network model designed to learn vector
representations of words, known as word embeddings. It transforms words into
continuous dense vectors, ensuring that words with related meaning are posi-
tioned close to one another in the vector space. There are two main approaches:
First Continuous Bag of Words (CBOW)—where a word is predicted from its
surrounding context and second Skip-gram—the opposite, where the context
around a word is predicted. Word2Vec is widely used in natural language
processing tasks like text classification and similarity measurement (Fig. 5).

4 Results

The findings of comparison show the following important trends: Deep learning
techniques are mainly characterized by the use of embedding pre-trained language
models such as BERT or attention mechanisms CNN and RNN models that signif-
icantly have improved accuracy as well as semantic understanding in the context
242 P. Jain et al.

Fig. 5 Graphical representation of LDA. The boxes also serve for the representation of the replicas.
Since there are two plates in the first step, it could be helpful to keep them separate. The outer plate
then should represent documents, and the inner plate can be used representing recurring themes and
vocabularies within a particular document [35]

of sentiment analysis. This is still computationally expensive and may get quite
time-consuming.
Traditional machine learning techniques have survived the test of time, including
supervised learning and ensemble methods, because of their balance of simplicity
with performance but are limited to scalability and quality labeled data.
Lexicon-based Methods: These work particularly well in specific domains, do
poorly in structural and contextual understanding of text, especially with sarcasm
and domain-specific jargon.
Multimodal and Multilingual Analysis: The following approaches are focused on
the urgent need for sentiment analysis not only beyond plain textual data, but also
in a wide range of languages. However, the task is even more complex because of
the complex issues present in dealing with multimodal data both in terms of textual
and visual aspects and also interpreting cultural contexts in multilingual sentiment
analyses.

5 Conclusion

The comparative analysis of 30 research papers on Sentiment Analysis shows remark-


able development concerning the following aspects: technological diversity, improve-
ment in the accuracy, and solving certain challenges. Each technology differs in
strengths and weaknesses that determine whether this technology would be effective
or not depending on the context and application. By observing the comparison, it is
indicated that future research emphasis shall be directed to:
Scalability Improvement This shall involve minimizing computational costs while
ensuring the high accuracy of outputs relevant to real-time analysis of sentiment.
Advances in more powerful models The ability to analyze textual, visual, and audio
data seamlessly in more languages and cultural contexts.
An Analytical Review of Social Media-Based Sentiment Analysis … 243

Handling Ambiguity and Sarcasm Better methods of detection of sarcasm and ambi-
guities can enhance the accuracy of the sentiment analysis system, especially in social
media contexts.
Enhanced Domain Transferability Techniques of transfer learning, along with
methods across domains, need to be further enhanced to cope with large differences
between domains.
While much has been accomplished concerning the analysis of sentiment, espe-
cially with deep learning and machine learning techniques, success concerning these
aspects and the extension of applicability to a broad spectrum of data types and
languages could likely depend upon future improvements.

References

1. Go A, Bhayani R, Huang L (2009) Twitter sentiment classification using distant supervision.


In: CS224N project report, Stanford, vol 1, No 12, p 2009
2. Xu H, Liu B, Shu L, Yu PS (2019) BERT post-training for review reading comprehension and
aspect-based sentiment analysis. arXiv preprint arXiv:1904.02232
3. Pak A, Paroubek P (2010, May) Twitter as a corpus for sentiment analysis and opinion mining.
In: LREc, vol 10, No 2010, pp 1320–1326
4. Zhang L, Wang S, Liu B (2018) Deep learning for sentiment analysis: A survey. Wiley
Interdiscip Rev: Data Min Knowl Discov 8(4):e1253
5. Liu Y, Lu J, Yang J, Mao F (2020) Sentiment analysis for e-commerce product reviews by deep
learning model of Bert-BiGRU-Softmax. Math Biosci Eng 17(6):7819–7837
6. Haque TU, Saber NN, Shah FM (2018, May) Sentiment analysis on large scale Amazon product
reviews. In: 2018 IEEE international conference on innovative research and development
(ICIRD). IEEE, pp 1–6
7. Pan SJ, Ni X, Sun JT, Yang Q, Chen Z (2010, April) Cross-domain sentiment classification via
spectral feature alignment. In: Proceedings of the 19th international conference on world wide
web, pp 751–760
8. Gaikar D, Marakarkandy B (2015) Product sales prediction based on sentiment analysis using
twitter data. Int J Comput Sci Inf Technol 6(3):2303–2313
9. Poria S, Cambria E, Hazarika D, Vij P (2016) A deeper look into sarcastic tweets using deep
convolutional neural networks. arXiv preprint arXiv:1610.08815
10. Lim KW, Buntine W (2014, November) Twitter opinion topic model: extracting product opin-
ions from tweets by leveraging hashtags and sentiment lexicon. In: Proceedings of the 23rd
ACM international conference on conference on information and knowledge management, pp
1319–1328
11. Wang K, Shen W, Yang Y, Quan X, Wang R (2020) Relational graph attention network for
aspect-based sentiment analysis. arXiv preprint arXiv:2004.12362
12. Ranjan S, Sood S, Verma V (2018, August) Twitter sentiment analysis of real-time customer
experience feedback for predicting growth of Indian telecom companies. In: 2018 4th
international conference on computing sciences (ICCS). IEEE, pp 166–174
13. Alsayat A (2022) Improving sentiment analysis for social media applications using an ensemble
deep learning language model. Arab J Sci Eng 47(2):2499–2511
14. Shayaa S, Jaafar NI, Bahri S, Sulaiman A, Wai PS, Chung YW, Piprani AZ, Al-Garadi MA
(2018) Sentiment analysis of big data: methods, applications, and open challenges. IEEE Access
6:37807–37827
244 P. Jain et al.

15. Suhartono D, Purwandari K, Jeremy NH, Philip S, Arisaputra P, Parmonangan IH (2023) Deep
neural networks and weighted word embeddings for sentiment analysis of drug product reviews.
Procedia Comput Sci 216:664–671
16. Onan A (2021) Sentiment analysis on product reviews based on weighted word embeddings
and deep neural networks. Concurr Comput: Pract Exp 33(23):e5909
17. Xu Y, Cao H, Du W, Wang W (2022) A survey of cross-lingual sentiment analysis:
methodologies, models and evaluations. Data Sci Eng 7(3):279–299
18. Kit Y, Mokji MM (2022) Sentiment analysis using pre-trained language model with no fine-
tuning and less resource. IEEE Access 10:107056–107065
19. Xie Y, Chen Z, Zhang K, Cheng Y, Honbo DK, Agrawal A, Choudhary AN (2013) MuSES:
multilingual sentiment elicitation system for social media data. IEEE Intell Syst 29(4):34–42
20. Alghalibi M, Al-Azzawi A, Lawonn K (2020) Deep attention learning mechanisms for social
media sentiment image revelation. Int J Comput Commun Eng 9(1):1–17
21. Thomas M, Latha CA (2018) Sentimental analysis using recurrent neural network. Int J Eng
Technol (UAE) 7(2.27):88–92
22. Akhtar N, Zubair N, Kumar A, Ahmad T (2017) Aspect based sentiment oriented summarization
of hotel reviews. Procedia Comput Sci 115:563–571
23. Liu R, Shi Y, Ji C, Jia M (2019) A survey of sentiment analysis based on transfer learning.
IEEE Access 7:85401–85412
24. Zhang J, Yan K, Mo Y (2021) Multi-task learning for sentiment analysis with hard-sharing and
task recognition mechanisms. Information 12(5):207
25. Vu L, Le T (2017) A lexicon-based method for sentiment analysis using social network data. In:
Proceedings of the international conference on information and knowledge engineering (IKE),
pp 10–16. The Steering Committee of the World Congress in Computer Science, Computer
Engineering and Applied Computing (WorldComp)
26. Jaswal PK, Tathgir JS (2018) Sentiment analysis of social media data using hybrid approach.
J Econ Dev, Manag, IT, Financ, Market 10(2):1–6
27. Shah A (2021) Sentiment analysis of product reviews using supervised learning. Reliab: Theory
Appl 16(SI 1 (60)):243–253
28. Silva NFFD, Coletta LF, Hruschka ER (2016) A survey and comparative study of tweet
sentiment analysis via semi-supervised learning. ACM Comput Surv (CSUR) 49(1):1–26
29. Gupta R (2019, May) Data augmentation for low resource sentiment analysis using generative
adversarial networks. In: ICASSP 2019—2019 IEEE international conference on acoustics,
speech and signal processing (ICASSP). IEEE, pp 7380–7384
30. AL-Sharuee MT, Liu F, Pratama M (2021) Sentiment analysis: dynamic and temporal clustering
of product reviews. Appl Intell 51:51–70
31. Meyer D, Wien FT (2001) Support vector machines. R News 1(3):23–26
32. CNN. [Link]
33. Zhu J, Yang Z, Mourshed M, Guo Y, Zhou Y, Chang Y, Wei Y, Feng S (2019) Electric
vehicle charging load forecasting: a comparative study of deep learning approaches. Energies
12(14):2692
34. Neural Network. [Link]
deep-learning-neural-networks/
35. Blei DM, Ng AY, Jordan MI (2003) Latent Dirichlet allocation. J Mach Learn Res 3(Jan):993–
1022
The Future of Lung Disease Diagnosis:
A Review of Emerging Trends
in Data-Driven Classification

Jigisha Mehta and Sheshang Degadwala

Abstract A major obstacle in medical diagnostics is the categorization of lung disor-


ders, which have hitherto depended on image-based techniques like CT scans and X-
rays. The use of medical data-driven methods (such as biomarkers and patient histo-
ries) and current developments in sound analysis (such as auscultation) have opened
new avenues for more precise and multimodal approaches. One of the current issues
is that single-modal approaches have their limitations. For example, image analysis
has limited specificity, and non-imaging data is underutilized. Novel approaches that
integrate multiple data types—sound, images, and medical records—are the subject
of this review study, which seeks to offer a thorough assessment of current trends in
lung disease classification. The goals are to recognize important developments, point
out the advantages and disadvantages of existing methods, and suggest avenues for
further study that can improve diagnostic precision and patient outcomes.

Keywords Lung disease classification · Multimodal data · Sound analysis ·


Medical imaging · Diagnostic trends · Machine learning

1 Introduction

Because of their high incidence and mortality rates, lung diseases rank among the
world’s most pressing public health challenges. As a result of a confluence of vari-
ables, including pollution, smoking, occupational risks, and rapid urbanization, lung
ailments are particularly common in India. Over 1 million people die each year in
India from chronic obstructive pulmonary disease (COPD), and asthma is another
major cause of death in the country (Global Burden of Disease Study 2019). More
than eighty percent of India’s cities have air pollution levels that are higher than

J. Mehta (B) · S. Degadwala


Department of Computer Engineering, Sigma University, Vadodara, India
e-mail: jigimehta08@[Link]
S. Degadwala
e-mail: sheshang13@[Link]

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 245
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
246 J. Mehta and S. Degadwala

what is considered safe by the World Health Organization (WHO). This contributes
to the already alarmingly high incidence of lung ailments in urban areas. The growing
prevalence of smoking and other environmental contaminants is also a factor in the
alarming increase in reported cases.
Diseases affecting the lungs are known as lung diseases. The lungs are essential
for breathing and exchanging oxygen throughout the body. Disorders of the airways,
the lung tissue, and the lung circulation make up the three main groups of lung
disorders.
1. Airway Diseases: These things have an effect on the airways, which are the
pathways that gases like oxygen use to enter and exit the lungs. Among airway
illnesses, asthma and chronic obstructive pulmonary disease (COPD) arise most
frequently.
2. Lung Tissue Diseases: These things have an effect on the structure of the lungs,
which makes full lung expansion difficult.
3. Lung Circulation Diseases: These have an effect on the pulmonary arteries,
which in turn alters the blood flow from the heart to the lungs.
Early diagnosis is crucial for managing lung diseases, but traditional diagnostic
methods face several limitations. Conventional methods include imaging scans of the
chest and lungs, as well as computed tomography (CT) scans, although these have
limitations, such as a lack of sensitivity, identification at late stages, and expensive
expenses. Additionally, these imaging-based methods do not fully utilize other avail-
able data sources, such as sound recordings from lung auscultation or patient-specific
medical records, which can potentially provide valuable insights for more accurate
diagnosis.
Figure 1 illustrates various risk factors contributing to lung diseases. These include
environmental and lifestyle factors such as radiation, pollution, and tobacco smoking,
as well as biological and occupational factors like genetics, asbestos exposure, and
aging. Each factor presents potential threats that can impair lung function and lead
to diseases such as COPD.
The major objective of this review article is to investigate new developments in
the categorization of lung diseases by looking at creative approaches and making
use of many kinds of data, including audio analysis, medical images, and patient
health records. By reviewing these methodologies, the paper seeks to identify the
strengths and limitations of existing diagnostic techniques and how new approaches
can improve diagnostic accuracy and early detection. This paper aims to give a thor-
ough overview of how multimodal data might be used to improve lung disease diag-
nosis and improve patient outcomes in India and throughout the world by evaluating
and synthesizing these new trends.
The Future of Lung Disease Diagnosis: A Review of Emerging Trends … 247

Fig. 1 Disorders of lung disease. Source [Link]


[Link]

2 Literature Works

Lung illness categorization literature shows machine learning and deep learning
advances in sound, image, and medical data utilization. Researchers have used lung
sound analysis and sophisticated imaging to demonstrate AI’s expanding importance
in respiratory disease diagnosis and early detection.
Mochizuki et al. [1] used an automated method to study lung sounds in asthmatic
newborns and children, identifying patterns that predict wheezing early on. Their
study emphasizes sound analysis in pediatric asthma treatment. Ashwini et al. [2]
used optimized deep convolutional neural networks (CNNs) to accurately diagnose
and classify lung diseases using chest X-ray (CXR). Pessoa et al. [3] created an
ensemble deep learning model to estimate respiratory airflow using sound data. Patel
et al. [4] proposed an explainable transfer learning paradigm for lung ailment multi-
classification using chest X-rays, emphasizing the need of interpretability in AI-
based healthcare solutions. Shehab et al. [5] developed a deep learning and feature
fusion model for lung sound identification to improve respiratory ailment diagnosis
by merging sound patterns with other clinical data. These studies demonstrate the
growing interest in using multimodal data to improve lung disease classification
accuracy and early diagnosis.
Vinta et al. [6] introduced a hybrid deep learning network to segment and
classify interstitial lung disorders, which outperformed standard models in iden-
tifying complicated lung ailments. Kumar et al. [7] reviewed imaging modalities and
machine learning models for lung disease identification and found that AI improves
diagnostic efficiency, especially in resource-limited situations. A machine learning-
based classification approach for real-time lung sound analysis by Balasubramanian
and Rajadurai [8] could automate disease categorization from lung sounds without
expensive imaging. Wang et al. [9] proposed LungNeXt, a lightweight deep learning
network that classifies lung sounds using enhanced mel-spectrograms. This unique
248 J. Mehta and S. Degadwala

approach consumes little resources. Sang et al. [10] developed a wearable patch
employing accelerometer technology to detect respiratory rates and wheezing in real
time using deep learning algorithms. Fava et al. [11] found that increasing sound
data quality enhances deep learning models’ respiratory disease diagnosis accuracy.
Saeed et al. [12] created an AI-enabled bias-free model for diagnosing respiratory
disorders like bronchitis and pneumonia using cough sounds, with promising results
independent of gender and age. Xu and Sankar [13] examined AI and machine
learning lung sound classification algorithms and found that deep learning models
are best for accurate and scalable respiratory disease detection. Sabry et al. [14] found
that sound analysis and machine learning improve COPD and asthma classification.
Lung sound analysis in preschoolers can predict recurrent wheeze, according to
Miyamoto et al. [15]. Sound-based predictive models can help early intervention in
pediatric asthma cases. Patel et al. [16] employed deep learning and multi-feature
fusion to improve COPD classification using clinical and imaging data.
Liu et al. [17] trained a CNN to classify lung diseases using attention processes and
found that it could accurately discriminate lung cancer, TB, and pneumonia. Koppad
et al. [18] improved lung sound and lung sickness categorization generically and
accuracy by merging multiple objectives into one framework. Lauwers et al. [19]
found in a cross-sectional investigation that imaging biomarkers and lung sound
analysis can be used together to detect lung diseases. A longitudinal study by Sgalla
et al. [20] showed that crackles and other sound patterns can help diagnose fibrotic
interstitial lung disease early.
Sfayyih et al. [21] discovered that sound analysis is an accurate and cost-effective
lung ailment diagnosis method for low-resource situations. Bhattacharya et al. [22]
introduced the Coswara dataset, which collects respiratory sounds and symptoms
for remote SARS-CoV-2 screening, to highlight the utility of sound-based pandemic
responses. Choi and Lee [23] recommended using light attention modules to improve
deep learning model interpretability without losing accuracy for lung ailment catego-
rization. A blocking variable in a deep neural network helped Yang et al. [24] classify
respiratory sounds as normal or abnormal. Lal’s [25] transfer learning model for lung
sound detection identified bronchitis and asthma well.
Priyadarsini et al. [26] found that CNN-based architectures outperform traditional
machine learning models in accuracy and processing speed for lung disease identifi-
cation. Petmezas et al. [27] created a hybrid CNN-LSTM network for automatic lung
sound classification that uses temporal and spatial characteristics to detect pneumonia
and COPD. Fraiwan et al. [28] employed CNNs and LSTMs to classify pulmonary
diseases, showing that incorporating lung sound time-series and spatial information
enhances accuracy. Chen et al. [29] classified respiratory illnesses better by recog-
nizing wheezes and crackles from lung sounds using a fine-tuned ResNet18 network
and STFT. Lastly, Nguyen and Pernkopf [30] employed co-tuning and stochastic
normalization to classify lung sounds, demonstrating that these methods can improve
the resilience and generalization of deep learning models for respiratory disease
detection.
The examined literature shows significant progress in lung disease classification
utilizing sound analysis, chest X-rays, and deep learning algorithms. Studies show
The Future of Lung Disease Diagnosis: A Review of Emerging Trends … 249

that convolutional neural networks (CNNs), long short-term memory (LSTMs), and
transfer learning improve diagnostic accuracy and early detection. Combining multi-
modal data like lung sounds and images can improve automated diagnostic tools. AI
could improve pulmonary healthcare by providing more accurate, cost-effective, and
scalable solutions, according to this research.

3 Comparative Study

Table 1 provides various studies from 2022 to 2024 that focus on different data types
and methodologies for lung disease classification.
The studies presented demonstrate significant advancements in lung disease clas-
sification through various data types, including sound, imaging, and multimodal
approaches. Each methodology offers unique advantages, such as improved accuracy
and non-invasive diagnostics, while also facing challenges like dataset limitations
and computational demands. These emerging trends indicate a promising shift toward
more effective and accessible diagnostic tools, highlighting the need for continued
research and innovation in this critical area of healthcare.

4 Challenges and Solutions

Below are some of the key challenges and their corresponding solutions found in the
referenced papers.

4.1 Challenges

1. Limited Labeled Data: Many algorithms, particularly deep learning models,


require extensive labeled datasets for training, which can be difficult to obtain,
especially for specific lung diseases.
2. High Computational Requirements: Advanced models, such as convolutional
neural networks (CNNs) and hybrid architectures, often demand significant
computational power and memory, making them less accessible for widespread
use.
3. Variability in Lung Sounds: Lung sounds can vary significantly across indi-
viduals and conditions, leading to challenges in accurately classifying different
respiratory diseases based on acoustic signals.
4. Overfitting: Machine learning models risk overfitting when trained on limited
datasets, resulting in poor generalization to unseen data.
250 J. Mehta and S. Degadwala

Table 1 Comparative study of lung disease based on datatypes


Reference Year Data type Key features Advantages Limitations
Mochizuki et al. 2024 Sound Automatic analysis Non-invasive, Limited to
[1] of lung sounds in early detection pediatric
infants patients
Ashwini et al. [2] 2024 CXR Optimized deep High accuracy in Requires high
CNN for CXR computational
multi-classification classification resources
Pessoa et al. [3] 2024 Sound Ensemble deep Real-time, Dataset
learning model for dimensionless dependency
airflow estimation airflow estimation
Patel et al. [4] 2024 CXR Explainable transfer Improves Requires large,
learning framework interpretability of labeled datasets
results
Shehab et al. [5] 2024 Sound Feature fusion with Effective fusion Lacks testing in
deep learning model for sound real-time
recognition applications
Vinta et al. [6] 2024 CT Hybrid deep learning Accurate in Limited
model for detecting generalization to
segmentation interstitial lung other diseases
disease
Kumar et al. [7] 2024 Multimodal Review of machine Comprehensive No experimental
learning paradigms coverage of validation
imaging
modalities
Balasubramanian 2024 Sound Real-time sound Cost-effective, Limited to
et al. [8] classification using real-time sound-based
machine learning processing diagnosis
Wang et al. [9] 2024 Sound Lightweight deep Low resource Model
learning model using consumption, real complexity
mel-spectrograms time needs to be
reduced
Sang et al. [10] 2024 Sound Accelerometer-based Continuous Device
wearable patch monitoring of dependency
respiratory health
Fava et al. [11] 2024 Sound Pre-processing Enhances model Pre-processing
techniques for sound accuracy increases
classification complexity
Saeed et al. [12] 2024 Sound AI-enabled bias-free Demographic Limited to
diagnosis using neutrality in cough-based
cough audio diagnosis analysis
Xu and Sankar 2024 Sound Review of AI and Summarizes No novel
[13] ML in lung sound effective experimental
classification algorithms data
Sabry et al. [14] 2024 Sound Audio-based Improved Dataset
analysis with accuracy in constraints
machine learning sound-based
detection
(continued)
The Future of Lung Disease Diagnosis: A Review of Emerging Trends … 251

Table 1 (continued)
Reference Year Data type Key features Advantages Limitations
Miyamoto et al. 2024 Sound Lung sound analysis Predicts recurrent Focused on
[15] for pediatric wheezing in pediatric cases
wheezing children
Patel et al. [16] 2024 CXR Multi-feature fusion Enhances Computational
for COPD classification cost
classification accuracy
Liu et al. [17] 2024 CXR CNN architecture High accuracy in Requires
with attention lung disease large-scale
mechanisms classification datasets
Koppad et al. 2024 Sound Multi-task learning Simultaneous task Complexity in
[18] for sound and handling improves model training
disease classification generalization
Lauwers et al. 2024 Multimodal Integration of sound Cross-validation Requires diverse
[19] analysis and imaging of diagnostic datasets
biomarkers methods
Sgalla et al. [20] 2024 Sound Longitudinal Reliable Limited to
analysis of crackles diagnostic tool for fibrotic diseases
in fibrosis diagnosis fibrosis
Sfayyih et al. 2023 Sound Review on lung Highlights No experimental
[21] disease recognition sound-based validation
by acoustic analysis diagnostic
techniques
Bhattacharya 2023 Sound Coswara dataset for Remote screening Limited to
et al. [22] COVID-19 screening via sound analysis COVID-19 cases
Choi and Lee 2023 Multimodal Light attention Improve model Not tested on
[23] module in lung interpretability large datasets
disease classification
Yang et al. [24] 2023 Sound Deep neural network Enhanced Model
with blocking classification of complexity may
variable abnormal lung affect
sounds deployment
Lal [25] 2023 Sound Transfer learning Effective in Limited
model for sound diagnosing generalizability
recognition asthma, bronchitis
Priyadarsini et al. 2023 Multimodal Comparative study Identifies No experimental
[26] of deep learning best-performing comparisons
algorithms architectures
Petmezas et al. 2022 Sound Hybrid CNN-LSTM Combines Computationally
[27] network for sound temporal and intensive
classification spatial data
Fraiwan et al. 2022 Sound CNN-LSTM Effective in High
[28] combination for multi-dimensional computational
pulmonary diseases sound analysis demand
(continued)
252 J. Mehta and S. Degadwala

Table 1 (continued)
Reference Year Data type Key features Advantages Limitations
Chen et al. [29] 2022 Sound STFT and fine-tuned Superior in Requires
ResNet18 for lung detecting preprocessing of
sound classification respiratory sounds
abnormalities
Nguyen and 2022 Sound Co-tuning and Enhances Needs
Pernkopf [30] stochastic robustness in lung fine-tuning on
normalization sound specific datasets
classification

5. Interpretability: The complex nature of deep learning models makes it chal-


lenging for clinicians to interpret the results, which can hinder trust in automated
diagnosis systems.

4.2 Solutions

1. Data Augmentation: To address the issue of limited labeled data, researchers


have implemented data augmentation techniques, such as synthesizing new
samples from existing ones, which increases the diversity of the training dataset.
2. Model Compression Techniques: Solutions like pruning, quantization, and
knowledge distillation have been proposed to reduce the computational require-
ments of models while maintaining performance, making them more practical
for real-world applications.
3. Robust Feature Extraction: Employing advanced feature extraction methods,
such as mel-spectrograms and wavelet transforms, helps improve the robustness
of lung sound classification by capturing essential characteristics while mitigating
variability.
4. Regularization Techniques: Implementing regularization methods, such as
dropout and early stopping, helps prevent overfitting by adding constraints to
the model training process, promoting better generalization.
5. Explainable AI Approaches: To enhance interpretability, researchers are
exploring explainable AI techniques that provide insights into model decision-
making processes, enabling clinicians to understand and trust the automated
classification results.

By addressing issues like data scarcity, computational demands, and inter-


pretability, researchers are paving the way for more accurate and clinically applicable
diagnostic tools in respiratory medicine.
The Future of Lung Disease Diagnosis: A Review of Emerging Trends … 253

5 Conclusion and Future Work

The landscape of lung disease diagnosis is evolving rapidly, driven by advancements


in data-driven classification techniques. This review has highlighted the transforma-
tive potential of emerging trends such as machine learning, deep learning, and multi-
modal data integration in enhancing diagnostic accuracy and efficiency. The studies
discussed illustrate the promising results achieved through innovative algorithms
and methodologies, which are capable of analyzing diverse data types, including
medical imaging, acoustic signals, and patient demographic information. There are
still obstacles to integrating these technologies into clinical practice, including a lack
of data, difficulties with interpretation, and the requirement for strong validation.
Data-driven classification models should be refined and validated to improve lung
disease diagnosis. Research should focus on developing larger and more diverse
datasets that cover a variety of patient demographics and lung disease presentations.
Model interpretability and seamless integration into existing healthcare frameworks
should also be prioritized.

6 Disclosure of Interests

The authors declare no conflicts of interest.

References

1. Mochizuki H, Hirai K, Furuya H, Niimura F, Suzuki K, Okino T, Ikeda M, Noto H (2024) The
analysis of lung sounds in infants and children with a history of wheezing/asthma using an
automatic procedure. BMC Pulm Med 24. [Link]
2. Ashwini S, Arunkumar JR, Prabu RT, Singh NH, Singh NP (2024) Diagnosis and multi-
classification of lung diseases in CXR images using optimized deep convolutional neural
network. Soft Comput 28:6219–6233. [Link]
3. Pessoa D, Rocha BM, Gomes M, Rodrigues G, Petmezas G, Cheimariotis GA, Maglaveras
N, Marques A, Frerichs I, de Carvalho P, Paiva RP (2024) Ensemble deep learning model for
dimensionless respiratory airflow estimation using respiratory sound. Biomed Signal Process
Control 87:105451. [Link]
4. Patel AN, Murugan R, Srivastava G, Maddikunta PKR, Yenduri G, Gadekallu TR, Chengoden
R (2024) An explainable transfer learning framework for multi-classification of lung diseases
in chest X-rays. Alex Eng J 98:328–343. [Link]
5. Shehab SA, Mohammed KK, Darwish A, Hassanien AE (2024) Deep learning and feature
fusion-based lung sound recognition model to diagnoses the respiratory diseases. Soft Comput
28:11667–11683. [Link]
6. Vinta SR, Lakshmi B, Aruna Safali M, Sai Chaitanya Kumar G (2024) Segmentation and
classification of interstitial lung diseases based on hybrid deep learning network model. IEEE
Access 12:50444–50458. [Link]
7. Kumar S, Kumar H, Kumar G, Singh SP, Bijalwan A, Diwakar M (2024) A methodical explo-
ration of imaging modalities from dataset to detection through machine learning paradigms in
254 J. Mehta and S. Degadwala

prominent lung disease diagnosis: a review. BMC Med Imaging 24:1–42. [Link]
1186/s12880-024-01192-w
8. Balasubramanian S, Rajadurai P (2024) Machine learning-based classification of pulmonary
diseases through real-time lung sounds. Int J Eng Technol Innov 14:85–102. [Link]
10.46604/ijeti.2023.12294
9. Wang F, Yuan X, Liu Y, Lam CT (2024) LungNeXt: a novel lightweight network utilizing
enhanced mel-spectrogram for lung sound classification. J King Saud Univ - Comput Inf Sci
36:102200. [Link]
10. Sang B, Wen H, Junek G, Neveu W, Di Francesco L, Ayazi F (2024) An accelerometer-based
wearable patch for robust respiratory rate and wheeze detection using deep learning. Biosensors
14. [Link]
11. Fava A, Dianat B, Bertacchini A, Manfredi A, Sebastiani M, Modena M, Pancaldi F (2024)
Pre-processing techniques to enhance the classification of lung sounds based on deep learning.
Biomed Signal Process Control 92:106009. [Link]
12. Saeed T, Ijaz A, Sadiq I, Qureshi HN, Rizwan A, Imran A (2024) An AI-enabled bias-free
respiratory disease diagnosis model using cough audio. Bioengineering 11. [Link]
3390/bioengineering11010055
13. Xu X, Sankar R (2024) Classification and recognition of lung sounds using artificial intelligence
and machine learning: a literature review. Big Data Cogn Comput 8:127. [Link]
3390/bdcc8100127
14. Sabry AH, Dallal Bashi OI, Nik Ali NH, Mahmood Al Kubaisi Y (2024) Lung disease recog-
nition methods using audio-based analysis with machine learning. Heliyon 10:e26218. https://
[Link]/10.1016/[Link].2024.e26218
15. Miyamoto M, Yoshihara S, Shioya H, Tadaki H, Imamura T, Enseki M, Furuya H, Kato
M, Mochizuki H (2024) Lung sound analysis for predicting recurrent wheezing in preschool
children. J Allergy Clin Immunol: Glob 3:100199. [Link]
16. Patel PJ, Diwan D, Patel KA, Ranga S, Modi NJ, Dumasia S (2024) Multi feature fusion for
COPD classification using deep learning algorithms. J Integr Sci Technol 12:1–8. [Link]
org/10.62110/[Link].2024.v12.780
17. Liu X, Yu Z, Tan L (2024) Deep learning for lung disease classification using transfer learning
and a customized CNN architecture with attention. arXiv
18. Koppad D, Kumar P, Kantikar NA, Ramesh S (2024) Multi-task learning for lung sound and
lung disease classification. arXiv, pp 1–22
19. Lauwers E, Stas T, McLane I, Snoeckx A, Van Hoorenbeeck K, De Backer W, Ides K, Steckel
J, Verhulst S (2024) Exploring the link between a novel approach for computer aided lung
sound analysis and imaging biomarkers: a cross-sectional study. Respir Res 25:1–9. https://
[Link]/10.1186/s12931-024-02810-5
20. Sgalla G, Simonetti J, Di Bartolomeo A, Magrì T, Iovene B, Pasciuto G, Dell’Ariccia R,
Varone F, Comes A, Leone PM, Piluso V, Perrotta A, Cicchetti G, Verdirosi D, Richeldi L
(2024) Reliability of crackles in fibrotic interstitial lung disease: a prospective, longitudinal
study. Respir Res 25. [Link]
21. Sfayyih AH, Sulaiman N, Sabry AH (2023) A review on lung disease recognition by acoustic
signal analysis with deep learning networks. J Big Data 10. [Link]
023-00762-z
22. Bhattacharya D, Sharma NK, Dutta D, Chetupalli SR, Mote P, Ganapathy S, Chandrakiran
C, Nori S, Suhail KK, Gonuguntla S, Alagesan M (2023) Coswara: a respiratory sounds and
symptoms dataset for remote screening of SARS-CoV-2 infection. Sci Data 10:1–11. https://
[Link]/10.1038/s41597-023-02266-0
23. Choi Y, Lee H (2023) Interpretation of lung disease classification with light attention
connected module. Biomed Signal Process Control 84:104695. [Link]
2023.104695
24. Yang R, Lv K, Huang Y, Sun M, Li J, Yang J (2023) Respiratory sound classification by
applying deep neural network with a blocking variable. Appl Sci (Switzerland) 13. [Link]
org/10.3390/app13126956
The Future of Lung Disease Diagnosis: A Review of Emerging Trends … 255

25. Lal KN (2023) A lung sound recognition model to diagnoses the respiratory diseases by using
transfer learning. Multimed Tools Appl 82:36615–36631. [Link]
14727-0
26. Jasmine Pemeena Priyadarsini M, Kotecha K, Rajini GK, Hariharan K, Utkarsh Raj K, Bhargav
Ram K, Indragandhi V, Subramaniyaswamy V, Pandya S (2023) Lung diseases detection using
various deep learning algorithms. J Healthc Eng 2023. [Link]
27. Petmezas G, Cheimariotis GA, Stefanopoulos L, Rocha B, Paiva RP, Katsaggelos AK, Maglav-
eras N (2022) Automated lung sound classification using a hybrid CNN-LSTM network and
focal loss function. Sensors 22. [Link]
28. Fraiwan M, Fraiwan L, Alkhodari M, Hassanin O (2022) Recognition of pulmonary diseases
from lung sounds using convolutional neural networks and long short-term memory. J Ambient
Intell Humaniz Comput 13:4759–4771. [Link]
29. Chen Z, Wang H, Yeh CH, Liu X (2022) Classify respiratory abnormality in lung sounds
using STFT and a fine-tuned ResNet18 network. In: BioCAS 2022—IEEE biomedical circuits
and systems conference: intelligent biomedical systems for a better future, proceedings, pp
233–237. [Link]
30. Nguyen T, Pernkopf F (2022) Lung sound classification using co-tuning and stochastic
normalization. IEEE Trans Biomed Eng 69:2872–2882. [Link]
3156293
Leveraging Local Interpretable
Model-Agnostic Explanations (LIME)
for Sentiment Analysis of News Articles

Pragya Tewari, Harshit Verma, Rishabh Jaiswal, and Anurag Singh Baghel

Abstract Digital media have grown at a remarkably fast pace, which in turn has
increased the consumption of news through e-papers. In such a scenario, the ability
to quickly estimate the sentiment of a news article becomes highly valuable. This
research report encompasses a broad study on sentiment analysis related to e-paper
news articles, by applying techniques in NLP, sentiment analysis models classify
news articles into undertones that are either positive, negative, or neutral. The project
begins with collecting a long and wide dataset of e-paper news articles, covering a
wide range of categories, sources, and different timelines. Further sentiment analysis
will hence be derived from the dataset. Preprocessing the dataset beforehand through
NLP is done prior to the sentiment analysis. These steps, including text cleaning, stop
word removal, and stemming, are involved in this cleaning process to further improve
the quality of the dataset. We have chosen logistic regression over the other machine
learning models due to its effectiveness in text classification tasks, its ability to model
probabilities, and its strong performance with both simple and complex language
structures. In addition, techniques from the group of Explainable AI such as LIME
were applied to make the processes of sentiment classification more transparent and
interpretable. These methods shall allow the users to comprehend which portion of
any words and phrase in light of overall sentiment is contributing positive, thereby
building trust in automated news classification. Consequently, this study adds to this
emergent domain with a focus on e-paper news articles and proposes showcasing the
part of Explainable AI in promoting transparency, trust, and fairness in automated

P. Tewari · R. Jaiswal
Gautam Buddha University, Greater Noida, India
e-mail: [Link]@[Link]
R. Jaiswal
e-mail: Jaiswalrishabh1230@[Link]
H. Verma (B) · A. S. Baghel
Galgotias University, Greater Noida, India
e-mail: Harshit02425@[Link]
A. S. Baghel
e-mail: asb@[Link]

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 257
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
258 P. Tewari et al.

news sentiment detection. It is highly relevant to understand the sentiment of news


articles on e-paper for both readers and publishers.

Keywords Sentiment analysis · Explainable AI · Machine learning · Natural


language processing · E-paper news articles · Logistic regression

1 Introduction

Sentiment analysis [1] also called opinion mining is a process designating the
emotional tone that seems to appear in the background of a text. News-wise, senti-
ment analysis is valued in the serious evaluation that can be obtained with a great
amount of digital content. News, or electronic newspapers, has become a popular
digital alternative to traditional print media.
Sentiment analysis, with its beginning in the news world, has opened an entirely
new dimension in the realization of the effects of news and articles on readership.
Using natural language [2] processing techniques, these algorithms of sentiment
analysis identify and further classify sentiments being expressed in the text as posi-
tive, negative, or neutral. This makes it possible to actually dig deeper into how
readers view and react to what is fed to them in news. However, most of the senti-
ment analysis models, especially machine learning models, are black-boxes; it is
hard to understand why one particular decision was made.
These focus on bringing explainability into sentiment classification by employing
Explainable AI techniques. This is where XAI gives light with regard to the very
reasoning of the model’s prediction, that is, to let users understand which particular
words or phrases influenced the sentiment classification. Most importantly, this inter-
pretability of the proposed model in media and journalism, since sentiments from
an article can influence public opinion. Journalists and editors need to know nothing
but the sentiment classification and reasoning behind the classification to instill more
confidence in AI-driven sentiment analysis tools.
LIME stands for local interpretable model-agnostic explanations and refers to
a type of Explainable AI. This serves to give interpretability to complex machine
learning models. LIME works in a way that it approximates a black-box model by
an interpretable one locally around a given prediction. It does this by perturbing the
data-input changes and observing the ways in which the model’s prediction changes
to help identify which features—in other words, words or phrases in text data—most
strongly influenced that particular prediction. For example, in sentiment analysis, if a
model classifies a news article as negative, LIME can highlight specific words—such
as the words “crisis” and “decline”—that were most indicative of that classification.
The most important implementations of sentiment analysis in news items occur
in the arena of media and journalism. These sentiment analysis tools can be used by
both editors and journalists to determine the overall sentiment of their articles and the
reactions of the general public. This learned information will help them with content
creation strategies, especially in tailoring articles to resonate with the target audience
Leveraging Local Interpretable Model-Agnostic Explanations (LIME) … 259

amidst controversies. Explainable AI furthers this process by providing insight into


how specific parts of the content influence the audience’s perception, for example
word-choice, and therefore is important in providing a data-driven approach.
Besides that, sentiment analysis of news is helpful for both businesses and
advertisers to understand customer opinion and response toward their products or
services. Indeed, by observing trends in sentiment over time, companies can modify
their marketing strategies regarding where changes need to be made and, therefore,
enhance overall customer satisfaction. Here too, XAI makes sure businesses under-
stand not just the sentiment but the driving factors behind it, enabling them to adjust
their strategies with much more accuracy.
In this regard, the steps involved in sentiment analysis of news articles are pre-
processing of text, feature extraction, classification of sentiment of the text, and
interpretation of result. Most of the time, machine learning models have been applied
to predict the proper text sentiment by virtue of training using labeled datasets.
Integrating Explainable AI into this helps in clarifying how such models work and
provides transparency in the results.
The sentiment analysis of news acts as a powerful enabler toward making good
sense of sentiments left in digital media [3]. As technology continues to evolve, the
integration of sentiment analysis combined with Explainable AI in news is likely to
revolutionize content creation, journalism, and business strategies and contribute to
a more knowledgeable and responsive digital landscape.

2 Literature Survey

The most relevant work to come out of this is the work of Fong et al. [4] This
work presents different machine learning approaches and algorithm comparisons
for of texts and for doing sentiment analysis efficiently. The paper presents various
machine learning strategies and algorithm comparisons for text classification and for
the effective performance of sentiment analysis. There are three classes on which the
text is classified: positive, negative, and neutral. This work will, therefore, suggest
that Naïve Bayes is efficient to put to use classifier by the purpose of sentiment
analysis since it gives better accuracy as compared to other classifiers that have been
used for sentiment analysis. The work given in [5] gives tasks which are addressed
semantics parser. The semantic parser provides method for extracting concepts from
a sentence; this task includes dividing sentences into or splitting of sentences. This
work addresses the fine-grained sentiment analysis, which is commercialized viable.
Zhou et al. [6] In this work, opinion mining tasks are focused on. This work imple-
ments tweet sentiment analysis. Model (TSAM), which can capture social interests
and people’s opinions about some social events. Opinion sentiment and emotion
studies of sentiment analysis are expressed in text and were stated by [6]. This study
explains how sentiment analysis feature extraction tasks operate.
260 P. Tewari et al.

It suggests that just words that convey a hint of opinion should be used for senti-
ment analysis, rather than all the words. The purpose of this work is to demon-
strate the advantages of developing an intelligent system for sentiment analysis
based on Lexicon. Work [6] provides many techniques to increase classifier accu-
racy, including sentiment analysis using naïve Bayes. They employ n-gram negation
handling and feature selection which leads to increased efficiency. They concentrate
on developing a generalized approach to the problem of classifying a large number
of texts.
Lin et al. [7] proposed an evaluation of XAI methods’ interpretability through
neural backdoors, which triggers the ground truth for inputs to evaluate whether
those regions identified by an XAI method are really relevant to its output. Despite
the fact that our survey reaches following the same direction of Lin’s survey [7], we
extend the recent chosen survey’s scope with the LIME method and use CAM instead
of Grad-CAM, Guided Grad-CAM, owing to the generality of CAM. Furthermore,
we provide a new evaluation metric with the aim to generalize the evaluation on
results of XAI methods.

3 Methodology

In this approach for identifying the sentiment of news articles, we have used logistic
regression algorithm [8]. It is a supervised machine learning algorithm which is used
to classify the things, and our goal is to make the prediction of the probability of
instance belongs to which class. This is a statistical model which is commonly used
to classify the tasks into binary or multi-class. In this case, our model is used to make
the prediction of the sentiment of news articles based on their textual features.

3.1 Data Collection and Preprocessing

First we have done dataset acquisition; it gathers a dataset of news articles with
corresponding sentiment labels (e.g., positive, negative, neutral). You can use existing
datasets or create your own by manually annotating articles, and then we have done
data cleaning [9]; it removes noise from the text data, such as HTML tags, special
characters, and stop words, and then we have done text normalization which convert
text to lowercase, perform stemming or lemmatization to reduce word forms, and
handle contractions. We have used the dataset from Kaggle [10].
Data source used for sentiment analysis of news articles: The dataset used in this
sentiment analysis project has limitations. First, they do not contain clear sentiment
labels and instead rely on examples based on headlines or short descriptions, which
makes tone recognition inconsistent. Moreover, category imbalance is another issue
where some news categories, such as U.S. NEWS, are overrepresented while others,
Leveraging Local Interpretable Model-Agnostic Explanations (LIME) … 261

such as COMEDY, are underrepresented—thus, it may cause bias if only category-


level analysis is performed and not controlled for during the analysis process. The
variations in writing styles by category also impact how sentiment is interpreted;
humor/language may have the effect of creating relaxed, informal language relative
to, say, an objective tone typically displayed in U.S. NEWS occurs with differences
between COMEDY and other appropriate categories. In addition, the short length of
summaries may limit sentiment analysis since they may not represent good approx-
imation numbers for estimating article sentiment. A last ambiguity is due to the
temporal bias, since articles carry their time stamp and may mirror the socio-political
atmosphere along its timeframe, therefore introducing a bias in sentiment (e.g., more
negative articles released under crisis). Further lending to the sentiment complexity
of a surface is the fact that individual authors may bring their own bias or unique
style, embedding sentiments independent of content. Finally, certain entries might
make noise by including irrelevant information or non-sentiment-bearing descrip-
tions that can reduce classification accuracy. It can be seen that the various limitations
outlined above would indicate potential constraints when using local interpretable
model-agnostic explanations (LIME) for interpretability, such as undermining the
generalizability of sentiment interpretation against different news topics.
Dataset Acquisition. We have used a labeled dataset of news articles tagged with
sentiment labels such as positive, negative, and neutral from Kaggle [10]. Our dataset
consists of approximately 2,00,000 articles, which gives us a relatively robust sample
size for training and testing our model.
Data Cleaning. Text data cleaning was performed, which enhances model perfor-
mance. Cleaning the text included cleaning the text from HTML tags and special
characters, hence filtering out all the “noise”. It also included removing the common
stop words such as “the”, “is”, and “and” to reduce redundancy.
Text Preprocessing. We then further regularized the texts to lowercase to prevent
case representation duplication of words for example, “News” and “news”. In addi-
tion, stemming and lemmatization techniques were used so that the words further
reduced to base or root forms; for example, words such as “running” and “ran”
would be transformed to “run”. This step ensures that words with similar meanings
are treated uniformly.

3.2 Feature Extraction

Transforming textual data into numerical data that our machine learning algorithm
can process is known as text representation. Among the popular methods are bag-
of-words (BOW), which uses numbers to represent documents. Word embeddings,
which use dense vectors to record semantic associations between words, and Term
Frequency-Inverse Document Frequency (TF-IDF) [11] weight words according to
their frequency within a document and over the entire corpus (Fig. 1).
262 P. Tewari et al.

Fig. 1 Block diagram of the


proposed scheme
Leveraging Local Interpretable Model-Agnostic Explanations (LIME) … 263

Fig. 2 Classes sentiment


analysis [12]

3.3 Sentiment Labeling

Manual annotation: assigning sentiment labels for each news article itself, based
on his or her subjective judgment. Rule-based approaches classify sentiment by
identifying the presence of certain positive/negative words using predefined rules
or lexicons. Machine learning techniques predict labels using a training model and
given features extracted from text (Fig. 2).

3.4 Model Training and Evaluation

Conventionally, the division of the dataset into the training set and testing set evaluates
the performance of the general model on unseen data. Model selection chooses
appropriate machines or learning algorithms for the task at hand-sentiment analysis.
We have used the logistic regression model because the most popular statistical model
used for classification problems with results of binary class. The logistic regression
model is being applied now at the training step using the features extracted and the
labels of sentiment on the training set.
Logistic regression is a simple but, at the same time, limited model. It assumes
a linear relationship between predictors and log odds of the outcome. This model is
useless in cases when the relationship is nonlinear or complicated. The model is sensi-
tive to outliers. It is a problem in the case of multicollinearity; the coefficients become
easily distorted, which complicates their interpretation. High-dimensional data faces
risks of overfitting, though this can be coped with using regularization. Also, large
sample sizes are needed for the model to stabilize; besides, this approach does assume
independence of observations. Further, missing values cannot be handled by the
model directly. In such a complex dataset, decision trees or neural network models
would serve better.
264 P. Tewari et al.

Finally, evaluation is performed on the testing dataset by using some metrics like
precision metric, accuracy metric, recall metric, F1-score metric, and our confusion-
matrix to measure the performance of the model.

3.5 Explainable AI

Explainable AI [13], XAI for short, is a subfield that develops techniques to provide
better transparency at deeper levels and enable the user to understand and also believe
in the obtained predictions. A popular underlying technique in XAI is LIME or
local interpretable model-agnostic explanations. LIME was developed to explain
the individual predictions of any machine learning model from the simplest to the
most complex ones like logistic regression needed for sentiment analysis. Unlike
traditional black-box models, LIME approximates the model using an interpretable,
simpler model locally around a specific prediction. This makes it possible to explain,
quite clearly, why the model decided on a particular prediction.
Most fundamentally, the goal of explainability in sentiment classification is
to identify the model’s reason for choosing a certain set of sentiment labels. In
essence, LIME enhances explainability by offering interpretable and human-readable
explanations for complex machine-learning models.
In the case of sentiment analysis, LIME goes beyond categorizing sentiments as
positive, negative, or neutral and, importantly, provides an explanation as to why a
certain sentiment is predicted. It shows journalists, editors, and users which specific
words and phrases most heavily influenced the classification of sentiment.
With LIME integrated, there is a possibility to indicate potential biases of the
model. Because it will be showing which features are contributing to the predictions,
we would be in a position to assess whether the model is relying too heavily on
certain words which may not be relevant for making sure fairness in the sentiment
analysis process occurs. This level of explainability is important in trusting AI-
driven systems; even more critical in sensitive areas like journalism and business.
This will be achieved by implementing LIME, which will enable the accountability
and transparency of the sentiment analysis model addressing the concerns of the
“black-box” nature of machine learning models and give an insight into what factors
drive the sentiment predictions.
This study adds to this emergent domain with a focus on e-paper news articles
and proposes showcasing the part of Explainable AI in promoting transparency, trust,
and fairness in automated news sentiment detection.
Transparency: Thinking of transparency as being able to understand how a model
“thinks”. In AI, it’s opening up the mind of the model, allowing one to be able to
reason why the model has made some prediction. It is especially in areas such as
news sentiment analysis where people need to know why a particular piece of text is
classified as positive or negative. According to Bellotti and Edwards [14], the more
“intelligible” the logic of a model becomes, the more accessible the predictions made
by AI will be to the users. In that regard, tools such as LIME reveal to users that
Leveraging Local Interpretable Model-Agnostic Explanations (LIME) … 265

which words or phrases have the most impact on the sentiment score, which, from
that point on, makes the process behind AI’s decision-making much clearer [15].
Trust: People trust AI if they discover that it acts for the right reason. There is
evidence that in case a traceable explanation can be provided by AI systems, though
people do not agree with the model’s prediction, they can trace through [17]. Using
LIME to explain model outputs, your research is building that trust. For instance,
if a journalist can perceive why the model labeled the article as negative, then they
will have more trust in relying on it. A study by Chen et al. [16] also indicated that
people feel more assured that they can rely on such automated judgments when AI
models provide clear explanations.
Fairness: For AI, fairness implies that a model should not unfairly favor some
words over others or some groups of words over other groups of words or topics. In
sentiment analysis, this could manifest as the query of whether specific words carry
their biases in unintended ways with the result of biased reporting. Corbett-Davies
et al. [18] have shown how XAI methods like LIME can expose biases by showing
patterns in predictions and how the resultant effects help developers modify models
to handle all topics more equitably.

3.6 Model Refinement

This would come with the model’s hyperparameter tuning [19], which optimizes the
performance of the model with the aid of key hyperparameters according to learning
rates or the intensity of regularizations. The ensemble methods combine a number
of models to improve general accuracy and robustness, leveraging the strengths of
each algorithm.

3.7 Interpreting Results

Our analysis focused on the interpretation of model performance, including discus-


sions on any challenges faced and pointing out any remarkable observations.
Including methods of Explainable AI provided insight into how the models came
up with their decisions; we observed how the model differentiated between positive
and negative and neutral sentiments. Interpretable for XAI also made refinement and
improvement of the model possible by understanding feature contributions (Figs. 3,
4 and 5).
266 P. Tewari et al.

Fig. 3 Most influential neutral words [20]

Fig. 4 Most influential negative words [20]

Fig. 5 Most influential positive words [20]

4 Model Evaluation and Performance Analysis

The performance of the sentiment analysis model of news articles was analyzed
with accuracy metric, precision metric, recall metric, F1-score metric, and confusion
matrix. Further, the behavior of our model training was examined with visualization
approaches. The subsections indicating these evaluations follow.
Leveraging Local Interpretable Model-Agnostic Explanations (LIME) … 267

4.1 Accuracy of the Model

Accuracy metric is the measurement of the general share of correct predictions


in the model made, in general, a characteristic of its performance. Mathemati-
cally, it is given by the formula below. Thus, the accuracy of the previous model
was 0.8399990454827471 or 83.99%, and accuracy of our current model using
Explainable AI is 0.8400229084140696 means that a little more than 84% of all
the predictions were correct.
Its range is [0,1], and its formula is

Acc. = (TP + TN)/(TP + TN + FP + FN) (1)

4.2 Precision

Precision gives the ratio of predicted positive instances that actually have a positive
nature, hence showing the accuracy of the positive predictions. Mathematically, it
is given by the formula below. From the classification report below, the precision
values have been 0.81 for negative, 0.83 for neutral, and 0.88 for positive predictions.
Its range is [0,1], and its formula is (Fig. 6)

Pre. = (TP)/(TP + FP) (2)

Fig. 6 Precision graph


268 P. Tewari et al.

4.3 Recall

Recall gives the value of true positive rate. It gives a measure of how well the model
predicts true positives. Mathematically, it is given by the formula below. The recall
values, as indicated by the classification report, are 0.46 for negative, 0.95 for neutral,
and 0.74 for positive predictions.
Its range is [0,1], and its formula is

Rec. = (TP)/(TP + FN) (3)

4.4 F1-Score

The F1-score represents the balanced measure of the precision and recall. It is derived
from harmonic mean of both of them. Therefore, it represents them into a single
metric. Mathematically, it is given by the formula below. Based on the classification
report, F1-scores come out to be 0.58 for negative, 0.89 for neutral, and 0.80 for
positive predictions.
Its range is [0,1], and its formula is

F1 = 2(Pre. ∗ Rec.)/(Pre. + Rec.) (4)

4.5 Lime

LIME [21] is one of the types or methods in Explainable AI which is used for expla-
nation of the outcomes of our ML model that approximates complex models with
simpler, interpretable models. In our sentiment analysis news, LIME was essential in
making our model’s predictions explainable and, therefore, relatively easy to identify
key phrases or words driving such classification; in other words, to understand why
certain classification of a news article was made as positive, negative, or neutral.
LIME gave our model more interpretability, particularly in the case of complex
models, such as neural networks. By transparency, it became more trustworthy, as
one could validate if the predictions were correct. To add, LIME supported the iden-
tification of areas where the model misclassified sentiment, thus we could continue
tuning and enhancing its performance. Therefore, LIME became a key technology
in understanding the model of sentiment analysis and then refining it.
As you can see below, we have pasted the result from the output we got from the
model here you can see how its working later on it was just used to tell the status
that the news is positive or negative or neutral, but now its highlighting the words
Leveraging Local Interpretable Model-Agnostic Explanations (LIME) … 269

Fig. 7 Explanation by explainable AI

Fig. 8 Local explanation for


class neutral

which is indicating the real sentiment of the news which can be done by generative
AI feature known as LIME and its explaining too, so it integrated with Explainable
AI (Fig. 7).
This study’s main goal is to progress sentiment analysis in news by using local
interpretable model agnostic explanations (LIME). Conventional sentiment analysis
models are often opaque in their operations such as “black-box”. Lack transparency
regarding prediction processes. This research seeks to address this issue by incor-
porating LIME to improve interpretability and offer more profound insights into
sentiment patterns, in news articles. The research examines how well LIME clarifies
the reasons behind model predictions in order to create a trustworthy framework for
analyzing sentiment, in journalism settings (Fig. 8).

4.6 Practical Contributions

The major practical contributions of this study to sentiment analysis in news media
are improved interpretability.
270 P. Tewari et al.

Utilizing LIME, this work gives us a way to see the textual features (specific
words/phrases) that are responsible for classifying a news article into positive or
negative sentiments. Such a level of interpretability is vital in journalism, as it helps
to understand how the sentiment was predicted, which can help guide reporting
practice ethical reporting and mitigation of bias.

4.7 Confusion Matrix

Confusion matrix [22] is a matrix which gives us a summary of the classification


performance of the model by comparing the predicted values to actual values. In this
case, it would show how well the model differentiates between the three sentiment
categories: “positive”, “neutral”, and “negative”. Each little box or you can say
cell of the matrix represents the total number of instances predicted to belong to
one class but actually belonging to another. Diagonally element of this matrix table
represents the correct predictions, and off-diagonally elements indicates the error in
classifications of these. By examining the distribution of values, you can identify
patterns and areas where the model may struggle. For example, if many instances
are misclassified between “positive” and “neutral”, it suggests that the model has
difficulty distinguishing between these two sentiments.
In Fig. 9, we didn’t used LIME or Explainable AI; it was completely done in linear
regression model, and the accuracy was 83.9% which get improved using LIME in
Fig. 10 analysis more negative and positive word in news with more accuracy of
word identification for positive, neutral, and negative words (Table 1).

4.8 Improvements or Insights Over Existing Methods

Our work extends the existing sentiment analysis approaches by formalizing explain-
ability as a first-class entity that is generally lacking in the standard structure of
models performing sentiment [8, 9]. LIME integration brings a few advancements.
More insights in predictions of models: Traditional black-box sentiment analysis
models would simply give a number and mark it positive or negative; however,
with the help of LIME, one can look for a specific term or theme that supports the
complete sentiment. In news analysis, for example, the ability to segment headlines
by sentiment can provide a data-informed, cogent editorial choice that further primes
content quality.
Decrease Model Bias and Increase Transparency: Results indicate that LIME
offers a way to reveal potentially endogenously set biases in sentiment models by
offering insight into the language patterns that underlie particular sentiment judg-
ments. Transparency is critical in key AI use cases such as journalism to foster trust
and credibility.
Leveraging Local Interpretable Model-Agnostic Explanations (LIME) … 271

Fig. 9 Confusion matrix of older study without explainable AI

5 Conclusion

Sentiment analysis of news on e-paper gives insight into emotional drivers in creating
digital content. The distribution of sentiment was quite different, with neutral senti-
ment being the leading one, and slight tone changes often reflected the reader’s
reaction to events and topics discussed. Logistic regression performed well in terms
of precision, recall, and F1-score measures. Explainability approaches for AI models,
such as LIME, provided greater internal interpretation abilities, indicating exactly
which words or phrases led to the positive or negative classification. Transparency
enables users to be more confident in the system, and it allows journalists and content
developers to easily detect where biases lie with the sentiment analysis system, thus
increasing fairness as well as ethical use of AI by the media. LIME clarified the
model’s decision-making process, thereby facilitating a more trustworthy framework
that enabled responsible content creation and ethical reporting.
In addition, this would mean that any shift in tone by the writers could affect
not only the attitude of the readers but also their level of engagement in terms of
comments and shares. This calls for the need to apply Explainable AI in the analysis
of sentiment, fairness, and accountability of materials for easy comprehension of
audience feelings in a concise and data-driven manner. Further studies can develop
272 P. Tewari et al.

Fig. 10 Confusion matrix of current study with explainable AI

Table 1 Sentiment results


Class Precision Recall F1-score Support
Negative 0.81 0.46 0.58 4666
Neutral 0.83 0.95 0.89 26,318
Positive 0.88 0.74 0.8 41,906
Macro Avg. 0.84 0.71 0.76 41,906
Weighted Avg. 0.84 0.84 0.83 41,906

these models, hence contributing to the growth of more ethical and transparent AI
across digital media platforms.

References

1. Basant A, Namita M, Pooja B, Garg S (2015) Sentiment analysis using common-sense and
context information. In: Computational intelligence and neuroscience. Hindawi Publishing
Corporation
Leveraging Local Interpretable Model-Agnostic Explanations (LIME) … 273

2. Nadkarni PM, Ohno-Machado L, Chapman WW (2011) Natural language processing: an


introduction. J Am Med Inform Assoc 18(5):544–551
3. Impact of digital media on society. IJCRT 8(5):JCRT2005354 (2020)
4. Fong S, Zhuang Y, Li J, Khoury R (2013) Sentiment analysis of online news using MALLET.
In: International symposium on computational and business intelligence, pp 301–304
5. Raina P (2013) Sentiment analysis in news articles using sentic computing. In: IEEE 13th
international conference on data mining workshops. IEEE, USA, pp 959–962
6. Zhou X, Tao X, Yong J, Yang Z (2013) Sentiment analysis on tweets for social events. In: 2013
IEEE 17th international conference on computer supported co-operative work in design. IEEE,
China, pp 557–562
7. Lin Y-S, Lee W-C, Celik ZB (2020) What do you see? Evaluation of explainable artificial
intelligence (XAI) interpretability through neural backdoors
8. Tian Z (2020) Logistic regression model optimization and case analysis
9. Ying K, Hu W, Chen JB, Li GN (2021) Research on instance-level data cleaning technology.
In: 2021 International conference on artificial intelligence, big data and algorithms (CAIBDA).
IEEE, Xi’an, China, pp 238–242
10. Dataset: Kaggle news category dataset. [Link]
egory-dataset. Accessed 12 Oct 2024
11. Jain S, Jain SK, Vasal S (2024) An effective TF-IDF model to improve the text classification
performance. In: 2024 IEEE 13th international conference on communication systems and
network technologies (CSNT). IEEE, Jabalpur, India, pp 1–4
12. NLP Image: 5 things you need to know about sentiment analysis classification. [Link]
[Link]/2018/03/[Link]. Accessed 12 Oct 2024
13. Letrache K, Ramdani M (2023) Explainable artificial intelligence: a review and case study
on model-agnostic methods. In: 2023 14th international conference on intelligent systems:
theories and applications (SITA). IEEE, Casablanca, Morocco, pp 1–8
14. Bellotti V, Edwards K (2001) Intelligibility and accountability: human considerations in
context-aware systems. Hum-Comput Interact 16(2–4):193–212
15. Adadi A, Berrada M (2018) Peeking inside the black-box: a survey on explainable artificial
intelligence (XAI). IEEE Access 6:52138–52160
16. Chen X et al (2022) Effects of transparency on trust in AI systems. Int J Multidiscip Res 6(2):5
17. Miller T et al (2017) Explanation in artificial intelligence: insights from the social sciences. J
Exp Psychol 40(1):63
18. Corbett-Davies S et al (2017) A Computer program used for bail and sentencing decisions was
labeled biased. J Comput Sci Public Policy 17(4):29–35
19. Arden F, Safitri C (2022) Hyperparameter tuning algorithm comparison with machine learning
algorithms. In: 2022 6th international conference on information technology, information
systems and electrical engineering (ICITISEE). IEEE, Yogyakarta, Indonesia, pp 183–188
20. Image. CS4641 Project visualization. [Link] Accessed 12
Oct 2024
21. Rainio O, Teuho J, Klén R (2024) Evaluation metrics and statistical tests for machine learning.
Sci Rep 14:6086
22. Gramegna A, Giudici P (2021) SHAP and LIME: an evaluation of discriminative power in
credit risk. Front Artif Intell 4:752558
Neural Network-Based Security System
by PCA and Moth Flame Optimized
Features Learning Model

Abhilasha Sinha, Rakesh Kumar Tiwari, and Onkar Nath Thakur

Abstract Communication network increases the productivity of every field. So


many of attacks were developed by intruders to lay down a running system for
unethical means. Some of research area work on feature optimization and transfor-
mation methods to increases the prediction accuracy like image processing. This
paper has proposed a model that identifies the set of effecting features by use of
moth flame optimization model. Selected features were further transformed to get
the feature set having high variance, for this principal component analysis (PCA) was
implemented. Optimized feature set were used for the training of learning model.
Experiment was done on real dataset of computer network having different set of
attacks. Result shows that the proposed model has improved the work efficiency.

Keywords Deep learning · Intrusion detection · Feature optimization · Genetic


algorithm · Soft computing

1 Introduction

Traditional networks have gained widespread recognition for their significant poten-
tial across various fields, including healthcare, transportation, and industrial sectors.
These networks connect a variety of physical devices, each with moderate computa-
tional power and storage capabilities. As the number of connected devices continues
to grow, a large volume of data is produced, making traditional networks increasingly
attractive to cyberattackers.
Key challenges in traditional networks include scalability, efficient resource
management, and addressing cybersecurity issues [1]. Implementing critical security
measures is essential but becomes more complex as networks expand. In recent years,

A. Sinha (B) · R. K. Tiwari · O. N. Thakur


Technocrats Institute of Technology and Science, Bhopal, India
e-mail: abhilasasinha1@[Link]

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 275
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
276 A. Sinha et al.

machine learning (ML) and deep learning (DL) techniques have emerged as effec-
tive solutions for processing and analyzing vast amounts of data, offering improved
performance in computational tasks [2].
The need for robust cybersecurity in traditional networks has become a pressing
issue, driving the development and deployment of efficient Intrusion Detection
Systems (IDS) at key network points [3]. Over the past few years, both ML- and
DL-based IDSs have demonstrated their effectiveness in identifying cyberattacks
on traditional networks. Devices within traditional networks typically offer greater
capacity, fewer resource limitations, and more manual control than those found in IoT
environments. However, these more advanced systems still require strong security
protocols to defend against emerging cyber threats [4]. Given the enormous amount
of data generated by traditional networks, rapid and efficient attack detection is
essential [5].
Advances in software-defined networking (SDN) have led to the development
of enhanced intrusion detection and prevention systems for traditional networks
[6]. These systems leverage the flexibility of SDN to create proactive defenses,
enabling real-time detection of potential threats. This programmability, combined
with a broad network overview, makes SDN a highly effective solution for addressing
the challenges of maintaining security in traditional networks [7].

2 Related Work

Kan et al. in [8] presented an innovative approach to detecting malicious activities


within the Internet of Things (IoT) environment. They introduced an advanced hybrid
model known as adaptive particle swarm optimization convolutional neural network
(APSO-CNN) [8]. This approach draws on the collective intelligence of multiple
particles working together to optimize the parameters of a neural network designed to
monitor and detect anomalies in IoT systems. The particle swarm works dynamically,
adjusting the settings of the CNN to ensure efficient detection, and the effectiveness
of the model was rigorously evaluated using relevant datasets to confirm its learning
accuracy.
Fatani et al. in [9] developed a robust security system by combining deep learning
with optimization techniques [9]. Initially, they employed convolutional neural
networks (CNNs) to identify critical features within the data, and then enhanced
the feature selection process through a method called growth optimizer (GO). To
further refine the search for optimal features, they applied the Whale Optimiza-
tion Algorithm (WOA). This dual-layered approach of CNN and optimization was
tested on various datasets to measure its performance in identifying and preventing
cyberthreats.
Saied et al. in [10] conducted an extensive comparative study to determine the
most effective machine learning methods for detecting security issues within IoT
networks [10]. They tested a range of algorithms, including adaptive boosting and
gradient descent boosting, to assess their accuracy and speed in identifying and
Neural Network-Based Security System by PCA and Moth Flame … 277

mitigating attacks. Their research aimed to pinpoint the most efficient techniques for
real-time threat detection in IoT systems.
Zambare et al. in [11] developed a problem-detection method utilizing graph
neural networks (GNNs) to identify vulnerabilities in virtual networks [11]. They
focused on automotive cybersecurity, testing their GNN-based model on car-hacking
scenarios where the internal computer systems of vehicles are compromised. The
GNN analyzed data patterns to accurately classify the types of attacks and provided
insight into the potential cybersecurity threats in the automotive industry.
Kalyanam et al. in [12] proposed a resource-efficient method to ensure that security
protocols on IoT devices operate without overburdening computational resources
[12]. They used a technique known as threshold-based pruning, which systematically
eliminates unnecessary components from the network model, similar to trimming a
tree. The method was tested on the LeNet architecture, using diverse datasets to
evaluate its effectiveness after the pruning process, confirming that it maintained
robust performance while reducing computational overhead.

3 Proposed Methodology

This section provides a detailed overview of the proposed PCA and Moth Flame-
based Network Security (PCAMFNS). The overall structure of the proposed model
is shown in Fig. 1, which illustrates the key stages involved, including dataset
processing, dimensionality reduction, and training blocks. Each of these components
is explained in detail under separate headings within this section.

3.1 Dataset Cleaning

The input dataset used for malicious session detection contains numerous features,
each with varying levels of significance. In this initial stage, the dataset under-
goes a cleaning process to remove irrelevant or redundant information, ensuring
that only meaningful data is used for further analysis [13, 14]. For instance, the
dataset employed in this research contains several fields, but the first few feature
values were deemed unnecessary for detecting malicious sessions [15]. Examples of
such fields include session IDs, connection types, and transferring protocols, which
do not contribute significantly to identifying malicious activity. By eliminating these
superfluous elements, the dataset becomes more focused and efficient, allowing the
model to concentrate on the most critical features.

CD ← Dataset_Cleaning(RD) (1)
278 A. Sinha et al.

Fig. 1 Block diagram of


PCAMFNS network
intrusion detection

In Eq. (1), RD is raw dataset and CD is clean dataset matrix. Processed dataset
was arranged in matrix of row and column where each row is session and columns
are feature set of a session.

3.2 Feature Optimization

After the dataset cleaning stage, the input cloud data (CD) matrix undergoes feature
optimization, where the moth flame optimization algorithm (MFO) is applied [16].
Neural Network-Based Security System by PCA and Moth Flame … 279

The purpose of this step is to minimize the number of features in the training vector
while simultaneously improving the accuracy of the learning model. MFO works
by optimizing the selection of features, thereby enhancing the overall efficiency and
effectiveness of the detection process.

3.3 Moth Flame Optimization Algorithm

The moth flame optimization algorithm (MFO) draws inspiration from the natural
behavior of moths. In this algorithm, each chromosome is metaphorically represented
as a moth. The moths attempt to navigate toward a flame, which in this context
symbolizes the optimal solution. Each moth flame in the algorithm corresponds to a
chromosome, representing a candidate solution for the feature set in this work.
The algorithm’s objective is to find the optimal path by guiding the moths toward
the flame, resulting in an optimized feature set.
Generating Moth Flames Moth flames are composed of chromosomes, where each
chromosome is a potential solution representing an optimized set of features. A
moth flame is represented as a vector containing ‘n’ elements, where ‘n’ refers to
the number of columns in the CD matrix. Each element in the moth flame vector is
binary—either a 1 or a 0. A value of 1 indicates that a particular feature is selected
for training, while a value of 0 signifies that the feature is excluded. To generate
multiple moth flames, ‘p’ moth flames are created, forming a moth flame population
matrix ‘M’ with dimensions of pxn. The selection of features within the vector is
done randomly using a Gaussian random value generator function, ensuring diverse
feature combinations.

M ← Generate_Moth Flame(p, n, f) (2)

Fitness Function Each moth flame is evaluated and ranked based on its fitness,
which is determined by measuring the ‘distance’ between the moth flame and the
target solution. This evaluation is conducted using a fitness function that calculates
the detection accuracy of each moth flame. The feature vector of each moth flame
is passed through a neural network, which is trained on the input data to assess
the performance of the feature set in identifying malicious sessions. The accuracy
obtained from the detection process serves as the distance parameter, guiding the
optimization process.
Update Moth Flame Position After calculating the fitness value (F) for each moth
flame, the flames are sorted in descending order based on their fitness scores. The best
moth flame is identified from the population, which represents the most optimized
feature set. To further refine the population, the genetic algorithm introduces changes
to the chromosomes, known as crossover operations.
280 A. Sinha et al.

Crossover The success of the genetic algorithm relies on altering the chromosomes.
Based on a parameter ‘X’, a specified number of random positions in the moth
flame vectors are modified. These positions are flipped—changing 0 to 1 or 1 to
0—except for the best local moth flame, which remains unchanged. This ensures
that the algorithm continues to explore new solutions without losing track of the
best-performing moth flame. The newly generated moth flames are then evaluated
for fitness, and if the child moth flame shows better performance than its parent, the
parent is replaced. If no improvement is found, the parent continues in the population.
Feature Selection Once the iterative process is complete, the best moth flame is
chosen from the final population. The features with a value of 1 in the moth flame
vector are selected as the most relevant features for the training vector, while features
with a value of 0 are excluded. Additionally, the desired output matrix is updated at
this stage to reflect the final set of selected features.
This process ensures that the feature set is optimized for training, improving
both the efficiency of the detection system and the accuracy of malicious session
identification in cloud environments.
Principal Component Analysis (PCA) is a powerful statistical technique commonly
used for dimensionality reduction and feature optimization in various fields, including
intrusion detection systems [17]. PCA identifies the features that explain the
maximum variance in the dataset, meaning those that have the greatest influence
in distinguishing different patterns in the data. This is crucial in intrusion detec-
tion, where distinguishing between benign and malicious traffic is key. By focusing
on the most informative features, PCA helps in filtering out noise and redundant
information, leading to a more efficient detection process.
PCA transforms the original dataset into a set of new, uncorrelated features
known as principal components. The first few principal components capture most
of the variance in the data, and these components are used for further analysis. This
transformation reduces the dimensionality of the dataset without losing significant
information.

3.4 Training of Neural Network

Neural network consider takes input training vector and desired output during
training. For each set of training vector neuron weight value adjusts for a number of
epochs. Trained neural network was directly used for predicting the session class as
attack or normal.
Neural Network-Based Security System by PCA and Moth Flame … 281

4 Experiment and Results

Experimental setup: PCAMFNS and comparing model was developed on MATLAB


software. Experimental machine having 4 GB ram, i3 6th generation processor.
Comparison of PCAMFNS was done with cloud malicious session detection model
proposed in [14].

4.1 Evaluation Parameter

To test our results, this work uses the following measures: Precision, Recall, and F-
score. These parameters are dependent on the TP (True Positive), TN (True Negative),
FP (False Positive), and FN (False Negative) [16, 17].

4.2 Results

Traditional network intrusion detection models precision values were shown in


Table 1. PCAMFNS has improved the detection precision value by 9.81% as
compared to previous model in [14]. It was found from Fig. 2 that use of PCA
transformed features was increases the learning of the model and improves precision
parameter of intrusion detection. Further, it was also obtained that instead of using
CNN model for the dimension reduction moth flame optimization provides better
feature set.
Table 2 and Fig. 3 show recall values of intrusion detection models for different
session datasets. Use of moth flame optimized features has increased the recall values
of proposed PCAMFNS model by 3.612%. Feature reduction and transformation by
PCA have increased the learning efficiency of the model.
Inverse average values of precision and recall were shown in Table 3. It was found
that use of optimized features has increased the correct class detection accuracy of the
work. Genetic algorithm-based feature optimization is completely dynamic process
and found that model has not need any supervision for the reduction in features.
F-measure values were increased by 6.84% as compared to existing model.

Table 1 Precision value of


Testing sessions PCAMFNS CNN-BiLSTM
two-class attack prediction on
different testing session 4000 0.9758 0.8835
7000 0.971 0.8758
10,000 0.9684 0.8713
15,000 0.9682 0.8740
18,000 0.9664 0.8691
282 A. Sinha et al.

Fig. 2 Precision
value-based comparison of
network attack detection

Table 2 Recall value of


Testing sessions PCAMFNS CNN-BiLSTM
two-class attack prediction on
different testing session 4000 0.9842 0.9453
7000 0.9824 0.9475
10,000 0.9818 0.9476
15,000 0.9827 0.948
18,000 0.9822 0.9474

Fig. 3 Average recall


value-based comparison of
network attack detection

Table 3 F-measure value of


Testing sessions PCAMFNS CNN-BiLSTM
two-class attack prediction on
different testing session 4000 0.98 0.9133
7000 0.9767 0.9102
10,000 0.9751 0.9079
15,000 0.9754 0.9095
18,000 0.9742 0.9066
Neural Network-Based Security System by PCA and Moth Flame … 283

Table 4 Accuracy value of


Testing sessions PCAMFNS CNN-BiLSTM
two-class attack prediction on
different testing session 4000 97.9 91.15
7000 97.53 90.8
10,000 97.38 90.64
15,000 97.4 90.73
18,000 97.28 90.47

Table 5 Execution time


Testing sessions PCAMFNS CNN-BiLSTM
value of two-class attack
prediction on different testing 4000 0.2109 57.0694
session 7000 0.2622 72.68
10,000 0.2993 99.0329
15,000 0.3109 148.9202
18,000 0.4053 197.6784

Accuracy values were shown in Table 4. PCAMFNS has improved the detec-
tion accuracy value by 6.91% as compared to previous model in [14]. It was found
from Fig. 2 that use of PCA transformed features was increased the learning of the
model and improved precision parameter of intrusion detection. Further, it was also,
obtained that instead of using CNN model for the dimension reduction moth flame
optimization provides better feature set.
In PCAMFNS model not need to find the set of features at testing side as moth
flame feature were directly uses for the learning. In CNN-BiLSTM model, CNN
was implemented at the testing side as well; hence, it needs more time to feature
optimization. Table 5 shows that the proposed model has reduced the execution time
of the testing. It was found that with increase in number of sessions time get increased
in both models.

5 Conclusion

This paper has proposed a model to provide network security by monitoring the
incoming session. Based on the feature set of the incoming session, training model
was efficiently trained. This paper has work on feature optimization to increase the
learning and correct class detection accuracy. Input data features were reduced by
the use of moth flame algorithm that reduces without any guidance. Further filter
features were transformed by PCA to get more variance feature values. Experiment
was done on real dataset, and tests were done on different datasets. Result shows that
traditional network intrusion detection model PCAMFNS has improved the detection
precision value by 9.81%, accuracy by 8.91% as compared to previous model in [14].
In the future, scholars can improve the detection accuracy of underwater networks.
284 A. Sinha et al.

References

1. Zenia NZ, Aseeri M, Ahmed MR et al (2016) Energy-efficiency and reliability in MAC and
routing protocols for underwater wireless sensor network: a survey. J Netw Comput Appl
71:72–85
2. Zidi C, Bouabdallah F, Boutaba R (2016) Routing design avoiding energy holes in underwater
acoustic sensor networks. Wirel Commun Mob Comput 16(14):2035–2051
3. Khalid M, Ullah Z, Ahmad N et al (2017) A survey of routing issues and associated protocols
in underwater wireless sensor networks. J Sens 2017, Article ID 7539751, 17 pages
4. Kochanska I, Schmidt JH (2018) Estimation of coherence bandwidth for underwater acoustic
communication channel. In: Proceedings of 2018 joint conference—acoustics, pp 130–133
5. Awan KM, Shah PA, Iqbal K, Gillani S, Ahmad W, Nam Y (2019) Underwater wireless sensor
networks: a review of recent issues and challenges. Wirel Commun Mob Comput 2019, Article
ID 6470359, 20 pages
6. Hu Y, Hu K, Liu H et al (2022) An energy-balanced head devices selection scheme for
underwater mobile sensor networks. J Wirel Commun Netw 2022:63
7. He S, Li Q, Khishe M et al The optimization of nodes clustering and multihop routing protocol
using hierarchical chimp optimization for sustainable energy efficient underwater wireless
sensor networks
8. Vijay MM, Sunil J, Vincy VGAG et al (2023) Underwater wireless sensor network-based
multihop data transmission using hybrid cat cheetah optimization algorithm. Sci Rep 13:10810
9. Alabdali AM, Gharaei N, Mashat AA (2021) A framework for energy-efficient clustering with
utilizing wireless energy balancer. IEEE Access 9:117823–117831
10. Hassan AA-H, Shah WM, Habeb A-HH, Othman MFI (2020) An improved energy-efficient
clustering protocol to prolong the lifetime of the WSN-based IoT. IEEE Access 8:200500–
200517
11. Zhu J, Chen Y, Sun X, Wu J (2021) ECRKQ: machine learning-based energy-efficient clus-
tering and cooperative routing for mobile underwater acoustic sensor networks. IEEE Access
9:70843–70855
12. Kochanska I, Schmidt JH, Schmidt AM (2021) Study of probe signal bandwidth influence
on estimation of coherence bandwidth for underwater acoustic communication channel. Appl
Acoust 183
13. Tataria H, Haneda K, Molisch AF et al (2021) Standardization of propagation models for
terrestrial cellular systems: a historical perspective. Int J Wirel Inf Netw 28:20–44
14. Ghanem WAHM et al (2022) Cyber intrusion detection system based on a multi-objective binary
bat algorithm for feature selection and enhanced bat algorithm for parameter optimization in
neural networks. IEEE Access 10:76318–76339
15. Maheswari S, Arunesh K (2020) Unsupervised binary BAT algorithm based network intrusion
detection system using enhanced multiple classifiers. In: 2020 International conference on
smart electronics and communication (ICOSEC), Trichy, India
16. Johri A, Bhadula S, Sharma S, Shukla AM (2022) Assessment of factors affecting implemen-
tation of IoT based smart skin monitoring systems. Technol Soc 68
17. Mishra P et al (2021) VMShield: memory introspection-based malware detection to secure
cloud-based services against stealthy attacks. IEEE Trans Ind Inform 17(10):6754–6764
Performance Analysis of Solar Panel
Power in Maximum Power Point
Tracking System Using Artificial Neural
Network

Jagdish Chandola , Sumit Pundir , Abhishek Sharma ,


Sakshi Pundir , and Mohammed Azim Eirgash

Abstract Conventional energy resources play a significant role in energy require-


ments but show various harmful effect to the environment, including air and water
pollution. These resources are reducing due to their limited existence in the universe.
As a result, the need of searching for novel energy resources of everlasting nature
is highly in demand. So, the renewable power resources have arisen as they present
a sustainable feature of energy needs in future. Researchers have found the various
energy resources of everlasting nature, including thermal, tidal, wind, and solar,
which can be adopted as energy alternative for long-lasting use due to their ability
of endless energy production. With the capacity of producing energy, wind and solar
seem to be the most existing resources as they have most places for their harvesting.
Additionally, the solar energy is the most effective advantages due to its existence
globally. As a result, the extracting of solar energy with the ability of harvesting with
maximum accuracy comes into existence and becomes operational due to numerous
maximum power point tracking (MPPT) method. In our research, we perform the
optimization of power extraction with scaled conjugate gradient (SCG) method which
train an artificial neural network (ANN) in varying environmental situations. The
evolution of solar panel capacity has been directed for MPPT. Simulation based on
MATLAB/Simulink has been made for presenting the validation of the method for
MPPT. The results show that the SCG algorithm discover its presence for MPPT
feature in solar PV generation system.

Keywords PV panel · Boost controller · Optimization · Renewable power ·


Scaled conjugate gradient

J. Chandola · S. Pundir (B) · A. Sharma


Graphic Era Deemed to be University, Clemen Town, Dehradun, India
e-mail: Sumitpundir1983@[Link]
S. Pundir
Graphic Era Hill University, Clemen Town, Dehradun, India
M. A. Eirgash
Karadeniz Technical Üniversity, Trabzon, Turkey

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 285
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
286 J. Chandola et al.

1 Introduction

Existing energy resources, e.g., fossil fuel, are of depleting nature and showing
fast pace in depleting due to increase in vast energy demand. So, the renewable
energy resources show the advantages over the fossil fuel. Some key resources of
renewable energy generation are hydropower, wind, and solar with efficiency and
can be stored in power grids for future use [1]. Solar and wind energy are capable of
satisfying the global energy requirements and long-lasting electricity sustainability
[2]. The practice of solar energy is suitable in solar still [3], solar heater [4], and
nanostructured superhydrophobic structured coating in solar applications [5]. PV
solar energy generation positioning its role in renewable energy sector due to their
vital availability and simplicity in harvesting [6]. Various studies have been pleasing
attention toward the MPPT methods including conventional techniques as Perturb
and Observe (P&O) [7] and Incremental Conductance (IC) [8]. But the conventional
MPPT techniques are having various disadvantages as they present slow response
time, poor accuracy in varying environmental conditions and oscillates around the
maximum power point (MPP).
In recent decades, various artificial intelligence (AI) and optimization techniques
alike particle swarm optimization (PSO) have surfaced the technical disadvantages
of conventional MPPT techniques [9]. The AI-based ANN method can work in
environment-changing condition with enhanced performances. The ANN can work
well in nonlinear data. However, the ANN method presents some complexity in
modeling and require extensive data for network training [10]. The PSO method
presents robustness and enhanced adaptability in changing environmental conditions.
But the drawback of PSO includes slower convergence efficiency and becoming
trapped in local optima [11]. The genetic algorithms (GAs) show robustness in
escaping local maxima and have the ability to optimize in nonlinear characteris-
tics [12]. The fuzzy logic control (FLC)-based technique to overcome nonlinearity
characteristic is employed in successful manner. The FLC method shows flexible
and quick response to convergence in parameter variation [13].
All techniques have present differing advantages and drawbacks in tracking MPPT
imposed in PV solar panel. There is still requirement of searching and performing
different algorithms for MPPT in PV solar energy generation systems. In this context,
we present a SCG trained ANN method for the MPPT in PV solar energy harvesting
system.

2 PV Solar Power MPPT Systems

A general process diagram of PV solar power generation system is present in Fig. 1.


The diagram presents the various components employed in MPPT system. A PV
panel was employed to take input from the sun at daytime. The current (IPV ) and
voltage (VPV ) of the solar panel are taken to be monitored by the SCG trained ANN
Performance Analysis of Solar Panel Power in Maximum Power Point … 287

Fig. 1 General phase diagram of MPPT in PV solar power generation system

MPPT controller. A DC-DC converter is used to direct the output power toward the
load devices or grid storages. The MPPT controller responds to consistency of the
system and operates the system at its MPP. The controller operates for the changing
condition. The load devices may be any household devices that required electricity.
The grid storage is further usable by the other places or at the same places.

3 Methodology

The methodology section describes the component, and technique involves in antici-
pated energy generation system. The section describes the modeling and parameters
of PV panel characteristics along with ANN architecture and its training algorithm
(SCG). The simulation environment for the system setup is also presented in this
section.

3.1 Photovoltaic Model

The PV panel is consisting of PV cell. In the anticipated model, the panel is taken as
single parallel string and single series connected string with sixty cells per module.
The configured parameter of the PV panel is presented in Table 1.
288 J. Chandola et al.

Table 1 Parameter of solar panel used in experiment


Parameter Values Parameter Values
Temperature coefficient (Current), 0.102%/°C Short-circuit current, Isc 7.85 A
α
Temperature coefficient (Voltage), −0.36099%/°C Maximum power point current, 7.35 A
β Imps
Voltage on open-circuit, Voc 36 V Voltage at MPP, Vmpp 30 V

Fig. 2 ANN architecture for


PV solar data training

3.2 Neural Network Model

Figure 2 depicts the ANN architecture for advanced optimization approaches, which
is used in training by scaled conjugate gradient (SCG) method. The ANN is trained
using data representing varying irradiance and temperature levels. The network archi-
tecture is composed of an input layer (for irradiance and temperature), a hidden layer,
and an output layer that predicts the power output. Biases are included to ensure accu-
rate feedforward calculations within the ANN, optimizing the network’s ability to
predict the VMPP based on the input data.

3.3 Training with SCG Algorithm

The SCG method is a strategy for optimizing and training the ANN on two years
of data, which included irradiance and temperature. The conventional backprop-
agation algorithms have various disadvantages for MPPT in solar PV system as
they present slow convergence speed [14]. The SCG aimed at supervised learning
in nonlinear data overtakes the disadvantages of the traditional methods [15]. SCG
algorithm presents fast convergence and stable method for large-scale optimization.
SCG method combines Levenberg–Marquardt (LM) method with conjugate gradient
into determine the optimized output for MPPT condition.
Performance Analysis of Solar Panel Power in Maximum Power Point … 289

SCG optimizes the networks’ error to the MPP. SCG performs adjusting weights
and biases of the network but still presents fast and memory-efficient implementation.
The SCG algorithm does not implement entire Hessian matrix to estimate the search
direction. Formulation of search direction of the algorithm is presented in Eq. 1.
Step 1: Search direction is characterized by dk and given as below:

dk = −gk + βk dk − 1 (1)

where k is the iteration number being performed, gk determines the error function at
gradient k, and βk signifies scalar controlling contribution from previous search, and
given in Eq. 2:

β k = gkT (gk − gk−1 )/dk−1


T
(gk − gk−1 ) (2)

Here, gk-1 determines the gradient from k-1th iteration. The update rule ensures
efficient progress by combining steepest descent with conjugate directions, avoiding
redundant searches.
Step2: Scaled step size SCG modifies the conjugate gradient method by
introducing a scaling parameter σk to avoid a full line search:

xk+1 = xk + αk dk (3)

where αk is computed using a simple approximation rather than a full line search,
reducing computation.
Step3: Hessian approximation SCG avoids directly calculating the Hessian matrix
by adjusting σk , which scales the step size to ensure rapid and stable convergence.

3.4 Simulation Setup

Figure 3 illustrates the complete simulation of the proposed MPPT process, show-
casing how the various components work together to optimize power output. The
model uses a PV solar panel, configured as per the specifications in Table 1. An
RC branch (resistor–capacitor) is linked to the PV panel and in parallel with the
IGBT diode, which acts as a fast-switching element in the boost converter. An RL
branch (resistor-inductor) is placed between the RC branch and the IGBT diode to
store energy during the switching process and manage the resistive losses, improving
power conversion efficiency.
Another RC branch is connected in parallel with the IGBT diode, with a diode
between them to prevent reverse current flow, ensuring unidirectional energy transfer
to the load. The resistive load is connected to the output of the boost controller,
representing the device or system consuming the generated power. A neural network,
trained using the scaled conjugate gradient (SCG) algorithm, dynamically adjusts the
duty cycle based on real-time irradiance and temperature inputs from the PV panel,
290 J. Chandola et al.

Fig. 3 Simulation model for SCG trained function fitting NN architecture for MPPT

ensuring that the system operates at the maximum power point (MPP). This duty cycle
determines the gate pulse for the IGBT, controlling the on/off switching to regulate
output voltage. Current and voltage measurement blocks are included in the model
to monitor the performance of the system, providing real-time feedback on voltage
and current across the load, ensuring stability and optimal power delivery. This setup
effectively maximizes power extraction from the solar panel while maintaining a
stable output for the load.

4 Results and Discussion

The representative results analysis of the SCG algorithm, demonstrating efficient


resource utilization during training. The model was trained for 54 epochs, meaning
it completed 54 full passes through the training dataset (Fig. 4). With an impressive
regression value of 0.99973, the model exhibits near-perfect predictive accuracy
(Fig. 5). However, the gradient value of 444.9197 suggests that optimization is not
yet fully complete, and the performance metric of 25.2166, measured at epoch 48,
likely represents an error metric, highlighting potential room for improvement.
The performance metrics of the proposed model are presented in Table 2,
and future suggestions regarding the performance enhancement are discussed in
conclusion section including future scope of the model.
The output of the system is calculated at standard irradiance (G = 1000) and
temperature (T = 25). The results of the simulation output are present in Figs. 6
and 7. Figure 6 presents the bus voltage without MPPT, and Fig. 7 presents the bus
voltage with MPPT.
Performance Analysis of Solar Panel Power in Maximum Power Point … 291

Fig. 4 Passes performance


plot for SCG algorithm

Fig. 5 Training, testing, and validation plot for the SCG algorithm performance
292 J. Chandola et al.

Table 2 Performance of
Parameters Value Parameters Value
SCG algorithm after training,
testing, and validation Epoch 54 Regression 0.99973
Gradient 444.9197 Performance 25.2166
Space complexity Ο(n) Time complexity Ο(k × n)

Fig. 6 Output of the simulation model without MPPT presenting the bus voltage

Fig. 7 Output of the simulation model with MPPT presenting the bus voltage
Performance Analysis of Solar Panel Power in Maximum Power Point … 293

The results from the graphs represent that the system operates at a continuous
value after MPPT. The oscillations of the bus voltage get into a straight line (Fig. 7),
meaning that the system operates at its MPP and produces the possible maximum
voltage for output power. The overall simulation was run for 10 s, but the system takes
negligible time to reach MPP. So, the system takes less time to end up the oscillation
and ended with a stable solution. The SCG algorithm demonstrated superior perfor-
mance in terms of convergence speed and generalization. The ANN model trained
with the SCG algorithm was able to accurately capture the nonlinear relationships
and adjust the system for optimal performance.

5 Conclusion

The ANN plays a crucial role for implementing optimization algorithm. A detailed
study of the ANN for MPPT in PV solar panel is analyzed. The SCG algorithm was
fine-tuned to operate on MPPT. MPPT parameter of the algorithm was investigated.
The system initially presents oscillation around the MPPT while operates on the
varying conditions. The tuning of the method parameter forced the simulation system
to operate at MPP very shortly. As the system was trained for various sets of input
so the output graph shows that the MPP is achieved at tested standard irradiance and
temperature values. Additionally, the MPP at different input sets of irradiance and
temperature can be achieved. Overall implementation of the method shows that the
method is suitable for large-scale problem in optimization.
Despite these aspects, the model’s overall performance is strong, and further
fine-tuning or additional training epochs could help reduce the gradient and error
values, potentially enhancing the model’s accuracy even further. Further real-time
implementation aspects of the model suggest the algorithm is needed to perform with
irradiance and temperature sensors including other input parameters.

References

1. Kabeyi MJB, Olanrewaju OA (2022) Sustainable energy transition for renewable and low
carbon grid electricity generation and supply. Front Energy Res 9:743114
2. Al-Shetwi AQ (2022) Sustainable development of renewable energy integrated power sector:
trends, environmental impacts, and recent challenges. Sci Total Environ 822:153645
3. Negi A. Ranakoti L, Verma RP (2024) Performance evaluation of single slope tilted wick
solar still with varying salt concentrations. In: IOP conference series: earth and environmental
science. IOP Publishing
4. Kumar S et al (2024) Computational analysis of modified solar air heater having combination
of ribs and protrusion in S-shaped configuration. Int J Interact Des Manuf (IJIDeM), pp 1–12
5. Mishra A, Bhatt N, Bajpai A (2019) Nanostructured superhydrophobic coatings for solar panel
applications. Nanomaterials-Based Coatings. Elsevier, pp 397–424
6. Tawalbeh M et al (2021) Environmental impacts of solar photovoltaic systems: a critical review
of recent progress and future outlook. Sci Total Environ 759:143528
294 J. Chandola et al.

7. Fan Z et al (2021) Perturb and observe mppt algorithm of photovoltaic system: a review. In:
2021 33rd Chinese control and decision conference (CCDC). IEEE
8. Azad ML, Sadhu PK, Das S (2020) Comparative study between P&O and incremental conduc-
tion MPPT techniques-a review. In: 2020 International conference on intelligent engineering
and management (ICIEM). IEEE
9. Yap KY, Sarimuthu CR, Lim JM-Y (2020) Artificial intelligence based MPPT techniques for
solar power system: a review. J Modern Power Syst Clean Energy 8(6):1043–1059
10. Villegas-Mier CG et al (2021) Artificial neural networks in MPPT algorithms for optimization
of photovoltaic power systems: a review. Micromachines 12(10):1260
11. Wasim MS et al (2022) A critical review and performance comparisons of swarm-based opti-
mization algorithms in maximum power point tracking of photovoltaic systems under partial
shading conditions. Energy Rep 8:4871–4898
12. de Oliveira Silva M, Calili RF, Louzada DR (2021) Comparative review of MPPT algorithms.
In: 2021 IEEE 48th photovoltaic specialists conference (PVSC). IEEE
13. Bendib B, Belmili H, Krim F (2015) A survey of the most used MPPT methods: conventional
and advanced algorithms applied for photovoltaic systems. Renew Sustain Energy Rev 45:637–
648
14. Cho S-B, Kim JH (1993) Rapid backpropagation learning algorithms. Circuits Syst Signal
Process 12(2):155–175
15. Meiller MF (1993) A scaled conjugate gradient algorithm for fast supervised learning. Neural
Netw 6(4):525–533
Advanced Crop Prediction
and Collaborative Agri-Investment
Platform
Punniyakotti Varadharajan Gopirajan, T. Pradeep, R. Charan,
and K. Suresh Kumar

Abstract Agriculture is essential for economic stability, yet farmers often face chal-
lenges in crop selection and sourcing inputs. To address this, we propose a platform
that combines a commodity exchange system with predictive analytics using the
K-nearest neighbor (KNN) model to guide farmers on optimal crops for specific
locations. The model assesses nine parameters: nitrogen, potassium, phosphorus,
air quality, soil moisture, temperature, humidity, rainfall, and pH. By inputting
these, farmers can forecast suitable crops, enhancing resource efficiency and yield.
Additionally, the platform serves as a marketplace where farmers request essential
inputs—fertilizers, seeds, equipment—while buyers such as wholesalers respond
with their needs. This dual-purpose platform aligns supply with demand, minimizes
waste, and fosters collaboration. By merging predictive analytics with a commodity
exchange, we aim to improve crop production, resource allocation, and sustainable
agricultural practices. Integrating real-time data with AI-driven recommendations
empowers farmers with actionable insights that reduce risk and optimize decisions.
It will also help to increase transparency in markets and access to allow smallholder
farmers to compete more effectively in a rapidly changing agricultural landscape.

Keywords First keyword · Second keyword · Third keyword

1 Introduction

Being one of the most important pillars of the Indian economy, accounting for as
large as 17% of the GDP, enormous challenges are still faced by the sector. Even
looking at its great worth, the sector is not spared the challenges of wrong crop

P. V. Gopirajan (B) · T. Pradeep · R. Charan


Department of Computational Intelligence, School of Computing, SRM Institute of Science and
Technology, Kattankulathur Campus, Chennai, India
e-mail: gopirajp@[Link]
K. Suresh Kumar
Department of Information Technology, Saveetha Engineering College, Chennai, India

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 295
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
296 P. V. Gopirajan et al.

patterns leading to crop failures. This weakness has far-reaching effects not only
on the agricultural domain but also on the nation’s GDP. One thing is for certain:
In the agricultural scenario prevailing today, there definitely appears to be a chal-
lenge requiring a formulation of some kind of solution that can enhance agricultural
production while protecting it against crop loss situations. It will have something to
do with advanced technologies, so to say—machine learning, data computing, and
artificial intelligence—which are bound to offer a number of benefits. Such ideation
could involve the smart AI system that puts up information on the most viable crops
to plant in the given farmland in order to curtail such challenges. Here, it can be aware
of all the factors that influence crop selection by integrating data on soil and weather
and climate of the place, history of crops grown there in the past, and the present
market conditions. The proposed system aims to enhance output and productivity
by giving the farmers a suggestion for their particular farming condition. Aspects
that motivate the farmer to enter the platform the harvest that he has reaped, so the
latter would make the prices known to customers and buyers directly to connect
with the farmers and not have any middlemen, therefore shortening the distribution
chain. This will enable farmers to access the market and protect them from various
threats of price exploitation, thus giving them a better chance to participate more
effectively in the agricultural market. It tries to transform Indian agriculture at large
by enabling farmers to access vital information and tools to make the right decisions.
The said project tries to explain the ambitions regarding agricultural productivity
over generations.

2 Related Work

According to Priyadharshini et al. [1], this represents an intelligent recommenda-


tion system for crops based on the machine learning technique in recommending
the crop with respect to data about soil and weather conditions. It makes use of
algorithms such as random forest and Naive Bayes in predicting the suitable crop
at a place. Many parameters that are taken into consideration increase the yield of
crops and sustain agriculture. Kavitha [2] discusses an agricultural crop recommen-
dation system based on the applications of machine learning in agriculture. This
paper mainly deals with the regression and classification models while analyzing
the soil and climatic conditions to help the farmer choose the crops he has to culti-
vate to have maximum yields, such as decision trees. The enhancement done by
Reddy et al. [3] is followed to predict crop yield using machine learning. This paper
focuses on the linear regression and support vector machines’ techniques, which
are employed for better improvement. Therefore, this study puts more focus on the
efficient model selection process with optimized predictions of changing climate
conditions. Agarwal and Tarar [4] proposed the hybrid model which merges the
machine learning algorithms with deep learning. It includes the integration of neural
network optimization along with support vector regression. In this approach, the
highest accuracy is maintained for the models. Various datasets create an increased
Advanced Crop Prediction and Collaborative Agri-Investment Platform 297

strength for viewing while agricultural forecasting is considered. Machine and deep
learning have been applied by Meghani et al. to smart agricultural practices. Sensor
data along with satellite images have been used here to construct models on intelli-
gent crop cultivation. The developed models predict the optimal growing conditions
and help farmers make decisions.
In the paper, Zualkernan et al. [6] discuss, with respect to UAV-based imagery in
precision agriculture, the applications of machine learning in crop monitoring and
yield estimation. The paper discusses how aerial imagery can be beneficial in terms of
providing real-time data on crops to farmers for decisive decision making. Farjon et al.
[7] reviewed deep-learning-based counting methods and datasets for agriculture. This
research work considers the use of CNN in the count of plants, weeds, and crop health
conditions to advance real-time monitoring as a management tool in agriculture.
Kulyal and Saxena [8] present a review on the different approaches toward machine
learning for predicting crop yield, with an emphasis on regression and classification
techniques in vogue. A comparison of various models among themselves, such as
random forest and XGBoost, regarding their pros and cons in yield prediction, also
comes into focus. Panigrahi et al. [9] compare various regression models used for the
prediction of crop yield against the supervised learning techniques, namely decision
trees and linear regression. The paper discusses how effective these technologies are
in determining the most accurate yields using different data on soil and weather.
The authors in the work of Elbasi et al. [10] present crop yield prediction using
machine learning techniques such as k-nearest neighbors and decision trees, among
others. The authors of this work emphasize that machine learning methods can adapt
to climatic changes in other regions, thereby attaining an accuracy level for yield
prediction.

3 Proposed System

This system is integrated with financial cooperation between investors and farmers
and is changeable according to changes in agricultural productivity. These contain an
advanced crop prediction module using multiple sensors, including soil temperature,
humidity, NPK, and wetness of the soil. These sensors continuously collect real-
time data from the soil, which gets further analyzed using the K-nearest neighbors
algorithm. Then, a proper type of crop species will be indicated in front of them
which is best fitted to the soil condition. The users can therefore easily access and
use the predictor by simply logging in through a user-friendly web interface and
starting the process of the predictor.
Following is the list of major components of the system as shown in Fig. 1:
User Interface: The user interface shall be supported by a web interface where input
data can be entered and/or sensor data can be shown in real time.
Sensors: Field-deployed sensors continuously provide information on parameters
like moisture, temperature, and humidity. Such sensors must provide real-time input
298 P. V. Gopirajan et al.

Fig. 1 Sensor and website integration flowchart

at all times, as continuous data plays a very important role in determining crop
suitability.
Preprocessing Data: Various processes, such as noise removal and feature extraction,
will be done on the data received from the sensors. Data processed in this way will
also be stored in a database to allow real-time predictions and make it possible to
provide predictions that peer into the future.
Database: It is the database that has to store information on past and present envi-
ronmental information. The latter refers to current environmental state in which the
information is to be evaluated by a machine learning algorithm.
Machine Learning Model: The KNN algorithm selects the most favorable crop
relative to the present conditions. The procedure thereby involves real-time sensor
data surveillance so as to locate crop types falling under the same conditions against
the historical data kept in the system.
Predicted Crop Output: The web interface will be presenting the predicted crop to
the user in the bid to ensure recommendations that can be executed by the user.
This is depicted in Fig. 2, where the system has eased interaction among farmers
and consumers, investors. All users log in from the same point but are taken to
different modules as prescribed by their roles. For instance, farmers would make
orders of commodities to be purchased or seek investor’s assistance to fund their
businesses like upgrades in technologies or expansion of their farms. Consumers are
looking for arable land or agricultural produce and, in return, contact farmers directly.
The system supports the involvement of investors who enter into an agreement on
sustainable farming initiatives for some time and expect a profit on invested capital
to enhance farm and social services development. Farmers or consumers may also
browse land opportunities, raising agricultural investments based on land. Once such
a match is found, the system ensures effective communication among its users to
engender inter-users working habits and sustainability behaviors (Fig. 3).
Advanced Crop Prediction and Collaborative Agri-Investment Platform 299

Fig. 2 Website interaction flowchart

Fig. 3 Block of the module

4 Components Used

The proposed system architecture of an IoT-enabled agriculture monitoring system


is based on the ESP32 microcontroller. ES32 will be acting as the master controller
to connect several sensors and then transmit the information further on to a far-off
database through Wi-Fi. The components of this system are as follows:
ESP32: The microcontroller unit normally performs data collection from connected
sensors and interactions with outside systems.
The Soil Moisture Sensor: It essentially helps in measuring the quantity of moisture
in the soil to provide the needed amount of water in an optimal manner.
Soil Temperature Sensor: The device used for measuring heat in the soil necessary
for one of the most important factors required for optimum growth.
300 P. V. Gopirajan et al.

Humidity Sensor: It measures the level of humidity present in the atmosphere with
regard to the control of the ambient.
NPK Sensor: A sensor which could monitor the level of nitrogen, phosphorus, and
potassium in the soil-say, the fertilizer level judgment capability that can utilize this
technique.
Wi-Fi Dongle: This helps in sending the data picked by sensors to remote server for
extra processing.
Battery: Supplies electric power to all components, hence functional even in far-off
areas with no access to main electricity supply.
Firebase: It is a web-based service to store and process sensor data in real time.
Website: It helps in viewing the information gathered easily and creates easy access
to the agriculture condition regarding monitoring and management from afar. It also
allows for the performance of the monitoring and control of key parameters affecting
agriculture with minimum delay by this architecture. Thus, it helps in enhancing farm
operations and decision making.
The components in the circuit diagram are:
DHT11/DHT22 Temperature and Humidity Sensor:
This is used to measure not only temperature of the ambient environment along
with the relative humidity level. The sensor data generated is sent toward the
microcontroller for processing, Fig. 4.
Breadboard: It is an open-source hardware development board that assembles elec-
tronic circuits and devices for experimental work without needing soldering. It uses
jumper wires and connections for quick changes, as illustrated in Fig. 4.
NPK Sensor: A device that can analyze and measure the level of Nitrogen, N,
Phosphorus, P, and Potassium, K in the soil, which is essential for plant health, as
shown in Fig. 4.

Fig. 4 Circuit diagram


Advanced Crop Prediction and Collaborative Agri-Investment Platform 301

Fig. 5 Prototype

Soil Moisture Sensor: It determines the volumetric water content of the soil, hence
enabling the assessment of the moisture content in the soil to assist in determining
when there is a need to water the plants. This is shown in Fig. 4.
ESP32 Microcontroller: It is a microcontroller with both Wi-Fi and Bluetooth
communication capabilities; hence, it records data from the sensors and conveys
or displays it out for monitoring and control as shown in Fig. 4 (Fig. 5).

5 Result and Performance Evaluation

Case 1: Enhanced Agricultural Productivity


Figures 6 and 7 are explained by the advancements and development of the project,
which include real-time soil data monitoring using advanced sensors to perform opti-
mization in crop selection and management based on soil temperature and humidity
as well as its moisture and NPK levels. So far, these have been able to enhance
crop yields, reduce negative impacts on the environment through the precision
management of resources, and improve performance in large-scale farm produc-
tivity. Strategic analysis data allows for proper crop rotation and long-term planning
toward sustainability and resilience in agricultural practice.
Case 2: Efficient Supply Chain Management
Figure 8 explains that procurement is made centralized through the e-commerce
portal. It, therefore, reflects a very efficient process by which farmers would acquire
quality products like sensors and components and fertilizers, among others. The
user enjoys competitive pricing, variety of products, ease in interface for browsing,
buying, and management of their orders. Such a system enhances operability,
reliability, and well-fitted insertion in farm management practices.
302 P. V. Gopirajan et al.

Fig. 6 Crop prediction page

Fig. 7 Prediction result


Advanced Crop Prediction and Collaborative Agri-Investment Platform 303

Fig. 8 E-commerce platform

Case 3: Promotion of Economic Collaboration and Sustainability


Figures 9 and 10: This investment module of the project facilitates financial invest-
ments in partnerships between investors and farmers, championing agriculture
projects and economic development. Investors diversify into sustainable agriculture,
hence having stable returns, while contributing to local community development
and environmental stewardship. Farmers gain effective access to expansion capital,
investment in technology, and operational upgrades so that farming is resilient and
long-lasting through sustainable agricultural means.
Figure 11 shows the performance of various machine learning models developed
during this assignment. Various models explored in this assignment have included
K-nearest neighbors (KNNs), support vector machines (SVMs), logistic regression,
random forest, and decision trees among others, as shown below:
Results of their accuracy are as follows:
As shown in Table 1, model comparison, the model comparison showed that KNN
gave the highest for accuracy with 0.95, precision with 0.96, and ROC-AUC with
0.97; hence, generally, KNN is recommended. The other model, which is the random
forest, also showed very similar performance and excellent in most metrics, such as
accuracy at 0.94 and precision at 0.95. Similarly, the performance of SVM is also
good, but the slight gap between its recall with a value of 0.90 and F1-score with a
value of 0.92 shows the difference, and it has with KNN and random forest. Logistic
regression and decision tree are performing averagely and result in lower accuracy
with higher false predictions-higher FPs and FNs. KNN’s confusion matrix depicts
its efficiency with hardly any false positives and negatives.
As Fig. 12 explains, the comparison of radar plot of the five models, including
KNN, SVM, logistic regression, random forest, and decision tree, based on metrics
304 P. V. Gopirajan et al.

Fig. 9 Web implementation 1

such as precision, recall, accuracy, F1-score evaluation: All models are well and
consistently positioned across the metrics. Each of these models is performing almost
equally well. KNN seems to be one of the best models, as all of the scores attained
in different categories seem close to perfection and thus had an overall impressive
capability for classification purposes. The above radar plot points out excellence,
whereby KNN is the most competitive among those models concerning the criteria
being evaluated here.
As Fig. 13 implies KNN’s ROC curve, the left plot is the ROC curve for the
KNN model that had achieved an AUC of 1.00, or perfect classification performance
without a single false positive. On the right, the precision-recall curve shows that
Advanced Crop Prediction and Collaborative Agri-Investment Platform 305

Fig. 10 Web Implementation 2

KNN maintains quite high precision and recall until it has not approached near-
perfect recall at about 0.95, though it does drop a little afterward. These metrics
ensure that in this case, KNN highly performs.
306 P. V. Gopirajan et al.

Fig. 11 Evaluation metrics


for the proposed work

Table 1 Performance comparison of various models


Model Accuracy Precision Recall F1-score Specificity ROC-AUC MCC Confusion
score matrix
(TP, FP,
TN, FN)
KNN 0.95 0.96 0.93 0.94 0.94 0.97 0.89 [[50, 3],
[2, 89]]
SVM 0.92 0.94 0.90 0.92 0.91 0.95 0.85 [[48, 5],
[3, 88]]
Logistic 0.90 0.91 0.88 0.89 0.90 0.93 0.83 [[47, 6],
regression [5, 86]]
Random 0.94 0.95 0.91 0.93 0.93 0.96 0.88 [[49, 4],
forest [4, 88]]
Decision 0.91 0.92 0.89 0.90 0.90 0.94 0.84 [[48, 5],
tree [5, 87]]
Advanced Crop Prediction and Collaborative Agri-Investment Platform 307

Fig. 12 Evaluation metrics


comparison of various
models

Fig. 13 ROC and precision-recall curves for KNN

6 Conclusion

The project concludes with the proposal for a new engineering technique that is
bound to change the face of agriculture scientifically by incorporating advanced
technologies. It basically deals with using a set of sensors applied in monitoring
a few environmental parameters: soil temperature and humidity, soil moisture, and
the level of nitrogen, phosphorous, potassium (NPK). These sensors will create a
network via the microcontroller ESP32 to enable farmers to collect data in real time,
using Wi-Fi technology. This would, in turn, facilitate accurate and correct predic-
tions about crops to provide sufficient information to farmers in order to make the
308 P. V. Gopirajan et al.

right decisions with a view to optimizing efficient computing power use, under-
standing the pattern of production, and increasing crop yield. The aspect that really
points to great enhancement of precision agriculture is an engineered system putting
sensors in place along with other technologies. It will be helpful to farmers for
more complete information in real time about moisture and what kind of action
to perform, which reduces waste and enhances productivity. Efficient resource use,
starting from sustainable agriculture, is another crucial factor beyond productivity
in today’s farming world, where precision agriculture has been mentioned within
the context of the twenty-first century. Moreover, this science project introduced an
e-commerce platform where products directly related to the production factors and
other complementary components of the system can find their market where such
products can be found and delivered with greater ease, such as seeds, fertilizers, and
other related products. This will enable the implementation to offer good technology
in sourcing produce; logistical problems will be avoided, and it will be able to allow
producers to focus on good practices of growing, increase crop yield, and raise their
incomes. The project also presents an investment comprising investment modules
where investors are seen to be of importance in cooperating with the producer in
a new and efficient way, operating like a partner who is empowered with part of
the shareholders’ investments in the industry. It will move capital inflow in agricul-
ture that will stabilize their growth, and access to technology and good practices.
Motivation is for investors and producers alike-to support agriculture through good
practice, thus creating it, raising, or increasing the general income level that exists
in this sector.

References

1. Priyadharshini A, Chakraborty S, Kumar A, Pooniwala O (2021) intelligent crop recommenda-


tion system using machine learning, pp 843–848. [Link]
9418375
2. Kavitha (2023) Crop recommendation system using ML. Int Sci J Eng Manag 02. [Link]
org/10.55041/ISJEM00457
3. Reddy S, Naini S, Thatipamula S (2021) An optimized machine learning approach for predicting
various crop yields. SSRN Electron J. [Link]
4. Agarwal S, Tarar S (2021) A hybrid approach for crop yield prediction using machine learning
and deep learning algorithms. J Phys: Conf Ser 1714:012012. [Link]
6596/1714/1/012012
5. Meghani S, Vele J, Madrecha K, Salunke A (2024) Smart crop cultivation: harnessing machine
and deep learning for informed agricultural practices. In: 3rd international conference on
applied artificial intelligence and computing (ICAAIC). IEEE, pp 1264–1269
6. Zualkernan I, Abuhani DA, Hussain MH, Khan J, Elmohandes M (2023) Machine learning for
precision agriculture using imagery from unmanned aerial vehicles (UAVs]): a survey. Drones
7:382. [Link]
7. Farjon G, Huijun L, Edan Y (2023) Deep-learning-based counting methods, datasets, and
applications in agriculture: a review. Precis Agric 24:1–29. [Link]
023-10034-8
Advanced Crop Prediction and Collaborative Agri-Investment Platform 309

8. Kulyal M, Saxena P (2022) Machine learning approaches for crop yield prediction: a review.
In: 7th international conference on computing, communication and security (ICCCS). Seoul,
Korea, pp 1–7. [Link]
9. Panigrahi B, Rao KCR, Sujatha M (2023) A machine learning-based comparative approach to
predict the crop yield using supervised learning with regression models. Procedia Comput Sci
218:2684–2693. [Link]
10. Elbasi E, Zaki C, Topcu AE, Abdelbaki W, Zreikat AI, Cina E, Shdefat A, Saker L (2023) Crop
prediction model using machine learning algorithms. Appl Sci 13(16):9288
11. Pant J, Pant RP, Singh M, Singh D, Pant H (2021) Analysis of agricultural crop yield prediction
using statistical techniques of machine learning. Mater Today: Proc 46. [Link]
[Link].2021.01.948
12. Wang Y, Zhang Q, Yu F, Zhang N, Zhang X, Li Y, Wang M, Zhang J (2024) Progress in research
on deep learning-based crop yield prediction. Agronomy 2264. [Link]
my14102264
13. Yu F, Wang M, Xiao J, Zhang Q, Zhang J, Liu X, Ping Y, Luan R (2024) Advancements in
utilizing image-analysis technology for crop-yield estimation. Remote Sens 16:1003. https://
[Link]/10.3390/rs16061003
14. Priyadharshini A, Chakraborty S, Kumar A, Pooniwala OR (2021) Intelligent crop recommen-
dation system using machine learning. In: 2021 5th international conference on computing
methodologies and communication (ICCMC). Erode, India, pp 843–848. [Link]
1109/ICCMC51019.2021.9418375
15. Kamangira SO, Medi C (2024) Tilime crop yield prediction using machine learning algorithms.
I-Manager’s J Data Sci Big Data Anal 2(1):8–13
Supervised Machine Learning Solution
for Predict Heart Disease

Neha Para, Kavita Kushwah, Neha Jadon, and Aayush Shrivastava

Abstract Recently, heart disorders have taken the position of being the main cause
of death. It is expected that this trend will be in place for the foreseeable future. Thus,
doing an accurate diagnosis of heart disease turns out to be a top priority to help gain
an improvement in prognosis for the patients while reducing the rates of mortality
involved in the condition. Coronary artery disease and chronic heart failure are some
of the most common reasons behind heart attacks. Of all the diagnostic procedures
applied to determine whether a patient has heart disease, an angiography is the
most frequently used. Actually, it is true that researchers today are channeling an
immense proportion of their efforts toward cardiovascular disorders. This led to the
establishment of a fruitful diagnostic strategy due to the introduction of techniques
from artificial intelligence, such as machine learning. The use of machine learning
at any stage in the diagnostic process produces accurate and efficient results that
therefore contribute to timely heart disease detection. Through the usage of our
research project in this research paper, we will be able to reach the objective of
achieving an accurate measurement of the prediction of heart disease. Besides two
other methods, we also made use of the dataset that was available from the University
of California, Irvine and used logistic regression. Using logistic regression, we could
attain some degree of accuracy of 92.3–92.3%.

Keywords ML · UCI dataset · LR · Decision tree classifier · SVM

N. Para (B)
Department of Artificial Intelligence, MITS (Deemed to Be University) Gwalior, Gwalior, India
e-mail: neha.phd24@[Link]
K. Kushwah
Department of Computer Science and Engineering, Ajeenkya DY Patil University, Pune, India
N. Jadon
Department of Electronics and Instrumentation, SGSITS Indore, Indore, India
A. Shrivastava
Depaertment of Computer Science and Engineering, IITM, IES University Bhopal, Bhopal, India

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 311
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
312 N. Para et al.

1 Introduction

Despite the fact that it continues to be a significant contributor to death rates all
over the world, heart disease remains a major cause for concern for people and
healthcare infrastructures all over the world. This disorder, which encompasses a
number of conditions affecting the heart and the blood vessels such as heart attacks,
coronary artery disease among others related to coronary artery disease, continues
to rise in prevalence. This disorder affects the heart and the blood vessels in various
conditions. Thus, for this scenario, proper prediction and early detection of heart
disease become of high importance due to their relevance. Since it allows speedy
therapies, which consequently has the ability to improve patient outcomes and prevent
later complications, prompt identification is important because it allows for rapid
therapies. The public health problem that is heart disease is an important one.
It accounts for the tragic loss of life of a staggering number of people every year.
In fact, the World Health Organization has reported that it is the cause of death on the
global level, causing a staggering 17.9 million fatalities in a year. A research from
the American College of Cardiology [3] came out saying heart disease is accountable
for nearly a third of all fatalities that occurred worldwide in 2019. The report also
accounted that the condition was also responsible for the deaths of about 10 million
men and 9 million women. Moreover, the YLDs from cardiac diseases have risen
between 1990 and 2019, from 17.7 million to 34.4 million. For this reason, there is
a need for general interventions seeking to prevent and control this endemic disease
[4]. Most of the time, the medical doctors are experts in their different fields of
specialty; however, they sometimes face challenges.
ML-based diagnostics might prove useful at times. This study aims at creating
a tool which, based on machine learning, can provide a more accurate diagnosis of
diseases. In the present fast-moving world, most individuals have an affinity and
tendency to give priority to their work and other obligations over their health, which
can become worse in due course. Maybe it is not true that seventy percent of Indians
are taking long-term health problems, and that twenty-five percent of people die
due to ejaculation. The main purpose of this study is to develop high-performance
machine learning prediction models toward accurate prediction of heart disease, and
one of the reasons behind doing this research is to achieve the latter. The primary goal
of this research endeavor is aimed at the development and improvement of machine
learning models for an effective prediction of cardiac disease. We shall begin the
process of laying our baseline by making use of a dataset already existing in the UCI
repository. The main target for us will be to come up with a novel machine learning
model designed with the intent of outperforming the accuracy of the methods that
are currently being adopted. We’ll have a highly detailed set of comparisons and
evaluations in order to draw an assessment about how our model performs compared
to the benchmarks we produce. First and foremost, this project aims at empowering
peoples making easily accessible tools on top of adopting machine learning tech-
niques making use of personal medical data in predicting the potential development
Supervised Machine Learning Solution for Predict Heart Disease 313

of heart disease. Using this instrument does not require an immediate consultation
with a medical professional.

2 Machine Learning

It works efficiently because of the assistance of testing and training methodologies


with machine learning. This approach includes direct training of the system from
data and experiences, which then enables the system to use the information it has
garnered when it is tested against a number of scenarios, as the particular algorithm
requires.
There are the three major categories of the machine learning algorithms exist.
A. Learning
This is the most often used method when categorization is difficult. If the task is
to train a model on an associated labeled dataset, then this method is unsupervised
learning. This process involves matching each individual piece of input data with
the target labels that are connected with it. This particular mode of education is
termed as “supervised learning.” The model can learn how to categorize new images
properly if it is provided with a labeled dataset of images that have associated labels-
for example, cat, dog, bird. This is one way in which the model can learn how to
classify new images.
B. Unsupervised Learning
It is unique, in the sense that the algorithm goes through its paces without there
being any form of direct supervision or input that had been labeled. In this process,
there is neither an instructor nor predetermined labels that may guide the learning
process itself. Instead, the algorithm conducts its self-analyzed nature on the dataset
for discovering patterns and relations among the data points. Based on these recently
developed relationships, the algorithm can classify new input data points that are
unknown in advance and then relate them to one of the clusters or groups it has
discovered. The unsupervised learning paradigm is also known to be “self-complete,”
meaning that the algorithm itself is solely responsible for how it perceives and classi-
fies input patterns. The unsourced learning auto-clustering procedure guarantees the
automatic classification of fruits such as mangoes, bananas, and apples from a given
dataset based on the similarity of the fruits. The locations of the newly obtained fruits
in the groups depend on the similarity between the fruits and the ones already existing
in the groups. Unsourced learning is also known as the process of data categoriza-
tion without receiving direct instruction. This type is helpful for feature grouping or
dimensionality reduction.
C. Reinforcement Learning
One such type of machine learning is reinforcement learning. This is a learning
type in which an agent learns through direct contact with its environment. The agent
314 N. Para et al.

will get good feedback for the acts that are helpful in nature, and bad for actions
that are not favorable in nature. This process is similar to trial and error. This is
an iterative system of feedback, making possible improved agent capabilities in its
decision-making, which subsequently allows it to make better-informed decisions
whenever it faces such events in the future.

3 Related Work

This article by Kavitha and others [1] gives a summarization of many researches that
focus on the prediction and analysis of cardiac disease. It is also a work on research
using the random forest approach for analyzing the Cleveland dataset. Feature selec-
tion techniques are used in both Chi Square and genetic algorithm approaches to reach
an accuracy level in this paper. A second experiment resorted to the application of
the PSO algorithm in a view of creating specific rules. Then the C 5.0 technique was
applied in the context of a binary classification. Now it is worth noting the promise
of the DL and backpropagation neural networks and models, which can allegedly
create predictions almost too good. Besides, different data mining techniques, such
as KNN, decision tree, neural network classification, and Bayesian algorithms have
been used to predict cardiac disease. For feature selection research investigations
based on genetic algorithm have been found quite accurate, especially when used in
combination with the decision tree model.
Archana Singh et al. In their paper [3], the authors Archana Singh and several
colleagues work with a topic dedicated to machine learning algorithm application
for cardiac disease diagnosis. They emphasize the necessity of developing more
accurate prediction systems since they play a significant part in enhancing our
understanding of cardiovascular disease and lowering the number of deaths that
occur. This study makes use of the UCI dataset for both the train and test datasets.
It also compares the performance of several machine learning approaches. These
methods are implemented using the Python programming language and the Anaconda
notebook, Jupyter.
As per Asmit Srivastava and co-authors [4], the attributes of a machine learning
model for cardiac condition prediction train and validate the model employing the
UCI heart prediction standard dataset, which contains fourteen heart-related features.
A predictive model must capture at least the interaction among the components and
engage with LR and RF. It is thereby intended to investigate, in this research study,
how machine learning can be applied in structured data analysis models and predictive
models.
Albert Mayan, in his work, discusses the critical importance of these diseases as
one of the major causes of death and a promising machine learning in the analysis of
medical data [5]. The results of the study conducted by Mayan reveal a new paradigm
of cardiovascular diseases prediction. Features and classification algorithms were
integrated into the solution; its accuracy was 88.7%.
Supervised Machine Learning Solution for Predict Heart Disease 315

Ambrish et al. [6] consider the prediction of cardiovascular disease through the
logistic regression technique applied to the dataset of the University of California,
Irvine. The aim is to create a procedure that is affordable and has the ability for early
detection. That would especially be useful for developing countries or nations with
low incomes. The present model is also found to have significantly higher accuracy
than previous machine learning approaches utilized for similar purposes. For each
age group, females were much more likely than males to be better off but this is not
the case when considering the 20–29-year-old cohort. Instead, men aged between
50 and 59 years reported a higher household income. Results are compared with
previous literature, and although logistic regression is often used as a predictive tool
for the intention of predicting morbidity of cardiovascular disease is rarely reported
in this area. Such research is of great importance because it would help to antedate the
costliest and, more importantly, early diagnosis of cardiac disease, which improves
the quality of life.
In prediction of risk of heart disease using medical datasets, Ambrish et al. and
Guruprasad et al. [9] present their work on a machine learning approach. Their
algorithm does its job by comparing inputs from the user to normal ranges already
known, then grading each variable based on that comparison. Beyond that, the report
includes an extensive literature review wherein it has critically analyzed various
methodologies and approaches that are presently being employed for the purpose of
heart disease forecasting. Of course, with an accuracy of more than 80 percent, it is
very significant that this model was developed.
In her article, Pooja Anbuselvan [10] discusses the application of data science
techniques, specifically data analysis and machine learning (ML), for heart problem
predictions. In this paper, the supervised learning approach with various models is
considered in an attempt to identify which model can be considered the most accu-
rate for heart disease prediction. Results of the study indicate that RF algorithm
produces the best accuracy, which is 86.89%, when compared to other algorithms
evaluated in this study. Significance of Additional Research: This study finds impor-
tance in conducting additional research to address the identified constraints and to
analyze other machine learning methodologies existing for the purpose of further
improvement of the proposed system.
Marbaniang et al. [11] besides the accuracy of the models improve, employ a
tailored dataset and takes into account two of the most critical health risk factors of
blood pressure and body mass index (BMI). His model includes various processes in
the prediction of disease, feature extraction, and data preprocessing. The algorithms
appear to indicate higher accuracy as the modified dataset has been merged with
the KNN classifier, and especially the RF classifier that indicates high precision.
Additionally, the program has shown efficient usage of time while computing in
both phases; that is, the training phase as well as in the testing phase.
Pasha et al. [12] are developing deep learning methods for the CVD predic-
tion. They have established that the standard machine learning methods, namely
SVM, KNN, and DT, are not that efficient for enormous datasets. Using an artifi-
cial neural network (ANN) in TensorFlow Keras attains a good accuracy of 85.24%
316 N. Para et al.

and overcomes the traditional methods for cardiovascular disease prediction in large
datasets.
They have successfully developed a hybrid machine learning approach by using
the crow search algorithm, together with data preparation and feature selection, which
is called HSPUCD. By using it, they are able to obtain efficiency in the proper
identification of feature from medical datasets. Here, for this purpose, they have
used methods like cross-validation and hyperparameter optimization through which
they have successfully predicted issues with a proper accuracy. Compared to the
methods used presently for data analysis, results yield less interference and increased
accuracy. The conclusions of this study are an indication that proper treatment of
cardiac datasets and data cleaning are essential. In a nutshell, this rigorous and
efficient methodology can easily reveal situations with and without the heart.
Apurv Garg and others [14] evaluated the use of machine learning algorithms
KNN and RF for predicting cardiovascular disease in relation to physical attributes.
In the case of RF, where development is carried out on several decision trees for
training, they found the expected accuracy encouraging which is very encouraging.
So, to ensure that the target property representing the presence of heart disease is
well-balanced, I conducted a deep examination against the dataset. For comparison,
the algorithm of the random forest, when used with 10 trees, obtained an accuracy
at 81.967%, whereas KNN obtained an accuracy at 86.885%. Since correctness in
the dataset was key, a count plot analysis was done on the attribute of interest for the
study. Using the public dataset from the University of California repository, Smita
et al.
Abdollahi [15] applied a set of multiple machine learning algorithms with the
aim of making an attempt at prediction for CVD. The findings indicate that accuracy
82.47% with area under the curve at 86.41% is outpacing KNN. The MLP that
it has looks very promising for automated identification of cardiovascular disease
and thus may be useful to diagnose other diseases. Correlation matrix plots were
used by the authors to describe interactions among individual characteristics and
incidence of cardiovascular disease, but it used scatter plots in order to identify
potential contributors to cardiovascular disease. The authors Smita et al. [16] used a
publicly available dataset found through the University of California data repository
to compare and contrast different machine learning algorithms used for the prediction
of cardiovascular disease (CVD).
For a multi-layer perceptron, the accuracy and area under curve were 82.47% and
86.41% respectively which is higher than that of a KNN algorithm. The proposed
MLP is well-suited for the diagnosis of cardiovascular illnesses. It may also be used
for the detection of other diseases. For unearthing probable factors contributing
to cardiovascular disease (CVD), scatter plots were utilized by the researchers.
However, the plots for correlation matrix indicated that some of the characteris-
tics correlated with each other and had a probability of developing CVD, except
Vardhana, etc. Application of ML to the Heart Disease Prediction Problem [17]:
To achieve higher accuracy, they first compare the different classifiers (DT, NB,
LR, SVM, and RF) and then they recommend an ensemble model. They emphasize
Supervised Machine Learning Solution for Predict Heart Disease 317

machine learning for cardiac illness and data preprocessing in the direction of an
attempt toward achieving the best possible output from the model, Indal et al.
A study [18] was performed as if to examine the several approaches like logistic
regression (LR) and K-nearest neighbors about how well these machine learning
methods performed for deciding whether a patient was or not suffering from cardiac
disease. In their work, they have developed a prediction system based on the medical
history of patients that compared to other classifiers belonging to the same class did
not perform better than Naive Bayes or any other related method in terms of accuracy.
The authors used three different data mining approaches: logistic regression (LR), K
nearest neighbor (KNN), and random forest (RF) classifiers. The authors were able
to reach a classification accuracy of 87.5%. The dataset used was from 304 patients,
where every patient had thirteen different medical characteristics. The report contains
detailed information about the investigation, including a collection of observations
from several sources. The research conducted by Rindhe and colleagues [19] focuses
on a strong and reliable approach that will diagnose heart ailments correctly.
That is possible thanks to the application of the machine learning algorithm and
method. To make a prediction of possible heart problems, the researchers use SVM,
RF, and artificial neural networks as their methods of choice. For gathering the
weather data which would be input fed into the model, the need of web scraping is
necessary, and this is also done with Python. These decision tree and random forest
algorithms are applied for the purpose of achieving the intended goal of classification.
In terms of attaining a higher accuracy, the random forest algorithm is better compared
to the decision tree method in terms of the accuracy that it achieves. To arrive at
forecasts regarding cardiac that are accurate in nature, the authors have focused
a great deal of attention toward the utilization of machine learning technologies.
According to Himanshu et al. [20], the Naïve Bayes method is reported to outperform
the K-nearest neighbor algorithm, which is also subject to overfitting, if low variance
is used in combination with strong biasness.
Using low variance and strong biasness, which can be applied for small-sized
datasets, can reduce the time involved in training as well as testing the model. Asymp-
totic errors can also be achieved along with reduced biasness if the size of the dataset
continues to expand. Even if the decision tree suffers from an overfitting problem,
this problem could be handled using some separable overfitting strategies. Using the
Cleveland Heart disease Dataset with a variety of different machine learning and data
mining techniques, Dilip et al. [20] provide a new way for recognizing heart illness.
This method can be used to identify heart illness. The best accuracy level among
these algorithms is by the KNN method, with an accuracy of 87%. The Jupyter
Notebook platform is deployed to determine how effective these algorithms are in
terms of diagnosis. Through the duration of this dataset, there are 303 records; there
also exist six instances where variables are missing. By the use of preprocessing
processes, it is possible to handle the missing variables, which finally result in the
development of a dataset consisting of 297 cases. There are both occurrences that are
positive for heart disease and instances that are negative for cardiac disease within
this dataset. This includes application of support vector machine with random forests
within the scope of the investigation.
318 N. Para et al.

4 Methodology

The first step of the process is a data acquisition retrieved from the UCI repository.
This dataset, which has a long history of use, had already been validated to be correct,
both by the University of California, Irvine (UCI) and other academics. Then, we
make great care in choosing the characteristics that are most pertinent to our special
investigation. After that, we proceed on the next step, which is the implementation of
the machine learning model and then the assessment of correctness for each model
applied (Fig. 1).
A. Data Collection
This dataset is from the repository at the University of California, Irvine; it contains
303 occurrences and 14 attributes. Many experts have verified that this is an accurate
dataset. In this project, we have used thirty percent of the data for testing and seventy
percent for training [6].
B. Attributes Selection
Attributes of the dataset represent the qualities used by the system. Attributes
which some include the heart rate of the individual, gender, and age. We aimed to
develop a model that would predict the presence of cardiac disease near to 1 or its
absence near to 0. This dataset for the study was imbalanced (Fig. 2). There were
165 positive cases of CVD and 138 negative cases. The number of total cases is 138.

Fig. 1 Methodology model


Supervised Machine Learning Solution for Predict Heart Disease 319

Fig. 2 Target

C. Preprocessing of Data

Because the excellence of a machine learning model highly depends on the quality
of the data used in its making, this procedure must be preceded by data preparation.
This initial phase involves a set of stages that are miserably important. One of these
stages is data cleaning, consisting of the elimination of either corrupted or nonexistent
data points and outliers. Further, in this respect, to make the dataset more viable for
model training, procedures of data transformation, resampling, and feature selection
have been adopted.
D. Balancing of Data
Whether a patient has heart disease (close to 1) or does not have heart disease
(near to 0), the goal is if that patient actually has heart disease or not. Figure 2: The
dataset was imbalanced because it diagnosed 165 patients with cardiac disease and
138 patients were put into the normal category.
E. Histogram
Apart from the same code that has been used in constructing the histogram of
attributes represented as “[Link](),” the histogram of attributes provides infor-
mation concerning the range of distribution of attributes making up the dataset
(Fig. 3).
320 N. Para et al.

Fig. 3 Histogram of attributes

5 Machine Learning Algorithms

A. Logistic Regression
For the purpose of classification, one of the forms of classifiers that are used under
supervised learning is known as logistic regression. In most cases, there exist two
classes of a target variable; 0 represents failure, while 1 represents success. This aids
in assessing the possibilities of a target variable (Fig. 4).
Supervised Machine Learning Solution for Predict Heart Disease 321

Fig. 4 Logistic regression

Another term for logistic regression is graphing a sigmoid curve to illustrate


the probability cut-off for classification. Features of Character: This algorithm will
be most appropriate in applications where the outcome is intelligible and can be
classified under different categories. Its main function is the representation of the
relationship between the independent variables and the possibility of the occurrence
of a category of outcome. The sigmoid curve represents the probabilities for being
classified when the cut-off threshold taken at which separation is assumed as 0.5.
B. Decision Tree Classifier
The usage of decision tree classifiers has been quite prominent in the field of
machine learning because these decision trees are user-friendly as well as adaptive.
For achieving the objectives of this tree diagram, divisions are done by using feature
values. In a flowchart, the branching node is a test which has been conducted on a
feature, while the branch explains the inferences which have been drawn from the
tests, and the leaf node contains information about a class label in isolation. The
classifier tree is the first of these types, used in making categorical decisions. The
second type is a regression tree. It is applied for making decisions that are continuous
in nature. To implement your action, you have to divide your dataset into smaller
groups depending on the values of features. In this way, you will obtain a branching
structure of if-then rules. At every stage, it appears that the algorithm is mindful to
decide on the features it will generate (Fig. 5).
C. Support Vector Machine
The support vector machine is all about producing wise decisions as well as
solving problems of regression that may come along in the classification of data.
They come up with the most sensible dividing line they could get after analyzing
your data and differentiating it from one group to another. To be precise in such
circumstances when they encounter novel cases, they make sure to hit the highest
possible separation of the groups. Kernels facilitate SVMs to work effectively on
322 N. Para et al.

Fig. 5 Decision tree

complicated and wavy boundaries, hence they are even smarter. In image recognition
cases, text categorization cases, and complex biological data analysis cases, they
happen to be the best tools around. In the next figure, we illustrate the largest margin
that has occurred between classes, and support vectors over the entire linear decision
boundary (Fig. 6).

Fig. 6 SVM
Supervised Machine Learning Solution for Predict Heart Disease 323

6 Results Analysis

A. Jupyter Notebook

With regard to this research, the Jupyter Notebook is important because it will
facilitate the development of interactive models and study the outcome to predict
cardiac disease. Because it can embed codes, visualizations, and also explanations,
experimenting with and fine-tuning machine learning models, such as SVM, decision
tree, and logistic regression, becomes easier. This is because it makes it easier to test
and improve such models.
B. Result
In this paper, we used four different independent supervised machine learning
classifiers. We used the number of different performance metrics-accuracy, preci-
sion, F1 score, and recall/sensitivity-to classify and choose the best from all those
classifiers that can predict heart disease better.
• Accuracy
The ratio of the number of correct predictions made by a classifier to the total
number of predictions generated by the model is the definition of this statistical
measure. It is possible to articulate it as follows, therefore:

Accuracy = (Correct Predictions)/(Total Predictions). (1)

The sum of true positives (TP) and true negatives (TN) is what constitutes correct
prediction, whereas the sum of false positives (FP) and false negatives (FN) is what
constitutes flawed prediction. Consequently, the equation presented above can be
rewritten as follows:

Accuracy = (TP + TN)/(TP + FP + FN + TN). (2)

• Sensitivity/Recall

It is a measure of the proportion of patients diagnosed with cardiac disease who


are the subject of the model. You may also hear it referred to as recall at times.
Furthermore, both are computed for a model, which is as follows:

Recall = (True Positive)/(True Positive + False Negative). (3)

It is crucial for the treatment of illnesses that a variety of tests have the ability to
reliably identify individuals who exhibit symptoms of illness. Therefore, the propor-
tion of people who test positive for the ailment compared to those who truly have it.
Complete recollection is able to identify every single person who has the illness.
324 N. Para et al.

• Precision

Precision refers to the proportion of those determined by the model to be suffering


from the disease who in fact have heart disease. As a consequence of this, the precision
may appear as follows for it:

Precision = (True Positive)/(True Positive + False Positive). (4)

• F1 Score

An explanation of it is that it is the harmonic mean of sensitivity, which is also


known as recall and precision, and it results in a single numerical value. The following
is an expression of the computation for this:

F1 Score = (2Precision + Recall)/(Precision + Recall). (5)

An extensive number of metrics, such as accuracy, precision, recall, and F1-


score, were evaluated as part of the evaluation process for each approach, as shown
in Table 1.
The best classifiers according to the research included logistic regression, decision
tree classifier, and support vector machine (SVM). The dataset revealed that logistic
regression had the highest accuracy if it was tested out to be at 92.3%. After that
was support vector machines, which attained 88%, and then decision tree classifier,
attaining 82.4%. In summary, the logistic regression and SVM both had about the
same precision and recall scores of about 92%. Conversely, the decision tree classifier
had a precision score of 90.7%, and recall score of 76.5%. “The” This is also a
description of the results that were replicated by the F1-score, which is a measure
that combines precision and recall. In classification, logistic regression performed
very well in all cases, which reflects that it had proven proper for the dataset and
provided a reasonable balance between precision and robustness. Robustness and
precision are two perks that come with classification.

Table 1 Result analysis


Classifier Accuracy (%) Precision (%) Recall (%) F1-score (%)
Logistic regression 92.3 92.3 94.1 93.2
Decision tree classifier 82.4 90.7 76.5 83
Support vector machine 88 85.7 94.1 89.7
Supervised Machine Learning Solution for Predict Heart Disease 325

7 Conclusion

In summary, there are three models: logistic regression, decision tree classifier, and
SVM to prove that it is possible to predict heart disease with reasonable accuracy on
the basis of all primary indications. Overall, the combination of logistic regression,
support vector machines, and decision tree classifiers produced the best result. Even
though all models showed their strengths and weaknesses in terms of accuracy and
completeness, performance differences were still apparent. For heart disease predic-
tion, logistic regression is the best choice for this dataset, inasmuch as it provides a
good balance of accuracy and robustness.

8 Future Work

In the future, we will perform an investigation of the method that is both more
extensive and additional. Key function evaluation through value determination of
estimators for increased accuracy through the use of random forest classifier as
well as working with the linear kernel that was implemented in SVM along with
other SVM classifier kernels for a clear and practical understanding of its operations
focused on improving prediction.

References

1. Kavitha M, Gnaneswar G, Dinesh R, Rohith Sai Y, Sai Suraj R (2021) Heart disease predic-
tion using hybrid machine learning model. In: 2021 6th international conference on inventive
computation technologies (ICICT), Jan 2021, pp 1–5. [Link]
2021.9358597
2. Das RC, Rahman M, Madhab C, Hossen H, Hossain M, Hasanv R (2023) Heart disease detection
using ML. In: 2023 IEEE 13th annual computing and communication workshop and conference
(CCWC), pp 979–984. [Link]
3. Singh A, Kumar R (2020) Heart disease prediction using machine learning algorithms. In:
International conference on electrical and electronics engineering (ICE3)
4. Srivastava A, Singh AK (2022) Heart disease prediction using machine learning. In: 2nd
international conference on advance computing and innovative technologies in engineering
(ICACITE), Apr 2022, pp 1–6. [Link]
5. Praneetha M, Sri Varsha M, Jesudoss A, Mayan A (2021) Cardiovascular disorder prediction
using machine learning. In: 5th international conference on intelligent computing and control
systems (ICICCS), Apr 2021, pp 1–6. [Link]
6. Steinbrunn Janosi WPM, Andras DR (1998) Heart disease. UCI machine learning repository.
[Link]
7. Ambrish G, Ganesh B, Ganesh A, Srinivas C, Dhanraj, Mensinkal K (2022) Logistic regression
technique for prediction of cardiovascular disease. Glob Transit Proc 3:127–130. [Link]
org/10.1016/[Link].2022.04.008
8. Samanvitha RVR, Prasanna SVL, Lokesh B, Sushmanth V (2023) Heart disease prediction
using machine learning algorithms. J Emerg Technol Innov Res 10(4):466–471
326 N. Para et al.

9. Guruprasad S, Mathias VL, Dcunha W (2021) Heart disease prediction using machine learning
techniques. In: 2021 5th international conference on electrical, electronics, communication,
computer technologies and optimization techniques (ICEECCOT). IEEE, pp 762–766. https://
[Link]/10.1109/ICEECCOT52851.2021.9707966
10. Anbuselvan P (2020) Heart disease prediction using machine learning techniques. 23 Int J Eng
Res Technol (IJERT) 9(11):515–518. ISSN: 2278-0181
11. Marbaniang IA, Choudhury NA, Moulik S (2020) Cardiovascular disease (CVD) predic-
tion using machine learning algorithms. In: IEEE 17th India council international conference
(INDICON), pp 2825–2830. [Link]
12. Pasha SN, Ramesh D, Mohmmad S, Harshavardhan A, Shabana (2020) Cardiovascular disease
prediction using deep learning techniques. In: IEEE 17th India council international conference
(INDICON), pp 2825–2830. [Link]
13. Sinha D, Sharma A, Sharma S (2022) A hybrid machine learning approach for heart disease
prediction using hyper parameter optimization. Research Square. [Link]
[Link]-2303591/v1
14. Garg A, Sharma B, Khan R (2021) Heart disease prediction using machine learning techniques.
12 IOP Conf Ser: Mater Sci Eng 1022:01204634. [Link]
012046
15. Pal M, Parija S, Panda G, Dhama K, Mohapatra RK (2022) Risk prediction of cardiovascular
disease using machine learning classifiers. Open Medicine 17:1100–111323. [Link]
10.1515/med-2022-05084
16. Vardhini V, Leela S, Varalakshmi S, Kumar VH, Kumar AS (2023) Heart disease prediction
using machine learning. J Eng Sci 14(04)
17. Jindal H, Agrawal S, Khera R, Jain R, Nagrath P (2021) Heart disease prediction using machine
learning algorithms. IOP Conf Ser: Mater Sci Eng 1022:Article 0120722. [Link]
1088/1757-899X/1022/1/012072
18. Rindhe BU, Ahire N, Patil R, Gagare S, Darade M (2021) Heart disease prediction using
machine learning. Int J Adv Res Sci Commun Technol (IJARSCT) 5(1). [Link]
48175/IJARSCT-1131
19. Sharma H, Rizvi MA (2017) Prediction of heart disease using machine learning algorithms: a
survey recent and innovation trends in computing and communication. IJRITCC
20. Princy RJP, Parthasarathy S, Jose PSH, Lakshminarayanan AR, Jeganathan S et al (2020)
On Prediction of cardiac disease using supervised machine learning algorithms. In: 4th
international conference on intelligent computing and control systems (ICICCS), Madurai,
India
Network and Security Solution Using
Artificial Intelligence-Based Graph
Search Algorithm for Drone
Transportation

Pramod Kumar Patel, Amarjeet Ghosh, Aishwarya Mishra, Rajesh Nema,


Dilip Kumar Gandhi, Neeraj Agarwal, and Ashish Raghuwanshi

Abstract This study reports the suitability of several algorithms to identify the best
course in a network of lines that is utilized to guide an unmanned aerial vehicle
(UAV) aircraft in a safety configuration. It utilizes a waypoint graph for this purpose,
founded on the GPS location of flight route movement locations in three dimensions.
It performs a comparative analysis of the drone path prediction using a dynamic
adjacency matrix that represents the shortest path in the graph. In the perspective of
the spot surveillance mode, evaluate the benefits and drawbacks of the two routing
methods.

Keywords Unmanned aerial vehicle · Algorithms · Networking · GPS node

1 Introduction

UAV or a drone, which is also known as unmanned aerial vehicles, or UAVs, are
becoming more and more common in our daily lives, and this trend is expected to
pick up speed over the next 10 years. In particular, drone use has grown in recent years,
and its uses have expanded to include military applications, last mile delivery, mobile

P. K. Patel (B) · A. Ghosh · R. Nema · A. Raghuwanshi


Department of Electronics and Communication Engineering, IES College of Technology, Bhopal,
India
e-mail: pk_patel05@[Link]
A. Mishra
Department of Computer Science and Engineering, IES College of Technology, Bhopal, India
D. K. Gandhi
Department of Electronics and Communication Engineering, Bansal College of Engineering,
Mandideep, India
N. Agarwal
Department of Mechanical Engineering, IES College of Technology, Bhopal, India

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 327
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
328 P. K. Patel et al.

imaging, and distant surveillance. Global projects aimed to expand the capabilities
and potential uses of drones to new industries are driven by the massive need for UAV-
based solutions. In this instance, among other things, we highlight several current
areas of innovation and study that are becoming more significant: use of drones for
target identification; study includes identifying objects, artificial neural network use,
immediate recognition, and recognizing items in pictures that remain static [1–3].
Further studies are on drone control, different ways to reduce battery usage, and
the development of precise mathematical frameworks [4]. The study is on enhanced
communication, safety, and dependability [5, 6]. Furthermore, studies in fields where
the occurrence of specific natural events, animal migration, fires, etc., are predicted
using memory circuits and artificial intelligence [7, 8].
Drones still need to increase their range using PV solar, which is still constrained
by their comparatively high power consumption, even with the widespread use of
this technology [9, 10]. Since drones are currently primarily powered by recharge-
able batteries, their flight time is often constrained to 15–20 min [9]. Furthermore,
the weather and the extra parts installed on the drones (sensors, camera, and load)
have a significant impact on how long the flight lasts. As a result, the drones must
be as optimized as possible in terms of the shortest flying path criterion [11–15].
Two techniques for determining a graph’s shortest path are covered in Sect. 2. The
experimental setup is illustrated in Sect. 3, and the test results are shown in the Sect. 4.

2 Method

The drone’s exact path through the aircraft must be specified in order to offer the
required security and protection. This path needs to be the best one and satisfy the
following requirements: moving the drones in the quickest amount of time (to use
the least amount of energy), covering the most viewing area, and promptly reshaping
a breakthrough or other infraction. Here, two algorithms are examined to find such
a path: two algorithms have been explored presented to identify such a path: the
most efficient by number of nodes algorithm and the Dijkstra algorithm, used to
determine the smallest route starting with one source to the next location. Also
assess this approach in a configuration that mirrors the actual regions in support of
which have acquired the GPS coordinate (Fig. 1).
The dataset describes the drone’s path as a collection of ocations known as
endpoints (GPS locations). Following the creation of that graph, the weights are
represented by the peaks, which are GPS locations is the separation between two
nearby locations. In order to determine the travel distance, the GPS signals have to
be in a format that corresponds to the elliptical 3-D space. Two algorithms have been
explored presented to identify such a path: the most efficient by number of nodes
algorithm and the Dijkstra algorithm is used to determine the smallest route from
source to next node and final destination. Also assess this approach in a configuration
that mirrors the actual region which required the GPS coordinates shown in Fig. 1.
Network and Security Solution Using Artificial Intelligence-Based … 329

Fig. 1 Process of graph


generation

The following procedure is used to change a point from an ellipsoidal space. At


a location relative to the Cartesian coordinate system, can be represented by (φ, Z,
h): Eq. 1 [9] provides the following expression for point (u, v, w): first, (N + h) = u.
secondly, v = (N + h).cos(φ) = cos(φ).cos(h) and third dimension w is (N. (1-e2) +
h) sin(φ), in which u, v, and w are Cartesian coordinates and φ, w, and h represent the
GPS point’s latitude, longitude, and altitude data. For the prime direction upward, N
is the radius of curve, while e represents eccentricity. Discusses a few constants in
[10].

x = (N + h).cos(ϕ).cos(h)
y = (N + h).cos(ϕ).sin(h) (1)
z = (N.(1 − e2) + h).sin(ϕ),

where φ, Z, h are the latitude, longitude, and altitude coordinates of the GPS point
and u, v, w are Cartesian coordinates. N is the radius of curvature in the prime vertical
and e is eccentricity. Some constants are discussed in [10]. The gap D connecting
the sites is subsequently calculated using the Euclidean distance method (Eq. 2). in
which location 1’s coordinates in Cartesian space are represented by u1 , v1 , and w1 ,
and Point 2’s coordinates are represented by u2 , v2 , and w2 (Fig. 2).
330 P. K. Patel et al.

Fig. 2 Ellipsoidal (φ,Z,h)


and Cartesian (x,y,z)
coordinates [16]


D= (u2 − u1 )2 + (v2 − v1 )2 + (w2 − w1 )2 (2)

where u1 , v2 , w3 are Cartesian coordinates of Point1, and u1 ,v2 , w3 -the coordinates


of Point2.
The dimension that defines the arc is the computed length. The matrix of adja-
cency is a representation of the network. A chart can be displayed using either the
matrix of connections or the adjacency array. The first approach is simpler to imple-
ment and monitor, but it has the drawback of typically being static and including
data concerning every two locations. This implies that an entirely new matrix must
be made whenever the data structure is altered. The additional method offers an
interactive graph transformation but is more complicated.
The first approach is applied in this work, and the dynamic matrix realization
is selected. With N being the number of vertices, it is depicted as a 2-D vector of
dimension N × N. It is thus depicted as a dynamic matrix that is subject to real-time
modification. All that has to be done is recreate the new linkages in the graph if there
are any changes. The graph is then put into the updated matrix. Optimization can be
achieved by making enough minor adjustments to the graph to search for differences
between the graphs and reflect those differences in the matrix.

2.1 Shortest Path Algorithm

The technique for the total amount of points on the fastest route connecting two
points [11]. The foundation of this approach is the application of breadth-first search
(BFS) algorithm. To implement a graph is provided to it the value for G(V, E). Two
vertices, i and j, and a graph, G (V, E), are provided. The minimal path between the
two vertexes will be determined by the minimum number of peaks criteria when the
BFS(i) is executed and we reach the top.
Network and Security Solution Using Artificial Intelligence-Based … 331

Fig. 3 Bhopal area map

2.2 Dijikstra Algorithm

The fundamental requirement is that the arcs f (I, y) have positive values. Information
is given in a grid with n vertices, G (V, E). The smallest path from vertex s to vertices
x is indicated with φ(s.x). Finding the least φ(s,x) + f(i,y) for each y, y = I, is
necessary to determine d(s.x). In this manner, each vertex x from the data structure,
that represents the maximum value for φ(s,x), is assigned an indefinite valued[x]. The
quantity is decreased throughout the method, and once it is finished, d[x] = φ(s,x).
Applying the Dijkstra algorithm, initially create the array’s value d [] by setting d[i]
= MAX_VAL of all vertices i that does not belong to adjacent to any of s, and d[i]
= A[s][i] for every adjacent iϵV of vertexes. When the algorithm d[x] = MAX_
VAL is finished, because when the procedure is finished, d[x] = MAX_VAL since
there is no path between the s and x vertexes. Configure the array’s dimension d[]
as follows: d[x] = A[s][x] for each vertex’s adjacent. D [x] = MAX_VAL for each
vertex x without an adjacent one of s. After the process is complete, d[i] = MAX_
VAL because there is no route that connects s and x vertexes. Secondly, explain
the group T, which at first consists of all graph vertices without s: V\{s} = T. In the
scenario where d[y] < MAX_VAL and T contains a vertex x: find the vertex yϵT with
the lowest d[y]. The Dijkstra method is a useful tool for figuring out the shortest paths
from a given vertex to every other vertex. Unique y or T: T = T \{j} for determine
d[x] = min (d[x], d[y] + A[y] [x] for all x ϵT (Fig. 3).

3 Experimental Setup

The fundamental requirement is that the arcs f (I, y) have positive values. Information
is given in a grid with n vertices, G (V, E). The smallest path from vertex s to vertices
x is indicated with φ(s.x). Finding the least φ(s,x) + f(i,y) for each y, y = I, is
necessary to determine d(s.x). In this manner, each vertex is illustrated in Fig. 3, the
332 P. K. Patel et al.

Table 1 Performance parameters


Characteristics Assessment
Maximum increase/decrease speed 5 m/s; 2 m/s
Maximum flight time (with no wind) 25 min (at a consistent 30 km/h)
Maximum flight distance (no wind) 15 km (at a consistent 55 km/h)
Drone battery 3600 mAh, 1800 mA, 3.7 V

two methods are suggested in a configuration that can be helpful in building a graph
using transport route nodes that will be protected. In order to achieve this, the sites
are chosen using a Google map as actual coordinates (latitude, longitude), and the
height (altitude) is physically established.
For the selection process, the Bhopal Zone is selected. The orange trajectory that
the drones will fly on is connected to the territorial area where the university is
situated by the red line. The drones had to traverse the whole border region using a
total of 15 waypoints that were chosen. Finding the least φ(s, x) + f (i, y) for each
y, y = I, is necessary to determine d(s. x). In this manner, each vertex the drone has
the ability to return and observe again because each location is connected to both
the previous and the next one. This indicates that the graph is an undirected, finite
graph. The drone’s performance assessment according to several variables is shown
in Table 1.
The fundamental requirement is that the arcs f (I, y) have positive values. Infor-
mation is given in a grid with n vertices, G (V, E). The smallest path from vertex s
to vertices x is indicated with φ(s.x). Finding the least φ(s, x) + f(i, y) for each y, y
= I, is necessary to determine d(s.x). In this manner, each vertex the graph is repre-
sented by a 15 × 15 adjacency matrix, and each vertex is a GPS point with the three
coordinates. The element’s value in this matrix is equal to the distance between the
nodes (GPS locations) if there is an arc connecting two vertices; if not, the element’s
value is 0. Since distances are integers that are positive and the Dijkstra algorithm
on the condition operates with positive weights of the arcs, that presenting style was
used. It eliminates the need for further processing and computation by allowing the
algorithm to interact effectively with the matrices.

4 Experimental Results

The graph used for the studies was made up of the path shown in Fig. 3. That is
accomplished by creating a 15 × 15 adjacency matrix using this chart. The begin-
ning vertex is zero, and the ending vertex is eight. Every algorithm was run 1000
times, and the duration of each run was noted. The experimental tests are run on
an Intel Pentium N3710 portable microprocessor. To put it briefly, both algorithms
calculated the same routes between any two vertices. It’s primarily due to the graph’s
simplicity with the fact that every point is situated at the identical elevation. Figure 4a
Network and Security Solution Using Artificial Intelligence-Based … 333

displays the histogram of durations and the rate of occurrence during the Dijkstra
algorithm’s implementation, while Fig. 4b displays the algorithm with the fewest
number of maxima. Table 2 shows the experimental results evaluated through drone
performance.
Algorithm Dijkstra Times frequencies are shown in Fig. 4a. Figure 4b: Algo-
rithm, the number of peaks on the shortest route for two vertices frequency versus
times. Since the Dijkstra algorithm identifies all routes from a single vertex to each
additional one in the network, users can observe that it performs more slowly than
the alternative approach. This makes it possible to preserve details about accessible
roads while having to look for a route again from the same place. The data that is

Fig. 4 a Algorithm Dijkstra Algorithm Dijkstra


Times versus frequencies. Times Vs. frequencies
b Algorithm shortest path
between two vertices by 1000 Times
FREQUNCY

number of peaks times Frequency


frequencies
500

0
1 4 7 10 13 16 19 22 25 28
TIMES
a
Algorithm Shortest Path
Times Vs. frequencies

1000
800
FREQUNCY

600
400
200
0
1 4 7 10 13 16 19 22 25 28
TIMES
b

Table 2 Experiment results


Task performed Average flight time (s) Standard deviation
Take off and climb (25 m) 7.95 0.14
Cruise segment (2500 m) 225 0.12
Drop and landing (25 m) 12.1 0.75
Transferring data (2 points) 90 –
Net flight on square cell 1208 0.551
334 P. K. Patel et al.

Table 3 Drone performance comparisons


Algorithm Drone delivery rate Latency (in hrs)
Power unit Power unit Power unit Power unit
recharge recharge recharge recharge
Square Triangular Square Triangular Square Triangular Square Triangular
Epidemic 0.166 0.209 0.135 0.146 0.81 0.72 2.28 2.13
Spray and 0.211 0.179 0.141 0.156 0.52 0.56 1.75 1.92
wait
PRoPHET 0.594 0.762 0.143 0.319 0.61 0.52 2.28 2.49
Maximum 0.646 0.743 0.135 0.261 0.52 0.47 1.72 1.90
prop
Maximum 0.203 0.271 0.139 0.160 1.08 0.71 2.11 1.80
delivery
TD-Drone 0.954 0.973 0.540 0.664 0.43 0.45 1.69 1.48
Dijkstra

required is thus obtained. Although the approach is with a small number of vertices
is much faster, because only retrieves what is required and only gives details on the
route between two locations. It is necessary to run the algorithm repeatedly whenever
another route needs to be found (Table 3).

5 Conclusion

Several shortest path identification methods that were deemed optimal for the appli-
cation we developed were examined in this study. It reports that both approaches
identify the graph’s optimal path determined on the outcome of the experiment. It
provides a summary of some findings regarding the benefits and drawbacks of the two
approaches in relation to our implementation case below. Advantages of the Dijkstra
algorithm include the best method for locating a route in a graph has the ability to
retain and analyze data by identifying every route to the other points and exceptional
suitability to be utilized with unmanned aircraft, as the shortest path actually repre-
sents the shortest route across the two locations, that is crucial for the idea of least
path minimum power usage. However it is a slower approach, the method’s speed
decreases if the data structure has too many changes (unobtainable vertex, route
alterations, inserting new vertexes, etc.). The algorithm quantity of spikes on the
shortest route connecting two vertices has benefits of quick algorithm, appropriate
for a graph that changes dynamically, however it only provides data regarding the
route between two vertexes; generally could lead to the same route as the Dijkstra
algorithm.
Network and Security Solution Using Artificial Intelligence-Based … 335

Acknowledgements Funding: The M.P. Council of Science and Technology, Bhopal has provided
financial support for this project work (Grant No. 3844/CST/R&D/Phy. & Engg. and Pharmacy/
2022–2023).

References

1. Lee J, Wang J, Crandall D, Sabanovi´ c Sˇ, Fox G (2017) Real-time object detection for
unmanned aerial vehicles based on cloud-based convolution neural networks. In: 2017 first
IEEE international conference on robotic computing (IRC), 10–12 April
2. Thiang IN, Lu Maw HMT (2016) Vision-based object tracking algorithm with AR. Drone. Int
J Sci Technology Res 5(06)
3. Patel PK et al (2023) A mini review on MEMS switches: design, fabrication, and applications.
In: 2023 3rd international conference on energy, power and electrical engineering (EPEE),
Wuhan, China, pp 96–103
4. Kryjak T, Komorkiewicz M, Gorgon M (2014) Real-time implementation of foreground object
detection from a moving camera using the ViBE algorithm. Comput Sci Inf Syst 11(4):1617–
1637
5. Yadav S, Patel P (2015) Constrained parameter analysis of Microstrip antenna. In: 2015 inter-
national conference on soft-computing and networks security (ICSNS), Coimbatore, India, pp
1–5
6. Huang Q, Mathematical modeling of quad copter dynamics. Rose-Hulman Institute of
Technology. [Link]
7. Patel PK, Malik M, Gupta TK (2022) Optimization techniques for reliable low leakage
GNRFET-based 9T SRAM. IEEE Trans Device Mater Reliab 22(4):506–516
8. Narkhedkar SG, Patel PK (2014) Recipe of speech compression using coiflet wavelet. In: 2014
international conference on contemporary computing and informatics (IC3I), Mysore, India,
pp 1135–1139
9. Patel PK et al (2023) Efficiency evaluation of solar PV array with improved performance. In:
2023 3rd international conference on energy, power and electrical engineering (EPEE), Wuhan,
China, pp 64–69
10. Ngoc G, Moon K-S, Lee S-H, Kwon K-R (2016) GIS map encryption algorithm for drone
security based on geographical features. In: 2016 international conference on computational
science and computational intelligence (CSCI), pp 1422–1423
11. Patel PK, Agrawal N, Raghuwanshi A, Nema RK (2024) Power and stability optimization of
high performance SRAM cell using jaya algorithms for low power drone. In: 2024 IEEE 3rd
international conference on electrical power and energy systems (ICEPES), Bhopal, India, pp
1–6
12. Akram R, Markantonakis K, Mayes K, Habachi O, Sauveron D, Steyven A, Chaumette S (2017)
Security, privacy and safety evaluation of dynamic and static fleets of drones. In: 2017 IEEE/
AIAA 36th digital avionics systems conference (DASC), pp 1–12
13. Carlon R (2015) Tracking tagged fish using a wave glider. OCEANS2015-MTS/IEEE
Washington, pp 1–5
14. Patel PK, Nema RK, Bhadhouriya P (2024) Drones based on multi-detection and mobilize
algorithm for observation and routing. In: 2024 IEEE international conference on big data and
machine learning (ICBDML), Bhopal, India, pp 109–114
15. Bhambu P, Kumar S (2016) levy flight based animal migration optimization algorithm. In:
2016 international conference on recent advances and innovations in engineering (ICRAIE),
pp 1–5
16. Yu R, Liu Y, Meng Y, Guo Y, Xiong Z, Jiang P (2024) Optimal configuration of heterogeneous
swarm for cooperative detection with minimum DOP based on nested cones. Drones 8(1):11
Artificial Bee Colony Algorithm
for Efficient Hyperparameter Tuning
in Alzheimer’s Disease Classification

Raghubir Singh Salaria and Neeraj Mohan

Abstract Alzheimer’s disease (AD) is a progressive neurodegenerative disorder


that necessitates early and accurate diagnosis for effective intervention. Traditional
diagnostic methods often fall short due to subjective interpretation and limited scal-
ability. Deep neural networks (DNNs) offer a promising alternative for automated
AD classification, yet the success of these models heavily depends on precise hyper-
parameter tuning. This paper presents a novel approach that leverages the Artificial
Bee Colony (ABC) algorithm to optimize the hyperparameters of a DNN model
tailored for Alzheimer’s disease classification using MRI images. The ABC algo-
rithm, inspired by the foraging behavior of honeybees, effectively balances explo-
ration and exploitation in the hyperparameter space, dynamically adjusting param-
eters such as learning rate, batch size, and layer configuration to achieve optimal
performance. The Alzheimer MRI Pre-processed Dataset, which includes 6400
images categorized into four classes based on dementia severity (non-demented,
very mild demented, mild demented, and moderate demented), was used to evaluate
the proposed model. The ABC-tuned DNN achieved a classification accuracy of
92.3%, significantly outperforming models tuned with grid search (88.5%), random
search (87.0%), and particle swarm optimization (PSO) (90.1%). Furthermore, the
ABC algorithm enhanced the model’s computational efficiency, reducing conver-
gence time to 70 epochs compared to 100 and 120 epochs for grid search and random
search, respectively. Detailed performance analysis across the classes demonstrated
high average precision (90.3%), recall (88.3%), and F1-score (89.2%), with espe-
cially strong results in distinguishing early stages of AD, such as non-demented
and very mild demented. This robust classification performance highlights the ABC
algorithm’s effectiveness in optimizing complex hyperparameter spaces in medical
imaging models, contributing to more accurate and efficient AD classification. The
ABC-tuned DNN model’s superior accuracy and convergence time make it a viable
tool for clinical applications where timely and precise AD diagnosis is essential.

R. S. Salaria (B) · N. Mohan


IKG Punjab Technical University, Kapurthala, Punjab, India
e-mail: rssalaria@[Link]

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 337
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
338 R. S. Salaria and N. Mohan

Keywords Alzheimer’s disease · Deep neural networks · Artificial bee colony ·


Particle swarm optimization · Hyperparameter tuning · Ant colony optimization

1 Introduction

The classification of Alzheimer’s disease (AD) using deep neural networks (DNNs)
has emerged as a promising area in medical diagnostics due to its potential for early
and accurate detection. Alzheimer’s disease, a progressive neurodegenerative condi-
tion, is characterized by memory loss, cognitive decline, and changes in behavior.
Early diagnosis is critical, as it enables timely intervention, potentially slowing
disease progression and improving patient quality of life. Traditional diagnostic
methods often rely on subjective evaluation, making it challenging to consistently
identify early stages. In this regard, machine learning and deep learning models,
particularly DNNs, have been explored extensively to automate and enhance diag-
nostic accuracy. However, the efficacy of a DNN model in AD classification depends
heavily on the proper configuration of its hyperparameters, such as learning rate,
batch size, number of layers, and neuron count per layer. Hyperparameter tuning,
a process that aims to optimize these settings, plays a pivotal role in enhancing
the performance of DNNs by reducing overfitting, improving generalization, and
ensuring efficient use of computational resources [1–3]. Properly tuned models
can achieve higher accuracy and robustness in classification tasks, especially when
handling complex medical data, such as Magnetic Resonance Imaging (MRI) scans
or Positron Emission Tomography (PET) images. Hyperparameter tuning methods
vary, ranging from traditional approaches like grid search and random search to more
advanced optimization techniques. Swarm intelligence (SI) has gained attention as
a robust and efficient method for hyperparameter optimization in DNNs for AD
classification. SI is a collective behavior observed in decentralized, self-organized
systems, such as ant colonies or bird flocks, which work collectively to find optimal
solutions. In the context of hyperparameter tuning, SI algorithms, including Particle
Swarm Optimization (PSO) and Ant Colony Optimization (ACO), have shown signif-
icant promise in navigating the high-dimensional and complex hyperparameter space
[4–6]. Unlike traditional methods, SI-based optimization dynamically adapts to the
hyperparameter landscape, improving the likelihood of finding globally optimal
solutions and avoiding local optima traps.
PSO, one of the most widely used SI algorithms, has been particularly effective in
tuning DNN hyperparameters for AD classification tasks. This algorithm simulates
the social behavior of particles (agents) in search of optimal positions within the
hyperparameter space. Each particle adjusts its trajectory based on both its experi-
ence and that of neighboring particles, gradually converging toward the best solution.
PSO’s flexibility and adaptability make it highly suitable for hyperparameter tuning,
especially in the case of nonlinear, high-dimensional search spaces typical of DNN
configurations [7]. Furthermore, swarm-based tuning methods enhance model reli-
ability and prediction accuracy by continuously refining parameter values during
Artificial Bee Colony Algorithm for Efficient Hyperparameter Tuning … 339

the training process, reducing the chances of overfitting and model instability. By
leveraging swarm intelligence, researchers have reported improvements in diag-
nostic classification accuracy, robustness, and computational efficiency, making these
approaches ideal for real-time clinical applications [8]. The collective and adaptive
nature of SI methods also facilitates multi-objective optimization, a key advantage
in complex medical domains where accuracy, speed, and resource usage are equally
crucial.
This paper contributes to the advancement of Alzheimer’s disease (AD) classifi-
cation by focusing on the optimization of hyperparameters in deep neural networks
(DNNs) through the application of Artificial Bee Colony (ABC) algorithms. The
primary contributions are as follows:
(a) ABC-Based Hyperparameter Optimization: We present a framework that
leverages the ABC algorithm to automatically and effectively tune the hyperpa-
rameters of a DNN model for Alzheimer’s disease classification. By simulating
the foraging behavior of honeybees, the ABC algorithm dynamically searches
for optimal hyperparameter configurations, leading to improved classification
accuracy and stability.
(b) Comparison with State-of-the-Art Techniques: The proposed ABC-tuned
DNN is systematically compared with other hyperparameter optimization
methods, including particle swarm optimization (PSO) and grid search. Experi-
mental results demonstrate that ABC provides superior results, achieving higher
accuracy, faster convergence, and lower computational costs in AD classification
tasks.
(c) Performance Evaluation on Medical Imaging Data: We validate the proposed
model on medical imaging data, specifically MRI and PET scans, commonly
used in Alzheimer’s diagnosis. The evaluation highlights ABC’s ability to
enhance the DNN’s predictive accuracy and robustness across diverse AD
stages, making it a valuable tool for clinical applications.
(d) Insights on Hyperparameter Sensitivity: This study offers insights into the
sensitivity of DNN hyperparameters, showcasing how different configurations
impact model performance in AD classification. This understanding is crucial
for developing efficient, adaptable models in healthcare diagnostics.

2 Related Work

Kaur et al. (2020) explored hyperparameter optimization techniques in deep learning


models, applying grid search to improve Alzheimer’s classification accuracy by
tuning parameters like learning rate and dropout rate [9]. This study highlighted
the challenges of exhaustive search techniques in high-dimensional spaces, where
computational costs are prohibitive, motivating the need for more efficient methods
such as swarm-based algorithms. Arafa et al. (2024) proposed a framework using
transfer learning and fine-tuning strategies for AD diagnosis on MRI images,
targeting early stages of the disease [10]. Their approach emphasized the benefits
340 R. S. Salaria and N. Mohan

of pre-trained models and transfer learning to reduce training time, though they did
not use swarm algorithms directly. This work laid foundational insights for adopting
hybrid models, combining transfer learning with optimization techniques like swarm
intelligence for enhanced model robustness. Abbas et al. (2023) introduced CAD-
ALZ, a classification system combining convolutional neural networks (CNNs) with
random forest classifiers optimized using a blockwise fine-tuning approach [11]. This
model achieves high accuracy by adjusting CNN layers and hyperparameters, demon-
strating how systematic hyperparameter tuning enhances performance. Although
CAD-ALZ does not directly utilize swarm intelligence, it provides insights into the
benefits of structured parameter adjustments for AD diagnosis. Fouladi et al. (2022)
implemented a deep neural network (DNN) framework optimized for Alzheimer’s
disease and mild cognitive impairment classification using EEG recordings [12].
The authors highlighted the benefits of systematic hyperparameter adjustment in
improving classification accuracy, and the study provides a baseline for integrating
swarm algorithms into hyperparameter tuning for EEG-based models.
Kaya et al. (2024) applied particle swarm optimization (PSO) for multiclass AD
classification using MRI data, demonstrating how PSO can optimize DNN param-
eters effectively [13]. The PSO algorithm dynamically explores the hyperparam-
eter space, yielding higher accuracy and robustness compared to static methods.
This study underscores PSO’s suitability for high-dimensional medical data and
serves as a key reference for swarm intelligence in AD classification. Alzheimer’s
Disease Neuroimaging Initiative (2018) conducted a comprehensive study on deep
learning applications in AD classification, focusing on early detection in mild cogni-
tive impairment (MCI) subjects using a combination of random forest and DNN
classifiers [14]. While the study used traditional optimization techniques, it under-
scores the importance of feature selection and hyperparameter tuning, motivating the
use of swarm intelligence to further enhance model performance. Palaniswamy et al.
(2022) explored convolutional neural networks (CNNs) for medical image classifi-
cation, using an optimization-based deep learning model that demonstrated improve-
ments in classification precision through hyperparameter tuning [15]. The authors
employed a simplified approach, but their work supports the potential of swarm
intelligence for improved efficiency and convergence in hyperparameter optimiza-
tion tasks. Bloch and Friedrich (2020) applied Bayesian optimization to fine-tune
parameters in random forest and XGBoost models for early AD diagnosis [16]. Their
approach effectively demonstrates how Bayesian and swarm-based optimizations can
yield significant improvements in classification accuracy, particularly when applied
to diverse datasets in early-stage AD detection. Sharma et al. (2022) proposed an
optimized DNN for AD classification using structural MRI data, focusing on global
hyperparameters like learning rate and batch size [17]. This study highlights the crit-
ical role of tuning these parameters to improve model convergence and robustness,
presenting a promising base for further work using swarm algorithms for hyperpa-
rameter selection. Ibrahim et al. (2023) employed particle swarm optimization (PSO)
for DNN tuning, focusing on brain tumor and AD detection [18]. PSO’s adaptability
and efficiency were noted to enhance DNN performance, particularly in optimizing
Artificial Bee Colony Algorithm for Efficient Hyperparameter Tuning … 341

the model for complex medical imaging data, showing swarm intelligence’s effective-
ness in this domain. Du et al. (2020) proposed a hybrid swarm intelligence approach
for hyperparameter tuning in DNNs, combining the strengths of PSO and genetic
algorithm (GA) to improve model robustness and accuracy [19]. The study’s find-
ings demonstrate that hybrid swarm approaches can outperform single-algorithm
methods, especially in complex datasets, highlighting an effective framework for
medical diagnostics.
Shaaban et al. (2021) introduced a model combining PSO and GA for hyperpa-
rameter optimization in deep learning applications [20]. This method demonstrated
improved convergence rates and classification accuracy over traditional methods,
suggesting that hybrid swarm algorithms offer practical advantages in hyperparam-
eter tuning for medical image classification. Mukherjee et al. (2020) utilized the
Artificial Bee Colony (ABC) algorithm for DNN tuning in image classification tasks,
showing that ABC achieves comparable or superior accuracy relative to other opti-
mization methods [21]. The adaptive nature of ABC was noted for handling high-
dimensional parameter spaces, which are common in medical diagnostics and AD
classification. Zhang et al. (2021) applied Ant Colony Optimization (ACO) to opti-
mize CNN hyperparameters for high-precision image classification [22]. ACO was
found to be efficient in navigating complex parameter spaces, achieving a balance
between exploration and exploitation. This study validates ACO’s potential as an
alternative to PSO and ABC in AD classification tasks.

3 Proposed Work

This research paper presents a deep neural network (DNN) framework for
Alzheimer’s disease (AD) classification, optimized through hyperparameter tuning
using the Artificial Bee Colony (ABC) algorithm. The proposed approach aims to
address the inherent challenges of optimizing high-dimensional and complex param-
eter spaces in DNN models tailored for medical diagnostics. The ABC algorithm,
inspired by the foraging behavior of honeybees, dynamically explores and exploits the
hyperparameter space, seeking optimal combinations of parameters such as learning
rate, batch size, and neuron count. This optimization is critical for achieving enhanced
model accuracy, faster convergence, and improved generalization, especially when
applied to complex medical imaging data like MRI and PET scans.
In contrast to traditional tuning methods, which may suffer from inefficiency (e.g.,
grid search) or susceptibility to local optima (e.g., random search), ABC provides
a balanced mechanism that adaptively searches the hyperparameter space based on
real-time feedback, facilitating both exploration and exploitation. Through the ABC
algorithm’s adaptive and decentralized nature, this paper proposes an optimized DNN
framework for AD classification that enhances classification performance, reduces
computational requirements, and offers a robust solution for early and precise AD
diagnosis.
342 R. S. Salaria and N. Mohan

First, a population of food sources is initialized, with each food source representing
a unique set of hyperparameters for a deep neural network (DNN) model. These initial
solutions are generated randomly within the defined hyperparameter search space.
The algorithm then enters an iterative process with three main phases: the employed
bee phase, the onlooker bee phase, and the scout bee phase. In the employed bee
phase, each employed bee generates a neighboring solution to its current hyper-
parameter configuration by a small random adjustment, thereby exploring nearby
configurations in the hyperparameter space. This adjustment, defined by a random
factor, is based on the difference between the current solution and another randomly
chosen solution. The fitness of the new configuration is evaluated by training the
DNN and assessing its performance. If the new configuration has better fitness, it
replaces the current solution; otherwise, a trial counter for the food source is incre-
mented, indicating that it did not yield an improvement. Next, in the onlooker bee
phase, bees probabilistically select food sources based on their relative fitness values
and generate new neighboring solutions around the selected food sources. This phase
focuses exploration on the most promising areas of the hyperparameter space, as fitter
configurations are more likely to be chosen. The fitness of each generated solution
is evaluated, and replacements occur if improvements are observed. If a food source
has not improved after a predefined number of trials (limit), the scout bee phase is
triggered, where scout bees replace the stagnating food source with a new random
solution, encouraging exploration of diverse areas in the hyperparameter space and
avoiding local optima. This adaptive process, iterating through these phases, leads to
the convergence of the algorithm toward an optimal set of hyperparameters, which
is ultimately returned as the best configuration for the DNN model’s performance.

4 Results and Discussion

For the evaluation of the proposed work, the following dataset is used.
The Alzheimer MRI Pre-processed Dataset is a well-curated collection of
Magnetic Resonance Imaging (MRI) scans designed for machine learning appli-
cations in Alzheimer’s disease (AD) diagnosis and staging. This dataset aggregates
6400 MRI scans from multiple sources, including repositories like UCI Machine
Learning and the National Library of Medicine (NLM) [10, 14–16], to provide a
comprehensive view of AD progression. Each image is standardized to 128 × 128
pixels, ensuring uniformity and consistency in model training. The dataset is divided
into four classes representing different stages of dementia severity:
• Mild Demented: 896 images, indicating early-stage Alzheimer’s signs.
• Moderate Demented: 64 images, showing progression from mild to moderate
dementia.
• Non-demented: 3200 images, serving as the control group without dementia.
• Very Mild Demented: 2240 images, representing the onset phase of Alzheimer’s.
Artificial Bee Colony Algorithm for Efficient Hyperparameter Tuning … 343

This dataset serves as a foundational resource for the development and validation
of machine learning models aimed at early detection, classification, and diagnosis of
Alzheimer’s disease.
The results of the proposed work are presented in two tables. The first table summa-
rizes the performance of the ABC-tuned DNN model in classifying Alzheimer’s
disease stages across four categories: Mild Demented, Moderate Demented, Non-
demented, and Very Mild Demented. This table includes key metrics such as preci-
sion, recall, F1-score, and misclassification rate for each class, enabling a detailed
analysis of the model’s effectiveness in distinguishing between different stages of
dementia (Figs. 1 and 2).

Performance Analysis

Very Mild Demented

Non-Demented

Moderate Demented

Mild Demented

0 10 20 30 40 50 60 70 80 90 100

Misclassification Rate (%) F1-Score (%) Recall (%) Precision (%)

Fig. 1 Classification of results

Comparitive Analysis
140
120
100
80
60
40
20
0
Accuracy (%) Convergence Precision (Avg Recall %)
(Avg %) F1-Score (Avg Time (epochs)
)
ABC-tuned DNN (Proposed)
search Grid Search Random pso-tuned DNN

Fig. 2 Comparative analysis


344 R. S. Salaria and N. Mohan

Following this, the second table compares the proposed ABC-tuned DNN model
with other hyperparameter tuning approaches, specifically grid search, random
search, and a particle swarm optimization (PSO)-tuned DNN model. The compar-
ison focuses on overall accuracy, convergence time, and average precision, recall,
and F1-score across all classes (Tables 1 and 2).
The proposed ABC-tuned DNN model achieved an accuracy of 92.3%, outper-
forming both grid search (88.5%), and random search (87.0%). Additionally, the
ABC algorithm significantly reduced convergence time, requiring only 70 epochs
compared to 100 for grid search and 120 for random search. The PSO-tuned DNN,
while competitive with an accuracy of 90.1%, was still slightly outperformed by the
ABC-tuned model in terms of both accuracy and efficiency. The ABC-tuned model
also excelled in precision, recall, and F1-score averages, showing 90.3%, 88.3%, and
89.2%, respectively. In contrast, PSO achieved slightly lower average scores, with
precision at 88.0% and F1-score at 87.5%. The reduced convergence time and supe-
rior classification performance of the ABC-tuned DNN illustrate the efficacy of the
ABC algorithm in efficiently tuning hyperparameters, providing both high accuracy
and computational efficiency, which is essential for real-time clinical applications in
Alzheimer’s disease diagnosis.

Table 1 Classification details of proposed work


Class Sample size Precision (%) Recall (%) F1-score (%) Misclassification
rate (%)
Mild demented 896 91 89 90 9
Moderate 64 85 78 81 22
demented
Non-demented 3200 93 95 94 5
Very mild 2240 92 91 91 9
demented

Table 2 Comparative analysis


Approach Accuracy (%) Convergence Precision Recall (Avg F1-score (Avg
time (epochs) (Avg %) %) %)
ABC-tuned 92.3 70 90.3 88.3 89.2
DNN
(Proposed)
Grid search 88.5 100 86.5 84.8 85.4
Random 87 120 85.1 83.5 84
search
PSO-tuned 90.1 85 88 87 87.5
DNN
Artificial Bee Colony Algorithm for Efficient Hyperparameter Tuning … 345

5 Conclusion

This paper introduced an Artificial Bee Colony (ABC)-tuned deep neural network
(DNN) framework for classifying Alzheimer’s disease (AD) stages. By optimizing
hyperparameters using the ABC algorithm, the model achieved a high classifica-
tion accuracy of 92.3%, outperforming grid search (88.5%), random search (87.0%),
and PSO-tuned models (90.1%). The ABC-tuned model also demonstrated faster
convergence, requiring only 70 epochs compared to 100 and 120 epochs for grid and
random search, respectively, thus enhancing computational efficiency. The model
showed strong performance across classes, particularly in distinguishing early AD
stages. Average precision, recall, and F1-scores of 90.3%, 88.3%, and 89.2%, respec-
tively underscore its robustness. These results highlight ABC’s potential in improving
DNN performance for medical imaging tasks. Future work could further enhance this
approach by exploring hybrid algorithms and advanced architectures, extending its
applicability to broader medical diagnostic fields.

References

1. Oommen DK, Arunnehr J (2023) Comprehensive analysis of machine learning algorithms to


detect Alzheimer’s disease using predictor factors. J Theor Appl Inf Technol 101(10):37–48
2. Oommen DK, Arunnehru J (2023) Alzheimer’s disease stage classification using a deep transfer
learning and sparse auto encoder method. Comput Mater Contin 76(1):123–134
3. Topsakal O, Lenkala S (2024) Enhancing Alzheimer’s disease detection through ensemble
learning of fine-tuned pre-trained neural networks. Electronics 13(17):3452
4. Ismail WN, PP FR, Ali MAS (2023) A meta-heuristic multi-objective optimization method for
Alzheimer’s disease detection based on multi-modal data. Mathematics 11(4):957
5. Yaqoob N, Khan MA, Masood S (2024) Prediction of Alzheimer’s disease stages based on
ResNet-self-attention architecture with Bayesian optimization and best features selection. Front
Comput Neurosci 14(8):1393849
6. Sabahat N, Akhtar A, Minhas S (2022) A deep longitudinal model for mild cognitive impairment
to Alzheimer’s disease conversion prediction in low-income countries. Wiley Online Library
45(2):1419310
7. Shamrat FMJM, Akter S, Azam S (2023) AlzheimerNet: an effective deep learning based propo-
sition for Alzheimer’s disease stages classification from functional brain changes in magnetic
resonance images. IEEE Access 13:2789–2802
8. Odusami M, Maskeliūnas R, Damaševičius R (2023) Explainable deep-learning-based diag-
nosis of Alzheimer’s disease using multimodal input fusion of PET and MRI Images. J Med
Biol Eng 45:567–579
9. Kaur S, Aggarwal H, Rani R (2020) Hyper-parameter optimization of deep learning model for
prediction of Parkinson’s disease. Mach Vis Appl 31(5):1013–1025
10. Arafa DA, Moustafa HED, Ali HA (2024) A deep learning framework for early diagnosis of
Alzheimer’s disease on MRI images. Multimed Tools Appl 77:1234–1248
11. Abbas Q, Hussain A, Baig AR (2023) CAD-ALZ: a blockwise fine-tuning strategy on convo-
lutional model and random forest classifier for recognition of multistage Alzheimer’s disease.
Diagnostics 13(1):167
12. Fouladi S, Safaei AA, Ghaderi F (2022) Efficient deep neural networks for classification of
Alzheimer’s disease and mild cognitive impairment from scalp EEG recordings. Cogn Comput
18:456–465
346 R. S. Salaria and N. Mohan

13. Kaya M, Çetin-Kaya Y (2018) A novel deep learning architecture optimization for multiclass
classification of Alzheimer’s disease level. IEEE Access 18:11732–11742. Alzheimer’s disease
neuroimaging initiative. Deep learning reveals Alzheimer’s disease onset in MCI subjects:
results from an international challenge. J Neurosci Methods 14(4):1725–1736
14. Initiative ADN (2018) Deep learning reveals Alzheimer’s disease onset in MCI subjects: results
from an international challenge. J Neurosci Methods 14(4):1725–1736
15. Palaniswamy T (2022) Hyperparameter optimization based deep convolution neural network
model for automated bone age assessment and classification. Displays 23:115–126
16. Bloch L, Friedrich CM (2020) Using Bayesian optimization to effectively tune random forest
and xgboost hyperparameters for early Alzheimer’s disease diagnosis. Eur Alliance Innov
23:175–185
17. Sharma R, Goel T, Murugan R (2022) An optimized deep learning network for prognosis
of Alzheimer’s disease using structural magnetic resonance imaging. In: IEEE Region 10
humanitarian technology conference, vol 12, pp 385–395
18. Ibrahim R, Ghnemat R, Abu Al-Haija Q (2023) Improving Alzheimer’s disease and brain tumor
detection using deep learning with particle swarm optimization. AI 4(3):30–40
19. Du C, Lu Z, Li Y (2020) A hybrid swarm intelligence-based approach for hyperparameter
optimization in deep learning models. Soft Comput 24(8):5657–5666
20. Shaabana AL, Fouad MM, Ahmed FH (2021) Hybrid particle swarm optimization and genetic
algorithm for hyperparameter tuning of deep learning models. Expert Syst Appl 144:113128
21. Mukherjee S, Chaudhuri B, Nag S (2020) Bee colony optimization for efficient hyperparameter
selection in deep neural networks for image classification. Comput Electr Eng 81:106524
22. Zhang X, Ma Y, Wang Q (2021) Ant colony optimization-based hyperparameter tuning of
convolutional neural networks for high precision image classification. Pattern Recogn Lett
144:10–16
Machine Learning Approaches
for Anomaly Detection in IoT:
A Comprehensive Study

Rajesh Rajaan , Baldev Singh, and Nilam Choudhary

Abstract The Internet of Things (IoT) encompasses a vast array of smart devices that
can collect, store, process, and transmit data. While IoT adoption has fueled signifi-
cant innovation across industries, homes, environments, and businesses, its inherent
vulnerabilities have raised concerns about its widespread deployment. Unlike tradi-
tional IT systems, securing the IoT is particularly challenging due to resource limita-
tions, device heterogeneity, and the distributed nature of the network. Implementing
host-based protection mechanisms, including antivirus and anti-malware software, is
not feasible due to these issues. A monitoring strategy like anomaly detection, both at
the device and network levels, becomes crucial beyond the organizational perimeter
in light of these difficulties and the particulars of IoT applications. As such, anomaly
detection systems are well-positioned to safeguard IoT devices more effectively than
conventional security methods. In this paper, the authors highlight a comprehen-
sive review of existing efforts to develop machine learning-based anomaly detection
solutions for IoT security and explore how blockchain-integrated anomaly detec-
tion systems can collaboratively train machine learning models to enhance anomaly
detection capabilities.

Keywords Smart devices · Internet of things (IoT) · Data collection · Data


processing · Vulnerabilities · Anomaly detection · Machine learning

1 Introduction

The Internet of Things (IoT) includes numerous smart devices that are capable of
collecting, storing, processing, and sharing data. IoT adoption has driven consid-
erable advancements in industries, homes, environmental monitoring, and business

R. Rajaan (B) · B. Singh


Vivekananda Global University, Jaipur, India
e-mail: raaj0028@[Link]
N. Choudhary
Swami Keshvanand Institute of Technology, Management & Gramothan, Jaipur, India

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 347
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
348 R. Rajaan et al.

operations, enhancing quality of life, productivity, and profitability. However, IoT-


related infrastructures, applications, and services also introduce significant secu-
rity risks and vulnerabilities, as emerging protocols and operations have consider-
ably expanded the potential attack surface [1]. For example, the Mirai botnet attack
leveraged IoT vulnerabilities to disrupt various websites and domain name systems
[2].
Securing IoT devices is challenging due to their heterogeneity, resource limita-
tions, and the decentralized nature of IoT networks, which are beyond the protection
scope of traditional perimeter-based security solutions. Existing solutions, such as
cloud-based security, suffer from centralization and latency issues. Furthermore, IoT
device manufacturers often prioritize rapid product deployment over robust secu-
rity, contributing to the lack of established security standards, which increases the
complexity of securing IoT environments. Beyond the boundaries of an organization,
monitoring tools like anomaly detection are crucial at the device and network levels
due to these difficulties and the distinct architecture of IoT applications.
An anomaly in IoT systems is defined as a pattern or sequence of patterns that
significantly deviates from typical behavior. Anomalies may be classified as point,
contextual, or collective, depending on the source [3]. A point anomaly indicates an
isolated data point deviating from normalcy, often seen as an outlier. A contextual
anomaly is one that is abnormal in a particular context, such as within a specific
time period, where the same data may appear normal in another setting. Contex-
tual anomalies are influenced by factors like time, space, or the specific application
domain. Collective anomalies involve groups of interconnected or sequential data
points, where the entire sequence is abnormal, though individual data points may
not appear anomalous. Although rare, anomalous events can have severe negative
consequences for businesses and governments relying on IoT applications [4].
To secure IoT and other information technology (IT) applications, intrusion detec-
tion systems (IDSs) have been developed to alert against unusual activities or poten-
tial attacks. IDSs fall into two primary categories: anomaly-based and signature-
based [5]. Anomaly-based IDSs can detect unknown attacks, including zero-day
attacks, by identifying deviations from normal behavior, while signature-based IDSs
require updated attack signatures, making them less effective against new threats.
Therefore, anomaly-based IDSs are well-suited for IoT security. Due to the large
volume of raw data generated by IoT devices, detecting suspicious activity can
be computationally expensive and prone to noise, making lightweight, distributed
anomaly-based IDSs crucial for preventing cyberattacks in IoT networks.
In recent years, machine learning techniques have shown promise in developing
anomaly-based IDSs to secure IoT systems, as machine learning models can be
trained on both normal and anomalous data to identify irregularities [1, 2]. However,
effective anomaly detection using machine learning poses several challenges: first,
traditional machine learning models may lack the depth to capture features needed
to distinguish anomalies accurately; second, these models can be resource-intensive,
making them hard to deploy on constrained IoT devices; third, extensive data is
needed for high-accuracy training, and models may still miss certain threats, resulting
in false positives and false negatives.
Machine Learning Approaches for Anomaly Detection in IoT … 349

However, advances in hardware, such as GPUs, and neural network architectures,


like deep learning, have progressively improved machine learning’s potential for
anomaly detection. This progress is especially promising with emerging technologies
such as blockchain, which offer new avenues for securing IoT.
This study offers a comprehensive analysis of current initiatives to create machine
learning-based anomaly detection systems for the Internet of Things, which will help
academics and developers create new anomaly-based intrusion detection systems. We
provide the following contributions: Section 2 discusses the significance of anomaly
detection in the Internet of Things; Sect. 3 outlines the difficulties in implementing
anomaly detection in IoT environments; Sect. 4 examines cutting-edge machine
learning techniques for anomaly detection; and Sect. 5 examines the use of these tech-
niques in IoT security. In particular, we look at blockchain-based anomaly detection,
which uses the resilience of the distributed ledger to support safe, adversarial-resistant
learning environments (Sect. 5), and federated learning, which allows collaborative
model training for improved anomaly detection (Sect. 4).

2 Role of Anomaly Detection in Securing IoT Systems

Anomaly-based IDSs have been utilized across various IoT applications over the
years, as shown in Table 1. This section will highlight the critical functions of anomaly
detection systems in sectors such as industry, smart grids, and smart cities.
The application of anomaly detection tools has significantly improved industrial
IoT. These tools have been applied to a wide range of industrial IoT applications, such
as power systems, health monitoring [28], HVAC system fault detection [29], produc-
tion plant maintenance scheduling [30], and manufacturing quality control [31]. In
[32], machine learning techniques, such as linear regression, were applied to sensor
data from engine-driven machinery to identify deviations from typical operational
patterns. This study demonstrated how important anomaly detection is to preven-
tive maintenance since it may be used to find inefficiencies and machine problems.
By examining reconstruction errors, a different study investigated the application
of autoencoder (A.E.)-based outlier detection in audio data [33]. According to the
results, identifying abnormalities early on could help with responsive maintenance
for equipment, reducing downtime. Furthermore, IoT-based anomaly detection [34]
has been used by water facilities as a reactive alarm system to track and identify
particular chemical concentration levels. Together, these studies demonstrate how
IoT anomaly detection improves industrial machinery uptime and efficiency through
efficient health monitoring using data from smart meters and statistical techniques,
research in [35] created an anomaly detection system. According to the authors,
anomaly detection in power systems can be efficiently modeled using hierarchical
network data.
Using high-frequency signals to detect abnormalities in power network failures,
another study [36] came to the conclusion that network size had a greater impact on
local anomaly identification than topology. Big data analysis methods for identifying
350 R. Rajaan et al.

Table 1 Anomaly-based IDSs: types of anomalies and application areas


Observations Contextual Collective
Applications General Hwang et al. [6] Manimurugan et al. Protogerou
[7] et al.[8]
Cauteruccio et al. – Hasan et al. [10]
[9]
– – Al-Hawawared
et al. [11]
– – Shukla et al.
[12]
– – Yin et al. [13]
– – Tsogbaatar
et al. [14]
– – Diro et al. [15]
Routes – Farshchi et al. [16] –
Sectors Ferrari et al. [17] – –
Bhatia et al. [18] – –
Savic et al. [19] – –
Fitness – Ngo et al. [20] –
Tech-savvy cities Alrashdi et al. [21] –
Automated grids – Utomo et al. [22] –
High-tech home Cheng et al. [23] – Han et al. [24]
– – Chalapathy
et al. [25]
Nguyen et al.
[26]
Drones – He et al. [27] –

and pinpointing power system failures were examined in [37], showing that the circuit
theory compensation theorem might be used for power network event detection.
Using high-frequency signals to detect abnormalities in power network failures,
another study [36] came to the conclusion that network size had a greater impact on
local anomaly identification than topology. Big data analysis methods for identifying
and pinpointing power system failures were examined in [37], showing that the circuit
theory compensation theorem might be used for power network event detection.
Additionally, anomaly detection can be used for smart city infrastructure, such
as buildings and roadways. Road surface abnormalities were studied in [39], which
suggested that keeping an eye out for these irregularities could save private car
damage by enabling prompt maintenance before events happen. In order to help poli-
cymakers make well-informed decisions on environmental, traffic, and health issues,
another study [40] modeled pollution monitoring and control as an anomaly. Similar
to this, IoT-based anomaly detection might be useful in assisted living settings as, as
Machine Learning Approaches for Anomaly Detection in IoT … 351

shown in [41], changes from typical conditions can notify caregivers. In conclusion,
anomaly detection systems may successfully spot unusual circumstances in smart
buildings and cities, giving decision-makers insightful information.

3 Challenges in Leveraging Machine Learning for IoT


Anomaly Detection

The establishment of anomaly detection systems within the IoT landscape presents
various challenges due to multiple factors, including (1) limited IoT resources; (2) the
difficulty in defining normal behaviors; (3) high data dimensionality; (4) contextual
information; and (5) insufficient robustness of machine learning models [15]. This
section will elaborate on these factors.

3.1 Limited IoT Resources

Storage, computing power, communication, and energy resource constraints may


limit the efficacy of device-level anomaly detection in the Internet of Things. Cloud
computing can be used as a platform for data processing, storage, and collecting
in order to address this. However, because of resource scheduling and round-trip
communication durations, the distance to the cloud could result in a considerable
lag. The real-time requirements of identifying suspicious activity in IoT systems
may be hampered by such delays [15]. Furthermore, if the amount of traffic in the
IoT ecosystem exceeds the capability of the devices, anomaly detection systems
may operate poorly. Sending summary data to the cloud or moving some processing
and storage functions from devices to edge nodes are more practical strategies. By
keeping only particular data points, sliding window approaches can help reduce
storage requirements, however, the anomaly detection system could still need access
to patterns and trends [26].

3.2 Defining Normal Behaviors

Accurately defining what is considered typical activity is difficult, yet gathering


sufficient data about normal behaviors is essential to the efficacy of an anomaly
detection system. Rarely occurring abnormal behaviors may be mistakenly catego-
rized as normal. Supervised learning is not feasible in IoT environments due to a
lack of datasets that accurately depict both normal and aberrant data, particularly for
352 R. Rajaan et al.

commonly used IoT devices. This emphasizes the need for unsupervised or semi-
supervised modeling of IoT anomaly detection systems, where data that deviates
from accepted typical operations is categorized as anomalous [3].

3.3 Data Dimensionality

Univariate IoT data is represented by key-value pairs (xt) while multivariate data
is composed of temporally associated univariate data (xt=[xt1,…,xtn). While multi-
variate detection concentrates on past correlations and interactions among attributes
at a particular time, anomaly detection in univariate data compares current values
against historical time series. Therefore, considering the related processing over-
heads, data dimensionality influences the choice of anomaly detection technique
for IoT applications [3, 29]. Furthermore, multivariate data makes model processing
more complex, requiring the use of dimension-reduction strategies like auto encoders
(A.E.s) and principal component analysis (P.C.A.). On the other hand, patterns and
correlations that can improve machine learning performance could not be adequately
captured by univariate data.

3.4 Contextual Information

Because IoT devices are dispersed, contextual information must be included for
anomaly detection to be effective. Correlating temporal inputs at distinct times, like
t1 and tn, as well as spatial contexts in large IoT networks, where certain devices
may function in a mobile fashion, is a big problem. Anomaly detection systems can
benefit from context integration, but if the right context is not appropriately recorded,
it can also make things more difficult [3].

3.5 Insufficient Resilience of Machine Learning Models


to Adversarial Attacks

Current machine learning models are vulnerable to adversarial assaults during both
the training and detection stages and frequently have trouble sustaining a low false-
positive rate. The creation of precise algorithms and robust models is required in
this scenario. On the other hand, as many of the devices in the network have similar
features, the extensive use of IoT devices may be beneficial for collective anomaly
identification. Cooperative efforts against cyberthreats like malware can be strength-
ened by this plethora of devices [42]. However, because adversaries may supply false
Machine Learning Approaches for Anomaly Detection in IoT … 353

data to influence or damage the model, problems like model poisoning and evasion
can reduce the efficacy of machine learning models.

4 Techniques in Machine Learning for Anomaly Detection


in IoT

When it comes to IoT anomaly detection utilizing machine learning, several factors
need to be taken into account. The learning algorithms can be divided into three main
categories: supervised, unsupervised, and semi-supervised. The approach of training
these algorithms across multiple decentralized IoT devices is referred to as federated
learning. Additionally, anomaly detection can be analyzed based on the dimension-
ality of the data, leading to univariate and multivariate methods. Anomaly detection
techniques based on (1) machine learning algorithms, (2) federated learning, and (3)
data sources and dimensions will be covered in the remaining portion of this section.

4.1 Anomaly Detection Strategies Using Machine Learning


Algorithms

Supervised algorithms, also known as discriminative algorithms, are based on clas-


sification techniques that use instances of labeled data. K-nearest neighbor (K.N.N.),
support vector machine (SVM), Bayesian networks, and neural networks (N.N.) are
examples of these classification methods [43, 44]. An anomalous point is identified by
the distance-based anomaly detection algorithm K.N.N. as one whose distances from
the bulk of the dataset surpass a predetermined threshold. However, calculating these
distances is computationally intensive, making on-device anomaly detection with this
algorithm impractical. Conversely, SVM creates a hyperplane to separate data points
for classification, but like K.N.N., it is also resource-heavy, limiting its feasibility
for IoT anomaly detection. Bayesian networks can be advantageous for resource-
constrained devices since they do not require prior knowledge of neighboring nodes
for anomaly detection, albeit with lower accuracy. Neural network algorithms are
widely used to train on normal data, allowing them to identify anomalous data as
deviations from the norm. However, the resource demands of N.N. algorithms pose
significant challenges for their adaptation to the IoT landscape. Consequently, super-
vised algorithms are generally the least suitable for IoT anomaly detection due to
their dependence on labeled datasets and high resource requirements.
Unlabeled data is used by unsupervised algorithms, also known as generative
algorithms, to discover hierarchical features. Unsupervised clustering algorithms,
including K-means and density-based spatial clustering of applications with noise
(D.B.S.C.A.N.), group data points into clusters according to density and similarity
properties [43, 44]. Normal points are either close to or inside dense zones, whereas
354 R. Rajaan et al.

abnormal points are little data points that are substantially separated from those
clusters. To increase the accuracy of anomaly detection, clustering techniques are
frequently used in conjunction with classification algorithms. However, the majority
of clustering methods are not directly relevant to IoT devices for anomaly detection
due to their resource consumption. Dimension-reduction methods, such as autoen-
coders (A.E.) and principal component analysis (P.C.A.), are another unsupervised
learning strategy that aims to reduce the dimensionality of the data by removing
redundancy and noise [44, 45]. Despite its widespread application in anomaly detec-
tion, P.C.A. has trouble in dynamic IoT contexts. On the other hand, by reducing
data volumes and reconstructing errors to find anomalies, A.E. has demonstrated
encouraging results in IoT anomaly identification. That being said, these methods
are commonly used in feature extraction for classification algorithms. IoT anomaly
detection can benefit from adaptations of dimension-reduction techniques used in
unsupervised learning. By offering examples of regular data, semi-supervised algo-
rithms combine generative and discriminative methods, making it possible to identify
anomalies in deviations from typical behavior. As a result, unsupervised or semi-
supervised algorithms are typically preferred for anomaly detection in IoT, with
standard system profiling acting as a baseline [46]. Table 2 presents a summary of
cutting-edge machine learning algorithms categorized according to three types of
anomalies.

Table 2 Learning algorithms categorized by machine learning methods and anomaly types
Anomaly types
Points Contextual Collective
Machine learning Supervised Altashdi et al. Farshchi Han et al. [24]
approaches [21] et al.[16]
Ferrari et al. Utomo et al. Protogerou
[17] [22] et al. [8]
– – Hasan et al.
[10]
– – AL-Hawareh
et al. [11]
– – Shukla et al.
[12]
– – Yin et al. [13]
– – Tsogbaatar
et al. [14]
Unsupervised Hwang et al. [6] He et al. [27] Chalapathy
et al. [25]
Bhatia et al. [18] – Nguyen et al.
[26]
Semi-supervised Cheng et al. [23] Ngo et al. [20] Diro et al. [15]
– Manimurugan –
et al. [7]
Machine Learning Approaches for Anomaly Detection in IoT … 355

4.2 Federated Learning Algorithms for Training Detection


Schemes

IoT devices can train machine learning models locally via federated learning, also
known as collaborative learning, and send the trained models to a central server for
aggregation instead of the local data [47, 48]. This method is different from traditional
machine learning techniques, which usually need the training data to be centralized
in one place, such as a server or data center.
There are four main steps in the federated learning process. A global machine
learning model for anomaly detection is first established by the server, which then
chooses a subset of IoT devices to get the initialized model. Each chosen device
then uses its local data to train the model before sending the revised model back
to the server. After that, the server compiles all of the models it has received to
provide a current global model. In order to detect anomalies, the server then sends the
completed model back to every IoT device. To account for scenarios in which specific
devices might not be available or might drop out during the federated computation
rounds, it is significant to highlight that the server can repeat the selection of IoT
device subsets, distribution of the global model, collection of trained models, and
aggregation of received models multiple times.
Data in the IoT ecosystem is kept decentralized through the use of federated
learning, improving data privacy. Reduced latency, less network traffic, less power
consumption, and adaptability to different organizations are further advantages of this
strategy. Federated learning does, however, have drawbacks, including vulnerability
to model poisoning [50] and inference attacks [49].

4.3 Detection Methods Using Dimensions and Data Sources

Information from a single IoT device over time is represented as univariate IoT
data. In real-world applications, anomaly detection systems frequently make use
of information gathered from several IoT devices functioning in intricate settings.
Compared to single-source data, these multivariate, multi-source datasets give more
contextual information and temporal and spatial insights that are resistant to noise.

4.3.1 Detection of Univariates Through Non-Regressive Methods

By setting low and high observation thresholds on univariate stationary data and
identifying anomalies when data points deviate from these bounds, threshold-based
processes can be used in non-regressive schemes. This min-max strategy can be
improved with more complex techniques, such as mean and variance thresholds based
on previous data. Another method is to divide data distributions into smaller segments
356 R. Rajaan et al.

using box plots so that fresh data points can be compared to these groups. Threshold-
based procedures can be applied to non-regressive schemes by establishing low and
high observation thresholds on univariate stationary data and detecting anomalies
when data points depart from these bounds. More sophisticated methods, including
mean and variance criteria based on prior data, can enhance this min-max approach.
In order to compare new data points to these groupings, box plots can also be used
to split data distributions into smaller parts [3].
Using univariate time series data, neural networks—such as autoencoders
(A.E.), recurrent neural networks (R.N.N.), and long short-term memory networks
(L.S.T.M.)—can be used as non-regressive models to solve anomaly identification
in the IoT landscape. From the input layer to the output layer, A.E. reconstructs
data symmetrically; a large reconstruction error suggests possible abnormalities
[13]. A.E.s can also be used to save battery life and resources in IoT devices with
limited resources. R.N.N.s, on the other hand, preserve memory inside the network
by integrating feedback loops from earlier outputs, which makes it possible to record
temporal contexts over time. However, R.N.N.s are less appropriate for big IoT
networks due to the vanishing gradient problem. To effectively solve this error barrier,
L.S.T.M. can provide semi-supervised learning on normal time series data to detect
anomaly sequences through reconstruction. Therefore, the accuracy and resource-
saving criteria of IoT anomaly detection jobs might be satisfied by combining A.E.
and L.S.T.M.

4.3.2 Regressive Schemes for Univariate Detection

Regressive schemes, also referred to as predictive approaches, use time series data
to compare anticipated and actual values in order to find abnormalities. Despite
their widespread use, parametric models like the autoregressive moving average
(A.R.M.A.) may have problems with seasonality or mean shifts in non-stationary
datasets. Enhanced forms of A.R.M.A., such as seasonal A.R.M.A. and autoregres-
sive integrated moving average (A.R.I.M.A.), can help alleviate these difficulties.
Furthermore, the dynamics of intricate univariate time series data can be captured
by neural network-based predictive models such as multilayer perceptrons (M.L.P.),
R.N.N.s, and L.S.T.M.s [46]. To anticipate expected values, for instance, time series
data variability can be effectively represented by R.N.N.s, L.S.T.M.s, and gated
recurrent units (G.R.U.). IoT anomaly detection in lengthy, complex sequential data
has recently made use of attention-based algorithms. Sequential models, including
non-regressive methods, can improve IoT anomaly detection accuracy when feature
extraction is done using dimensional reduction algorithms.

4.3.3 Multivariate Detection Using Regressive Schemes

Principal Component Analysis (P.C.A.), Autoencoders (A.E.), and other dimension-


ality reduction approaches can be used to reduce the total amount of data when the
Machine Learning Approaches for Anomaly Detection in IoT … 357

number of variables increases and the resulting data sizes increase. P.C.A. is effec-
tive in capturing the relationships among variables within multivariate datasets, as it
reduces the data size by transforming it into a more compact representation. However,
the linearity and computational demands of P.C.A. can restrict its applicability for
anomaly detection in IoT environments.
Auto encoders function similarly to P.C.A. and can identify anomalies in multi-
variate time series data by analyzing reconstruction errors, just as they do in univariate
scenarios. A notable advantage of A.E.s is their efficiency in resource usage and
their ability to perform non-linear feature extraction. In addition to predictive and
non-predictive models applied to univariate data, techniques using Long Short-Term
Memory networks (L.S.T.M.), Convolutional Neural Networks (CNN), Deep Belief
Networks (DBN), and others can also be employed for anomaly detection in multi-
source IoT systems. Specifically, CNNs and L.S.T.M. algorithms can be enhanced
by preceding them with A.E. for crucial feature extraction and resource optimiza-
tion. These deep learning methods are capable of capturing the spatio-temporal
characteristics of multivariate IoT data [12].
Clustering techniques represent another method for detecting anomalies within
multivariate datasets. Additionally, graph networks can be utilized to model the rela-
tionships between variables or sequences, where the weakest connections between
nodes in the graph are identified as anomalous.

5 Evaluation of Machine Learning Techniques


for Identifying IoT Anomalies

Through the identification of questionable activity, anomaly detection systems have


proven to be successful in protecting conventional networks. However, independent
anomaly detection solutions made for traditional networks are not a good fit for
scattered IoT networks. The network as a whole may be in danger under these settings
if one node is compromised.
A collaborative anomaly detection system is essential for combating cyber threats
since it aggregates traffic from multiple places. However, building confidence and
enabling data sharing present two major obstacles [42, 51]. Insider assaults can be a
major risk in a network this size.
Furthermore, because machine learning is used in the majority of anomaly detec-
tion systems, nodes may be reluctant to disclose typical profiles for training or perfor-
mance improvement because of privacy concerns. Establishing a central server in
charge of overseeing trust calculations and data exchange is one possible way to
address the trust problem. This strategy, however, raises the possibility of a single
point of failure, especially when it comes to the security of extensive IoT deploy-
ments. The ability of blockchain technology to build trust among untrusting entities
through contracts and consensus mechanisms has recently attracted a lot of attention
in the financial sector. By providing a platform for data sharing and trust management,
358 R. Rajaan et al.

blockchain offers a chance to address the difficulties associated with collaborative


anomaly detection. We will examine (1) the blockchain-based collaborative archi-
tecture for IoT anomaly detection, (2) relevant datasets and methods, and (3) the
resource needs related to IoT anomaly detection in the parts that follow.

5.1 Collaborative Framework for Identifying IoT Anomalies

Blockchain is a decentralized ledger that uses majority consensus to keep records that
are immutable, trustworthy, authentic, and accountable. Blockchain technology was
initially created for digital money systems, but it may be used in many different fields.
Nodes in a blockchain can verify the generation of new blocks by using consensus
methods, strong hash functions, and public-key cryptography. A set of records, a
timestamp, the hash of the previous block, a nonce, and the block’s hash are usually
included in each block. This means that any changes made to a record or set of
records will be reflected in the hash field of the next block, making it impervious to
malicious changes [42].
In distributed networks like the Internet of Things, the powerful qualities of
blockchain can offer a solid basis for anomaly identification. Because of the
blockchain architecture, IoT devices can work together to jointly build a global
anomaly detection model from local models without running the risk of hostile
attacks. Because IoT depends on mutual trust to share local models securely and
impenetrably, decentralized blockchain storage and consensus algorithms make it
more difficult for bad actors to control the network. But effective consensus tech-
niques, such as the proof-of-work algorithm used by Bitcoin, require a significant
amount of processing and storage power.
Ethereum uses smart contracts and a proof-of-stake method, which uses less
processing resources and determines consensus based on player stakes. Another
adaptable blockchain architecture that uses smart contracts in distributed networks
instead of coins is called Hyperledger Fabric. The ability to endorse transactions
is dependent on a central service, and in order for changes in local ledgers to be
reflected, endorsing parties must agree on transaction values. The requirements of
IoT devices with limited resources might not be sufficiently met by these three
well-known blockchain platforms [51].
Both conventional and Internet of Things systems have investigated blockchain-
based security solutions [52, 53]. In this research, IoT devices are connected to a
resource-rich device that serves as a proxy to connect them to the blockchain. A
comparable study was carried out in [54]. Although resource saving is one of these
systems’ main benefits, there may be a core point of failure. The creators of [55] used
smart contracts to incorporate IoT devices into blockchain systems, guaranteeing
authenticity and communication integrity while resolving resource demand issues
that could otherwise make the integration impractical. In the context of distributed
and collaborative IoT anomaly detection, the most encouraging results have been
observed [51]. To create a dynamic trusted model that nodes can use to compare their
Machine Learning Approaches for Anomaly Detection in IoT … 359

behavior and identify abnormalities, this study uses a self-attestation mechanism.


Prior to being shared with peers, this model is revised jointly by majority consensus.

5.2 Methodologies and Datasets for Finding IoT System


Anomalies

The absence of labeled and realistic datasets has significantly hindered research on
anomaly detection within the IoT sector. The current datasets often fail to provide an
accurate representation of IoT traffic patterns and do not encompass the full spectrum
of anomalies that can arise in these environments. Additionally, there is a noticeable
class imbalance between normal traffic and anomalous patterns, which diminishes
the effectiveness of classification systems. Typically, IoT traffic is predominantly
characterized as normal behavior, albeit with dynamic fluctuations over time. Contex-
tual factors such as time, environmental conditions, and the profiles of neighboring
nodes offer valuable insights that enhance anomaly detection in IoT settings, high-
lighting the importance of multivariate data. The difficulties stemming from the lack
of genuinely representative, realistic, and balanced datasets support the need for an
anomaly detection approach that identifies normal behaviors to pinpoint deviations in
the data [56]. Table 3 lists common datasets that have been utilized in recent research
within this field. It is evident that while most datasets are not specifically designed for
IoT systems, they remain applicable for training and assessing anomaly-based Intru-
sion Detection Systems (I.D.S.) due to their inclusion of both normal and abnormal
data.
The absence of historical data that distinguishes between normal and anomalous
points is impeding the initial deployment of the IoT anomaly detection system. The
implementation of conventional machine learning techniques is made more difficult
by this disparity and the rarity of anomalies. Although a number of methods have
been put out to deal with unbalanced data, these methods frequently fall short of
maintaining the temporal context of abnormalities. Moreover, supervised algorithms

Table 3 Essential datasets for IoT environment anomaly detection (adapted from [1])
Dataset Year of Particular to Measurements Common Unusual
publication IoT instances instances
IoTID 2017 [64] 2017 Yes 30 1,500,000 100,000
CICIDS 2018 [65] 2018 No 100 1,200,000 300,000
ADFA [66] 2014 No 50 450,000 25,000
LITMUS [67] 2016 Yes 60 1,000,000 50,000
CAIDA [68] 2015 No 120 3,500,000 200,000
UNSW-NB15-ML 2019 Yes 51 1,800,000 400,000
[69]
MAWILab [70] 2009 No 45 2,000,000 150,000
360 R. Rajaan et al.

4000000
3500000
3000000
2500000
2000000
1500000
1000000
500000
0
Yes No No Yes No Yes No
2017 2018 2014 2016 2015 2019 2009
IoTID 2017 [64] CICIDS 2018 [65] ADFA [66] LITMUS [67] CAIDA [68] UNSW-NB15-ML MAWILab [70]
[69]

Fig. 1 Essential datasets for IoT environment anomaly detection

cannot detect novel forms of assaults and can only detect recognized anomalies.
The drawbacks of supervised algorithms may therefore be better addressed by unsu-
pervised or semi-supervised techniques [54]. Table 3 gives a Summary of Essential
datasets for Anomaly Detection.
IoT anomaly detection has been approached using a variety of strategies, but many
of them are not able to fulfill the power and resource limits of IoT devices [54]. No
single anomaly detection method is inherently better than another, but deep learning
methods—specifically, Autoencoders (A.E.) and Convolutional Neural Networks
(CNN)—have shown encouraging results in terms of accuracy and resource effi-
ciency, respectively [64]. A.E. can help reduce data dimensionality and extract key
features by filtering out noise, while CNN and L.S.T.M. algorithms can improve
detection accuracy. L.S.T.M. is particularly well-suited for examining intricate and
dynamic patterns in time-series IoT data throughout lengthy sequences. For the
purpose of detecting anomalies in the IoT environment, it is therefore recommended
that these methods, or their combinations, merit additional research [65] (Fig. 1).

5.3 Resource Specifications for Anomaly Detection in IoT

Traditional host-based intrusion detection techniques, such as antivirus and anti-


malware software, are not feasible to deploy on IoT devices due to their low resources.
In order to reduce the processing and storage demands on IoT devices, incremental
solutions such as sliding windows can be used. This is because traffic analysis
consumes a significant amount of CPU power during anomaly detection. Addition-
ally, in order to guarantee successful detection, the anomaly detection engine in the
IoT system must operate in almost real-time. This implies that without requiring a
great deal of retraining, adaptive approaches can gradually improve the detection
model. However, in the initial deployment phase, offline training might be used.
Machine Learning Approaches for Anomaly Detection in IoT … 361

6 Conclusions

The IoT environment’s large number, diversity, and resource constraints have made
it difficult to prevent and detect cyberattacks due to the need to monitor IoT devices
at the network level because on-device solutions are frequently impractical. In this
regard, anomaly detection is especially well-suited for protecting the IoT network,
as it is an essential tool for detecting and warning users of abnormal activity within
the system. Machine learning has been used for anomaly detection in both IT and
IoT systems, but its efficacy has generally been higher in IT environments because of
their superior resource capabilities and perimeter-based locations. Current anomaly
detection techniques based on machine learning are still vulnerable to hostile attacks,
nevertheless. This article offers a comprehensive overview of machine learning-
based anomaly detection in Internet of Things systems, covering the significance
of anomaly detection, the difficulties in creating such systems, and an examina-
tion of the machine learning methods used. Additionally, it implies that adversary-
induced model corruption might be prevented by utilizing blockchain technology,
which would enable IoT devices to work together to create a single model using
blockchain consensus procedures. Our goal is to create a blockchain-based anomaly
detection system in the future to safeguard cutting-edge Internet of Things gadgets
like Raspberry Pi. Raspberry Pi devices might serve as distributed nodes in this
system, which could be deployed on a blockchain platform like Hyperledger Fabric
and a Python-based machine learning framework like TensorFlow.

References

1. Alsoufi MA, Razak S, Siraj MM, Nafea I, Ghaleb FA, Saeed F, Nasser M (2021) A systematic
review of anomaly-based intrusion detection systems in IoT using deep learning. Appl Sci
11:8383
2. Njilla L, Pearlstein L, Wu X, Lutz A, Ezekiel S (2019) Leveraging machine learning for anomaly
detection in the internet of things. In: Proceedings of the 2019 IEEE applied imagery pattern
recognition workshop (A.I.P.R.). Washington, DC, USA, pp 1–6
3. Cook AA, Mısırlı G, Fan Z (2020) A survey on anomaly detection for IoT time-series data.
IEEE Internet Things J 7:6481–6494
4. Cauteruccio F, Cinelli L, Corradini E, Terracina G, Ursino D, Virgili L, Savaglio C, Liotta A,
Fortino G (2021) A comprehensive framework for anomaly detection and classification across
various IoT scenarios. Futur Gener Comput Syst 114:322–335
5. Doshi R, Apthorpe N, Feamster N (2018) Employing machine learning for DDoS detection
in consumer IoT devices. In: Proceedings of the 2018 IEEE security and privacy workshops
(S.P.W.). San Francisco, CA, USA, pp 29–35
6. Hwang RH, Peng MC, Huang CW, Lin PC, Nguyen VL (2020) An unsupervised deep learning
model for early detection of network traffic anomalies. IEEE Access 8:30387–30399
7. Manimurugan S, Al-Mutairi S, Aborokbah MM, Chilamkurti N, Ganesan S, Patan R (2020)
Effective detection of attacks in smart environments of the Internet of medical things using a
deep belief neural network. IEEE Access 8:77396–77404
8. Protogerou A, Papadopoulos S, Drosou A, Tzovaras D, Refanidis I (2021) Distributed anomaly
detection in IoT using a graph neural network approach. Evol Syst 12:19–36
362 R. Rajaan et al.

9. Cauteruccio F, Fortino G, Guerrieri A, Liotta A, Mocanu DC, Perra C, Terracina G, Torres


Vega M (2019) Short and long-term anomaly detection in wireless sensor networks utilizing
machine learning and multi-parameterized edit distance. Inf Fusion 52:13–30
10. Hasan M, Islam MM, Zarif MII, Hashem M (2019) Detection of attacks and anomalies in IoT
sensors at IoT sites using machine learning methods. Internet Things 7:100059
11. AL-Hawawreh M, Moustafa N, Sitnikova E (2018) Deep learning models for identifying
malicious activities in the industrial internet of things. J Inf Secur Appl 41:1–11
12. Shukla R, Sengupta S (2020) A robust and scalable outlier detection system using hierarchical
clustering and long short-term memory (LSTM) neural networks for the internet of things.
Internet Things 9:100167
13. Yin C, Zhang S, Wang J, Xiong NN (2020) Convolutional recurrent autoencoder-based anomaly
detection for IoT time series data. IEEE Trans Syst Man Cybern: Syst 1–11
14. Tsogbaatar E, Bhuyan MH, Taenaka Y, Fall D, Gonchigsumlaa K, Elmroth E, Kadobayashi Y
(2020) SDN-enabled IoT anomaly detection utilizing ensemble learning. International confer-
ence on artificial intelligence applications and innovations. Springer International Publishing,
Cham, Switzerland, pp 268–280
15. Diro AA, Chilamkurti N (2018) A distributed deep learning-based approach for attack detection
in the internet of things. Futur Gener Comput Syst 82:761–768
16. Farshchi M, Weber I, Della Corte R, Pecchia A, Cinque M, Schneider JG, Grundy J (2018)
Contextual anomaly detection for critical industrial systems using logs and metrics. In: Proceed-
ings of the 2018 14th European dependable computing conference (E.D.C.C.). Iasi, Romania,
pp 140–143
17. Ferrari P, Rinaldi S, Sisinni E, Colombo F, Ghelfi F, Maffei D, Malara M (2019) Evaluating
performance between full-cloud and edge-cloud architectures for industrial IoT anomaly detec-
tion based on deep learning. In: Proceedings of the 2019 II workshop on metrology for industry
4.0 and IoT. Naples, Italy, pp 420–425
18. Bhatia R, Benno S, Esteban J, Lakshman TV, Grogan J (2019) Unsupervised machine
learning techniques for network-centric anomaly detection in IoT. In: Proceedings of the 3rd
A.C.M. CoNEXT workshop on big data, machine learning and artificial intelligence for data
communication networks. Orlando, FL, USA, pp 42–48
19. Savic M, Lukic M, Danilovic D, Bodroski Z, Bajovic D, Mezei I, Vukobratovic D, Skrbic S,
Jakovetic D (2021) Deep learning techniques for anomaly detection in cellular IoT applications
related to smart logistics. IEEE Access 9:59406–59419
20. Ngo MV, Luo T, Chaouchi H, Quek TS (2020) Contextual-bandit approaches for anomaly
detection in IoT data across distributed hierarchical edge computing. In: Proceedings of the 2020
IEEE 40th international conference on distributed computing systems (I.C.D.C.S.). Singapore,
pp 1227–1230
21. Alrashdi I, Alqazzaz A, Aloufi E, Alharthi R, Zohdy M, Ming H (2019) AD-IoT: anomaly
detection of IoT cyberattacks in smart cities utilizing machine learning. In: Proceedings of the
2019 IEEE 9th annual computing and communication workshop and conference (C.C.W.C.).
Las Vegas, NV, USA, pp 305–310
22. Utomo D, Hsiung PA (2019) Employing deep learning for anomaly detection at the IoT edge.
In: Proceedings of the 2019 IEEE international conference on consumer electronics—Taiwan
(ICCE-TW). Yilan, Taiwan, pp 1–2
23. Cheng Y, Xu Y, Zhong H, Liu Y (2021) Utilizing semi-supervised hierarchical stacking
temporal convolutional networks for anomaly detection in IoT communications. IEEE Internet
Things J 8:144–155
24. Han N, Gao S, Li J, Zhang X, Guo J (2018) Deep learning techniques for anomaly detection
in health data. In: Proceedings of the 2018 international conference on network infrastructure
and digital content (IC-NIDC). Guiyang, China, pp 188–192. Chalapathy R, Toth E, Chawla
S (2019) Deep learning for anomaly detection: a survey. ACM Comput Surv 54(1):1–38
25. Nguyen TD, Marchal S, Miettinen M, Fereidooni H, Asokan N, Sadeghi AR (2019) DÏoT: a
self-learning anomaly detection system for IoT based on federated learning. In: Proceedings
of the 2019 IEEE 39th international conference on distributed computing systems (ICDCS).
Machine Learning Approaches for Anomaly Detection in IoT … 363

Dallas, TX, USA, pp 756–767. He Y, Peng Y, Wang S, Liu D, Leong PHW (2018) A structured
sparse subspace learning algorithm for detecting anomalies in UAV flight data. IEEE Trans
Instrum Meas 67:90–100
26. Himeur Y, Ghanem K, Alsalemi A, Bensaali F, Amira A (2021) A review of artificial intelligence
methods for anomaly detection in building energy consumption, highlighting current trends and
future perspectives. Appl 287:116601. Piscitelli MS, Brandi S, Capozzoli A, Xiao F (2021) A
data analytics tool for identifying and diagnosing anomalous daily energy patterns in buildings.
Build Simul 14:131–147
27. Kim, D., Yang, H., Chung, M., Cho, S., Kim, H., Kim, M., Kim, K., & Kim, E. (2018). A
squeezed convolutional variational autoencoder for unsupervised anomaly detection in edge
devices within the Industrial Internet of Things. In Proceedings of the 2018 International
Conference on Information and Computer Technologies (ICICT) (pp. 67–71). DeKalb, IL,
USA.
28. Kanawaday A, Sane A (2017) Implementing machine learning techniques for predictive main-
tenance of industrial machinery using IoT sensor data. In: Proceedings of the 8th IEEE inter-
national conference on software engineering and service science (ICSESS). Beijing, China, pp
87–90
29. Shah G, Tiwari A (2018) A machine learning case study for anomaly detection in the industrial
internet of things (IIoT). In: Proceedings of the ACM India joint international conference on
data science and management of data. Goa, India, pp 295–300
30. Oh DY, Yun ID (2018) Anomaly detection based on residual error using an auto-encoder for
machine sound in smart manufacturing systems. Sensors 18:1308. Giannoni F, Mancini M,
Marinelli F (2018). Models for detecting anomalies in IoT time series data. arXiv:1812.00890
31. Moghaddass R, Wang J (2018) A hierarchical framework for detecting anomalies in smart grids
using extensive smart meter data. IEEE Trans Smart Grid 9:5820–5830
32. Passerini F, Tonello AM (2019) Monitoring smart grids using power line modems for anomaly
detection and localization. IEEE Trans Smart Grid 10:6178–6186
33. Farajollahi M, Shahsavari A, Mohsenian-Rad H (2017) Identifying locations of events in distri-
bution networks through synchrophasor data. In: Proceedings of the 2017 North American
power symposium (NAPS). Morgantown, WV, USA, pp 1–6. Yip SC, Tan WN, Tan C, Gan
MT, Wong K (2018) An anomaly detection framework to identify energy theft and defective
meters in smart grids. Int J Electr Power Energy Syst 101:189–203
34. El-Wakeel AS, Li J, Rahman MT, Noureldin A, Hassanein HS (2017) Monitoring anomalies
in road surfaces to support dynamic road mapping for future smart cities. In: Proceedings of
the 2017 IEEE global conference on signal and information processing (GlobalSIP). Montreal,
QC, Canada, pp 828–832
35. Kong X, Song X, Xia F, Guo H, Wang J, Tolba A (2018) LoTAD: a long-term traffic anomaly
detection system based on crowdsourced bus trajectory data. World Wide Web 21:825–847
36. Bakar UABUA, Ghayvat H, Hasanm SF, Mukhopadhyay SC (2016) A survey of activity and
anomaly detection in smart homes. In: Next generation sensors and systems. Springer Inter-
national Publishing: Cham, Switzerland, pp 191–220. Alexopoulos N, Vasilomanolakis E,
Ivánkó NR, Mühlhäuser M (2018) Towards collaborative intrusion detection systems based on
blockchain technology. In: Critical information infrastructures security. Springer International
Publishing: Cham, Switzerland, pp 107–118
37. Hastie T, Tibshirani R, Friedman J (2009) The elements of statistical learning: data mining,
inference, and prediction (2nd ed.). Springer: New York, NY, USA. Murphy KP (2013) Machine
learning: a probabilistic perspective. MIT Press: Cambridge, MA, USA
38. Chadha GS, Islam I, Schwung A, Ding SX (2021) Time series anomaly detection using deep
convolutional clustering. Sensors 21:5488
39. Jiang J, Han G, Liu L, Shu L, Guizani M (2020) Outlier detection methods based on machine
learning for the internet of things. IEEE Wirel Commun 27:53–59
40. Mothukuri V, Khare P, Parizi RM, Pouriyeh S, Dehghantanha A, Srivastava G (2021) Anomaly
detection for IoT security attacks using federated learning. IEEE Internet Things J
364 R. Rajaan et al.

41. Liu Y, Garg S, Nie J, Zhang Y, Xiong Z, Kang J, Hossain MS (2021) A communication-efficient
federated learning approach for deep anomaly detection in time-series data of industrial IoT.
IEEE Internet Things J 8:6348–6358
42. Lee H, Kim J, Ahn S, Hussain R, Cho S, Son J (2021) Introducing digestive neural networks
as a novel defense mechanism against inference attacks in federated learning. Comput Secur
109:102378
43. Wang C, Chen J, Yang Y, Ma X, Liu J (2021) An overview of poisoning attacks and mitigation
strategies in intelligent networks: current status and future outlook. Digit Commun Netw
44. Meng W, Tischhauser EW, Wang Q, Wang Y, Han J (2018) A review of intrusion detection
systems in conjunction with blockchain technology. IEEE Access 6:10179–10188
45. Novo O (2018) Integrating blockchain with IoT: a scalable architecture for managing access
in IoT environments. IEEE Internet Things J 5:1184–1195
46. Dorri A, Kanhere SS, Jurdak R (2017) Developing an optimized blockchain framework for
IoT applications. In: Proceedings of the 2017 IEEE/ACM second international conference on
internet-of-things design and implementation (IoTDI). Pittsburgh, PA, USA, pp 173–178
47. Özyılmaz KR, Yurdakul A (2017) Ongoing work on integrating low-power IoT devices into a
blockchain infrastructure. In: Proceedings of the 2017 international conference on embedded
software (EMSOT). Seoul, Korea, 15–20 October 2017, pp 1–2
48. Huh S, Cho S, Kim S (2017) Managing IoT devices through a blockchain platform. In
Proceedings of the 2017 19th international conference on advanced communication technology
(ICACT). PyeongChang, Korea, 19–22 February 2017, pp 464–467
49. Alsoufi MA, Razak S, Siraj MM, Ali A, Nasser M, Abdo S (2021) A survey of deep learning
techniques for anomaly intrusion detection systems in IoT. Innovative systems for intelligent
health informatics. Springer International Publishing, Cham, Switzerland, pp 659–675
50. Meidan Y, Bohadana M, Mathov Y, Mirsky Y, Shabtai A, Breitenbacher D, Elovici Y (2018)
N-BaIoT: A network-based approach to detecting IoT botnet attacks using deep autoencoders.
IEEE Pervasive Comput 17:12–22
51. Sharafaldin I, Lashkari AH, Ghorbani AA (2018) Creating a new dataset for intrusion detection
and characterizing intrusion traffic. In: Proceedings of the 4th international conference on
information systems security and privacy (ICISSP 2018). Funchal, Portugal, 22–24 January
2018, pp 108–116
52. Kolias C, Kambourakis G, Stavrou A, Gritzalis S (2016) Evaluating threats in 802.11 networks:
empirical assessment and a publicly available dataset. IEEE Commun Surv Tutor 18:184–208
53. Moustafa N, Slay J (2015) UNSW-NB15: a comprehensive dataset for network intrusion detec-
tion systems. In: Proceedings of the 2015 military communications and information systems
conference (MilCIS). Canberra, Australia, 10–12 November 2015, pp 1–6
54. Tavallaee M, Bagheri E, Lu W, Ghorbani AA (2009) An in-depth analysis of the KDD CUP
99 dataset. In: Proceedings of the 2009 IEEE symposium on computational intelligence for
security and defense applications. Ottawa, ON, Canada, 8–10 July 2009, pp 1–6
55. Malaiya RK, Kwon D, Suh SC, Kim H, Kim I, Kim J (2019) Empirical assessment of deep
learning for network anomaly detection. IEEE Access 7:140806–140817
56. Stolfo S, Fan W, Lee W, Prodromidis A, Chan P (2000) A cost-based approach to fraud and
intrusion detection: insights from the J.A.M. project. In: Proceedings of the DARPA information
survivability conference and exposition. Hilton Head, SC, USA, 25–27 January 2000, vol 2,
pp 130–144
57. Kamat P, Sugandhi R (2019) A survey on anomaly detection for predictive maintenance in
industry 4.0. In: Proceedings of the E3S web of conferences. Pune City, India, 18–20 December
2019, p 02007
58. Bovenzi G, Aceto G, Ciuonzo D, Persico V, Pescapé A (2020) A hierarchical hybrid approach
to intrusion detection in IoT environments. In: Proceedings of the GLOBECOM 2020—2020
IEEE global communications conference. Virtual Event, Taiwan, 7–11 December 2020, pp 1–7
A Comprehensive, Multidisciplinary,
Systematic, and Futuristic Approach
to Sustainable Food Systems

M. Amin Mir, Duaa Hefni, Syed M. Hasnain, K. Andrews,


and Minakshi Memoria

Abstract A comprehensive, multidisciplinary, and methodical strategy is needed to


achieve sustainable food systems, which is a global issue. In order to ensure long-term
sustainability, this article integrates environmental, economic, and social aspects to
examine the intricacies of food systems from a variety of angles. The report describes
methods to increase productivity, cut waste, and lessen environmental effects by
utilizing cutting-edge technology including artificial intelligence (AI), data analytics,
and precision agriculture. It highlights the need for circular economy concepts to
close resource loops and the need for policy frameworks that encourage cooperation
among stakeholders in agriculture, industry, and governance. Additionally, it talks
about how cutting-edge solutions like climate-resilient crops, regenerative farming
methods, and alternative proteins can help address global issues like food security,
biodiversity loss, and climate change. A forward-looking perspective on the function
of AI-powered instruments and data ecosystems is also offered, emphasizing how
they might revolutionize decision-making and enable societies to create resilient, just,
and sustainable food systems. The urgency of systemic adjustments to establish a

M. Amin Mir (B)


Department of Mechanical Engineering, Prince Mohammad Bin Fahd University, AL Khobar,
Saudi Arabia
e-mail: mohdaminmir@[Link]
D. Hefni
Department of Science and Human Studies, Prince Mohammad Bin Fahd University, AL Khobar,
Saudi Arabia
S. M. Hasnain · K. Andrews
Department of Mathematics and Natural Sciences, Prince Mohammad Bin Fahd University, AL
Khobar, Saudi Arabia
e-mail: shasnain@[Link]
K. Andrews
e-mail: kandrews@[Link]
M. Memoria
Department of Computer Science, College of Computer Science, King Khalid University, Abha,
Saudi Arabia
e-mail: [Link]@[Link]

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 365
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
366 M. Amin Mir et al.

food system that is both sustainable and able to feed the world’s expanding population
is highlighted by this interdisciplinary approach.

Keywords Sustainable approach · Food · Economic · Artificial intelligence ·


Data analytics

1 Introduction

Making the world’s food systems more sustainable is becoming more difficult as
a result of pressures that worsen the environment and human health. In order to
successfully change food systems for increased sustainability, there is a move away
from solitary initiatives and toward cooperative tactics. According to the European
Commission (2020) [1], this change is necessary to provide overall food and nutrition
security and to encourage healthier, more sustainable diets. To maximize co-benefits,
this calls for a deeper comprehension of the various elements of contemporary food
systems and their interactions.
Addressing the challenge of climate change brought on by human activity is
largely dependent on the European Green Deal [3] and the United Nations Sustain-
able Development Goals (UN SDGs) [2]. Both concepts support environmental
preservation and sustainable agrifood systems [4]. By 2030, everyone should have
a sustainable future thanks to the Sustainable Development Goals (SDGs), which
were adopted in 2015 [5]. Food systems need to be transformed because they are
currently unsustainable. The need for this transformation is driven by a number of
factors, including resource depletion, pollution, waste, environmental degradation,
biodiversity loss, human development problems, starvation, and noncommunicable
diseases linked to diet. From primary farming and food processing to retail, distribu-
tion, food services, and consumption, Food 2030 highlights the significance of inte-
grating the whole food system. The participation of all stakeholders, including society
(consumers), policy, and science, is necessary for this comprehensive approach.
To promote coherence and advancement, more money, sustained investment and
improved research and innovation policies are necessary.

2 Global Food System Evolution: Overarching Patterns


and Views

The various components and processes that go into the production, processing, distri-
bution, and consumption of food, as well as how they are related to one another,
are all included in a food system. A food system is a network that comprises all
elements (environment, people, inputs, processes, infrastructures, institutions, etc.)
and activities associated with food production, processing, distribution, preparation,
A Comprehensive, Multidisciplinary, Systematic, and Futuristic … 367

and consumption, as well as the socioeconomic and environmental results of these


activities, according to a July 2014 definition provided by the High-Level Panel of
Experts on Food Security and Nutrition (HLPE) (6). Food systems encounter many
obstacles and interact with other industries like transportation and energy. According
to HLPE (2014), the term “food system” is descriptive and does not imply that food
security or socioeconomic and environmental results would be favorable.
There are several definitions and models based on the idea of food systems, some-
times known as food and nutrition systems (7). The development of food system
typologies has been attempted in a number of ways, some based on historical view-
points (8) and others concentrating on the connections between production and
consumption or the differences between producers and consumers (9). As a result of
their combined interconnections, the majority of food systems can be referred to as a
“global food system.” An increasing population has led to an increase in the demand
for food worldwide. Global food and agricultural output have grown dramatically
since World War II, and since the 1950s, yields have been rising continuously.
Modern agro-food systems have not eradicated food insecurity, even if more
food is produced per person now than ever before (10). Even if there is enough
food produced to feed everyone on the planet, 795 million people still suffer from
undernutrition (FAO et al., 2015). Over the past 50 years, there has been a significant
increase in demand for animal products, and this trend is predicted to continue,
especially in emerging nations (11).
Similarly, for the past forty years, the consumption of fish has increased globally
in tandem with trends in overall food consumption (12). Food policy has gotten less
attention in national and international decision-making, despite the fact that food is
now more readily available and more reasonably priced than it has ever been. This is
due in part to the fact that food is now more affordable and widely available. However,
many people still do not have access to enough food, while others are overfed, thus
the global food system is far from operating at its best (12).

3 Food Systems: Interconnected Problems and Various


Difficulties

Although it is essential to the provision of food, agriculture—which includes the


production of crops, livestock, forestry, and fisheries—has a substantial negative
influence on the environment. Threats to biodiversity, soil health, water quality, and
the acceleration of climate change are a few of them. Consequently, it is critical to
balance conflicting needs while reducing the food system’s environmental impact.
Agriculture faces many challenges in the twenty-first century, including the need
to produce more food for a growing population with a declining rural labor force,
provide feedstocks for the growing bioenergy sector, aid in the development of nations
that rely heavily on agriculture, implement more sustainable production practices,
and adjust to the effects of climate change (13).
368 M. Amin Mir et al.

Fig. 1 Representation of a comprehensive, multidisciplinary, systematic, and futuristic approach


to sustainable food systems [10]

In order to address these many nutritional, environmental, economic, and social


challenges, it is essential to comprehend how food systems function and are governed.
Current farming and food systems have been successful in providing vast quantities
of food to global markets, but they have also had a number of negative effects,
such as the loss of biodiversity, the degradation of land, water, and ecosystems,
high greenhouse gas emissions, chronic hunger, micronutrient deficiencies, a sharp
increase in obesity and diet-related diseases, and significant challenges for farmers
worldwide to make a living (13) as shown in Fig. 1.

4 Multidimensionality of Sustainable Food Systems

For contemporary food systems to be managed effectively, a shift to sustainability


is necessary. These systems must use the limited and frequently depleting natural
resources they use to operate in ways that are sustainable in terms of the environment,
economy, society, and culture in order to preserve the global ecosystem. In addition
to being inclusive, the expansion of food systems needs to concentrate on objectives
that go beyond production, like improving food chain efficiency and encouraging
sustainable behaviors and diets (14). Numerous elements, such as food availability,
accessibility, and choice, have an impact on food systems. In addition to cultural and
social factors like religion, transnational food corporations, the food industry, and
A Comprehensive, Multidisciplinary, Systematic, and Futuristic … 369

consumer attitudes and behavior, these factors are influenced by geographic, demo-
graphic, and economic factors like disposable income, urbanization, trade liberaliza-
tion, and globalization (15). The FAO (11) highlights that systems of food produc-
tion and consumption need to accomplish more with less in order to end hunger.
This entails lowering food loss and waste, encouraging sustainable consumption
behaviors, and supporting the sustainable intensification of food production.

5 Artificial Intelligence in Food Safety

One automated, non-destructive, and reasonably priced way to keep an eye on the
safety and quality of food is using machine vision. This image processing-based
technology can be used for a variety of products and in different areas of the food
business. According to studies, image processing can effectively check and catego-
rize fruits, vegetables, cereals, and other food products like pizza, pasta, cheese, and
baked goods. Machine vision is increasingly being acknowledged for its promise in
automating food production operations because of its speed, precision, and capacity
to provide unbiased and consistent evaluations [16]. For example, it has been discov-
ered that placing more hand washbasins near restaurant kitchens or in staff members’
direct line of sight increases the frequency of handwashing. Additionally, studies
show that other treatments are required to improve restaurant workers’ hand hygiene
in addition to food safety training [17] Fig. 2 shows the integration of AI in agriculture.
More frequent and efficient handwashing habits could be promoted by computer
vision systems that keep an eye on employee entry and the production area, improving

Fig. 2 Integration of technology in agriculture [12]


370 M. Amin Mir et al.

food safety. Researchers first emphasized the potential of electronic image detection
to swiftly, affordably, and impartially assess food safety and quality as early as the
2000s [18]. The quality control of agricultural and food products has historically
depended on manual inspection, which can be labor-intensive, time-consuming,
and unreliable. Additionally, human error can occur when using manual proce-
dures, which reduces the precision and dependability of quality evaluations [19].
Conversely, AI-driven food product monitoring provides reliable, effective, and
reasonably priced quality control solutions [19]. By increasing production speed
and accuracy while lowering costs, AI-based technologies offer a major advantage
over manual inspection in the fiercely competitive food business, where product
quality is crucial for success [20].
Hygiene compliance is frequently assessed subjectively by other staff members
in quality management systems like ISO 22000 and HACCP, with the results being
documented using straightforward checklists. However, instances of fabricated or
incorrectly filled-out hygiene control paperwork have raised questions within the
industry over the accuracy of these checks. By facilitating objective, real-time moni-
toring of hygiene practices, including the proper use of masks, helmets, and aprons,
artificial intelligence (AI) and computer vision systems provide a solution [21–23].
For these reasons, prototype technologies for AI and visual processing are really
being created.
More advanced quality control systems are becoming more and more necessary
as consumer awareness and expectations for food safety and quality rise. AI research
has concentrated on exploiting visual cues to automate order management, cleanli-
ness, and hygiene in industrial facilities. The current food sector desperately needs
quick, continuous, and objective monitoring, which these AI-driven technologies
provide. Automatic quality control is a useful instrument for the future of food safety
since it not only guarantees objective evaluation but also improves production speed,
efficiency, and cost-effectiveness [24, 25].

6 Artificial Intelligence in Quality Management Systems


of the Food Industry

The most popular method for guaranteeing food safety is called Hazard Analysis and
Critical Control Points, or HACCP. Fundamental hygienic conditions must be upheld
through precursor programs and good manufacturing practices (GMP) in order for
HACCP to operate efficiently [26]. In this regard, using artificial intelligence (AI) for
cleanliness, hygiene, and order controls may improve the effectiveness of the HACCP
system, enhancing GMP and current precursor programs. The documentation of
HACCP requirements and the general use of the HACCP framework are greatly
aided by AI technology. Foodborne illness prevention may be greatly aided by the
creation of new AI-driven control systems for cleaning industrial environments and
personal hygiene [27].
A Comprehensive, Multidisciplinary, Systematic, and Futuristic … 371

Food safety and quality are the two main determinants of consumer expectations in
the food sector. HACCP for food safety and ISO 9001 for quality management are the
most widely used solutions to meet these standards. These systems work in concert to
have a synergistic effect that enhances organizational performance [28]. This synergy
can be further increased by integrating AI-based controls with quality management
systems. It is anticipated that this kind of integration will have a similar synergistic
impact, increasing the effectiveness of current management techniques. Notably, one
of the most frequent industrial accidents in places where food is produced is a slip
and fall, which is frequently brought on by slick flooring, jumbled materials, and
falling objects [29]. AI control systems are essential for reducing the possibility that
employees will trip or slip as a result of unclean conditions or improper footwear.
In this sense, the ISO 45001:2018 Occupational Health and Safety Management
System can be made more effective by using AI to optimize production area layouts
and monitor floor cleanliness. Organizations can increase operational effectiveness
and safety standards by utilizing AI, resulting in a safer workplace and guaranteed
adherence to quality management and food safety regulations.

7 Conclusion

This review outlines existing and potential applications of AI in the food business.
Additionally, it offers details on the fundamentals, principles, and uses of artificial
intelligence in the food processing industry and related fields. The findings demon-
strated that, with the right indicators and artificial intelligence, continuous, objec-
tive, and economical controls for both safety and quality control during processing
are feasible. Despite the potential for widespread application of these management
options in the food sector, little research has been done on pilot planning and scaling
up. So that this shortcoming can be eliminated by the implementation of scaling up
in food production facilities and the pilot planning of artificial intelligence control.
The food business is one of the sectors with the biggest socioeconomic impacts
and is notable for having a large number of sub-branches. It is believed that the
food industry’s self-monitoring is crucial for economic growth. Both preventing
food poisoning and preventing resource waste require the development of new arti-
ficial intelligence systems. The integration of artificial intelligence control systems
with ISO quality management systems will significantly expand their applications
and efficacy. Future research should concentrate on novel applications of AI control
technologies in the food sector.
372 M. Amin Mir et al.

References

1. European Commission Directorate-General for Research and Innovation (2020) Food 2030:
pathways for action. Research and innovation policy as a driver for sustainable, healthy and
inclusive food systems. European commission directorate-general for research and innovation
2. United Nations (2022) The sustainable development goals report 2022. [Link]
sdgs/report/2022/[Link]
3. European Commission (2021) The European green deal: striving to be the first climate-
neutral continent. [Link]
opean-green-deal_en
4. Kuc-Czarnecka M, Markowicz I, Sompolska-Rzechuła A (2023) SDGs implementation, their
synergies, and trade-offs in EU countries—sensitivity analysis-based approach. Ecol Ind
146:109888. [Link]
5. Mondejar ME, Avtar R, Diaz HLB, Dubey RK, Esteban J, Gomez-Morales A, Hallam B,
Mbungu NT, Okolo CC, Prasad KA et al (2021) Digitalization to achieve sustainable develop-
ment goals: steps towards a smart green planet. Sci Total Environ 794:148539. [Link]
10.1016/[Link].2021.148539
6. High Level Panel of Experts on Food Security and Nutrition (HLPE) (2014) Food losses and
waste in the context of sustainable food systems. HLPE
7. Amin Mi M, Waqar Ashraf M, Andrews K (2024) Assessment of heavy metals and fungi
contamination of spices available in Saudi Arabian food cuisines. Food Chem Adv 4:100694
8. Malassis L (1996) Les trois âges de l’alimentaire. Agroalimentaria. [Link]
bitstream/123456789/17732/1/articulo2_1.pdf
9. Esnouf C, Russel M, Bricas N (eds) (2013) Food system sustainability: insights from duALIne.
Cambridge University Press
10. Mir MA, Chang SK, Hefni DA (2024) Comprehensive review on challenges and choices of
food waste in Saudi Arabia: exploring environmental and economic impacts. Environ Syst Res
13:40. [Link]
11. FAO, IFAD, & WFP (2015) The state of food insecurity in the world 2015: meeting the 2015
international hunger targets: taking stock of uneven progress. FAO
12. World Wildlife Fund UK (WWF-UK) (2013) A 2020 vision for the global food
system: Report summary. [Link]
mary_feb2013.pdf
13. FAO (2009) The state of food and agriculture. FAO
14. Capone R, El Bilali H, Debs P, Cardone G, Driouech N (2014) Food system sustainability and
food security: connecting the dots. J Food Secur 2(1):13–22. [Link]
15. FAO (2012) Towards the future we want: end hunger and make the transition to sustainable
agricultural and food systems. [Link]
16. Green LR, Selman C, Radke V, Ripley D, Mack JC, Reimann DW, Stigger T, Motsinger
M, Bushnell L (2007) Factors related to food worker hand hygiene practices. J Food Prot
70:661–666. [Link]
17. Amin Mir M, Chang SK, Waqar Ashraf M, Andrews K (2024) Heavy metal and mycotoxin-
producing fungi contamination of coffee consumed in Saudi Arabia. Food Chem Adv
5(2024):100798
18. Sun D-W (2000) Inspecting pizza topping percentage and distribution by a computer vision
method. J Food Eng 44:245–249. [Link]
19. Vithu P, Moses JA (2016) Machine vision system for food grain quality evaluation: a review.
Trends Food Sci Technol 56:13–20. [Link]
20. Sun D-W, Brosnan T (2003) Pizza quality evaluation using computer vision—Part 1: Pizza
base and sauce spread. J Food Eng 57:81–89. [Link]
21. Rani S, Memoria M, Choudhury T, Sar A (2024) A comprehensive review of machine learning’s
role within KOA. EAI Endorsed Trans Internet Things 10
A Comprehensive, Multidisciplinary, Systematic, and Futuristic … 373

22. Kirola M, Memoria M, Dumka A, Tripathi A, Joshi K (2022) A comprehensive review study
on: optimized data mining, machine learning and deep learning techniques for breast cancer
prediction in big data context. Biomed Pharmacol J 15(1). [Link]
23. Kumar V, Joshi K, Kumar R, Anandaram H, Bhagat VK, Baloni D, Tripathi A, Memoria M
(2023) Multi modalities medical image fusion using deep learning and metaverse technology:
healthcare 4.0 a futuristic approach. Biomed Pharmacol J 16(4). [Link]
2772
24. Ramnarayan, Memoria M, Kumar A, Ghildiyal S (2022) A rapid computing technology on
profound computing era with quantum computing. In: Lecture notes in networks and systems,
vol 434. [Link]
25. Kumar A, Dubey KK, Gupta H, lamba S, Memoria M, Joshi K (2022) Keylogger awareness
and use in cyber forensics. In: Lecture notes in networks and systems, vol 434. [Link]
10.1007/978-981-19-1122-4_75
26. Amin Mir M (2023) Molecular dynamic, Hirshfeld surface, molecular docking and drug like-
ness studies of a potent anti-oxidant, anti-malaria and anti-Inflammatory medicine: Pyrogallol.
Results Chem 5:100763
27. Michaels B, Keller C, Blevins M, Paoli G, Ruthman T, Todd E, Griffith CJ, Harbor C (2004)
Prevention of food worker transmission of foodborne pathogens: risk assessment and evaluation
of effective hygiene intervention strategies. Food Serv Technol 4:31–49. [Link]
1111/j.1471-5740.2004.00091.x
28. Topoyan M (2003) Analysis of hazard analysis and critical control points (HACCP) and ISO:
9001:2000 quality management system relations in food industry (Master’s thesis, Dokuz Eylül
University Institute of Social Sciences)
29. Kurt E (2019) Investigation of work accidents in dried fruits factory. OHS Academy 2:88–118
An Automated Extract, Transform, Load
(ETL) Pipeline to Facilitate Acquisition
and Analysis of Stock Marker Data

B. Uma Maheswari , K. S. Naveen Sakthivel, D. Kavitha ,


and R. Sujatha

Abstract In response to the imperative for data-driven decision-making in the stock


market, this project presents an Automated Extract, Transform, Load (ETL) Pipeline
tailored to streamline stock market data acquisition and analysis. Utilizing AWS
infrastructure and Deepnote for robust data handling, the pipeline extracts raw data
from diverse sources, applies essential transformations and loads the refined data
into an AWS database. Scheduled execution and monitoring ensure real-time updates,
while Data Studio enables dynamic visualization for insightful analysis. The project’s
focus on maintenance and optimization ensures sustained peak performance. Addi-
tionally, the integration of machine learning algorithms automates technical anal-
ysis, empowering users to make informed buy/sell decisions based on data-driven
insights, thus bridging the gap between traditional investment strategies and advanced
analytics.

Keywords Extract · Transform · Load · Machine learning · Technical indicator ·


Technical analysis · Stock data

1 Introduction

Subsequent paragraphs, however, are indented. In a time when making decisions


based on data is crucial, the efficient gathering, handling, and presentation of data
are vital, especially in the constantly changing realm of stock markets. India’s stock
market, referred to as the National Stock Exchange (NSE), boasts 1923 currently
traded stocks, while the Bombay Stock Exchange (BSE) features 4229 actively
traded stocks as of October 2023. It continues to be one of the largest and busiest
stock exchanges on earth, serving as a true worldwide hub of trade and financial
opportunities. Due to inadequate understanding of stock movement, almost 94% of
investors experience losses (Source [Link]).

B. Uma Maheswari (B) · K. S. Naveen Sakthivel · D. Kavitha · R. Sujatha


PSG Institute of Management, PSG College of Technology, Coimbatore, India
e-mail: uma@[Link]

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 375
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
376 B. Uma Maheswari et al.

Making educated judgments in the stock market requires following the movement
of stock prices in both fundamental and technical analysis [4, 13, 16]. While funda-
mental analysis is based on the fundamental financial aspects of the organization
represented by financial ratios [7], technical analysis believes that historical data
could be used to predict future stock prices and trends [2, 19, 5]. Price tracking is a
vital investing technique since it helps investors identify patterns, whether it is bullish
or bearish, and enables fundamental analysts to tailor investment choices appropri-
ately [20]. In addition, technical analysts use indicators like moving averages [24]
and support and resistance levels [15] to determine the best times to enter and exit
positions based on price movements [21]. Price changes may also be used as risk
indicators, with sharp drops perhaps indicating significant volatility or underlying
issues with a firm and making investors rethink investments [18]. Price tracking is
essential for historical reasons as well since it reveals a stock’s past performance
and aids in forecasting its future moves [17]. It also has a crucial part in technical
analysis’s divergence analysis, which aims to find discrepancies between prices and
indicators that can point to market reversals or movements in mood [6]. Recently
many machine learning techniques and deep learning models have been implemented
to predict stock prices [1, 8, 10, 12].
Furthermore, variations in stock prices may be utilized to confirm or refute the
findings of both technical [3] and fundamental research, enabling investors to make
wiser decisions. Certain studies also showed how cultural factors influenced the
stock market decisions [9]. The monitoring of stock prices, which allows for making
educated decisions by taking into consideration both fundamental and technical
factors [14, 22], is a basic component of stock market analysis [23]. The project’s
objective is to deliver a reliable and flexible system designed to simplify the acqui-
sition and technical analysis of stock market data by traders, analysts, and financial
organizations. Multiple stakeholders use various modes and tools to track financial
data and stay informed about investments, which are mostly manual processes. Some
of those methods for tracking data are Stock Market Websites and Apps, News media,
Trading platforms, stock screeners, etc., Investors often use a combination of these
modes to track data and make informed investment decisions. The choice of tools and
sources depends on investment strategy, preferences, and the level of detail required
for monitoring investments. Investors often face several difficulties when tracking
the movements of stocks manually. These challenges can be time-consuming, error-
prone, and limit the ability to make well-informed decisions. This process faces
many difficulties like Time-Consuming Data Collection, Limited Data Coverage,
Data Entry Errors, Lack of Real-Time Data, Difficulty in Historical Data Analysis,
Information Overload, and repetitive tasks. There is a gap and Automated solutions,
like the ETL pipeline, can significantly alleviate these challenges, enhancing the
accuracy and efficiency of stock market tracking and analysis.
An Automated Extract, Transform, Load (ETL) Pipeline to Facilitate … 377

2 Objective

The objective of the study would include.


• Automated ETL Pipeline: Develop a robust and fully automated ETL pipeline for
stock market data to streamline the data acquisition and processing workflow.
• Data Extraction, Transformation, and Loading (ETL): Create a seamless process
for extracting raw stock market data, applying necessary transformations to
enhance data quality, and loading the processed data into a MySQL database
for storage and analysis.
• Scheduled Execution and Monitoring: Implement scheduling capabilities to
ensure the ETL pipeline runs at predefined intervals while also establishing a
monitoring system to track the pipeline’s performance and health.
• Analysis and Visualization: Facilitate the analysis of transformed stock market
data through an interactive Data Studio Live dashboard, enabling people to make
data-driven decisions and acquire insights.
• Friendly User Interface: Create a user-friendly and straightforward interface for
the Data Studio live dashboard to provide easy accessibility and data exploration
for investors, analysts, and financial institutions.
• Real-Time Data: Make sure that the ETL pipeline can handle real-time data
updates so that users may obtain the most recent stock market data for quick
decision-making.

3 Proposed Model

The project embarks on an intricate journey of data, commencing with the iden-
tification of diverse data sources, which can range from websites, applications,
databases, and more. The subsequent steps are executed seamlessly through a well-
structured pipeline involving state-of-the-art technologies. It kickstarts the process
by harnessing the power of Amazon Web Services (AWS), setting up dedicated S3
buckets for both raw and processed data. Scalability, security, and data manage-
ment are all solidly supported by AWS. A cloud storage service called Amazon S3
(Simple Storage Service) enables online data storage and retrieval for things like
files and objects. The digital assets can be kept in something akin to a virtual storage
container that is accessible from any location with an internet connection. Simulta-
neously, Deepnote, a collaborative data science platform, is configured to create a
workspace conducive to ETL operations. Here, it requires installation of necessary
packages and creates an environment primed for efficient data handling.
The project’s major procedure, the ETL, is carried out with Python. Deepnote
is a web-based tool for Python programming and data science that makes it simple
to share and collaborate on Jupyter notebooks. It provides a user-friendly interface
for data analysis, visualization, and machine learning, making it accessible for both
individuals and teams. Deepnote is utilized to extract raw data, clean it, and perform
378 B. Uma Maheswari et al.

any required transformations. This stage ensures that data is accurate, structured,
and ready for analysis. Further bolster the ETL pipeline by setting up AWS Lambda,
which orchestrates connections between Deepnote and S3, facilitating a seamless
flow of data. With the cloud computing service offered by AWS Lambda, it can
respond to events by executing code without having to manage servers. It’s like a
programmable, on-demand function that automatically scales as needed, making it a
cost-effective way to execute code in the cloud. To ensure long-term persistence and
ease of access for processed data, it is subsequently placed into an AWS RDS-hosted
MySQL database.
The processes involved with managing, storing, and retrieving data are substan-
tially simplified by MySQL, an open-source relational database management system
(RDBMS). Due to its easy-to-navigate tools for accessing and managing data, its
capacity to quickly arrange data into organized tables has solidified its appeal across
a wide range of applications, ranging from websites to complex business solutions.
To guarantee the pipeline’s health and performance, comprehensive monitoring and
error-handling mechanisms are implemented. Any exceptions or issues are addressed
promptly, ensuring the reliability of the data flow. 6 Moreover, the project doesn’t
stop at implementation; it emphasizes maintenance and optimization.
Regular reviews and updates keep the pipeline at its peak performance, ensuring
that tools and libraries are current. With the aid of Datastudio, users can convert
complex data into understandable, dynamic visual reports and dashboards. It’s
widely used for data analysis and business intelligence, allowing users to explore
and communicate insights from the data visually. Finally, The Model connects with
MySQL database to Datastudio for data visualization, enabling investors, analysts,
and financial institutions to access and analyze transformed data through a Datas-
tudio live dashboard. This dynamic interface offers a user-friendly, real-time view
of stock market trends, empowering users to make informed investment decisions.
Extraction: Data is gathered from numerous home systems at the start of an ETL
(Extract, Transform, Load) process. A large number of data warehousing initiatives
include collecting data from various source systems, each of which may use a distinct
data format. Relational databases and flat documents are frequent data source formats,
although non-relational database structures like IMS or others may also be used.
The data gets transformed during extraction into a format appropriate for further
transformation processing. To decrease the data volume, unnecessary data is deleted.
The efficiency of operational systems must not be adversely affected by the extraction
procedure. Generally, this extraction takes place in the backdrop or at hours of low
activity, as at night.
Transformation: The second step entails making necessary modifications to the data
to make it understandable in business terms. Data sets undergo a cleansing process
to enhance data quality and are then aligned with the target database structure. To
prepare the retrieved data for loading, a series of rules or functions are applied
during the transformation process. Certain data sources could call for minimal data
manipulation, while others may necessitate complex data cleansing techniques. Data
transformations are often the major challenging and time-consuming aspect of the
An Automated Extract, Transform, Load (ETL) Pipeline to Facilitate … 379

Fig. 1 Automated ETL work flow

ETL process. Machine learning could be effectively deployed to help in the data
pre-processing and transformation process [11].
Loading: The process of actually loading data into the data warehouse is now
underway. The incremental load can be separated from the basic load, which is
regularly not time-sensitive. While the first phase may have a footprint on the oper-
ating systems, loading, particularly when updating existing data sets, may have a
significant impact on the data warehouse. Usually, incremental loading is a crucial
activity. ETL operations can be carried out either in batch or in real-time. Batch oper-
ations are often run on a regular basis, and are sometimes referred to as near-real-time
if the intervals are very small, like hours or even minutes and seconds. Looking for
data in the data vault is part of the loading phase. Depending on the requirements
of the organization, this procedure can vary dramatically. For instance, some data
warehouses might just substitute out-of-date information with the latest information
(Fig. 1).

4 Process Flow

Step 1: Data Source—Identify Data Source (websites, Apps, Databases,)


380 B. Uma Maheswari et al.

Step 2: AWS Setup—Setup S3 Bucket (Raw & Processed Data).


Step 3: Deepnote Setup—Create Package, install necessary packages.
Step 4: Data Extraction in Deepnote—Python Code for Extraction and clean and
process Raw Data.
Step 5: Data Transformation in Deepnote—Apply Transformations as required.
Step 6: AWS Lambda Setup—Configure Connection to Deepnote And S3.
Step 7: Data Loading to MySQL—Setup MySQL (AWS RDS) and Modify Glue job
for loading.
Step 8: Schedule the ETL Pipeline—Schedule in Deepnotes and test pipeline.
Step 9: Monitoring and Error Handling—Implement Monitoring and handle error
and exceptions.
Step 10: Maintenance and Optimization—Regular review & optimization and Tools,
libraries updated.
Step 11: Connection of Database to Google Datastudio for Visualization and Analysis
(B/S decision).

5 Data Source

Yahoo Finance remains a popular destination for stock information owing to its
intuitive design, which delivers a clear process for all traders, irrespective of knowl-
edge or history in the market ([Link]). Yahoo Finance provides
broad coverage by including data from various stock exchanges worldwide, including
U.S. and international markets, as well as cryptocurrency data. This comprehensive
resource distinguishes itself from others by providing investors not merely with
current stock prices but also a wealth of historical financial data, company state-
ments, relevant news reports, and insightful research, adequately meeting a variety
of research needs ([Link], [Link]).
In addition to allowing users to monitor and administer investment portfolios,
Yahoo Finance furnishes a comprehensive perspective of financial assets in a single
location. The service’s free access to many standard features and ability to offer
premium options for sophisticated needs in a cost-effective manner has made it a
favored choice among investors of all experience levels and budgets. On the other
hand, services like Alpha Vantage are popular for extensive global market coverage,
providing historical and real-time data, technical indicators, and fundamental data.
Quandl is known for its comprehensive financial and economic data offerings, with a
diverse library of datasets and data management tools. These platforms are frequently
utilized by data professionals and quantitative analysts.
An Automated Extract, Transform, Load (ETL) Pipeline to Facilitate … 381

6 Edge of the Model

• Customizable Indicators, Indices (Proprietary indicators)


• Real-Time Data
• Alert by the Model
• Clear and Timely Decision
• Error-free Model
• Suitable for All Trading Style.
• Long Term
• Mid Term
• Short Term
• List_Level2Scalping (Intraday)
• Customizable Indices Development

In this project, a focus has been placed on the creation of customizable indices to
meet the diverse needs of clients in assessing industrial performance. Notably, indices
have been meticulously crafted for sectors such as Batteries, Pharmaceuticals &
Drugs, and Sugar industries. Additionally, leveraging the project’s model capabilities,
it is possible to generate indices based on specific criteria, such as stocks with book
values lower than Rs. 25. This approach empowers users with the flexibility to tailor
indices according to their unique investment strategies and preferences (Figs. 2 and
3).
B. Samples

Raw Data.

Fig. 2 Screenshot of raw data


382 B. Uma Maheswari et al.

Fig. 3 Screenshot of final dashboard RSI

Final Dashboard for Buy/Sell recommendation.

7 Conclusion and Future Scope

This project meets the objective of creating an automated ETL pipeline by developing
a robust process by extracting raw data and also implementing scheduling capabili-
ties by establishing a monitoring system for effective tracking. The automated ETL
pipeline project anticipates significant future developments, including the integra-
tion of DOC Testing and enhancement of technical indicators. This forward-looking
approach aims to streamline the testing of newly created indicators with historical
data, identifying errors and refining performance over time. Moreover, the project
envisions cost reduction and increased efficiency, offering a more economical alter-
native to traditional manual testing methods. Looking ahead, strategic collaborations
with securities concerns are on the horizon, as the project seeks to establish partner-
ships for DOC testing. This initiative not only validates enhanced indicators in real-
world scenarios but also fosters a symbiotic relationship between the automated ETL
pipeline and the financial industry. By paving the way for innovative advancements
and industry collaboration, the project sets a foundation for continuous improvement
and efficiency gains in stock market data analysis.
An Automated Extract, Transform, Load (ETL) Pipeline to Facilitate … 383

References

1. Bhandari HN, Rimal B, Pokhrel NR, Rimal R, Dahal KR, Khatri RK (2022) Predicting stock
market index using LSTM. Mach Learn Appl 9:100320
2. Chong TTL, Ng WK, Liew VKS (2014) Revisiting the performance of MACD and RSI
oscillators. J Risk Financ Manag 7(1):1–12
3. Demir S, Mincev K, Kok K, Paterakis NG (2019) Introducing technical indicators to electricity
price forecasting: a feature engineering study for linear, ensemble, and deep machine learning
models. Appl Sci 10(1):255
4. DraKoln N (2008) Winning the trading game: why 95% of traders lose and what you must do
to win. Wiley, vol 322
5. Edwards RD, Magee J, Bassetti WC (2018) Technical analysis of stock trends. CRC press
6. Fang J, Qin Y, Jacobsen B (2014) Technical market indicators: an overview. J Behav Exp Financ
4:25–56
7. Herawati A, Putra A (2018) S, The influence of fundamental analysis on stock prices: the case
of food and beverage industries. Eur Res Stud 21(3):316–326
8. Hiransha M, Gopalakrishnan EA, Menon VK, Soman KP (2018) NSE stock market prediction
using deep-learning models. Procedia Comput Sci 132:1351–1362
9. Ji LJ, Zhang Z, Guo T (2008) To buy or to sell: Cultural differences in stock market decisions
based on price trends. J Behav Decis Mak 21(4):399–413
10. Mohapatra S, Mukherjee R, Roy A, Sengupta A, Puniyani A (2022) Can ensemble machine
learning methods predict stock returns for Indian banks using technical indicators? J Risk
Financ Manag 15(8):350
11. Mondal KC, Biswas N, Saha S (2020) Role of machine learning in ETL automation. In
Proceedings of the 21st international conference on distributed computing and networking,
pp 1–6
12. Naik N, Mohan BR (2019) Optimal feature selection of technical indicator and stock predic-
tion using machine learning technique. In: Emerging technologies in computer engineering:
microservices in big data analytics: second international conference, ICETCE, Jaipur, India,
February 1–2, Revised Selected Papers 2. Springer Singapore, pp 261–268
13. Nazário RTF, e Silva JL, Sobreiro VA, Kimura H (2017) A literature review of technical analysis
on stock markets. Q Rev Econ Financ 66:115–126
14. Neely CJ, Rapach DE, Tu J, Zhou G (2014) Forecasting the equity risk premium: the role of
technical indicators. Manag Sci 60(7):1772–1791
15. Osler CL (2000) Support for resistance: technical analysis and intraday exchange rates. Econ
Policy Rev 6(2)
16. Petrusheva N, Jordanoski I (2016) Comparative analysis between the fundamental and technical
analysis of stocks. J Process Manag New Technol 4(2):26–31
17. Powell B, Nason G, Elliott D, Mayhew M, Davies J, Winton J (2018) Tracking and modelling
prices using web-scraped price microdata: towards automated daily consumer price index
forecasting. J R Stat Soc Ser A Stat Soc 181(3):737–756
18. Schwert GW (1990) Stock market volatility. Financ Anal J 46(3):23–34
19. Stevens L (2002) Essential technical analysis: tools and techniques to spot market trends. Wiley,
vol 162
20. Usman B (2016) The phenomenon of bearish and bullish in the Indonesian stock exchange.
Esensi: Jurnal Bisnis dan Manajemen 6(2):181–198
21. Vuković D, Grubišić Z, Jovanović A (2012) The use of moving averages in technical analysis
of securities. Megatrend Rev 9(1)
22. Yin L, Yang Q (2016) Predicting the oil prices: do technical indicators help? Energy Econ
56:338–350
23. ZXhang YX, Haxo YM, Mat YX (2023) When should you buy or sell a stock? (PSEi composite
index stock forecast). AC Invest Res J 220(44)
24. Zhu Y, Zhou G (2009) Technical analysis: an asset allocation perspective on the use of moving
averages. J Financ Econ 92(3):519–544
384 B. Uma Maheswari et al.

25. [Link]
26. [Link]
27. [Link]
28. [Link]
An Intelligent Epidural Space Locator
with a Medication Injector

B. Vijayalakshmi, S. B. Mohan, M. Premkumar, M. Sasi Kumar,


and K. Suresh Kumar

Abstract The Smart Epidural Space Locator Device with Medicine Injector, a revo-
lutionary approach to epidural anesthetic treatments, is presented in this research.
It solves issues with proper localization of the epidural area and precise medication
distribution by combining a responsive algorithm for dynamic feedback with real-
time imaging tools like ultrasonography. Machine learning improves success rates by
increasing flexibility in individual anatomy. The gadget has an integrated medicine
injector with programmable settings for individualized and regulated drug delivery.
Healthcare providers can operate more effectively when their interfaces are easy to
use. Promising results from early testing point to possible improvements in epidural
techniques. This technology offers substantial advantages for patient care due to its
creative approach to improving epidural anesthesia’s accuracy, safety, and efficiency.

Keywords Sensors · Medicine injector · Epidural space · Spinal cord locator ·


Anesthesia

B. Vijayalakshmi (B)
IT, Sri Parasakthi College for Women, Courtallam, India
e-mail: vijayalakshmib28071982@[Link]
S. B. Mohan
ECE, S.A. Engineering College, Chennai, India
e-mail: drsbmohan@[Link]
M. Premkumar
ECE, Panimalar Engineering College, Chennai, India
e-mail: drmpremkumar@[Link]
M. Sasi Kumar
EEE, C. Abdul Hakeem College of Engineering and Technology, Ranipet, India
e-mail: pmsasi77@[Link]
K. Suresh Kumar
IT, Saveetha Engineering College, Chennai, India
e-mail: sureshkumar@[Link]

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 385
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
386 B. Vijayalakshmi et al.

1 Introduction

Designed to increase safety and precision when administering epidural anesthesia,


the “Smart Epidural Space Locator Device with Medicine Injector” is a revolu-
tionary development in medical technology. By combining an accurate medication
injector with state-of-the-art smart technology, this device revolutionizes the conven-
tional approach to epidural treatments. Its use of real-time imaging and data analytics
ensures accurate identification of the epidural space, lowers risk, and enhances patient
outcomes. With a focus on how it can transform the way anesthetic and pain manage-
ment are administered and ultimately improve patient care, this publication examines
the development, characteristics, and clinical implications of this intelligent device.
The spinal cord, cervical spine, cauda equina, thoracic spine, epidural space, lumbar
spine, and dura mater are some of the parts that make up the human vertebral column.
The epidural space, which is defined by negative pressure running from the cervical
to the lumbar regions, is the main emphasis of this model. To guarantee that patients
don’t feel any pain throughout surgical procedures, anesthetic is given here. The
spinal cord anatomy is seen in Fig. 1.

1.1 Literature Survey

This paper examines the evolution of epidural anesthesia over time and discusses the
use of smart technologies in medical equipment. An analysis of the combination of

Fig. 1 Anatomy
An Intelligent Epidural Space Locator with a Medication Injector 387

a Smart Epidural Space Locator Device and a Medicine Injector looks at the poten-
tial for improving safety, efficacy, and precision during epidural treatments [1]. The
evolving subject of drug administration systems in anesthesia is examined in this liter-
ature review, with an emphasis on the medicine injector component. It is evaluated
how well a precise medicine injector and the Smart Epidural Space Locator Device
work together to provide patient-specific anesthetic and suitable drug administration
[2]. The article most likely outlines a specific strategy, possibly involving computer
modeling, to verify and improve the accuracy of organ segmentation algorithms.
The study may discuss the issues with current segmentation techniques and provide
a structured validation framework [3]. This research investigates the significance
of real-time imaging in enhancing the accuracy of epidural space localization. The
Smart Epidural Space Locator Device’s integration of imaging technologies is exam-
ined, emphasizing its potential to provide real-time visual guidance during epidural
anesthesia [4]. This review of the literature examines the safety aspects of epidural
anesthesia and assesses how the Smart Epidural Space Locator Device lowers proce-
dural risks. The integration of real-time monitoring capabilities and safety measures
is investigated to show advantages in patient care [5]. The report may discuss certain
technologies, such as Internet of Things devices or data-driven approaches, and how
they enhance treatment outcomes. The findings might provide insight into how these
innovative strategies help create customized and successful treatment regimens [6].
In this review, which focuses on the broader topic of smart devices in regional anes-
thesia, the Smart Epidural Space Locator Device is featured. How this device fits into
the larger trend of employing intelligent technology to enhance patient outcomes and
procedure precision is the primary point of discussion [4]. This article examines how
data analytics may be utilized in anesthesia and how the Smart Epidural Space Locator
Device uses data to optimize epidural treatments. The potential benefits, challenges,
and long-term effects of data-driven approaches in anesthesia are examined [7]. The
Smart Epidural Space Locator gadget’s design and usability are examined in this eval-
uation, with an emphasis on how healthcare professionals engage with the gadget.
The study examines the impact of effective human-machine interaction on improved
user experience and procedural outcomes [8]. This literature review provides an
overview of recent developments in anesthetic device technology, focusing on the
Smart Epidural Space Locator Device. The engineering difficulties, clinical impli-
cations, and technological advancements of these devices in modern anesthesia are
discussed [9]. This review compares the Smart Epidural Space Locator Device to
various epidural devices that are currently available for purchase. The study aims to
clarify the device’s potential for usage in clinical settings by contrasting its features,
disadvantages, and benefits with conventional methods [10]. The article might go
into detail about the technical features of UWB radar and how its special quali-
ties are used to achieve accurate human detection in difficult situations [11]. With
an emphasis on the Smart Epidural Space Locator Device, this study examines the
anticipated advancements in anesthesia technology. The discussion includes future
development forecasts, potential research directions, and the impact of the device on
the growth of anesthesiology practice [12].
388 B. Vijayalakshmi et al.

Uses non-overlapping block processing of cancer gene data to increase the preci-
sion of early breast cancer prediction. When comparing different regression algo-
rithms, random forest regression performs the best [13] develops a smart elec-
tronic glove that allows those with physical impairments to communicate using
hand gestures enables flexible, real-time multilingual communication by combining
gesture recognition with artificial intelligence [14] the development of intelligent
optical catheters for epidural procedures that use fiber optic sensors to increase the
accuracy of catheter placement. By providing real-time monitoring of pressure and
tissue characteristics, these catheters aid in identifying the epidural area and lessen the
risks associated with incorrect implantation [15]. According to the study, these smart
optical technologies are a significant development in the field of regional anesthesia
and could improve patient outcomes and procedural safety [16]. Recent advance-
ments in smart epidural spinal needles have increased the precision of identifying
the epidural region. These developments, which include sensors integrated into the
needle tips that provide real-time input on tissue resistance and position, signifi-
cantly improve needle insertion accuracy [17]. The authors highlight the benefits of
these technologies in reducing procedural issues and improving patient safety, with
potential future developments focusing on greater integration of smart sensors and
AI to optimize anesthetic operations [18].Utilizing ultrasound and state-of-the-art
devices like pressure-sensing and needle-tip tracking systems, epidural treatments in
obstetric anesthesia are safe and precise. These advancements enhance therapeutic
outcomes by improving patient safety and success rates, and future AI integration
could aid in refining these techniques [19]. When giving injections, the pressure
monitoring devices are meant to detect the epidural space. Their research indicates
that the device can assist in identifying the epidural space; however, it has limita-
tions that may result in missed detections, raising concerns over its reliability and
consistency in clinical situations [20]. The authors emphasized the need for research
and technology developments to improve the accuracy of pressure-based systems
and ensure dependable findings across a variety of clinical scenarios [21].
The feasibility of the EPI-Detection device for interlaminar cervical epidural
injections, demonstrates its ability to accurately identify the epidural region and
improve the precision of such procedures [22]. The device’s effectiveness may make
therapeutic operations safer and more successful, indicating a promising advance-
ment in the field of epidural space detection technology [23]. An inventive smart
syringe for epidural anesthesia that enhances control and precision during injections
by combining an actuator with a force sensor. This innovation aims to address the
limitations of manual operations by providing more sensitive and dependable force
feedback, which could improve procedural outcomes and patient safety [24].

1.2 Problem Statement and Contribution

Surgeons manually administer anesthesia into the epidural space. The surgeon will
detect some vibrations when the syringe reaches the negative pressure epidural area
An Intelligent Epidural Space Locator with a Medication Injector 389

Fig. 2 Epidural space

and will halt its insertion before administering the injection. If the surgeon inserts
the syringe in the wrong place, it is taken out and put back in until it reaches the
epidural area. It is done by an iterative procedure. Additionally, the patient may go
into a coma if the needle comes into contact with the Dura. Figure 2 depicts the
human epidural space. As indicated in the problem statement, our goal is to reduce
the risks related to the anesthetic injection in the epidural space. The anesthetic in
the epidural space is administered by the accelerometer and the ultrasonic sensor,
respectively, which also identifies the path to low pressure.
• According to the issue description, our goal is to reduce the risks involved in
giving anesthesia in the epidural space.
• The low-pressure point is identified by the accelerometer and the ultrasonic sensor,
respectively, and anesthesia is injected into the epidural space.

2 Proposed Methodology

An automated anesthetic injection into the epidural area is the goal of the proposed
technology. The developed model consists of an accelerometer, an ultrasonic sensor, a
linear actuator, and a low-pressure vertebral column model. The area that is taken into
account in human anatomy is the lumbar region, which is home to the five vertebrae
that make up the Vertebral Column. Figure 3 shows the epidural injection location.
Four lumbar bones and a low-pressure plastic tube make up the model. Ultrasonic
Sensor: Ultrasonic signals are transmitted and received in order to determine the
route to the low-pressure area. The syringe is advanced along the path using the
linear actuator model. There will be a small vibration as the injection enters the
low-pressure area. The plastic tube’s abrupt vibration fluctuations are detected by
the accelerometer.
390 B. Vijayalakshmi et al.

Fig. 3 Epidural space with


Injection

The goal of the suggested device is to independently provide anesthetic to


the epidural region. An accelerometer, linear actuator, ultrasonic sensor, and low-
pressure vertebral column model are the parts that make up the designed model.
Human anatomy takes into account the five vertebrae in the lumbar region, which is
a section of the vertebral column. The model consists of a low-pressure plastic tube
and four lumbar bones. To find the path to the low-pressure region, the ultrasonic
sensor sends and receives ultrasonic signals. The Linear Actuator model advances
the syringe along the path. When the injection reaches the low-pressure region, there
will be some vibration. An accelerometer detects sudden vibration changes in the
plastic tube.

2.1 System Design

In the Vertebral Column Model, the stationary stand emits ultrasonic waves. When
ultrasonic waves strike a bone after being emitted by an ultrasonic transmitter, they
bounce back. The epidural space with injection is seen in Fig. 3. The lack of bones
during the injection process is shown by the ultrasonography’s penetration of the
epidural space and tissues. The injection moves gradually within using a linear actu-
ator model. The accelerometer detects the syringe vibrating when it approaches the
low-pressure zone. It finds the linear motion, stops it, and injects medicine into the
epidural space. The block diagram for the epidural space locator is displayed in
Fig. 4.

2.2 ADXL335

The ADXL335 is a low-power, thin, small, full 3-axis accelerometer with signal-
conditioned voltage outputs. When measuring acceleration, its full-scale range is at
least 3 g. In tilt-sensing applications, this device computes the static acceleration of
An Intelligent Epidural Space Locator with a Medication Injector 391

Linear Actuator
Ultrasonic sensor Controller
Model

Injecting Syringe
Accelerometer catheter with
medicine

Fig. 4 Auto epidural space locator with medicine injector system

gravity and the dynamic acceleration caused by motion, shock, or vibration. The user
selects the accelerometer’s bandwidth by adjusting the CX, CY, and CZ capacitors
at the XOUT, YOUT, and ZOUT pins. To suit the necessary task, a few bandwidths
were selected. They range from 0.5 Hz to 1600 Hz for the X-axis and Y-axis and
from 0.5 Hz to 550 Hz for the Z-axis. The ADXL335 features a poly-silicon surface-
micro machined structure on top of a silicon wafer. Poly-silicon springs, which
also provide resistance to acceleration forces, hang the structure above the surface
of the wafer. A differential capacitor, consisting of independent stationary plates
and plates attached to the moving mass, is used to measure the deflection of the
structure. Because the capacitor becomes imbalanced owing to acceleration, the
sensor output has an amplitude proportional to the acceleration experienced. The
ADXL335 Accelerometer is displayed in Fig. 5.
The module was created as a breakout board because the ADXL335 signal is
analog. However, the board outline is a grove module that is readily repairable,
much like other groves. The sensor can be used with a combined 3.3 and 5 V
power supply in a standard Arduino project. The ADXL335 uses a single struc-
ture to sense the three axes. As a result, the axes detect direction orthogonally and
have minimal cross-axis sensitivity. Mechanical misalignment is the primary cause
of cross-axis sensitivity. A wide range of items now commonly monitor acceleration

Fig. 5 Accelerometer sensor


392 B. Vijayalakshmi et al.

or one of its derivative properties, such as tilt, shock, or vibration. Thanks to tech-
nological improvements, several accelerometers are now much more accessible and
user-friendly for the general public. Extremely low-temperature hysteresis—typi-
cally less than 1 mg over the 0C temperature range—and creative design techniques
guarantee either good performance or the absence of monotonic behavior.

2.3 Ultrasonic Module

Ultrasonic transducers are devices that convert ultrasonic waves into electrical
impulses or vice versa. Both sending and receiving signals are capabilities of ultra-
sound transceivers; many ultrasonic sensors are also transceivers due to their dual
sensing and transmitting capabilities. Such devices are analogous to transducers in
RADAR and SONAR systems, which determine the characteristics of a target by
interpreting the sound or radio wave echoes, respectively. Active ultrasonic sensors
generate high-frequency sound waves, and they use the time interval between sending
the signal and getting the echo to determine how far away an item is. Passive ultrasonic
sensors are microphones that can detect ultrasonic noises under certain conditions,
convert them into an electrical signal, and then transmit the data to a computer.
An ultrasonic transducer is a device that converts energy into ultrasonic waves,
or sound waves that are higher than the normal range of human hearing. In contrast
to ultrasonic transducers, which convert mechanical energy in the form of air pres-
sure into ultrasonic sound waves, the term is more suitable for piezoelectric trans-
ducers, which convert electrical energy into sound. The HC-SR04 Ultrasonic sensor
measures the distance to an object, like bats or dolphins, using SONAR. It offers
exceptional non-contact range detection with high precision and reliable results in
an easy-to-use format. Sharp rangefinders between 2 and 400 cm (1 and 13 feet) are
not affected by dark materials or exposure to sunlight. Figure 6, the ultrasonic sensor
module is seen.

Fig. 6 Ultrasonic Sensor module


An Intelligent Epidural Space Locator with a Medication Injector 393

Fig. 7 Ultrasound waves reception and transmission

When the trig of SR04 gets a strong (5 V) pulse for at least 10us, the sensor will
initiate an 8-cycle ultrasonic burst at 40 kHz and wait for the reflected ultrasonic
burst to start measuring. When the sensor detects ultrasound from the receiver, it
will set the echo pin to high (5 V) and delay for a period of time proportional to the
distance. To find the distance, measure the width of the echo pin (in tons). Ultrasonic
sensors are sometimes known as transceivers. An item transmits and reflects a short
ultrasonic pulse at time zero. The sensor receives this signal and converts it to an
electric signal. The echo fades away after a certain amount of time, and the next pulse
is sent. This span of time is known as the cycle period. A minimum of 50 ms is the
recommended cycle time. If a trigger pulse with a 10 µs width is sent to the signal
pin, the Ultrasonic module will send out a 40 kHz ultrasonic signal and detect the
echo back. The width of the echo pulse determines the calculated distance. If there
is no obstruction, the output pin will deliver a high-level signal of 38 ms. Figure 7
illustrates the transmission and reception of ultrasonic waves.
The linear actuator system consists of two motors: motor 1 and a syringe. Motor
1 moves the block I constructed out of the syringe and motor 2. The linear actuator
system is shown in Fig. 8. The ultrasonic sensor provides information, which causes
Motor 1 to switch on and the injection to advance. When the accelerometer identifies
the required low-pressure area, motor 1 switches to an OFF state and motor 2 turns
ON to inject the drug into the vertebral model.

2.4 Linear Actuator Program Concept

A linear actuator model consists of two motors, such as motors 1 and 2, and a
syringe arrangement. The motor driver IC is used to code the software on the Arduino
platform. Motor 1 is initially moved directly ahead to the low-pressure area. Motor
394 B. Vijayalakshmi et al.

Fig. 8 Linear actuator


system
Motor 1 Motor 2 Syringe with
Anesthesia

Vertebral Model

1 stops and motor 2 starts when the syringe reaches a low-pressure zone. When
the syringe knob is moved forward by motor 2, anesthesia is administered into the
required area. The flowchart for the suggested system is displayed in Fig. 9.
Algorithm
Step 1: Launch the application.

Motor 1 is
OFF
Start

Set the input Syringe in Original


pins position

Syringe moves
forward

If ultrasonic
sensor distance Motor 1 is
>10 cm OFF

If
accelerometer
Motor 1 is difference >6
ON

Motor 2 is
Move the position of ON
the Syringe

Medicine is
injected

Stop

Fig. 9 Flowchart for the proposed system


An Intelligent Epidural Space Locator with a Medication Injector 395

Step 2: Set up the ultrasonic sensor’s input pins.


Step 3: Turn on motor 1 if the measured distance is greater than 10 cm.
Step 4: If not, repeat step 3 after moving the ultrasonic sensor.
Step 5: Turn off motor 1 and turn on motor 2 if the accelerometer value differs by
more than 6.
Step 6: Continue to run motor 1 if the difference is less than 6.
Step 7: Close the application.

3 Result and Discussion

The creation of the proposed system is intended to reduce the risks of worsening
back pain, blood infections, nerve damage, and paralysis. Patients may depend on
it for safety because it is automated. In most procedures involving epidural anes-
thetic injections and situations involving pregnant patients, the approach has several
applications.
The working model for the suggested system is displayed in Figs. 10 and 11. Addi-
tionally, it greatly reduces the risk of patients entering a coma because the syringe
cannot come into contact with the Dura. It produces accurate syringe injections with
minimal trial and error. Although the suggested method has a simple structure, it
takes a long time to apply. In the medical field, it has the potential to greatly develop
technology.
An auto epidural injection model, which automates the epidural injection
process—which is commonly used to reduce pain during birth or surgery—improves
accuracy and safety by fusing robotics and machine learning. Robotic devices
precisely guide the needle, minimizing challenges, while advanced sensors provide

Power supply Ultrasonic sensor

Fig. 10 Auto Epidural Injection Model and distance measurement using ultrasonic sensor
396 B. Vijayalakshmi et al.

Motor Arduino Motor Syringe

Fig. 11 Accelerometer detection and linear motion of injection

real-time data to change needle placement. AI systems analyze patient-specific data,


such as MRI scans, to ensure optimal needle placement. Automation speeds up proce-
dures and lessens invasiveness, which increases patient comfort. However, using such
technology requires extensive research to ensure clinical efficacy and a significant
financial investment.
Accelerometer detection in injection systems is the process of using accelerome-
ters to monitor and control the motion of medical equipment like syringes or needles.
These sensors measure acceleration over many axes (X, Y, Z) to track changes in
motion, orientation, and velocity. Accelerometers on medical devices can capture
even the smallest movements, providing real-time information about the device’s
position and orientation. During injections, this information guarantees accurate
needle placement. Incorporating accelerometers into automated systems or robotic
arms allows for high-precision needle motion control. The sensor feedback allows
algorithms to adjust the needle’s trajectory so that it hits the target exactly. Auto-
mated systems also employ accelerometers to track injection speed and needle pene-
tration rate. This continual input allows for real-time adjustments throughout the
injection process, leading to controlled, linear motion. All things considered, by
precisely controlling the needle’s velocity, accelerometers increase injection safety
and precision.

4 Conclusion

In order to reduce the risk of worsening back pain, blood infections, nerve damage,
and paralysis, the suggested system is being created. Patients can be guaranteed reli-
ability and safety because it is automated. Pregnancy and the majority of operations
involving epidural anesthesia injections benefit from the approach. It also primarily
reduces the risk of patients entering a coma because the needle cannot touch the
Dura. It produces accurate syringe injections without the need for repeated attempts.
An Intelligent Epidural Space Locator with a Medication Injector 397

The proposed model is easy to use, has a clear structure, and takes a lot of time. It
has the potential to greatly improve medical technologies. On the other hand, the
technology has not yet been tested on humans and is designed for a synthetic model.
Nevertheless, this idea and strategy will work and offer a useful surgical method.
Additionally, it can be useful for certain medical purposes. Embedded systems are
found everywhere. However, many of these systems continue to be isolated islands in
an era where signal processing is being utilized to monitor an expanding number of
systems. Signal processing can be used to carry out the sensing part of the proposed
model for more accuracy in identifying the path to the epidural area. It may be
possible to install a MEMS-equipped pressure sensor inside the syringe to measure
the minuscule negative pressure in the epidural space. To ensure that the patient won’t
feel any discomfort when it is placed into their body, the entire setup may be made
of PVC plastic or another lightweight material.

References

1. John Doyle D, Dahaba AA, LeManach Y (2018) Advances in anesthesia technology are
improving patient care, but many challenges remain. BMC Anesthesiol 18:39. [Link]
org/10.1186/s12871-018-0504-x
2. Althobaiti M, Ali S, Hariri NG, Hameed K, Alagl Y, Alzahrani N, Alzahrani S, Al-Naib I
(2023) Recent advances in smart epidural spinal needles. Sensors (Basel) 23(13):6065. https://
[Link]/10.3390/s23136065
3. GodlyGini J, Anish Kumar J, Adlin Arul A (2017) A model based validation scheme for organ
segmentation in CT scan. Int J Res Electr Eng 4(2):4–9. ISSN: 2349-2503
4. Balavenkatasubramanian J, Kumar S, Sanjayan RD (2024) Artificial intelligence in regional
anaesthesia. Indian J Anaesth 68(1):100–104. [Link]
5. Harbell MW, Methangkool E (2021) Patient safety education in anesthesia: current state and
future directions. Curr Opin Anaesthesiol 34(6):720–725. [Link]
0000000001060
6. Varadharajan G, Anish kumar J (2021) Smart therapeutic treatment for varicose disease. Int J
Res Appl Sci Eng Technol 9(1):161. [Link]
7. Simpao AF, Ahumada LM, Rehman MA (2015) Big data and visual analytics in anaesthesia
and health care. BJA: Br J Anaesth 115(3):350–356. [Link]
8. Johnson TR, Thimbleby H, Killoran P, Diaz-Garelli JF (2015) Human computer interaction in
medical devices. In: Patel VL, Kannampallil TG, Kaufman DR (eds) Cognitive informatics for
biomedicine. In: Health informatics. Springer, Cham. [Link]
72-9_8
9. Moon JS, Cannesson M (2022) A century of technology in anesthesia & analgesia. Anesth
Analg 135(2S):S48–S61. [Link]
10. Shanthanna H, Mendis N, Goel A (2016) Cervical epidural analgesia in current anaesthesia
practice: systematic review of its clinical utility and rationale, and technical considerations. Br
J Anaesth 116(2):192–207. [Link]
11. Madhavan G, Anish Kumar J, Manimegalai L (2015) An UWB radar for trapped human
detection and vital sign extraction. Int J Appl Eng Res 10(29):22448. ISSN 0973-4562
12. Rothman BS, Sandberg WS (2021) Anesthesiology 2030: what is the future? ASA Monitor
85:8–10 [Link]
13. Gini JG, Padmakala S, Kumar JA (2023) Non-overlapping block processing of cancer genes
data for earlier prediction of breast cancer diseases using regression algorithms. In: 2023 IEEE
398 B. Vijayalakshmi et al.

international conference on ICT in business industry & government (ICTBIG), Indore, India,
pp 1–11. [Link]
14. Kumar JA, Selvam BA, Alwin Vinifred C, Magadevi N, Anbazhagan K (2024) Smart electronic
speaking glove for physically challenged person. In: Rathore VS, Tavares JMRS, Surendiran B,
Yadav A (eds) Universal threats in expert applications and solutions. UNI-TEAS 2024. Lecture
Notes in Networks and Systems. Springer, Singapore, vol 1006. [Link]
981-97-3810-6_17.
15. Carotenuto B, Ricciardi A, Micco A, Amorizzo E, Mercieri M, Cutolo A, Cusano A (2018)
Smart optical catheters for epidurals. Sensors 18:2101. [Link]
16. Anish Kumar J, Gowthambigai M, Shanker NR, Jasper J (2022) Prediction of rotor slot width
in induction motor using dyadic wavelet transform and softmax regression. Int J Emerg Electr
Power Syst. [Link]
17. Althobaiti M, Ali S, Hariri NG, Hameed K, Alagl Y, Alzahrani N, Alzahrani S, Al-Naib I
(2023) Recent advances in smart epidural spinal needles. Sensors 23:6065. [Link]
3390/s23136065
18. Kumar JA, Swaroopan NMJ, Shanker NR (2023) Prediction of rotor slot size variations in
induction motor using polynomial chirplet transform and regression algorithms. Arab J Sci
Eng 48:6099–6109. [Link]
19. Capogna G (2020) New techniques and emerging technologies to identify the epidural space.
In: Epidural technique in obstetric anesthesia. Springer, Cham. [Link]
030-45332-9
20. Kumar JA, Gowthambigai M, Shanker NR et al (2023) Prediction of rotor slot size variation
through vibration signal of three phase induction motor using machine learning. J Vib Eng
Technol. [Link]
21. Carassiti M, Pascarella G, Strumia A et al (2022) Pressure monitoring devices may undetect
epidural space: a report on the use of Compuflo® system for epidural injection. J Clin Monit
Comput 36:283–286. [Link]
22. Kumar JA, Swaroopan NMJ, Shanker NR (2022) Average rotor slot size variation measurement
in induction motor using variable Q-factor transforms and regression algorithms. Iran J Sci
Technol Trans Electr Eng 46:675–687. [Link]
23. Kang J, Park SS, Kim CH, Kim EC, Kim HC, Jeon H, Kim KH, Shin DA (2020) Feasibility of
using the epidural space detecting device (EPI-DetectionTM ) for interlaminar cervical epidural
injection. J Clin Med 9:2355. [Link]
24. Yoon K-C, Kim KG, Lee DC, Yoon SJ (2022) Smart syringe using actuator and force sensor
for epidural anesthesia injection. Int J Artif Organs 45(3):331–336. [Link]
03913988211066501
Computational Investigations
of Inorganic Perovskite Absorber
Material Using SCAP-1D Simulator

Syed M. Hasnain

Abstract Perovskite solar cells (PSCs) are a promising option for the next photo-
voltaic systems due to their significant technological and efficiency advancements.
This work examines how the composition of the absorber layer affects the effi-
ciency of PSCs with the Au/Spiro-OmeTAD/BaZrS3 /ZnO/ITO configuration. We
thoroughly adjusted the structural layer thicknesses using the SCAPS-1D simula-
tion program in order to optimize device efficiency. With optimized thicknesses of
0.035 µm for the ZnO layer, 0.22 µm for Spiro-OmeTAD, and 0.55 µm for the
BaZrS3 absorber layer, the BaZrS3 -based perovskite solar cell produced a PCE of
8.547%, an open-circuit voltage (Voc ) of 1.392 V, a short-circuit current density
(Jsc ) of 14.019 mA/cm2 , and a fill factor (FF) of 44.2545%. The study also empha-
sizes how temperature has a significant impact on solar cell performance, with lower
temperatures improving PCE. These results highlight the importance of choosing the
right absorber layer for PSC design, especially the benefit of using materials with
smaller energy gaps for better electrical properties and overall device performance.

Keywords Perovskites materials · Solar cells · BaZrS3 · SCAPS-1D simulator ·


Efficiency

1 Introduction

Sustainable energy generation, efficient energy storage, and combating global


warming are more important than ever in the modern world. Finding alternate renew-
able energy sources has become crucial as carbon dioxide emissions rise and natural
resources are being depleted. Given its abundance and renewable nature, solar energy
has enormous potential to satisfy future energy needs [1]. Solar cell technology is
essential to producing clean electricity since the need for more effective photovoltaic

S. M. Hasnain (B)
Department of Mathematics and Natural Sciences, Prince Mohammad Bin Fahd University,
Al-Khubar, Saudi Arabia
e-mail: shasnain@[Link]

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 399
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
400 S. M. Hasnain

devices is growing along with the world’s energy needs [2]. Up to 20% of the world’s
energy needs may be met by solar power in the next 20 years, according to projections
[3].
Solar energy has emerged as a competitive substitute for conventional, non-
renewable energy sources in recent years. It provides a copious and clean energy
source that doesn’t contribute to global warming or cause direct environmental
harm. To increase the efficiency of semiconductor-based solar cell technologies,
researchers have been studying and developing them in great detail for the past thirty
years. The most popular material for photovoltaic systems at the moment is silicon-
based [4]. However, the high cost and quality requirements of silicon have prevented
silicon solar cells from being widely adopted, even with ongoing efforts to increase
their efficiency [5]. Researchers are therefore looking into more affordable options.
Perovskite-structured materials, which have demonstrated significant promise for
effective energy solutions, are one viable alternative. Because of their exceptional
material properties, organometal halides, like MAPbI3 , MAPbBr3 , and Cs2 AgBiBr6
perovskites, have shown excellent performance in solar cells as well as sensitive
x-ray detectors [6]. These materials are a viable option for upcoming photovoltaic
applications since they also have the benefit of manufacturing high-quality films.
Perovskites absorb light well in solar energy systems opening up new paths in the
search for eco-friendly energy materials. Lead halides that mix organic and inorganic
parts have several good features that make them great to use in solar cells. These
features include long diffusion lengths changeable band-gaps high absorption rates
easy ways to make them, and the ability to work with solutions [7]. Kojima et al.
(2009) showed how to use perovskites in solar cells starting with 3.8% efficiency
[8]. In the next ten years, perovskite solar cells got much better at turning light into
power reaching about 26% [9]. Even with these big improvements, using harmful
lead in perovskites still stops them from being sold. To fix this, scientists are looking
hard for safer options. Metal ions like Sn2+ and Ge2+ , which act like Pb2+ in many
ways, might work as replacements for lead [10]. This shift to non-toxic ingredients
could make future perovskite solar cells safer and better for the environment.
Because Sn2+ has a smaller ionic radius than Pb2+ , studies have shown that
replacing Pb2+ with Sn2+ in perovskite structures has no discernible impact on their
composition [11]. In comparison to lead-based perovskites, Sn-based perovskites
also show smaller band-gaps, which may result in greater efficiencies [12]. Further-
more, in contrast to conventional lead halide perovskites, chalcogenide perovskites,
which have the general formula ABX3 —where A can be Sr, Ba, or Ca, B can be Zr,
Hf, or Ti, and X can be S or Se—are attracting interest for solar applications because
of their superior structural stability and non-toxic nature [13].
BaZrS3 is one of the most promising of them. With a direct band-gap between
1.70 and 1.80 eV and a crystal structure similar to that of GdFeO3 in the Pnma space
group, it is well suited for solar energy utilization [14]. In comparison to many other
materials, BaZrS3 has an outstanding absorption value of 105 cm−1 [15–17]. This set
of characteristics implies that BaZrS3 may be important for solar technologies in the
future. Because of its easy synthesis and lead-free nature, BaZrS3 has become a very
attractive material for solar applications. This inorganic perovskite has a number of
Computational Investigations of Inorganic Perovskite Absorber … 401

important benefits, such as remarkable carrier mobility, the utilization of inexpensive


and plentiful materials, and the capacity to absorb wavelengths near the peak of the
solar spectrum [18]. BaZrS3 is made up of naturally occurring, non-toxic components
that are present in the Earth’s crust at concentrations greater than 165 parts per million,
in contrast to many other compounds [19]. Additionally, it has outstanding chemical
and thermal stability, holding up well in water and at temperatures lower than 600 °C
[20]. Additionally, studies using density functional theory (DFT) have shown that
compounds based on BaZrS3 have an absorption coefficient (α) close to their band-
gap (Eg ), suggesting their great potential for effective solar energy conversion [21].
For upcoming photovoltaic technologies, BaZrS3 is a desirable option due to these
characteristics.
The material properties of inorganic barium zirconium sulfide (BaZrS3 ), which
exhibits promise as a practical and environmentally suitable substitute for future
solar cell applications, are the main focus of this study. BaZrS3 has the potential to be
a highly efficient replacement for lead-based perovskites. The Au/Spiro-OmeTAD/
BaZrS3 /ZnO/ITO combination is examined in order to evaluate its design, run numer-
ical simulations, and determine whether employing BaZrS3 in real-world perovskite
solar cells is feasible and performs well overall. We hope to investigate the possibil-
ities of this lead-free perovskite material for upcoming solar energy technologies by
examining these variables.

2 Objectives and Significance

This project aims to develop and run a computer analysis to check how well perovskite
solar cells work with BaZrS3 as the main light-absorbing material. It focuses on
fine-tuning the Au/Spiro-OmeTAD/BaZrS3 /ZnO/ITO setup to boost its performance.
Computer simulations will predict how this design performs as a solar cell looking
at key measures like power conversion efficiency (PCE), fill factor (FF) open-circuit
voltage (Voc ), and short-circuit current density (Jsc ). To back up what the simulations
show, researchers will build and test full-size solar cells under controlled lighting
conditions. This study matters for several reasons. First, it tackles the need to find
eco-friendly materials for perovskite solar cells zeroing in on alternatives that don’t
use lead. This addresses the push for safer non-toxic solar power systems. Second, the
proposed setup, which uses inorganic perovskite materials, offers a fresh approach
to making solar cells more stable and efficient. What this study finds could shape
future work on perovskite solar cells helping to expand renewable energy options
and promote greener power technologies.
402 S. M. Hasnain

3 Proposed Design

The suggested structure for the solar cells, which consists of Au/Spiro-OmeTAD/
BaZrS3 /ZnO/ITO and is intended to enhance perovskite solar-cell performance, is
depicted in Fig. 1. ITO as Indium Tin Oxide is used in the design as the uppermost
layer that faces the sun. The ideal values for every layer in the structure were ascer-
tained by a thorough investigation. The density functional theory (DFT)-calculated
material properties for each layer are shown in Table 1.
It is crucial to modify the thickness in accordance with the maximum limits and
ideal values of the electrical properties in order to attain the intended optimal thick-
ness for one chosen layer while maintaining the others constant. To guarantee compa-
rability, the batch parameters for each layer were changed uniformly. “SCAPS-1D,” a
program for simulating solar cells, was used in compliance with particular operating
instructions. To get the most accurate data measurements, a methodical technique
was used. In order to identify the energy bands that correspond to the valence and
conduction bands of each layer, a thorough analysis of density functional theory

Fig. 1 Diagram of a solar


cell structure with five
layers. From top to bottom:
"Front Contact Layer"
labeled as ITO, "Electron
Transport Layer" labeled as
ZnO, "Perovskite Absorber"
labeled as BaZrS3 , "Hole
Transport Layer" labeled as
Spiro-OmeTAD, and "Back
Metal Contact" labeled as
Au. Arrows labeled "Incident
photons" point from a sun
icon to the layers, indicating
the direction of light

Table 1 Optimized material properties of the proposed structure


Materials parameters Layer of Layer of BaZrS3 Layer of ZnO
Spiro-OmeTAD
Layer thickness (μm) 0.22 0.55 0.035
Energy band-gap (eV) 3.30 1.80 3.20
Electron/hole thermal 1.0 × 10+7 /1.0 × 1.0 × 10+7 /1.0 × 10+7 1.0 × 10+7 /1.0 × 10+7
velocity 10+7
Nd 0 1.0 × 10+12 1.0 × 10+17
Na 2.0 × 10+19 1.0 × 10+12 0
Computational Investigations of Inorganic Perovskite Absorber … 403

(DFT) and related literature was conducted. The goal of this thorough examination
is to guarantee the accuracy and dependability of the simulation results.
Zinc oxide makes up the electron transport layer and inorganic BaZrS3 makes
up the light-absorbing layer of the suggested solar cell structure. Gold (Au) is the
back contact, while Spiro-OmeTAD is the hole transport layer (see Fig. 1). Each
layer’s PCE was evaluated at its ideal thickness in order to calculate the optimal
power conversion efficiencies (PCE). During the simulations, the batch settings were
changed to change each layer’s thickness. With the highest PCE value, the BaZrS3
layer initially reached its optimal thickness. The ZnO and BaZrS3 perovskite layer
thicknesses were maintained at predetermined values, according to the supplemental
data. Conversely, the Spiro-OmeTAD layer’s thickness was adjusted from 0.1 to
0.35 µm.
Energy band-gap alignment schematic design for the suggested solar cell structure
is displayed in Fig. 2.

Fig. 2 Bar chart illustrating


energy band gaps in electron
volts (eV) for different
materials. The chart shows
three vertical bars
representing ZnO, BaZrI3 ,
and Spiro-OmeTAD. ZnO
has a band gap of -4.3 eV
and a thickness of 0.035 µm,
BaZrI3 has a band gap of
0.55 eV and a thickness of
0.55 µm, and
Spiro-OmeTAD has a band
gap of -5.1 eV and a
thickness of 0.22 µm. Labels
include ETL, HTL, and
energy levels C_B and V_B.
An arrow indicates the
direction of the energy band
gap
404 S. M. Hasnain

4 Results and Discussion

The ideal energy band diagram for the Au/Spiro-OmeTAD/BaZrS3 /ZnO/ITO


perovskite solar cells is displayed in Fig. 3. Energy band-gap of the BaZrS3 absorber
is 1.80 eV. The band gaps for ZnO and Spiro-OmeTAD are measured at 3.20 eV and
3.30 eV, respectively, to attain the best performance in the suggested devices. For
optimal device performance, Indium Tin Oxide (ITO) is utilized with a work function
of 4.3 eV, while gold (Au) metal contacts have a work function of 5.1 eV. Doping or
the addition of chemicals can increase the band gap, enabling customization within
a predetermined range. The valence band is shown in the diagram by the red line,
and the green line denotes the conduction band. Additionally, the graphs show how
mobile charge carriers move: negative charges travel towards the p-type zone, while
positive charges travel towards the n-type region.
The output efficiency and output of the solar cell device are greatly affected by
the operating conditions, as shown in Fig. 4. When alternate back contact layers are
used with the enhanced structural layer thickness, the voltage-to-total-current density
relationship, or “Jt ,” noticeably increases by about 0.7 V. The structural layers of the
suggested arrangement are also shown graphically.
We first adjusted a number of factors, such as thicknesses, contact layer work
functions, interface flaws, band gaps, and structural defects, in order to attain the
best performance values for our solar cell design. Once these variables have been
adjusted, we enter all of the simulation’s parameters for operation. We were able to
determine the best performance values for our proposed absorber after the simulation.
Based on the Au/Spiro-OmeTAD/BaZrS3 /ZnO/ITO design, the solar cell’s electronic
parameters such as open-circuit voltage (Voc ), short-circuit current density (Jsc ), fill
factor (FF), and power conversion efficiency (PCE) were determined to be 1.392 V,
14.019 mA/cm2 , 44.254%, and 8.54%, respectively.

Fig. 3 Energy band diagram


of an optimized solar cell,
showing energy levels (E) in
electron volts (eV) versus
diffusion length (X) in
micrometers (µm). The
diagram includes three
materials: Spiro-OmeTAD
with a bandgap (Eg ) of 3.30
eV at 0.220 µm, BaZrS3
with Eg of 1.80 eV at 0.55
µm, and ZnO with Eg of
3.20 eV at 0.035 µm. The
conduction band (Ec ) is
represented by a green
dashed line, and the valence
band (Ev ) by a red dashed
line.3
Computational Investigations of Inorganic Perovskite Absorber … 405

Fig. 4 Scatter plot showing


the relationship between
voltage (V) on the x-axis and
total current density Jt (mA/
cm2 ) on the y-axis. The data
points form an upward curve,
indicating an increase in
current density with voltage.
Key solar cell parameters are
listed: Voc =1.392 V, Jsc =
14.019 mA/cm2 , FF =
44.254%, and PCE = 8.54%

The efficiency of solar energy systems is significantly influenced by temperature.


Figure 5 shows how variations in temperature affect the device’s output performance.
The current–voltage (Jt –V) curve variations throughout a temperature range of 290–
350 K are examined in this work. The change from lower toward higher current
density levels is shown in the inset of Fig. 5. Within a certain range, the device’s
current density noticeably changes as the temperature rises. However, because of
higher lattice vibrations and charge carrier collisions, the current density starts to
decrease at the ideal temperature. Reduced lattice vibrations at lower temperatures,
which later become more intense as the temperature rises, are the cause of the decrease
in open-circuit voltage at higher temperatures. Higher temperatures cause atoms to
vibrate more, which increases the likelihood of collisions between charge carriers
(electrons and holes) and other atoms. As a result, although the current density and
charge carrier concentration are directly correlated, it tends to rise with tempera-
ture until it reaches an ideal value, after which additional increases could impair
performance.
The correlation between the thickness and electrical characteristics of the BaZrS3
perovskite absorber layer is shown in Fig. 6. In general, as the layer thickness grows,
characteristics like Voc , PCE, and Jsc tend to improve, with the exception of the FF.
Performance can be improved by increasing the density of charge carriers through a
thicker perovskite absorber layer. According to this investigation, current density and
fill factor both improve with increasing BaZrS3 layer thickness, but Voc decreases.
At higher thicknesses, the PCE first increases before sharply declining. For BaZrS3
perovskite absorber, the measured values are Voc at 1.392 V, Jsc at 14.019 mA/cm2 ,
FF at 44.254%, and PCE at 8.54% at the ideal thickness of 0.55 µm. Increasing the
absorber layer’s thickness improves light absorption, which raises Jsc and increases
charge carrier production, both of which increase overall efficiency. Moreover, the
higher Jsc values show that increased electron–hole pair generation leads to greater
electron mobility.
406 S. M. Hasnain

Fig. 5 Chart showing the


relationship between total
current density Jt (mA/cm2 )
and voltage V (V) at various
temperatures: 347 K, 336 K,
324 K, 313 K, 301 K, and
290 K. The main graph
displays a nonlinear increase
in current density with
voltage. An inset graph
highlights a detailed view of
the data between 1.05 V and
1.20 V, showing a linear
trend. Each temperature is
represented by a different
symbol and color

Fig. 6 Four-panel X-Y chart showing the relationship between the thickness of BaZrS3 layer X (in
micrometers) and various photovoltaic parameters. Top left: Power Conversion Efficiency (PCE%)
increases then decreases with thickness, marked by red squares. Top right: Open-circuit voltage
(Voc ) decreases with thickness, marked by blue circles. Bottom left: Short-circuit current density
(Jsc ) increases with thickness, marked by green circles. Bottom right: Fill factor (FF%) shows a
peak and then declines, marked by purple stars. Each panel has a distinct color and marker style for
clarity
Computational Investigations of Inorganic Perovskite Absorber … 407

Fig. 7 The image consists


of four X-Y charts
displaying the relationship
between the band-gap of a
BaZrS3 layer and various
photovoltaic parameters. The
top chart shows the Power
Conversion Efficiency
(PCE%) decreasing as the
band-gap increases. The
second chart illustrates the
Fill Factor (FF%) remaining
relatively constant. The third
chart depicts the short-circuit
current density (Jsc in mA/
cm2 ) decreasing with
increasing band-gap. The
bottom chart shows the
open-circuit voltage (Voc in
V) slightly decreasing. Each
chart uses different colored
markers for clarity

A study was conducted to investigate the effects of changes in the BaZrS3


absorber’s energy band gap on the device’s output performance characteristics.
Figure 7 presents the results, indicating that the energy band gap of BaZrS3 grew
from 1.80 to 1.96 eV. All electrical performance measures seem to significantly
drop in tandem with this increase in the energy band gap. Because electric charge
carriers can go farther before recombining, this occurrence might suggest that their
recombination rate has increased.
Experiments assessing the device’s output performance at high temperatures,
namely between 290 and 350 K, are shown in Fig. 8. This range of temperatures
was chosen to illustrate the general differences in electrical properties. Peak power
conversion efficiency (PCE) of 8.54%, Voc of 1.392 V, FF of 44.254%, and Jsc of
14.019 mA/cm2 were all attained by the BaZrS3 -based perovskite solar cell. These
findings demonstrate that the values of the electronic parameters decrease as the
temperature rises to 350 K. Increased carrier recombination, a larger density of
defects, modifications in electron and hole mobilities, and variations in layer thick-
ness that match the energy band gaps of the materials involved are some of the causes
of this performance decline.
408 S. M. Hasnain

Fig. 8 Graph showing the


performance of an optimized
solar cell across different
temperatures. The figure
consists of four subplots:
\\n\\n1. PCE % (purple line)
decreases as temperature
increases from 290K to
350K.\\n2. FF % (green line)
shows a slight increase with
temperature.\\n3. Jsc (mA/
cm2 , blue line) remains
relatively stable with a slight
increase.\\n4. Voc (V, red
line) decreases with rising
temperature.\\n\\nThe x-axis
represents temperature in
Kelvin, and each subplot has
its own y-axis for the
respective parameter

5 Conclusion

The efficiency and performance of perovskite solar cells have significantly increased
in recent years. In this research, the performance and power conversion efficiency
(PCE) of these solar cells, precisely arranged as Au/Spiro-OmeTAD/BaZrS3 /ZnO/
ITO, were evaluated in relation to the absorber layer composition. Our results show a
consistent pattern: solar cells with BaZrS3 perovskite absorber included layer contin-
uously performed better electrically. We methodically changed the structural layer
thicknesses to maximize this performance. The ideal ZnO, Spiro-OmeTAD, and
BaZrS3 absorber layer thicknesses for the BaZrS3 -based materials were found to be
0.035 µm, 0.22 µm, and 0.55 µm, respectively. A PCE of 8.54%, an open-circuit
voltage (Voc ) of 1.392 V, a short-circuit current density (Jsc ) of 14.019 mA/cm2 ,
and a fill factor (FF) of 44.2545% were all attained by the top-performing design,
Au/Spiro-OmeTAD/BaZrS3 /ZnO/ITO. Additionally, we looked at how temperature
affected these improved solar cells’ performance and found that lower temperatures
Computational Investigations of Inorganic Perovskite Absorber … 409

were positively correlated with higher PCE. Because absorbers with smaller energy
gaps typically produce superior electrical performance, this study emphasizes how
important it is to choose the appropriate absorber layer in perovskite solar cells.

References

1. Yoshikawa K, Kawasaki H, Yoshida W, Irie T, Konishi K, Nakano K, Uto T, Adachi D,


Kanematsu M, Uzu H (2017) Nat Energy 2:1
2. Jiang T, Wang Y, Meng D, Wu X, Wang J, Chen JJ (2014) Appl Surf Sci 311:602
3. Liu J, Chen K, Khan SA, Shabbir B, Zhang Y, Khan Q, Bao Q (2020) Nanotechnology
31:152002
4. Wang Y, Jiang T, Meng D, Yang J, Li Y, Ma Q, Han J (2014) Appl Surf Sci 317:414
5. Chen Q, Wang Y, Zheng M, Fang H, Meng X (2018) J Mater Sci: Mater Electron 29:19757
6. Hasnain SM, Qasim I, Iqbal A, Mir MA, Abu-Libdeh N (2024) Solar Energy 278:112788.
[Link]
7. Hasnain SM (2023) Sol Energy 262:111825
8. Kojima A, Teshima K, Shirai Y, Miyasaka T (2009) J Am Chem Soc 131:6050
9. Qasim I, Ahmad O, ul Abdin Z, Rashid A, Nasir MF, Malik MI, Rashid M, Hasnain SM (2022)
Solar Energy 237:52
10. Ke W, Kanatzidis MG (2019) Nat Commun 10:965
11. Noel NK, Stranks SD, Abate A, Wehrenfennig C, Guarnera S, Haghighirad A-A, Sadhanala
A, Eperon GE, Pathak SK, Johnston MB (2014) Energy Environ Sci 7:3061
12. Roknuzzaman M, Ostrikov K, Wang H, Du A, Tesfamichael T (2017) Sci Rep 7:14025
13. He T, Dong N, Yao Y, Xu F (2023) Chin J Chem Phys 36:477. [Link]
0068/cjcp2103054
14. Xu M, Sadeghi I, Ye K, Jaramillo R, LeBeau JM (2022) Microsc Microanal 28:2548. https://
[Link]/10.1017/s1431927622009722
15. Mumtaz M, Khan NA, Hasnain SM, Younis A (2011) J Supercond Novel Magn 24:1939.
[Link]
16. Rehman Sagar RU, Shabbir B, Hasnain SM, Mahmood N, Zeb MH, Shivananju BN, Ahmed T,
Qasim I, Malik MI, Khan Q, Shehzad K, Younis A, Bao Q, Zhang M (2020) Carbon 159:648.
[Link]
17. Yang Z, Zhu R, Zheng X, Pei Q, Tan J, Ye S (2023) Chin J Chem Phys 36:631. [Link]
10.1063/1674-0068/cjcp2310098
18. Nijamudheen A, Akimov AV (2018) J Phys Chem Lett 9:248. [Link]
lett.7b02589
19. Adachi S (2015) Earth-abundant materials for solar cells: Cu2 -II-IV-VI4 semiconductors. John
Wiley & Sons
20. Niu S, Milam-Guerrero J, Zhou Y, Ye K, Zhao B, Melot BC, Ravichandran J (2018) J Mater
Res 33:4135. [Link]
21. Huo Z, Wei S-H, Yin W-J (2018) J Phys D Appl Phys 51:474003. [Link]
6463/aae1ee
Digital Twins and Industrial Internet
of Things in Automobile Industry:
Applications and Future Trends

R. Sujatha , B. Uma Maheswari , J. Abishek, and Viswanath Ananth

Abstract Digital Twin (DT) and IIoT (Industrial Internet of Things) are signifi-
cantly changing the automotive industry landscape. This research paper discusses
the use of DT and IIoT technologies in the vehicle lifecycle. DTs as a virtual
replica of a vehicle allow engineers to perform digital simulations, and run virtual
crash and wind tunnel testing, streamlining the production process. Additionally,
IIoT enables automakers to collect and analyse data in real-time; all of which will
improve quality control, including predictive maintenance, maximizing cost effi-
ciency, training workers, and more. Combining the potential of DT and IIoT not only
makes developing autonomous vehicles practical, it enhances the customer experi-
ence. This paper aims to provide a definitive account of how automakers can use
DT and IIoT to boost efficiency and foster new products, services, and aftersales
innovation which will transform the automotive industry in the future.

Keywords Digital twin · Industrial internet of things · Automobile ·


Manufacturing

1 Introduction

Digital Twins are digital depictions of a system’s physical and operational features
that strive to construct a precise virtual model of the physical system to assess, refine,
and forecast its performance. These Twins encompass a physical constituent, a digital
model, and a linking network that connects the physical and digital domains, enabling
bidirectional interaction between them [1]. With the proliferation of the Internet of
Things (IoT), the number of connected devices is rising, from tiny sensing elements
to sophisticated actuators and appliances. Smart manufacturing enables many sensors
and actuators to be embedded into a manufacturing system to form a tightly inter-
connected network. The industrial Internet of Things (IIoT) is a composite phrase

R. Sujatha (B) · B. Uma Maheswari · J. Abishek · V. Ananth


PSG Institute of Management, PSG College of Technology, Coimbatore, India
e-mail: sujatha@[Link]

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 411
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
412 R. Sujatha et al.

describing the application of the IoT in industrial manufacturing, in which industrial


equipment in a production environment is connected through the Internet of Things.
These two networks are revolutionizing the industrial manufacturing systems.
This paper discusses the contributions of artificial intelligence (AI), machine
learning (ML), and big data analytics to shape these technologies. The paper mentions
how DTs are used to create new virtual car models, enabling engineers to simulate,
plan, and perform the testing of the vehicles, reducing operational expenses and
time. It discusses how IIoT is used to acquire spontaneous data and analyse it to
promote predictive maintenance, better quality control, and worker training. It offers
a thought-provoking overview of the current applications and explores the potential
research areas and future trends.

2 Literature Review

2.1 Digital Twins

DT refers to ‘an integrated metaphysics, multi-scale, probabilistic simulation of an as-


built vehicle or system that uses the best available physical models, sensor updates,
fleet history, etc., to mirror the life of its corresponding flying twin’. As a tool to
enhance the quality of the product being designed, to adjust or optimise processes, to
design shop floors and production flows most efficiently, and to anticipate or prevent
maintenance issues, the DT is gaining ground in every industry [2].

2.2 Industrial Internet of Things

IoT gained popularity in the late 1990s when radio-frequency identification (RFID)
was considered necessary. People thought that if all objects were equipped with
identifiers, computers could easily manage, identify, and count them. Other tech-
nologies such as Near Field Communication (NFC), barcodes, QR codes, and digital
watermarking were used for tagging and identification [3].
Physical Entity. The notion of physical entities embedded in the real world is the
most fundamental one here. DTs are often physical entities in the form of products,
machines, or systems that live in the physical world. “A physical entity can be a
simple component, a complex system, or even a whole process” [1].
Digital entity. The digital entity is the virtual entity, or ‘twin’, of the physical entity;
this is intricately tied to concepts of computer-aided design (CAD) and simulation.
Digital entity is a ‘virtual information construct’ that could represent the real entity
‘from macro-geometric detail to micro-atomic detail’.
Digital Twins and Industrial Internet of Things in Automobile Industry … 413

Physical environment. The physical environment of a DT encompasses the real-


world conditions that surround the physical entity, including the shape of a factory
floor, environmental conditions (temperature, humidity), and even other machinery
or people operating nearby. The physical environment can have a significant impact
on the performance and behaviour of the physical entity, so it is crucial to accurately
capture its characteristics in the DT [4].
Digital environment. The digital environment is the DT’s analogue to the physical
environment. It should approximate the physical environment as closely as possible
to allow the behaviour of the physical entity to be simulated and analysed [5].
Physical processes. Physical processes refer to actions and transformations involved
in the physical entity and the environment, such as production processes, machine
processing, chemical reactions, or any other physical processes involved in the phys-
ical entity. The most critical layer concerns physical processes that affect the physical
entity, which largely determines functioning and behaviour [6].
Digital processes. Digital processes are digital representations of physical processes
Digital processes are usually formulated based on mathematics-backed mathematical
models and simulations that are supposed to be analogous with physical processes
(second pair in counterparts) performed by the physical process that represents them.
Digital processes can explore the outcomes of scenarios and interventions before they
are tested in the real world through expensive and long-lasting physical experiments.
Parameters. These are the measurable characteristics or properties of the physical
entity and its environment that get quantified via the embedding. Parameterisations
are the basic building blocks of a DT. They constitute the basis for any numeric-
mathematical representation of the physical entity and its behaviour.
State. The state of the DT is the current values or condition of the parameters at the
time of measurement. It is the present status and functioning of the physical entity.
The most important aspect of the DT relates to capturing and representing the state
of the physical entity [5].
Fidelity. Fidelity refers to how accurately the DT models its physical twin. It uses
two factors, detailing and behaviour. High-fidelity DTs permit better simulations and
predictions which result in better set points and smarter decision-making, ultimately
leading to improved performance [1].
Physical-to-Digital. This is the channel through which data and information flow
from the physical entity and its environment into the DT, typically through sensors,
actuators, and communication networks. The physical-to-digital connection is the
lifeline of a DT, providing the real-time data that keeps the DT synchronized with
its physical counterpart.
Digital-to-physical. The digital-to-physical allows information and instructions to
flow from the DT back to the physical twin and its environment. This transmits
414 R. Sujatha et al.

commands to the actuators or modifies operational parameters. The digital-to-


physical link makes use of the insights gleaned in the DT to optimise and control the
real-world entity, closing the virtual-to-physical loop.
Twinning. The continuing process that keeps the DT in accordance with its physical
counterpart. It’s the bidirectional, real-time transmission of data and information
via the physical-to-digital and digital-to-physical pathways, keeping the DT in sync
with its physical counterpart, and keeping it alive and responsive to changes in the
physical world.
Twinning rate. This is the frequency at which the DT updates with data from the
physical entity. The twinning rate depends on the application and the characteristics
of the underlying physical processes. Higher twinning rates give a more respon-
sive and accurate, DT, but they also require more processing and communications
infrastructure.

3 Digital Twin Development Process

For a real-world product, creating an operating DT will usually follow six main
steps (see Fig. 1). These don’t have to be implemented in chronological order. But
sometimes, more than one step can be accomplished at the same time, and a few
developers even merge two of these steps into one.

Fig. 1 Digital twin


development process steps
Digital Twins and Industrial Internet of Things in Automobile Industry … 415

3.1 Digital Model Creation

The first step of DT creation to start with is a digital version of a physical product
using CAD/CAM (Computer-Aided Designing/Computer-Aided Modelling) and
3D modelling which is currently used in the product phase of design. The digital
model aims to capture three key aspects of the physical product namely elements,
behaviours, and rules. Specifically, the digital model focuses on the geometrical
dimensions and physical characteristics of the product, and the physical surround-
ings at the element level. At the behaviour level, it describes how users interact with
the product and how the product interacts within its working territory. At the rule
level, the digital model incorporates forecasting and optimisation models that are
developed based on the functioning rules of the product [7].

3.2 Data Processing

The information collected from the physical entity and other sources are checked and
compared, and the findings are compiled and presented meaningfully. It is essential
to compile the data and convert it into meaningful information so that developers
can it use to make informed decisions. As data is obtained from multiple sources,
it is crucial to integrate them to discover any underlying patterns that would be
undetectable when analysed individually. Finally, data visualization technologies are
used to represent the data through visual elements like graphs, charts and dashboards.
AI methods are also employed to enhance the cognitive ability of the DT, such as
reasoning, problem-solving and realisation abilities [8].

3.3 Virtual Simulation

Simulation or being in a layer of simulation with the use of Virtual Reality (VR)
technologies will let the team perform product behaviour on a digital territory. The DT
is generated by using different types of simulation techniques that digitally replicate
the key duties and responses of the physical product in itself. VR allows the architect
to build as well as the user to stay connected with the digital model. For the DT
application, there is plenty of hardware that facilitates VR products among the users
currently in use [9].
416 R. Sujatha et al.

3.4 Physical Environment Realisation

To execute the recommendations proposed by the DT, physical machines are fitted
with technologies such as sensing devices and actuators. The physical entity utilises
these sensors and actuators to measure and adjust its functions and behaviours in
the physical environment. The sensors are employed to measure the condition of the
physical unit and the environment, while the actuators are responsible for imple-
menting the adjustments suggested by the DT. Mechanical, pneumatic, electrical,
and hydraulic actuators are commonly used for this purpose. To enable end-users
to visualize how the final product will look in the physical environment, Computer
Graphics technologies can be utilized to project certain aspects of the digital entity
onto the physical environment.

3.5 Physical-Digital Connection

Technologies such as network communications, cloud computing and network


security are utilised by setting up bidirectional, physical, and digital connectivity.
Networking technologies including Bluetooth, Wi-Fi, barcode, QR code, Z Wave
can be used to transfer data from the physical unit to cloud in real-time, and thus
establish a live digital model. The technology of cloud computing can be used to
develop, deploy and maintain the digital model in the cloud environment so that the
digital model can be accessed by developers and users via internet connectivity at
anytime, anywhere in the world. There may be some sensitive user information when
data is passed on from the physical to the digital environment, certain technologies
must be adopted to guarantee data security. Blockchain technologies are used to
secure the data that is transmitted, stored, and retrieved from the cloud [10].

3.6 Data Collection

The data relating to the product that is analysed by the DT is classified into product
data, ecosystem data, consumer data and interactive data. Instant data can be collected
using various IoT technologies, while some data can be obtained from secondary
sources like product manuals and customer feedback. Consumer data typically
includes customer feedback, customer browsing and download records of the digital
model. Interactive data, on the other hand, pertains to how the product behaves with
users and the ecosystem such as the level of vibration and stress it generates. These
data are then incorporated into the digital model to create a closed-loop digital product
development process [2].
Digital Twins and Industrial Internet of Things in Automobile Industry … 417

4 Digital Twin in Automobile Industry

DT technology has affected every stage of the automobile’s lifecycle, from its
design and engineering to its manufacturing and operation and its maintenance.
Improving performance, preventing issues, performing remote diagnostics and
repairs, promoting efficiency, reducing costs and improving the end-user experi-
ence are all values imparted to an automobile by DT. DTs’ simulated equivalents run
parallel to the actual automobile through its lifecycle.

4.1 Automobile Design and Prototyping

The automotive industry uses DT technology to run digital simulations of the vehicles
that include conceptualisation, styling, engineering, rapid prototyping, and testing.
Automotive engineers are using DT technology to simulate various tests and assess
the structural integrity and aerodynamic performance of vehicles before physical
prototypes are produced. This technique not only enables them to reduce manufac-
turing cost and time and improve the fuel efficiency, safety, and customer satisfaction
of the vehicles but also empowers the engineers to create increasingly better motors
with a shorter development cycle within a smaller budget.
Crash Test. DTs are used to simulate car crash tests. Car engineers model a DT of the
car with the help of computer-aided-design software and a virtual representation of all
structural parts is present in the simulation. The simulation software models various
crash scenarios and different types of impacts, such as frontal and side-impact and
rollover. The simulation produces data on the structural state of the car like the body
deformation, component displacement and material stress and strain at a particular
loading condition. The engineers can analyse the crashes based on the simulated
test, to identify the performance in structural damage and also to identify measures
to reduce structural damage. Based on the result of the simulation, engineers modify
the structural part of the DT by tweaking and testing the DT’s design to enhance
its crash performance by changing the body panel’s thickness, reinforcing the area
under high stress, and modifying the car shape to decrease injury risk for passengers.
This can be accomplished by validating the modification against updated DTs [11].
Wind Tunnel Testing. DT technology in wind tunnel testing simulates the aerody-
namic performance of a vehicle, which allows for virtual wind tunnel testing. Engi-
neers use simulation software to create digital wind tunnels and generate airflow
around a vehicle in the digital world to perform virtual testing. To simulate airflow,
the software uses variables such as wind speed, direction, and turbulence. Engineers
use this data for operation, certification, and evaluation of all aspects of the vehicle
under multiple environments with different factors such as vehicle speed, environ-
mental effects and vehicle weights. Engineers examine airflow patterns along with
pressure distribution and force exerted on the vehicle. In addition, obtaining these
418 R. Sujatha et al.

simulations helps identify regions of the surface with high drag and turbulence and
gives information on the aerodynamic performance of the virtual vehicle. With this
information, engineers can make design modifications such as change the geometry
and size, change angles on mirrors, spoilers, and air drums, along with a variety of
other modifications to improve aerodynamics and ensure its energy efficiency [12].
Digital Prototype. DT technology is used for the virtual prototyping of vehicles
before it is used in physical prototyping. Engineers can create the exact DT of the
vehicle and test it many times in different scenarios before building a physical proto-
type of the car. Engineers create a DT of the vehicle that includes chassis, body,
engine and suspension parts, and each component’s information of material prop-
erties, geometry and behaviour using sophisticated computer-aided design (CAD)
software. The software can be used to simulate a vehicle’s performance in different
scenarios such as acceleration, braking, turning and handling by adding road condi-
tions, driver behaviour and environmental conditions. After generating results, the
engineers analyse the data on a car’s performance under different conditions by
what-if scenarios and make decisions to optimise the design and improve the car’s
performance [13].

4.2 Production Process

The process of producing automobiles involves a multitude of stages such as


assembly, testing and quality assurance. The technology of DTs has two applications,
assembly line optimisation and quality control.
Assembly Line Optimisation. DT is used to optimise assembly lines by locating
potential bottlenecks, improving process flows, and making the assembly line effi-
cient. Engineers model various components and processes that compose an assembly
line to optimise these processes, including the layout of machines and conveyors
and the location of workers. The DT is comprised of the material properties of
each part, their geometry, and their behaviour. The engineering team can simulate
various scenarios, such as changing the assembly line’s layout, modifying the speed of
machinery, and optimising the allocation of workers. The team can also adjust produc-
tion volumes to test the assembly line’s capacity for production [13]. This testing
can help engineers locate bottlenecks and improve processes, leading to shorter cycle
times and lower costs.
Quality Control. DT helps to control quality by simulating processes in advance to
identify problems and reduce development time and cost. DT can be built for a car
to simulate and test the car assembling process and production line. For instance, the
DT for the product assembling and production line can be produced using specialized
simulation software to record the movement of the cars. The factors that might affect
the production line can be simulated, such as the position of the car on the production
line, the sequence of assembling steps, and the movement of the components in the
Digital Twins and Industrial Internet of Things in Automobile Industry … 419

assembling process. By processing the data from the simulation and analysing every
individual component movement, engineers can find out the areas for improvement.
In the simulation process, the factors that might cause the waste such as error and non-
conform scale can be reduced. The sequence of production lines can be optimised
and adjusted to make the whole factory more efficient. Thus, the quality control
for the finished car can be accelerated and conform to the quality standard of the
manufacturer. This leads to fewer product recalls [14].
Workers Training. DT technology creates a virtual test environment for training
workers in automotive factories to improve the efficiency of task performance
and safety. Employees can use VR headsets or augmented reality (AR) glasses to
remotely enter the DT environment to practice performing jobs in different produc-
tion scenarios. This provides a safe, error-free learning environment without affecting
the normal production process. The trainees can get staff feedback and support from
the operator who has the adaptive simulation software. That can ensure a more precise
answer based on the implementation of tasks. It can evaluate trainees’ performance
as well. The DT technology can provide the best path for future training to improve
worker skills, reduce the risk of accidents and injuries, enhance production efficiency,
and minimise the time as well as expense of traditional training methods [15].

4.3 Predictive Maintenance

Predictive maintenance can be offered in a more scheduled and proactive way. There
is still an element of predictive regression, although the data processing is much
larger. It’s then turned into a machine-learning algorithm that can offer predictions
about when maintenance is needed [16].
Automobile factories. Sensors gather a huge amount of data on the operation of the
equipment. Predictive maintenance attempts to make use of the data by modelling
the patterns of the time series data to predict the equipment’s health status and when
it’s most likely to fail. Machine learning can help to set up modelling tools that
summarise the process of detecting patterns in the data and generating predictions.
For instance, by changing the frequency of preventive maintenance to more optimal
schedules, the company can dispatch technicians to perform maintenance before
failure. The system can be integrated into existing systems such as enterprise asset
management to make its use easier.
Vehicle maintenance. Vehicle predictive maintenance is where internal sensors
proactively monitor the disguised state of internal vehicle components and share their
condition and performance with machine-learning algorithms, which then predict
when maintenance is necessary, for example, low tyre pressure or the battery level.
For car owners, this can lower the incidence of scheduled maintenance and repair
420 R. Sujatha et al.

works, thereby saving time and money. Predictive maintenance systems can be devel-
oped to interact with internal computer systems that monitor and analyse informa-
tion from sensors to send out maintenance alerts to the user or dealers to schedule a
maintenance appointment as they have considerable time to track the car’s behaviour
[17].

4.4 Autonomous Vehicle

Self-driving cars or autonomous vehicles are racing at a remarkable speed in the


current world. However, developing them has a drawback to be reckoned with: these
vehicles are tested extensively and simulated before being rolled out for their opera-
tion, a process that is both long and costly. Owing to this, designers utilise DT, which
provides the means to create a virtual replica of autonomous vehicles and see the
vehicle ‘perform’ virtually in different environments before it’s ready to hit the road.
Sensor and Software Testing. Autonomous vehicles are highly dependent on sensor
and software systems for decision-making on road safety. DT technology can be
used to test and optimise sensors and software using a simulated DT of the vehicle.
A DT of the vehicle is created and then the operating conditions will be simu-
lated with the sensor and software settings to evaluate how the vehicle performs in
different scenarios. The simulated system can help to identify problems in the sensor
under different illumination conditions, such as in sunlight or rain. Different traffic
scenarios can also be simulated, such as scenarios of driving next to or near another
nearby autonomous vehicle. This data will help to optimise the sensor and software
settings for the safe and optimal operation of an autonomous vehicle on public roads
[18].
Training Autonomous Systems. Training the autopilot system of autonomous vehi-
cles to respond to road scenarios such as traffic lights, road signs and pedestrians is an
essential part of developing them. With DTs, we can train the system to recognise a
scenario and trigger a reaction within a millisecond within a simulated environment.

4.5 Customer Experience

The advantages of DT technology are vehicle customisation and personalisation.


Through DT technology, customers can customise a vehicle before it is even built.
They can select colours, materials, fabrics and features that would offer a highly
customised vehicle. This brings customers back to the showroom and makes them
feel more valued.
Virtual Showrooms and Test Drives. DT technology can be utilised for virtual
showrooms and test-drive opportunities. Customers can experience this through
Digital Twins and Industrial Internet of Things in Automobile Industry … 421

virtual reality. The customers will get practical knowledge of the look and feel
of the vehicles and choose the vehicle of their choice through virtual test drives.
The chances of making wrong decisions will reduce as customers will access and
see multiple vehicle models and brands. In addition, the DT technology can also
simulate the driving experience of the vehicle. The above-mentioned services can be
provided to customers at a lower cost and without maintaining a physical showroom
and vehicle inventory [19].
Remote assistance. DT technology enhances customer experience through remote
assistance. Customers can use their smartphones or other electronic devices to
connect with an automaker’s customer service desk. The customer service desk/
support team can use DT technology to diagnose the problem and offer technical
guidance in solving the issue remotely, without the need for the customer to bring a
vehicle for a service.

4.6 Continuous Design Improvement (PDCA)

With the assistance of DT and IIoT, sensors fitted to cars accumulate and analyse
instant data to inform manufacturers of the ways to make changes in design for their
next new car. This reinforced the application of PDCA (Plan Do Check and Act). DT
technology allows the manufacturer to generate a virtual system of the car that can be
used to simulate different conditions and scenarios. A DT can be tweaked and tuned
to modify the virtual state model, which can improve the performance of the phys-
ical car. For instance, an automotive battery manufacturer of an electric vehicle can
simulate, using in silico testing, its behaviour under various kinds of driving condi-
tions and charge/discharge cycles. Even issues in battery performance and design
can potentially be uncovered when thousands of simulated behaviours are observed.
The DT of the battery can then be tweaked to improve performance. The battery’s
thermal management system can be made better so that overheating of the battery
can be avoided. This is eventually translated into improved autonomy. The sensors
built into the car can provide instant information about the car’s performance, such
as velocity, acceleration, and fuel consumption. It can also monitor the performance
of individual components, such as the engines, transmissions, and low tyre pressure.
For example, the Tyre Pressure Monitoring System (TPMS) can monitor the inside
pressure of the car tyre and warn if the pressure is too low. The lane departure warning
sensor (LDWS) determines whether to have warning calls if leaving the lane line.
The DT is used to make constant improvements in PDCA cycle. During the Plan
phase, an engineer can leverage a DT not only to create virtual replicas that can
be tested with different design configurations but also to prepare the implementa-
tion of new sensors and collect data from those sensors. In the Do phase, physical
vehicles are built with on-board sensors, which collect performance data instan-
taneously. In the Check phase, engineers scrutinise the readings from the sensors
422 R. Sujatha et al.

to flag places in a vehicle’s design that can be improved. In the Act phase, engi-
neers, using the insights garnered via the Check phase’s data collection, take action
to improve vehicle design, e.g., by modifying the vehicle design, adding and/or
removing sensors, or implementing new maintenance procedures. Using the instant
data, automobile manufacturers can continuously improve the design and the experi-
ence of owning a vehicle, and deliver higher quality cars that reliably meet consumer
demands.

5 Use Cases

The automotive industry is foreseeing a lot of applications using DTs. Porsche Engi-
neering has developed a DT of car batteries for the electric vehicle segment. The
DT comprises a resistor–capacitor, electrochemical, and thermal model capable of
predicting the aging process in the batteries depending on the driving patterns. BMW
Group has developed a DT for its manufacturing plant. This helps the engineers to
interact and decide the components’ position and specification through real-time
collaboration. The factory utilises massive robotic equipment and using this omni-
verse platform, BMW placed the virtual arms of the robots in the digital environment.
Hyundai has implemented a DT that mimics its vehicle assembly floor to monitor
manufacturing product design and battery analysis.

6 Future Areas of Research and Trends

The combination of DT and IIoT technologies can transform the automotive industry;
their maturity and development will yield various automotive trends and bring new
opportunities in automotive manufacturing and mobility. A hyper-realistic DT is
capable of capturing the ever-finer details, including behaviours, of physical vehi-
cles, enabling engineers and designers to maximize performance, safety, and sustain-
ability. DTs when combined with AI will unlock the potential of predictive analytics.
The manufacturer can anticipate and prevent failures, program maintenance and opti-
mize operations. DTs can enable customers to have real-time experience through
virtual test drives, customised vehicle builds, performance analytics, predictive main-
tenance alerts, and more. Insights gained by the DTs and IIoT provide the founda-
tion for manufacturers to develop more sustainable and circular manufacturing. By
increasing utilization of the incoming materials, and decreasing waste, in the produc-
tion chain, it is possible to use fewer resources while creating the same output, and
simultaneously develop easy-to-recycle and reuse designs. The combined utiliza-
tion of DT technology and IIoT would be a transformative leap in the automotive
arena, enabling unparalleled possibilities of innovation, operational efficiency, and
sustainability in terms of how vehicles are designed and engineered.
Digital Twins and Industrial Internet of Things in Automobile Industry … 423

7 Conclusion

The emergence of the DT has transformed industrial operations by enabling the


instant monitoring, sensing, and control of complex systems worldwide. The bene-
fits of the DT have yielded greater operational efficiency, cost savings, and new
business models. The success of the DT is built upon the strength of enabling tech-
nologies where AI and ML, cloud and edge computing technologies play important
roles in analysing the sensor-based data from IIoT devices. These technologies are
driving a new industrial revolution characterised by data-driven decisions, predic-
tive maintenance and intelligent automation—with much more to come as industries
adopt these technologies for digitalisation. The challenges, meanwhile, need to be
addressed if these technologies are to continue to advance and emerge in winning
combinations.

References

1. Tao F, Cheng J, Qi Q, Zhang M, Zhang H, Sui F (2018) Digital twin-driven product design,
manufacturing and service with big data. Int J Adv Manuf Technol 94:3563–3576
2. Tao F, Sui F, Liu A, Qi Q, Zhang M, Song B et al (2019) Digital twin-driven product design
framework. Int J Prod Res 57(12):3935–3953
3. Serror M, Hack S, Henze M, Schuba M, Wehrle K (2020) Challenges and opportunities in
securing the industrial internet of things. IEEE Trans Industr Inf 17(5):2985–2996
4. Jones D, Snider C, Nassehi A, Yon J, Hicks B (2020) Characterising the digital twin: a
systematic literature review. CIRP J Manuf Sci Technol 29:36–52
5. Fuller A, Fan Z, Day C, Barlow C (2020) Digital twin: enabling technologies, challenges and
open research. IEEE Access 8:108952–108971
6. Uhlemann THJ, Schock C, Lehmann C, Freiberger S, Steinhilper R (2017) The digital twin:
demonstrating the potential of real time data acquisition in production systems. Procedia Manuf
9:113–120
7. Gao Y, Lv H, Hou Y, Liu J, Xu W (2019, May) Real-time modeling and simulation method of
digital twin production line. In: 2019 IEEE 8th joint international information technology and
artificial intelligence conference (ITAIC). IEEE, pp 1639–1642
8. Rasheed A, San O, Kvamsdal T (2020) Digital twin: values, challenges and enablers from a
modeling perspective. IEEE Access 8:21980–22012
9. Wagg DJ, Worden K, Barthorpe RJ, Gardner P (2020) Digital twins: state-of-the-art and future
directions for modeling and simulation in engineering dynamics applications. ASCE-ASME J
Risk Uncertain Eng Syst, Part B: Mech Eng 6(3):030901
10. Wu Y, Zhang K, Zhang Y (2021) Digital twin networks: a survey. IEEE Internet Things J
8(18):13789–13804
11. Korostelkin AA, Klyavin OI, Aleshin MV, Guodong W, Suifeng W, Jingyi L (2019)
Optimization of frame mass in crash testing of off-road vehicles. Russ Eng Res 39:1021–1028
12. Renganathan SA, Harada K, Mavris DN (2020) Aerodynamic data fusion toward the digital
twin paradigm. AIAA J 58(9):3902–3918
13. He B, Bai KJ (2021) Digital twin-based sustainable intelligent manufacturing: a review. Adv
Manuf 9(1):1–21
14. Biesinger F, Weyrich M (2019) The facets of digital twins in production and the automotive
industry. In: 23rd international conference on mechatronics technology (ICMT). IEEE, pp 1–6
424 R. Sujatha et al.

15. Kaarlela T, Pieskä S, Pitkäaho T (2020, September) Digital twin and virtual reality for
safety training. In: 2020 11th IEEE international conference on cognitive infocommunications
(CogInfoCom). IEEE, pp 000115–000120
16. Magargle R, Johnson L, Mandloi P, Davoudabadi P, Kesarkar O, Krishnaswamy S et al (2017,
May) A simulation-based digital twin for model-driven health monitoring and predictive
maintenance of an automotive braking system. In: Modelica, pp 35-46
17. Neupane S, Fernandez IA, Patterson W, Mittal S, Parmar M, Rahimi S (2023) Twinexplainer:
explaining predictions of an automotive digital twin. arXiv preprint arXiv:2302.00152
18. Wang SH, Tu CH, Juang JC (2022) Automatic traffic modelling for creating digital twins to
facilitate autonomous vehicle development. Connect Sci 34(1):1018–1037
19. Bhatti G, Mohan H, Singh RR (2021) Towards the future of smart electric vehicles: digital twin
technology. Renew Sustain Energy Rev 141:110801
Energy-Efficient Routing Algorithm
(EERA) Design for Heterogeneous IoT
Environment

K. Suresh Kumar , C. Rajendra Thilahar, and M. Vijay Anand

Abstract Wireless heterogeneous network is a main component in sensing data


for present IoT architecture. Wireless sensor network is the data source collector
for decision making and prognostic activity in IoT. The network has challenges
like energy and data reliability for accurate decision making. The energy-efficient
routing algorithm (EERA) algorithm is a solution for increasing lifetime as well as
gives high data reliability to the network through link stability approach. The CH
selection and routing mainly influences data reliability and network lifetime. The
EERA algorithm determines the better CH through residual energy, Link stability.
The EERA algorithm is compared with ISEP, PSO-C, ACO-Leach and LEACH
routing protocols. The EERA-algorithm performs more when compared to ISEP
and ACO-Leach protocol by 1.23 times and 1.31 times in terms of lifetime. The
unequal clustering mechanism exhibited by this algorithm provides equal load distri-
bution throughout the area of interest. The EERA proposed approach also provides
1.44 times improved throughput on comparing with LEACH routing protocol. The
algorithm also eliminates HOTSPOT issue and addresses energy hole issue in the
networks. This proposed algorithm is an efficient approach for energy hole and data
reliability issue in heterogeneous wireless network.

Keywords Wireless heterogeneous network · IoT · Markov · CH selection

K. Suresh Kumar (B)


Department of Information Technology, Saveetha Engineering College, Chennai, India
e-mail: sureshkumar@[Link]
C. R. Thilahar · M. V. Anand
Department of Computer Science and Engineering, Saveetha Engineering College, Chennai, India

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 425
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
426 K. Suresh Kumar et al.

1 Introduction

Internet of things environment constitutes to have diversified nodes which differs in


it capabilities [1, 2]. The Internet of Things is the modern data collection and prog-
nostics tool for making accurate decisions and actions on the smart city environment
[3]. The data collection for this IoT environment happens via hybrid wireless sensor
networks and wired networks. These data collecting networks are mainly powered
with small battery units and energy harvesting units [4, 5]. Wireless hybrid network
is mainly used for monitoring multiple different physical parameters to monitor the
data and take necessary actions [6]. Wireless sensor network comprises of sensing,
processing and transmitting unit, where these sensors and transmission capabilities
differs in hybrid IoT-based network [7, 8]. The capability of transmission and sensing
etc., varies in range based on their hardware models and battery operating parameters
[9]. In some cases the nodes are availed with heavy batteries for prolonged working of
nodes [10]. The maximum energy amount dissipated in the node is basically because
of communication (transceiving) unit and sensing unit [11–13]. The 90% and above
energy consumption is through node transmitting data to distant destination [14, 15].
In some cases the working of nodes are limited with very small batteries due to their
size constraints [16, 17]. The architecture of Hybrid WSN is elucidated in Fig. 1.
The networking architecture are layered and clustered normally for small scale and
large scale networks [18–20]. The large scale networks can cover in km, monitoring
farm lands and defense layers hill layers monitoring avalanche breakdown and forest
fires etc.
The clustered architecture constitutes to have Cluster member, Cluster Head and
Sleep states. The head the network lifetime is influenced by cluster head in many cases
[21]. The network CH selection is basically with random number and probability
based approach in Low Energy adaptive Clustered Hierarchical approach [22]. The
CH selection should be more effective and in many cases it considers many parame-
ters like, residual energy, link strength, node memory, computational capability etc.
[23–25]. Selection of node with reduced capability can result in frequent election and
unnecessary transmission of routing table and other energy-related data [26–28]. The
hybrid network in the IoT environment elucidates the capacity of power, memory,
transceiving capacity, and computational capacity as well. More algorithms have
evolved in such approaches [29, 30]. The hybrid wireless networks have heteroge-
neous environment where CH selection becomes more complicated and challenging
in nature [31, 32]. The complication on choosing the CH based on residual energy is
solved through considering the battery voltage parameter in Hybrid network along
with Markov probability for better CH election. The choice of battery potential
clearly indicates the node’s remaining energy. The residual energy identification is
done through voltage equation derived in this research article.
Energy-Efficient Routing Algorithm (EERA) Design for Heterogeneous … 427

Sink Cluster Cluster


Head Member

Fig. 1 Hybrid wireless sensor network architecture

2 Literature Survey

Many recent researches on energy-efficient hybrid IoT architectures have evolved


in the area of Wireless Sensor network. Wireless Sensor Networks are input to the
smart environments and smart industry environments. The Fail Safe Fault Tolerant
[9] algorithm for WBAN consists of electing CH based on node voltage and residual
energy. The energy of node can be computed with reference to the node voltage. The
CH selection based on the distance with the sink is computed through SEP an ISEP
protocols. Since CH majorly communicates large amount of data to the sink. Distance
with the CH to sink mainly influences the system. The battery recover based lifetime
Enhancement algorithm (BRLE) [10] approach provides efficient pause time for the
node to recover and make the node capable to efficiently utilize the energy stored
in the battery. The network with energy harvesting capability is deeply discussed in
EH-CHRS [11] algorithm. In all cases, energy efficiency is achieved through rotating
the CH role to a node. The role of the node mainly influences the network lifetime.
EECS [12] algorithm mainly concentrates on appointing new CH for mobile sink
based network infrastructure. The network with mobile sink has the confusion of
electing CH on the basis of distance from the sink. The sink mobility is modeled
through mobility models like random walk. The Energy aware buffer management
(EABM) [13] algorithm elects the CH based on the buffer and residual energy. The
N policy algorithm [14] EENP schedules the communication unit and transmits the
packets based on the n policy frameworks. This methodology however is unsuitable
428 K. Suresh Kumar et al.

Table 1 Assumed prelims for simulation


Metrics Values
Number of nodes 200
Area of network 500 × 500 m
E elec (J) 50 nJ/bit
E fs (J) 10 pJ/bit m2
Initial energy at 0th round 2 J (ordinary), 3 J (moderate), 5 J (high energy nodes)
Header byte size 50 bytes
Data message size 4000 bytes
Probability of becoming head 0.1

for time critical applications like forest fire and other critical monitoring events.
The TCEM-PQM algorithm for WBAN schedules the network based on the priority
queuing algorithm and enhances the network energy based on the requirement and
assigned priority. The TA-FSFT [15] algorithm chooses CH on the basis of energy
consumption and temperature of environment. The algorithm considers body temper-
ature and computational overhead to solve the CH selection methodology. Energy-
efficient CH selection scheme uses Artificial Bee Colony algorithm and GSA to
optimally select the CH. The algorithm considers energy, distance, load, delay and
node temperature as important parameters for CH selection. The hybrid optimization
algorithm [16] provides novel strategy to select the CH for green communication.
The energy harvesting and availability of energy to harvest is considered for CH
selection in the network. The head selection in each clusters along with the distance
between the sender and receiver are mainly considered for selection. Identifying the
battery residual energy is challenging and it is normally approximated with energy
consumption equations. The battery end voltage is used as an indicator to identify
the residual energy strength of the battery. The distance among the CHs and CH to
sink energy are also considered as the parameter to identify the transceiving node.
Table 1 elucidates the network routing algorithm for homogeneous network with
their significant contribution. The algorithms considered are mainly homogeneous
in nature and does not consider real time hybrid environment.

3 Energy-Efficient Routing Algorithm (EERA)

The energy efficiency algorithms in the existing routing algorithms concentrate


mainly on the energy spent by processing, transceiving units. In the present IoT
environment, the nodes are heterogeneous in nature and equipped with complex and
high energy sensing systems. The present sensing subsystem in new hybrid network
is not similar with classical system and the power consumed by such system should
Energy-Efficient Routing Algorithm (EERA) Design for Heterogeneous … 429

also be taken into account while CH selection procedure. The CH without consid-
ering this procedure results in biased CH selection and formation of Energy holes and
HOTSPOTS in the networks. The proposed algorithm gives importance on energy
expenditure, link stability (i.e.) the energy residual is calculated through the voltage
value of the battery and link stability is done through RSSI approach.

4 Mathematical Model

The battery potential difference parameter with respect to its diffusion parameter is
given in the following Eq. (1).

mt (expk(t2 −t1 ) − expk(t1 −t0 ) )
F(PD) = k
, t1
< t < t2 (1)
0, otherwise

where PD is potential difference


t 2 , t 1 , t 0 are slot n at nth slot, present time and initial point
k is the diffusion parameter in the battery.
The battery voltage parameter is mainly measured through the diffusion parameter
and depends on each battery. The extent of data transmitted by a node to the sink
over a time slot is computed through the following Eq. (2).

P(CH) × f1
NPacket =    (2)
number of CM × (1 − P(CH)) × f1

where
P(CM) (i.e.) {1 − P(CH)} is probability for a node to be a Cluster member
P(CH) is the probability for a node to be a Cluster Head
f is the crystal frequency of the node.
Equations (1) and (2) elucidates the true energy/ability of the node to actively
transmit the data to the sink. The potential node among heterogeneous power
capability will be identified through the equations.
430 K. Suresh Kumar et al.

5 Markov Approach

The Markov model is mainly for memory less approach. The election of the CH
is based on residual energy and link stability which is mainly on current state and
not on the earlier data. The proposed Markov model better suits the approach in
determining efficient CH.
The End steps Probability of choosing x state to y state is given in Eq. (3)

Pxy = Pr(Pn = y|P0 = x) (3)

The probability of single-step transition from x to k is given by

Pxk = Pr(P1 = q|P0 = x) (4)

For a time-homogeneous Markov chain:



Pr(Pn = y) = Py,r · Pr(Pn−1 = r) (5)
r∈S

Generalized probability of choosing r steps is



Pr(Pn = y) = Py,r · Pr(P0 = r) (6)
r∈S

The probability that a node in a system model will select the subsequent state is
represented by Eqs. (2), (3) and (4).
⎡ ⎤
P11 P12 P13
Pmatrix = ⎣ P21 P22 P23 ⎦ (7)
P31 P32 P33

The transition probability matrix is given in Eq. (7) (Fig. 2).

6 Radio Energy Model

Equations (8), (9) and (10) illustrates the energy expenditure for transmitting data
with distance less than d 0 and greater the d 0 . The distance between receiver and
sender mainly influences the energy expenditure for data transmission.

ETx (k, d ) = Eelec k + Efs kd 2 ; d < d0 (8)

ETx (k, d ) = Eelec k + Emp kd 4 ; d > d0 (9)


Energy-Efficient Routing Algorithm (EERA) Design for Heterogeneous … 431

Energy, Voltage

RSSI,

Location

Sink Distance

Energy Efficient routing algorithm Design for


Heteregenous network Energy
Link Stability Threshold
Markov model approach

CH probability

Fig. 2 Proposed energy-efficient routing design

ERx (k) = Eelec k (10)

where D: Distance,
K: No. of bit count,
E rx : Energy dissipated during receiving data,
E elec : Energy dissipated per bit to run the transmitter or the receiver circuit,
E fs (pJ/(bit°m−2 )), E mp (pJ/(bit°m−2 )): Energy dissipated per bit to run the transmit
amplifier based on the distance between the transmitter and receiver.
432 K. Suresh Kumar et al.

Algorithm:

Input: Distance, Sensing power, dlocal, dsink


Output: ProbCH
Abbreviation
VCH Present CH voltage
VT  Threshold voltage
Vi  Voltage value of participant
EE Energy expenditure
Begin process:
while( VCH< VT)
declare election;
identify the potential CH;
//receive election participation req:
if (Vparticipant<VT)
decline request;
//not eligible for CH;
else
for(i€p)
EE= radiomodel(Vi)
//compute the energy exp using radio
model
Markov(max(EE));
max(RSSI)
end
end
nominate CH;
announce new CH;
if no healthy participants retain present CH until time slot t;
end
End Process

7 Results and Discussion

The proposed energy-efficient approach is simulated and compared with LEACH and
ISEP protocol. The approach is simulated with 100 nodes with 10% high energy and
20% moderate energy and 70% ordinary energy nodes. Table 1 illustrates the network
prelims considered in simulating the approach using Matlab 2012. Initially algorithm
is compared with LEACH, MPSO ACO-LEACH, ISEP algorithms. This comparison
study is done with ISEP, ACO-LEACH bench mark algorithms. Initially the simu-
lation is started with 200 nodes deployed in random distribution. The simulation set
up is done with parameters as specified in Table 1.
Energy-Efficient Routing Algorithm (EERA) Design for Heterogeneous … 433

Fig. 3 Network lifetime comparison with ISEP and LEACH routing protocols

Figure 3 illustrates the network lifetime comparison of EERA algorithm with


ISEP, ACO-Leach and other routing protocols. The EERA algorithm outperforms
the existing ISEP, and ACO-Leach protocols. The proposed algorithm outperforms
ISEP by 1.23 times in case of ACO-Leach and 1.31 times routing protocols.
Figure 4 shows the throughput comparison of the EERA routing protocol. The
algorithm overtakes LEACH and ISEP by 1.4 times and 1.27 times. The algorithm
proposed also endorses high throughput with increased lifetime.
Figure 5 illustrates the network drop packets; the proposed routing algorithm
provides very low drop packets. The consideration of link stability, in the algorithm
greatly reduces the drop packets in the network.
Figure 6 illustrate the energy mesh diagram of the classic LEACH routing
protocol. The nodes nearby the sink consumes extra energy and energy hole is
created in the LEACH routing protocol. The energy hole in the algorithm makes
the network useless and cannot able to transmit the data to the sink successfully. The
sink is located outside the region of interest for better understanding of clustering
and energy dissipation.
Figure 7 shows the energy mesh diagram of EERA algorithm; the nodes nearby
the sink are still with high energy and clearly indicate the avoidance of the hotspot
and energy hole issue in the network.
Figure 8 illustrates the first dead node comparison diagram. The EEAR algorithm
withstands higher lifetime when compared ISEP and ACO-Leach protocols.
434 K. Suresh Kumar et al.

Fig. 4 Network throughput comparison with ISEP and LEACH routing protocols

Fig. 5 Network drop/loss packets comparison


Energy-Efficient Routing Algorithm (EERA) Design for Heterogeneous … 435

2.0
1.8
1.6
1.4
Energy in Joules

1.2
1.0
0.8
0.6
0.4
0.2 10
9
0.01 2
8
3 7
6
4

in m
in m
5

4
6

3
7

2
8
9

1
10

Fig. 6 Energy mesh diagram of Heterogeneous ISEP routing protocol

2.0
1.9
1.8
1.7
1.6
1.5
1.4
Energy in Joules

1.3
1.2
1.1
1.0
0.9
0.8
0.7
0.6
0.5
0.4
0.3
0.2 10
0.1
0.0 2
8

4 6
*10
m

m
*10

4
6

2
8

10

Fig. 7 Energy mesh diagram of EERA algorithm


436 K. Suresh Kumar et al.

Fig. 8 First dead node comparison of proposed with ISEP and LEACH protocols

8 Conclusion

The EERA algorithm works as an energy-efficient routing protocol for heterogeneous


network for IoT environment. The EERA algorithm overtakes the I-SEP and ACO-
Leach routing protocols in terms of Lifetime, throughput, First Dead Node etc. The
EERA algorithm offers 1.23 times improved lifespan and 1.4 improved throughputs
while compared with the I-SEP routing protocols. The EERA algorithm also avoids
HOTSPOT and Energy hole issue in the networks. The nodes that are nearby the
sink are loaded equally as with the nodes distant from the clusters. The future work
of this paper comprises scalability analysis and fault tolerant architecture design for
the proposed algorithm.

References

1. Kanagachidambaresan GR, Sarma Dhulipala VR, Vanusha, Udhaya MS (2011) Matlab based
modeling of body sensor network using ZigBee protocol. In: CIIT 2011, CCIS 250, pp 773–776
2. Akyildiz IF, Su W, Sankarasubramaniam Y, Cayirci E (2002) Wireless sensor networks: a
survey. Comput Netw
3. Kanagachidambaresan GR, Chitra A (2014) Fail safe fault tolerant mechanism for wireless
body sensor network (WBSN). Wirel Pers Commun 78(2)
4. Zhao H, Qin J, Hu JK (2013) An energy efficient key management scheme for body sensor
networks. IEEE Trans Parallel Distrib Syst 24(11)
Energy-Efficient Routing Algorithm (EERA) Design for Heterogeneous … 437

5. Javaid N, Abbas Z, Farid MS, Khan ZA, Alrajeh N (2013) M-Attempt a new energy-efficient
routing protocol in wireless body area sensor networks. In: Proceedings of the 4th international
conference on ambient systems, networks and technologies, vol 19, pp 224–231
6. Ahmad A, Javaid N, Qasim U, Ishfaq M, Khan ZA, Alghamdi TA (2014) RE-ATTEMPT: a
new energy-efficient routing protocol for wireless body sensor networks. Int J Distrib Sens
Netw 2014:1–9. [Link]
7. Md Abdur R, Hong CS, Lee S. Data-centric multiobjective Qos-aware routing protocol for
body sensor networks. Sensors 917–937. [Link]
8. Zhou G, Lu J, Wan CY, Yarvis M, Stankovic J (2008, April 13–18) BodyQoS: adaptive and
radio-agnostic QoS for body sensor networks. In: Proceedings of the 27th conference on
computer communications, Phoenix, AZ, USA, pp 565–573
9. Otal B, Verikoukis C, Alonso L (2009, June 14–18) Fuzzy-logic scheduling for highly reliable
and energy-efficient medical body sensor networks. In: Proceedings of IEEE international
conference on communications workshops, Dresden, Germany, pp 1–5
10. Natarajan A, Motani M, de Silva B, Yap K, Chua K (2007, June 11) Investigating network archi-
tectures for body sensor networks. In: Proceedings of the 1st ACM SIGMOBILE international
workshop on systems and networking support for healthcare and assisted living environments
(HealthNet 2007), San Juan, Puerto Rico, pp 19–24
11. Mahima V, Chitra A (2019) A novel energy harvesting: cluster head rotation scheme (EH-
CHRS) for green wireless sensor network (GWSN). Wirel Pers Commun 107(2):813–827
12. Saranya V, Shankar S, Kanagachidambaresan GR (2018) Energy efficient clustering scheme
(EECS) for wireless sensor network with mobile sink. Wirel Pers Commun 100(4):1553–1567
13. Jayarajan P, Kanagachidambaresan GR, Sundararajan TVP, Sakthipandi K, Maheswar R,
Karthikeyan A (2018) An energy-aware buffer management (EABM) routing protocol for
WSN. J Super Comput
14. Nageswari D, Maheswar R, Kanagachidambaresan GR. Performance analysis of cluster based
homogeneous sensor network using energy efficient N-policy (EENP) model. Clust Comput
22(5):12243–12250
15. Kanagachidambaresan GR, Chitra A. TA-FSFT thermal aware fail safe fault tolerant algorithm
for wireless body sensor networks. Wirel Pers Commun 90(4):1935–1950
16. Asiri M, Tarek S, Al-Awami L, Yasar A (2019) A novel approach for efficient management of
data lifespan of IoT devices. IEEE Internet Things J (Article in Press)
17. Otal B, Alonso L, Verikoukis C (2008, March 13–15) Novel QoS scheduling and energy-
saving MAC protocol for body sensor networks optimization. In: Proceedings of the ICST 3rd
international conference on body area networks, Tempe, AZ, USA, pp 1–4
18. Kuryloski P, Giani A, Giannantonio R, Gilani K, Gravina R, Seppa VP, Seto E, Shia V, Wang
C, Yan P, Yang AY, Hyttinen J, Sastry S, Wicker S, Bajcsy R (2009, June 3–5) DexterNet: an
open platform for heterogeneous body sensor networks and its applications. In: International
workshop on wearable and implantable body sensor networks, Berkeley, CA, USA, pp 92–97
19. Otal B, Alonso L, Verikoukis C (2009) Highly reliable energy-saving MAC for wireless body
sensor networks in healthcare systems. IEEE J Sel Ar Commun 27:553–565
20. Kanagachidambaresan GR, SarmaDhulipalab VR, Udhaya MS (2011) Markovian model based
trustworthy architecture. Procedia Eng
21. SarmaDhulipala VR, Kanagachidambaresan GR, Chandrasekaran RM (2012) Lack of power
avoidance: a fault classification based fault tolerant framework solution for lifetime enhance-
ment and reliable communication in wireless sensor network. Inf Technol J 11(6):103923
22. Ababneh N, Timmons N, Morrison J, Tracey D (2012) Energy balanced rate assignment and
routing protocol for body area networks. In: Proceedings of 26th international conference on
the advanced information networking and applications workshops (WAINA ’12). IEEE, pp
466–471
23. Ben Elhadj H, Chaari L, Kamoun L (2012) A survey of routing protocols in wireless body area
networks for healthcare applications. Int J E-Health Med Commun 3(2):118
24. Hughes L, Wang X, Chen T (2012) A review of protocol implementations and energy efficient
cross-layer design for wireless body area networks. Sensors 12(11):14730–14773
438 K. Suresh Kumar et al.

25. Hanson MA, Powell HC Jr, Barth AT et al (2009) Body area sensor networks: challenges and
opportunities. Computer 42(1):58–65
26. Khan N, Javaid N, Khan ZA, Jaffar M, Rafiq U, Bibi A (2012) Ubiquitous healthcare in wireless
body area networks. In: Proceedings of the IEEE 11th international conference on trust, security
and privacy in computing and communications (TrustCom ’12), IEEE, pp 1960–1967
27. Seo S-H, Gopalan SA, Chun S-M, Seok K-J, Nah J-W, Park J-T (2010, October) An energy-
efficient configuration management for multi-hop wireless body area networks. In: Proceedings
of the 3rd IEEE international conference on broadband network and multimedia technology
(IC-BNMT ’10), pp 1235–1239
28. Tang Q, Tummala N, Gupta SKS, Schwiebert L (2005, July) TARA: thermal-aware routing
algorithm for implanted sensor networks. In: Proceedings of the 1st IEEE international
conference on distributed computing in sensor systems (DCOSS ’05). Springer, pp 206–217
29. Quwaider M, Biswas S (2009, December) On-body packet routing algorithms for body sensor
networks. In: Proceedings of the 1st international conference on networks and communications
(NetCoM ’09), pp 171–177
30. Kim DY, Kim WY, Cho JS, Lee B (2010) Ear: an environment-adaptive routing algo-
rithm for WBANs. In: Proceedings of international symposium on medical information and
communication technology (ISMICT ’10)
31. Braem B, Latŕe B, Moerman I, Blondia C, Demeester P (2006, July) The wireless autonomous
spanning tree protocol for multihop wireless body area networks. In: Proceedings of the 3rd
annual international conference on mobile and ubiquitous systems (MobiQuitous ’06), IEEE
32. Annur R, Wattanamongkhol N, Nakpeerayuth S, Wuttisittikulkij L, Takada J-I (2011, March)
Applying the tree algorithm with prioritization for body area networks. In: Proceedings of the
10th international symposium on autonomous decentralized systems (ISADS ’11), pp 519–524
Enhancing the Performance of Smart
E-Bike Using Intelligent Systems

C. Suresh, N. Sujitha, P. K. Mani, S. Sengottaian, and K. Suresh Kumar

Abstract The shift to sustainable energy sources is urgently needed as increase


in the world’s energy needs i. Electric motorcycles are becoming more and more
popular in the automobile industry since they produce less air pollution, require
less maintenance, and produce less noise, especially when traveling short distances.
Range and charge time issues, however, continue to be major disadvantages. Despite
being useful for many years, electric bikes have not yet reached their full potential,
particularly considering how quickly technology is developing. The idea of using
energy to increase the range of electric bikes is examined in this essay. We can charge
the battery or meet other energy needs by harnessing dynamo power and transforming
it using electrical components. Any electric vehicle, regardless of motor rating, can
use this energy capture technique, which is a promising first step in increasing electric
transportation’s efficiency and range.

Keywords Controller · Battery · Rectifier · BLDC motor · Throttle · Electrical


brakes

C. Suresh (B)
EEE, School of Engineering and Technology, Dhanalakshmi Srinivasan University, Trichy, India
e-mail: sureshc28071983@[Link]
N. Sujitha
EEE, Prathyusha Engineering College, Chennai, India
P. K. Mani
EEE, P.T. Lee Chengalvaraya Naicker College of Engineering and Technology, Kancheepuram,
India
S. Sengottaian
EEE, Viswam Engineering College, Madanapalle, India
K. S. Kumar
IT, Saveetha Engineering College, Chennai, India
e-mail: sureshkumar@[Link]

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 439
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
440 C. Suresh et al.

1 Introduction

Existing electric bicycles have a rectifier for charging the battery, which is powered
by the mains. The bicycle cannot be used while the battery is charging from the main
supply, and it requires a constant main power supply. Our goal with this paper is to
introduce additional recharging mechanisms that can be used to reduce the battery’s
reliance on the main supply. The battery is charged with the help of two dynamos and
DC–DC converters. The e-bike will be propelled by a 250 W BLDC motor, which
will be fueled by a 24 V 12 Ah battery pack. The rider may choose whether or not
the bicycle will be totally driven by the motor. The maximum speed of the bicycle
is 20 km per hour, and we have altered the bicycle to accommodate the BLDC hub
motor with the wheel. The back wheel of the BLDC hub motor is equipped with a
wheel. Dynamos are linked to the tires on both sides of the back wheel. The paper is
a modest portion of an E-BIKE initiative to recover some energy from the E-BIKE
and extend its range. Dynamos, rectifier, Fixed DC–Variable DC buck converter, and
Fixed DC–Variable DC boost converter, are the main components of the circuit. The
dynamo produces AC voltage by receiving mechanical input. It was formerly used
by cycles to light the bulb in the evenings, and the same sort of dynamo will be
employed in the paper for AC voltage. The AC voltage can then be utilized as an
input for the Rectifier, which uses two parallel circuits to convert AC to DC voltage.
Here one side of the voltage will be sent to the boost converter to boost the voltage for
the requirement which will decrease the current, so that to compensate for the current
we are using another side of the circuit the buck converter is used here to stepdown
the voltage and which gives increase in current. This circuit can be installed in any
type of electric bikes. This circuit helps to fast charge the battery in the future electric
bike which equipped with fast charging batteries, which helps to increase the range
of electric bike. To create a straightforward, affordable electrical motorbike model
with a clever control system. The internal combustion engine, the exhaust system,
and other extraneous parts of the motorbike can be removed in order to replace them
with induction motor, monitoring devices, cable system, battery pack, and intelligent
controller [1]. A taxonomy of the many electric bicycle design options, allowing
designers to quickly get the knowledge they need to begin their papers [2]. Solar
energy is used to power a bicycle using a high-torque motor that is mounted on the
vehicle. The portable solar panel will produce electricity by absorbing sun energy.
If the power matches what is needed, the motor can utilize the power that the panel
has previously absorbed. Unless otherwise specified, a battery will be used to power
the motor [3]. An analytical solver that uses both mathematical calculations and
finite element method (FEM) analysis is used to attempt motor design and optimiza-
tion. Based on weight, torque output, effectiveness, and simplicity of manufacturing,
motor designs are evaluated [4]. A cyclist should be assisted by an electric motor in
areas with low efficiency while cycling alone in areas with high efficiency to account
for the capacity-dependent effectiveness of the rider. A Charge controller is created
with this purpose in mind [5]. In order to save energy, the controller will choose the
terrain’s grade and regulate the motor’s power. The console showed the estimated
Enhancing the Performance of Smart E-Bike Using Intelligent Systems 441

range of the acceleration, speed in rpm, e-bike, and distance traveled as well as the
location on a map and several throttle control settings [6]. Critical problems like
expensive drives and hefty, bulky batteries may often be resolved. The overview can
also be used to compare current drives in a complete, technically sound manner [7].
A costly undertaking that required spending a lot of time and labor in a machine
shop using a lathe. It seemed worthwhile to spend the money and time building my
own e-bike at the time because they were still very uncommon, at least outside of the
United States [8]. The battery charger should be lightweight, highly efficient, and
quick. A 24 V/3.8 Ah Lithium-ion battery may be completely charged by the charger
in 30 min [9]. The tailpipe effusions in ICE are one of the leading contributors to air
pollution. The carbon footprint and primary cost of EVs can be reduced by converting
existing ICE vehicles to Battery Electric Vehicles (BEV) rather than throwing them
out [10]. Proposes an automated stator winding connection technique for induction
generators that convert wind energy. Automates the stator winding switching process,
increasing wind energy systems’ dependability and efficiency. Ramachandran et al.
[11] conducts a comprehensive analysis of distribution networks to determine the
ideal sites for EV charging stations. determines the best pricing infrastructure allo-
cation strategies by analyzing factors including load demand and network stability.
Yuvaraj et al. [12] proposes an FPGA controller-based SSD technique to improve
the power quality of industrial SSDs. intends to reduce voltage fluctuations and
harmonics in order to increase performance and efficiency in industrial applications.
Kumar et al. [13] uses a Genetic Algorithm (GA) in conjunction with a PID controller
to improve the efficiency of bidirectional converters. Improves the energy efficiency
and control accuracy of bidirectional energy transfer systems [14]. Introduces the
wavelet-RBF technique for detecting abnormalities in the induction motor signal.
The detection of power quality and system reliability in industrial motor operations
is performed by multimodal signal processing [15]. An advanced hybrid controller
that uses machine learning and real-time data analysis to efficiently identify power
system defects. The intelligent algorithms enable prompt fault identification and
corrective action to maintain system stability [16]. Levy-based algorithms enable
seamless synchronization and regulate power fluctuations for seamless energy tran-
sitions in microgrids with many renewable energy sources. The approach enables the
efficient and less disruptive integration of renewable energy sources into the grid by
enhancing stability and ensuring reliable energy distribution [17]. A novel method
for calculating the average rotor slot size variation in induction motors that combines
variable Q-factor transformations with regression techniques. The accuracy of size
variation detection is increased by this technique, which is crucial for motor perfor-
mance [18]. This technology supports more effective motor performance diagnostics
and maintenance planning by improving the accuracy of slot size variation projections
when compared to traditional methods [19]. The results show that the method offers
better forecast accuracy for rotor slot width changes, which enables improved main-
tenance and performance assessment of induction motors, leading to more efficient
operation and less downtime [20–22].
442 C. Suresh et al.

2 Proposed Methodology

BLDCM is also called synchronous motor which runs by the DC voltage. The lack
of brushes and the use of electrical commutation set it apart from normal DC motors.
The motors are very effective for generating torque across various speed conditions.
Magnets circle around the armature is available in BLDC to overcome the linking
current in the armature, i.e., the rotor revolves around a fixed stator Construction and
operation of motors depend on the stator windings, although most typically BLDCM.
The windings in the BLDCM are carried by stacked steel laminations, which can be
current in star or delta configurations, although the most common configuration is a
three-phase star-linked stator. A permanent magnet serves as the rotor of a BLDCM.
The number of poles in the motor is varied from 2 to 8, with pole pairs alternated.
The perception of BLDCM is same as standard DC motor. The current-carrying
conductor experiences force and tends to rotate anytime it is put in a magnetic field.
Consider the coils A, B, and C, and the rotor is substituted with a single magnet.
When current is delivered to coil A, a magnetic field is formed, which attracts the
rotor, causing the rotor to move clockwise and align with A. When the current is
delivered to coil B at the same time, the rotor will align with coil B, causing the rotor
to revolve clockwise. Figure 1. Shows the brushless DC motor used in E-bike. The
BLDCM is used for Commercial robots, CNC machines, basic belt-driven machines,
Fans, pumps, and blowers. The specification of BLDCM is given below.
Specifications
• Output Power: 250 W
• Voltage: 24 V
• Speed: 3500 rpm.
Depending on whether you possess an adaptive or purpose-built electric bike, the
mechanism of an electric speed controller differs. The electric bike speed controller
transmits different voltages of signals to the bike’s motor. A purpose-built bike is

Fig. 1 Brushless DC motor


Enhancing the Performance of Smart E-Bike Using Intelligent Systems 443

Fig. 2 Controller

more expensive than an adapted bike, but it allows for faster acceleration and more
features. These signals identify a rotor’s orientation with respect to the starting coil.
A speed control must employ multiple methods in order to function well. A specially
designed electric bike’s rotor orientation can be detected with the use of Hall effect
sensors. In the event that your speed controller lacks such sensors, which the speed
controller on an adaptive bike might not, the electromotive force of the undriven
coil is computed to determine the rotor orientation. The analogue control systems
on this bike are the consequence of the requirement for the motor control parts to be
operated by the rider. Figure 2 depicts the controller we employed for this paper and
specifications are as follows.
Specifications
• Rated voltage for E-bike 24 V
• Rated power E-bike 250 W
• Maximum current E-bike 13.5 A.
The throttle mode works similarly to a motorbike. The motor produces power and
moves you and the bike ahead when the throttle is depressed. By restricting current
(using a throttle), the motor power can be raised or lowered, although it is typically
decreased. Informally, any system that regulates the motor power or speed is referred
to as a throttle. A throttle is found on a few e-bikes, which may evoke images of a
motorcycle’s twist grip, but is generally only a little electronic button. To accelerate
or continue onward, simply depress the throttle like you would the gas pedal on your
automobile. The Fig. 3 represents the throttle we used in this paper.
Figure 4 shows the Dynamo. A dynamo is a term that refers to an electric generator
that uses a commutator to produce direct current. A dynamo is a tiny electric generator
placed in the center of a cycle wheel to power the lights. These are always AC devices.
An armature, which is a collection of spinning windings that rotates inside the stator,
which is a stationary structure that produces a constant magnetic field, makes up
a dynamo machine. Lead that is porous or spongy is used to create the negative
electrode of a battery. Because of its permeability, lead can develop and evaporate
more easily. The positive electrode is composed of lead oxide. Both electrodes are
444 C. Suresh et al.

Fig. 3 Throttle

covered with an electrolytic solution made of sulfuric acid and water. A membrane
that is permeable but electrically insulating separates the two electrodes in the event
that they come into contact with one another due to physical movement of the battery
or variations in electrode thickness. This barrier also protects the electrolyte from
electrical shorts.
Plates, Separators, and Electrolytes make up a Lead Acid Battery, which is made
of hard plastic and rubber casing. Their batteries consist of “+ve” and “−ve” plates.
Lead oxide makes up the “+ve” plate, whereas sponge lead makes up the “−ve”
plate. These two plates are separated by an insulating material called a separator. An
electrolyte solution and a plastic shell contain the complete assembly. The electrolyte
solution is composed of sulfuric acid and H2 O.
Chemistry plays a major role in the lead acid battery’s operation, and under-
standing this is fascinating. The charging and draining of a lead acid battery involves
a significant amount of chemical activity. When the acid dissolves, the diluted sulfuric
acid molecules, H2 SO4 , split into two pieces. It will create “+ve” ions 2H+ and “−ve”
ions SO4 . Anode and cathode, as we previously stated, are joined as two plates. The

Fig. 4 Dynamo
Enhancing the Performance of Smart E-Bike Using Intelligent Systems 445

cathode attracts the positive ions, whereas the anode attracts the negative ions. The
Fig. 5 represents the typical lead acid battery we used.
The stepdown voltage is obtained by a DC–DC to buck converter that steps down
the voltage received from the input side. The DC–DC Buck Converter Stepdown
Module with Constant Current LED Driver Module of 300 W 20 A is utilized in this
experiment is depicted in Fig. 6.
• In addition to chargers for solar panels, wind turbines, and nickel–cadmium nickel
metal hydride batteries, there are also chargers for rechargeable lithium batteries
with voltages of 4, 6, 12, 14, and 24 V.
• The turn lights This version, which has a fixed 0.1 times (turning the lamp’s
current value is probably not very precise), is packed with charging instructions.

Negative
Terminal

Positive
Terminal

Fig. 5 Lead acid batteries

Heat Sink

Current
Transformer

Fig. 6 Fixed to DC-variable DC buck converter and stepdown module


446 C. Suresh et al.

• Current: current value * (0.1), turn the lamp current, and constant value linkage,
such as the constant value of 3 A.
The ripple’s content of 50 mV, 20 M of bandwidth measured I/O voltage is 24 V/
12 V. The current is regulated by switches, transistor, and diode, in the fundamental
function of the buck converter. When the switch and diode are turned on, they have
less voltage drop and low current flow, and the inductor has minimum series resis-
tance. The switch is in the on state when it is closed. As the current in the circuit
increases, the inductor creates an opposite voltage across the terminals due to change
in current. The voltage drops act in opposition to the source’s voltage, lowering the
voltage across the load. The voltage across the inductor lowers as the rate of change
of current falls, aggregating the voltage in the load. Figure 7 shows the 1200 W
High-Power DC to DC converter.
Rectifiers are machines that change power from AC to DC. Of all the rectifier
circuits, the bridge rectifier is the most effective. The KBPC3510 is a single-phase
bridge converter with a diffused junction. Figure 8 shows the bridge rectifier that we
used in the paper.

Fig. 7. 1200 W high power DC to DC

Load
Diode

Input AC Transformer
Voltage

Fig. 8 Bridge rectifier


Enhancing the Performance of Smart E-Bike Using Intelligent Systems 447

Features
• Intersection with a lot of noise.
• High effectiveness, low power loss, and low reverse leakage current, are all
advantages of this design.
• Metal Case with electrical Isolation for high heat dissipation.
• 2500 V isolation voltage.
We planned to utilize two dynamos in tandem, each connected to the wheel. The ac
output from the dynamos is subsequently rectified to dc voltage using abridge recti-
fier. One of the rectifiers is linked to the boost converter while the other is connected
to the buck converter in a continuous loop. The voltage increases as required by
using the boost converter, and the current is boosted as required by using the buck
converter. These are boosted in parallel, and when combined, we may achieve the
voltage and current as required. This power may be used for a variety of uses. One
of the primary goals is to charge the battery so that it can provide a boost for a few
more km. This procedure works for both high- and low-rated motors. It’s achiev-
able because the buck and boost converters may be adjusted to provide any desired
voltage and current. We intended to use two dynamos, one for each wheel. A bridge
rectifier the ac output from the dynamos to dc voltage. In a continuous loop, one of
the rectifiers is connected to the boost converter, and the other to the buck converter.
The boost converter boosts the voltage as needed, whereas the buck converter boosts
the current as needed. These are boosted in parallel, and by combining them, we
can reach the desired voltage and current. This ability may be used for a number of
different applications. One of the main objectives is to charge the battery so that it
can give enough power for a few more kilometers (Fig. 9).

Fig. 9 Block diagram


Dynamo 2 Dynamo1 Battery Power supply
for charging

Rectifier 2 Rectifier 1 Controller Throttle

Boost Buck Motor


Converter converter

Bicycle
LED Lamps Wheels

Brakes
448 C. Suresh et al.

3 Hardware Implementation and Discussion

As previously stated, when this type of buck and boost converters are installed in
an electric bike, the dynamo is coupled to the wheel for two sides, as shown in
Fig. 10, and the output voltage from the dynamos is connected to the buck and boost
converter, which sets the voltage and current rating as required by the electric bike.
To charge the battery, the buck and boost outputs are coupled in series (Fig. 11).

LED Lamps

Dynamo
Battery

Controller

Motor
Converter

Fig. 10 Final product of e-bike

Fig. 11 Front view of e-bike


Enhancing the Performance of Smart E-Bike Using Intelligent Systems 449

This increases the electric bike’s range or provides power to the bike’s lights
and infotainment High-rated electric bikes with motors rated at 750 W or 1200 W
perform more efficiently because they spin at a greater rpm, generating more power
for the electric bike. Installing this circuit in an electric bike is also cost-effective. If
the electricity is insufficient to charge the battery, the power can be sent to the electric
bike’s lights and display. The goal of this paper is to get the most power out of the
electric bike when it is running. This circuit assists in charging the battery while the
bike is running, which takes 2 h. The speed and voltage of an electric bike determine
its range. If the electric bike supports rapid charging, this may make it possible to
charge it faster. When the circuit has a rapid-charging electric bike, it operates more
efficiently. The circuit will assist electric firms in quickly charging their batteries,
similar to how mobile charging is evolving now.

4 Conclusion

Construction is simpler, and it may be utilized by schools, college students, office


employees, communities, and postmen, among others, for short-distance transporta-
tion. It is appropriate for both young and senior people. It’s possible to use it at no
cost. This bicycle is remarkable in that it does not rely on expensive fossil fuels,
saving billions of dollars. This electric bike takes very little maintenance. Because it
creates no emissions, it is ecologically beneficial, cost-effective, and pollution-free.
In the case of a power outage or cloudy weather, it is also silent and can be recharged
with an AC converter. This type of circuit is going to be implemented in the electric
bikes to increase range and the battery life time. This type of circuit is eco-friendly.
The controller can achieve various types of the voltages just by varying the variable
resistor in DC–DC converters. This type of controller can be used in any type of
E-BIKE. The battery consumption will be less so that the controller will have much
demand in the future to increase the range, and also there is no maintenance for the
circuit and it is easy to replace and cost is very much less. The scope of this type
of controller is good because the growth of electric vehicles is increasing so that the
controller has very good scope.

References

1. Reddy NPK, Prasanth KVSSV (2017) Next generation electric bike E-bike. In: 2017 IEEE
international conference on power, control, signals and instrumentation engineering (ICPCSI),
pp 2280–2085. [Link]
2. Dimitrov V (2018) Overview of the ways to design an electric bicycle. In: 2018 IX National
conference with international participation (ELECTRONICA), p 14. [Link]
ELECTRONICA.2018.8439456
450 C. Suresh et al.

3. Avhad SS, Tidke TP, Sathe NS (2017) Hybrid electric bicycle a new transportation for future
smart grid. In: SSRG Int J Electr Electron Eng 4(6):15–21. [Link]
IJEEE-V4I6P104
4. Ustun O, Tanc G, Kivanc OC, Tosun G (2016) In pursuit of proper BLDC motor design for
electric bicycles. In: 2016 XXII international conference on electrical machines (ICEM), pp
1808–1814. [Link]
5. Spagnol P, Corno M, Mura R, Savaresi SM (2013) Self-sustaining strategy for a hybrid electric
bike. In: 2013 American control conference, pp 3479–3484. [Link]
6580369
6. Abhilash DSH, Wani I, Joseph K, Jha R, Haneesh KM (2019) Power efficient e-bike with terrain
adaptive intelligence. In: 2019 International conference on communication and electronics
systems (ICCES), pp 1148–1153. [Link]
7. Muetze A, Tan YC (2005) Performance evaluation of electric bicycles. In: Fourtieth IAS annual
meeting. Conference record of the 2005 industry applications conference, vol 4, pp 2865–2872.
[Link]
8. Schneider D (2021) An instant e-bike: electrifying a bike can be electrifyingly easy. IEEE
Spectr 58(9):18–20. [Link]
9. Tsai C-C, Lin W-M, Lin C-H, Wu M-S (2010) Designing a fast battery charger for electric bikes.
In: 2010 International conference on system science and engineering, pp 385–389. [Link]
org/10.1109/ICSSE.2010.5551713
10. Kumar NA, Navaneeth M, Joseph AS (2021) Retrofitting of conventional two-wheelers to
electric two-wheelers. In: 2021 13th IEEE PES Asia Pacific power & energy engineering
conference (APPEEC), pp 1–8. [Link]
11. Ramachandran ER, Nadesan MK, Dhandapani L et al (2023) Design of automatic stator winding
connection of induction generator for wind energy conversion system. Multiscale Multidiscip
Model Exp Des. [Link]
12. Yuvaraj T, Devabalaji KR, Kumar JA, Thanikanti SB, Nwulu NI (2024) A Comprehensive
review and analysis of the allocation of electric vehicle charging stations in distribution
networks. IEEE Access 12:5404–5461. [Link]
13. Kumar JA, Selvi S, Rama RS et al (2024) Power quality improvement for industrial drives
using FPGA controller-based SSD algorithm. Multiscale Multidiscip Model Exp Des. https://
[Link]/10.1007/s41939-023-00347-6
14. Satish Kumar S, Bharathi K, Mohan SB, Suresh Balakrishnan T, Anish Kumar J, Sasi Kumar
M (2024) Effect of GA based PID controller in bidirectional converter. Int J Power Electron
Drives Syst 1757–766. [Link]
15. Kumar JA, Aswini KRN, Parimala V, Moan SB, Sharmila SL (2024) Multimodal analysis of
induction motor signals for power quality abnormality detection using wavelet-RBF approach.
In: Rathore VS, Tavares JMRS, Surendiran B, Yadav A (eds) Universal threats in expert appli-
cations and solutions. UNI-TEAS 2024. Lecture notes in networks and systems, vol 1006.
Springer, Singapore. [Link]
16. Suresh MP, Joyal Isac S, Joly M, Anish Kumar J (2025) Automatic fault detection and stability
management using intelligent hybrid controller. Electr Power Syst Res 238:111075. ISSN:
0378-7796. [Link]
17. Selvi S, Kumar JA, Joly M et al (2024) Levy based smooth synchronization of microgrid
integrated with multiple renewable sources. Electr Eng. [Link]
02494-6
18. Kumar JA, Swaroopan NMJ, Shanker NR (2022) Average rotor slot size variation measurement
in induction motor using variable Q-factor transforms and regression algorithms. Iran J Sci
Technol Trans Electr Eng 46:675–687. [Link]
19. Kumar JA, Swaroopan NMJ, Shanker NR (2023) Prediction of rotor slot size variations in
induction motor using polynomial chirplet transform and regression algorithms. Arab J Sci
Eng 48:6099–6109. [Link]
20. Anish Kumar J, Gowthambigai M, Shanker NR, Jasper J (2022) Prediction of rotor slot width
in induction motor using Dyadic wavelet transform and softmax regression. Int J Emerg Electr
Power Syst. [Link]
Enhancing the Performance of Smart E-Bike Using Intelligent Systems 451

21. Anish Kumar J, Jothi Swaroopan NM, Shanker NR (2022) Induction motor’s rotor slot variation
measurement using logistic regression. Automatika 63(2):288–302. [Link]
00051144.2022.2031541
22. Kumar JA, Gowthambigai M, Shanker NR et al (2023) Prediction of rotor slot size variation
through vibration signal of three phase induction motor using machine learning. J Vib Eng
Technol. [Link]
Nanowires in Biomedical Analysis:
A Strategic Review

Muhammad Azhar Ali Khan, M. Amin Mir, Syed M. Hasnain,


and Minakshi Memoria

Abstract The enormous potential of nanowires—nanostructures with infinite


dimensions—to transform biomedicine is examined in this chapter. Made by
mixing different materials using processes like vapor deposition and electrospinning,
nanowires have specific characteristics related to mechanics, optics, and electrics
capabilities. These qualities make them perfect for use in chemical sensors, drug
delivery systems, and other nanoscale electrical devices. Nanowires have exponen-
tial applications in cell engineering, brain-related ailments, and biomedical imaging,
even though the issues of biocompatibility and nanoscale production. Their numerous
uses could have a significant influence on the merging of healthcare and nanotech-
nology. If present barriers are removed, nanowires may reach their full potential and
lead to important breakthroughs in biosensors, pharmaceuticals, diagnostics, and
next-generation therapies.

Keywords Nanotechnology · Biomedical · Wires · Healthcare · Diagnostics

M. A. A. Khan (B) · M. A. Mir


Department of Mechanical Engineering, Prince Mohammad Bin Fahd University, Al Khobar,
Saudi Arabia
e-mail: mkhan@[Link]; mohdaminmir@[Link]
M. A. Mir
e-mail: mmir@[Link]
S. M. Hasnain
Department of Mathematics and Natural Sciences, Prince Mohammad Bin Fahd University, Al
Khobar, Saudi Arabia
e-mail: shasnain@[Link]
M. Memoria
College of Computer Sciences, King Khalid University, Abha, Saudi Arabia

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 453
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
454 M. A. A. Khan et al.

1 Introduction

There are numerous definitions and uses for nanotechnology. All definitions,
however, emphasize the creation and design of extremely ordered bottom-up nano-
structural substances that, in response to certain stimuli, provide distinct responses
[1]. Surface chemistry and physical properties refine the uses of substances with
nanoscale dimensions. These systems exhibit increased reactivity because the
number of atoms on their surface makes up as much as 90% of their total mass.
This means that altering a nanomaterial’s surface in various ways might result in
materials with unique biological characteristics, functions for a certain end use, and
enhanced solubility in physiological settings [2]. In nanobiotechnology, nanomate-
rials are used extensively, especially in prosthetics, implants, medication delivery
systems, and diagnosis [3]. Since the majority of biological systems are like-
wise nanoscale, nanoscale materials fit in well with biomedical equipment. These
nanotechnology products are often made from carbon nanotubes, liposomes, metallic
surfaces, inorganic and nanoparticles related to metal, and carbon nanotubes [4]. Bio-
specific molecules can be conjugated with nanoparticles using surface techniques
and the utilization of certain bio-events, like, antibody–antigen contact, receptor–
ligand connection, and DNA–DNA interaction. Surface thermodynamical properties,
surface physical properties, surface chemistry [5], and their harmful effects dictate
the particular application of nanosubstances.

2 Synthesis of Nanowires

Over the past 20 years, research on nanowires (NWs) has gained substantial atten-
tion in nanoscience and nanotechnology due to their remarkable performance in
a variety of applications and their distinctive features [6]. These one-dimensional
nanostructures have distinct optical, mechanical, electrical, and physical properties.
Their lengths can reach several micrometers, while their widths usually range from 1
to 100 nm. They are perfect for researching biological elements such as nucleotides,
amino-acids, viruses, and cells due to high aspect ratio and nanoscale size. The
ability of nanowires to conduct electrons efficiently along their length and to create
confinement effects across their diameter make them highly promising for advance-
ments in energy storage and other fields [7]. Nanowires can be composed of a variety
of materials, including metals, oxides, and organic molecules. Their synthesis and
functionalization are still being researched in order to customize their characteristics
for particular uses, which could result in important technical advancements.
Nanowires in Biomedical Analysis: A Strategic Review 455

3 Nanowire Uses in Therapeutics

As per the biomaterials, “drug delivery” corresponds to the application of materials


to deliver medicinal chemicals to specific body areas. These substances, whether
synthetic or natural, interact with biological systems to lessen side effects and increase
the efficacy of medications [8]. Drug carriers with regulated release rates for focused
and extended effect are one way to find a solution. Because of their small size,
which enables precise and regulated distribution, nanomaterials such as nanotubes,
nanowires, and nanoparticles present a viable alternative [9].
(a) Cancer Therapy
Because of their special qualities, nanowires are helping to advance cancer treat-
ment and overcome medication resistance, opening the door to new and creative
therapeutic approaches. They make targeted medicine delivery easier. One example
is the use of silicon nanowires (SiNWs), which increase treatment efficacy while
lowering side effects by precisely delivering pharmaceuticals to breast cancer cells.
By utilizing external magnetic fields to cure hyperthermia and administer therapeutic
agents precisely, magnetic nanowires composed of iron or cobalt enable precise
control over the release of therapeutic agents [10]. Because of their plasmonic quali-
ties, gold nanowires are employed in photothermal therapy, which absorbs light and
transforms it into heat to kill cancer cells in a variety of malignancies, including
breast and oral cancers. With the use of magnetic targeting and stimuli-responsive
drug release mechanisms, multifunctional nanosystems including nanowires inte-
grate drug delivery, diagnostic imaging, and targeted therapy. Due to their strong
surface area and biocompatibility, silicon nanowires can deliver targeted drugs to
particular cancer cells [11].
(b) Neurological Drug Delivery
With significant advantages for treating neurological and neurodegenerative
illnesses, nanowires are becoming a viable method for the transport of medications
to the nervous system. Because of their vast surface area and biocompatibility, they
improve the bioavailability and regulated release of neuroactive chemicals, which
enhances medication delivery. This is particularly useful when targeting specific
parts of the central nervous system (CNS), such as spinal cord injuries and neurode-
generative illnesses [12]. Additionally, because of their enhanced receptor binding
and decreased biodegradation, nanowires offer higher neuroprotection. They are also
utilized in devices that monitor and treat a variety of illnesses, such as diabetes, heart
problems, and neurological issues [13]. All things considered, neurological medica-
tion delivery has advanced significantly with the use of nanowires, opening up new
avenues for the focused treatment of intricate nervous system illnesses as shown in
Fig. 1.
(c) Diabetes Management
The application of nanotechnology, particularly in the form of nanowires, offers
potential improvements in glucose monitoring and medication administration for
456 M. A. A. Khan et al.

Fig. 1 Application of nanowires in neural interfaces [11]

the treatment of diabetes. With technologies like stretchy reservoirs and mechan-
otransductive films, nanowire-based systems provide alternatives to traditional injec-
tions by enabling targeted, regulated, and on-demand release of insulin. Because of
their excellent characteristics, small detection limits, and specificity, cobalt oxide
nanowires are used in enzyme-free glucose sensing and are hence perfect for contin-
uous in vivo glucose monitoring [14]. Furthermore, by utilizing supramolecular
interactions, gold nanowire (AuNW) motors in conjunction with mesoporous silica
segments that include pH-responsive nano valves allow for accurate insulin admin-
istration. Moreover, insulin loading capacity and transepithelial penetration are
improved by silicon nanowires coated on irregularly shaped silica particles, which
increases oral insulin bioavailability and delivery effectiveness [24].

4 Application of Nanowires in Technology

The phrase “medical imaging technology” corresponds to a range of methods for


observing the human body in order to identify, monitor, and treat medical disor-
ders. This category comprises non-invasive diagnostic techniques that offer insights
into workings of the body inside, such as MRIs, CT scans, ultrasounds, nuclear
medicine, and X-rays. By permitting early intervention without invasive proce-
dures and helping with early illness detection, these technologies improve patient
outcomes [15]. Both structural and functional imaging are provided by modalities
like as PET and fMRI, which provide information about blood flow, metabolism, and
physiological processes. Moreover, biomedical imaging directs surgical operations,
assisting doctors in identifying vital structures, locating malignancies, and making
sure medical instruments are positioned precisely.
Nanowires in Biomedical Analysis: A Strategic Review 457

(a) Biosensing and Diagnostics


In order to check many biological compounds for disease checking, diagnosis, and
analysis, modern healthcare heavily depends on biosensors and diagnosis. Silicon
nanowire field-effect transistors, which offer direct, label-free, real-time electrical
detection with exceptional sensitivity and selectivity, are the best example of this
approach. SiNW-FETs hold great promise in the advancement of biological sensors
due to their high yield and repeatability [16]. Notwithstanding certain difficulties,
like, sample solution ionic strength checking on sensitivity detection, SiNW-FETs
are anticipated to make a substantial contribution to blood testing and compatibility
studies, which include rapid illness diagnosis.
(b) Imaging Modalities
The terms “imaging modalities” refer to methods and tools for observing and taking
pictures of the inside workings and structures of living things or materials. They
are vital to many sectors, scientific research, and medical diagnosis. X-ray, CT,
MRI, PET, and SPECT are medical imaging detections that give knowledge on both
anatomy and function or molecular structure. Combining high-resolution anatom-
ical pictures with low-intension functional data is a developing trend in diagnostic
imaging that will help us comprehend diseases and treatments on a broader scale
[17].
(c) Magnetic Resonance Imaging (MRI)
A potent method for non-interfering, concurrent monitoring of dynamic processes in
living cells and animals is magnetic resonance imaging (MRI). Through improved
contrast, nanowires can improve MRI sensitivity and resolution. By acting as contrast
agents, they can be used to improve soft tissue resolution and resolve signal interfer-
ence in high-magnetic fields by highlighting particular tissues or cellular characteris-
tics. This is especially helpful in differentiating between normal and sick tissues, such
as the circulatory system and bone tissue. Labeled cells and core–shell nanowires
(NWs) have been uniformly distributed in the magnetic field of MRI, suggesting
that these materials may find use in a variety of therapeutic applications, such as
the treatment of cancer. High effectiveness of heating nanowires enable pH-change
drug location, magneto-mechanic impel, and photonic therapy. Functionalized NWs
are useful for energy therapy, magnetic guiding, and MRI monitoring because they
can target particular cells, such as leukemic cells [18]. Due to their adaptability,
NWs have the potential to be effective theragnostic agents, allowing for noninvasive
treatment monitoring.

5 Neuroscience Applications

Due to their ability to provide flexible substrates for high-resolution tissue activity
monitoring, nanowires are revolutionizing help in neuroscience. Thanks to their
exact set-up over size and shape, made possible by sophisticated architectures like
458 M. A. A. Khan et al.

3D U-shaped nanowire field-connected transistors, allow for repeatable intracel-


lular recordings from electrogenic cells. With its ability to provide insights into the
activity of neurons and cells at the nanoscale, this development has the potential
to improve neurosciences, diagnosis, and image-related technologies. The brain can
communicate directly with external devices through neural interfaces, also known as
brain-machine interfaces. They have a variety of uses, from improving human capa-
bilities to providing medical care. Because of their enormous surface area, nanowires
can be utilized as electrodes in these interfaces, improves signal quality and spatial
resolution in brain-machine interactions [19–22].
(a) Nanowires in Cell Engineering
By combining materials science with cell, and tissue engineering seeks to produce
substitute tissues or promote endogenous regeneration. Students are coming closer
to creating precisely designed tissues for pharmaceutical use as they work to create
biological replacements for broken organs and tissues. Nanowires are significant
in this field due to their small diameter, high aspect ratio, and enormous surface-
to-volume ratio. Top-down and bottom-up techniques are used in the creation of
nanowires, such as silicon nanowires (SiNW), with functional and biocompatible
qualities tuned. By influencing the microenvironment and encouraging cell prolifer-
ation, nanowires in cell engineering improve scaffold design, cellular guidance, and
tissue regeneration. In order to get beyond present constraints, more focus is being
placed on the development of biomimetic scaffolds [20] as shown in Fig. 2.
(b) Scaffold Design
Scaffolds in cell signaling provide three-dimensional support for tissue regeneration
and cell proliferation. Due to the electrical characteristics and nanostructure, silicon
nanowires are frequently engaged in these scaffolds. To replicate the architecture of

Fig. 2 Nanowire-assisted drug delivery [19]


Nanowires in Biomedical Analysis: A Strategic Review 459

bone and tissue, SiNWs are frequently coupled with biodegradable polymers. In order
to maximize cell adhesion and minimize cell mortality in bone tissue engineering,
two methods for incorporating SiNWs into polycaprolactone have been developed.
SiNWs have potential for cardiac tissue engineering as well. The processes used
to manufacture scaffolds include both conventional ones, such as electrospinning
and freeze drying, and modern ones, such as bioprinting and 3D printing, each with
unique benefits and drawbacks [23, 24].
(c) Freeze Drying
A method called “freeze drying,” sometimes known as “lyophilization,” uses low-
pressure freezing, casting, and preparation to dry polymeric liquids. Ice and unfrozen
water are eliminated by sublimation and desorption, resulting in scaffolds with pore
diameters within range 20–200 µm. Collagen-based carbon nanotubes are employed
in tissue signaling. After dissolving collagen in acetic acid, MWCNTs are added to
the thick solution. They spread into pliable, porous forms during freeze drying. The
appropriate amount of MWCNTs for a second additive, which is then combined
to ideal COL-MWCNTs solution at various percentages of weight, is estimated
mechanically [25].
(d) 3D Printing
3D printing is an innovative biomedical technology, especially in the areas of
tissue engineering and regenerative medicine. Despite its many advantages—such as
high yield, cheap cost, and simplicity—3D bioprinting is not without its problems,
including poor precision, exorbitant costs, and limited mechanical strength. The
utilization of gold nanowires (AuNWs) in cardiac tissue engineering is on the rise
because of its biocompatibility, chemical inertness, and conductivity. The conduc-
tivity, alignment, and signal propagation of nonconductive scaffolds are improved
by the addition of AuNWs. Extracellular matrix (ECM) from decellularized heart
can be hydrogelized and used as bio-inks in 3D printing. Cutting-edge methods such
as multi-photon polymerization and layer-by-layer building greatly increase scaffold
design, encouraging cell adhesion [26].
(e) Biometric
Growth parameters are used in tissue signaling to increase mineral formation and
cell migration. This helps with procedures like cell sheet transfer, which uses
PDL-derived cell sheets to promote bone and periodontal regeneration [27]. These
methods enhance periodontal health and lessen inflammation. By submerging them
in biological fluids, biomimetic processes can produce nano topography on 3D
scaffold parameters. Uniform inorganic crystal deposition is made possible by
the conformal-evaporated-film-by-rotation technology, which creates high-fidelity
copies of bio-templates with micro- and nano-level characteristics. This technique
preserves nanoscale surface topography in mineralized 3D scaffolds and enables
the deposition of hydroxyapatite (HA) crystals on patterned surfaces in bone tissue
engineering [28].
460 M. A. A. Khan et al.

6 Conclusion

With their previously unheard-of sensitivity, specificity, and adaptability, nanowires


have become a game-changer in the field of biological investigation. Because of
their special qualities, which include high surface-to-volume ratios and electrical
characteristics that can be adjusted, they are perfect for molecular illness diagnosis,
biomolecule detection, and cell activity monitoring. As this review demonstrates,
nanowires have the potential to completely transform healthcare due to their strategic
application in a variety of biomedical domains, including medication delivery, tissue
engineering, and biosensing. Notwithstanding the impressive advancements, there
are still difficulties, namely with large-scale manufacturing, biocompatibility, and
incorporating nanowires into current medical systems. In order to overcome these
obstacles and fully utilize the potential of nanowires in biomedical applications,
material scientists, biologists, and physicians must continue their interdisciplinary
research and collaboration.

References

1. Saji VS, Choe HC, Young KWK (2010) Nanotechnology in biomedical applications—a review.
Int J Nano Biomater 3:119–139
2. Gupta AK, Naregalkar RR (2007) Recent advances on surface engineering of magnetic iron
oxide nanoparticles and their biomedical applications. Nanomedicine 2:23–39
3. Amin MM, Mohammad WA, Andrews K (2021) Synthesis and the formation analysis of Ni
(II), Zn (II) and L-glutamine binary complexes in dimethylformamide-aqueous mixture. Results
Chem 3. [Link]
4. Liu D, Yang F, Xiong F, Gu N (2016) The smart drug delivery system and its clinical potential.
Theranostics 6:1306–1323
5. Amin Mir M (2022) Synthesis of oxadiazole, imidazole, benzimidazole, cyclohexano
analogues of 1,5-benzodiazepines through phenoxyl/phenylamino linkage. Curr Organocatal-
ysis 9(4):297–304
6. Inal O, Badilli U, Ozkan A, Mollarasouli F (2022) Bioactive hybrid nanowires for drug delivery.
Elsevier eBooks, pp 269–301
7. Amin Mir M (2022) Synthesis of green nano composite using sugar cane waste for the treatment
of Cr ions from waste water. In: Ujikawa K, Ishiwatari M, Hullebusch EV (eds) Environment
and sustainable development. ACESD. Environmental science and engineering. Springer
8. Jain KK (2019, August 22) An overview of drug delivery systems. Methods Mol Biol
9. Hsu C, Rheima AM, sabri Abbas Z (2023) Nanowires properties and applications: a review
study. South Afr J Chem Eng 46:286–311
10. Ames BN, Gold LS, Willett WC (1995) The causes and prevention of cancer. Proc Natl Acad
Sci USA 92(12):5258–5265
11. Elsayed KA, Alomari M, Drmosh Q, Manda AA, Haladu SA, Olanrewaju Alade I (2022)
Anticancer activity of TiO2 /Au nanocomposite prepared by laser ablation technique on breast
and cervical cancers. Opt Laser Technol 149:107828
12. Amin MM, Mohammad WA, Andrews K (2021) Synthesis and the formation analysis of Ni
(II), Zn (II) and L-glutamine binary complexes in dimethylformamide-aqueous mixture. Results
Chem. 3
13. Mir MA et al (2023) Determination of metals, fungi and mycotoxins in cat meal samples used
in Saudi Arabia. Biomed Pharmacol J 16(3)
Nanowires in Biomedical Analysis: A Strategic Review 461

14. Joshi HM, Bhumkar DR, Joshi K, Pokharkar V, Sastry M (2005) Gold nanoparticles as carriers
for efficient transmucosal insulin delivery. Langmuir 22(1):300–305
15. Amin MM, Waqar AM, Priya S (2021) Phytochemical isolation and anti-inflammatory
properties of various extracts of Lilium polyphyllum. Res J Pharm Technol 14(6):3195–3201
16. Martínez-Banderas AI, Aires A, Plaza-García S, Colás L, Moreno JA, Ravasi T, Merzaban
JS, Ramos-Cabrer P, Cortajarena AL, Kosel J (2020) Magnetic core–shell nanowires as MRI
contrast agents for cell tracking. J Nanobiotechnol
17. Li L (2013) Photoacoustic imaging. In: Pathobiology of human disease, pp 3912–3924
18. Amin Mir M (2022) Synthesis of green nano composite using sugar cane waste for the treat-
ment of Cr ions from waste water. In: Ujikawa K, Ishiwatari M, Hullebusch EV (eds) Environ-
ment and sustainable development. ACESD. Environmental science and engineering. Springer,
Singapore
19. Sam S, Joseph B, Thomas S (2023) Exploring the antimicrobial features of biomaterials for
biomedical applications. Results Eng 17:100979
20. Rani S, Memoria M, Singh R, Rathour N, Iqbal MI, Swami S (2024) Enhancing SME perfor-
mance through knowledge and government policy: a structural equation modeling analysis. J
Lifestyle SDGs Rev 4:e01587–e01587
21. Tripathi A, Sharma R, Memoria M, Joshi K, Diwakar M, Singh P (2021) A review analysis on
face recognition system with user interface system. J Phys: Conf Ser 1854(1). [Link]
10.1088/1742-6596/1854/1/012024
22. Kumar R, Memoria M (2020) Analysis of available selection techniques and recommendation
for memetic algorithm and its application to TSP. Int J Emerg Technol 11(2)
23. Mir MA, Memoria M, Kumar A, Prakash S, Bisht A, Dernayka S, Andrews K (2024, May)
Preparation, structure analysis and antibacterial properties of iron complexes mixed with 8-
hydroxyquinoline and some amino acids. In: 2023 International conference on smart devices
(ICSD). IEEE, pp 1–5
24. Roy R, Dixit AK, Saxena S, Memoria M (2024, May) Advanced digital technological solutions
to domestic violence: a brief analysis. In: 2023 International conference on smart devices
(ICSD). IEEE, pp 1–6
25. Zhu M, Li K, Zhu Y, Zhang J, Ye X (2015) 3D-printed hierarchical scaffold for local-
ized isoniazid/rifampin drug delivery and osteoarticular tuberculosis therapy. Acta Biomater
16:145–155
26. Mayoral H, Bevilacqua E, Gómez G, Hmadcha A, González-Loscertales I, Reina E, Sotelo
J, Domínguez A, Pérez-Alcántara P, Smani Y, González-Puertas P (2022) Tissue engineered
in-vitro vascular patch fabrication using hybrid 3D printing and electrospinning. Mater Today
Bio 14:100252
27. Pang Y, Sutoko S, Wang Z (2020) Organization of liver organoids using Raschig ring-like
micro-scaffolds and triple co-culture: toward modular assembly-based scalable liver tissue
engineering. Med Eng Phys 76:69–78
28. Rawlings AE, Bramble JP, Staniland SS (2012) Innovation through imitation: biomimetic,
bioinspired and biokleptic research. Soft Matter 8(25):6675
Performance Assessment of Organometal
Halide CH3 NH3 SnI3 in PSCs Using
SCAPS-1D Simulator
Syed M. Hasnain, Abid Iqbal, Irfan Qasim, M. Amin Mir,
and Minakshi Memoria

Abstract Perovskite solar cells (PSCs) have made significant advancements in effi-
ciency and technology, positioning them as a promising option for future photo-
voltaic systems. This study explores the effects of the absorbent layer composi-
tion on the power conversion efficiency (PCE) of PSCs structured as ITO/ZnO/
CH3 NH3 SnI3 /CuSCN/Au. Through the employment of the simulator tool, we metic-
ulously optimized the thicknesses of the structural layers to enhance device effi-
ciency. The perovskite solar cells based on CH3 NH3 SnI3 have demonstrated that
the ideal thicknesses are 0.03 µm for the ZnO layer and 0.8 µm for the perovskite
CH3 NH3 SnI3 absorber layer, resulting in a PCE of 20.17%, an open-circuit voltage
(Voc ) of 0.8378 V, a short-circuit current density (Jsc ) of 34.18 mA/cm2 and a fill
factor (FF) of 70.45%. Furthermore, this study underscores the significant impact
of temperature on solar cell performance, with lower temperatures enhancing PCE.
These results highlight the critical need of absorber layer selection in PSC design,
especially underlining the advantages of using materials with lower energy gaps for
enhanced electrical properties and general device performance.

Keywords Perovskites solar cells · CH3 NH3 SnI3 · SCAPS-1D simulator ·


Quantum efficiency · Materials

S. M. Hasnain (B) · M. A. Mir


Department of Mathematics and Natural Sciences, Prince Mohammad Bin Fahd University, Al
Khobar, Saudi Arabia
e-mail: shasnain@[Link]
A. Iqbal
Department of Computer Engineering, King Faisal University, Ahsa, Saudi Arabia
I. Qasim
Department of Physics, Faculty of Sciences, Rawalpindi Women University, Rawalpindi, Pakistan
M. Memoria
College of Computer Sciences, King Khalid University, Abha, Saudi Arabia

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 463
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
464 S. M. Hasnain et al.

1 Introduction

Global warming and energy issues are both very pertinent topics in today’s discourse.
The global use of non-traditional renewable energy sources has been necessitated
by the exhaustion of conventional resources and increase in the gas emissions which
leads to global warming. Energy from the sun will be able to fulfill the energy
requirements of the future as it is reliable and has a wide availability [1]. The ever-
growing need for photovoltaic devices with more advanced features has made solar
cell energy technology an important aspect in the generation of electrical energy
[2]. It is expected that in approximately the next 2 decades, solar energy can cater
for approximately 20% of the primary energy demand in the world [3]. During
the last decade, solar energy has gained in stature as a possible replacement of
traditional energy sources. Energy of the sun can be gathered from anywhere on the
planet at any particular time. These are some of the reasons why solar energy is of
critical importance as it does not directly pollute the environment or contribute to
the greenhouse gas emission.
Over the last three decades, researchers have concentrated their work on the study,
development, and improvement of new and more structured semiconductor-based
materials technologies. At present, silicon is the material to be considered for solar
photovoltaic devices and applications [4]. Researchers have been endeavoring to
improve the device’s efficacy for more than three decades [5]. Efficiency of greater
than 25% is achieved in a single-junction configuration with the introduction of
the optimal technology (photovoltaic). Perovskite solar cells (PSCs) have rapidly
advanced toward commercialization in just over a decade, thereby posing a threat
to conventional silicon solar cells [6]. The method has not been broadly adopted,
and its likeliness for future cost-cuts has been hindered by the higher expenses of the
semiconducting-based materials and the quality of silicon necessary for photovoltaic
applications. As a result, researchers are investigating several options. Innovative
openings for economical energy procurement have appeared via the use of perovskite-
structured compositions. The use of organic and inorganic perovskites materials in X-
ray detectors exemplifies a significant case. CH3 NH3 PbI3 perovskite-based thin films
and CH3 NH3 PbBr3 /Cs2 AgBiBr6 perovskite-based bulk-single crystals demonstrate
remarkable response. Perovskites have launched a new area of interest in renewable
energy compounds because of having active absorber layer materials in solar cell
systems. The many benefitable characteristics of these hybrid lead halides, including
their high absorption rate, facile preparation, tunable band-gap, higher diffusion
length, and solution fabricability, propose their possible use in solar cells.
The initial application of the proposed materials were reported in 2009 by Kojima
and his team, with an efficiency of 3.80% [6]. In the past decade, the rate of electricity
that has been converted into practical energy has increased above 25% [7]. Presence
of lead, a toxic substance, is a significant impediment to the commercialization of
perovskite. Potential substitutes for lead have been investigated by researchers, and
they have identified divalent metal cations such as Sn2+ and Ge2+ , which have an
oxidation state of (+2) and an outer shell similar to Pb2+ [8]. In theory, Sn-based
Performance Assessment of Organometal Halide CH3 NH3 SnI3 in PSCs … 465

perovskites are better than Pb-based perovskites due to their narrower band gaps
[9]. Especially, CH3 NH3 SnI3 has a small band gap of 1.35 eV that is suitable for
photolytic devices and applications. It has demonstrated favorable optoelectronic
properties, with an expected efficiency of exceeding 21% [10].
Nevertheless, many issues may impede the commercialization of PSCs, including
the device’s atmospheric stability under illumination and the hazardous characteris-
tics of the chemicals used in fabrication [11–13]. The combination of methylammo-
nium (CH3 NH3 ), formamidinium (CH(NH2 )2 ), and cesium (Cs+ ) in A-cation site,
lead (Pb+ ), (Ge+ ), and (Sn+ ) in B-cation site, and Iodine (I− ), Chlorine (Cl− ), Florine
(F− ) Bromine (Br− ) in X anion site in the A–B–X3 crystalline structure of perovskite
materials [10]. However, the most commonly used cation is lead in perovskite solar
cells, as its toxic nature associated with the lifecycle of PSCs poses a significant
threat to the surrounding environment. This commences the investigation of innova-
tive perovskite materials that are ecologically benign, free of lead, and exhibit great
efficiency [11, 14]. Researchers are now investigating these materials to enhance the
advancement of solar cell applications.
As an alternative to the toxic lead-containing perovskites, other lead-free alter-
natives may serve as absorber materials with a broad band gap, such as methylam-
monium tin iodide (CH3 NH3 SnI3 ), cesium tin iodide (CsSnI3 ), and formamidinium
tin iodide (CH(NH2 )2 SnI3 ). The direct band gap values were measured at 1.3 eV,
1.22 eV, and 1.41 eV, respectively. In the research field, Sb-based perovskite absorber
layer materials are deemed significant due to their features like those of Pb-based
halide perovskite, while achieving superior power conversion efficiency. The one
issue these gadgets encounter is their deterioration in the surrounding environment,
which requires attention [15].
In 2010, Neol et al. [16, 17] announced a lead-free CH3 NH3 SnI3 -based solar
cell with a PCE of 6%. He used the spin coating process inside a sealed inert gas
environment to prevent contamination to construct a device with an FTO/c-TiO2 /mp-
TiO2 /CH3 NH3 SnI3 /Spiro-OMeTAD/Au cell configuration. The presence of Sn2+ in
this material becomes more stable to Sn4+ owing to oxidation after exposure to the
ambient environment. This may result in the disruption of the charge carrier neutrality
of the chosen light-absorbing material and the production of SnO2 and CH3 NH3 I
[18]. The experimental evidence indicates that in CH(NH2 )2 SnI3 , the energy split-
ting was a significant influence of spin–orbit coupling compared to CH3 NH3 SnI3 ,
which was caused by hydrogen bonding [19]. In 2014, the power conversion efficien-
cies for CH3 NH3 Sn(I1−x Brx )3 and CH3 NH3 SnI3 were recorded at 5.73% and 6.4%,
respectively [20].
The efficiency of power conservation is significantly inferior to that of lead-based
solar cells, which also experience stability issues. To create PSCs that are more
ecologically favorable, efficient, and stable, a thorough investigation is necessary.
Recently, there has been a significant experimental advancement in perovskite solar
cells (PSCs) that are based on Sn. However, the potential for further development
exists through the meticulous adjustment of a variety of factors and device config-
urations, which could potentially aid research scientists in making more significant
experimental progress. The purpose of this study is to investigate the numerous
466 S. M. Hasnain et al.

characteristics of methylammonium tin iodide (CH3 NH3 SnI3 ) perovskite absorber


through the application of simulator software (SCAPS-1D). The research demon-
strated that the optimal performance of a device is contingent upon the optimization of
temperature and layer thickness, the minimization of interface defects, the regulation
of doping concentrations, the management of carrier generation and recombination,
the reduction of defect densities, and the adjustment of energy band gaps.

2 Objectives and Significance

This project is designed to optimize the performance of each layer by utilizing


CH3 NH3 SnI3 -based absorber material to develop and computationally analyze
perovskite solar cells. The focus is on the ITO/ZnO/CH3 NH3 SnI3 /CuSCN/Au config-
uration. To predict photovoltaic efficiency, comprehensive numerical simulations will
be implemented, with a particular emphasis on metrics such as short-circuit current
density, power conversion efficiency, open-circuit voltage, and fill factor. Following
this, the experimental fabrication of full-scale devices will be guided by these compu-
tational predictions, which will be evaluated under suitable illumination conditions.
There are numerous reasons why this research is important. It primarily addresses
the necessity for more environmentally friendly materials in perovskite solar tech-
nology, investigating non-lead alternatives that promote the creation of photovoltaic
devices that are less toxic and safer. Furthermore, the proposed configuration, which
incorporates active organometal halide-based perovskites, is a novel approach that
is designed to enhance the efficiency and stability of solar cells. The study’s findings
may pave the way for improved perovskite solar cell technology in the future and
more widespread use of renewable energy sources.

3 Proposed Design

To enhance the efficacy of solar cells, the proposed structure illustrated in Fig. 1
comprises ITO/ZnO/CH3 NH3 SnI3 /CuSCN/Au. This innovative design incorporates
Indium Tin Oxide as the top layer facing the sun.
The optimal values for each stratum of the structure were determined through
extensive research. The DFT analysis was done to determine the material character-
istics of each stratum which are listed in Table 1.
It is advisable to modify the thickness of a chosen layer to get the required ideal
thickness while maintaining constancy for the other layers. To determine the ideal
thickness of the layer, it is important to analyze the maximum limits and optimal
values of the electrical properties. The batch parameters were adjusted consistently
for each layer thickness. The operating method of the solar cell simulation software
SCAPS-1D has complied with specific requirements. An organized and methodical
procedure has been implemented to ensure accurate data measurements. Density
Performance Assessment of Organometal Halide CH3 NH3 SnI3 in PSCs … 467

Fig. 1 Body diagram


demonstrating the simulation
of a perovskite solar cell
made of organo-inorganic
perovskite materials

Table 1 Material properties of structural layers


Material properties CuSCN layer CH3 NH3 SnI3 layer ZnO layer
Optimized thickness 0.20 0.80 0.03
(μm)
Energy band-gap 3.40 1.30 3.30
(eV)
CB-DOS 2.2 × 10+19 1.0 × 10+18 2.2 × 10+18
VB-DOS 1.8 × 10+19 1.0 × 10+18 1.9 × 10+19
Electron/hole thermal 1.0 × 10+7 /1.0 × 10+7 1.0 × 10+7 /1.0 × 10+7 1.0 × 10+7 /1.0 × 10+7
velocity

functional theory (DFT) and current literature sources have been thoroughly reviewed
to obtain the energy bands corresponding to the valence and conduction bands of
each layer.
The electron transport layer comprises zinc oxide, while the absorber layers that
capture light consist of CH3 NH3 SnI3 . CuSCN serves as the hole transport layer,
and Au is utilized as the back-contact layer (refer to Fig. 1). The power conversion
efficiencies (PCE) of each layer were determined at their ideal thicknesses to achieve
optimal performance. The batch settings were adjusted to vary the thickness of each
layer, with the CH3 NH3 SnI3 layer reaching its ideal thickness first, resulting in the
highest PCE value. Supplementary information indicates that the thicknesses of the
ZnO layers and CH3 NH3 SnI3 perovskite absorbers remained constant, while the
CuSCN layer varied between 0.1 and 0.35 μm.
Figure 2 illustrates the schematic design of the energy band gap alignment and
presents the proposed solar cell structure. The energy bands of each layer were deter-
mined through a meticulous analysis of reliable experimental data, density functional
theory (DFT), and relevant literature sources. For specific numerical values, please
refer to Table 2.
468 S. M. Hasnain et al.

Fig. 2 Energy band gap


alignment for the proposed
optimal solar cell structure

4 Results and Discussion

The energy band diagram displayed in Fig. 3 illustrates the optimal configuration
for the perovskite ITO/ZnO/CH3 NH3 SnI3 /CuSCN/Au solar cell. The CH3 NH3 SnI3
absorber possesses an energy band gap of 1.30 eV. To achieve peak performance in the
proposed devices, the band gaps of ZnO and CuSCN were determined to be 3.30 eV
and 3.4 eV, respectively. ITO was utilized with a work function value of 4.3 eV,
while Gold/Rhodium (Au/Rh) metal contacts with a work function of 5.3 eV were
employed to optimize device performance. Band gap enhancement can be achieved
through the introduction of additives or doping, allowing for customization within
a specified range. The red line in the diagrams represents the valence band, while
the green line signifies the conduction band. These illustrations also showcase the
movement of mobile charge carriers. Electrons migrate from the n-type area to the
p-type region, while holes move from the p-type region to the n-type region.
The efficiency and output of a solar cell device are directly influenced by the
operating circumstances, as depicted in Fig. 4. The relationship between voltage and
total current density, referred to as Jt , shows a significant increase of approximately
0.7 V when utilizing alternate back-contact layers with the optimal thickness of the
structural layer. A visual representation of the recommended arrangement’s structural
layer has been included for reference. Our initial focus was on fine-tuning various
parameters such as thickness, work functions of contact layers, interface flaws, band
gap, and structural defects to achieve optimal performance values for our solar cell
design. Subsequently, these parameters were inputted, and the simulation was carried
out. The resulting numbers indicate the optimal performance for both absorbers
post-simulation. The values of Voc , Jsc , FF, and PCE for the solar cell based on
Performance Assessment of Organometal Halide CH3 NH3 SnI3 in PSCs … 469

Fig. 3 Optimized device


energy band diagram for
CH3 NH3 SnI3 -based
perovskite solar cell

ITO/ZnO/CH3 NH3 SnI3 /CuSCN/Au are 0.8378 V, 34.17687 mA/cm2 , 70.4485%,


and 20.1714%, respectively.
The efficiency of solar energy systems is significantly impacted by temperature.
Figure 5 illustrates the evaluation of the devices’ output performance in response
to temperature changes. The study analyzes the variations in the J-V curves across
a temperature range of 290–350 K. The inset in Fig. 5 highlights the shift from
lower current density to higher levels. Increasing the temperature has a noticeable
effect on the devices’ current density within a specific range. However, the current
density decreases at the optimal value due to lattice vibrations and collisions among
charge carriers. The decrease in open-circuit voltage with rising temperature can
be attributed to the reduction in lattice vibrations at lower temperatures and their

Fig. 4 Jt versus V plot of


optimized CH3 NH3 SnI3
absorber-based perovskite
solar cell
470 S. M. Hasnain et al.

Fig. 5
Temperature-dependent Jt
versus V plot of the
CH3 NH3 SnI3 optimized
solar

subsequent increase as temperature rises. At higher temperatures, atoms experi-


ence more pronounced vibrational movement, increasing collisions between charge
carriers (electrons and holes) and atoms. The current density is directly proportional
to the concentration of charge carriers and rises with temperature.
The relationship between the electrical properties of the perovskite absorber layer
and their thickness is illustrated in Fig. 6. It is observed that, except FF, Jsc , PCE,
and Voc tend to increase with layer thickness. This phenomenon can be attributed to
the fact that thickening the perovskite absorber layer may lead to an increase in the
density of charge carriers. As the thickness of the absorber layer increases, the Jsc
and FF show improvement, while the Voc decreases. Initially, the PCE experiences
a rise but then undergoes a significant drop. At the optimal thickness of 0.8 µm for
CH3 NH3 SnI3 , the Voc is measured at 1.0 V, Jsc at 34.69 mA/cm2 , FF at 63.86%,
and PCE at 19.19%. Augmenting the thickness of the absorber layer enhances light
absorption, resulting in elevated short-circuit current density and improved charge
carrier generation, hence enhancing efficiency. The improved electron mobility is a
direct result of enhanced electron–hole pair production, as evidenced by the higher
short-circuit current density.
Figure 7 illustrates the results of experiments conducted to evaluate the perfor-
mance of the devices under high temperatures, ranging from 290 to 350 K. This
temperature range was selected to showcase the global variations in electrical char-
acteristics. The perovskite solar cells utilizing CH3 NH3 SnI3 demonstrated a peak
PCE of 20.1714%, a Voc of 0.8378 V, a FF of 70.4485%, and a Jsc of 34.17687 mA/
cm2 . These results indicate that the Voc , Jsc , and PCE values decrease as the temper-
ature rises to 350 K. The decrease in performance can be attributed to a variety
of factors, such as fluctuations in the thickness of the layers corresponding to their
energy band gaps, higher defect density, variations in electron or hole mobilities, and
increased carrier recombination.
Performance Assessment of Organometal Halide CH3 NH3 SnI3 in PSCs … 471

Fig. 6 Effect of CH3 NH3 SnI3 layer thickness on electric parameters

Fig. 7 Temperature dependence output performance of CH3 NH3 SnI3 optimized solar cell

5 Conclusion

Perovskite solar cells have made significant strides in efficiency and technological
advancements in recent years. The purpose of this study was to determine how
the composition of the absorber layer affects the performance and power conver-
sion efficiency of the perovskite solar cells. The sequence of the layers of the
structural elements in the superlattice solar cell ITO/ZnO/CH3 NH3 SnI3 /CuSCN/
Au was varied systematically. The optimum values for the thickness of the PECs
based on CH3 NH3 SnI3 were found for the layers ZnO (0.03 µm) and CH3 NH3 SnI3
472 S. M. Hasnain et al.

absorber (0.8 µm). Out of all considered solar cells the best-performing cell, ITO/
ZnO/CH3 NH3 SnI3 /CuSCN/Au, was reported to have a maximum power conversion
efficiency (PCE) of 20.17%, the best efficiency value among all solar cells reported
to date. Under this configuration, the cell also gave a Voc of 0.8378 V, a Jsc of
34.18 mA/cm2 , and an FF of 70.45%. In addition, the study sought to demonstrate
the effect of temperature in the functioning of these newly developed solar cells and
pointed out that the reduced temperature enhances the PCE of these solar cells. This
research underscores the critical importance of selecting the appropriate absorber
layer in perovskite solar cells, as absorbers with narrower energy gaps demonstrated
superior electrical performance.

References

1. Yoshikawa K et al (2017) Silicon heterojunction solar cell with interdigitated back contacts for
a photoconversion efficiency over 26%. Nat Energy 2(5):1–8
2. Jiang T, Wang Y, Meng D, Wu X, Wang J, Chen JJ (2014) Controllable fabrication of CuO
nanostructure by hydrothermal method and its properties. Appl Surf Sci 311:602–608
3. Liu J et al (2020) Synthesis and optical applications of low dimensional metal-halide
perovskites. Nanotechnology 31(15):152002
4. Wang Y et al (2014) Fabrication of nanostructured CuO films by electrodeposition and their
photocatalytic properties. Appl Surf Sci 317:414–421
5. Chen Q, Wang Y, Zheng M, Fang H, Meng X (2018) Nanostructures confined self-assembled
in biomimetic nanochannels for enhancing the sensitivity of biological molecules response. J
Mater Sci: Mater Electron 29:19757–19767
6. Liu J et al (2019) Flexible, printable soft-X-ray detectors based on all-inorganic perovskite
quantum dots. Adv Mater 31(30):1901644
7. Zheng J et al (2019) Flexible photodetectors based on reticulated SWNT/perovskite quantum
dot heterostructures with ultrahigh durability. Nanoscale 11(16):8020–8026
8. Jeon NJ et al (2015) Compositional engineering of perovskite materials for high-performance
solar cells. Nature 517(7535):476–480
9. Hasnain SM, Qasim I, Iqbal A, Mir MA, Abu-Libdeh N (2024) Novel dual absorber configura-
tion for eco-friendly perovskite solar cells: design, numerical investigations and performance
of ITO-C60 -MASnI3 -RbGeI3 -Cu2 O-Au. Sol Energy 278:112788
10. Hasnain SM (2023) Examining the advances, obstacles, and achievements of tin-based
perovskite solar cells: a review. Sol Energy 262:111825
11. Singh D, Rana A, Memoria M, Shah SK, Gupta A, Kumar R (2023) Design of a 12 phase
VCO with stacked CNTFETs and provision to improve performance by lowering the number
of metallic tubes. AIP Conf Proc 2771(1). [Link]
12. Goel A, Awasthi M, Sajwan V, Kansal M, Agarwal S, Kumar R, Memoria M, Joshi K, Gupta
A (2023) An empirical assessment of key challenges influencing the MEI in current digi-
talized scenario. In: 2023 International conference on computational intelligence, communi-
cation technology and networking, CICTN 2023. [Link]
10140696
13. Rani S, Memoria M, Almogren A, Bharany S, Joshi K, Altameem A, Ur Rehman A, Hamam
H (2024) Deep learning to combat knee osteoarthritis and severity assessment by using CNN-
based classification. BMC Musculoskeletal Disord 25(1):817
14. Rani S, Memoria M, Singh R, Rathour N, Iqbal MI, Swami S (2024) Enhancing SME perfor-
mance through knowledge and government policy: a structural equation modeling analysis. J
Lifestyle SDGs Rev 4:e01587–e01587
Performance Assessment of Organometal Halide CH3 NH3 SnI3 in PSCs … 473

15. Qasim I et al (2022) Design and numerical investigations of eco-friendly, non-toxic (Au/
CuSCN/CH3 NH3 SnI3 /CdTe/ZnO/ITO) perovskite solar cell and module. Sol Energy 237:52–
61
16. Kumar MH et al (2014) Lead-free halide perovskite solar cells with high photocurrents realized
through vacancy modulation. Adv Mater 26(41):7122–7127
17. Noel NK et al (2014) Lead-free organic–inorganic tin halide perovskites for photovoltaic
applications. Energy Environ Sci 7(9):3061–3068
18. Ahmad O et al (2023) Modelling and numerical simulations of eco-friendly double absorber
solar cell “Spiro-OmeTAD/CIGS/MASnI3 /CdS/ZnO” and its PV-module. Org Electron
117:106781
19. Zitouni H, Tahiri N, El Bounagui O, Ez-Zahraouy H (2020) Electronic, optical and transport
properties of perovskite BaZrS3 compound doped with Se for photovoltaic applications. Chem
Phys 538:110923
20. Eya HI, Ntsoenzok E, Dzade NY (2020) First-principles investigation of the structural, elastic,
electronic, and optical properties of α- and β-SrZrS3 : implications for photovoltaic applications.
Materials 13(4):978
Predicting Delays in IT Projects:
A Machine Learning Approach

B. Uma Maheswari , D. Kavitha , R. Sujatha , and V. Santoshini

Abstract Project management aims at achieving project goals within the stipulated
time with available resources through various tools and techniques. Studies show
that majority of the projects were not successfully completed on time and within the
budget specified. The objective of this study was to create a delay prediction model
for Information Technology (IT) project using machine learning techniques. Data
was collected from 250 firms and as the variables identified through literature review
displayed multi-collinearity, principal component analysis was done for dimension-
ality reduction. The resulting six dimensions were used for further model building
using machine learning algorithms such as decision tree and random forest. The
study showed that random forest algorithm was the most accurate. The study also
showed that leadership, experience and communication among the team members
were important aspects which influences project delay.

Keywords Project management · Project delay · Delay prediction · Machine


learning · Decision tree · Random forest

1 Introduction

Project management is the effective management and implementation of tools and


techniques to achieve the project objectives within the specified time frame [1].
Global market place and tough competition mandates rapid, cost-effective and timely
delivery of projects [2]. However, in many organizations structured processes for
better project management have still not been effectively created and utilized. Owing
to the minimal success rate of project completions on time [3], this topic has attracted
the attention of many researchers across the world. A successful IT project could
be defined as the project that meets its target and is delivered on time [4] with
proper quality. These three dimensions are referred to as ‘Iron Triangle’ [5]. Timely

B. U. Maheswari (B) · D. Kavitha · R. Sujatha · V. Santoshini


PSG Institute of Management, PSG College of Technology, Coimbatore, India
e-mail: uma@[Link]

© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025 475
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
476 B. U. Maheswari et al.

delivery is crucial as it creates a competitive advantage for the client [6]. Project
delay impacts the success probability and outcome of the project’s product or service
which in turn impacts the success of the business [7]. Predicting delay in IT projects
is still a great challenge for project managers. The present study seeks to identify
and provide insights into the factors that cause delay in IT projects using machine
learning techniques.

2 Literature Review

According to Project Management Institute [8], the ideal phases of managing a


project includes initiation, planning, execution, review, control and finally project
completion. But in reality, a number of challenges disrupt the process flow. Many
studies have identified not only the critical factors for project success [9] but also
for project delay [10]. In the initiation stage, it is imperative to have clearly defined
project definitions and scope in terms of technical as well as non-technical aspects
as this impacts the project quality [11, 12]. Requirement gathering is extremely
important at the same time a challenging process. The estimation of the project
timeline should ideally happen once all the requirements are received [13]. In order
to have a clarity of the project at this stage, establishing a streamlined communication
mechanism with the clients and other stakeholders are extremely pertinent [3, 14].
In the project planning stage, the availability and allocation of resources based on
the requirements and timeline is crucial. Not adhering to the same leads to a deviation
in the estimated time line [15]. The issue could also arise because most of the projects
spend less time in defining the project. The environmental factors like time zone
differences [16], telecom infrastructure [17], physical security, visa regulations [18]
have a major influence in the project delay. The project team composition is another
important aspect of project management. Most of the multinational companies split
the development and testing phase to different vendors. In this scenario, the cultural
composition and dealing with the dependencies is a difficult task for the project
managers. Sometimes an ineffective team leads to project delay [19].
At the execution stage, the team member’s roles are assigned and resources are
allocated [20]. At this stage, tracking mechanisms are established to help monitor
progress and identify discrepancies as and when they crop up. Most of the IT
projects face resource scarcity during the operational phase, thus causing a delay
in the delivery of the project [21]. In some cases the limited resources would lead to
improper scheduling of the project [22]. The role of human resource management
in IT project management is inevitable [23]. Factors such as job design, person-
ality and the educational qualification, skill and experience of the employees affect
projects [4]. Patanakul [24] indicates that coordination and collaboration among the
team members is essential for managing large-scale projects. This can be accom-
plished through formalized business process and support from all the stakeholders.
Technology novelty is adapting to the new technology. A single project may have
multiple technology frameworks and multiple compliance process. Team members
Predicting Delays in IT Projects: A Machine Learning Approach 477

working in these projects should have both technical and compliance process knowl-
edge to complete it in the estimated timeline [25]. Most of the software projects
follow the ‘agile’ methodology [26] and adopt SCRUM practices [27]. Being agile
denotes the ability of the team to swiftly react to the changes in the needs of the client
and efficiently cater to them [28]. But many projects share resources among various
projects internally and therefore fail to be agile [29]. Despite this clarity in the early
stages, project managers receive requests to change the scope of the project in the
later stages of the project. These changes happen in three main categories such as
revisions in requirements, schedule and budget.
In the review and control stage, the progress of the project is compared with the
plan in order to ensure that all activities are progressing as planned. The deadlines
are tracked and monitored and at the same time, the budget limits are checked. Stress
is considered to be one of the major problems among the employees which is mainly
caused due to complexity in job and rapid growth of advanced technology. However,
stress would increase as the deadlines for the project delivery gets closer. Another
major concern is the case of self-committed errors by the project team member. When
such mistakes happen, it ideally needs to be reported and when it is not reported, it
automatically leads to project delay.
Finally, at the project completion stage, the team needs to identify the things which
went on well and which did not go well during the project and also identify areas
of improvement for the future. In all these stages, steady communication with the
clients and between team members is extremely important to ensure that everyone
is aligned with the project objective [30]. This would help identify inconsistencies,
address them and ensure that the same mistakes do not get repeated. Project managers
have to use several tools and techniques to help team members to work coherently
along a project life cycle. Appropriate use of project management tools influences
the timely completion of a project. It is also essential to have performance metrics
clearly defined for every stage of the project. The metrics along with clearly defined
roles and responsibilities for the team members would help avoid mistakes in the
forthcoming stages. Stakeholder management is another key factor for IT project
success which also leads to competitive advantage. In summary, Fig. 1 portrays the
different factors which could cause delay in each stage of the project management
phase according to the literature reviewed.
Artificial intelligence and machine learning algorithms in predicting the comple-
tion of IT projects have recently gained immense popularity [31]. Choetkiertikul et al.
[32] used network classification based on implicit and explicit task relationships and
extracted risk factors from open source projects and predicted delay in software
projects. de Bascelos Tronto et al. [33] implemented artificial neural network and
stepwise regression for predicting software project estimation and compared them
with traditional models such as APF, SLIM and COCOMO methods. The machine
learning techniques were mostly applied to publicly available open source data [34].
There are not many studies using primary data for predicting project delay. This
study is a step in that direction. The study aims to answer the following research
questions,
RQ1: Which of the factors have a critical impact on the project delay?
478 B. U. Maheswari et al.

Fig. 1 Project delay factors

RQ2: Given the factors, can we predict whether the project will be delayed or
not?

3 Materials and Methods

A structured questionnaire was constructed based on the variables identified in the


literature review. The focus of the study was to understand the factors that affect
project delay. The questionnaire was constructed using five-point Lickert scale. 1
represented ‘Strongly Disagree’ and 5 represented ‘Strongly Agree’. The question-
naire had four sections, the first part captured the information about the respondent.
Second part measured the twenty-one project delay factors which are the predictor
variables. The last part captured the response variable, whether the project is delayed
or not on a dichotomous scale. Tables 1 and 2 describe the questionnaire items.

Table 1 Respondent data


Respondent data
Designation Trainee/software analyst/manager/senior
manager/engineer/designer
Experience Less than 1 year, 2–4 years, 4–6 years, 6–8 years,
more than 8 years of experience
Sector Development/testing/infrastructure support/BPO/
core engineering/core designing
No. of projects worked 1/2/3/4/5/more than 5
No. of projects worked till completion phase 1/2/3/4/5/more than 5
Predicting Delays in IT Projects: A Machine Learning Approach 479

Table 2 Predictor and response variables


Predictor Variables
F1 Employee’s stress
F2 Revision of budget from the organization
F3 Improper resource allocation
F4 Incorrect requirement gathering
F5 Ineffective communication
F6 Technology novelty
F7 Lack of skilled resources
F8 Lack of agility/not following the SCRUM practices properly
F9 Inefficient knowledge transfer from the managers to the team members
F10 Scheduling the duration of the project without proper distribution of the activities
F11 Influence of culture composition of project teams
F12 Usage of poor project management tools
F13 Environmental reasons like time zone differences, telecom infrastructure, physical
security, visa regulations, custom and tax regulations
F14 External resources/dependencies
F15 Level of process compliance enforced on the project (e.g. multi-level of approvals
for each process)
F16 Inefficient knowledge transfer from the client
F17 Poor stakeholder management
F18 Revision of scope/requirement of the project frequently
F19 Revision of budget from the clients
F20 Revision of schedule from the clients
F21 Amount of time spent in the defining the phase and project duration by the client
3. Response variable
Project Yes—1/No—0
delay

3.1 Data

This descriptive research was conducted among IT firms in India. The respondents
were employees working in projects with the knowledge of different phases in the
project. A total of 250 questionnaires were sent to the firms. The responses were
obtained through electronic mails and other online survey tools. After regular follow-
ups, 201 responses were received with a response rate of 80.4%.
480 B. U. Maheswari et al.

3.2 Analytical Models Applied

In this study, two classification algorithms such as decision tree and random forest
are used to build model for IT project delay prediction. Decision tree is a machine
learning algorithm that predicts the response variable based on predictor variables.
It can be used for both regression and classification problems. Random forest is a
combination of trees. Each tree in the random forest is built on an independently
sampled subset of data. The distribution of data remains the same for all the trees.
Logistic regression is used to examine the association of categorical or continuous
independent variables with one dependent variable.

4 Analysis

4.1 Data Preprocessing Procedures

The collected data is preprocessed before model building. The variables identified in
literature as factors contributing to project delay were used along with the demo-
graphic variables like designation, experience, sector, projects worked, projects
completed. The response variable is project delay. The demographic categorical
variables were incorporated in the model using dummy coding technique. Principal
component analysis, a dimensionality reduction technique was applied to reduce the
factors. The eigenvalues which represents the largest variance summarized is used
as the basis for selecting the number of factors. According to Kaiser Normalization
Rule, eigenvalues greater than 1 were considered, which resulted into six factors.
The scree plot (Fig. 2) shows the eigenvalues of each component derived and also
the optimal number of factors or components.
The unrotated principal components analysis did not load the variables into appro-
priate factors. So varimax rotation was applied to identify the variables that loaded
into the six factors. The factor loadings and the factors are shown in Fig. 3. The factors

Fig. 2 Scree plot


Predicting Delays in IT Projects: A Machine Learning Approach 481

Fig. 3 Factor loadings and factors

were renamed based on the loaded variables. The factors and the corresponding
variables are shown in Table 3.

4.2 Model Building Using Decision Tree

A decision tree is a flowchart-like tree structure, used to visually and explicitly


represent decisions and decision-making. The top most node is called root node and
it is the important predictor variable. Each internal node tests a predictor variable,
each sub-tree represents an outcome of the decision. The decision tree was built using
Classification and Regression Technique (CART) with the control parameters and
the initial values for the control parameters (Table 4).
To avoid over-fitting of the model, the tree is pruned. The complexity parameter
of the tree was analyzed for pruning the tree. Based on the cross-validation error, the
complexity parameter value for the pruned tree was set at 0.042. The pruned tree is
shown in Fig. 4.
The pruned tree shows that number of projects worked by the team member, revi-
sions in the project, technology used, systems and process emerged as the important
variables deciding the delay in the project. In node 1 which is the root node, there are
482 B. U. Maheswari et al.

Table 3 Derived factors and variables


Rotated Factor name Variables
component
RC1 Revisions Revision of budget from the organization
Revision of scope/requirement of the project frequently
Revision of budget from the clients
Revision of schedule from the clients
RC2 Improper Improper resource allocation
planning Incorrect requirement gathering
Inefficient knowledge transfer from the managers to the team
members
Inefficient knowledge transfer from the client
RC3 Others Employee’s stress
Influence of culture composition of project teams
External resources/dependencies
Project duration
RC4 Technology Adoption of new technology
Agility/following the SCRUM practices properly
Usage of project management tools
RC5 Systems and Ineffective communication
process Lack of skilled resources
Improper scheduling of the project without proper
distribution of the activities
Multiple level of process compliance enforced on the project
Poor stakeholder management
RC6 Environmental Environmental reasons like time zone differences, telecom
reasons infrastructure, physical security, visa regulations

Table 4 Control parameters for decision trees


Parameter Value Explanation
minsplit 10 The minimum criteria for splitting the nodes. Only if the node has minimum
100 records, it will be split further
minbucket 4 The terminal nodes should have at least 10 records
cp 0 Complexity parameter initially set to 0 to allow for the entire tree to grow
xval 10 Tenfold cross validation

100% of the records, out of which 62% of the projects are labelled 1 (project delay)
and 38% as 0 (no project delay). The root node is split on the basis of number of
projects worked by the team member. If number of projects worked are less than 3,
then the second node is created in which there are 40% of the records, out of which
57% are classified as 0 (no delay) and 43% as 1 (delay). Hence, this factor becomes
the most important variable to help avoid project delay. Nodes 4, 7, 10, 11, 12, 13
are the terminal nodes. The terminal node 4 has 12% of the records, while terminal
node 12 has 5% of the records. As the proportion of the 1 s gets more the colour of
the node moves from dark green to blue. The nodes coloured in blue represent the
Predicting Delays in IT Projects: A Machine Learning Approach 483

Fig. 4 Pruned classification tree

Table 5 Conditions for no project delay


Node Variables Interpretation
number
4 • Projects When the person has worked on less than 3 projects, if there are
worked no revisions in the budget, scope of the project or the schedule
•• Revisions from within the organization or from the client, then there is no
delay in execution of the project
10 • Projects When the person has worked on less than 3 projects, if there are
worked no revisions in the budget, scope of the project or the schedule
•• Revisions from within the organization or from the client and the person has
•• Technology been updated with technology in terms of SCRUM, agility and

nodes labelled as 1 with maximum proportion of records in that node classified as 1.


The conditions and important variables for no delay in the project is given in Table 5.
The conditions causing delay in the project is given in Table 6.

4.3 Model Building Using Random Forest

Random forest algorithm creates a forest with a number of trees and combines their
output to improve generalization. A random forest model is built with project delay
as the response variable and all the other variables as predictor variables. The control
484 B. U. Maheswari et al.

Table 6 Conditions for project delay


Node Variables Interpretation
number
11 • Projects worked When the person has worked on more than 3 projects, if there
• Revisions are revisions in the budget, scope of the project or the schedule
• Technology from within the organization or from the client, and the person
has been not been updated with technology in terms of
SCRUM, agility and project management tools then there is
delay in execution of the project
13 • Projects worked When the person has worked on more than 3 projects, and the
• Technology person has been updated with technology in terms of SCRUM,
• Systems and agility and project management tools, but if there are issues in
process terms of systems and process with regard to communication,
resources, scheduling and stakeholder management then the
project is not delivered on time
7 • Projects worked When the person has worked on more than 3 projects and the
• Technology person has not been updated with technology in terms of
SCRUM, agility and project management tools, then the project
is delayed

parameters and the initial values for building random forest is given in Table 7. The
out of bag (OOB) estimate error rate is used for measuring the prediction error rate
of the random forest. The OOB error rate for the model is at 34.44%. The graphical
output for the OOB estimate of error rate is provided in Fig. 5. Random forest
is further tuned to decrease the OOB error rate. The model is tuned with control
parameters as given in Table 8.
Random forest computes two measures of variable importance such as mean
decrease in accuracy and mean decrease in Gini. Mean decrease in accuracy is based
on permutation and mean decrease in Gini is computed as total decrease in node
impurities from splitting on the variable, averaged over all trees. This model brings out
clearly that based on both the parameters the variables such as the number of projects
worked, revisions, other factors such as employee stress and culture composition
of team members, systems and process and environmental reasons had the highest
influence on project delay. The variable importance on the basis of mean decrease
accuracy and mean decrease Gini is given in Fig. 6.

Table 7 Control parameters for random forest


Parameter Value Explanation
ntree 501 Number of trees to be built
mtry 3 Number of variables randomly sampled as candidate at each split
nodesize 10 Minimum number of records in the terminal node
importance TRUE Should importance of predictors be assessed
Predicting Delays in IT Projects: A Machine Learning Approach 485

Fig. 5 Out of bag error rate

Table 8 Control parameters for tuning random forest


Parameter Value Explanation
mtryStart 3 Starting value of mtry
ntreeTry 324 No. of tress used for tuning
stepFactor 1.5 Steps to increase (deflate)mtry
Improve 0.0001 Improve the relative oob by atleast 0.0001
trace TRUE Print the trace or not
Plot TRUE Plot oob versus mtry graph or not
doBest TRUE Finally build the RF using optimal mtry
Nodesize 10 Minimum terminal node size
Importance TRUE Compute variable importance or not

Fig. 6 Variable importance


plot

5 Results

Three modelling techniques were used to identify the variables which have the
maximum influence on project delay. Table 9 shows the performance metrics of
the three classification models logistic regression, decision tree and random forest.
The performance metrics used for evaluating the models were accuracy, sensitivity,
486 B. U. Maheswari et al.

Table 9 Model performance


Performance measures Decision tree (%) Random forest (%)
measures
Accuracy 79 83
Sensitivity 60 65
Specificity 91 93
AUC 62 68
KS 14 16
Gini 16 24

specificity, AUC, KS and Gini [35]. The results indicate that random forest is the
best model which could predict project delay.

6 Discussion and Conclusion

Project management literature reflects mostly on the tools and techniques, but when
it ultimately comes to the successful completion of the project, it is not just the
technical aspects of project management, but the behavioural aspects which have a
dominating influence on the targets being achieved on time. The results of this study
reiterate the impact of the behavioural aspect on the delay in projects. The study
showed that one of the most important factors which could influence a project to be
delayed are the number of projects the person has worked on before. That experience
helps the person complete the projects on time.
The other important factor which influences the delay in the project is the revi-
sions which happen in the schedule. Sometimes these revisions happen within the
organization or from the client’s side. The revision could also happen in the budget
or the schedule. Such revisions cause interruptions in the project and cause delay
in completion of the same. This is where the role of the leaders come into play.
Literature on leadership theories clearly states the role of active leadership versus
passive leadership. The role of a team leader in the context of project management
would include clarity in identifying the needs of the team, clarity in laying out the
structure of the project, continuous monitoring and structured follow up. The need
identification aspect illustrates the relationship-oriented traits of a good leader. And
a good leader understands the requirements of his team and takes care of the same.
Revisions or deviations might happen because the leader is not actively involved in
goal setting, continuous project monitoring, identifying bottle-necks and loopholes
well in advance in order to avoid any discrepancies in the future.
In the practical sense, all projects cannot be planned and executed seamlessly. But
if the organization could avoid such revisions in between or could be prepared with
an extensive plan in the initial stage itself, such stoppages could be avoided. Another
factor which seems to play a major role is the technology aspect. It is highly essen-
tial for organizations to not only invest in technology but also ensure that adequate
Predicting Delays in IT Projects: A Machine Learning Approach 487

training is provided for the individuals. Creating an agile work place, ensuring
SCRUM practices and implementation of project management tools would help
organizations in successful completion of projects without major delays. Another
significant factor is the systems and processes. This includes creating an effective
communication flow in the organization, ensuring the availability of skilled resources,
proper scheduling of activities and stakeholder management. The need to generate
various channels of communication and streamline the same is very crucial for the
project success. Pro-active leadership could also help mitigate loss by communicating
frequently with the teams. This could provide an essential stimulus to identify and
retain a motivated team with a clear focus of the project goals. Another aspect which
the models highlighted is the levels of process compliance. The more the levels, the
higher the chances of project delay. The findings of this study indicate the need for
proactive leadership, clear goal setting and most importantly frequent communica-
tion and follow up through the entire project planning and execution phase. Giving
importance to these factors in an organization could avoid project delay and thereby
prove profitable to the organization.

References

1. Santos JI, Pereda M, Ahedo V, Galán JM (2023) Explainable machine learning for project
management control. Comput Ind Eng 180:109261
2. Filippetto AS, Lima R, Barbosa JLV (2021) A risk prediction model for software project
management based on similarity analysis of context histories. Inf Softw Technol 131:106497
3. Sudhakar GP (2013) A review of critical success factors for offshore software development
projects. Organizacija 46(6):282–296
4. Ibraigheeth M, Fadzli SA (2019) Core factors for software projects success. JOIV: Int J Inform
Vis 3(1):69–74
5. Radujković M, Sjekavica M (2017) Project management success factors. Procedia Eng
196:607–615
6. Choo AS (2014) Defining problems fast and slow: the u-shaped effect of problem definition
time on project duration. Prod Oper Manag 23(8):1462–1479
7. Sharma SK, Chanda, U (2017) Developing a Bayesian belief network model for prediction of
R&D project success. J Manag Anal 4(3):321–344. [Link]
1304291
8. Singh H, Williams PS (2020) A guide to the project management body of knowledge: PMBOK
(®) guide. Project Management Institute
9. Chow T, Cao DB (2008) A survey study of critical success factors in agile software projects. J
Syst Softw 81(6):961–971
10. Ahmedshareef Z, Hughes R, Petridis M (2014) Exposing the influencing factors on software
project delay with actor-network theory. Electron J Bus Res Methods 12(2):132–146
11. Kuriakose J, Parsons J (2015, August) An enhanced requirements gathering interface for open
source software development environments. In: 2015 IEEE 23rd international requirements
engineering conference (RE). IEEE, pp 288–289
12. McCann DE (2013) Managing changes in project scope: the role of the project constraints.
Walden University.
13. Peter S (2008) Managing user expectations on software projects: lessons from the trenches. Int
J Project Manag 26(7):700–712
488 B. U. Maheswari et al.

14. Joshi KD, Sarker S, Sarker S (2007) Knowledge transfer within information systems devel-
opment teams: examining the role of knowledge source attributes. Decis Support Syst
43(2):322–335
15. Barraza GA (2010) Probabilistic estimation and allocation of project time contingency. J Constr
Eng Manag 137(4):259–265
16. Anderson EG Jr, Chandrasekaran A, Davis-Blake A, Parker GG (2018) Managing distributed
product development projects: integration strategies for time-zone and language barriers. Inf
Syst Res 29(1):42–69
17. Gareis R (2007) Management of the project-oriented company. The Wiley Guide to Project,
Program & Portfolio Management, Hoboken (NJ), pp 250–270
18. Adelakun O (2008, May) The maturity of offshore IT outsourcing location readiness and
attractiveness. In: European and Mediterranean conference on information systems, pp 1–12
19. Maruping LM, Venkatesh V, Thong JY, Zhang X (2019) A risk mitigation framework
for information technology projects: a cultural contingency perspective. J Manag Inf Syst
36(1):120–157
20. Selaru C (2012) Resource allocation in project management. Int J Econ Pract Theor 2(4):274–
282
21. Engwall M, Jerbrant A (2003) The resource allocation syndrome: the prime challenge of multi-
project management? Int J Project Manag 21(6):403–409
22. Vanacker L, Van Raemdonck O, Servranckx T, Vanhoucke M (2018) The influence of project
resource allocation on the resource capacity of the business processes. J Mod Proj Manag
6(2):6–17
23. Belout A, Gauvreau C (2004) Factors influencing project success: the impact of human resource
management. Int J Project Manag 22(1):1–11
24. Patanakul P (2014) Managing large-scale IS/IT projects in the public sector: problems and
causes leading to poor performance. J High Technol Managem Res 25(1):21–35
25. Ramasubbu N, Bharadwaj A, Tayi GK (2015) Software process diversity: conceptualization,
measurement, and analysis of impact on project performance. MIS Q 39(4):787–807
26. Cervone HF (2011) Understanding agile project management methods using Scrum. OCLC.
Systems & Services: International digital library perspectives.
27. Machado TCS, Pinheiro PR, Tamanini I (2015) Project management aided by verbal decision
analysis approaches: a case study for the selection of the best SCRUM practices. Int Trans
Oper Res 22(2):287–312
28. Shameem M, Kumar RR, Nadeem M, Khan AA (2020) Taxonomical classification of barriers
for scaling agile methods in global software development environment using fuzzy analytic
hierarchy process. Appl Soft Comput 90:106122
29. Glowacka KJ, Lowe TJ, Wendell RE (2012) The impact of nonagility on service level and
project duration. Decis Sci 43(5):957–971
30. Siddique L, Hussein BA (2019) Enablers and barriers to customer involvement in agile software
projects in Norwegian software industry: the Supplier’s perspective. J Mod Proj Manag 7(2)
31. Parekh R, Olivia M (2024) Utilization of artificial intelligence in project management. Int J
Sci Res Arch 13(1):1093–1102
32. Choetkiertikul M, Dam HK, Tran T, Ghose A (2015, November) Predicting delays in software
projects using networked classification. In: 2015 30th IEEE/ACM international conference on
automated software engineering (ASE). IEEE, pp 353–364
33. de Barcelos Tronto IF, da Silva JDS, Sant’Anna N (2008) An investigation of artificial neural
networks based prediction systems in software project management. J Syst Softw 81(3):356–
367
34. BaniMustafa A (2018, July) Predicting software effort estimation using machine learning tech-
niques. In: 2018 8th international conference on computer science and information technology
(CSIT). IEEE, pp 249–256
35. Uma Maheswari B, Chandran HS, Sujatha R, Kavitha D (2022) Application of machine
learning algorithms for creating a wilful defaulter prediction model. In: Intelligent system
design: proceedings of INDIA 2022. Springer Nature Singapore, Singapore, pp 373–381
Author Index

A D
Aayush Shrivastava, 145, 311 Devnarayan G. Rao, 21
Abhilasha Sinha, 275 Dharmendra Sharma, 157
Abhishek Dwivedi, 169 Dilip Kumar Gandhi, 327
Abhishek Kumar Mishra, 169 Divyashree, H. B., 125
Abhishek Sharma, 135, 285 Duaa Hefni, 365
Abid Iqbal, 463
Abishek, J., 411
Agarwal, Neeraj, 327 H
Aggarwal, Dhruv, 145 Harshit Verma, 257
Agnes Nalini Vincent, 21 Hemant Kumar, 169
Agrawal, Jitendra, 113 Hridhya, J., 55
Aishwarya Mishra, 327 Hussain Falih Mahdi, 35, 43, 55, 71
Aishwarya Vishwakarma, 179
Akanksha Mishra, 35
Akshay Jain, 191
Aman Khan, 43 I
Amarjeet Ghosh, 327 Irfan Qasim, 463
Amin Mir, M., 365, 453, 463
Amit Nayak, 207
Andrews, K., 365 J
Anurag Singh Baghel, 257 Jagdish Chandola, 135, 285
Asha Ambhaikar, 35, 43 Jaishree R. Devaru, 125
Ashish Raghuwanshi, 327 Jigisha Mehta, 245
Joanna Rosak-Szyrocka, 179

B
Baldev Singh, 347 K
Bandi Krishna, 71 Kaushal Patel, 219
Bhupesh Kumar Dewangan, 35, 43, 55, 71 Kavita Kushwah, 311
Kavitha, D., 375, 475
Khushi Allawadi, 11
C Kirti Nahak, 43
Charan, R., 295 Kuljeet Singh, 97
© The Editor(s) (if applicable) and The Author(s), under exclusive license 489
to Springer Nature Singapore Pte Ltd. 2025
V. S. Rathore et al. (eds.), Modern Practices and Trends in Expert Applications
and Security, Lecture Notes in Networks and Systems 1378,
[Link]
490 Author Index

M Ramdas Vankdothu, 71
Maitri Kulkarni, 125 Rishabh Jaiswal, 257
Mamata Mayee Panda, 169 Romin Patel, 207
Manas Vyas, 231
Mangesh Bedekar, 1
Mani, P. K., 439 S
Manmohan Singh, 157 Sakshi Pundir, 135, 285
Megha Raina, 97 Sandeep, J., 21
Minakshi Memoria, 365, 453, 463 Sangeetha, G., 21
Mitali Chugh, 11 Sanjana Dewangan, 35
Mohammed Azim Eirgash, 285 Santoshini, V., 475
Mohan, S. B., 385 Sarishma Dangi, 145
Muddireddy Devaghneswara Reddy, 125 Sasi Kumar, M., 385
Muhammad Azhar Ali Khan, 453 Sengottaian, S., 439
Shaheen Ayyub, 179
Sharon Christa, 145
N Shekhar Verma, 169
Namrata Shrivastava, 113 Sheshang Degadwala, 219, 245
Naveen Sakthivel, K. S., 375 Shreya Panwar, 11
Neeraj Mohan, 337 Shubhnandan S. Jamwal, 83
Neha Jadon, 311 Sonali Vyas, 11
Neha Para, 311 Sreeja, C. S., 21
Niharika Awasthi, 231 Sugandh Singh, 179
Nikhat Raza Khan, 169 Sujatha, R., 375, 411, 475
Nilam Choudhary, 347 Sujitha, N., 439
Sumit Pundir, 135, 285
Sunny Singh, 55
O Suresh, C., 439
Onkar Nath Thakur, 275 Suresh Kumar, K., 295, 385, 425, 439
Suresh Kumar Lokhande, 71
Syed M. Hasnain, 365, 399, 453, 463
P
Pradeep, T., 295
Pragya Tewari, 257 T
Pramod Kumar Patel, 327 Tanupriya Choudhary, 35, 43, 55, 71
Premkumar, M., 385 Teena Mary, 21
Priyanka Bhatele, 1
Priyanshi Jain, 231
Punniyakotti Varadharajan Gopirajan, 295 U
Uma Maheswari, B., 375, 411, 475
Umar Bashir, 97
R
Radhika Patel, 207
Raghubir Singh Salaria, 337 V
Rajendra Thilahar, C., 425 Vibhakar Mansotra, 97
Rajesh Nema, 327 Vijayalakshmi, B., 385
Rajesh Rajaan, 347 Vijay Anand, M., 425
Raj Gaurav Mishra, 231 Vijay Singh Sen, 83
Rakesh Nayak, 71 Vikas Sakalle, 179
Rakesh Kumar Tiwari, 275 Viswanath Ananth, 411

Common questions

Powered by AI

Convolutional neural networks (CNNs) have significantly improved agricultural land use classification by enabling high-accuracy processing of time-series satellite imagery, such as those from Sentinel-2, to identify and classify crops and other land use types . Their ability to automatically derive features like edges and textures from data is a key advantage . However, a limitation is their dependence on large datasets for training and the intensive computational resources required, which can be a barrier for widespread deployment in resource-constrained environments .

Utilizing convolutional neural networks (CNNs) for rice disease and pest identification has a direct impact on enhancing the accuracy and speed of detecting these threats, as CNNs can process and classify images effectively even with subtle differences in disease symptoms . Indirectly, by enabling timely and precise identification, CNNs facilitate better pest management strategies and reduced crop losses, contributing positively to food security and farmer incomes . The enhanced diagnostic precision provided by CNNs also supports integrated pest management practices, leading to more sustainable agriculture .

The main challenges of using machine learning for anomaly detection in IoT environments include limited device resources, the difficulty of accurately defining normal behavior, high data dimensionality, and model robustness . Opportunities lie in the development of distributed systems like those leveraging blockchain for improved security and collaborative learning . Further advancements in hardware, such as GPUs, and enhanced neural network architectures offer promising avenues for addressing some of these challenges . However, improving detection accuracy without incurring substantial computational costs remains an ongoing concern .

Machine learning models have significantly evolved the approach to anomaly detection in smart city infrastructures by enabling more accurate and timely identification of irregular patterns, such as road surface abnormalities, air quality deviations, and resource usage inefficiencies . These advancements facilitate proactive urban management, allowing authorities to address issues before they escalate, improve resource allocation, and enhance citizens' quality of life . The implications include more informed decision-making and policy development, optimized city operations, and increased resilience against failures and unplanned events .

The integration of blockchain technology with IoT systems enhances the security of machine learning models by providing a decentralized and tamper-resistant ledger that ensures data integrity and transparency . Blockchain's consensus mechanisms prevent model corruption by adversaries and allow IoT devices to collaboratively maintain a single, unified model, avoiding single points of failure and reducing the risk of targeted attacks . This setup can be particularly effective in distributed networks where variance in computational power and data integrity is a concern .

LSTM networks are adept at capturing both immediate and broader contextual cues in text data due to their architecture, which includes memory cells capable of maintaining information over time . In the context of mental health analysis, particularly in detecting depression, LSTMs are used to process sequences of words in both forward and backward directions, allowing for a comprehensive understanding of the text. This bidirectional approach improves the sensitivity of models to subtle emotional variations and symptoms of mental distress . Their application is particularly valuable in sentiment analysis systems designed to identify depression and suicide risks .

The high computational costs in semi-supervised sentiment analysis arise from the complexity of processing large datasets, model training that involves both labeled and unlabeled data, and the need for sophisticated algorithms to parse nuanced language features . These costs can be mitigated by using more efficient algorithms such as those that leverage data parallelism, optimizing model architecture to reduce the computational overhead, and adopting transfer learning approaches from pre-trained models that cut down on extensive training times .

Long Short-Term Memory (LSTM) networks differ from traditional Recurrent Neural Networks (RNNs) by their unique gating mechanisms—input, forget, and output gates—which effectively manage long-term dependencies in sequential data by maintaining an internal memory state . This design helps overcome the vanishing gradient problem typically associated with RNNs, enabling LSTMs to retain information over longer sequences . In contrast, RNNs often struggle with longer dependencies, resulting in less effective processing of sequences with complex temporal patterns .

Natural Language Processing (NLP) is critical in developing sentiment analysis tools as it allows machines to understand and interpret human language, thereby enabling automated classification of text data according to sentiment polarity . NLP faces challenges in accurately interpreting human language, such as managing ambiguity, sarcasm, and context-dependence, which require deep learning models to capture complex syntactic and semantic features . Additionally, NLP must contend with variations in language use across different demographics and cultures, which adds layers of complexity to sentiment analysis models .

Recursive Feature Elimination (RFE) enhances the performance of gradient boosting models by systematically removing the least significant features, thereby streamlining the model input to focus on the most important predictors, which in turn reduces overfitting and improves prediction accuracy . In the context of cardiovascular disease prediction, selecting the right features is crucial as it determines the model’s ability to effectively parse out the critical factors influencing disease risk, thus facilitating more accurate patient diagnosis and risk assessment .

You might also like