0% found this document useful (0 votes)
18 views10 pages

AI Methods for Heart Disease Detection

This research article analyzes the effectiveness of various artificial intelligence methods for detecting heart diseases, utilizing a dataset of 4,238 records with 16 patient characteristics. The study evaluates seven machine learning algorithms, finding that Logistic Regression achieved the highest accuracy of 85.5%. The findings highlight the potential of machine learning to enhance early diagnosis and management of heart disease.

Uploaded by

rustychain
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
18 views10 pages

AI Methods for Heart Disease Detection

This research article analyzes the effectiveness of various artificial intelligence methods for detecting heart diseases, utilizing a dataset of 4,238 records with 16 patient characteristics. The study evaluates seven machine learning algorithms, finding that Logistic Regression achieved the highest accuracy of 85.5%. The findings highlight the potential of machine learning to enhance early diagnosis and management of heart disease.

Uploaded by

rustychain
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Intelligent Methods in Engineering Sciences 2(4): 115-124 2023

International
INTELLIGENT METHODS Open Access

IN ENGINEERING SCIENCES December, 2023

[Link] e-ISSN 2979-9236

Research Article [Link]

A Detailed Analysis of Detecting Heart Diseases Using Artificial Intelligence


Methods
Kenan ERDEM a , Muslume Beyza YILDIZ b , Elham Tahsin YASIN c ,
Murat KOKLU d, *
a
Department of Cardiology, Faculty of Medicine, Selcuk University, Konya, TURKIYE
b
Department of Computer Engineering, Selcuk University, Konya, TURKIYE
c
Graduate School of Natural and Applied Sciences, Selcuk University, Konya, TURKIYE
d
Department of Computer Engineering, Selcuk University, 42250 Selcuklu, Konya, Türkiye

ARTICLE INFO ABSTRACT


Article history: Hearts are crucial for maintaining a healthy lifestyle and are ranked high among the organs that
Received 30 November 2023 need special care. Globally, heart disease is one of the leading causes of death, posing a
Accepted 28 December 2023 considerable public health challenge in low-income countries in particular. Early diagnosis and
Keywords: the identification of risk factors are critical when dealing with these diseases, as early symptoms
Artificial Neural Networks, are often not evident. Heart disease can be caused by several factors, including smoking, poor diet,
Classification, Heart disease
diagnosis, Logistic Regression,
stress, a lack of physical activity, and excessive alcohol consumption. During the diagnosis
Random Forest. process, doctors may encounter various challenges, including vague symptoms, misleading test
results, and other medical complications. It is currently possible to diagnose heart disease more
accurately and effectively using machine learning algorithms. The present study examines seven
different machine learning algorithms on a dataset consisting of 4,238 records and 16 different
patient characteristics. Among the classification models, Naive Bayes, Decision Trees, Random
Forests, Support Vector Machines (SVM), Artificial Neural Networks (ANNs), K Nearest
Neighbors, and Logistic Regressions yielded 78.9%, 79.9%, 83.9%, 70.9%, 83.7%, 83.4%, and
85.5%, respectively.
This is an open access article under the CC BY-SA 4.0 license.
([Link]

1. INTRODUCTION screenings, identification of risk factors, and lifestyle


changes are crucial steps in effectively combating heart
The heart is one of the fundamental organs of the body,
disease [1-3].
and its significance extends beyond physical health. A
Jagtap et al.'s study is based on data collected from
healthy heart is among the most crucial factors
medical research conducted by Kaggle and the Cleveland
determining an individual's quality of life. Therefore, the
Foundation (from the University of California, Irvine).
heart plays a vital role in sustaining the body's vital
Seventy-five percent of the entries in the dataset are used
functions. Heart disease ranks among the most common
for training, while the remaining 25% are allocated for
causes of death worldwide. A healthy heart is a critical
testing purposes. Among the Support Vector Machine
factor in determining the quality of life. Epidemiological
(SVM), Logistic Regression, and Naive Bayes algorithms,
data consistently places heart diseases as the primary cause
it was observed that SVM exhibited the highest accuracy,
of death for many years. This trend is evident in both
with an efficiency of 64.4% [4].
developed and developing countries. A significant
Singh & Kumar used machine learning algorithms,
characteristic of heart diseases is that symptoms in the
specifically KNN, SVM, DT, and LR, in their study to
early stages can be mild or vague, often progressing
predict heart disease. They utilized the UCI dataset for
unnoticed. Hence, the diagnosis, identification of risk
training and testing. Seventy-three percent of the dataset
factors, and effective treatment methods are of great
was used for training, and the remaining 37% was
importance in preventing and managing heart disease.
allocated for testing. KNN emerged as the most successful
Unhealthy lifestyle habits such as smoking, poor diet,
model with an accuracy of 87% [5].
stress, a sedentary lifestyle devoid of physical activity, and
Dutta et al. utilized data from the National Health and
excessive alcohol consumption are significant factors that
Nutrition Examination Survey (NHANES) to predict the
increase the risk of heart disease. Regular health
* Corresponding Author: mkoklu@[Link]
Erdem et al., Intelligent Methods in Engineering Sciences 2(4): 115-124, 2023

occurrence of Coronary Heart Disease (CHD) in their Forest, and Support Vector Machine in their project to
study. They employed NHANES data from 1999–2000 to predict heart disease in patients. They utilized the UCI
2015–2016 in their project. Using a Convolutional Neural dataset consisting of 303 samples and 14 input features.
Network (CNN) architecture, they achieved a With the Support Vector Classifier, they achieved an
classification power of 77% for accurately identifying the accuracy of 84.0% [13].
presence of CHD in a test dataset and 81.8% for accurately This section comprehensively covers the topic of heart
classifying the absence of CHD cases. The balanced disease. Additionally, it focuses on how artificial
accuracy of the model was determined to be 79.5% [6]. intelligence techniques, particularly machine learning
Srivastava & Choubey utilized the Cleveland Heart methods, can be applied to the classification and diagnosis
Disease dataset from UCI in their study to detect heart of heart diseases. Throughout this section, numerous
disease. They employed machine learning algorithms, studies and research findings related to recent
including K-Nearest Neighbors, Support Vector developments in the field of cardiology are referenced.
Machines, Decision Trees, and Random Forests. Among Table 1 includes previously published studies on heart
these, the K-Nearest Neighbors model achieved the diseases.
highest accuracy, with an accuracy rate of 87% [7].
Nikam et al. used a dataset comprising 12 rows and Tablo 1. Summary of Previously Published Studies on Heart
70,000 columns (patient records) in their study. After Diseases.
removing similar records, they utilized the remaining
Methods Dataset size Accuracy References
68,975 patient records. To determine which technique
more accurately predicted cardiovascular disease, they 303 samples
Support Vector
and 14 %64.4 [4]
employed various algorithms such as Neural Networks, Machine
features
Decision Tree Classifier, K-Nearest Neighbors, Logistic 303 samples
K-Nearest
Regression, Naive Bayes, XGB Classifier, and LGBM and 14 %87 [5]
Neighbor
Classifier. The Decision Tree Classifier yielded the highest features

accuracy, with a rate of 73.12% [8]. Convolutional 37079


%79.5 [6]
Pasha et al. analyzed various algorithms such as Support Neural Networks patients
Vector Machines (SVM), K-Nearest Neighbors, and
Decision Trees in their articles. Among these, Artificial 303 samples
K-Nearest
Neural Network (ANN) achieved the highest accuracy, and 14 %87 [7]
Neighbor
features
with a rate of 85.24% [9].
Rubini et al. conducted a comparative analysis of 68,975
machine learning techniques, including Naïve Bayes, Decision Tree patient %73,12 [8]
records
Logistic Regression, Support Vector Machine (SVM), and
Random Forest (RF), for the classification of Artificial Neural
- %85.24 [9]
cardiovascular diseases in their articles. They Network

demonstrated that Random Forest achieved the highest


accuracy at 84.81%, establishing it as the most accurate Random Forest 14 features %84.81 [10]
and reliable algorithm among the tested methods [10].
Garg et al. employed two supervised machine learning 303 samples
K-Nearest
and 14 %86.885 [11]
algorithms, namely K-Nearest Neighbors (K-NN) and Neighbor
features
Random Forest, in their articles. The prediction accuracy Logistic
obtained with the Random Forest algorithm was 81.967%. Regression, 303 samples %88.52
Random Forest, and 14 %88.52 [12]
On the other hand, the prediction accuracy achieved by K- XGBoost features %88.52
Nearest Neighbors (K-NN), which outperformed Random Algorithms
Forest, was 86.885% [11].
Vayadande et al. utilized the Kaggle heart dataset, 2. Materials and Methods
which comprises 303 rows with a total of 14 feature There are various algorithms for the classification [14,
attributes, in their study. According to their findings, 15] process of heart disease detection, and the results can
Logistic Regression, Random Forest, and XGBoost vary for different datasets. Therefore, selecting the most
algorithms achieved a higher accuracy of 88.52% suitable classifier based on the used data is crucial for
compared to other methods such as Naive Bayes (NB), K- obtaining accurate classification results. For the detection
Nearest Neighbors (K-NN), Support Vector Machine of heart disease, models were trained using Naive Bayes,
(SVM), Multi-Layer Perceptrons, Artificial Neural Decision Tree, Random Forest, Support Vector Machine,
Network, Decision Tree, and Cat Boost [12]. K-Nearest Neighbor, Logistic Regression, and Artificial
Rindhe et al. used Artificial Neural Network, Random Neural Network algorithms. The steps used throughout

- 116 -
Erdem et al., Intelligent Methods in Engineering Sciences 2(4): 115-124, 2023

this project are outlined in Figure 1.

2.1. Dataset
The dataset is named "Heart Disease Dataset," and it has
been obtained from Kaggle. Originally published by Mirza
HASNINE [16], this dataset comprises 16 different patient
features. In total, it consists of 4,238 records. The values
and value ranges of the features in this dataset are
presented in Table 2. The dataset provides a concise
overview of heart disease, covering its definition,
symptoms, statistics, risk factors, cardiac rehabilitation, a
quiz, public health initiatives, and additional resources.
Key points include the prevalence of conditions like
Coronary Artery Disease, symptoms of heart attacks and
heart failure, alarming statistics, crucial risk factors, the
importance of cardiac rehabilitation, and efforts by
organizations like the CDC. The information is sourced
from reputable institutions like the American Heart
Association and the National Heart, Lung, and Blood
Institute.
2.2. Performance Metric and Confusion Matrix
The confusion matrix is a matrix used to evaluate the
performance of classification algorithms [17]. This matrix
visually represents correct and incorrect classifications by
comparing the values predicted by a model with the actual
values. The confusion matrix assists in calculating
important metrics for a model in classification problems,
such as precision, specificity, accuracy, and F1 score. It is
used to understand how accurately the model predicts each
class and identify the types of errors made, providing
guidance for improving the model [18]. A binary
classification problem can be expressed as shown in Table
3 [19]:
Figure 1. General flow diagram of the project.

Table 2. Values and value ranges of features in the dataset.


Features Values Features Values Features Values
Blood Pressure
Systolic Blood
Gender Male/Female Medications 0/1 83.5-295
Pressure (SysBP)
(BPMeds)
Diastolic Blood
Age 32-70 Prevalent Stroke Yes/No 48-142.5
Pressure (DiaBP)
Graduate
postgraduate Prevalent Body Mass Index
Education 0/1 15.54-56.8
Primaryschool Hypertension (BMI)
Uneducated Empty
Current Smoker 0/1 Diabetes 0/1 Heart Rate 44-143
Number of
0-70 Total Cholesterol 107-696 Glucose 40-394
Cigarettes Per Day
Heart Stroke Yes/No

- 117 -
Erdem et al., Intelligent Methods in Engineering Sciences 2(4): 115-124, 2023

TN represents the number of True Negatives, which is Table 4. Table of Performance Measurements, Formulas, and
Evaluation Conditions.
the number of data points that the model correctly predicts
as negative. FP represents the number of False Positives, Performance
Formula Description
which is the number of data points that the model wrongly Metrics
The
predicts as positive when they are actually negative. FN summation
represents the number of False Negatives, which is the rate of
number of data points that the model wrongly predicts as correct
Accuracy (𝑇𝑁 + 𝑇𝑃)⁄(𝑇𝑁 + 𝐹𝑃 + 𝑇𝑃 + 𝐹𝑁) predictions
negative when they are actually positive. TP represents the is the
number of True Positives, which is the number of data number of
samples
points that the model correctly predicts as positive. evaluated.
It is used to
Table 3. Confusion matrix measure
positive
Predicted Class

Actual Class patterns


correctly
Positive Negative
predicted
Positive 𝑇𝑃 𝐹𝑃 Precision 𝑇𝑃⁄(𝑇𝑃 + 𝐹𝑃) from the
total
Negative 𝐹𝑁 𝑇𝑁 number of
Performance metric is a measurement tool used to prediction
forms in a
assess the effectiveness and success of a system, model, or positive
process [20, 21]. These metrics are employed to evaluate class.
It is used to
the degree of success of a specific task, compare results, or measure the
identify improvement opportunities. The success of the proportion
Recall-
classification methods used in this research was measured 𝑇𝑃⁄(𝑇𝑃 + 𝐹𝑁) of correctly
Sensitivity
classified
with Table 4, which includes performance metrics, positive
formulas, and evaluation conditions. Classification results patterns.
Represents
were assessed according to these criteria, and the the
effectiveness of the model was evaluated [18, 22]. harmonic
mean
F1-score (2 ∗ 𝑇𝑃)⁄(2 ∗ 𝑇𝑃 + 𝐹𝑃 + 𝐹𝑁)
between
Recall and
Precision
values.

2.3. Cross Validation


Cross-validation, is a method used to objectively
evaluate the performance of machine learning models [23].
The dataset is divided into training and testing data, and
the model is trained and evaluated on different subsets of
data. This helps identify overfitting issues, assess
generalization capabilities, and obtain more reliable results
[21]. Cross-validation is a crucial tool in machine learning
projects to better predict the real-world performance of a
model. [24].
2.4. Machine Learning Algorithms
2.4.1. Naive Bayes -NB
Naive Bayes (NB) classifier is also referred to as the
"Independent Feature Model." It is based on Bayes'
theorem and serves as a simple probabilistic classifier with
a strong independence hypothesis. Essentially, the NB
classifier is used to predict the probability of an object or
data sample belonging to a specific class. In other words,
it assumes that the presence/absence of a specific class
feature is independent of the presence of another class
feature. NB classifiers typically work in supervised
learning and are highly suitable for high-dimensional input
situations [25]. The diagram of a Naive Bayesian classifier
for two classes, one for being heart disease positive and the
- 118 -
Erdem et al., Intelligent Methods in Engineering Sciences 2(4): 115-124, 2023

other for being negative, is depicted in Figure 2. commonly used, simple method, particularly for
processing medical datasets [26]. For the training
examples of dataset D, trees are created based on high-
entropy inputs [27]. These trees are constructed simply and
quickly using a top-down recursive divide-and-conquer
(DAC) approach. The tree pruning process is applied to
remove irrelevant examples from D [28]. A diagram of a
Decision Tree classifier is shown in Figure 3.
𝒎

𝑬𝒏𝒕𝒓𝒐𝒑𝒚 = − ∑ 𝒑𝒊𝒋 𝐥𝐨𝐠 𝟐 𝒑𝒊𝒋


𝒋=𝟏 (5)

Figure 2. Diagram of a two-class Naive Bayesian classifier.

Step 1: Let's assume D represents the training set, and


each record is represented by an n-dimensional feature
vector, denoted by 𝑋 = (𝑥1 + 𝑥2 … + 𝑥𝑛 ), which implies
predicting n measurements from n attributes (let's say from
A1 to An).
Step 2: Consider the number of classes m for prediction
(denote them as C1, C2, ..., Cm).
According to Bayes' theorem:

𝑷(𝑿|𝑪𝒊 ) ∗ 𝑷(𝑪𝒊 )
𝑷(𝑪𝒊 |𝑿) = Figure 3. Decision Tree diagram.
𝑷(𝑿) (1)
Step 3: Since P(X) is constant for each class, 𝑃(𝑋|𝐶𝑖 ) ∗ 2.4.3. Random Forest (RF)
𝑃(𝐶𝑖 ) should be maximized for each class. Random Forest (RF) classifier creates multiple decision
Step 4: Afterward, conditional independence of class is trees during the training phase and forms a class with an
assumed. average prediction. Each tree is trained with a subset of
examples consisting of independently randomly chosen
𝑷(𝑿|𝑪𝒊 ) = 𝑷(𝒙𝟏 |𝑪𝒊 ) ∗ 𝑷(𝒙𝟐 |𝑪𝒊 ) … … 𝑷(𝑿𝒎 |𝑪𝒊 ) (2) samples from the training dataset. In this process, trees use
Step 5: To predict class X, 𝑃(𝑋|𝐶𝑖 )𝑃(𝐶𝑖 ) is calculated randomly selected features depending on the input data.
for each class (𝐶𝑖 ). During the classification process, each tree independently
Naive Bayes classifier predicts the class label as (𝐶𝑖 ) if votes for the most popular class for the input vector, and
X is predicted to belong to class (𝐶𝑖 ) . the results are combined to make the classification. This
method is a logical strategy to achieve more reliable and
𝑷(𝑿|𝑪𝒊 )𝑷(𝑪𝒊 ) > 𝑷(𝑿|𝑪𝒋 )𝑷(𝑪𝒋 ) (3) effective classification results through the combination of
different trees. The number of features and the number of
𝒇𝒐𝒓 𝟏 ≤ 𝒋 ≤ 𝒎, 𝒋 ≠ 𝒊 (4) trees to be grown are two user-defined parameters required
to create a random forest classifier. At each node, only the
2.4.2. Decision Tree-DT selected features are considered for the best split [29, 30].
Decision Tree is a classification algorithm that can A diagram of a two-class Random Forest classifier, one for
handle numerical and categorical data, creating tree-like being heart disease positive and the other for being
structures. This algorithm facilitates the analysis of data by negative, is shown in Figure 4.
representing it in a graphical tree structure. It is a

- 119 -
Erdem et al., Intelligent Methods in Engineering Sciences 2(4): 115-124, 2023

Figure 4. Diagram of a two-class Random Forest.

2.4.4. Support Vector Machine (SVM)


Support Vector Machine (SVM) is a machine learning
algorithm that performs well, especially with small
datasets. Its main objective is to find the separation
hyperplane that best separates the data. By representing
data points as vectors in space, it seeks to separate classes
with the hyperplane that has the widest margin, and
support vectors are crucial in this process [31].
Additionally, it successfully classifies non-linear data by
mapping them into a high-dimensional feature space using
kernel functions, allowing it to handle complex datasets
[32, 33]. Figure 5 depicts the structure of a two-class
Support Vector Machine.

Figure 5. Diagram of a two-class Support Vector Machine.

2.4.5. Artificial Neural Network (ANN)


Artificial Neural Networks (ANN) is an artificial
intelligence model that mimics the functioning of
biological neurons. This structure consists of artificial
neurons organized in layers. It performs tasks such as data
processing, pattern recognition, and prediction. Input data
is multiplied by weights, processed using activation
functions, and produces the output [34]. Figure 6 illustrates
the diagram of a two-class artificial neural network.

- 120 -
Erdem et al., Intelligent Methods in Engineering Sciences 2(4): 115-124, 2023

the sample data points representing the dataset and


iteratively updates the model parameters [39]. After
training is complete, the logistic function is used to predict
the class of a new data point, and class labels are assigned
based on probability values above a threshold (usually
0.5). Logistic Regression is not only a simple and effective
classification method but also performs well with high-
dimensional data and is preferred for its interpretability.
However, it may not be sufficient on its own for linearly
inseparable data and can be extended with kernel methods
to address nonlinear problems [40, 41]. For illustrative
purposes, a typical two-class logistic regression is
Figure 6. Diagram of a two-class Artificial Neural Network provided in Figure 8.
2.4.6. K-Nearest Neighbor (KNN)
The K-Nearest Neighbors (KNN) algorithm is a simple
and popular learning method used in the field of machine
learning for classification and regression problems [35,
36]. In the case of classification, the algorithm determines
the k nearest sample data points using features
representing the data points in space to classify a new data
point, predicting the class based on the majority class of
these k neighbors. In regression, it predicts the target value
of a new data point by taking the average of the target
variables of the k nearest neighbors. KNN is preferred due
to its simple implementation and low cost, but its
performance may decrease with large datasets and high-
dimensional data. Additionally, performance should be
optimized with the proper choice of k value and data
preprocessing methods [37, 38]. The diagram of a two- Figure 8. Diagram of binary logistic regression.
class K-Nearest Neighbors classifier used to distinguish
3. Experimental Results
individuals with and without heart disease is shown in
Figure 7. In this section, the results obtained with the Naive
Bayes, DT, RF, SVM, ANN, KNN, and LR classification
methods on a dataset with a total of 4238 records and 16
different patient features are presented. The confusion
matrices for all models used are provided in Table 5.

Figure 7. Diagram of a two-class K-Nearest Neighbors.

2.4.7. Logistic Regression (LR)


Its main goal is to use a logistic function to separate data
points into two or more classes. The logistic function
predicts class labels by transforming input data into a
probability value range, typically between 0 and 1. The
algorithm is trained using feature vectors and their
corresponding class labels. During the training phase, the
model attempts to approximate a logistic function that fits
- 121 -
Erdem et al., Intelligent Methods in Engineering Sciences 2(4): 115-124, 2023

Table 5. Confusion matrices of classification models Table 6. Performance measurement results of the models.
Naiv
Confusion Matrix Performanc e SV AN KN
DT RF LR
e Metrics Baye M N N
s
Algorithms Predicted Class 79. 83. 85.
Accuracy 78.9 70.9 83.7 83.4
9 9 5
Healthy Diseased 77. 78. 83.
Heart Heart Precision 79.9 74.5 78.3 77.8
8 0 4
Patient Patient 79. 83. 85.
Recall 78.9 70.9 83.7 83.4
Healthy 9 9 5
3111 483
Heart Patient 78. 79. 80.
Naive Bayes F1-score 79.4 72.6 79.1 79.3
Diseased 7 2 2
411 233
Heart Patient
Healthy
3244 350
Heart Patient According to Table 6, the highest classification
DT
Diseased accuracy value belongs to the LR model. The lowest
503 141
Heart Patient
classification accuracy value is attributed to the SVM
Healthy
3509 85 model. Other performance metrics also parallel the
Heart Patient
RF
Diseased
596 48
classification accuracy values of the classification models.
Heart Patient
The comparative display of all model performances is
Actual Class

Healthy
2865 729 shown in the graph in Figure 9.
Heart Patient
SVM
Diseased
505 139 According to Figure 9, the highest classification
Heart Patient
accuracy is for the LR model, while the lowest
Healthy
3483 111 classification accuracy is for the SVM model. Due to the
Heart Patient
ANN
Diseased different learning styles of each algorithm based on the
579 65
Heart Patient data, there may be variations in performance metrics. The
Healthy rankings of classification accuracies of classification
3474 120
Heart Patient
KNN models may vary depending on the data. The rankings
Diseased
582 62 shown in Figure 9 were obtained for this dataset.
Heart Patient
Healthy
3573 21
Heart Patient
LR
Diseased Figure 9. Performance table of models used in the diagnosis of
594 50 heart disease
Heart Patient

The confusion matrices in Table 5 shape the main


findings of this research. The confusion matrix allows us
to evaluate the performance of classification models in
detail. In Table 5, the highest TP value is 3573, belonging
to the LR model. The lowest TP value is 2865, associated
with the SVM model. TP and TN values are of great
importance in disease detection. These values are the
primary determinants of success. It is expected that the
ratio of these values to the entire data is at the maximum
level. Table 6 displays the performance metrics of the
classification models. 4. Results and Findings
The obtained results are important in evaluating the
performance of various machine learning algorithms used
for the diagnosis of heart disease. This study examines the
effectiveness of these algorithms through classification
experiments conducted on a large dataset with 4,238
records and 16 different patient features. The results
indicate that different algorithms achieve different levels
of accuracy. The highest accuracy rate, 85.5%, is achieved
by the LR (Logistic Regression) model. Other algorithms
such as RF (Random Forest) and ANN (Artificial Neural
Network) also show good results with accuracies of 83.9%

- 122 -
Erdem et al., Intelligent Methods in Engineering Sciences 2(4): 115-124, 2023

and 83.7%, respectively. These results demonstrate the [7] K. Srivastava and D. K. Choubey, "Heart disease prediction
using machine learning and data mining," International
potential usefulness of machine learning algorithms in the Journal of Recent Technology and Engineering, vol. 9, no. 1,
diagnosis of heart disease and their ability to assist in pp. 212-219, 2020, doi:
making accurate diagnoses. [Link]
[8] A. Nikam, S. Bhandari, A. Mhaske, and S. Mantri,
The KNN model has an accuracy of 83.4%, indicating "Cardiovascular disease prediction using machine learning
successful classification of the data. However, it has a models," in 2020 IEEE Pune Section International Conference
slightly lower accuracy compared to RF and ANN models. (PuneCon), 2020: IEEE, pp. 22-27, doi:
[Link]
The DT model achieves an accuracy of 79.9%, showing [9] S. N. Pasha, D. Ramesh, S. Mohmmad, and A. Harshavardhan,
that it classifies the data more successfully than Naive "Cardiovascular disease prediction using deep learning
Bayes and SVM models. techniques," in IOP conference series: materials science and
engineering, 2020, vol. 981, no. 2: IOP Publishing, p. 022006,
The Naive Bayes model has an accuracy of 78.9%. On doi: [Link]
the other hand, the accuracy of the SVM model is [10] P. Rubini, C. Subasini, A. V. Katharine, V. Kumaresan, S. G.
determined to be 70.9%. SVM draws attention with Kumar, and T. Nithya, "A cardiovascular disease prediction
using machine learning algorithms," Annals of the Romanian
particularly low accuracy on this dataset, indicating a Society for Cell Biology, vol. 25, no. 2, pp. 904-912, 2021.
mismatch with certain features of this dataset. Similarly, [11] A. Garg, B. Sharma, and R. Khan, "Heart disease prediction
the Naive Bayes also achieving low accuracy suggests that using machine learning techniques," in IOP Conference Series:
Materials Science and Engineering, 2021, vol. 1022, no. 1: IOP
this algorithm may not be an ideal choice for this dataset. Publishing, p. 012046, doi: [Link]
These results highlight that the success of machine 899X/1022/1/012046.
[12] K. Vayadande et al., "Heart Disease Prediction using Machine
learning projects depends on the characteristics of the Learning and Deep Learning Algorithms," in 2022
dataset and the accurate selection of the algorithm. International Conference on Computational Intelligence and
Sustainable Engineering Solutions (CISES), 2022: IEEE, pp.
Acknowledgments 393-401, doi:
[Link]
We would like to thank the Scientific Research [13] B. U. Rindhe, N. Ahire, R. Patil, S. Gagare, and M. Darade,
Coordinatorship of Selcuk University for their support "Heart disease prediction using machine learning," Heart
with the project titled “Diagnosis and Classification of Disease, vol. 5, no. 1, 2021, doi:
[Link]
Heart Disease with Artificial Intelligence Techniques” [14] M. Koklu, H. Kahramanli, and N. Allahverdi, "A new accurate
numbered 23401163. and efficient approach to extract classification rules," Journal
of the Faculty of Engineering and Architecture of Gazi
Data Availability University, vol. 29, no. 3, pp. 477-486, 2014.
[15] M. Koklu, H. Kahramanli, and N. Allahverdi, "A new approach
The dataset can be accessed through the following link: to classification rule extraction problem by the real value
[[Link] coding," International Journal of Innovative Computing,
Information and Control, vol. 8, no. 9, pp. 6303-6315, 2012.
disease-dataset/data]. [16] M. Hasnine. Heart Disease Dataset. [Online]. Available:
[Link]
References dataset
[1] R. Das, I. Turkoglu, and A. Sengur, "Effective diagnosis of [17] R. Butuner, I. Cinar, Y. S. Taspinar, R. Kursun, M. H. Calp,
heart disease through neural networks ensembles," Expert and M. Koklu, "Classification of deep image features of lentil
systems with applications, vol. 36, no. 4, pp. 7675-7680, 2009, varieties with machine learning techniques," European Food
doi: [Link] Research and Technology, vol. 249, no. 5, pp. 1303-1316, 2023.
[2] A. Javeed, S. Zhou, L. Yongjian, I. Qasim, A. Noor, and R. [18] M. Hossin and M. N. Sulaiman, "A review on evaluation
Nour, "An intelligent learning system based on random search metrics for data classification evaluations," International
algorithm and optimized random forest model for improved journal of data mining & knowledge management process, vol.
heart disease detection," IEEE access, vol. 7, pp. 180235- 5, no. 2, p. 1, 2015, doi:
180243, 2019, doi: [Link]
[Link] [19] N. B. Harikrishnan. "Confusion Matrix, Accuracy, Precision,
[3] K. Erdem and A. Duman, "Pulmonary artery pressures and Recall, F1 Score Binary Classification Metric." Analytics
right ventricular dimensions of post-COVID-19 patients Vidhya. [Link]
without previous significant cardiovascular pathology," Heart matrix-accuracy-precision-recall-f1-score-ade299cf63cds
& Lung, vol. 57, pp. 75-79, 2023, doi: (accessed.
[Link] [20] Y. S. Taspinar, "Light weight convolutional neural network and
[4] A. Jagtap, P. Malewadkar, O. Baswat, and H. Rambade, "Heart low-dimensional images transformation approach for
disease prediction using machine learning," International classification of thermal images," Case Studies in Thermal
Journal of Research in Engineering, Science and Management, Engineering, vol. 41, p. 102670, 2023, doi:
vol. 2, no. 2, pp. 352-355, 2019. [Link]
[5] A. Singh and R. Kumar, "Heart disease prediction using [21] I. Cinar, Y. S. Taspinar, R. Kursun, and M. Koklu,
machine learning algorithms," in 2020 international "Identification of Corneal Ulcers with Pre-Trained AlexNet
conference on electrical and electronics engineering (ICE3), Based on Transfer Learning," in 2022 11th Mediterranean
2020: IEEE, pp. 452-457, doi: Conference on Embedded Computing (MECO), 2022: IEEE, pp.
[Link] 1-4, doi: 10.1109/MECO55406.2022.9797218.
[6] A. Dutta, T. Batabyal, M. Basu, and S. T. Acton, "An efficient [22] M. Koklu, S. Sarigil, and O. Ozbek, "The use of machine
convolutional neural network for coronary heart disease learning methods in classification of pumpkin seeds (Cucurbita
prediction," Expert Systems with Applications, vol. 159, p. pepo L.)," Genetic Resources and Crop Evolution, vol. 68, no.
113408, 2020, doi: 7, pp. 2713-2726, 2021.
[Link] [23] Y. S. Taspinar, M. Koklu, and M. Altin, "Fire Detection in

- 123 -
Erdem et al., Intelligent Methods in Engineering Sciences 2(4): 115-124, 2023

Images Using Framework Based on Image Processing, Motion [33] V. Jakkula, "Tutorial on support vector machine (svm)," School
Detection and Convolutional Neural Network," International of EECS, Washington State University, vol. 37, no. 2.5, p. 3,
Journal of Intelligent Systems and Applications in Engineering, 2006.
vol. 9, no. 4, pp. 171-177, 2021, doi: [34] R. Katarya and S. K. Meena, "Machine learning techniques for
[Link] heart disease prediction: a comparative study and analysis,"
[24] D. Berrar, "Cross-Validation," vol. 1, ed, 2019, pp. 542-545. Health and Technology, vol. 11, pp. 87-97, 2021, doi:
[25] A. N. Repaka, S. D. Ravikanti, and R. G. Franklin, "Design and [Link]
implementing heart disease prediction using naives bayesian," [35] I. Ozkan, M. Koklu, and R. Saraçoğlu, "Classification of
in 2019 3rd International conference on trends in electronics pistachio species using improved k-NN classifier," Health, vol.
and informatics (ICOEI), 2019: IEEE, pp. 292-297, doi: 23, p. e2021044, 2021, doi: 10.23751/pn.v23i2.9686.
[Link] [36] Y. S. Taspinar, M. Koklu, and M. Altin, "Identification of the
[26] Y. S. Taspinar, M. Koklu, and M. Altin, "Classification of english accent spoken in different countries by the k-nearest
flame extinction based on acoustic oscillations using artificial neighbor method," International Journal of Intelligent Systems
intelligence methods," Case Studies in Thermal Engineering, and Applications in Engineering, vol. 8, no. 4, pp. 191-194,
vol. 28, p. 101561, 2021, doi: 10.1016/[Link].2021.101561. 2020, doi: [Link]
[27] I. Cinar and M. Koklu, "Determination of Effective and [37] N. Absar et al., "The efficacy of machine-learning-supported
Specific Physical Features of Rice Varieties by Computer smart system for heart disease prediction," in Healthcare, 2022,
Vision In Exterior Quality Inspection," Selcuk Journal of vol. 10, no. 6: MDPI, p. 1137, doi:
Agriculture and Food Sciences, vol. 35, no. 3, pp. 229-243, [Link]
2021. [38] K. S. K. Reddy and K. Kanimozhi, "Novel Intelligent Model
[28] S. Mohan, C. Thirumalai, and G. Srivastava, "Effective heart for Heart Disease Prediction using Dynamic KNN (DKNN)
disease prediction using hybrid machine learning techniques," with improved accuracy over SVM," in 2022 International
IEEE access, vol. 7, pp. 81542-81554, 2019, doi: Conference on Business Analytics for Technology and Security
[Link] (ICBATS), 2022: IEEE, pp. 1-5, doi:
[29] M. G. El-Shafiey, A. Hagag, E.-S. A. El-Dahshan, and M. A. [Link]
Ismail, "A hybrid GA and PSO optimized approach for heart- [39] R. Kursun, I. Cinar, Y. S. Taspinar, and M. Koklu, "Flower
disease prediction based on random forest," Multimedia Tools recognition system with optimized features for deep features,"
and Applications, vol. 81, no. 13, pp. 18155-18179, 2022, doi: in 2022 11th Mediterranean Conference on Embedded
[Link] Computing (MECO), 2022: IEEE, pp. 1-4.
[30] M. Pal, "Random forest classifier for remote sensing [40] S. Ambesange, A. Vijayalaxmi, S. Sridevi, and B. Yashoda,
classification," International journal of remote sensing, vol. 26, "Multiple heart diseases prediction using logistic regression
no. 1, pp. 217-222, 2005, doi: with ensemble and hyper parameter tuning techniques," in 2020
[Link] fourth world conference on smart trends in systems, security
[31] K. Tutuncu, I. Cinar, R. Kursun, and M. Koklu, "Edible and and sustainability (WorldS4), 2020: IEEE, pp. 827-832, doi:
poisonous mushrooms classification by machine learning [Link]
algorithms," in 2022 11th Mediterranean Conference on [41] P. Schober and T. R. Vetter, "Logistic regression in medical
Embedded Computing (MECO), 2022: IEEE, pp. 1-4, doi: research," Anesthesia and analgesia, vol. 132, no. 2, p. 365,
10.1109/MECO55406.2022.9797212. 2021, doi: [Link]
[32] S. I. Ayon, M. M. Islam, and M. R. Hossain, "Coronary artery
heart disease prediction: a comparative study of computational
intelligence techniques," IETE Journal of Research, vol. 68, no.
4, pp. 2488-2507, 2022.

- 124 -

You might also like