AI Methods for Heart Disease Detection
AI Methods for Heart Disease Detection
International
INTELLIGENT METHODS Open Access
occurrence of Coronary Heart Disease (CHD) in their Forest, and Support Vector Machine in their project to
study. They employed NHANES data from 1999–2000 to predict heart disease in patients. They utilized the UCI
2015–2016 in their project. Using a Convolutional Neural dataset consisting of 303 samples and 14 input features.
Network (CNN) architecture, they achieved a With the Support Vector Classifier, they achieved an
classification power of 77% for accurately identifying the accuracy of 84.0% [13].
presence of CHD in a test dataset and 81.8% for accurately This section comprehensively covers the topic of heart
classifying the absence of CHD cases. The balanced disease. Additionally, it focuses on how artificial
accuracy of the model was determined to be 79.5% [6]. intelligence techniques, particularly machine learning
Srivastava & Choubey utilized the Cleveland Heart methods, can be applied to the classification and diagnosis
Disease dataset from UCI in their study to detect heart of heart diseases. Throughout this section, numerous
disease. They employed machine learning algorithms, studies and research findings related to recent
including K-Nearest Neighbors, Support Vector developments in the field of cardiology are referenced.
Machines, Decision Trees, and Random Forests. Among Table 1 includes previously published studies on heart
these, the K-Nearest Neighbors model achieved the diseases.
highest accuracy, with an accuracy rate of 87% [7].
Nikam et al. used a dataset comprising 12 rows and Tablo 1. Summary of Previously Published Studies on Heart
70,000 columns (patient records) in their study. After Diseases.
removing similar records, they utilized the remaining
Methods Dataset size Accuracy References
68,975 patient records. To determine which technique
more accurately predicted cardiovascular disease, they 303 samples
Support Vector
and 14 %64.4 [4]
employed various algorithms such as Neural Networks, Machine
features
Decision Tree Classifier, K-Nearest Neighbors, Logistic 303 samples
K-Nearest
Regression, Naive Bayes, XGB Classifier, and LGBM and 14 %87 [5]
Neighbor
Classifier. The Decision Tree Classifier yielded the highest features
- 116 -
Erdem et al., Intelligent Methods in Engineering Sciences 2(4): 115-124, 2023
2.1. Dataset
The dataset is named "Heart Disease Dataset," and it has
been obtained from Kaggle. Originally published by Mirza
HASNINE [16], this dataset comprises 16 different patient
features. In total, it consists of 4,238 records. The values
and value ranges of the features in this dataset are
presented in Table 2. The dataset provides a concise
overview of heart disease, covering its definition,
symptoms, statistics, risk factors, cardiac rehabilitation, a
quiz, public health initiatives, and additional resources.
Key points include the prevalence of conditions like
Coronary Artery Disease, symptoms of heart attacks and
heart failure, alarming statistics, crucial risk factors, the
importance of cardiac rehabilitation, and efforts by
organizations like the CDC. The information is sourced
from reputable institutions like the American Heart
Association and the National Heart, Lung, and Blood
Institute.
2.2. Performance Metric and Confusion Matrix
The confusion matrix is a matrix used to evaluate the
performance of classification algorithms [17]. This matrix
visually represents correct and incorrect classifications by
comparing the values predicted by a model with the actual
values. The confusion matrix assists in calculating
important metrics for a model in classification problems,
such as precision, specificity, accuracy, and F1 score. It is
used to understand how accurately the model predicts each
class and identify the types of errors made, providing
guidance for improving the model [18]. A binary
classification problem can be expressed as shown in Table
3 [19]:
Figure 1. General flow diagram of the project.
- 117 -
Erdem et al., Intelligent Methods in Engineering Sciences 2(4): 115-124, 2023
TN represents the number of True Negatives, which is Table 4. Table of Performance Measurements, Formulas, and
Evaluation Conditions.
the number of data points that the model correctly predicts
as negative. FP represents the number of False Positives, Performance
Formula Description
which is the number of data points that the model wrongly Metrics
The
predicts as positive when they are actually negative. FN summation
represents the number of False Negatives, which is the rate of
number of data points that the model wrongly predicts as correct
Accuracy (𝑇𝑁 + 𝑇𝑃)⁄(𝑇𝑁 + 𝐹𝑃 + 𝑇𝑃 + 𝐹𝑁) predictions
negative when they are actually positive. TP represents the is the
number of True Positives, which is the number of data number of
samples
points that the model correctly predicts as positive. evaluated.
It is used to
Table 3. Confusion matrix measure
positive
Predicted Class
other for being negative, is depicted in Figure 2. commonly used, simple method, particularly for
processing medical datasets [26]. For the training
examples of dataset D, trees are created based on high-
entropy inputs [27]. These trees are constructed simply and
quickly using a top-down recursive divide-and-conquer
(DAC) approach. The tree pruning process is applied to
remove irrelevant examples from D [28]. A diagram of a
Decision Tree classifier is shown in Figure 3.
𝒎
𝑷(𝑿|𝑪𝒊 ) ∗ 𝑷(𝑪𝒊 )
𝑷(𝑪𝒊 |𝑿) = Figure 3. Decision Tree diagram.
𝑷(𝑿) (1)
Step 3: Since P(X) is constant for each class, 𝑃(𝑋|𝐶𝑖 ) ∗ 2.4.3. Random Forest (RF)
𝑃(𝐶𝑖 ) should be maximized for each class. Random Forest (RF) classifier creates multiple decision
Step 4: Afterward, conditional independence of class is trees during the training phase and forms a class with an
assumed. average prediction. Each tree is trained with a subset of
examples consisting of independently randomly chosen
𝑷(𝑿|𝑪𝒊 ) = 𝑷(𝒙𝟏 |𝑪𝒊 ) ∗ 𝑷(𝒙𝟐 |𝑪𝒊 ) … … 𝑷(𝑿𝒎 |𝑪𝒊 ) (2) samples from the training dataset. In this process, trees use
Step 5: To predict class X, 𝑃(𝑋|𝐶𝑖 )𝑃(𝐶𝑖 ) is calculated randomly selected features depending on the input data.
for each class (𝐶𝑖 ). During the classification process, each tree independently
Naive Bayes classifier predicts the class label as (𝐶𝑖 ) if votes for the most popular class for the input vector, and
X is predicted to belong to class (𝐶𝑖 ) . the results are combined to make the classification. This
method is a logical strategy to achieve more reliable and
𝑷(𝑿|𝑪𝒊 )𝑷(𝑪𝒊 ) > 𝑷(𝑿|𝑪𝒋 )𝑷(𝑪𝒋 ) (3) effective classification results through the combination of
different trees. The number of features and the number of
𝒇𝒐𝒓 𝟏 ≤ 𝒋 ≤ 𝒎, 𝒋 ≠ 𝒊 (4) trees to be grown are two user-defined parameters required
to create a random forest classifier. At each node, only the
2.4.2. Decision Tree-DT selected features are considered for the best split [29, 30].
Decision Tree is a classification algorithm that can A diagram of a two-class Random Forest classifier, one for
handle numerical and categorical data, creating tree-like being heart disease positive and the other for being
structures. This algorithm facilitates the analysis of data by negative, is shown in Figure 4.
representing it in a graphical tree structure. It is a
- 119 -
Erdem et al., Intelligent Methods in Engineering Sciences 2(4): 115-124, 2023
- 120 -
Erdem et al., Intelligent Methods in Engineering Sciences 2(4): 115-124, 2023
Table 5. Confusion matrices of classification models Table 6. Performance measurement results of the models.
Naiv
Confusion Matrix Performanc e SV AN KN
DT RF LR
e Metrics Baye M N N
s
Algorithms Predicted Class 79. 83. 85.
Accuracy 78.9 70.9 83.7 83.4
9 9 5
Healthy Diseased 77. 78. 83.
Heart Heart Precision 79.9 74.5 78.3 77.8
8 0 4
Patient Patient 79. 83. 85.
Recall 78.9 70.9 83.7 83.4
Healthy 9 9 5
3111 483
Heart Patient 78. 79. 80.
Naive Bayes F1-score 79.4 72.6 79.1 79.3
Diseased 7 2 2
411 233
Heart Patient
Healthy
3244 350
Heart Patient According to Table 6, the highest classification
DT
Diseased accuracy value belongs to the LR model. The lowest
503 141
Heart Patient
classification accuracy value is attributed to the SVM
Healthy
3509 85 model. Other performance metrics also parallel the
Heart Patient
RF
Diseased
596 48
classification accuracy values of the classification models.
Heart Patient
The comparative display of all model performances is
Actual Class
Healthy
2865 729 shown in the graph in Figure 9.
Heart Patient
SVM
Diseased
505 139 According to Figure 9, the highest classification
Heart Patient
accuracy is for the LR model, while the lowest
Healthy
3483 111 classification accuracy is for the SVM model. Due to the
Heart Patient
ANN
Diseased different learning styles of each algorithm based on the
579 65
Heart Patient data, there may be variations in performance metrics. The
Healthy rankings of classification accuracies of classification
3474 120
Heart Patient
KNN models may vary depending on the data. The rankings
Diseased
582 62 shown in Figure 9 were obtained for this dataset.
Heart Patient
Healthy
3573 21
Heart Patient
LR
Diseased Figure 9. Performance table of models used in the diagnosis of
594 50 heart disease
Heart Patient
- 122 -
Erdem et al., Intelligent Methods in Engineering Sciences 2(4): 115-124, 2023
and 83.7%, respectively. These results demonstrate the [7] K. Srivastava and D. K. Choubey, "Heart disease prediction
using machine learning and data mining," International
potential usefulness of machine learning algorithms in the Journal of Recent Technology and Engineering, vol. 9, no. 1,
diagnosis of heart disease and their ability to assist in pp. 212-219, 2020, doi:
making accurate diagnoses. [Link]
[8] A. Nikam, S. Bhandari, A. Mhaske, and S. Mantri,
The KNN model has an accuracy of 83.4%, indicating "Cardiovascular disease prediction using machine learning
successful classification of the data. However, it has a models," in 2020 IEEE Pune Section International Conference
slightly lower accuracy compared to RF and ANN models. (PuneCon), 2020: IEEE, pp. 22-27, doi:
[Link]
The DT model achieves an accuracy of 79.9%, showing [9] S. N. Pasha, D. Ramesh, S. Mohmmad, and A. Harshavardhan,
that it classifies the data more successfully than Naive "Cardiovascular disease prediction using deep learning
Bayes and SVM models. techniques," in IOP conference series: materials science and
engineering, 2020, vol. 981, no. 2: IOP Publishing, p. 022006,
The Naive Bayes model has an accuracy of 78.9%. On doi: [Link]
the other hand, the accuracy of the SVM model is [10] P. Rubini, C. Subasini, A. V. Katharine, V. Kumaresan, S. G.
determined to be 70.9%. SVM draws attention with Kumar, and T. Nithya, "A cardiovascular disease prediction
using machine learning algorithms," Annals of the Romanian
particularly low accuracy on this dataset, indicating a Society for Cell Biology, vol. 25, no. 2, pp. 904-912, 2021.
mismatch with certain features of this dataset. Similarly, [11] A. Garg, B. Sharma, and R. Khan, "Heart disease prediction
the Naive Bayes also achieving low accuracy suggests that using machine learning techniques," in IOP Conference Series:
Materials Science and Engineering, 2021, vol. 1022, no. 1: IOP
this algorithm may not be an ideal choice for this dataset. Publishing, p. 012046, doi: [Link]
These results highlight that the success of machine 899X/1022/1/012046.
[12] K. Vayadande et al., "Heart Disease Prediction using Machine
learning projects depends on the characteristics of the Learning and Deep Learning Algorithms," in 2022
dataset and the accurate selection of the algorithm. International Conference on Computational Intelligence and
Sustainable Engineering Solutions (CISES), 2022: IEEE, pp.
Acknowledgments 393-401, doi:
[Link]
We would like to thank the Scientific Research [13] B. U. Rindhe, N. Ahire, R. Patil, S. Gagare, and M. Darade,
Coordinatorship of Selcuk University for their support "Heart disease prediction using machine learning," Heart
with the project titled “Diagnosis and Classification of Disease, vol. 5, no. 1, 2021, doi:
[Link]
Heart Disease with Artificial Intelligence Techniques” [14] M. Koklu, H. Kahramanli, and N. Allahverdi, "A new accurate
numbered 23401163. and efficient approach to extract classification rules," Journal
of the Faculty of Engineering and Architecture of Gazi
Data Availability University, vol. 29, no. 3, pp. 477-486, 2014.
[15] M. Koklu, H. Kahramanli, and N. Allahverdi, "A new approach
The dataset can be accessed through the following link: to classification rule extraction problem by the real value
[[Link] coding," International Journal of Innovative Computing,
Information and Control, vol. 8, no. 9, pp. 6303-6315, 2012.
disease-dataset/data]. [16] M. Hasnine. Heart Disease Dataset. [Online]. Available:
[Link]
References dataset
[1] R. Das, I. Turkoglu, and A. Sengur, "Effective diagnosis of [17] R. Butuner, I. Cinar, Y. S. Taspinar, R. Kursun, M. H. Calp,
heart disease through neural networks ensembles," Expert and M. Koklu, "Classification of deep image features of lentil
systems with applications, vol. 36, no. 4, pp. 7675-7680, 2009, varieties with machine learning techniques," European Food
doi: [Link] Research and Technology, vol. 249, no. 5, pp. 1303-1316, 2023.
[2] A. Javeed, S. Zhou, L. Yongjian, I. Qasim, A. Noor, and R. [18] M. Hossin and M. N. Sulaiman, "A review on evaluation
Nour, "An intelligent learning system based on random search metrics for data classification evaluations," International
algorithm and optimized random forest model for improved journal of data mining & knowledge management process, vol.
heart disease detection," IEEE access, vol. 7, pp. 180235- 5, no. 2, p. 1, 2015, doi:
180243, 2019, doi: [Link]
[Link] [19] N. B. Harikrishnan. "Confusion Matrix, Accuracy, Precision,
[3] K. Erdem and A. Duman, "Pulmonary artery pressures and Recall, F1 Score Binary Classification Metric." Analytics
right ventricular dimensions of post-COVID-19 patients Vidhya. [Link]
without previous significant cardiovascular pathology," Heart matrix-accuracy-precision-recall-f1-score-ade299cf63cds
& Lung, vol. 57, pp. 75-79, 2023, doi: (accessed.
[Link] [20] Y. S. Taspinar, "Light weight convolutional neural network and
[4] A. Jagtap, P. Malewadkar, O. Baswat, and H. Rambade, "Heart low-dimensional images transformation approach for
disease prediction using machine learning," International classification of thermal images," Case Studies in Thermal
Journal of Research in Engineering, Science and Management, Engineering, vol. 41, p. 102670, 2023, doi:
vol. 2, no. 2, pp. 352-355, 2019. [Link]
[5] A. Singh and R. Kumar, "Heart disease prediction using [21] I. Cinar, Y. S. Taspinar, R. Kursun, and M. Koklu,
machine learning algorithms," in 2020 international "Identification of Corneal Ulcers with Pre-Trained AlexNet
conference on electrical and electronics engineering (ICE3), Based on Transfer Learning," in 2022 11th Mediterranean
2020: IEEE, pp. 452-457, doi: Conference on Embedded Computing (MECO), 2022: IEEE, pp.
[Link] 1-4, doi: 10.1109/MECO55406.2022.9797218.
[6] A. Dutta, T. Batabyal, M. Basu, and S. T. Acton, "An efficient [22] M. Koklu, S. Sarigil, and O. Ozbek, "The use of machine
convolutional neural network for coronary heart disease learning methods in classification of pumpkin seeds (Cucurbita
prediction," Expert Systems with Applications, vol. 159, p. pepo L.)," Genetic Resources and Crop Evolution, vol. 68, no.
113408, 2020, doi: 7, pp. 2713-2726, 2021.
[Link] [23] Y. S. Taspinar, M. Koklu, and M. Altin, "Fire Detection in
- 123 -
Erdem et al., Intelligent Methods in Engineering Sciences 2(4): 115-124, 2023
Images Using Framework Based on Image Processing, Motion [33] V. Jakkula, "Tutorial on support vector machine (svm)," School
Detection and Convolutional Neural Network," International of EECS, Washington State University, vol. 37, no. 2.5, p. 3,
Journal of Intelligent Systems and Applications in Engineering, 2006.
vol. 9, no. 4, pp. 171-177, 2021, doi: [34] R. Katarya and S. K. Meena, "Machine learning techniques for
[Link] heart disease prediction: a comparative study and analysis,"
[24] D. Berrar, "Cross-Validation," vol. 1, ed, 2019, pp. 542-545. Health and Technology, vol. 11, pp. 87-97, 2021, doi:
[25] A. N. Repaka, S. D. Ravikanti, and R. G. Franklin, "Design and [Link]
implementing heart disease prediction using naives bayesian," [35] I. Ozkan, M. Koklu, and R. Saraçoğlu, "Classification of
in 2019 3rd International conference on trends in electronics pistachio species using improved k-NN classifier," Health, vol.
and informatics (ICOEI), 2019: IEEE, pp. 292-297, doi: 23, p. e2021044, 2021, doi: 10.23751/pn.v23i2.9686.
[Link] [36] Y. S. Taspinar, M. Koklu, and M. Altin, "Identification of the
[26] Y. S. Taspinar, M. Koklu, and M. Altin, "Classification of english accent spoken in different countries by the k-nearest
flame extinction based on acoustic oscillations using artificial neighbor method," International Journal of Intelligent Systems
intelligence methods," Case Studies in Thermal Engineering, and Applications in Engineering, vol. 8, no. 4, pp. 191-194,
vol. 28, p. 101561, 2021, doi: 10.1016/[Link].2021.101561. 2020, doi: [Link]
[27] I. Cinar and M. Koklu, "Determination of Effective and [37] N. Absar et al., "The efficacy of machine-learning-supported
Specific Physical Features of Rice Varieties by Computer smart system for heart disease prediction," in Healthcare, 2022,
Vision In Exterior Quality Inspection," Selcuk Journal of vol. 10, no. 6: MDPI, p. 1137, doi:
Agriculture and Food Sciences, vol. 35, no. 3, pp. 229-243, [Link]
2021. [38] K. S. K. Reddy and K. Kanimozhi, "Novel Intelligent Model
[28] S. Mohan, C. Thirumalai, and G. Srivastava, "Effective heart for Heart Disease Prediction using Dynamic KNN (DKNN)
disease prediction using hybrid machine learning techniques," with improved accuracy over SVM," in 2022 International
IEEE access, vol. 7, pp. 81542-81554, 2019, doi: Conference on Business Analytics for Technology and Security
[Link] (ICBATS), 2022: IEEE, pp. 1-5, doi:
[29] M. G. El-Shafiey, A. Hagag, E.-S. A. El-Dahshan, and M. A. [Link]
Ismail, "A hybrid GA and PSO optimized approach for heart- [39] R. Kursun, I. Cinar, Y. S. Taspinar, and M. Koklu, "Flower
disease prediction based on random forest," Multimedia Tools recognition system with optimized features for deep features,"
and Applications, vol. 81, no. 13, pp. 18155-18179, 2022, doi: in 2022 11th Mediterranean Conference on Embedded
[Link] Computing (MECO), 2022: IEEE, pp. 1-4.
[30] M. Pal, "Random forest classifier for remote sensing [40] S. Ambesange, A. Vijayalaxmi, S. Sridevi, and B. Yashoda,
classification," International journal of remote sensing, vol. 26, "Multiple heart diseases prediction using logistic regression
no. 1, pp. 217-222, 2005, doi: with ensemble and hyper parameter tuning techniques," in 2020
[Link] fourth world conference on smart trends in systems, security
[31] K. Tutuncu, I. Cinar, R. Kursun, and M. Koklu, "Edible and and sustainability (WorldS4), 2020: IEEE, pp. 827-832, doi:
poisonous mushrooms classification by machine learning [Link]
algorithms," in 2022 11th Mediterranean Conference on [41] P. Schober and T. R. Vetter, "Logistic regression in medical
Embedded Computing (MECO), 2022: IEEE, pp. 1-4, doi: research," Anesthesia and analgesia, vol. 132, no. 2, p. 365,
10.1109/MECO55406.2022.9797212. 2021, doi: [Link]
[32] S. I. Ayon, M. M. Islam, and M. R. Hossain, "Coronary artery
heart disease prediction: a comparative study of computational
intelligence techniques," IETE Journal of Research, vol. 68, no.
4, pp. 2488-2507, 2022.
- 124 -