0% found this document useful (0 votes)
13 views10 pages

An Explainable Machine Learning Framework

The document presents a study on an explainable machine learning framework for early stroke detection, utilizing various machine learning techniques to analyze patient datasets. The proposed ensemble model achieved a high accuracy of 99.90% in predicting stroke features, emphasizing the importance of factors such as age, BMI, and glucose levels. The research aims to assist healthcare professionals in diagnosing and preventing strokes through improved predictive accuracy and feature analysis.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
13 views10 pages

An Explainable Machine Learning Framework

The document presents a study on an explainable machine learning framework for early stroke detection, utilizing various machine learning techniques to analyze patient datasets. The proposed ensemble model achieved a high accuracy of 99.90% in predicting stroke features, emphasizing the importance of factors such as age, BMI, and glucose levels. The research aims to assist healthcare professionals in diagnosing and preventing strokes through improved predictive accuracy and feature analysis.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

International Multilingual Journal of Science and Technology (IMJST)

ISSN: 2528-9810
Vol. 8 Issue 5, May - 2023

An Explainable Machine Learning Framework


for Early Detection of Stroke
Md. Jamal Uddin1, Prodip Kumar Sarker2, Md. Nesarul Hoque1, Sakifa Aktar1, Md. Martuza Ahmad1
1
Department of Computer Science and Engineering
1
Bangabandhu Sheikh Mujibur Rahman Science and Technology University, Gopalganj-8100, Bangladesh
2
Department of Computer Science and Engineering
Begum Rokeya University, Rangpur, Bangladesh

Abstract— A stroke is a critical neurological struggle with difficulties. These problems can include
defect of the brain's blood vessels that occurs complications with attention, concentration, and
when the blood supply to a portion of the brain memory; struggling to speak or understand speech;
struggles or stops depriving brain cells of oxygen. emotional issues such as depression; loss of balance
It yields various forms of physical imbalance. It is or capability to walk; loss of sensation on the affected
one of the leading causes of illness and mortality side of the body; and having trouble eating food [1, 2].
worldwide. 20–25% of stroke survivors have Following the World Stroke Organization report, one in
substantial impairment, which has been linked to four adults over 25 will experience a stroke in their
an increased mortality risk. Recognizing the lifespan [3]. In addition, 12.2 million persons will suffer
numerous stroke warning signs early can prevent their first stroke in 2023, resulting in 6.5 million deaths.
a stroke from occurring. In this study, we The report also notes that over 110 million individuals
developed an ensemble learning-based machine have suffered a stroke worldwide. It also negatively
learning architecture capable of analyzing stroke impacts the patients’
patient datasets and precisely predicting and families, friends, workplaces, and social circumstances
recognizing stroke features. At first, a stroke [4]. Furthermore, contrary to prevalent opinion, it can
dataset is collected, and then the Synthetic occur at any stage of life, gender, or physical
Minority Oversampling Technique (SMOTE) is condition.
used to balance it. Then, we implemented several
Patients may be at risk for stroke for multiple
machine learning techniques, such as Decision
causes. Following to the National Heart, Lung, and
Tree, Naive Bayes, K-Nearest Neighbors, Random
Blood Institute, high blood pressure, heart and blood
Forest, Extreme Gradient Boosting, Multilayer
vessel diseases, diabetes, smoking, brain aneurysms,
Perceptron, Ada Boost, and our proposed
family history or genetics, and other complications may
Ensemble framework. After optimizing
be the leading causes of stroke [5]. To reduce the risk
hyperparameters, our proposed framework
of stroke, it is essential to routinely track blood
demonstrated the highest accuracy (99.90%)
pressure, exercise consistently, maintain a healthy
among all machine learning classifiers. We
weight, stop smoking and taking alcohol, and consume
identified Age, BMI, and Average Glucose Level,
a healthy, low-fat diet [6, 7].
Heart Disease as significant stroke indicators
using machine learning (information gain, Typically, the medical dataset includes patient
correlation, and Relief F) and statistical feature symptoms and physical conditions. Recent years have
selection techniques. The SHapley Additive seen the emergence of machine learning (ML) as a
exExplanations (SHAP) method is utilized to cutting-edge approach for healthcare prognosis and
determine the influence of each attribute on the diagnosis that can classify medical information into
model outcome. We believe that our proposed specific class labels, such as sick or non-sick [8]. Such
framework can assist physicians and clinicians in a strategy has enabled the successful implementation
prescribing and detecting a potential stroke early of ML for increased diagnostic accuracy and efficiency
on. [9–11]. A growing number of studies over the past
decade have investigated the applicability of the ML
Index Terms—Stroke; Machine Learning; Feature
Selection; Feature Importance; SHAP;
models for predicting stroke [12, 13]. The authors of
[14–16] experimented on the same dataset containing
I. INTRODUCTION 5,110 patients with eleven clinical features and one
target attribute. However, Dritsas and Trigka [14]
A stroke occurs when the regular supply of blood to considered 3,254 participants over 18 years old. In the
the brain is interrupted or blocked suddenly, or the preprocessing phase, they handled missing values and
blood vessels are leaked or ruptured in the brain. removed noisy contents. Then, they manipulated
Without sufficient blood circulation, brain cells SMOTE (synthetic minority over-sampling technique)
progressively perish, resulting in varying degrees of method to balance the significant (non-stoke) and
complications. Even though some patients can recover minor (stoke) classes. Finally, they applied various ML
after suffering a stroke, depending on the severity of models, where the stacking of Logistic Regression
the stroke, a significant number of patients continue to

[Link]
IMJSTP29120877 6295
International Multilingual Journal of Science and Technology (IMJST)
ISSN: 2528-9810
Vol. 8 Issue 5, May - 2023

(LR), Naive Bayes (NB), Random forest (RF), identify two types of stroke: ischemic stroke (IS) and
RepTree, and J48, outperformed with 98% accuracy hemorrhage stroke (HE). At first, the authors collected
and 97.4% F-measure. The authors did not consider 507 distinct patient records between the age group 35
any deep learning models in this prediction system. In and 90, where 91.52% of them are infected by IS,
[15], the authors removed an irrelevant feature and whereas 8.48% are by HE. Then, they extracted 22
substituted the missing values with the mean value of significant features from those records using a base-
a column. They used the under sampling method to form generator and a stemmer algorithm. Lastly, they
balance the dataset, where the total rows reached 498 exploited several ML models like Decision Tree (DT),
(249 for the stroke class and 249 for the non-stroke Support Vector Machine (SVM), Artificial Neural
class). They got the highest accuracy of 82% by Networks (ANN), RF, and others, where ANN
employing the NB classifier. The authors did not apply outperformed the others with 95.3% accuracy and a
neural network based models in this study. lower standard deviation (14.69). This study reveals
Furthermore, they were concerned about the detection that stroke is more common in men than women and
performance because of limited trained and test data. those aged 40 to 60. The authors did not employ any
Rahman et al. [16] filled the missing values with the technique to handle the highly imbalanced dataset.
most frequent values of the feature column, applied
Thought out the analysis of the existing research
the MinMaxScaler method to normalize the features,
works, we have detected several issues, such as
utilized Principal Component Analysis (PCA) technique
dataset completeness and balancing, feature
to reduce the feature dimensions, and employed a
selection analysis, detection performance, and others.
random over-sampling strategy for balancing the
In this research, first, we collect a renowned dataset
dataset in the pre-processing stage. After that, they
from a popular data repository platform [Link].
applied various ML and deep neural network (DNN)
After that, we made the following contributions by
models, where the RF showed excellent accuracy of
focusing on all the above issues:
99% than the other models. In this investigation, the
authors performed less work to analyze the most  To handle the completeness and the highly
significant factors strongly related to stroke imbalanced nature of the dataset.
occurrence. Dev et al. [17] worked on a comparatively  To figure out the important features of stroke
large dataset. The authors investigated various factors risk factors.
presented in the Electronic Health Record (EHR)  To propose an ensemble ML model for
records of 29,072 patients, where only 548 entries are automated stroke prediction with outstanding
associated with stroke condition, while the remaining performance.
28,524 are not of to stroke nature. They sorted out four  To rank and analyze, the risk factors of stroke
significant factors, including age, average glucose using ML techniques.
status, hypertension, and heart disease, using
Learning Vector Quantization (LVQ) model for stroke II. MATERIALS AND METHODS
prediction. They employed a random down sub- In this research, using a stroke dataset, we
sampling approach to handle the bias of the majority employed statistical and ML techniques to determine
class (not stroke). They achieved an optimistic output
for their four selected features of 78% accuracy with a
low miss rate of 19% by utilizing a neural network (NN)
model. However, the performance score is still lacking
for treatment and precluding measures for an
individual. Liu et al. [18] worked on another bigger
dataset for predicting cerebral stroke containing eleven
input features and 43,400 samples, of which only 783
patients experienced stroke. Besides the highly
imbalanced nature, the dataset is incomplete because
of the missing value of some fields. As the outliers and
noisy samples, the authors filtered those patient
instances that have aged below 25 and a BMI (body
mass index) value is more than 60%. Then they
employed the random forest regression (RFR) method
to attribute the missing values. Using an automated
hyperparameter optimization (AutoHPO)-based DNN
model, they obtained the lowest false negative rate of
19.1% with an accuracy of 71.6%. They used XGBoost Fig. 1: The workflow of stroke prediction
(XGB) and RF models to get significant factors
regarding the occurrence of strokes. The authors did the most prominent indicators connected to stroke.
not utilize DNN methods in the feature important Next, ML models were utilized to identify early-stage
analysis. Additionally, there is still enough space to strokes. Figure 1 depicts a comprehensive pictorial
enhance the stroke prediction system. Govindarajan et representation of the workflow.
al. [19] developed a stroke classification system to

[Link]
IMJSTP29120877 6296
International Multilingual Journal of Science and Technology (IMJST)
ISSN: 2528-9810
Vol. 8 Issue 5, May - 2023

A. Data Collection and Description  Information Gain (IG) depends on the entropy
We obtained the dataset from the publicly concept, which measures the impureness or
accessible Kaggle data repository, and it contains uncertainty of a data set [22].
twelve features: ID, Gender, Age, Hypertension, Heart  Pearson’s correlation (PC) calculates its value
disease, Marital Status, Work Type, Residence Type, for the variable in the class to determine the
Average Glucose Level, BMI, Smoking Status, and value of an attribute [23].
Stroke. This dataset includes 5,110 records, of which  Relief F determines the value of a feature by
249 (4.88%) are stroke patients and 4,861 (95.12%) continually sampling an instance and
are not. The average age of patients was 43.21 years, evaluating the given attribute’s value for the
ranging from 82 years to 8 months. Male patients nearest instances of the same and different
were 2,115, and females were 2,994. The mean BMI classes [24].
was 28.85, the maximum was 97.6, and the minimum
E. Machine Learning Model
was 10.3. Hypertension and cardiovascular disease
prevalence were extremely low (9.75% and 9.74%, In this study, we employed a number of ML
respectively). The average blood glucose level was algorithms, including Decision-Tree (DT), Naive Bayes
106.14, with a maximum of 271.74 and a minimum of (NB), K-Nearest Neighbors (KNN), Random Forest
55.12. Table I provides a detailed description of the (RF), Extreme Gradient Boosting Machine (XGB),
dataset. Figure 2 depicts the correlation between each Multilayer Perceptron (MLP), Ada Boost (AB), and
attribute. We did not eliminate any features because Ensemble Method (EM). We searched for the best-
there is no significant correlation between them. performing models to predict stroke. We illustrated
each model below:
B. Dataset Balancing Technique
The SMOTE creates samples of the minority Decision-Tree (DT) technique has rapidly become
population that are artificial. It trains a classifier by a popular ML tool for classification and regression
synthesizing a balanced synthetic training set across problems. The algorithm’s pervasive approval and use
classes. With SMOTE, the distribution of the dataset can be attributed to its remarkable resemblance to
is more uniform as more instances are added to the human thought. It can be used as a non-parametric
minority class. Consequently, ML models can profit supervised learning classifier, expanding its versatility.
from acquiring knowledge from a more specific data A limited number of “nodes” within a decision tree
set. It has found pervasive application in fields where represent potential steps, while “leaf nodes” define the
imbalanced datasets are standard, such as identifying outcomes of those actions.
fraud, health care diagnosis, and text classification Naive Bayes (NB) is a simple yet effective
[20]. classifier that can effectively forecast the outcomes of
C. Feature Transformation Method many challenging problems in a brief period. It
employs Bayes’ theorem and operates under the
Standard Scaler [21] rescales the characteristics assumption of a distinct variable [25].
with a mean of zero and a variance of one. It is used
when the ranges of the characteristics of the input K-Nearest Neighbors (KNN) is a straightforward,
dataset are extensive. The standard normal nonparametric supervised learning technique that
distribution is the consequence of subtracting attribute looks for similar features in the training set [26]. It
values by their mean and dividing by standard frequently employs the Euclidean, Manhattan, and
deviation (σ). Minkowski distance procedures to differentiate
between new input and preexisting knowledge.
𝑥 − 𝑥̅
𝑆𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑆𝑐𝑎𝑙𝑎𝑟(𝑥) = (1) Random Forest (RF) [27] is a compilation of
𝜎
decision trees derived from multiple samples that are
Where 𝑥 is the original feature value, 𝑥̅ is the mean employed to enhance the accuracy of a dataset. Each
of the feature values. Standard scalar is frequently decision tree was trained on a sample of data
employed in numerous ML algorithms to prepare and selected at random with replacement using the
normalize the input features, including linear bagging method. The output of a RF is a composite
regression, LR, SVM, and NN. prediction equation that is derived from the outputs of
D. Feature Ranking Method multiple decision trees. It can be applied to both
regression and classification problems [28].
Feature ranking is a method for evaluating the
significance or relevance of a dataset’s features. It eXtreme Gradient Boosting (XGBoost) technique
attempts to evaluate the features based on their combines gradient enhancement and boosting to
contribution to an ML model’s predictive ability. In our produce exceptional results. It is an effective and
investigation, we used information gain, person scalable variant of the Gradient Boosting Method
correlation, and Relief F techniques for ranking (GBM) that can perform a variety of tasks, including
features. regression, classification, and ranking [29].

[Link]
IMJSTP29120877 6297
International Multilingual Journal of Science and Technology (IMJST)
ISSN: 2528-9810
Vol. 8 Issue 5, May - 2023

TABLE I: Dataset Description

ID Attribute Name Feature Description Count/Average Value


Male: 2115(41.39%)
1 Gender Gender of the patient Female: 2994(58.60%)
Other: 1(00.01%)
Average: 43.21
Max: 82
2 Age Age of the patient
Min: 0.08
Median: 45
Yes: 498(9.75%)
3 Hypertension Whether or not the respondent has hypertension
No: 4612(90.25)
Yes: 498(9.74%)
4 Heart disease Whether the respondent has heart disease or not
No: 4612(90.26)
Yes: 3353(65.61%)
5 Ever married Whether the respondent is married or single
No: 1757(34.39%)
Govt. job: 657(12.85%)
It can be govt. job, never worked, private or self- Never worked: 709(13.88%)
6 Work type
employed Private: 2925(57.24%)
Self-employed: 819(16.03%)
Rural: 2514(49.19%)
7 Residence type Residence can be rural or urban
Urban: 2596(50.81%)
Average: 106.14
Average Max: 271.74
8 Average glucose level in blood
glucose level Min: 55.12
Median: 91.88
Average: 28.85
Max: 97.6
9 BMI Body mass index
Min: 10.3
Median: 27.78
Formerly smoked: 885(17.32%)
Smoking status can be formerly smoked, never Never smoked: 1892(37.03%)
10 Smoking status
smoked, smokes or unknown Smokes: 789(15.45%)
Unknown: 1544(30.22%)
Yes: 249(4.88%)
11 Stroke The patient has a stroke or not
No: 4861(95.12%)

increased robustness, and enhanced precision. It can


Multilayer Perceptron (MLP) is a popular form of
aid in mitigating individual model biases and
artificial neural network (ANN) in ML. It is a neural
enhancing model performance overall.
network with multiple layers of interrelated nodes
called neurons or units. It is renowned for its ability to F. Hyperparameter Optimization Technique
discover intricate data relationships and patterns.
Grid search is a straightforward but exhaustive
They can handle a vast array of challenges in
technique for tuning the hyperparameters of each
domains, such as classification, regression, and even
classifier. It investigates all possible hyperparameter
sequence related tasks.
value combinations within the specified grid. The ML
Adaptive Boosting (AB) is a technique popularized algorithm is trained with altered parameters to
by a boosting algorithm [30]. This method aims to identify the feature with the highest degree of
integrate multiple weak classifiers into a single robust precision. This entire process is a cycle within which
classifier. assessments are continued.
Ensemble Method (EM) combines XGB and RF G. Performance Evaluation Metrics
and incorporates the advantages of both algorithms
In this research, we evaluated the effectiveness of
to enhance the overall predictive performance. By
classifiers using a variety of assessment metrics [31],
integrating the assets of multiple models, ensemble
including accuracy [32], kapa statistics, precision,
techniques can offer improved generalization,
recall, F1-score, AUC, and log- loss. Using the

[Link]
IMJSTP29120877 6298
International Multilingual Journal of Science and Technology (IMJST)
ISSN: 2528-9810
Vol. 8 Issue 5, May - 2023

Fig. 2: The correlation of each feature of stroke dataset

Characteristic (ROC) Curve’s Area Under the Curve


(AUC). It is a crucial rating criterion.
following equations, evaluation metrics are
calculated. F1-score: Precision and recall are weighted
averages to provide the F1 score. The mathematical
Accuracy: Accuracy is the proportion of accurate
equation of the F1-score is
forecasts relative to the total number of forecasts.
This metric performs optimally when each class 𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛∗𝑅𝑒𝑐𝑎𝑙𝑙
𝐹1 − 𝑆𝑐𝑜𝑟𝑒 = 2 (6)
𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛+𝑅𝑒𝑐𝑎𝑙𝑙
contains approximately the same number of samples.
The accuracy is expressed as Log Loss: It is a reliable indicator of classification
(𝑇𝑃+𝑇𝑁) work-place productivity. The Log Loss probability-
𝐴𝑐𝑐𝑢𝑟𝑎𝑐𝑦 = (𝑇𝑃+𝑇𝑁+𝐹𝑃+𝐹𝑁) (2) based metric. Lower values for Log Loss indicate
more precise predictions. An ideal classifier would
Kapa Statistics: It is a measure of how well two exhibit zero Log Loss. The mathematical equation of
evaluators agree on a value. In the context of ML Log Loss is:
model evaluation metrics, this value is the difference
1
between the predicted and observed output. 𝐿𝑜𝑔 𝐿𝑜𝑠𝑠 = − ∑𝑁 𝑦 log(𝑝(𝑦𝑖 )) + (1 − 𝑦𝑖 )𝑙𝑜𝑔(1 −
𝑁 𝑖=1 𝑖

𝐾𝑎𝑝𝑎 𝑠𝑡𝑎𝑡𝑖𝑠𝑡𝑖𝑐𝑠 =
1−(1−𝑝0 )
(3) 𝑝(𝑦𝑖 )) (7)
1−𝑝𝑒
Here, TP is True Positive, FP is False Positive,
Precision: Precision is defined as the proportion of TN is True Negative and TP is True Positive. In log
accurately predicted outcomes to actual outcomes. loss, y represents the level of the target variable, p(y)
Precision is expressed as: represents the projected probability, and q represents
𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛 =
𝑇𝑃
(4) the actual log loss. And in the Kappa-Statistics 𝑝0 is
𝑇𝑃+𝐹𝑃 among raters agreement of relative observed, and 𝑝𝑒
Recall: Recall is the proportion of accurate is a chance agreement for hypothetical probability.
predictions to all the actual positive outcomes. The III. RESULTS
equation of recall is:
𝑇𝑃
We implemented the ML classifiers DT, NB, KNN,
𝑅𝑒𝑐𝑎𝑙𝑙 = (5) RF, XGB, MLP, AB, LR, and EM in our work. At
𝑇𝑃+𝐹𝑁
Google Collaboratory, the experimental task was
AUC-ROC: The efficacy of our model can be performed with Python sci-kit-learn. In this study,
demonstrated using the Receiver Operating prediction models were developed using a 10-fold

[Link]
IMJSTP29120877 6299
International Multilingual Journal of Science and Technology (IMJST)
ISSN: 2528-9810
Vol. 8 Issue 5, May - 2023

cross-validation procedure. We used the IG, PC, and most remarkable performance on the primary data
Relief F feature selection methodologies to determine set.
the significance of each feature. The SHAP summary
In the case of the balanced dataset (Table IV), XGB
graph was generated using Python’s shap module.
calculated the highest accuracy (95.60%), kappa
Several evaluation metrics, including precision, recall,
statistics (91.20%), auc-roc (95.60), F1-score
AUC-ROC, F1-score, and log loss, are applied to
(95.59%), and lowest log loss (1.5868). The EM and
validate the experimental results.
KNN had the highest precision (96.08%) and recall
A. Identifying crucial traits of stroke using statistical (98.17%). Similarly, the RF and EM produced well
and machine learning techniques across all evaluation metrics. Overall, XGM
performed better than competing classifiers.
The stroke dataset utilized the Chi-square test to
identify the most significant factors that caused the In relation to hyperparameter tuning of classifiers
stroke. Figure 3 depicts our findings. The most (Table V), the EM determined the maximum accuracy
significant indicators in descending order are Age, (99.90%), kappa statistics (99.79%), recall (99.90%),
Heart Disease, Average Glucose Level, Hypertension, auc-roc (99.90%), F1-score (99.90%), and the least
and Marital Status. amount of log loss (0.0371). The KNN demonstrated
the utmost precision (one hundred percent). The
Using the IG, PC, and Relief F methods, we
XGB and KNN additionally showed excellent
calculated the feature importance for the stroke
performance across all evaluation metrics. Overall,
dataset to identify the risk factors for stroke prediction.
EM performed better than competing classifiers.
Table II displays the numerical outcomes.
IG determined that the highest feature importance
value is 0.5129 for age, 0.4467 for BMI, and 0.0836
for Work Type. The most significant attributes, Age
(0.2453), Heart Disease (0.1349), and Average
Glucose Level (0.1319), were manipulated by PC
techniques. Age (0.2818), Work Type (0.0972), and
BMI (0.0927) were determined to be the essential
characteristics by the Relief F method. We also
calculated the average value of the IG, PC, and
Relief F methodologies, as shown in Figure 4. Age,
Body Mass Index, and Average Glucose Level were
observed as the most significant characteristics.
B. Exploring Discriminatory Stroke Identification
Factors
Figure 5 illustrates the order of importance of Fig. 3: The feature significance in which a larger bubble
SHAP values within stroke datasets. The evaluation represents a greater importance
of these values was conducted using the XGB, which
performed excellently. Age, BMI, Average Glucose
Level, and Smoking Status were the most critical TABLE II: Feature Importance using machine learning
discriminatory characteristics for detecting stroke at technique
an early stage. In contrast, Heart Disease, Relief
Hypertension, and Work Type constituted the least ID Attribute Name Info Gain Correlation
F
significant discriminatory factors. 1 Gender 0.0270 0.0089 0.0180
C. Classification of Stroke Using Machine Learning 2 Age 0.5129 0.2453 0.2818
Algorithms 3 Hypertension 0.0066 0.1279 0.0022
4 Heart disease 0.0000 0.1349 0.0046
In the main dataset (Table III), EM calculated the
5 Ever married 0.0085 0.1083 0.0262
maximum accuracy (95.07%) and minimal log loss
6 Work type 0.0836 0.0671 0.0972
(1.7775). The KNN, RF, XGB, MLP, and AB all
demonstrated approximately 95.00% accuracy. NB 7 Residence type 0.0253 0.0155 0.0119
demonstrated the highest kappa statistics (17.01%), Average glucose
8 0.0709 0.1319 0.0831
recall (42.17%), auc-roc (65.29%), and F1-score level
(22.91%). XGB performed with the highest accuracy 9 BMI 0.4467 0.0350 0.0927
(20.29%). Overall, the NB classifier demonstrated the 10 Smoking status 0.0715 0.0198 0.0769

[Link]
IMJSTP29120877 6300
International Multilingual Journal of Science and Technology (IMJST)
ISSN: 2528-9810
Vol. 8 Issue 5, May - 2023

Fig. 4: Feature ranking using machine learning techniques

Fig. 5: Analysis of Shapley values for the stroke dataset

TABLE III: Performance Analysis of Different Classifiers in Main Dataset

Evaluation Classifiers
Metrics DT NB KNN RF XGB MLP AB Ensemble
Accuracy 0.9088 0.8616 0.9474 0.9491 0.9432 0.9485 0.9499 0.9507
Kapa Stat. 0.0956 0.1701 0.0061 0.0097 0.0683 0.0152 0.0044 0.0060
Precision 0.1322 0.1572 0.0833 0.1333 0.2029 0.1500 0.1111 0.2000
Recall 0.1566 0.4217 0.0080 0.0080 0.0562 0.0120 0.0040 0.0040
AUC-ROC 0.5520 0.6529 0.5018 0.5027 0.5225 0.5043 0.5012 0.5016
F1-score 0.1434 0.2290 0.0147 0.0152 0.0881 0.0223 0.0078 0.0079
Log Loss 3.2870 4.9869 1.8974 1.8339 2.0455 1.8551 1.8057 1.7775

[Link]
IMJSTP29120877 6301
International Multilingual Journal of Science and Technology (IMJST)
ISSN: 2528-9810
Vol. 8 Issue 5, May - 2023

TABLE IV: Performance Analysis of Different Classifiers in Balanced Dataset

Evaluation Classifiers
Metrics DT NB DT RF DT MLP DT Ensemble
Accuracy 0.8948 0.7974 0.8926 0.9343 0.9560 0.8362 0.8437 0.9522
Kapa Stat. 0.7895 0.5947 0.7852 0.8685 0.9120 0.6725 0.6873 0.9043
Precision 0.8792 0.7414 0.8332 0.9118 0.9567 0.8073 0.8224 0.9608
Recall 0.9152 0.9134 0.9817 0.9615 0.9552 0.8834 0.8766 0.9428
AUC-ROC 0.8948 0.7974 0.8926 0.9343 0.9560 0.8362 0.8437 0.9522
F1-score 0.8969 0.8184 0.9014 0.9360 0.9559 0.8436 0.8486 0.9517
Log Loss 3.7927 7.3036 3.8706 2.3690 1.5868 5.9022 5.6353 1.7240

TABLE V: Performance Analysis of Different Classifiers Based on Hyperparameter Tunning

Evaluation Classifiers
Metrics DT NB DT RF DT MLP DT Ensemble
Accuracy 0.8524 0.7997 0.9840 0.9001 0.9930 0.8921 0.8736 0.9990
Kapa Stat. 0.7048 0.5995 0.9679 0.8002 0.9860 0.7842 0.7472 0.9979
Precision 0.8154 0.7450 1.0000 0.8583 0.9940 0.8674 0.8603 0.9990
Recall 0.9111 0.9115 0.9679 0.9584 0.9920 0.9257 0.8920 0.9990
AUC-ROC 0.8524 0.7997 0.9840 0.9001 0.9930 0.8921 0.8736 0.9990
F1-score 0.8606 0.8199 0.9837 0.9056 0.9930 0.8956 0.8759 0.9990
Log Loss 5.3202 7.2184 0.5784 3.5999 0.2521 3.8891 4.5564 0.0371

TABLE VI: Comparative analysis of the proposed model with the other prior studies

Approach Accuracy Kapa Stat. Precision Recall AUROC F1-score Log Loss

Stacking [14] 0.9800 0.9740 0.9740 0.9890 0.9740

NB [15] 0.8200 0.7920 0.8570 0.8230


Ensemble
0.9990 0.9979 0.9990 0.9990 0.9990 0.9990 0.0371
(Proposed Model)
Glucose Level, that are the same for both log-based
IV. DISCUSSION relationships and ML techniques. Our research
Several studies utilizing stroke datasets have been indicates that important characteristics are adequate
conducted, but stroke prediction still requires for identifying a stroke, which will make more
substantial refinement. In our investigation, we accessible the execution of stroke diagnosis.
collected a stroke dataset and used the SMOTE Table VI compares the proposed model with
method to balance it. After transforming the features pertinent prior findings. Dritsas et al. [14] employed
using Standard Scalar, we applied DT, NB, KNN, RF, stacking approach to achieve the highest levels of
XGB, MLP, AB, and Ensemble classifiers. The grid accuracy (98.00%), precision (97.40%), recall
search approach was then utilized for tuning the (97.40%), AUCROC (98.90%), and F1-score
hyperparameter of each classifier. Here, we found (97.40%). In a separate study, Sailasya et al. [15]
that NB, KNN, XGB, MLP, AB, and Ensemble utilized NB to obtain the best accuracy (82.0%),
classifiers improved performance. In contrast to other precision (79.20%), recall (85.70%), and F1-score
classifiers, the ensemble method produced the most (82.30%). In our proposed framework, however, the
accurate results. Additionally, we used SHAP to implementation of trait balancing, transformation, and
explain the ML model’s output. We also ranked the hyperparameter optimization yielded the highest
features using IG, PC, and Relief F feature ranking accuracy (99.90%), KS (99.70%), precision (99.90%),
techniques. recall (99.90%), AUCROC (99.90%), F1-score
Our findings imply several crucial and relevant (99.90%), and log loss (0.0371).
characteristics for early stroke diagnosis. Depending
on the log-based relationship, the essential V. CONCLUSION
characteristics are Age, Heart Disease, Average
Glucose Level, Hypertension, and Marital Status. A stroke is a life-threatening condition that
Age, BMI, average glucose level, work type, and demands immediate care. The dataset was
smoking status are the most significant preprocessed in this study, and ML and statistical
characteristics in the case of ML models. In addition, methods were used to identify critical stroke patient
we identified critical metrics, such as Age and Mean diagnostic characteristics. Age, BMI, and average

[Link]
IMJSTP29120877 6302
International Multilingual Journal of Science and Technology (IMJST)
ISSN: 2528-9810
Vol. 8 Issue 5, May - 2023

glucose level are the three most important risk [9] A. N. Richter and T. M. Khoshgoftaar, “A review
factors for stroke. It was also found that our proposed of statistical and machine learning methods for
ensemble framework exhibited a high level of modeling cancer risk using structured clinical
classification accuracy (99.90%), KS (99.70%), data,” Artificial intelligence in medicine, vol. 90,
precision (99.90%), recall (99.90%), AUCROC pp. 1–14, 2018.
(99.90%), F1-score (99.90%), and log loss (0.0371), [10] C. R. Pereira, D. R. Pereira, S. A. Weber, C.
which suggests that our findings can be utilized for Hook, V. H. C. De Al- buquerque, and J. P. Papa,
computer-assisted medical diagnosis to assist “A survey on computer-assisted parkinson’s
healthcare professionals and physicians in examining disease diagnosis,” Artificial intelligence in
stroke in a cost-effective manner. Our research medicine, vol. 95, pp. 48–63, 2019.
enables the early identification of patients with a high [11] A. Kaya, “Cascaded classifiers and stacking
risk of stroke who require additional examinations methods for classification of pulmonary nodule
and treatment prior to the progression of the disease. characteristics,” Computer Methods and
This investigation’s ultimate objective is to enhance Programs in Biomedicine, vol. 166, pp. 77–89,
the ML architecture using deep learning techniques. 2018.
In order to assess the predictive potential of deep [12] O. R. Shishvan, D.-S. Zois, and T. Soyata,
learning algorithms for stroke incidence, we will also “Machine intelligence in healthcare and medical
acquire image data from CT and MRI imaging of the cyber physical systems: A survey,” IEEE Access,
brain. vol. 6, pp. 46 419–46 494, 2018.
[13] C. Colak, E. Karaman, and M. G. Turtay,
DATA AVAILABILITY
“Application of knowledge discovery process on
The dataset is publicly available. the prediction of stroke,” Computer methods and
programs in biomedicine, vol. 119, no. 3, pp.
CONFLICTS OF INTEREST 181– 85, 2015.
The authors declare that they have no conflicts of [14] E. Dritsas and M. Trigka, “Stroke risk prediction
interest. with machine learning techniques,” Sensors, vol.
22, no. 13, p. 4670, 2022.
REFERENCES [15] G. Sailasya and G. L. A. Kumari, “Analyzing the
[1] B. Delpont, C. Blanc, G. Osseby, M. Hervieu performance of stroke prediction using ml
B`egue, M. Giroud, and Y. B ́ejot, “Pain after classification algorithms,” International Journal of
stroke: a review,” Revue neurologique, vol. 174, Advanced Computer Science and Applications,
no. 10, pp. 671–674, 2018. vol. 12, no. 6, 2021.
[2] S. Kumar, M. H. Selim, and L. R. Caplan, [16] S. Rahman, M. Hasan, and A. K. Sarkar,
“Medical complications after stroke,” The Lancet “Prediction of brain stroke using machine
Neurology, vol. 9, no. 1, pp. 105–118, 2010. learning algorithms and deep neural network
[3] “Learn about stroke,” World Stroke Organization, techniques,” European Journal of Electrical
accessed: 20 Mar 2023. [Online]. Available: Engineering and Computer Science, vol. 7, no. 1,
[Link] world-stroke-day pp. 23–30, 2023.
campaign/why-stroke-matters/learn-about-stroke. [17] S. Dev, H. Wang, C. S. Nwosu, N. Jain, B.
[4] T. Elloker and A. J. Rhoda, “The relationship Veeravalli, and D. John, “A predictive analytics
between social support and participation in approach for stroke prediction using machine
stroke: a systematic review,” African Journal of learning and neural networks,” Healthcare
Disability, vol. 7, no. 1, pp. 1–9, 2018. Analytics, vol. 2, p. 100032, 2022.
[5] “Causes and risk factors,” National Heart, Lung, [18] T. Liu, W. Fan, and C. Wu, “A hybrid machine
and Blood Institute, accessed: 20 Mar 2023. learning approach to cerebral stroke prediction
[Online]. Available: [Link] based on imbalanced medical dataset,” Artificial
health/stroke/causes intelligence in medicine, vol. 101, p. 101723,
[6] J. D. Pandian, S. L. Gall, M. P. Kate, G. S. Silva, 2019.
R. O. Akinyemi, B. I. Ovbiagele, P. M. Lavados, [19] P. Govindarajan, R. K. Soundarapandian, A. H.
D. B. Gandhi, and A. G. Thrift, “Prevention of Gandomi, R. Patan, P. Jayaraman, and R.
stroke: a global perspective,” The Lancet, vol. Manikandan, “Classification of stroke disease
392, no. 10154, pp. 1269–1278, 2018. using machine learning algorithms,” Neural
[7] V. L. Feigin, B. Norrving, M. G. George, J. L. Computing and Applications, vol. 32, pp. 817–
Foltz, G. A. Roth, and G. A. Mensah, “Prevention 828, 2020.
of stroke: a strategic global imperative,” Nature [20] M. J. Uddin, M. M. Ahamad, P. K. Sarker, S.
Reviews Neurology, vol. 12, no. 9, pp. 501–512, Aktar, N. Alotaibi, S. A. Alyami, M. A. Kabir, and
2016. M. A. Moni, “An integrated statistical and
[8] I. Yoo, P. Alafaireet, M. Marinov, K. Pena clinically applicable machine learning framework
Hernandez, R. Gopidi, J.- F. Chang, and L. Hua, for the detection of autism spectrum disorder,”
“Data mining in healthcare and biomedicine: a Computers, vol. 12, no. 5, p. 92, 2023.
survey of the literature,” Journal of medical [21] M. M. Ahamad, S. Aktar, M. J. Uddin, T.
systems, vol. 36, pp. 2431– 2448, 2012. Rahman, S. A. Alyami, S. Al-Ashhab, H. F.

[Link]
IMJSTP29120877 6303
International Multilingual Journal of Science and Technology (IMJST)
ISSN: 2528-9810
Vol. 8 Issue 5, May - 2023

Akhdar, A. Azad, and M. A. Moni, “Early-stage [28] S. Wan, Y. Liang, Y. Zhang, and M. Guizani,
detection of ovarian cancer based on clinical data “Deep multi-layer perceptron classifier for
using machine learning approaches,” Journal of behavior analysis to estimate parkinson’s
Personalized Medicine, vol. 12, no. 8, p. 1211, disease severity using smartphones,” IEEE
2022. Access, vol. 6, pp. 36 825–36 833, 2018.
[22] T. Akter, M. S. Satu, M. I. Khan, M. H. Ali, S. [29] A. Natekin and A. Knoll, “Gradient boosting
Uddin, P. Lio, J. M. Quinn, and M. A. Moni, machines, a tutorial,” Frontiers in neurorobotics,
“Machine learning-based models for early stage vol. 7, p. 21, 2013.
detection of autism spectrum disorders,” IEEE [30] R. Rojas et al., “Adaboost and the super bowl of
Access, vol. 7, pp. 166 509–166 527, 2019. classifiers a tutorial introduction to adaptive
[23] G. Fang, P. Xu, and W. Liu, “Automated boosting,” Freie University, Berlin, Tech. Rep,
ischemic stroke subtyping based on machine 2009.
learning approach,” IEEE Access, vol. 8, pp. 118 [31] M. M. Ahamad, S. Aktar, M. J. Uddin, M.
426– 118 432, 2020. Rashed-Al-Mahfuz, A. Azad, S. Uddin, S. A.
[24] S. M. Hasan, M. P. Uddin, M. Al Mamun, M. I. Alyami, I. H. Sarker, A. Khan, P. Li`o et al.,
Sharif, A. Ulhaq, and G. Krishnamoorthy, “A “Adverse effects of covid-19 vaccination:
machine learning framework for early-stage machine learning and statistical approach to
detection of autism spectrum disorders,” IEEE identify and classify incidences of morbidity and
Access, 2022. postvaccination reactogenicity,” in Healthcare,
[25] K. P. Murphy et al., “Naive bayes classifiers,” vol. 11, no. 1. MDPI, 2022, p. 31.
University of British Columbia, vol. 18, no. 60, pp. [32] T. Akter, M. H. Ali, M. I. Khan, M. S. Satu, M. J.
1 8, 2006. Uddin, S. A. Alyami, S. Ali, A. Azad, and M. A.
[26] L. E. Peterson, “K-nearest neighbor,” Moni, “Improved transfer-learning-based facial
Scholarpedia, vol. 4, no. 2, p. 1883, 2009. recognition framework to detect autistic children
[27] M. J. Vowels, “Trying to outrun causality with at an early stage,” Brain Sciences, vol. 11, no. 6,
machine learning: Limitations of model p. 734, 2021.
explainability techniques for identifying predictive
variables,” arXiv preprint arXiv:2202.09875,
2022.

[Link]
IMJSTP29120877 6304

You might also like