Machine Learning for Heart Disease Prediction
Machine Learning for Heart Disease Prediction
A PROJECT REPORT
submitted
in the partial fulfillment of the requirements for the award of the degree of
BACHELOR OF TECHNOLOGY
in
By
SATHWIKA KISTIPATI (19B81A05D7)
MASKU NELSON SHARON PREKSHA (19B81A05F0)
SATWIK YARABATI (19B81A05H8)
Mrs. G. Sandhya
Assistant Professor
i
ACKNOWLEDGEMENT
The satisfaction that accompanies the successful completion of any task would be incom-
plete without the mention of the people who made it possible and whose encouragement
and guidance has been a source of inspiration throughout the course of the project.
It is a great pleasure to convey our profound sense of gratitude to our principal Dr. K.
Ramamohan Reddy, Vice-Principal Prof. L. C. Siva Reddy, Dr. A. Vani Vathsala,
Head of CSE Department, CVR College of Engineering, for having been kind enough to
arrange necessary facilities for executing the project in the college.
We deem it a pleasure to acknowledge our sense of gratitude to our project guide Mrs. G.
Sandhya under whom we have carried out the project work. His incisive and objective
guidance and timely advice encouraged us with constant flow of energy to continue the
work.
We wish a deep sense of gratitude and heartful thanks to the management for providing
excellent lab facilities and tools. Finally, we thank all those whose guidance helped us in
this regard.
SATHWIKA KISTIPATI(19B81A05D7)
MASKU NELSON SHARON PREKSHA(19B81A05F0)
SATWIK YARABATI(19B81A05H8)
ii
DECLARATION
We the undersigned solemnly declare that the project report titled ‘PERFORMANCE
EVALUATION OF DIFFERENT MACHINE LEARNING TECHNIQUES FOR PRE-
DICTION OF CARDIOVASCULAR DISEASE’ is based on our own work carried out
during the course of our study under the supervision of Mrs. G. Sandhya, CSE Dept.,
CVR College of Engineering.
We assert the statements made and conclusions are drawn are an outcome of our research
work. We further certify that
1. The work contained in the report is original and has been done by us under the general
supervision of our supervisor.
2. The work has not been submitted to any other Institution for any other degree/di-
ploma/certificate in this university or in the any other University of India or abroad.
3. We have followed the guidelines provided by the university in writing the report.
iii
CERTIFICATE
This is to certify that the project entitled “Performance Evaluation Of Different Machine
Learning Techniques For Prediction Of Cardiovascular Disease ” being submitted by
Sathwika Kistipati (19B81A05D7), Masku Nelson Sharon Preksha (19B81A05F0), Satwik
Yarabati (19B81A05H8) in partial fulfillment for the award of Bachelor of Technology in
Com- puter Science and Engineering to the CVR College of Engineering , is a record of bona
fide work carried out by them under my guidance and supervision during the year 2020-2021.
The results embodied in this project work have not been submitted to any other University or
Institute for the award of any degree or diploma.
iv
Abstract
Machine Learning has vast applications in different domains and healthcare industry is one of
the most popular ones. Machine Learning can play an essential role in prediction of various
diseases. Such information, if predicted well in advance, can provide important intuitions to
doctors who can then adapt their diagnosis and dealing per patient basis. Heart related diseases
or Cardiovas- cular Diseases (CVDs) are the main reasons for a huge number of deaths in the
world over the last few decades and has emerged as the one of the most life-threatening
diseases, not only in India but also in other parts of the world. So, there is a need of a reliable,
accurate and feasible system to diagnose such diseases in time for proper treatment based on the
patients medical his- tory. So, This project aims at predicting the possibility of Heart Disease in
people using Machine Learning algorithms. In this project we perform the comparative analysis
of various classifiers and propose a best fit, precise and efficient model for heart disease
prediction in order to enhance medical care and reduce cost.
v
TABLE OF CONTENTS
Table of Contents
Page No.
List of Tables Vi
Abbreviations Ix
1 Introduction 1
1.1 Motivation 2
2 Literature Survey 5
vi
4.1 Proposed Methods 12
5.4 Modules 37
6.1 Conclusion 44
7 References 46
vii
List of Figures
viii
List of Abbreviations
ML Machine Learning
ix
1. INTRODUCTION
Predicting a disease which depends on several factors is indeed not an easy task. To predict
whether a heart is healthy or not involves studying the variation in several contributing factors.
The complexity and the ambiguity can be reduced to a greater extent by processing the data and
executing the algorithm in a proper way. Several techniques and methods have evolved in the
field of Machine Learning for handling the datasets that contribute to the development of
predic- tion and forecasting ML models. We employed the Cleveland dataset of UCI repository
for tech- nical and experimental analysis of the project. This dataset contains the heart data of
303 individ- uals which are classified into 14 columns, this part is discussed in the later sections
of the paper. Heart Disease is predicted based on the variation of these attributes with respect to
the normalcy. There are many cases and evidences that Machine Learning has proven to be very
efficient in tackling the real-world problems and generating optimal and efficient solutions. ML,
in today’s world is being deployed to develop numerous applications. The fascination of the
Machine Learn- ing and Artificial Intelligence for the people is very high and this leads to the
development of various optimal applications for various industries.
Our motivation to work on this project is to put in the efforts of studying the in-depth
applications of the Machine Learning models and learn the evolution of ML in the field of
healthcare. We propose an approach to generate the efficient output from the model by
comparing the attributes such as accuracy, sensitivity, precision of various algorithms. Many
papers and projects are pub- lished which were focused on improving accuracy and efficiency of
heart disease prediction. Our objective is to take the reference of such papers and develop our
own method of model as an application of the subject in graduate level. Here, we did a
comparative analysis of the Machine Learning models.
We employed Logistic Regression, Decision Tree, Naïve Baye’s Theorem, K-NN, Neural Net-
work, Support Vector Machine, Random Forest models and subjected them to the datasets for
analysis. Based on the comparison of the accuracy achieved by the models, the output is
generated from the model that has high accuracy. The programming language we employed for
the project is Python. Python has proven to be very efficient in developing such applications
1
with the help of
2
its vast number of libraries and functions. The code for the models is written on a file and the
application of python is used for developing the comparison analysis
1.1 MOTIVATION
In the last fifteen years, HDs become still endured the leading causes of death. (WHO,
2019). In United States, Millions of humans are having a HD every year so that the HD takes
placed the biggest killer of people in the world. According to analyzation of WHO, twelve
million people are death due to HD in worldwide. One person dies almost every 34 seconds
from HD. (Patel Jaymin, 2016) Diagnosis of HDs is an essential task and yet intricate task to
perform accu- rately and efficiently in the hospital and clinic. These things are motivated to
build a HDPS ap- plication using ML algorithms. This proposed system can reserve problem by
accurately predict- ing the presence of HD in the patient.
3
1.2 PROBLEM STATEMENT
With the consideration of WHO statistical facts, the most powerful causes of death glob-
ally are a HD. It seemed to the negligence of patients as well as doctors to increase a HD
patient. Some of the difficulties to execute the doctor’s decision and lack of application to
clearly diag- nosis of HD become the cause of human death. Regarding the above issues, we are
proposing a HDPS using set of 7 algorithms, that is one of best solutions to efficiently and
accurately predict the HD patients. The proposed system eliminates the various testing of HD
and supports the de- cision making of doctors. This system can accept a singleton query and
display the clear output of the presence of HD level. This system is useful for any hospital and
clinic to evaluate the patient getting HD. It is reduced the number of tests and provide an
efficient output of patient HD. It supports to make the decision of doctors that consulates with
their patients easily.
The proposed system will support the healthcare systems as well as health-related application to
expand their services with efficiently and accurately providing results. It mitigates the time to
checkup of doctors.
Our motivation to work on this project is to put in the efforts of studying the in-depth
applications of the Machine Learning models and learn the evolution of ML in the field of
healthcare. We propose an approach to generate the efficient output from the model by
comparing the attributes such as accuracy, sensitivity, precision of various algorithms. Many
papers and projects are pub- lished which were focused on improving accuracy and efficiency of
heart disease prediction. Our objective is to take the reference of such papers and develop our
own method of model as an application of the subject in graduate level. Here, we did a
comparative analysis of the Machine Learning models.
We employed Logistic Regression, Decision Tree, Naïve Baye’s Theorem, K-NN, Neural Net-
work, Support Vector Machine, Random Forest models and subjected them to the datasets for
analysis. Based on the comparison of the accuracy achieved by the models, the output is
generated from the model that has high accuracy. The programming language we employed for
the project is Python. Python has proven to be very efficient in developing such applications
with the help of its vast number of libraries and functions. The code for the models is written on
4
a file and the application of python is used for developing the comparison analysis.
5
1.3 PROJECT OBJECTIVE
To design and implement Machine Learning based algorithm using techniques such as
Regression and Classification, for predicting the heart disease of patients using the given
attributes.
The proposed work makes an attempt to detect these heart diseases at early stage to avoid dan-
gerous consequences.
6
2. LITERATURE SURVEY
Author Names: Rairikar, A., Kulkarni, V., Sabale, V., Kale, H., & Lamgunde,
Description: In this work, three data mining classification algorithms like Random Forest, Deci-
sion Tree and Naïve Bayes are addressed and used to develop a prediction system in order to
analyse and predict the possibility of heart disease. The main objective of this significant
research work is to identify the best classification algorithm suitable for providing maximum
accuracy when classification of normal and abnormal person is carried out. Thus prevention of
the loss of lives at an earlier stage is possible. The experimental setup has been made for the
evaluation of the performance of algorithms with the help of heart disease benchmark dataset
retrieved from UCI machine learning repository. It is found that Random Forest algorithm
performs best with 81% precision when compared to other algorithms for heart disease
prediction.
Description: This paper proposes a HDPS based on three different data mining techniques. The
various data mining methods used are Naive Bayes, Decision tree (J48), Random Forest and
WEKA API. The system can predict the likelihood of patients getting a heart disease by using
medical profiles such as age, sex, blood pressure, cholesterol and blood sugar. Also, the perfor-
mance will be compared by calculation of confusion matrix. This can help to calculate accuracy,
precision, and recall. The overall system provides high performance and better accuracy.
7
Title: Prediction and Analysis the occurrence of Heart Disease using data mining techniques.
overcome the problem of diagnosis of HD. It improved the existence methodology by choosing
Naïve Bayes, J48, and SVM for predicting the occurrence of HD for early automatic diagnosis
in short time in order to support the qualities of services and reduce costs to save the life of
patient has HD or not. The comparison of analysis in the dataset is used WEKA software.
Author Name: Shetty, Deeraj, Kishor Rit, Sohail Shaikh, and Nikita Patil
Description: The goal of the data mining methodology is to think data from a data set and
change it into a reasonable structure for further use. Our examination concentrates on this part
of Medical conclusion learning design through the gathered data of diabetes and to create smart
therapeutic choice emotionally supportive network to help the physicians. The primary target of
this exami- nation is to assemble Intelligent Diabetes Disease Prediction System that gives
analysis of diabe- tes malady utilizing diabetes patient's database. In this system, we propose the
use of algorithms like Bayesian and KNN (K-Nearest Neighbor) to apply on diabetes patient's
database and analyze them by taking various attributes of diabetes for prediction of diabetes
disease.
8
2.2 LIMITATIONS OF EXISTING WORK
The system also needs to be accurate and efficient in the prediction of heart disease. Various
machine learning algorithms were used for disease prediction; some of them are SVM, Naïve
Bayes etc. But the results state that there may be some improvements to be done on terms of
accuracy and the prediction rate and the false positive rate. Medical Misdiagnosis is a serious
risk in the healthcare industry. If this persists, it might instill fear amongst the patients. Some
other techniques can replace previously applied techniques such as SVM and Naïve Bayes.
Also, the study states that the dataset also can be improved by using some methods over it. To
improve the quality of the input to the proposed system. Using KNN, we have certain weakness
such as they suffer from high dimensionality and overfitting. Naive Bayes assumes that all
features are inde- pendent, which rarely happens in real life. Most of these studies are theoretical
analysis at the macro level and lack quantitative investigation.
9
3. SOFTWARE & HARDWARE SPECIFICATIONS
What is SRS?
Problem/Requirement Analysis:
The process is order and more nebulous of the two, deals with understand the problem,
the goal and constraints.
Requirement Specification:
Here, the focus is on specifying what has been found giving analysis such as
representation, Specification languages and tools, and checking the specifications are addressed
during this activity.
The requirement phase terminates with the production of the validate SRS docu-
ment. Producing the SRS document is the basic of this phase.
Role of SRS:
The purpose of the SRS is to reduce the communication gap between the clients and the develop-
ers. SRS is the medium though which the client and user needs are accurately specified.
It forms the basis of software development. A good SRS should satisfy all the
parties involved in the system.
1
Purpose:
The purpose of this document is to describe all external requirements for the E-
learning System. It also describes the interfaces for the system.
Scope:
This document is the only one that describes the requirements of the system. It is
meant for the use by the developers, and will also by the basis for validating the final deliver
system. Any changes made to the requirements in the future will have to go through a formal
change approval process. The developer is responsible for asking for clarifications, where
necessary, and will not make any alternations without the permission of the client.
Overview:
The SRS begins the translation process that converts the software Requirements into
the language the developers will use. The SRS draws on the Use Cases from the user
Requirement Document and analyses the situations from a number of perspectives to discover
and eliminate inconsistencies, ambiguities and omissions before development progresses
significantly under mistaken assumptions.
Requirements:
Requirement Analysis:
Taking into account the comparative analysis stated in the previous section we could start speci-
fying the requirements that our website should achieve. As a basis, an article on all the different
1
requirements for software development was taken into account during this process. We divide
the requirements in 2 types: functional and nonfunctional requirements.
Functional requirement should include function performed by a specific screen outline work-
flows performed by the system and other business or compliance requirement the system must
meet.
The functional specification describes what the system must do, how the system does it is de-
scribed in the design specification.
Functional Requirements:
1) Loading Dataset
2) Data cleaning
3) Data Transformation
4) Data Visualization
5) Choosing Algorithm
6) Train Dataset
7) Validate Dataset
8) Record Accuracy
9) Create Model
10) Predict Outcome
Describe user-visible aspects of the system that are not directly related with the functional
behav- ior of the system. Non-Functional requirements include quantitative constraints, such as
response time (i.e. how fast the system reacts to user commands.) or accuracy (.e. how precise
are the systems numerical answers.).
Portability
Reliability
1
Usability
Time Constraints
Error messages
Actions which cannot be undone should ask for confirmation
Responsive design should be implemented
Space Constraints
Performance
Standards
Ethics
Interoperability
Security
Privacy
Scalability
1
4. PROPOSED SYSTEM DESIGN
The objective of this research is to design a robust machine learning algorithm to predict heart
disease. The prediction of heart disease is performed using Ensemble of machine learning algo-
rithms. This is to boost the accuracy achieved by individual machine learning algorithms.
Methods: Heart Disease Prediction System is developed where the user can input the patient de-
tails and the prediction for the particular patient is made using the model developed. The model
will predict the output to be either normal or risky. Regression Trees (CART), Support Vector
Machines (SVM), K-Nearest Neighbors (KNN) and Naïve Bayes classifier are used as base
learn- ers. These algorithms are combined using random forest as the Meta classifier.
Logistic Regression:
Regression means to predict the value using the input data. Regression models are used
to predict a continuous value. It is mostly used to find the relationship between the
variables and forecasting (functional dependency). Regression models differ based on
the kind of relationship between dependent and independent variables. Logistic
Regression is a tech- nique to establish CATEGORICAL classification of data. Figure
4.4 represents the sche- matic of the Logistic Regression Model. The equation is
generally stated as:
E(y) = P(y=1/ x1, x2,.................xp)
1
Figure 4.1.1: Schematic of Logistic Regression
1
Naïve Baye’s Theorem:
Naïve Bayes classifier is a supervised algorithm. It is a simple classification technique
using Bayes theorem. It assumes strong (Naive) independence among attributes. Bayes
theorem is a mathematical concept to get the probability. The predictors are neither
related to each other nor have correlation to one another. All the attributes independently
con- tribute to the probability to maximize it. Naïve Bayes is a simple, easy to
implement, and efficient classification algorithm that handles non-linear, complicated
data. However, there is a loss of accuracy as it is based on assumption and class
conditional independence. The equation is represented as follows.
K- Nearest Neighbors :
The K- Nearest Neighbors algorithm is a supervised classification algorithm method. It
classifies objects dependent on nearest neighbor. It is a type of instance-based learning.
KNN usually works by just trying to see to which class is the new feature near to and it
just puts it to the class closest to that point. K-NN algorithm is simple to carry out
without creating a model or making other assumptions. This algorithm is versatile and is
used for classification, regression, and search. Even though K-NN is the simplest
algorithm, noisy and irrelevant features affect its accuracy. Figure depicts the schematic
representation of K-NN algorithm:
1
Figure 4.1.3: KNN Classifier
Random Forest:
It is a popular machine leaning algorithm that belongs to the supervised ML technique. It
is based on the concept of ensemble learning, which is a process of combining multiple
classifiers to solve a complex problem and to improve the performance of the model.
Ran- dom forest is a classifier that contains a number of decision trees on various
subsets of the given dataset and takes the average to improve the predictive accuracy of
that dataset. The figure depicts the flow of Random Forest classifier.
1
Support Vector Machine:
The goal of the SVM algorithm is to create the best line or decision boundary that can
segregate n-dimensional space into classes so that we can easily put the new data point
in the correct category in future. This best decision boundary is called a hyperplane.
SVM chooses the extreme points that help in creating the hyperplane. These extreme
cases are called as support vectors. Figure depicts the general representation of SVM.
XGBoosting:
XGBoost which stands for Extreme Gradient Boosting, is a scalable, distributed
gradient- boosted decision tree (GBDT) machine learning library. It provides parallel
tree boosting and is the leading machine learning library for regression, classification,
and ranking prob- lems.
It’s vital to an understanding of XGBoost to first grasp the machine learning concepts
and algorithms that XGBoost builds upon: supervised machine learning, decision trees,
en- semble learning, and gradient [Link] machine learning uses algorithms
to train a model to find patterns in a dataset with labels and features and then uses the
trained model to predict the labels on a new dataset’s features.
1
Decision Tree:
Decision tree algorithm falls under the category of supervised learning. They can be
used to solve both regression and classification problems.
Decision tree uses the tree representation to solve the problem in which each leaf node
corresponds to a class label and attributes are represented on the internal node of the
tree. We can represent any Boolean function on discrete attributes using the decision tree
UML DIAGRAMS
UML stands for Unified Modeling Language. UML is a standardized general-purpose mod-
eling language in the field of object-oriented software engineering. The standard is
managed, and was created by, the Object Management Group.
The goal is for UML to become a common language for creating models of object oriented
computer software.
1
4.2 USE CASE DIAGRAM:
A use case diagram in the Unified Modeling Language (UML) is a type of behavioral diagram
defined by and created from a Use-case analysis. The main purpose of a use case diagram is to
show what system functions are performed for which actor. Roles of the actors in the system
can be depicted.
2
4.3 SEQUENCE DIAGRAM:
Sequence Diagrams Represent the objects participating the interaction horizontally and time ver-
tically.
2
4.4 ACTIVITY DIAGRAM:
Activity diagrams are graphical representations of Workflows of stepwise activities and actions
with support for choice, iteration and concurrency. In the Unified Modeling Language, activity
diagrams can be used to describe the business and operational step-by-step workflows of com-
ponents in a system. An activity diagram shows the overall flow of control.
2
4.5 CLASS DIAGRAM:
In software engineering, a class diagram in the Unified Modeling Language (UML) is a type of
static structure diagram that describes the structure of a system by showing the system's
classes, their attributes, operations (or methods), and the relationships among the classes. It
explains which class contains information.
Component diagrams are used to model the physical aspects of a system. Physical aspects are
the elements such as executables, libraries, files, documents, etc. which reside in a node.
Component diagrams are used to visualize the organization and relationships among
components in a system.
It does not describe the functionality of the system but it describes the components used to make
those functionalities.
2
Component diagrams can also be described as a static implementation view of a system. Static
implementation represents the organization of the components at a particular moment.
A single component diagram cannot represent the entire system but a collection of diagrams is
used to represent the whole.
2
4.7 DEPLOYMENT DIAGRAM
There may be more steps involved, depending on what specific requirements you have, but
below are some of the main steps:
The proposed system is built around conventional three-tier architecture. The three-tier
architecture for web development allows programmers to separate various aspects of the
solution design into modules and work on them separately. That is, a developer who is best at
one part of development, say UI development need not worry about the implementation levels
so much. It also allows for easy maintenance and future enhancements. The three-tiers of the
solution include:
The Layout:
This tier is at the uppermost layer and is closely bound to the user, i.e., the users of the
system interact with it through this tier.
2
The business-tier:
This tier is responsible for implementing all the business rules of the organization. It
operates on the data provided by the users through the web-tier and the data stored in the
underlying data-tier. So in a way this tier works on data from the web-tier and the data-tier
in order to perform task for the users in agreement with the business rules of the
organization.
The data-tier:
This tier contains the persist able data that is required by the business tier to operate on. Data
plays a very important role in the functioning of any organization. Thus, persisting of such
data is very important. The data tier performs the job of persisting the data.
About python:
The Python programming language is an Open Source, cross-platform, high level, dynamic, in-
terpreted language.
The Python 'philosophy' emphasizes readability, clarity and simplicity, whilst maximizing the
power and expressiveness available to the programmer. The ultimate compliment to a Python
programmer is not that his code is clever, but that it is elegant. For these reasons Python is an
excellent 'first language', while still being a powerful tool in the hands of the seasoned and
cynical programmer.
Python is a very flexible language. It is widely used for many different purposes. Typical uses
include :
Web application programming with frameworks like Zope, Django and Turbogears
System administration tasks via simple scripts
Desktop applications using GUI toolkits like Tkinter or wxPython (and recently
Windows Forms and IronPython)
Creating windows applications, using the Pywin32 extension for full windows integration
and possibly Py2exe to create standalone programs
23
Scientific research using packages like Scipy and Matplotlib.
jupyter notebook:
The Jupyter Notebook is an open-source web application that allows you to create and share
doc- uments that contain live code, equations, visualizations and narrative text. Uses include:
data cleaning and transformation, numerical simulation, statistical modeling, data visualization,
ma- chine learning, and much more.
jupyter Notebooks are a powerful way to write and iterate on your Python code for data analysis.
Rather than writing and re-writing an entire program, you can write lines of code and run them
one at a time. Then, if you need to make a change, you can go back and make your edit and
rerun the program again, all in the same window.
Jupyter Notebook is built off of IPython, an interactive way of running Python code in the
termi- nal using the REPL model (Read-Eval-Print-Loop). The IPython Kernel runs the
computations and communicates with the Jupyter Notebook front-end interface. It also allows
Jupyter Notebook to support multiple languages. Jupyter Notebooks extend IPython through
additional features, like storing your code and output and allowing you to keep markdown notes.
2
5. IMPLEMENTATION AND TESTING:
Data Processing is a task of converting data from a given form to a much more usable and
desired form i.e. making it more meaningful and informative. Using Machine Learning
algorithms, math- ematical modeling and statistical knowledge, this entire process can be
automated. The output of this complete process can be in any desired form like graphs, videos,
charts, tables, images and many more, depending on the task we are performing and the
requirements of the machine. This might seem to be simple but when it comes to really big
organizations like Twitter, Facebook, Administrative bodies like Paliament, UNESCO and
health sector organizations, this entire pro- cess needs to be performed in a very structured
manner.
2
Figure 5.1: Data Preprocessing
Collection:
The most crucial step when starting with ML is to have data of good quality and accuracy. Data
can be collected from any authenticated source like [Link], Kaggle or UCI dataset reposi-
tory.
A huge amount of capital, time and resources are consumed in collecting data. Organizations or
researchers have to decide what kind of data they need to execute their tasks or research.
Preparation:
The collected data can be in a raw form which can’t be directly fed to the machine. So, this is a
process of collecting datasets from different sources, analyzing these datasets and then
construct- ing a new dataset for further processing and exploration. This preparation can be
performed either manually or from the automatic approach. Data can also be prepared in
2
numeric forms also which
2
would fasten the model’s learning.
Input:
Now the prepared data can be in the form that may not be machine-readable, so to convert this
data to readable form, some conversion algorithms are needed. For this task to be executed, high
computation and accuracy is needed.
Processing:
This is the stage where algorithms and ML techniques are required to perform the instructions
provided over a large volume of data with accuracy and optimal computation.
Output:
In this stage, results are procured by the machine in a meaningful manner which can be inferred
easily by the user. Output can be in the form of reports, graphs, videos, etc
Storage:
This is the final step in which the obtained output and the data model data and all the useful
information are saved for the future use.
• Pre-processing refers to the transformations applied to our data before feeding it to the algorithm.
• Data Preprocessing is a technique that is used to convert the raw data into a clean data set. In
other words, whenever the data is gathered from different sources it is collected in raw format
which is not feasible for the analysis.
• For achieving better results from the applied model in Machine Learning projects the format of
the data has to be in a proper manner. Some specified Machine Learning model needs
information in a specified format, for example, Random Forest algorithm does not support null
values, there- fore to execute random forest algorithm null values have to be managed from the
original raw data set.
• Another aspect is that data set should be formatted in such a way that more than one Machine
Learning and Deep Learning algorithms are executed in one data set, and best out of them is
chosen.
Rescale Data
When our data is comprised of attributes with varying scales, many machine
learn- ing algorithms can benefit from rescaling the attributes to all have the
same scale.
We can rescale your data using scikit-learn using the MinMaxScaler class.
We can transform our data using a binary threshold. All values above the
threshold are marked 1 and all equal to or below are marked as 0.
We can create new binary attributes in Python using scikit-learn with the
Binarizer class.
Standardize Data
2
5.2 DATA CLEANING
Data cleaning is one of the important parts of machine learning. It plays a significant part in
building a model. Data Cleaning is one of those things that everyone does but no one really talks
about. Professional data scientists usually spend a very large portion of their time on this step.
If we have a well-cleaned dataset, we can get desired results even with a very simple algorithm,
which can prove very beneficial at times.
This includes deleting duplicate/ redundant or irrelevant values from your dataset. Duplicate ob-
servations most frequently arise during data collection and irrelevant observations are those that
don’t actually fit the specific problem that you’re trying to solve.
3
Redundant observations alter the efficiency by a great extent as the data repeats and may
add towards the correct side or towards the incorrect side, thereby producing unfaithful
results.
Irrelevant observations are any type of data that is of no use to us and can be removed
directly.
2. Fixing Structural errors
The errors that arise during measurement transfer of data or other similar situations are called
structural errors. Structural errors include typos in the name of features, same attribute with dif-
ferent name, mislabeled classes, i.e. separate classes that should really be the same or
inconsistent capitalization.
Outliers can cause problems with certain types of models. For example, linear regression models
are less robust to outliers than decision tree models. Generally, we should not remove outliers
until we have a legitimate reason to remove them. Sometimes, removing them improves perfor-
mance, sometimes not. So, one must have a good reason to remove the outlier, such as
suspicious measurements that are unlikely to be the part of real data.
Missing data is a deceptively tricky issue in machine learning. We cannot just ignore or remove
the missing observation. They must be handled carefully as they can be an indication of
something important. The two most common ways to deal with missing data are:
After properly completing the Data Cleaning steps, we’ll have a robust dataset that avoids many
of the most common pitfalls. This step should not be rushed as it proves very beneficial in the
further process.
3
5.3 MACHINE LEARNING PROCESS
Machine Learning algorithm is trained using a training data set to create a model. When new
input data is introduced to the ML algorithm, it makes a prediction on the basis of the model.
The prediction is evaluated for accuracy and if the accuracy is acceptable, the Machine Learning
algorithm is deployed. If the accuracy is not acceptable, the Machine Learning algorithm is
trained again and again with an augmented training data set.
The Machine Learning process involves building a Predictive model that can be used to find a
solution for a Problem Statement. To understand the Machine Learning process let’s assume
that you have been given a problem that needs to be solved by using Machine Learning.
Step 7: Predictions
3
Input Design and Output Design
Input Design
The input design is the link between the information system and the user. It comprises the
devel- oping specification and procedures for data preparation and those steps are necessary to
put trans- action data in to a usable form for processing can be achieved by inspecting the
computer to read data from a written or printed document or it can occur by having people
keying the data directly into the system. The design of input focuses on controlling the amount
of input required, control- ling the errors, avoiding delay, avoiding extra steps and keeping the
process simple. The input is designed in such a way so that it provides security and ease of use
with retaining the privacy. Input Design considered the following things:
Objectives
1. Input Design is the process of converting a user-oriented description of the input into a com-
puter-based system. This design is important to avoid errors in the data input process and show
the correct direction to the management for getting correct information from the computerized
system.
2. It is achieved by creating user-friendly screens for the data entry to handle large volume of
data. The goal of designing input is to make data entry easier and to be free from errors. The
data entry screen is designed in such a way that all the data manipulates can be performed. It
also provides record viewing facilities.
3. When the data is entered it will check for its validity. Data can be entered with the help of
screens. Appropriate messages are provided as when needed so that the user will not be in maize
of instant. Thus the objective of input design is to create an input layout that is easy to follow
3
Output Design
A quality output is one, which meets the requirements of the end user and presents the
information clearly. In any system results of processing are communicated to the users and to
other system through outputs. In output design it is determined how the information is to be
displaced for im- mediate need and also the hard copy output. It is the most important and direct
source information to the user. Efficient and intelligent output design improves the system’s
relationship to help user decision-making.
1. Designing computer output should proceed in an organized, well thought out manner; the
right output must be developed while ensuring that each output element is designed so that
people will find the system can use easily and effectively. When analysis design computer
output, they should Identify the specific output that is needed to meet the requirements.
3. Create document, report, or other formats that contain information produced by the system.
The output form of an information system should accomplish one or more of the following ob-
jectives.
Feasibility Study
The feasibility of the project is analyzed in this phase and business proposal is put
forth with a very general plan for the project and some cost estimates. During system
analysis the feasibility study of the proposed system is to be carried out. This is to ensure
that the proposed system is not a burden to the company. For feasibility analysis, some
understanding of the major requirements for the system is essential.
3
Three key considerations involved in the feasibility analysis are
ECONOMICAL FEASIBILITY
TECHNICAL FEASIBILITY
SOCIAL FEASIBILITY
Economical Feasibility
This study is carried out to check the economic impact that the system will have on the organi-
zation. The amount of fund that the company can pour into the research and development of the
system is limited. The expenditures must be justified. Thus the developed system as well within
the budget and this was achieved because most of the technologies used are freely available.
Only the customized products had to be purchased.
Technical Feasibilty
This study is carried out to check the technical feasibility, that is, the technical
requirements of the system. Any system developed must not have a high demand on the
available technical resources. This will lead to high demands on the available technical
resources. This will lead to high demands being placed on the client. The developed system
must have a modest re- quirement, as only minimal or null changes are required for
implementing this system.
Social Feasibility
The aspect of study is to check the level of acceptance of the system by the user. This
includes the process of training the user to use the system efficiently. The user must not feel
threatened by the system, instead must accept it as a necessity. The level of acceptance by the
users solely depends on the methods that are employed to educate the user about the system and
to make him familiar with it. His level of confidence must be raised so that he is also able to
make some constructive criticism, which is welcomed, as he is the final user of the system
3
Software Development Life Cycle:
The Systems Development Life Cycle (SDLC), or Software Development Life Cycle in systems
engineering, information systems and software engineering, is the process of creating or altering
systems, and the models and methodologies use to develop these systems.
Analysis gathers the requirements for the system. This stage includes a detailed study of the
busi- ness needs of the organization. Options for changing the business process may be
considered. Design focuses on high level design like, what programs are needed and how are
they going to interact, low-level design (how the individual programs are going to work),
interface design (what are the interfaces going to look like) and data design (what data will be
required). During these phases, the software's overall structure is defined. Analysis and Design
are very crucial in the whole development cycle. Any glitch in the design phase could be very
expensive to solve in the later stage of the software development. Much care is taken during this
phase. The logical system of the product is developed in this phase.
Implementation
In this phase the designs are translated into code. Computer programs are written using a
conven- tional programming language or an application generator. Programming tools like
Compilers, In-
3
terpreters, and Debuggers are used to generate the code. Different high level programming lan-
guages like PYTHON 3.6, Anaconda Cloud are used for coding. With respect to the type of ap-
plication, the right programming language is chosen.
Testing
In this phase the system is tested. Normally programs are written as a series of individual
modules, this subject to separate and detailed test. The system is then tested as a whole. The
separate mod- ules are brought together and tested as a complete system. The system is tested to
ensure that interfaces between modules work (integration testing), the system works on the
intended platform and with the expected volume of data (volume testing) and that the system
does what the user requires (acceptance/beta testing).
Maintenance
Inevitably the system will need maintenance. Software will definitely undergo change once it is
delivered to the customer. There are many reasons for the change. Change could happen
because of some unexpected input values into the system. In addition, the changes in the system
could directly affect the software operations. The software should be developed to
accommodate changes that could happen during the post implementation period.
System Design
System design is transition from a user oriented document to programmers or data base
personnel. The design is a solution, how to approach to the creation of a new system. This is
composed of several steps. It provides the understanding and procedural details necessary for
implementing the system recommended in the feasibility study. Designing goes through logical
and physical stages of development, logical design reviews the present physical system, prepare
input and output specification, details of implementation plan and prepare a logical design
walkthrough.
The database tables are designed by analyzing functions involved in the system and format of
the fields is also designed. The fields in the database tables should define their role in the
3
system. The unnecessary fields should be avoided because it affects the storage areas of the
system. Then in
3
the input and output screen design, the design should be made user friendly. The menu should
be precise and compact.
Software Design
1. Modularity and partitioning: software is designed such that, each system should consists of
hierarchy of modules and serve to partition into separate function.
4. Shared use: avoid duplication by allowing a single module be called by other that need the
function it provide
5.4 MODULES:
Data Source:
As I mentioned before the main data sets were collected from the Kaggle machine learning
repos- itory which is open data resource for the data mining and for the predictive analytics
purposes. The acquired data source was a CSV file
Data cleaning
The data in the CSV files need to be checked whether it has any missing values, as the data
source have missing values every attribute have checked using the filters and null values are
removed which helps to increase the accuracy level.
Technology Implementation:
The Real estate data is processed with the help of Python programming and Jupyter IDE and the
numerical data is normalized using the script so the data will equally distributed and variables
cannot dominate each other. It helps to bring the values of the attribute on a common scale.
3
Implementation of Random Forest:
Random Forest, data is been split into test and train 75 percent data is used for the test data and
remaining for the train data since I am using the Random forest. The prices have been predicted
and its shown in the graph. Here the average output of several trees has been taken in order to
provide more accuracy
Evaluation:
The idea of a regression is to predict a real value which means number in regression model we
can compute the several values
Visualization:
Here the test data and the predicted price are loaded into the tableau and it's clear that there is a
linear trend between the actual and predicted values so the random forest model performs well
and the accuracy percentage is more when we compared with the others.
Libraries:
3. matplotlib: To create charts using pyplot, define parameters using rcParams and color them
with [Link]
4. warnings: To ignore all warnings which might be showing up in the notebook due to
past/fu- ture depreciation of a feature
6. StandardScaler: To scale all the features, so that the Machine Learning model better adapts
to the dataset
4
SAMPLE CODE:
corr = [Link]()
[Link](figsize = (15,15))
[Link](corr, annot = True)
corr()
sns.set_style('whitegrid')
[Link](x = 'target', data = data)
dataset = [Link]()
[Link]()
X = [Link](['target'], axis = 1)
y = dataset['target']
[Link]
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size = 0.2, random_state = 42)
model = RandomForestClassifier(n_estimators=20)
[Link](X_train, y_train)
pred = [Link](X_test)
pred[:10]
confusion_matrix(y_test, pred)
print(f"Accuracy of model is {round(accuracy_score(y_test, pred)*100, 2)}%")
claasifier=RandomForestClassifier(n_jobs=-1, n_estimators=400,bootstrap= False,crite-
rion='gini',max_depth=5,max_features=3,min_samples_leaf= 7)
[Link](X_train, y_train)
confusion_matrix(y_test, [Link](X_test))
print(f"Accuracy is {round(accuracy_score(y_test, [Link](X_test))*100,2)}%")
[Link](classifier, open('[Link]', 'wb'))
4
5.5 WEBSITE IMPLEMENTATION:
4
Figure 5.5.3 Result Screen
4
5.6 RESULT & DISCUSSION
The aim of this project is to know whether the patient has heart disease or not. The records in
the datasets are divided into training set and test sets. After preprocessing the data, data mining
clas- sification technique namely Random Forest were applied. This section shows the results of
those classification model done using Python Programming. The results are generated for both
training datasets and test data sets.
Results: The predictions of classifier are combined using random forest algorithm. The accuracy
is lifted from 85.53% to 87.64% which is an impressive improvement on accuracy.
4
Test cases
4
6. . CONCLUSION AND FUTURE SCOPE
6 . 1 CO NC LU SIO N
This project is developed for the prediction of heart disease using the heart data of patients. The
results and the outputs of the project show that the most accurate algorithm among the deployed
algorithms are Random Forest and Logistic Regression with 90.16% and 89.1% of accuracy
scores respectively. All the accuracy scores, as mentioned are compared and efficient one is
cho- sen among them. The target variable output of the model is generated based on the
accuracy scores of all the algorithms.
As the Random Forest builds multiple decision trees, it ensures to compute the output from the
major trees that contribute to produce the output accurately. The trees and the outputs are en-
hanced through proper training. On the other hand, the likelihood of the event that is about to
happen is statistically predicted precisely using the Logistic Regression approach. By estimating
the probabilities of the variables, the outputs are predicted.
Hence, in regard to the subject, we conclude that the Machine Learning is the most emerging
and beneficial in almost all the sectors and fields of studies. It finds its use in almost all the
applica- tions of medical and biomedical fields. The models can be altered in the way they work
for dif- ferent kinds of datasets. By using these kinds of techniques and computer predicted
values, we can classify the disease fast with reduced cost. This project helped us enrich our
knowledge of several Machine Learning Algorithms and the way ML organizes the synaptic
signals for classi- fication and progression of the data. We have learned how to develop ML
algorithmic models and use its functionalities.
4
6 . 1 FU TU RE S COPE
This work can be extended and enhanced by deploying even more efficient algorithms and em-
ploying the mechanisms and techniques of Deep Learning. Dynamic webpages and Mobile Ap-
plications with enhanced features such as Mobile Phone communication can be created and de-
ployed for easier and effectively enhanced access. While operating in the real time, with the pa-
tient’s consent, we can collect the contact number of the patient and use it for applications such
as communicating about the health care advices and suggestions. If a patient’s heart is predicted
as unhealthy, the doctors and the hospitals contacts nearby can be sent to the patient o the
mobile number in case of any emergency. Easier communication access can be provided. This
extension is just an indefinite conception. It can be enhanced, promoted and executed as per the
convenience in any form.
4
7. . REFERENCES
[1]. Rairikar, A., Kulkarni, V., Sabale, V., Kale, H., & Lamgunde, A. (2017, June). Heart
disease prediction using data mining techniques. In 2017 International Conference on Intelligent
Compu- ting and Control (I2C2) (pp. 1-8). IEEE.
[2]. Gandhi, Monika, and Shailendra Narayan Singh. "Predictions in heart disease using tech-
niques of data mining." In 2015 International Conference on Futuristic Trends on
Computational Analysis and Knowledge Management (ABLAZE), pp. 520-525. IEEE, 2015.
[3]. Chen, A. H., Huang, S. Y., Hong, P. S., Cheng, C. H., & Lin, E. J. (2011, September).
HDPS: Heart disease prediction system. In 2011 Computing in Cardiology (pp. 557-560). IEEE.
[4]. Aldallal, A., & Al-Moosa, A. A. A. (2018, September). Using Data Mining Techniques to
Predict Diabetes and Heart Diseases. In 2018 4th International Conference on Frontiers of
Signal Processing (ICFSP) (pp. 150- 154). IEEE.
[5]. Sultana, Marjia, Afrin Haider, and Mohammad Shorif Uddin. "Analysis of data mining
tech- niques for heart disease prediction." In 2016 3rd International Conference on Electrical
Engineer- ing and Information Communication Technology (ICEEICT), pp. 1-5. IEEE, 2016.
[6].Al Essa, Ali Radhi, and Christian Bach. "Data Mining and Warehousing." American
Society for Engineering Education (ASEE Zone 1) Journal (2014).
[7].Shetty, Deeraj, Kishor Rit, Sohail Shaikh, and Nikita Patil. "Diabetes disease prediction
using data mining." In 2017 International Conference on Innovations in Information, Embedded
and Communication Systems (ICIIECS), pp. 1-5. IEEE, 2017.
[8]. Methaila, Aditya, Prince Kansal, Himanshu Arya, and Pankaj Kumar. "Early heart disease
prediction using data mining techniques." Computer Science & Information Technology Journal
(2014): 53-59.
[9].[Link]
[10].[Link]
4
APPENDIX: SOURCE CODE
#Importing libraries
import numpy as np
import pandas as pd
import [Link] as plt
%matplotlib inline
data = pd.read_csv('[Link]')
[Link]()
[Link]()
[Link]().sum()
[Link]()
corr
sns.set_style('whitegrid')
[Link](x = 'target', data = data)
dataset = [Link]()
[Link]()
X = [Link](['target'], axis = 1)
y = dataset['target']
[Link]
pred = [Link](X_test)
pred[:10]
5
from [Link] import confusion_matrix
confusion_matrix(y_test, pred)
## Hyperparameter Tuning
search_clfr.fit(X_train, y_train)
params = search_clfr.best_params_
score = search_clfr.best_score_
print(params)
print(score)
[Link](X_train, y_train)
confusion_matrix(y_test, [Link](X_test))
import pickle
[Link](classifier, open('[Link]', 'wb'))
The proposed heart disease prediction system supports decision-making by providing healthcare professionals with precise predictions of heart disease presence based on a patient's data. It processes data through machine learning models, thereby offering a less ambiguous diagnostic tool that complements traditional clinical methods. By reducing the need for extensive diagnostic tests, it allows healthcare professionals to make informed decisions more swiftly . This decision support tool enhances diagnostic accuracy and efficiency within hospitals and clinics, thereby addressing a critical gap in existing health services .
Data preprocessing is crucial in the development of heart disease prediction models as it enhances the quality and accuracy of the data used by models. Preprocessing involves cleaning the data to remove missing or irrelevant values, normalizing datasets for better model training, and transforming the data to fit the model requirements. Such steps ensure that the machine learning algorithms can analyze data more effectively, leading to more accurate predictions .
The project performs a comparative analysis of various machine learning algorithms, including Logistic Regression, Decision Trees, Naïve Bayes, K-NN, Support Vector Machine, Random Forests, and Neural Networks. These algorithms were evaluated based on their accuracy rates, where the model showcasing the highest accuracy was selected for prediction tasks. This comparison underlines the capability of machine learning to address complex disease prediction tasks like heart disease, highlighting key differences in model performance .
The main motivations for implementing a heart disease prediction system using machine learning include addressing the high mortality rate due to heart disease, which remains a leading cause of death globally. There is a need for accurate and efficient diagnosis in hospitals and clinics, which current systems and practices lack . The project aims to improve decision-making by healthcare professionals and reduce the number of tests required to diagnose heart disease, thus saving time and resources .
Predicting heart disease using machine learning models comes with challenges such as handling data complexity due to the presence of numerous contributing factors and the potential for ambiguity. Further, achieving high model accuracy while maintaining sensitivity and precision is difficult. The models must be trained effectively to ensure they generalize well to new data, and this requires a robust preprocessing and validation process . The misuse or misunderstanding of algorithm outputs by medical practitioners could lead to issues in clinical decision-making .
In the context of heart disease prediction models, sensitivity refers to the true positive rate—how effectively the model identifies actual cases of heart disease. Precision, on the other hand, measures the proportion of true positive cases out of all cases predicted as positive by the model. Both metrics are essential for evaluating a model's performance, ensuring that it not only identifies all possible cases but also minimizes false alarms, an aspect crucial for medical diagnostics .
Python is significant in the development of the heart disease prediction model due to its efficiency in handling large datasets and its comprehensive libraries and functions that facilitate machine learning tasks. Its versatility allows for easy integration and implementation of models like Decision Trees and Support Vector Machines in a cost-effective manner . Python's extensive library support, such as NumPy and SciPy, aids in the quick deployment and testing of machine learning models .
The proposed system improves the efficiency of diagnosing heart diseases by utilizing machine learning algorithms that compare accuracy, sensitivity, and precision across multiple models such as Logistic Regression, K-NN, and Neural Networks. By selecting the model with the highest accuracy, it provides clear, consistent outputs to support clinical decision-making. This system streamlines the testing process, potentially reducing the need for multiple diagnostic tests by predicting the presence of heart disease through data analysis .
The Software Development Life Cycle (SDLC) plays a pivotal role in the development of the heart disease prediction system by providing a structured framework that guides each phase—requirement analysis, design, implementation, testing, and maintenance. This organized approach ensures that all components of the prediction system are developed systematically, with a strong emphasis on testing and validation to meet user requirements and ensure system reliability . It also allows for iterative improvements based on feedback, aligning system functionalities with healthcare needs .
Ensuring that the heart disease prediction model is socially feasible is important because it affects user acceptance and integration within clinical settings. This is achieved through effective training and education of users, which helps mitigate resistance by familiarizing healthcare professionals with the system's operations and benefits. Addressing users' concerns and increasing their confidence in the system fosters constructive feedback and improves the model's applicability in real-world scenarios . Social feasibility ensures that the system is seen as supportive rather than a hindrance to existing healthcare practices .