Heart Disease Prediction Using Machine Learning
Heart Disease Prediction Using Machine Learning
INTRODUCTION
We know that Heart is the important part of our body. Life is itself dependent on efficient
working of heart. It is a world known fact that heart is the most essential organ in human
body if that organ gets affected then it also affects the other vital parts of the body. There
are many factors which increases risk of Heart disease. Some of them are:
• Family history of heart diseases.
• Smoking.
• Cholesterol.
• High blood pressure.
• Obesity.
• Lack of physical exercise.
As World Health Organization has estimated that 12 million deaths occur worldwide, every
year due to the Heart diseases. In 2008, 17.3 million people died due to Heart Disease.
Over 80% of deaths in world are because of Heart disease. WHO estimated by 2030,
almost 23.6 million people will die due to Heart disease Predication should to be done to
reduce risk of Heart disease. Diagnosis is usually based on signs, symptoms and physical
examination of a patient. As all the doctors are predicting heart disease by learning and
their experience. The diagnosis of disease is a difficult task in medical field. Predicting
Heart disease from various symptoms is a big issue which may lead to unpredictable
effects. Healthcare industry generates large amounts of complex data about patients,
disease diagnosis, electronic patient records etc. This large amount of data to be processed
and analyzed for knowledge extraction that enables for cost-savings and decision making.
Only human intelligence alone is not enough for proper diagnosis. As we are facing many
difficulties, to improve the accuracy of diagnosis and to reduce the diagnosis time, we have
developed an efficient and reliable Decision Support System for Heart Disease using data
mining and machine learning techniques.
Machine learning is a field of computer science that gives computer systems the ability to
"learn" with data, without being programmed. It is an application of Artificial Intelligence
that provides systems the ability to learn and improve from experience without being
1
programmed. The primary aim is to allow the computers learn without human
intervention or help and adjust action. Machine learning algorithms are often categorized
as supervised, unsupervised and Reinforcement.
Supervised Learning: In Supervised learning, you train the machine using data which is
well labelled. A supervised learning algorithm learns from labelled training data, helps
you to predict outcomes for unforeseen data.
An artificial neural network learning algorithm, or neural network, or just neural net, is a
computational learning system that uses a network of functions to understand and
translate a data input of one form into a desired output, usually in another form. The
concept of the artificial neural network was inspired by human biology and the way
neurons of the human brain function together to understand inputs from human senses.
Neural networks are just one of many tools and approaches used in machine learning
algorithms. The neural network itself may be used as a piece in many different machine
learning algorithms to process complex data inputs into a space those computers can
understand.
Neural networks are being applied to many real-life problems today, including speech and
image recognition, spam email filtering
Deep learning is computer software that mimics the network of neurons in a brain. It is a
subset of machine learning and is called deep learning because it makes use of deep
neural networks. The machine uses different layers to learn from the data. The depth of
the model is represented by the number of layers in the model. Deep learning is the new
state of the art in term of AI. In deep learning, the learning phase is done through a neural
network. A neural network is an architecture where the layers are stacked on top of each
other.
2
A Convolutional Neural Network is a Deep Learning algorithm which can take in an input
image, assign importance (learnable weights and biases) to various aspects/objects in the
image and be able to differentiate one from the other. The pre-processing required in a
convolutional neural network is much lower as compared to other classification
algorithms. While in primitive methods filters are hand-engineered, with enough training,
Convolution Nets have the ability to learn these filters/characteristics.
Recurrent Neural Network (RNN) is a type of Neural Network where the outputs from
previous step are fed as input to the current step. In traditional neural networks, all the
inputs and outputs are independent of each other, but in cases like when it is required to
predict the next word of a sentence, the previous words are required and hence there is a
need to remember the previous words. Thus RNN came into existence, which solved this
issue with the help of a Hidden Layer. The main and most important feature of RNN is
Hidden state, which remembers some information about a sequence.
Coming to classification in the project we use neural network namely Sequential model
which is a variation of Recurrent Neural Network for analysis.
3
Chapter 2
PROBLEM STATEMENT
The heart is very important part of human body. Which pumps blood into the entire body.
If circulation of blood in body is inefficient the organs like brain suffer and if heart stops
working death occurs within minutes. Life is completely dependent on working of the
heart. The term Heart disease refers to disease of heart & blood vessel system within it.
To predict the presence or absence of heart disease in human body using machine
learning techniques on data set provided by using sequential model and observing the
attributes which have more impact in predicting the heart disease.
4
Chapter 3
LITERATURE SURVEY
1. Dilip Roy Chowdhury [Link] [1] represented the use of artificial neural networks in
predicting neonatal disease diagnosis. We observed that the proposed technique
involves training a Multilayer Perceptron with a Back-propagation learning algorithm
to recognize a pattern for the diagnosing and prediction of neonatal diseases. The
Back- propagation algorithm was used to train the ANN architecture and the same has
been tested for the various categories of neonatal disease. Various cases of different
sign and symptoms parameter have been tested in this model. This shows that ANN
based prediction of neonatal disease and also improves the diagnosis accuracy with
stability.
2. Vanisree K [Link] [2] has been proposed a Decision Support System for diagnosis of
Congenital Heart Disease. We saw that the proposed system is designed and developed
by using MATLAB’s GUI feature with the implementation of Back propagation
Neural Network. The Back propagation Neural Network used in this study is a multi-
layered Feed Forward Neural Network, which is trained by a supervised Delta
Learning Rule. The dataset used in this study are the signs, symptoms and the results
of physical evaluation of a patient.
3. Milan Kumari [Link] [3] proposed research contains data mining classification
techniques like RIPPER classifier, Decision Tree, and Support Vector Machine (SVM)
which are analyzed on cardiovascular disease dataset. Performance of these techniques
is compared through sensitivity, specificity, accuracy, error rate, True Positive Rate
and False Positive Rate. 10-fold cross validation method was used to measure the
unbiased estimate of these prediction models. The analysis shows that out of these
classification models SVM predicts cardiovascular disease or heart disease with least
error rate and highest accuracy. SVM is a discriminative classifier formally defined by
a separating a hyperplane. In other words, given labeled training data, the algorithm
outputs an optimal hyperplane which categorizes new examples
4. A proficient methodology for the extraction of significant patterns from the heart
disease warehouses for heart attack prediction has been presented by Shantakumar
5
[Link] [Link] [4]. Initially, the data warehouse is pre-processed in order to make it
suitable for the mining process. Once the preprocessing gets over, the heart disease
warehouse is clustered with the help of the K-means clustering algorithm.
Consequently, the frequent patterns applicable to heart disease are mined with the help
of the maximal frequent itemset algorithm (MAFIA) from the data extracted. In
addition, the patterns vital to heart attack prediction are selected on basis of the
computed significant weightage. The neural network is trained with the selected
significant patterns for the effective prediction of heart attack.
5. Niti Guru [Link] [5] proposed a system that uses neural network for prediction of heart
disease, blood pressure and sugar. A set of 78 records with 13 attributes are used for
training and testing. He suggested supervised network for diagnosis of heart disease
and trained it using back propagation algorithm, the doctor using the system will find
that unknown data from training data and generate list of possible disease from which
patient can suffer.
6. Gudadhe M [Link] [6] has proposed a decision support system for heart disease
classification based on Support Vector Machine and Multilayer perceptron(MLP)
neural network architecture. For training of MLP, Back propagation algorithm is used
which is the famous learning algorithm for training of MLP. Support Vector Machine
classifies the heart disease data into two classes which shows presence of heart disease
or absence of heart disease with great accuracy. In this algorithm, we plot each data
item as a point in n-dimensional space (where n is number of features you have) with
the value of each feature being the value of a particular coordinate and perform
classification by finding hyperplane.
7. Senthil Kumar Mohan [7] the hybrid method which is combination of Linear
regression (LR), Multivariate Adaptive Regression Splines (MARS) and ANN. The
proposed method effectively reduced the set of critical attributes in the dataset and rest
of the attributes are input for ANN. The heart disease datasets are used to demonstrate
the efficiency of the development of the hybrid approach and prediction with neural
network. Neural networks are generally regarded as the best tool for prediction of
various diseases like heart disease and brain disease. We generate the results using
hybrid method, which produces good performance in the prediction of heart disease.
6
8. Pahulpreet Singh Kohli [8] The dataset contains 303 instances, from which 2 instances
for a number of vessels colored attribute and 4 instances for that attribute are missing,
which are filled by their mean value for the dataset respectively. The prediction
attributeconsists of 5 classes ranging from integer value 0 - 4 where 0 indicate absence
and the integer value from 1 - 4 indicate the presence of heart disease and encoded the
prediction attribute to class 0 and 1 to indicate absence or presence of heart disease
respectively. We have adopted backward selection method. In this, we begin with all
the attributes of the model, followed by their elimination based on p-value. This helps
to determine the significance of the results while performing null hypothesis test in
statistics. The attributes with the p-value greater than 0.05 were deleted and the model
was refitted with the remaining variables. This process was iterated multiple times until
every existing variable for the model was at a significant level.
9. Dinesh Kumar G [9] it is presumed that albeit most analysts are utilizing diverse
classifier methods, for example, Neural system, SVM, KNN and two fold
discretization with Gain Ratio Decision Tree in the conclusion of coronary illness,
applying Naïve Bayes and Decision tree with data pick up counts gives better
outcomes in the finding of coronary illness and better exactness when contrasted with
different classifiers. The expectation of cardiovascular illness by methods for a few
machine learning calculations is going on. Many research papers have executed
different machine learning calculations, for example, Naive Bayes, Random Forest,
Gradient boosting, Logical Regression and Support Vector Machine for anticipating
cardiovascular illness. It depicts how these machine learning calculations are utilized
to foresee the pneumonia ailment.
10. Humar Kahramanli and Novruz Allahverdi [10] obtained an accuracy of 87.4% by
using a hybrid neural network that combines a fuzzy neural network (FNN) with an
artificial neural network (ANN). In [Link] [Link] [11] an Intelligent Heart
Disease Prediction System (IHDPS) was presented. The IHDPS system uses data
mining techniques, such as for example DT, NB and Neural Network (NN). The tests
performed by the authors indicated that the NB model achieved the best performance
in terms of correct predictions (86.12%). The second-best model was NN with 86.12%
and the third DT with a score of 80.4% of correct predictions. In our current research,
7
our goal has been to highlight a comparison of almost all the aforementioned
techniques that have been employed in previous studies, and of their use in
combination, in order to be able to select the most appropriate prediction method.
8
Chapter 4
SOFTWARE REQUIREMENTS AND SPECIFICATIONS
Systems analysis is the detailed study of a system into its component pieces to study how
those component pieces internet and work. Software Requirement Specification is the
starting point of the software developing activity. As system grew more complex it
became evident that the goal of the entire system cannot be easily comprehended. Hence
the needs for the requirement phase arouse. The software project is initiated by the client
needs The SRS is the means of translating the ideas of the minds of clients (the input into
a formal document). The purpose of the software requirement specification is to reduce
the communication gap between the clients and the developers. Software Requirement
Specification is the medium through which the client and our needs are accurately
specified. It forms the basis of software development. A good SRS should satisfy all the
parties involved in the system.
4.1 Purpose
The purpose of this project is to develop a method for the predicting occurrence of heart
disease to a patient. Machine learning is used here to improve the accuracy.
4.2 Scope
This project predicts the positive or negative remark of the movie reviews. Given any
movie dataset that contains review of the movie in the form of sentence then by polarity it
gives whether it is positive or negative remark.
4.3 Objectives
9
4.4 Existing System
The existing system consists of balanced random forest as a classification technique for
predicting the presence of the heart disease. The proposed Gini index feature selection
addresses the issues of uneven distribution of prior class probability and global goodness
of a feature in two stages. First, it transforms the samples space into a feature specific
normalized samples space without compromising the intra class feature distribution. In
the second stage of the framework, it identifies the features that discriminates the classes
most by applying Gini coefficient of inequality. Also, Balanced Random Forest
Algorithm is used for classification which handles missing values using median for
numerical values or mode for categorical values. In this the accuracy rate is low.
Decision tree is used as this problem is supervised learning where the data is continuously
split according to a certain parameter. One advantage of tree-based methods is that they
have no assumptions about the structure of the data and are able to pick up non-linear
effects if given sufficient tree depth. Over fitting is a problem while binding a decision
tree model. Decision tree can create complex trees that do not generalise well, and
decision trees can be unstable because small variations in the data might result in a
completely different tree being generated.
SVM is a supervised machine learning algorithm which can be used for classification or
regression problems. It uses a technique called the kernel trick to transform your data and
then based on these transformations it finds an optimal boundary between the possible
outputs. SVM’s are very good when we have no idea on the data. It works well only for
the unstructured and semi structured data.
The Main idea of this project is to predict whether the patient suffers from heart disease
or not. And also predicting the risk of heart disease that is patient it is at high risk or low
risk. The user enters the appropriate input values from his/her health report. After this, the
historical dataset is uploaded and then transform the uploaded data into structured data
with the help of a data cleaning and data imputation process. Then the neural network
model is implemented on the input values and on the bases of this heart disease is
predicted. By using the neural network model, we increase the prediction heart disease
with more accuracy than existing system and in this we adding some new attributes in
10
data set like whether the patient smokes or not and if the patient have obesity or not
which helps in increasing the accuracy of the project.
4.6 REQUIREMENTS
Functional requirement defines a function of a software system or its component and how
the system must behave when presented with specific inputs or conditions. These may
include calculations, data manipulation and processing and other specific functionality that
define what system must accomplish. First, we pre-process the data by filling the missing
values, removing outliers, scaling converting categorical to numerical Next, Building the
Model for the stored pre-processed data. Next, validating and testing will be applied for
the model. Finally, Predicting occurrence of heart disease to a patient.
Performance:
Performance is measured in terms of the output provided by the application. Using neural
networks algorithm, the performance is increased than the models used in the existing
system.
Fault tolerance:
The data which is unnecessary is not used by the model because of pre-processing
techniques. The model is iterated for number of times and the best one is evaluated for
less faults.
11
Cost effectiveness:
The total cost required for developing the system is less when compared to the existing
systems as we used open source environment.
Flexibility:
The model is more flexible with less error rate and the developed system can work in any
operating system but it requires RAM of 25GB and internal storage disk space of 512GB.
Robustness:
Robustness is the ability of a computer system to cope with errors during execution and
cope with erroneous input. The system should be able to handle the errors during the
execution process.
Hardware requirements:
• Processor i7
• RAM 25GB
• Hard disk 512GB
• OS – Windows or Linux
• Graphics - Nvidia 4GB
Additional requirements:
• Cloud Servers - Colab notebooks execute code on Google's cloud servers.
• High Speed Net - If required take use of high speed LANS.
[Link] Python:
We have to install the Latest versions of Python 3.3+ for other libraries working. The
libraries required to run the associated code are:
• Pandas
• opencv2
• Jupiter
[Link] Keras:
Keras was created to be user friendly, modular, easy to extend, and to work with Python.
The Keras offers the advantages of broad adoption, support for a wide range of
12
production deployment options, integration with at least five backend engines and strong
support for multiple GPUs and distributed training. Keras provides the VGG16 pre-
trained model directly. Keras will download the model weights from the Internet, which
are about 500 Megabytes.
[Link] Anaconda:
We have to install the version Anaconda 5.0.1 for distribution of the Python and R
programming languages on Linux, Windows, and Mac OS X and enabling us to:
• Quickly download 1,500+ Python/R data science packages.
• Develop and train machine learning and deep learning models with scikit-learn
Tensor Flow, and Theano.
13
Chapter 5
SYSTEM DESIGN
System design is the process or art of defining the architecture, components, modules,
interfaces, and data for a system to satisfy specified requirements. One could see it as the
application of systems theory to product development. There is some overlap and synergy
with the disciplines of system analysis, system architecture and system engineering.
System architecture is a conceptual model that defines the structure, behaviour and more
views of a system. The prediction of heart disease system architecture is organized in the
following way that outputs are generated in the efficient output about the presence of
heart disease.
14
5.2 UML DESIGN
UML stands for Unified Modelling Language. Taking SRS document of analysis as input
to the design phase UML diagrams have been drawn. The UML is one part of the
software development method. The UML is process independent, although optimally it
should be used in a process that should be driven, architecture-centric, iterative, and
incremental. The UML is language for visualizing, specifying, constructing, documenting
the articles in a software intensive system.
A modelling language is a language whose vocabulary and rules focus on the conceptual
and physical representations of the system. A modelling language such as the UML is
thus a standard language for software blueprints.
The UML is a graphical language, which consists of all interesting systems. There are
also different structures that can transcend what can be represented in a programming
language. There are different diagrams in UML namely
• Use Case diagram
• Class diagram
• Collaboration diagrams
• State chart diagram
• Activity diagram
Use Case is used during requirement elicitation and analysis to represent the functionality
of the system. Use Case describes a function provided by the system that yields a visible
result for an actor. The identification of actors and use cases result in the definition of the
boundary of the system i.e., differentiating the tasks accomplished by the system and the
tasks accomplished by its environment. The actors are outside the boundary of the system,
whereas the Use cases are inside the boundary of the system. Use Cases describe the
behaviour of the system as seen from the actor’s point of view. It describes the function
provided by the system as a set of events that yield a visible result for the actor.
The use case diagram is for modelling the behaviour of the system and used to represent
the static view of the system (Fig.5.2.1). The actor here is user.
15
Fig: 5.2.1 Use Case diagram
16
5.2.3. State Diagram:
A State chart diagram describes a state machine. State machine can be defined as a
machine which defines different states of an object and these states are controlled by
external or internal events (Fig.5.2.3).
17
5.2.4. Collaboration Diagram:
This is an interaction diagram, which represents the structural organization of objects that
send and receive messages (Fig.5.2.4). It consists of set of roles, connectors that connect
the roles and the messages sent and receive by those roles.
18
Chapter 6
In the World, past 10 years Heart Disease is the serious cause of death. Heart disease
diagnosis is a complex task which requires much experience and knowledge. A doctor
way of predicting Heart disease is by examining number of medical tests such as ECG,
Stress Test, and Heart MRI etc. Nowadays, Health care industry contains huge amount of
heath care data, which contains hidden information. This hidden information is useful for
making effective decisions. To predict the patient risk level of having heart disease,
computer-based information along with data mining classification techniques like Naïve
Bayes, KNN, Decision tree, Neural Networks are used. In this report, a Heart Disease
Prediction system (HDPS) is being developed using data mining and machine learning
techniques.
The Sequential model API is a way of creating deep learning models where an instance of
the Sequential class is created and model layers are created and added to [Link] example,
the layers can be defined and passed to the Sequential as an array. The Sequential model
API is great for developing deep learning models in most situations, but it also has some
limitations. For example, it is not straightforward to define models that may have multiple
different input sources, produce multiple output destinations or models that re-use layers.
The layers in the model are connected pairwise. This is done by specifying where the
input comes from when defining each new layer. A bracket notation is used, such that
after the layer is created, the layer from which the input to the current layer comes from is
specified. We can create the input layer as above, then create a hidden layer as a Dense
that receives input only from the input layer.
The sequential API allows you to create models layer-by-layer for most problems. It is
limited in that it does not allow you to create models that share layers or have multiple
inputs or outputs. In Sequential model we use the add() function to add layers to our
model. We will add one or more layers and an output layer. Dense is the layer type. Dense
is a standard layer type that works for most cases. In a dense layer, all nodes in the
previous layer connect to the nodes in the current layer. Activation is the activation
function for the layer. An activation function allows models to take into account nonlinear
relationships. The functional API allows you to create models that have a lot more
flexibility as you can easily define models where layers connect to more than just the
19
previous and next layers. In fact, you can connect layers to (literally) any other layer. As a
result, creating complex networks such as Siamese networks and residual networks
become possible.
The Sentiment analysis for movie reviews system is developed in the following modules:
• Data Pre-processing
• Evaluate models
There are 303 records in the dataset and contains 14 continuous attributes. The goal is to
predict the presence of heart disease in the patient.
Here are the 14 attributes from the dataset along with their descriptions. These attributes
have been narrowed down to total of 14 in the dataset from the original set of 76 .
Age : The person’s age in years
sex : The person’s sex (1 = male, 0 = female)
cp : The chest pain experienced
trestbps: The person’s resting blood pressure
chol : The person’s cholesterol measurement in mg/dl
fbs : The person’s fasting blood sugar (> 120 mg/dl, 1 = true; 0 = false)
restecg : Resting electrocardiographic measurement ventricular hypertrophy
thalach : The person’s maximum heart rate achieved
exang : Exercise induced angina (1 = yes; 0 = no)
oldpeak: ST depression induced by exercise relative to rest
slope : The slope of the peak exercise ST segment
20
ca : The number of major vessels (0–3)
thal : A blood disorder called thalassemia
target: Heart disease (0 = no, 1 = yes)
Since Python 3.0, strings are stored as Unicode, i.e. each character in the string is
represented by a code point. So, each string is just a sequence of Unicode code points. For
efficient storage of these strings, the sequences of code points are converted into set of
bytes. The process is known as encoding. There are various encodings present which
treats a string differently. The popular encodings being, utf-8, ascii, etc.
Using string's encode() method, you can convert unicoded strings into any encodings
supported by python. By default, Python uses utf-8 encoding. By default, encode()
method doesn't require any parameters.
Latin-1, also known as ISO-8859-1, is a similar encoding. Unicode code points 0–255 are
identical to the Latin-1 values, so converting to this encoding simply requires converting
code points to byte values; if a code point larger than 255 is encountered, the string can't
be encoded into Latin-1. It converts the sentences that are in different languages into
Latin-1 format.
Here we take different datasets that contains movie reviews and they are merged into one
dataset with the required columns.
Data Pre-processing:
Data preprocessing is a data mining technique that involves transforming raw data into an
understandable format. Real-world data is often incomplete, inconsistent, and/or lacking
in certain behaviors or trends, and is likely to contain many errors. Data preprocessing is a
proven method of resolving such issues.
In Real world data are generally incomplete: lacking attribute values, lacking certain
attributes of interest, or containing only aggregate data. Noisy: containing errors or
outliers. Inconsistent: containing discrepancies in codes or names.
21
Steps in Data Preprocessing
Random forests is a supervised learning algorithm. It can be used both for classification
and regression. It is also the most flexible and easy to use algorithm. A forest is
comprised of trees. It is said that the more trees it has, the more robust a forest is. Random
forests creates decision trees on randomly selected data samples, gets prediction from
each tree and selects the best solution by means of voting. It also provides a pretty good
indicator of the feature importance.
Random forests has a variety of applications, such as recommendation engines, image
classification and feature selection. It can be used to classify loyal loan applicants,
identify fraudulent activity and predict diseases. It lies at the base of the Boruta algorithm,
which selects important features in a dataset.
22
Decision Tree Classifier:
Decision Tree algorithm belongs to the family of supervised learning algorithms. Unlike
other supervised learning algorithms, the decision tree algorithm can be used for
solving regression and classification problems too.
The goal of using a Decision Tree is to create a training model that can use to predict the
class or value of the target variable by learning simple decision rules inferred from prior
data(training data).
In Decision Trees, for predicting a class label for a record we start from the root of the
tree. We compare the values of the root attribute with the record’s attribute. On the basis
of comparison, we follow the branch corresponding to that value and jump to the next
node.
Decision trees classify the examples by sorting them down the tree from the root to some
leaf/terminal node, with the leaf/terminal node providing the classification of the
example.
Each node in the tree acts as a test case for some attribute, and each edge descending
from the node corresponds to the possible answers to the test case. This process is
recursive in nature and is repeated for every subtree rooted at the new node.
23
Support Vector Machines(SVM):
“Support Vector Machine” (SVM) is a supervised machine learning algorithm which can
be used for both classification or regression challenges. However, it is mostly used in
classification problems. In the SVM algorithm, we plot each data item as a point in n-
dimensional space (where n is number of features you have) with the value of each
feature being the value of a particular coordinate. Then, we perform classification by
finding the hyper-plane that differentiates the two classes very well.
In the SVM classifier, it is easy to have a linear hyper-plane between these two classes.
But, another burning question which arises is, should we need to add this feature
manually to have a hyper-plane. No, the SVM algorithm has a technique called
the kernel trick. The SVM kernel is a function that takes low dimensional input space and
transforms it to a higher dimensional space i.e. it converts not separable problem to
separable problem. It is mostly useful in non-linear separation problem. Simply put, it
does some extremely complex data transformations, then finds out the process to separate
the data based on the labels or outputs you’ve defined.
24
Sequential Model:
Keras is a user-friendly neural network library written in Python. We have sequential and
functional. The sequential API allows you to create models layer-by-layer for most
problems. It is limited in that it does not allow you to create models that share layers or
have multiple inputs or outputs. In Sequential model we use the add() function to add
layers to our model. We will add one or more layers and an output layer. Dense is the
layer type. Dense is a standard layer type that works for most cases. In a dense layer, all
nodes in the previous layer connect to the nodes in the current layer. Activation is the
activation function for the layer. An activation function allows models to take into
account nonlinear relationships. The functional API allows you to create models that have
a lot more flexibility as you can easily define models where layers connect to more than
just the previous and next layers. In fact, you can connect layers to (literally) any other
layer. As a result, creating complex networks such as Siamese networks and residual
networks become possible.
25
Obtain Confusion matrix:
26
Chapter 7
In the field of machine learning and specifically the problem of statistical classification, a
confusion matrix, also known as an error matrix.
It gives us insight not only into the errors being made by a classifier but more importantly
the types of errors that are being made.
These all standard of correctness are depend upon the Actual and Predict class, so we
draw a 2x2 confusion matrix to know more about them.
27
Class1 Class2
Predicted Predicted
Class 1 TP FN
Actual
Class 2 FP TN
Actual
True Positive (TP): These values are correctively predicted positive that means value of
both actual class and predicted class are YES.
True Negative (TN): These values are correctively predicted negative that means value of
both Actual class and predicted class are NO.
False Positive (FP): When value of actual class is NO and value of predicted class is YES.
False Negative (FN): When value of actual class is YES and value of predicted class is
NO. False Positive and False Negative classes occur when actual class contradicts with
predicted class.
28
F1-Score: It is the weighted average of Precision and Recall. Therefore this score takes
both false negatives and false positives into account.
7.2TESTING
Testing is the process of deliberately trying to cause failures in a system in order to detect
any defects that might be present. It is the time where we analyze whether all the user
requirements are met or not by providing different test cases. A test case in software
engineering normally consists of a unique identifier, requirement references from a design
specification, preconditions, events, a series of steps (also known as actions) to follow,
input, output and it validates one or more system requirements and generates a pass or
fail. In order to make sure that the system does not have errors, the different levels of
testing strategies that are applied at differing phases of software development are:
• Unit Testing.
• Integration Testing.
• Regression Testing.
• Alpha Testing.
• Beta Testing.
• System Testing.
• Performance Testing.
29
7.3 TYPES OF TESTING
Unit Testing:
Integration Testing:
a) Black box testing: In this we ignore the internal working mechanism and focuses
only on the output. It is used for validation purposes.
b) White box testing: This focuses on the way output is achieved. It’s mainly used
for verification purposes.
Regression testing:
It is done whenever a new module is added to the system. It makes sure that the system
works fine even after adding the new module.
Alpha testing:
It is a kind of acceptance testing. This type of testing is done before releasing the product
and it is carried out within the organization.
Beta testing:
Beta testing is done by end-users at one or more customer sites. The beta version is
released to few customers to test in real time environment.
30
System Testing:
System testing ensures that the system works fine at different operating conditions. In this
we just focus on input and output of the system without bothering about the internal
working.
Performance Testing:
The Performance test ensures that the output is produced within the time limits, and the
time taken by the system for compiling, giving response to the users and request being
send to the system for to retrieve the results.
31
7.4 RESULT ANALYSIS
The dataset is trained and tested with the model, also converting the dataset imbalanced
data to the balanced data. The following diagram shows the confusion matrix indicating
that sequence model is the efficient model when compared with remaining models.
32
7.5 TEST CASES
The developed sentiment analysis for movie review system accuracy, mean, f1-score and
the performance is evaluated against the test cases mentioned in the below table 6.1.
Some test cases like developing model, evaluating model are to be tested carefully
because the overall system accuracy and confusion matrix is based on them. The test
cases like pre-processing the data set define the efficiency of the model. Overall the test
cases in the below table are obtained in a positive way.
Test Objective: To test whether.
Table: 7.5 Test Case Output
33
result
34
Chapter 8
CONCLUSION
To predict the patient risk level of having heart disease, computer-based information along
with data mining classification techniques like Naïve Bayes, KNN, Decision tree, Random
forest ,SVM, Neural Networks are used. In this report, a Heart Disease Prediction system
(HDPS) is being developed using data mining and machine learning techniques. We used
neural networks to find if the patient has the heart disease. The binary result which is
obtained by using Sequential model using the neural networks techniques present along with
it.
35
Chapter 9
REFERENCES
[1] Dilip Roy Chowdhury, Mridula Chatterjee & R. K. Samanta, “An Artificial Neural
Network Model for Neonatal Disease Diagnosis”, International Journal of Artificial
Intelligence and Expert Systems (IJAE), Volume (2): Issue (3), 2011, pp. 96-106.
[2] Vanisree K, Jyothi Singaraju, “Decision Support System for Congenital Heart Disease
Diagnosis based on Signs and Symptoms using Neural Networks”, International Journal
of Computer Applications (0975 – 8887) Volume 19– No.6, April 2011, pp. NA.
[5] Niti Guru, Anil Dahiya, Navin Rajpal, “Decision Support System for Heart Disease
Diagnosis Using Neural Network”, Delhi Business Review, Vol. 8, No. 1, January-June
2007, pp. NA.
[6] Gudadhe M, Wankhade K, Dongre S. “Decision support system for heart disease based
on support vector machine and Artificial Neural Network”. Computer and
Communication Technology (ICCCT), 2010 International Conference on; 2010, pp. 741–
745
[7] Senthil Kumar, Gautam Srivastava “Effective Heart Disease Prediction Using Hybrid
Machine Learning Techniques”, date of publication June 19, 2019, pp. 1-1.
[8] Pahulpreet Singh Kohli, Shriya Arora. “Application of Machine Learning in Disease
36
Prediction”, 2018 4th International Conference on Computing Communication and
Automation (ICCCA), pp. 1-4.
[10] H. Kahramanli and N. Allahverdi, “Design of a hybrid system for the diabetes and heart
diseases,” Expert systems with applications, vol. 35, no. 1, pp. 82–89, 2008.
[11] S. Palaniappan and R. Awang, “Intelligent heart disease prediction system using data
mining techniques,” in IEEE/ACS International Conference on the Computer Systems
and its all the Applications that are present in it. IEEE, 2008, pp. 108–115.
37
APPENDIX I: SAMPLE CODE
import numpy as np
import pandas as pd
import os
print([Link]("/content/drive/My Drive/Dataset"))
df=pd.read_csv('/content/drive/My Drive/Dataset/[Link]')
[Link](5)
[Link]([Link],[Link]).plot(kind="bar",figsize=(10,5),color=['tomato','indigo' ])
[Link]('Chest Pain Type')
[Link](rotation = 0)
[Link]('Frequency of Disease or Not')
38
[Link]()
[Link](figsize=(16,8))
[Link](2,2,1)
[Link](x=[Link][[Link]==1],y=[Link][[Link]==1],c='blue')
[Link](x=[Link][[Link]==0],y=[Link][[Link]==0],c='black')
[Link]('Age')
[Link]('Max Heart Rate')
[Link](['Disease','No Disease'])
[Link](2,2,2)
[Link](x=[Link][[Link]==1],y=[Link][[Link]==1],c='red')
[Link](x=[Link][[Link]==0],y=[Link][[Link]==0],c='green')
[Link]('Age')
[Link]('Cholesterol')
[Link](['Disease','No Disease'])
[Link](2,2,3)
[Link](x=[Link][[Link]==1],y=[Link][[Link]==1],c='cyan')
[Link](x=[Link][[Link]==0],y=[Link][[Link]==0],c='fuchsia')
[Link]('Age')
[Link]('Resting Blood Pressure')
[Link](['Disease','No Disease'])
[Link](2,2,4)
[Link](x=[Link][[Link]==1],y=[Link][[Link]==1],c='grey')
[Link](x=[Link][[Link]==0],y=[Link][[Link]==0],c='navy')
[Link]('Age')
[Link]('ST depression')
[Link](['Disease','No Disease'])
[Link]()
39
chest_pain=pd.get_dummies(df['cp'],prefix='cp',drop_first=True)
df=[Link]([df,chest_pain],axis=1)
[Link](['cp'],axis=1,inplace=True)
sp=pd.get_dummies(df['slope'],prefix='slope')
th=pd.get_dummies(df['thal'],prefix='thal')
rest_ecg=pd.get_dummies(df['restecg'],prefix='restecg')
frames=[df,sp,th,rest_ecg]
df=[Link](frames,axis=1)
[Link](['slope','thal','restecg'],axis=1,inplace=True)
classifier = Sequential()
40
# Adding the second hidden layer
[Link](Dense(output_dim = 11, init = 'uniform', activation = 'relu'))
model_accuracy = [Link](data=[rdf_ac,ac],
index=['RandomForest','Neural Network'])
fig= [Link](figsize=(8,8))
model_accuracy.sort_values().[Link]()
[Link]('Model Accracy')
43