Deep Learning for Malaria Detection
Deep Learning for Malaria Detection
CHAPTER-1
INTRODUCTION
1.1Overview
In 2015, the World Health Organization (WHO) reported that around 438 thousand
people died due to malaria parasites, and 620 thousand people died in 2017, and
approximately 300-500 million people were affected by this disease yearly. South-East
Asia, the Eastern Mediterranean, the Western Pacific, and the Americas have all been
recognized as high-risk regions by the WHO. Malaria is a dangerous and deadly disease
that is triggered from the bite of female anopheles mosquitoes that host plasmodium
parasites. There are around 400 types of anopheles; among them, 30 types mainly act as
parasite carrieres. To be a parasite carrier, the female anopheles mosquito has to bite a
malaria-infected person. The life cycle of the anopheles mosquito consists of four stages:
egg, larva, pupa, and adult. To bring the eggs into the adult stage, they have to feed on
human blood. In this feeding process, malaria spreads all around. There are five
Plasmodium parasite species, including Plasmodium falciparum, P. vivax, P. ovale, P.
malariae, and P. knowlesi. Among them, P. falciparum and P. vivax are the most
dangerous. It takes 10 to 15 days or even more to remain in hibernation after a carrier
mosquito bite. After the hibernation period, it seizures the red blood cells (RBC) and
reduces the number of RBC, revealing different indications of malaria, such as fever,
chills, nausea, and vomiting. Malaria should be diagnosed as soon as possible because it
can rapidly turn into a severe stage and is life-threatening. It cannot be passed from one
person to another, however malaria can be transmitted from mother to fetus, contracted
through blood transfusions or sharing injections. This disease can be spread in hot, humid
climates near natural water sources where Anopheles mosquitoes transmit deadly
diseases.
The advent of artificial intelligence (AI), specifically deep learning, offers promising
solutions to these challenges. Deep learning, a subset of machine learning, employs
artificial neural networks to automatically learn and identify complex patterns in data.
This technology has gained significant attention for its ability to analyze medical images
and improve diagnostic precision. In the context of malaria, deep learning models can be
trained to detect infected red blood cells from microscopic images, enabling automated,
rapid, and reliable diagnosis.
Deep learning-based systems offer several advantages. They reduce the reliance on
skilled technicians, minimize human error, and enable rapid processing of large volumes
of data. Additionally, these systems can be deployed in resource-constrained
environments using mobile or cloud-based platforms, enhancing accessibility and
scalability.
Despite their potential, challenges remain in implementing deep learning for malaria
detection. These include the need for large, diverse, and annotated datasets, addressing
class imbalance in datasets, and ensuring model robustness across varied imaging
conditions. Furthermore, the integration of AI-driven tools into existing healthcare
systems requires regulatory approval, user training, and ongoing maintenance.
In conclusion, deep learning has emerged as a transformative tool in the detection of
malaria. By automating and enhancing diagnostic accuracy, it holds the promise of
improving disease management and outcomes, particularly in endemic regions. Continued
research and collaboration between AI specialists, healthcare professionals, and
policymakers will be crucial to realizing the full potential of this technology in combating
malaria.
Role of Deep Learning
Deep learning, a subset of artificial intelligence (AI), offers a powerful approach to
address these challenges. By leveraging neural networks capable of automatically
identifying patterns in complex datasets, deep learning has demonstrated remarkable
success in medical imaging tasks. Its application in malaria detection involves analyzing
microscopic images of blood smears to identify infected cells, offering a potential
solution to overcome the limitations of traditional methods.
Proposed Solution Framework
Malaria remains a major public health concern, affecting millions of people annually,
particularly in tropical and subtropical regions. It is caused by Plasmodium parasites
transmitted through the bites of infected female Anopheles mosquitoes. Timely and
accurate detection of malaria is essential for effective treatment and management of the
disease.
1.3Objectives
• To detect the infected area in blood images we are using the water-shed image
segmentation techniques.
1.4Scope:
• This can be a valuable tool in the fight against malaria, as it can help to rapidly
and accurately diagnose cases, allowing for timely treatment and help to prevent
the spread of the disease.
• In this project, images are good enough. The effectiveness of this framework with
low quality is excluded from the scope of this project.
CHAPTER-2
2 LITERATURE SURVEY
Literature survey describes about the existing work on the given project. It deals with
the problem associated with the existing system and also gives user a clear knowledge on
Before building our application, the following system is taken into consideration:
1 Title: Developing an The SEIRS model The model It will work only for
agent-based model for represented the text data it is not
simulating the spread of malaria suitable for image
dynamic spread of by simulating the data.
plasmodium vivax interaction among
malaria Anopheles
mosquitoes,
Author: N. M.
humans, and the
Gharakhanlou, M. S.
environment. In the
Mesgari, and N.
model, the
Hooshangi
probability of
Year:2019 malaria
6 Title: Malaria Parasite Convolution Neural The proposed CNN • Its accuracy is
Classification using Network (CNN) setup-1 with kernel less.
Deep Convolutional models size 3 x 3 and pool
• It consumes
Neural Network size of 2 x 2
more training
achieved an
Author: Abhik time.
accuracy of 96%.
Paul; Rubul Kumar
Bania
Year: 2021
7 Title: Deep Learning The customized CNN CNN model in • Accuracy is less
for Smartphone-Based discriminating than 95%.
Malaria Parasite between positive
• No
Detection in Thick (parasitic) and
recommendatio
Blood Smears negative image
n of medicine.
patches in terms of
Author: Feng
the following • It will not able
Yang; Mahdieh
performance to detect the
Poostchi; Hang
indicators: types.
Yu; Zhou
accuracy (93.46%
Zhou; Kamolrat
± 0.32%).
Silamut
Year:2020
2.2Related Works
Several research efforts have been directed towards leveraging deep learning, particularly
Convolutional Neural Networks (CNNs), for the detection of malaria in microscopic
blood smear images. The following are significant contributions in this field:
EfficientNet-Based Detection
Research published in Scientific Reports utilized the EfficientNet deep learning
architecture for malaria detection. This approach demonstrated high accuracy in
identifying malaria-infected red blood cells while optimizing computational efficiency,
making it suitable for deployment in resource-constrained settings.
CNN for Classification of Blood Smears
A study available on arXiv explored a CNN-based model to classify microscopic images
into parasitized and uninfected categories. The research highlighted the adaptability of the
model to various devices, emphasizing its potential for use in remote and low-resource
environments.
Analysis of Deep Learning Algorithms
CHAPTER-3
3.1.1 Drawbacks
• Masud et al. proposed leveraging deep CNN for real-time detection of malaria
from RBC images. They developed a custom CNN using cyclical stochastic
gradient descent (SGD) as an optimizer and achieved an accuracy of 97.30%.
• Rosado et al. have used SVM to look at how to identify malaria parasites and
white blood cells using smartphones.
3.2.1Advantages
The SRS talks about the item however not the venture that created it, consequently
the SRS serves as a premise for later improvement of the completed item. The SRS may
need to be changed, however it does give an establishment to proceed with creation
assessment. In straightforward words, programming necessity determination is the
beginning stage of the product improvement action.
The SRS means deciphering the thoughts in the brains of the customers – the
information, into a formal archive – the yield of the prerequisite stage. Subsequently the
yield of the stage is a situated of formally determined necessities, which ideally are
finished and steady, while the data has none of these properties.
This section describes the functional requirements of the system for those
requirements which are expressed in the natural language style.
3. System will read and preprocess the extract features and train the model using
CNN.
5. System will read and preprocess the extract features and predict the type of
malaria using CNN
6. System will recommend the medicine as per the predicted malaria disease.
These are requirements that are not functional in nature, that is, these are constraints
within which the system must work.
The program must be self-contained so that it can easily be moved from one
Computer to another. It is assumed that network connection will be available on
the computer on which the program resides.
The system shall achieve 100 per cent availability at all [Link] system shall be
scalable to support additional clients and volunteers.
Maintainability.
• RAM :8 GB RAM
INTEL CORE i5
Intel Core is a brand name that Intel uses for various mid-range to high-end
consumer and business microprocessors. As of 2015 the current line up of Core
processors included the Intel Core i7, Intel Core i5, and Intel Core i3. 5th generation
Intel® Core™ i5 processors empower new innovations like Intel® Real Sense™
technology—bringing you features such as gesture control, 3D capture and edit, and
innovative photo and video capabilities to your devices. Enjoy stunning visuals, built-in
security, and an automatic burst of speed when you need it with Intel® Turbo Boost
Technology 2.0.
[Link] RAM
When you load up an application on to your computer it loads into your available
RAM memory. It is very quick type of memory. The more programs you load up, the
more RAM is taken up. At the point where you have loaded up enough apps to take up all
your free available physical RAM, your OS will create a swap-file on your hard drive.
This file is used as a reserve for all additional apps you run.
The trouble with that is that hard drives are a lot slower to read and write from
than RAM memory is. Therefore, your computer will perform much slower at that point.
Although new generation of SSD hard drives are much faster than your traditional
spinning drive, it is still best to have enough RAM available. If you are using Windows
and want to want to know how much RAM you are using up, you can right click on task
bar, then select start "Task Manager" and on the "performance" tab you will see a green
bar indicating "Memory".
A hard disk drive (HDD), hard disk, hard drive or fixed disk is a data storage
device used for storing and retrieving digital information using one or more rigid ("hard")
rapidly rotating disks (platters) coated with magnetic material. The platters are paired
with magnetic heads arranged on a moving actuator arm, which read and write data to the
platter surfaces. Data is accessed in a random-access manner, meaning that individual
blocks of data can be stored or retrieved in any order rather than sequentially.
• Technology : Python
• IDE : PythonIDLE
• Tools : Anaconda
[Link] Python:
Python interpreters are available for many operating systems. CPython, the
reference implementation of Python, is open source software[30] and has a community-
based development model, as do nearly all of Python's other implementations. Python and
CPython are managed by the non-profit Python Software Foundation.
History
Python was conceived in the late 1980s by Guido van Rossum at Centrum
Wiskunde&Informatica (CWI) in the Netherlands as a successor to the ABC language
(itself inspired by SETL, capable of exception handling and interfacing with the Amoeba
operating system. Its implementation began in December 1989. Van Rossum's long
influence on Python is reflected in the title given to him by the Python community:
Benevolent Dictator For Life (BDFL) – a post from which he gave himself permanent
vacation on July 12, 2018.
Python 2.0 was released on 16 October 2000 with many major new features,
including a cycle-detecting garbage collector and support for Unicode.
Python 3.0 was released on 3 December 2008. It was a major revision of the
language that is not completely backward-compatible. Many of its major features were
backported to Python 2.6.x[37] and 2.7.x version series. Releases of Python 3 include the
2to3 utility, which automates (at least partially) the translation of Python 2 code to Python
3. Python 2.7's end-of-life date was initially set at 2015 then postponed to 2020 out of
concern that a large body of existing code could not easily be forward-ported to Python 3.
In January 2017, Google announced work on a Python 2.7 to Go transcompiler to
improve performance under concurrent workloads.
Python uses dynamic typing, and a combination of reference counting and a cycle-
detecting garbage collector for memory management. It also features dynamic name
resolution (late binding), which binds method and variable names during program
execution.
Python's design offers some support for functional programming in the Lisp
tradition. It has filter(), map(), and reduce() functions; list comprehensions, dictionaries,
and sets; and generator expressions. The standard library has two modules (itertools and
functools) that implement functional tools borrowed from Haskell and Standard ML.
Readability counts
Rather than having all of its functionality built into its core, Python was designed
to be highly extensible. This compact modularity has made it particularly popular as a
means of adding programmable interfaces to existing applications. Van Rossum's vision
of a small core language with a large standard library and easily extensible interpreter
stemmed from his frustrations with ABC, which espoused the opposite approach.
Indentation
The assignment statement (token '=', the equals sign). This operates differently
than in traditional imperative programming languages, and this fundamental mechanism
(including the nature of Python's version of variables) illuminates many other features of
the language. Assignment in C, e.g., x = 2, translates to "typed variable name x receives a
copy of numeric value 2". The (right-hand) value is copied into an allocated storage
location for which the (left-hand) variable name is the symbolic address. The memory
allocated to the variable is large enough (potentially quite large) for the declared type. In
the simplest case of Python assignment, using the same example, x = 2, translates to
"(generic) name x receives a reference to a separate, dynamically allocated object of
numeric (int) type of value 2." This is termed binding the name to the object. Since the
name's storage location doesn't contain the indicated value, it is improper to call it a
variable. Names may be subsequently rebound at any time to objects of greatly varying
types, including strings, procedures, complex objects with data and methods, etc.
Successive assignments of a common value to multiple names, e.g., x = 2; y = 2; z = 2
result in allocating storage to (at most) three names and one numeric object, to which all
three names are bound. Since a name is a generic reference holder it is unreasonable to
associate a fixed data type with it. However at a given time a name will be bound to some
object, which will have a type; thus there is dynamic typing.
The if statement, which conditionally executes a block of code, along with else
and elif (a contraction of else-if).
The for statement, which iterates over an iterable object, capturing each element to
a local variable for use by the attached block. The while statement, which executes a
block of code as long as its condition is true. The try statement, which allows exceptions
raised in its attached code block to be caught and handled by except clauses; it also
ensures that clean-up code in a finally block will always be run regardless of how the
block exits.
The with statement, from Python 2.5 released on September 2006,[59] which
encloses a code block within a context manager (for example, acquiring a lock before the
block of code is run and releasing the lock afterwards, or opening a file and then closing
it), allowing Resource Acquisition Is Initialization (RAII)-like behavior and replaces a
common try/finally idiom.
Python does not support tail call optimization or first-class continuations, and,
according to Guido van Rossum, it never will. However, better support for coroutine-like
functionality is provided in 2.5, by extending Python's generators. Before 2.5, generators
were lazy iterators; information was passed unidirectionally out of the generator. From
Python 2.5, it is possible to pass information back into a generator function, and from
Python 3.3, the information can be passed through multiple stack levels.
Expressions
Some Python expressions are similar to languages such as C and Java, while some
are not:
Addition, subtraction, and multiplication are the same, but the behavior of division
differs. There are two types of divisions in Python. They are floor division and integer
division. Python also added the ** operator for exponentiation.
From Python 3.5, the new @ infix operator was introduced. It is intended to be used by
libraries such as NumPy for matrix multiplication.
Python uses the words and, or, not for its boolean operators rather than the
symbolic &&, ||, ! used in Java and [Link] has a type of expression termed a list
comprehension. Python 2.4 extended list comprehensions into a more general expression
termed a generator [Link] functions are implemented using lambda
expressions; however, these are limited in that the body can only be one expression.
Python makes a distinction between lists and tuples. Lists are written as [1, 2, 3],
are mutable, and cannot be used as the keys of dictionaries (dictionary keys must be
immutable in Python). Tuples are written as (1, 2, 3), are immutable and thus can be used
as the keys of dictionaries, provided all elements of the tuple are immutable. The +
operator can be used to concatenate two tuples, which does not directly modify their
contents, but rather produces a new tuple containing the elements of both provided tuples.
Thus, given the variable t initially equal to (1, 2, 3), executing t = t + (4, 5) first evaluates
t + (4, 5), which yields (1, 2, 3, 4, 5), which is then assigned back to t, thereby effectively
"modifying the contents" of t, while conforming to the immutable nature of tuple objects.
Parentheses are optional for tuples in unambiguous contexts.
Python has a "string format" operator %. This functions analogous to printf format
strings in C, e.g. "spam=%s eggs=%d" % ("blah", 2) evaluates to "spam=blah eggs=2". In
Python 3 and 2.6+, this was supplemented by the format() method of the str class, e.g.
Strings delimited by single or double quote marks. Unlike in Unix shells, Perl and
Perl-influenced languages, single quote marks and double quote marks function
identically. Both kinds of string use the backslash (\) as an escape character. String
interpolation became available in Python 3.6 as "formatted string literals".
Triple-quoted strings, which begin and end with a series of three single or double
quote marks. They may span multiple lines and function like here documents in shells,
Perl and Ruby.
Raw string varieties, denoted by prefixing the string literal with an r. Escape
sequences are not interpreted; hence raw strings are useful where literal backslashes are
common, such as regular expressions and Windows-style paths. Compare "@-quoting" in
C#.
Python has array index and array slicing expressions on lists, denoted as a[key],
a[start:stop] or a[start:stop:step]. Indexes are zero-based, and negative indexes are relative
to the end. Slices take elements from the start index up to, but not including, the stop
index. The third slice parameter, called step or stride, allows elements to be skipped and
reversed. Slice indexes may be omitted, for example a[:] returns a copy of the entire list.
Each element of a slice is a shallow copy.
The eval() vs. exec() built-in functions (in Python 2, exec is a statement); the
former is for expressions, the latter is for statements.
Methods
Methods on objects are functions attached to the object's class; the syntax
[Link](argument) is, for normal methods and functions, syntactic sugar for
[Link](instance, argument). Python methods have an explicit self parameter to
access instance data, in contrast to the implicit self (or this) in some other object-oriented
programming languages (e.g., C++, Java, Objective-C, or Ruby).
Typing
Python uses duck typing and has typed objects but untyped variable names. Type
constraints are not checked at compile time; rather, operations on an object may fail,
signifying that the given object is not of a suitable type. Despite being dynamically typed,
Python is strongly typed, forbidding operations that are not well-defined (for example,
adding a number to a string) rather than silently attempting to make sense of them.
Python allows programmers to define their own types using classes, which are
most often used for object-oriented programming. New instances of classes are
constructed by calling the class (for example, SpamClass() or EggsClass()), and the
classes are instances of the metaclass type (itself an instance of itself), allowing
metaprogramming and reflection.
Before version 3.0, Python had two kinds of classes: old-style and new-style. The
syntax of both styles is the same, the difference being whether the class object is inherited
from, directly or indirectly (all new-style classes inherit from object and are instances of
type). In versions of Python 2 from Python 2.2 onwards, both kinds of classes can be
used. Old-style classes were eliminated in Python 3.0.
The long term plan is to support gradual typing and from Python 3.5, the syntax of
the language allows specifying static types but they are not checked in the default
Mathematics
Python has the usual C language arithmetic operators (+, -, *, /, %). It also has **
for exponentiation, e.g. 5**3 == 125 and 9**0.5 == 3.0, and a new matrix multiply @
operator is included in version 3.5.[80] Additionally, it has a unary operator (~), which
essentially inverts all the bits of its one argument. For integers, this means ~x=-x-1. Other
operators include bitwise shift operators x << y, which shifts x to the left y places, the
same as x*(2**y) , and x >> y, which shifts x to the right y places, the same as x//(2**y) .
Python 2.1 and earlier use the C division behavior. The / operator is integer
division if both operands are integers, and floating-point division otherwise. Integer
division rounds towards 0, e.g. 7/3 == 2 and -7/3 == -2.
Python 2.2 changes integer division to round towards negative infinity, e.g. 7/3 ==
2 and -7/3 == -3. The floor division // operator is introduced. So 7//3 == 2, -7//3 == -3,
7.5//3 == 2.0 and -7.5//3 == -3.0. Adding from __future__ import division causes a
module to use Python 3.0 rules for division (see next).
Python provides a round function for rounding a float to the nearest integer. For
tie-breaking, versions before 3 use round-away-from-zero: round(0.5) is 1.0, round(-0.5)
is −1.0. Python 3 uses round to even: round(1.5) is 2, round(2.5) is 2. Python allows
boolean expressions with multiple equality relations in a manner that is consistent with
Python has extensive built-in support for arbitrary precision arithmetic. Integers
are transparently switched from the machine-supported maximum fixed-precision
(usually 32 or 64 bits), belonging to the python type int, to arbitrary precision, belonging
to the Python type long, where needed. The latter have an "L" suffix in their textual
representation. (In Python 3, the distinction between the int and long types was
eliminated; this behavior is now entirely contained by the int class.) The Decimal
type/class in module decimal (since version 2.4) provides decimal floating point numbers
to arbitrary precision and several rounding modes. The Fraction type in module fractions
(since version 2.6) provides arbitrary precision for rational numbers.
Due to Python's extensive mathematics library, and the third-party library NumPy
that further extends the native capabilities, it is frequently used as a scientific scripting
language to aid in problems such as numerical data processing and manipulation.
Libraries
Python's large standard library, commonly cited as one of its greatest strengths,
provides tools suited to many tasks. For Internet-facing applications, many standard
formats and protocols such as MIME and HTTP are supported. It includes modules for
creating graphical user interfaces, connecting to relational databases, generating
pseudorandom numbers, arithmetic with arbitrary precision decimals, manipulating
regular expressions, and unit testing.
Some parts of the standard library are covered by specifications (for example, the
Web Server Gateway Interface (WSGI) implementation wsgiref follows PEP 333), but
most modules are not. They are specified by their code, internal documentation, and test
suites (if supplied). However, because most of the standard library is cross-platform
Python code, only a few modules need altering or rewriting for variant implementations.
As of March 2018, the Python Package Index (PyPI), the official repository for
third-party Python software, contains over 130,000 packages with a wide range of
functionality, including:
• Web frameworks
• Multimedia
• Databases
• Networking
• Test frameworks
• Automation
• Web scraping
• Image processing
Development environments
Implementations
CHAPTER-4
SYSTEM DESIGN
The system “design” is defined as the process of applying various requirements and
permits it physical realization. Various design features are followed to develop the system
the design specification describes the features of the system, the opponent or elements of
the system and their appearance to the end-users.
This image depicts a process flow diagram for identifying malarial parasites using image
processing and a Convolutional Neural Network (CNN). Let's break down each step :
Dataset: This is the starting point. It refers to the collection of images used for training
and testing the model. These images likely consist of microscopic images of blood
samples, some containing malarial parasites and others not. A robust and diverse dataset
is crucial for the model's accuracy and generalization ability.
Data Pre-processing: This stage prepares the images for further processing. It often
involves several steps:
0.1
Blood Cell Efficient Result
Dataset Detection of
malaria
using
Level 0 Describes the overall process of this project. we are passing Blood cell dataset as
a input the system will efficiently detects malaria using the customized CNN model.
1.1 1.2
encod
Clean Features
features
Level 1 Describes the first stage process of this project. we are passing website dataset as
a input the system will perform the preprocess and extract the important features.
Level-2:
Features 2.1
CNN Model
2.1
Trained Data
Read data 2.3
Classificati
Re-
on
No Train
Yes
Result
Level 2 Describes the final stage process of this project. we are passing extracted features
from level 1and trained data as a input the system will detect malaria using CNN model
from the given blood cell images.
This diagram is a use case diagram illustrating the interaction between a user and
a system designed for malaria prediction and medicine recommendation. The diagram
shows two actors: the "User" and the "System." The ovals represent use cases, which are
specific functionalities the system provides. The arrows indicate the relationships between
the actors and the use cases, showing which actor initiates or participates in which use
case.
The process begins with the "User" initiating the "Load Image" use case. This implies the
user uploads or selects an image, presumably a microscopic image of a blood sample, to
be analyzed by the system. The "System" then performs the "Read Image" use case,
which involves reading and interpreting the image data. Following this, the "System"
carries out "Preprocess," where the image is prepared for analysis through techniques like
noise reduction, color correction, or resizing. Next, the "System" executes "Feature
Extraction," where relevant features are extracted from the preprocessed image. These
features are then fed into the "Apply ELM Model" use case. ELM stands for Extreme
Learning Machine, a type of machine learning algorithm. The ELM model, trained
beforehand, analyzes the extracted features to "Predict Malaria," outputting a prediction
of whether malaria is present in the sample. Finally, based on the prediction, the "System"
performs the "Recommended Medicine" use case, suggesting appropriate medication if
malaria is detected. The arrow pointing from "Recommended Medicine" back to the
"User" indicates that the system provides this recommendation to the user.
This class diagram models a system for malaria prediction using image processing and an
Extreme Learning Machine (ELM) model. It outlines the interactions between key
components of the system. A User provides one or more input images to the System. The
System receives these images and utilizes methods like readInputdata() and readdataset()
to handle them. The System then passes the image to the preprocessor, which performs
preprocessing operations using the preprocessing() method, generating one or more
preprocessed images. These preprocessed images are then passed to the feature Extract
class, which extracts relevant features using the ExtractedFeature() method. The extracted
features, along with trained data, are then fed into the ELM Model. The ELM Model has
methods for building a Convolutional Neural Network (CNN) using Build CNN and
training the model using Train(). The trained data from the ELM Model, along with the
extracted features, is used by the Prediction class to make a prediction using the predict()
method. Finally, both the ELM Model's trained data and the predictions are used by the
Performance analysis class to evaluate the model's performance using metrics like
precision and a performance matrix, calculated by the Test() method. The diagram
This activity diagram illustrates the process of malaria detection using blood image
analysis and an Extreme Learning Machine (ELM) model. The process begins with the
simultaneous handling of "Training Blood Image Data" and "Test Blood Image Data,"
indicating these datasets are processed concurrently. Once both datasets are ready, the
flow merges, and the "Preprocess" activity takes place, preparing the images for further
analysis through techniques like noise reduction and color correction. Following
preprocessing, "Extract Image Features" extracts relevant characteristics from the images,
converting them into numerical representations. These extracted features from the training
data are then used in the "Train using ELM Model" activity, where the ELM model is
trained to recognize patterns associated with malaria. Subsequently, the "Performance
This sequence diagram meticulously details the interaction between a User and a System
designed for malaria prediction, offering a granular view of the process. The interaction
commences when the User initiates the process by sending a "Load Dataset" message to
the System. This action signifies the user uploading or selecting a dataset of blood smear
images for analysis. Upon receiving this message, the System's lifeline becomes active,
initiating a series of operations.
The first operation is "Read," during which the System reads and loads the image data
from the provided dataset into its memory. This step involves accessing the image files,
decoding them, and preparing them for subsequent processing. Following the "Read"
operation, the System proceeds with "Pre-process." This stage is crucial for enhancing the
quality of the images and preparing them for feature extraction. Preprocessing techniques
CHAPTER- 5
IMPLEMENTATION
[Link] Collection
2. Data Preprocessing
The malaria dataset has been collected from the Kaggle. It contains 27,558
images of RBC. The dataset contain equal amount of malaria infected and uninfected
samples hence it is a balance dataset.
5.2.3 PRE-PROCESSING
Image pre-processing is a crucial step for this type of study because model outcome
highly depends on pre-processing techniques applied. It makes the learning process
smooth.
Procedure Preprocess()
def segment(path,status):
im0 = [Link](path)
gray1 = [Link](im0,cv.COLOR_BGR2GRAY)
#[Link]('image1',gray1)
blur = [Link](gray1,(5,5),0)
ret, thresh =
[Link](blur,0,255,cv.THRESH_BINARY_INV+cv.THRESH_OTSU)
kernel = [Link]((3,3),np.uint8)
closing = [Link](thresh,cv.MORPH_CLOSE,kernel,iterations = 5)
#[Link]('image',closing)
opening = [Link](closing,cv.MORPH_OPEN,kernel,iterations = 3)
#[Link]('image2',opening)
sure_bg = [Link](opening,kernel,iterations = 2)
#[Link]('image5',sure_bg)
dist_transform = [Link](opening,cv.DIST_L2,5)
#[Link]('image3',dist_transform)
sure_fg = np.uint8(sure_fg)
#[Link]('image4',unknown)
#Marker labelling
#[Link]('image6',sure_fg)
markers = markers+1
markers[unknown==255] = 0
#[Link]('image7',unknown)
markers = [Link](im0,markers)
#[Link]('image8',im0)
gray2 = [Link](im0,cv.COLOR_BGR2GRAY)
In this module we are applying image morphological function to extract image features.
Code snippet
image = imread(image_path)
try:
# Using KAZE, cause SIFT, ORB and other was moved to additional module
alg = cv2.KAZE_create()
kps = [Link](image)
dsc = [Link]()
To reduce the training time complexity caused by the repeated model parameter tuning
process. The CNN architecture is simple and does not have iterative parameter tuning
that makes the training process faster and achieves adequate performance in disease
classification.
def build_model(input):
model = Sequential()
[Link](Dense(128,input_shape=(input[1],input[2])))
[Link](MaxPooling1D(pool_size=2, padding='valid'))
[Link](MaxPooling1D(pool_size=1, padding='valid'))
[Link](Dropout(0.2))
[Link](Flatten())
#[Link](Dropout(0.2))
[Link](loss='mse',optimizer='adam',metrics=['accuracy'])
return model
Prediction:
model = load_model('results/CNN.h5')
df = pd.read_csv("[Link]")
x = [Link][ : , :-1].values
label_encoder = [Link]()
y = [Link][:, -1:].values
y=label_encoder.fit_transform(y)
print("y==",y)
[Link](X_train, y_train)
prediction1 =[Link](X_test)
System Testing
7.1.1Unit testing
Unit testing involves the design of test cases that validate that the internal program
logic is functioning properly, and that program inputs produce valid outputs. All decision
branches and internal code flow should be validated. It is the testing of individual
software units of the application .it is done after the completion of an individual unit
before integration. This is a structural testing, that relies on knowledge of its construction
and is invasive. Unit tests perform basic tests at component level and test a specific
business process, application, and/or system configuration. Unit tests ensure that each
unique path of a business process performs accurately to the documented specifications
and contains clearly defined inputs and expected results.
7.1.3Functional test
Test Cases
Expected Output The system should be read path and display the path
Expected Output It Should show the alert Message select input image
Expected Output The algorithm should extract the image important features
Expected Output The algorithm should predict the Disease name as per
historical data collected and recommend treatment
Input Import all valid libraries sklearn, tensorflow and keras libraries
Conclusion:
In this project a novel CNN has been presented for the automated diagnosis of
malaria from tiny blood samples. The result shows that the detection accuracy improves
when falsely labeled images are removed. Due to several morphological techniques, the
malaria parasites are highlighted. A lightweight watershed model has been trained on
preprocessed cell images for extracting the most informative features. Eventually, CNN
successfully differentiate between infected and uninfected samples from these features.
The proposed parasitic inflator, watershed capability to extract features, and CNN ability
to generalize make the proposed approach more effective. The efficacy of the proposed
framework is clear from its magnificent outcomes; whereas, the CNN has obtained
97.79% accuracy and 97.88% recall for the original dataset. Also, it has acquired a
promising score of 99.66% accuracy and 99.55% recall for the updated dataset.
Future Enhancement:
We have a plan to work on the object detection models with CNN feature
extractor and to increase the accuracy of malaria detection.
REFERENCES:
[1] W. H. Organization, Malaria microscopy quality assurance manual-version 2. World
Health Organization, 2016.
[3] L. M. Barat, N. Palmer, S. Basu, E. Worrall, K. Hanson, and A. Mills, “Do malaria
control interventions reach the poor? a view through the equity lens,” American Journal
of Tropical Medicine and Hygiene, vol. 71, no. 2, pp. 174–178, 2004.
[5] M. P. Singh, K. B. Saha, S. K. Chand, and L. L. Sabin, “The economic cost of malaria
at the household level in high and low transmission areas of central india,” Acta tropica,
vol. 190, pp. 344–349, 2019.