0% found this document useful (0 votes)
3 views136 pages

Python Machine Learn

The document is a book titled 'Python Machine Learning from Scratch' by Jonathan Adam, aimed at beginners to provide fundamental knowledge of machine learning concepts and applications using Python. It covers various topics including supervised and unsupervised learning algorithms, Python programming basics, and deep learning techniques. The book is designed for a diverse audience including newcomers to computer science, professionals, and educators, and emphasizes practical understanding without overwhelming technical details.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views136 pages

Python Machine Learn

The document is a book titled 'Python Machine Learning from Scratch' by Jonathan Adam, aimed at beginners to provide fundamental knowledge of machine learning concepts and applications using Python. It covers various topics including supervised and unsupervised learning algorithms, Python programming basics, and deep learning techniques. The book is designed for a diverse audience including newcomers to computer science, professionals, and educators, and emphasizes practical understanding without overwhelming technical details.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

e sf

@&} : ‘

7 hime,
, i & t

* \
>oY »

Al SCIENCES
Jonathan Adam
Digitized by the Internet Archive
in 2022 with funding from
Kahle/Austin Foundation

[Link]
PYTHON MACHINE LEARNING
FROM SCRATCH
Machine Learning Concepts and
Applications for Beginners

Jonathan Adam

Alo Gi 5
How to contact us

Please address comments and questions concerning this book


to our customer service by email at:

contact([Link]

Our goal ts to provide high-quality books for your technical learning in


computer science subjects.

Thank you so much for buying this book.

eel
L

AL SCIENCES
Table of Contents
Pla WerOl CONS once
es cassette
oeyiasiucttes eoesece sesseste: ill
Prom Al Sciences PubMsher, ss. .[Link]« ROCCE EEE 3
DEC
LA CCiiss censesagstvesscooss (ooasescinderdesPiet onss5streaesessasse sans. canys 5

Imtroduction...........ccccccccceees cere pate ctacse donk tt. sea ssuadeses S850 8


Do You Really Need to Know Statistics & Python?............:.::00 9

tno ier eodhinemitisletdyc ccccoacsscresurerorancteonmatacd occa 3)

Pychonminables Quick Proto


ty plier. giile tees Aidit
cece S)

PythonsHas Anvesome Selentifie Libraries: .[Link]


ness 10

Nieastor Stidy to Exp lone iscagitecia


a ee Ma itp

Machine Pearnitrg 2.0.22. cccisersssorssesosacses esters Deaseeesateresweces


Witat de NiAchine Learning o,c.-caesccscccenessstorrsoereretenancaneseeasasasrecess 12

Supetvised Learning Algorithms .,.....s..[Link] 13

Unsupervised Learning Algorithms .............cccsescccceeseeeeeeeeeeeesanees 15

Semi-supervised Learning Algorithms ............::ccccsseseeceeeeeerseeeees Me

Reinforcement Learning Algorithms ............::ccccesscsscsccssssssseeees 19

Overtitane and Umderiitting o..-.:<ccncactsenercecsevatacvaraeeersraneevssuterser 20

MCOTE COMMESS vices cgeceechoesasasvecrevsser <asdsvotscessansverenndetuceseeessericeseesneress P&S)

he Bias-Vatiance Utade-Ott.-[Link] 24

Feature Extraction and Selection ..............000sssssssecssesccssecaroenscones 27

Why Machine Learning is Popular.............-cesssessseerecessreeenaeeees 27

Overview of Python Programming Language .............. 29


The Python Programming Language ........ccceesseeseseeessserseeees 29

11
Common Python Syntax........cccccccccccesssessstcensceceecessersenensneneeneeaeses 31

Python Data Structures ..........csscccceesssececeecnsecceceesseseesseesessereceseees 36

Python for Scientific Computing..........csccsssnccessreerereecccssesoees 38

R@STOS 910115 a9. cccaee cree vstteve etanenssaitnccpeccataoeds


Parteceset ete eee
Introduction to Labels and Features............ccssseseeessceseeereseeeseeeres 42

PPG ATUPES sar. cnccactcncctsaveessccoccacuensteccetesenrecvaqvrss <Pesndvecee-usudnesacBezaseys! 42

A Regression Example: Predicting Boston Housing Prices ...... 43

Steps To Carry Out Analysis ..........cecscssesscsscscesessssessnscronansssocesens 44

[ni port, LAD PATieS!iccccs-accesnvetxens>onadevseveetsesose-eateevauesseecedyeesaercs


testseve 44

How to:fotecast and) P&edict cc isacsesrssceesscoeseosvaeatoeees Misret Se 47

Classification... <cusszdsssesc-ssoontteetsses4sstossenad-odsehesed semeseesAD


Multi-Class Classification: <iissc.k:.cveosecensess
asus conse i¢ivuces sncktsaceser ss50

Popular Classification Algorithms. ..........[Link] 52

Introduction to K Nearest Neighbots ............ccsssessececeeeeeeeeeneeeces 52

How to create and test the K Nearest Neighbor classifier......... 56

Introduction to Support Vector Machine.................0006+ 64


How to create and test the Support Vector Machine (SVM)
CIASGIAIER: ro ser ocycuveenseatac eontensnapteer estes oeowkeene cacy sctee mu cneres taster eee 66

Clustering ........ Beads sais Souaneseseastescee


fees uecensuuneessccatnnesaie: 69
Introduction: to, Chis tenting acc ossssescsptaaccassusassessstennsseacecavacaccntoateas 71

Example of [Link]
ossesugstanessv atvchcerrassave/uni esse 12

Running K-means with Scikit-Leatn............


cs cssescessesrceceesseees 72

Introduction to Deep Learning using TensorFlow....... 80


Deep Learning Compared to Other Machine Learning
APPLOaChes ivccscsaccausevoc canes icodussoustsvese cenyvarteuseseaaen eid pepeeaeeenaeantees 85

Applications of Deep Learwing s...[Link]-s


coeaces nese teereeess tee 86

IV
Python Deep Learning Frameworks ...............:cccssscssseecsenencessnees 87

dTUSEAIL AL ORG GEU LONG ve cxsexsacosecastaystetascsesvteot


cinesdsessenncesneessvamheneeetsets 88

Installing TensotFlow on Windows ...........:ccsscccsssscessscsssccssseeees 88

Installing dLemsOre low LAX wscioancsis cox s0ssassiatcsacteswersscerenecancees 88

fnstalling Vensorklow otf macOS Giscssc<tsssccveisespssesetesviventes


svsveess« 89

How to Create a Neural Network Model ..............:ccccssssscsesseeeeee 90

How to run the Neural Network using TensorFlow .... 93


OW CO! SCL Ue ACA ereesvene aunenreterstsissearscnaratesauuhs Visieksitstvenscsnent 93

Elow touteainiand test the Gatas.,..[Link].-cesescsevesvescnsaveoceveccccnsessoeest 94

Case Studies with Real Data .......... Sreestea tretee cea eeseees ee LUO
Beatle CHER VOCCIIN Ge no ccscensnnessaspacesonsessepeeesseanroreannetsrssasuccncts 105

SG CHINETIE ANIALVIGIG o coe vata ren ca canect-c ish iv ondsborerscantasecscpannuncsacecerenetes 110

Conclusion............... Batitea
iene sesemeaeee TheAwoseueoe saassetee ea 119

MEH AMIC VOU


ose seecteses tou ate mesuncasseqeendaceves ssencsscecgteastsacedsents 120
Sources & References...... Gee ay ore eet
Software, libraries, & programming language ...........eeeeeeeeeeeeee 121

DVALASCUSherrece seus ca cones tecees Sees evssseenccoanes cveustecedecatcesavess sess vresasceee’ 121

Online books, tutorials, & other references............cccccseeesesseeeeee 121

PUica ers seco crete ee. yncs


eeoeeeeescceseaveenccieesencetesnnanatsa
recs124
© Copyright 2016 by AI Sciences LLC
All rights reserved.
First Printing, 2016

Edited by Davies Company


Ebook Converted and Cover by Pixels Studio
Publised by AI Sciences LLC

ISBN-13: 978-1725929982
ISBN-10: 1725929988 ‘

The contents of this book may not be reproduced, duplicated or


transmitted without the direct written permission of the author.

Under no circumstances will any legal responsibility or blame be held


against the publisher for any reparation, damages, or monetary loss
due to the information herein, either directly or indirectly.
Legal Notice:

You cannot amend, distribute, sell, use, quote or paraphrase any part
ot the content within this book without the consent of the author.

Disclaimer Notice:

Please note the information contained within this document is for


educational and entertainment purposes only. No warranties of any
kind are expressed or implied. Readers acknowledge that the author
is not engaging in the rendering of legal, financial, medical or
professional advice. Please consult a licensed professional before
attempting any techniques outlined in this book.

By reading this document, the reader agrees that under no


circumstances is the author responsible for any losses, direct or
indirect, which are incurred as a result of the use of information
contained within this document, including, but not limited to, errors,
omissions, or inaccuracies.
From AI Sciences Publisher

Al Sciences sae Al Sciences

ene.
Learning
Fundamentals
= oe
With Python —
AN INTRODUC TION FOR BEGINNERS : i STEP IBY oFE 2 flu EWI
Wi HERERAS AN DPYTORGH:
ORC

AI SCIENCES

Chad Pan = Chao Pan

Data Science Dente Apeaee


WOM Scratch HOM SCI QiCh
wilh ee eS with Python —
STEP BY. STEP: GUIDE BES STEER BYES ER GUIDE
—=
nas Peters Morgan =
=
Peters Morgan
[Link]

EBooks, free offers of eBooks and online learning courses.

Did you know that AI Sciences offers free eBooks versions of


evety book published? Please subscribe to our email list to be
aware about our free eBook promotion. Get in touch with us
at [Link] for more details.

eG
RUS@IPINGCES
BOOKS
At [Link] , you can also read a collection of free
books and received exclusive free eBooks.
Preface

“Some people call this artificial intelligence, but the reality is this technology will enhance us. So
instead ofartifical intelligence, I think we'll augment our intelligence.”

Ginmt Rometty

The main purpose of this book ts to provide the reader with


the most fundamental knowledge of machine learning with
Python so that they can understand what these are all about.

Book Objectives

This book will help you:

e Have an appreciation for machine learning and deep learning


and an understanding of their fundamental principles.
e Have an elementary grasp of machine learning concepts and
algorithms.
e Have achieved a technical background in machine learning
and also deep learning

Target Users

The book designed for a variety of target audiences. The most


suitable users would include:

eNewbies in computer science techniques and machine


learning
¢ Professionals in machine learning and social sciences
e Professors, lecturers or tutors who are looking to find better
ways to explain the content to their students in the simplest
and easiest way
eStudents and academicians, especially those focusing on
machine learning practical guide using R
Is this book for me?

If you want to smash machine learning from scratch, this book


is for you. Little programming experience is required. If you
already wrote a few lines of code and recognize basic
programming statements, you'll be OK.

Why this book?


This book is written to help you learn machine learning using
Python programming. If you are an absolute beginner in this
field, you'll find that this book explains complex concepts in
an easy to understand manner without math or complicated
theorical elements. If you are an experienced data scientist, this
book gives you a good base from which to explore machine
learning application.

Topics are carefully selected to give you a broad exposure to


machine learning application. While not overwhelming you
with information overload.

The example and cases studies are carefully chosen to


demonstrate each algorithm and model so that you can gain a
deeper understand of machine learning. Inside the book and in
the appendices at the end of the book we provide you a
convenient references.

You can download the source code for the project and
other free books at:
http: / /[Link]/code
Your Free Gift

As a way of saying thanks you for your purchase, AI Sciences


Publishing Company offering you a free eBook in Machine
Learning with Python written by the data scientist Alain
Kaufmann.

Alain Kaufmann

It is a full book that contains useful machine learning


techniques using python. It is 100 pages book with one bonus
chapter focusing in Anaconda Setup & Python Crash Coutse.
AI Sciences encourage you to print, save and share. You can
download it by going to the link below or by clicking in the
book cover above.

http: / /[Link]/free-books

If you want to help us produce more material like this,


then please leave an honest review on amazon. It really
does make a difference.
Introduction

The importance of machine learning and deep learning is such


that everyone regardless of their profession should have a fair
understanding of how it works. Having said that, this book is
geared towards the following set of people:

@ Anyone who 1s intrigued by how algorithms arrive at


predictions but has no previous knowledge of the field.
® Software developers and engineers with a strong
programming background but seeking to break into
the field of machine learning.
@ Seasoned professionals in the field of artificial
intelligence and machine learning who desire a bird’s
eye view of current techniques and approaches.

While this book would seek to explain common terms and


algorithms in an intuitive way, it would however not dumb
down the mathematics on whose foundation these techniques
ate based. ‘There would be little assumption of prior knowledge
on the part of the reader as terms would be introduced and
explained as required. We would use a progressive approach
whereby we start out slowly and improve on the complexity of
our solutions.

To get the most out of the concepts that would be covered,


readers are advised to adopt a hands on approach which would
lead to better mental representations.

Finally, after going through the contents in this book and the
accompanying examples, you would be well suited to tackle
problems which pique your interests using machine learning
and deep learning models.

Do You Really Need to Know Statistics & Python?

For many using machine learning for day to day tasks, Python
is the programming language of choice. They are many reasons
for this bias towards Python. Let’s have a look at some of
them:

Python 1s Beginner Friendly:


The syntax of Python programming language is intuitive and
easy to learn for beginners and advanced professionals alike.
Python uses regular English words in a way that makes Python
code readable and easy to understand. Unlike some other
programming languages, Python does not use curly braces but
instead relies on indentation to separate blocks of code. The
classic “Hello World!” example in Python is just one line of
code

print (“Hello World!”)

Python Enables Quick Prototyping:


Python is often the preferred choice when you want to go from
idea to implementation in as short a time as possible. The
reason this 1s usually the case is because Python is a general
programming language that is dynamically typed. What that
means is that you do not need to compile your code before you
run it. You can rather use an interactive development cycle
whereby your write code, try things out and see what works,
then iterate progressively towards an optimal solution.

Teams seeking to bring products to market would use Python


to iterate quickly and use feedback from users as a loop to
improve the model. The advantage of being an early player in
a market cannot be overemphasized especially with data
products where more users mean more data and more data
means a better performing model.

Python Has Awesome Scientific Libraries:


Python has a host of world class libraries for scientific
computation which are mature and have been used in
production for many years. These libraries have popular
machine learning models implemented and are easily extensible
providing an array of tools for more experienced users to do
So.

You can get started building your own machine learning


models using these libraries in Python by calling various
components and assembling it into a stack peculiar for your
learning task. However, beyond this point it is desirable to get
an understanding of the mathematical underpinnings of these
models as it would help you tweak your models to achieve
better results.

An investment in studying the mathematics of machine


learning would lead to huge payoffs long term as models would

10
no longer be a black box and you can peel behind the curtain
to have a look when things don’t work as expected.

Areas of Study to Explore


There are three main areas of study to explore when pursuing
a deeper understanding of the mathematical models that
comprise the landscape in machine learning. The most
important area is probability and statistics, then linear algebra
and Calculus.

So strictly speaking, you do not need to be a statistician to get


started using machine learning to derive real world value, but
it does bestow an advantage when you understand the
mathematics behind why models work the way they do.

1]
Machine Learning
What is Machine Learning?

Machine learning has recently been attracting attention in the


media for several reasons, mainly because it has achieved
impressive results in various cognitive tasks such as image
classification, natural language understanding, customer churn
prediction etc. However, it has been regarded as some sort of
magic formula that is capable of predicting the future, but what
really is machine learning. Machine learning at its simplest form
is all about making computers learn from data by improving
their performance at a specific task through experience. Similar
to the way humans learn by trying out new things and learning
from the experience, machine learning algorithms improve
their capability by learning patterns from lots of examples. The
performance of these algorithms, generally improves as they
are exposed to more data (experience). Machine learning 1s
therefore a branch of artificial intelligence that aims to make
machines capable of performing specific tasks without being
explicitly programmed. What this means is that these
algorithms are not rule-based, the entire learning process 1s
constructed in such a way as to minimize or completely
eliminate human intervention.

Machine learning algorithms are typically used for a wide range


of learning ptoblems such as classification, tegression,
clustering, similarity detection etc. Many applications used in
the real world today are powered by machine learning.
Applications such as personal assistants on mobile phones use
machine learning algorithms to understand voice commands

12
“a
spoken in natural language, mobile keyboards predict the next
word a user is typing based on previous words, email clients
offer a smart reply feature whereby the content of an email is
scanned and appropriate responses are generated, e-commerce
applications offer recommendation to users based on previous
purchases and spending habits etc. Nearly every industry
would be impacted by machine learning as most processes can
be automated given that there is enough training data available.
Machine learning algorithms mostly excel in tasks where there
is a clear relationship between a set of inputs and outputs
which can be modelled by training data. Although machine
learning is a rapidly improving field, there is as of now no
notion of general intelligence of the form displayed by humans.
This is because models trained on one task cannot generalize
the knowledge gleaned to perform another task, that 1s
machine learning algorithms learn narrow verticals of tasks.

There are three main branches of machine learning namely -


supervised learning, unsupervised learning and reinforcement
learning. In some cases a fourth branch is mentioned - semi-
supervised learning but this is really just a special instance of
supervised learning as we would see in the explanations below.

Supervised Learning Algorithms

Supervised learning is by far the most common branch of


machine learning. Most of the real world value currently in the
field of machine learning can be attributed to supervised
learning. Supervised learning algorithms are those machine
learning algorithms which are trained with labelled examples.

13
It would be remembered that we defined machine learning as
making algorithms that learn from data (examples) without
being explicitly programmed. The main intuition to understand
when dealing with supervised learning algorithms is that, they
learn through the use of examples that are clearly annotated to
show them what they are supposed to learn. The algorithms
therefore try to find a mapping representation from inputs to
outputs using the labels as a guide. “Supervised” in the name
of these types of algorithms, point to the fact that the labels or
targets provide supervision throughout the learning process. It
is therefore possible for the algorithm to check its prediction
against actual values stored in the labels. It then uses this error
information (how far off its prediction was from the actual
label) to slowly improve its performance with each iteration.
The targets in a supervised learning problem can be seen as a
supervisor providing feedback to the algorithm on areas where
it can improve its performance. The two main applications of
supervised learning algorithms are classification and
regression.

Classification involves training a learning algorithm to correctly


separate examples into predefined categories or classes. The
classes ate usually chosen ahead of time by a human expert
with domain knowledge in the field where the learning
problem is posed. The examples that are used to train the
model are clearly labelled to indicate the category they belong
to. During training, the supervised learning algorithm, uses the
labels to guide its learning and at test time it is capable of
correctly predicting the categories of new examples. A popular
example of classification is spam detection where an email is
correctly identified to belong to one of two classes - spam or

14
not spam. Depending on what is predicted, appropriate action
could be taken such as shifting spam emails to a spam folder
while relevant emails are sent to a user’s inbox.

Regression is a learning problem where the algorithm is


interested in predicting a single real value number. Regression
is used where a single numeric entity is to be predicted. An
example of regression would be predicting the age of a person
given a profile picture or predicting the salary of an individual
given information about the individual such as level of
education, work experience, age, country of residence etc. It
would be observed that in both cases the final prediction 1s a
single number.

Supervised learning algorithms are easier to train when


compated to unsupervised or reinforcement learning
algorithms. ‘his is because the presence of labels simplify the
learning problem since there exist a clear way of determining
performance during training. However, it should be noted that
most supervised learning problems can also be modelled as
unsupervised if we get rid of labels. Datasets for supervised
learning are more expensive to acquite as it requires meticulous
human annotation of examples. The fact that most data in the
world today are in an unlabelled form makes the research of
unsupervised learning algorithms particularly important.

Unsupervised Learning Algorithms

Unsupervised learning involves learning directly from raw data.


This type of learning takes place without the presence of a

15
supervisor in the training loop in form of labels. Unsupervised
learning algorithms are free to explore the underlying data
distribution and come up with patterns that best describe the
entire dataset. The training process is not guided by humans
through labelled examples and as such unsupervised learning
algorithms ate more powerful as they can discover patterns
which domain experts may not have thought of. It is however
still the job of domain experts to understand the patterns so
discovered and explain them because unsupervised learning
algorithms do not truly have the sense of reasoning which we
would ascribe to humans. Unsupervised learning algorithms
merely use the data distribution or its latent (hidden)
representations to unearth insights which may be in the form
clusters, groups or distributions.

There are many applications of unsupervised learning


algorithms such as clustering, dimensionality detection,
generative models etc. Clustering is one of the popular
implementations of unsupervised learning. It involves the
automatic discovery of groups (clusters) of data points from
raw data. Members of a group share similar features, that 1s
they are alike. They can be thought of as belonging to the same
type of entity whereas as a group they are dissimilar to other
groups. Groups usually have semantic meaning which can be
further explored to understand the dataset. An example of
clustering would involve grouping customers of a tv streaming
service into the type of shows that they watch. Users with
similar interests would generally be found in the same group.
This is a very powerful application as new tv shows could be
recommended to users based on other users who share their

16
interests, leading to greater engagement on the platform and
increased revenues.

Dimensionality reduction is a machine learning technique that


reduces the number of attributes (features) fed in a model to
only the most relevant ones which drive predictions. It has
been observed empirically, that models with greater number of
features or dimensions perform worse on generalization. ‘That
is to say that their performance suffers in the real world. By
reducing the number of dimensions, a model can learn from
informative features which enables it to sdevelop valid
representations about the data that aids prediction. Principal
component Analysis (PCA) 1s probably the most popular
dimensionality reduction technique and 1s an example of an
unsupervised learning algorithm. PCA reduces the dimensions
of data by identifying those axes that contain the most
variability. What that means is that it discovers principal
components which offer the most discriminative features.

Generative models are another popular instance of


unsupervised learning algorithms. They have made recent
headlines because of their ability to artificially generate
photographs and works of art that look realistic. Generative
Adversarial Networks (GANs) currently produce state of the
art results across many image generation benchmarks and are
among the most popular examples of generative models.

Semi-supervised Learning Algorithms

lig
Semi-supervised learning algorithms are a special case of
supervised learning algorithms. In semi-supervised learning,
while there isn’t an explicit label, there exist an implicit
heuristic which serves as a supervisor in the training loop.
Semi-supervised models do not contain any external source of
labels but only rely on input features. However, the learning
task is set up in such a way that supervision still takes place in
the form of extraction of pseudo-labels from inputs through a
heuristic algorithm. A popular example of semi-supetvised
learning algorithms are autoencoders. Let us look at an
example to expand our understanding.

Input Reconstructed input


Py 4

|
Latent Space |
Representation fou |
|
|

In the autoencoder above, the learning task 1s to reduce the


dimensions of the input into a smaller latent space representing
the most important hidden features, then reconstruct the input
from this lower dimensional space. So given an input, example
an image, an autoencoder shrinks the image into a smaller
18
latent representation that still contains most of the information
about the image, then reconstructs the original input image
from this low dimensional space. Even if there are no explicit
labels, it would be observed that the input serves as the
supervisor since the learning task is to reconstruct the input.
Once such a model is trained to compress features into a
smaller dimension, the compressed features can serve as the
starting point of a supervised learning algorithm similar to
dimensionality reduction using PCA. The first part of the
network that reduces the dimensions of the input is called an
encoder while the second part that scales the encoded features
back to the full size input is called the decoder. .

Reinforcement Learning Algorithms

In reinforcement learning there are three main components, an


agent, an environment and actions. The goal of reinforcement
learning is to train an intelligent agent that is capable of
navigating its environment and performing actions that
maximizes its chances of arriving at some end goal. Actions
carried out by the agent change the state of the environment
and rewards or punishment may be issued based on the actions
taken by the agent. The challenge is for the agent to maximize
the accumulated rewards at the end of a specific period so that
it can actualize an end goal (objective).

nS)
Observation
Reward Action

Agent

In the schematic diagram of reinforcement learning above, an


agent (the reinforcement learning) interacts with the world
(environment) through actions. The environment provides
observations and rewards to the agent based on the kind of
action taken by the agent. The agent uses this feedback to
improve its decision making process by learning to carry out
actions associated with positive outcomes.

Overfitting and Underfitting

Overfitting and underfitting jointly form the central problem


of machine learning. When training a model we want to
improve its optimization by attaining better performance on
the training set. However, once the model is trained, what we
cate about is generalization. Generalization in a nutshell deals
with how well a trained machine learning model would

20
perform on new data which it has not seen, that is data it was
not trained on. In other words, how well can a model
generalize the patterns it learnt on the training set to suit real
world examples, so that it can achieve similar or better
performance. This is the crux of learning. A model should be
able to actually learn useful representations from data that
improves test time performance and not merely memorize
features as memorization is not learning.

We say a model has overfit to a training set when it has failed


to learn only useful representations in the data but has also
adjusted itself to learn noise in order to get an artificially high
training set accuracy. Underfitting means that the model has
not used the information available to 1t but has only learnt a
small subset of representations and has thrown away majority
of useful information, thereby making it to make unfounded
assumptions. The ideal situation is to find a model that neither
underfitts nor overfitts but exhibits the right balance between
optimization and generalization. This can be done by
maintaining a third set of examples known as the validation set.
The validation set is used to tune (improve) the performance
of the model without overfitting the model to the training set.
Other techniques for tackling overfitting includes applying
regularization which punishes more complicated models and
acquiring more training examples. Underfitting can be stymied
by increasing the capacity of the learning algorithm so that it
can take advantage of available featutes.

21
| ta
x - I
XO 0 Xo Oo X0X' 0
iy XO ly: O ~
es Ax 0 Xx 0
| Xoy | anes oe | LX
asm ei. arenecd 7 Rane » a e

The plots above show three simple line based classification


models. The first plot separates classes by using a straight line.
However, a straight line is an overly simplistic representation
for the data distribution and as a result it misclassified many
examples. The straight line model is clearly underfitting as it
has failed to use majority of the information available to it to
discover the inherent data distribution.

The second plot shows an optimal case where the optimization


objective has been balanced by generalization criterion. Even
though the model misclassified some points in the training set,
it was still able to capture a valid decision boundary between
both classes. Such a classifier is likely to generalize well to
examples which it was not trained on as it has learnt the
discriminative features that drive prediction. The last plot
illustrates a case of overfitting. The decision boundary 1s
convoluted because the classifier is responding to noise by
trying to correctly classify every data point in the training set.
The accuracy of this classifier would be perfect on the training
set but it would perform horribly on new examples because it

eu)
optimized its performance only for the training set. The trick
is to always choose the simplest model that achieves the
greatest performance.

Correctness

To evaluate a machine learning algorithm, we always specific


measutes of predictive performance. These metrics allow us to
judge the performance of a model in an unbiased fashion. It
should be noted that the evaluation metric chosen depends on
the type of learning problem. Accuracy is a popular evaluation
metric but it is not suitable for all learning problems. Other
measures for evaluation include recall, precision, sensitivity,
specificity, true positive rate, false positive rate etc. The
evaluation used should be in line with the goals of the
modelling problem. ,

To ensure fidelity of reported tnetrics, a rule of thumb ts that


models should never be trained on the entire dataset as any
evaluation reported by metrics is likely skewed because the
model’s performance when exposed to new data is
unascertained. The dataset should be divided into train,
validation and test splits. The model 1s trained on the training
set, the validation set is reserved for hyperparameter tuning for
best performing models and the test set is only used once at
the conclusion of all experimentation.

A confusion matrix is widely used as a simple visualization


technique to access the performance of classifiers in a
supervised learning task. It is a table where the rows represent

pa8
the instances in the actual class (ground truth) while the
columns represents predictions. The order may be reversed in
some cases. It 1s called a confusion matrix because it makes it
easy to see which classes the model is misclassifying for
another, that is which classes confuse the model.

Predicted class

P N

True False
P | Positives Negatives
(TP) (FN)
Actual
Class
False True
N Positives Negatives
(FP) (TN)

The examples which the model correctly classified are on the


diagonal from the top left to bottom right. False negatives are
positive classes which the classified wrongly predicted as
negatives while false positives are negative instances which the
classifier wrongly thought were positives. Several metrics like
true positive rate, false positive rate, precision etc are derived
from items in the confusion matrix.

The Bias-Variance Trade-off

The bias of a model is defined as the assumptions made by the


model to simplify the learning task. A model with high bias

24
makes assumptions which are not correlated by the data. This
lead to errors because predictions are usually some way off
from actuals. Variance on the other hand is how susceptible a
model is to noise in the training data. How widely does the
performance on the model vary based on the data it 1s
evaluated on. A good machine learning algorithm should strive
to achieve low bias and low variance. Bias and variance are
related to overfitting and underfitting earlier encountered. A
model with high bias is underfitting the training data because
it has made simplistic assumptions instead of learning from
information available. Similarly, a model with high variance 1s
overfitting, because it has modelled noise and ‘as a result, its
performance would vary widely across the training set,
validation set and test set.

Let us look use a dart analogy to further explain these concepts.

pas
Low Variance High Variance

Low
Bias

Bias
High

The top left image represents a model that has low bias and
low variance. This is the ideal model as it has learnt to hit the
target (produce correct classification) and usually hits the target
most of the time (does not vary with each throw). The image
at the top right shows a model that exhibits high variance and
low bias. Even if it does not make a lot of assumptions, its
ptedictions are spread all over the board which means its
performance varies widely (high variance). The image on the
bottom left depicts a model with high bias and low variance.
The shots are not all over the board but in a specific location.
This location is however far from the target meaning the model
is biased because of simplistic assumptions. Finally, the image
on the bottom right shows a model with high bias and high
variance. The shots on the board vary widely and are far away

26
from the target. This is the worst kind of model as it hasn’t
learnt any useful representation.

Feature Extraction and Selection

Feature extraction involves performing transformation on


input features that produce other features that are more
analyzable and informative. Feature extraction may occur by
combining original features to create new features which are
better suited for the modelling problem. This is similar to
feature engineering where we create new features to fed into a
model. An example of feature extraction is Principal
Component Analysis (PCA).

Feature selection is choosing a subset of features from the


original input features. The features selected are those that
show the most correlation with the target variable, that is those
features that drive the predictive capability of the model. Both
feature extraction and feature selection leads to dimensionality
reduction. The main difference between them is that feature
extraction is a transformation that creates new features
whereas feature selection chooses only a subset of available
features. Since feature selection removes certain features, it is
always advisable to first do feature extraction on a dataset, then
select the most important predictors via feature selection.

Why Machine Learning is Popular

The popularity of machine learning in the artificial intelligence


community and society in general is mainly because machine
learning techniques have proven to be highly successful in
various niches providing business value for a slew of

27
operations from fraud detection to speech recognition to
recommender systems. It is embedded in products we use
every day. When you buy a product from Amazon and you are
given a list of suggestions of other products that go with it,
that’s machine learning in action. When you open your mailbox
and emails are automatically classified into folders based on
their similarity, those are machine learning models doing the
work behind the scenes. Even when you use your credit card
online and your transaction is successful, a machine learning
model approved your transaction as being normal and not
fraudulent.

In light of all these, companies and organizations have poured


in more money into developing better performing models
through research and collaboration between industry and
academia.

It wasn’t always the case that machine learning was the darling
of the computer science community, however in recent years
three factors have conspired to give it an exalted place.

e With the dawn of the internet, more data was being


collected and this lead to the age of big data.
e Computational resources became faster and cheaper
with the arrival of Graphics Processing Units (GPUs),
which were originally developed by the gaming
industry to render graphics but are well suited for
parallel computation.
@ Better algorithms were developed which lead to better
accuracy of models.
Overview of Python Programming
Language
The Python Programming Language

Python is a general purpose programming language language


used in web development, scientific computing, system
administration, software development etc. Python favors
readability and has an English like syntax. One of the main
differences between Python and other programming languages
like Java, C++, PHP etc is that it does not make use of curly
braces to define scope. Blocks of code, called suites in Python
are delineated using whitespace. Python is also a dynamically
typed language which means variable types do not have to be
declared before variables are used. It makes use of duck typing
which makes certain assumptions about the type of a variable
from its content. Duck typing in programming is coined from
the popular phrase - “If it walks like a duck and quacks like a
duck, then it’s a duck”. What this means in essence is that if an
object exhibits certain properties, then its type can be deduced.
This is a powerful assumption as it allows code to be written
interactively and defers type checking to when the program is
actually run (runtime).

Python is a very expressive programming language and is


favored by programmers and scientists because it encourages
quick prototyping and significantly reduces development time
when compared to other high level programming languages. In
data science, Python enables fast iteration of data science
projects and because Python is a general purpose programming
language, prototype models can be more easily integrated into

29
production workflows as there is usually no need to switch to
another programming language. In this way, one language can
be used for the entire stack, from prototyping to deployment.
It should be noted however, that this generally depends on the
type of application and Python especially for scientific
computing usually reference lower level extensions in faster
programming languages like C or C++. Python can be seen as
the ideal interface to work across a slew of tasks effectively and
efficiently.

Python was created by Guido van Rossum in 1991 and has


undergone several iterations. There are currently two major
versions of Python - Python 2 and Python 3. At the time of
this writing the development of Python 2 has been
discontinued so it 1s advised to use Python 3 for all new
projects. For this reason, the examples we would come across
in this book all assume a Python 3 environment.

There are several ways to set up a local Python development


environment such as installing Python natively on a computer,
setting up a virtual environment using a tool like virtualenv or
using a bundled scientific distribution like Anaconda. For the
purposes of this book, we would leverage the immensely
popular scientific distribution Anaconda because it contains
many prepackaged Python scientific libraries, some of which
we would use throughout this book. Anaconda is cross
platform and is available for the major operating systems -
Windows, Linux, macOS etc. You can download the
Anaconda installer for your operating system by going to their

30
download page ( ‘ [Link]/downloads) and
following the installation instructions.

Once Anaconda has been installed on your computer, new


Python packages can be installed by using Anaconda’s package
manager, “conda” or using Python’s native package manager -
“pip”. It should also be noted that packages should be
upgraded or deleted using the appropriate package manager
through which the package was installed.

Here are examples instructions on installing packages from the


y

terminal for both cases:

$ conda install package name # installation via conda package


manager

$ pip install package


Sates
name# installation via pip

Similarly, an installed package can be updated using the


following commands.

$ conda update package name # update via conda package


manager

$ pip install --upgrade package_name # upgrade via pip

Common Python Syntax

In this section, a brief overview of basic Python syntax is


presented. As would become evident, indentation is important
as all lines of code in the same suite (block) must be indented
by an equal number of characters. By convention, this is usually

31
4 white spaces. Let us honour traditions and start with a simple
Hello world! example in Python.

print (Hello World!)

The code above outputs the string “Hello World!” to the


screen.

Next, we look at variable assignment. Variables can be seen as


containers that point to an entity or stored value. Entities are
assigned to a variable using the equality operator. The value on
the right hand side is put into the container on the left hand
side. A variable usually has a name and calling the variable by
its name references the stored object.

a= 3

=) 4

c=at+b

print('The value of a is {}, while the value of


is) Sle) Wj), flavel Veleyswine fyb vel ELEY {ip cumepantcne
(ey. oy {2))))

The value of a is 3, while the value of b is 4, and their sum c is 7

The code above assigns an integer with a value of 3 to the


variable named a, it also assigns 4 to b and finally computes the
sum of a and b and stores it in a new variable c. It should be
noted from the above piece of code that we never explicitly
defined the types of variables we created, rather the type
information was gotten from the kind of entity the object
ey
contained. There are mainly types of mathematical operations
available in Python apatt from addition used above. A good
approach is to familiarize yourself with the Python
documentation as the standard Python library contains many
useful utilities. The documentation for Python 3 can be
accessed at https:/ [Link]/3.

Moving forward, Python supports the use of conditionals to


determine which suite of code to execute. Regular conditionals
such as if, else if and else are available in Python. One thing to
note is that else if is expressed as elif in Python. Let us look at
a simple example below.

a = 200

b = 33

alia Joy Be Elg

print("b is greater than a")

elif a ==

print("a and b are equal")

else:

print("a is greater than b")

a is greater than b

The code snippet above uses conditionals to determine which


suite of code to run. Suites use whitespace indentation for
separation and the output printed to the screen is determined

=f
by the evaluation of the conditional in line with the declared
variables contained therein.

Another important Python syntax are loops. Loops are used


for repeating a block of code several times. They may be used
in conjunction with conditionals.

for x in range(2):

primes)

The code above prints 0 and 1 to the screen. Python indexes


start from 0 and the range function in Python is non inclusive.
What that means 1s that the last value of a range is not included
when it is evaluated. Loops are a very useful construct in
Python and generally are in the form shown above. There 1s
also another form known as while loops but for loops are used
more often.

This brings us to the concept of a function. A function is a


block of code which has been wrapped together and performs
a specific task. A function is usually named but may be
anonymous. In Python, a function is the major way we write
reusable chunks of code. Let us look at a simple function
below.

def my function (planet) :

print('Hello ' + planet)

34
A function is defined using the special keyword def. A function
may accept arguments or return a value. To call a function
(execute it), we type the function name followed by a
parenthesis containing parameters if the function expects
arguments, else we call it with empty parentheses.

my function ('Earth!")

Hello Earth!

Comments in Python are ignored by the interpreter and can be


used to explain code or for internal documentation. ‘There are
two types of comments in Python. The first uses the pound or
hash symbol which the second is known as a docstring and
uses 3 quotation marks.

# a single line comment using pound or hash


symbol

A multi-line

comment in Python
Ua

print('Comments in Python!')

35
Python Data Structures

Data structures are how data 1s collectively stored for easy


access and manipulation. There are several data structures in
Python which enables quick prototyping. Data structures are
the format in which data is stored and usually includes the
kinds of operations or functions that can be called on the data.
The most popular data structure in Python are lists. Lists can
contain different types of data and are ordinal. Lists are
synonymous to arrays in other programming languages. Other
data structures includes a tuple which is a collection that
cannot be modified, a set, which is an immutable list with
unique values and a dictionary, which 1s a key-value pair data
structure. Let us look at how to create each of them below.

my list = ['apple', 'banana', 4, 20]

print(my list)

['apple', 'banana', 4, 20]

Lists can also be defined using the list constructor as shown


below.

another list = list(('a', 'b’, ives Ih).))

print (another list)


Tuples are immutable, this means that we cannot change the
values of a tuple, trying to do so would result in an error. Below
is how tuples are created.

my tuple = (1, 2, 3, 4)

print
(my tuple)

print (type (my tuple) )

(2, oy oo
sclass: "tuple‘>

Using the inbuilt type function gives us the type of an object.

Sets are unordered collections that can contain only unique


values. Sets ate created using curly braces as shown below.

my set = {1, 1, 2, 2, 2, ‘three"}

print
(my set)

print
(type (my _set) )

{'three', 1, 2}
<class ‘set'>

37
In the example above, notice that all duplicate entries are
removed when the set is created and there is no concept of
ordering.

A dictionary is a collection of key value pairs that are unordered


and can be changed. Dictionaries are created using curly braces
with each key pointing to its corresponding value.

mye diets {tive aone’,, '2's two’, "3's ' three” }

print(my dict)

print (type (my dict) )

le heme ee Se es}
<class ‘dict'>

There are other data types in Python but these are by far the
most commonly used ones. To understand more about these
data structures and which operations that can be performed on
them, read through the official Python documentation.

Python for Scientific Computing

One of the reasons for the rapid adoption of Python by the


scientific community is because of the availability of scientific
computing packages and the relative ease of use as most
scientists are not professional programmers. This has in turn
lead to better algorithms being implemented in many Python
scientific packages as the community has evolved to support
several packages. Another treason for the widespread adoption

38
of Python in data science and in the larger scientific community
is because Python is a well designed language and is useful
across several tasks, so users do not need to learn a new
programming language when confronted with a new task but
can tather leverage Python’s rich ecosystem of libraries to
perform their tasks. Python 1s also easy to pick up so users can
learn to extend libraries to support the functionality that they
desire. This forms a virtuous cycle as libraries become more
mature and support a wider range of adopters.

Scipy, also known as scientific python contains several


packages that build on each other to provide a rich repository
of scientific tools. Numpy or numerical Python enables
numerical computation lke matrix operations, Fourier
transforms, random number operations etc. The Scipy library
contains modules that can be used for signal processing,
optimization, statistics etc, while matplotlib provides access to
a powerful plotting package that can be used to produce high
quality 2-dimensional and 3-dimensional plots. Other libraries
in the wider ecosystem are Pandas, Scikit-Learn, Jupyter
notebooks etc. We would look at each of these package in
more depth in the next section.
40
Regression
Regression is a statistical modelling technique whereby we are
majorly interested in predicting the value of variable. The value
to be predicted is normally a real value, which 1s a positive or
negative number. This number may be a whole number in
which case it is referred to as an integer ot a number with
several decimal places in which case it is referred to as a floating
point number.

The nature of regression problems is that we are trying to find


how the value of a dependent variable changes with respect to
one or more independent variables. In a nutshell, what we want
to know is how much a_ variable say y depends on a set of
other variables say x, w such that we can learn to predict the
value of y once we know the values of the variables it depends
on.

Our task is therefore to model this relationship in such a way


that it would hold true for a majority of examples in our data.

The main intuition to get from this section 1s that regression


always produces a single value hence it is best applied to
learning problems where we requite a single real valued
number. A good example is if we want to build a model that
takes in information about a person such as their age,
nationality, profession etc and we want to predict their
expected income for a year. Our output would be a single value
and regression would be well positioned to solve this problem.

4]
Introduction to Labels and Features

The two main branches of machine learning you would easily


come actoss are supervised and unsupervised learning.
Supervised learning as the name implies deals with teaching
machine learning models with the help of data that is clearly
labelled, that is data that has been annotated by a human to
show what is to be learnt by the algorithm. The labels serve as
a “supervisor” to the algorithm, teaching it during training by
providing information on which samples it got correct or
wrong.

The labels are regarded as the ground truth, the actual outcome
that was observed from a particular data point.

Unsupervised learning on the other hand does not include any


labelled outcomes. ‘The job of the algorithm is to learn from
the raw data in order to come up with patterns which would
provide insights on the data explored. The main difference is
that there isn’t explicit feedback during training in the form of
labelled examples. A famous application of unsupervised
learning is in clustering. Clustering involves using a machine
learning algorithm to categorize data points into groups
(clusters) based on similarity of features.

Features

A feature is a characteristic of an observed data point in a


dataset. There may be more than one feature in a dataset and
it is fairly common to come across a large number of features.
Features are usually measurable and represent a specific axis of
explanation for the data. The quality of features selected has a
direct impact on quality of models as models learn using

42
features that ate informative in order to arrive at a final
prediction.

Features that best describe the data should always be chosen


as such features have a high discriminative tendency which
helps the machine learning model classify outputs and
predictions.
A good approach is to always choose an optimum number of
features, not too much as in such a case, many features would
be uninformative which leads to overfitting of the model and
not too little that the model underfits thereby failing to learn
anything.
The technique of choosing the right number of features is an
instance of hyperparameter tuning in that we try out several
options and settle for the best one.
When we have too many features and we are not sure which
features ate important and which ones are not, an algorithm
that can help use find the most relevant features which drive
the discriminability is called Principal Component Analysis
(PCA). PCA is a dimensionality reduction technique which
reduces the number of dimensions of our data to a small
number that best describes [Link] features.

A Regression Example: Predicting Boston Housing


Prices

To get a good understanding of the concepts discussed so far,


we would introduce a running example in which we are faced
with a regression problem, predict the price of houses in the

43
Boston suburbs given information about such houses in the
form of features.

The dataset we would use can be found at:


https: / ‘/[Link] [Link]/inde x. php /p/rdataset/source / file /m
aster/csv/MASS/[Link]

Steps To Carry Out Analysis

To carry out analysis and build a model we first need to identify


the problem, perform exploratory data analysis to get a better
sense of what is contained in our data, choose a machine
learning algorithm, train the model and finally evaluate its
performance. The following steps would be carried out in a
hands on manner below, so the reader 1s encouraged to follow
along.

Import Libraries:

import numpy as np

import pandas as pd

import [Link] as plt

Numpy is a popular Python numerical computing library


which suppotts vectorized implementation of calculations and
significantly speeds up computation time. Pandas ts a library
used to manipulate data in data frames while Matplotlib is used
for plotting graphs and data visualizations.

Let’s load the data by using the read_csv method on the Pandas
library and passing it the location of our data.

44
dataset = pd.read_csv('[Link]')

The next step is to look at what our data contains

[Link]
(5)

Unnamed:0 crim zn indus chas nox rm age dis rad tax ptratio black Istat medv

0 1 0.00632 180 231 0 0.538 6575 652 4.0900 1 296 3 396.90 498 240

1 2 0.02731 0.0 7.07 0 0.469 6421 78.9 4.9671 2 242 178 306.90 914 216

2 3 0.02729 00 7.07 0 0.469 7.185 611 4.9671 2 242 178 30283 403 347

3 4 0.03237 0.0 218 0 0458 6.998 458 6.0622 3 222 187 304.63 294 334

4 5 0.06905 0.0 218 O 0.458 7.147 54.2 60622 3 222 187 396.90 533 362

This shows individual observations as rows. A row represents


a single data point while columns represent features. There are
13 features because the last column medv 1s the regression
value that we ate to predict (median value of ownet-occupied
homes in $1000s) and Unmamed: 0 column is a sort of identifier
and is not informative. Each feature represents a subset of
information example crim means per capita crime rate by
town while rm average number of rooms pet dwelling.

Next we run

dataset. shape

This gives the shape of dataset which contains 506


observations. We first need to separate our columns into our
independent and dependent variables

45
[Link](['Unnamed: 0', 'medv'], axis=1)

dataset[ 'medv' ]

We would need to split our dataset into train and test splits as
we want to train our model on the train split, then evaluate its
performance on the test split.

from [Link] selection import


train test split

x train, x_test, y train, y test =


train test split(X, y, test_size=0.3)

x_train, x_test contains our features while y_train, y_test are


the prediction targets for the train split and test splits
respectively. test_size=0.3 means we want 70% of data to be
used for training and 30% for the testing phase.
The next step is to import a linear regression model from the
Scikit-Learn library. Scikit-Learn is the defacto machine
learning library in Python and contains out of the box many
machine learning models and utilities.
Linear regression uses equation of a straight line to fit our
patametets.

# importing the model


from [Link]
model import
LinearRegression

regressor = LinearRegression()

46
The above code imports the linear regression model and
instantiates an object from it.

[Link](x_train,y
train)

This line of code fits the data using the fit method. What that
means is that it finds appropriate values for the independent
parameters that explains the data.

How to forecast and Predict

To evaluate our model, we use the test set to know whether


our model can generalize well to data it wasn’t trained on.

y_pred = [Link]
(x test,y test)

from [Link] import mean_squared error

mse = mean squared error(y test, y pred)

The predict method called on the regressor object returns


predictions which we use to evaluate the error of our model.
We use mean squared error as our metric. Mean Squared Error
(MSE) measures how far off our predictions ate from the real
(actual) values. The model obtains an MSE of 20.584.

Finally, we plot a graph of our output to get an idea of the


distribution.

[Link](y test, y pred)

47
[Link]("Prices: $Y_i$")

[Link] ("Predicted prices: e\hat{Y} 25")

[Link]("Prices vs Predicted prices: $Y_i$ vs


S\hat{Y} is")

Prices vs Predicted prices: Yj vs Y,

‘Predicted
prices:

Prices:Y;

We can see from the scatter plot above that predictions from
our model are close to the actual house prices hence the
concentration of points.

48
Classification

In machine learning most learning problems can be modelled


as a classification problem. A classification problem is one
whose core objective is to learn a mapping function from a set
of inputs to one or more discrete classes. Discrete classes are
sometimes teferred to as labels and both terms are often used
interchangeably.

A class ot label can be understood as a category that represents


a particular quantity, therefore what classification algorithms
do is to identify the category that an example fits into. If the
classification problem 1s posed 1n such a way that there are two
distinct classes, we have a binary classification problem. In a
case where we have more than two classes (labels), the learning
problem is referred to as multi-class classification indicating
that observations could fall into any of the n classes. The final
type of classification is where a sample may belong to several
categories that is it has more than one label and in such a
situation we would be dealing with a multi-label classification
task.

To get a better mental picture of classification let's look at the


image below:
Cc

49
x2 2
eaoe Surviving
2 companies
s o
& Fe Fs) x
& o z
re) x *
eo 9 86 o b
e o a x
x
_ « so «6 a : es
s
20 é ‘ x
¥ x
& o ‘ ‘ . -
2 a. = Ez * x
x x x

Failing ca x = x
companies x x
®

From the plot above we can see that there are two features that
describe the data X;and X2. What a classification task seeks to
do is divide the data into distinct categories such that there is a
decision boundary that best separates classes. In this example
we have two classes - falling companies and surviving
companies, a data point which represents a company can only
belong to one of those categories, falling or surviving. It is as a
result clear that this is a binary classification example because
there are only two classes.

Another point to note from the diagram is that the classes are
linearly separable, that is they can be separated by a straight
line. In other problems, this might not be possible and there
are more robust machine learning algorithms that handle such
instances.

Multi-Class Classification

50
It is important that we have a good understanding of
classification based on the number of classes that we want to
predict as classification has many real world applications. ‘To
further improve out intuition let us analyse the image below:

The data is projected onto a two dimensional plane to enable


visualization. There are three classes represented by red
squares, blue triangles and green circles. There are also three
decision boundaries that separates the data points into three
sections with the color of the class projected on the
background. What we have 1s a classic multi-class classification
example with three classes (0, 1, 2), there are also some
misclassified points however these are few and appear mostly
close to the decision boundaries.

To evaluate this model, we would use accuracy as our


evaluation metric. The accuracy of our model is determined by
the number of samples our classifier predicted corrected to the
number of samples it misclassified. Accuracy 1s usually a good

OM
metric for classification tasks but bear in mind that there are
other metrics such as precision and recall that we may wish to
explore based on how we intend to model our learning task.

Popular Classification Algorithms

Some machine learning models deliver excellent results with


classification tasks, in the next section we would have an
indepth look at a couple of them. The process involved in
running classification tasks are fairly standard across models.

First, we need to have a dataset that has a set of inputs and a


corresponding set of labels. The reason this is important is
because classification is a supervised learning task where we
need access to the ground truth (actual classes) of observations
in order to minimize the error in misclassification from our
models. So to end up with a good model, we train the classifier
on samples and provide the true values in a feedback
mechanism forcing the classifier to learn and reduce its error
tate on each iteration or pass through the training set (epoch).
Examples of models used for classification are logistic
regression, decision trees, random forests, k-nearest neighbor,
support vector machine etc. We use the last two models as case
studies to practice classification.

Introduction to K Nearest Neighbors

To understand the k-nearest neighbor algorithm, we first need


to understand nearest neighbor. Nearest neighbor algorithm is
an algorithm that can be used for regression and classification
tasks but is usually used for classification because it is simple
and intuitive.
At training time, the nearest neighbor algorithm simply
memorizes all values of data for inputs and outputs. During
test time when a data point is supplied and a class label 1s
desired, it searches through its memory for any data point that
has features which are most similar to the test data point, then
it returns the label of the related data point as its prediction. A
Nearest neighbor classifier has very quick training time as it is
just storing all samples. At test time however, its speed 1s
slower because it needs to search through all stored examples
for the closest match. The time spent to receive a classification
prediction increases as the dataset increases.

The k-nearest neighbor algorithm 1s a imedificanon of the


nearest neighbor algorithm in which a class label for an input
is voted on by the k closest examples to it. That is the predicted
label would be the label with the majority vote from the
delegates close to it. So a k value of 5 means, get the five most
similar examples to an input that is to be classified and choose
the class label based on the majority class label of the five
examples.

Let us now look at an example image to hone our knowledge:

D2
a ne
Training instance | Se
a | pees Class 1}
a ia a — i

Distance N\ Ke} |
Class 2)
ata. - iL.

|
\
®
New example
to classify Zo

The new example to be classified 1s placed in the vector space,


when k = 1, the label of the closest example to it is chosen as
its label. In this case, the new example is categorized as
belonging to class 1. When k = 1, k-nearest neighbor algorithm
is reduced to nearest neighbor algorithm.

From the image, when k = 3, we choose the 3 closest examples


to the new example using a similarity metric known as the
distance measure. We see that two close examples predict the
class as being class 2 (red triangle) while the remaining example
predicts the class to be class 1 (blue square). The predicted class
of the new data point 1s therefore class 2 because it has the
majority vote.
The distance metric used to measure proximity of examples
may be L1 or L2 distance. L1 distance is the sum of the
absolute of the difference between two points and 1s given by:

54
d,(P.q) =» IP,—q,l
|

L1 distance is sometimes referred to as the Manhattan distance.

L2 is an alternative distance measure that may be used. It is the


Euclidean distance between two points using a straight line. It
is given as:

[Spa
dj([Link]= I> (9g; —P,))?
\ i=]

A value of k = 1 would classify all training examples correctly


since the most similar example to a point would be itself. This
would be a sub-optimal approach as the classifier would fail to
learn anything and would have no power to generalize to data
that it was not trained on. Abetter solution is to choose a value
of k in a way that it performs well on the validation set. The
validation set is normally used to tune the hyperparameter k.
Higher values of k has a smoothing effect on the decision
boundaries because outlier classes are swallowed up by the
voting pattern of the majority. Increasing the value of k usually
leads to greater accuracy initially before the value becomes too
large and we reach the point of diminishing returns where
accuracy drops and validation error starts rising.
—— Validation error

.\
|
| 5074

39
\\ ——

bse, ar a
\

|
| \ -
\ ———

| |
hel geen au
| |
}
|
-— =

0 10 20 30 40 50 60
K- Value

The optimal value for k is the point where the validation error
is lowest.

How to create and test the K Nearest Neighbor


classifier

We would now apply what we have learnt so far to a binary


classification problem. The dataset we would use is the Pima
Indian Diabetes Database which is a dataset from the National
Institute of Diabetes and Digestive and Kidney Diseases. The
main purpose of this dataset 1s to predict whether a patient has
diabetes or not based on diagnostic measurements carried out
on patients. The patients in this study were female, of Pima
Indian origin and at least 21 years old.
The dataset can be found at:

// [Link]/uciml/pima-indians-diabetes-
database/data

56
Since we ate dealing with two mutually exclusive classes, a
patient either has diabetes or not, this can be modelled as a
binary classification task and for the purpose of our example
we would use the k-nearest neighbor classifier for
classification.

The first step is to import the libraries that we would use.

import numpy as np

import pandas as pd

import [Link] as plt

Next we load the data using Pandas.

dataset = pd.read_csv('[Link]')

As always what we should do 1s get a feel of our dataset and


the features that are available.

[Link]
(5)

Pregnancies Glucose BloodPressure SkinThickness Insulin BMI DiabetesPedigreeFunction Age Outcome

0 6 148 72 35 0 33.6 0.627 50 1

1 1 85 66 29 0 266 0.351 31 0

2 8 183 64 0 0 233 0.672 32 1

3 1 89 66 23 94 28.1 0.167 21 0

4 0 137 40 3168 43.1 2.288 33 si

We see that we have 8 features and 9 columns with Outcome


being the binary label that we want to predict.

a7
‘To know the number of observations in the dataset we run

dataset. shape

This shows dataset contains 768 observations.

Let’s now get a summary of the data so that we can have an


idea of the distribution of attributes.

dataset. describe()

Pregnancies Glucose BloodPressure SkinThickness Insulin BMI DiabetesPedigreeFunction Age Outcome

count —768.000000 768.000000 768,000000 768,000000 768,000000 768,000000 768,000000 768,000000 768.000000

mean 3.845052 120,894531 69.105469 20.536458 79.799479 31.992578 0.471876 33,240885 0.348958

std 3.369578 31.972618 19.355807 15,952218 115.244002 7.884160 0.331329 14.760232 0.476951

min 0.000000 0,000000 0,000000 0.000000 0.000000 0.000000 0.078000 21,000000 0,000000

25% 1.000000 99.000000 62,000000 0,000000 0.000000 27.300000 0.243750 24.000000 0,000000

50% 3.000000 117.000000 72,000000 23,000000 30.500000 32000000 0.372500 29.000000 0.000000

75% 6,000000 140.250000 80,000000 32,000000 127.250000 36,800000 0.626250 41,000000 1.000000

max —17,000000 199.000000 122.000000 99,000000 846,000000 67.100000 2.420000 81.000000 1.000000

The count row shows a constant value of 768.0 across features,


it would be remembered that this is the same number of rows
in our dataset. It signifies that we do not have any missing
values for any features. The quantities mean, std gives the mean
and standard deviation respectively across attributes in our
dataset. The mean is the average value of that feature while the
standard deviation measures the variation in the spread of
values.

58
Before going ahead with classification, we check for
correlation amongst our features so that we do not have any
redundant features

corr = [Link]() # data frame


correlation function

fig, ax = [Link](figsize=(13, 13))

[Link] (corr) # color code the rectangles


by correlation value

[Link] (range(len([Link])),
[Link]) # draw x tick marks

[Link] (range(len([Link])),
[Link]) # draw y tick marks

Pregnancies Glucose BloodPressure iSkinThickness Insulin BMI DiadetesPedigneeFunction Age Outcome

Pregnancies

Glucose

BloodPressur €

SkinTivekness

Insulin

DiabetesPedigreeFunction

Outcome

oS
The plot does not indicate any 1 to 1 correlation between
features, so all features are informative and _ provide
discriminability.

We need to separate our columns into featutes and labels

features = [Link](['Outcome'], axis=1)

labels = dataset['Outcome']

We would once again split our dataset into training set and test
set as we want to train our model on the train split, then
evaluate its performance on the test split.

from [Link] selection import


train test split

features train, features test, labels train,


labels test = train test split(features, labels,
test_size=0.25)

features_train, features_test contains the attributes while


labels_train, labels_test are the discrete class labels for the
train split and test splits respectively. We use a test_size of 0.25
which indicates we want to use 75% of observations for
training and reserve the remaining 25% for testing.

The next step is to use the k-nearest neighbor classifier from


Scikit-Learn machine learning library.

# importing the model


from [Link] import
KNeighborsClassifier

60
classifier = KNeighborsClassifier()

The above code imports the k-nearest neighbor classifier and


instantiates an object from it.

[Link](features train, labels train)

We fit the classifier using the features and labels from the
training set. To get predictions from the trained model we use
the predict method on the classifier, passing in features from
the test set.

pred = [Link]édict (features test)

In order to access the performance of the model we use


accuracy as a metric. Scikit-Learn contains a utility to enable us
easily compute the accuracy of a trained model. To use it we
import accuracy_score from metrics module.

from [Link] import accuracy score

accuracy = accuracy score(labels test, pred)

print('Accuracy: {}'.format (accuracy) )

We obtain an accuracy of 0.74, which means the predicted label


was the same as the true label for 74% of examples.

6]
Hete is the code in full:

# import libraries
import numpy as np

import pandas as pd

import [Link] as plt

# read dataset from csv file

dataset = pd.read_csv('[Link]')

# display first five observations

[Link]
(5)

# get shape of dataset, number of observations,


number of features

dataset. shape

# get information on data distribution


[Link]()

# plot correlation between features

corr = [Link]() # data frame


correlation function

fig, ax = [Link](figsize=(13, 13))

[Link] (corr) # color code the rectangles


by correlation value

[Link]
(range (len([Link])),
[Link]) # draw x tick marks

[Link] (range (len([Link])),


[Link]) # draw y tick marks
# create features and labels

features = [Link](['Outcome'], axis=1)

labels = dataset['Outcome']

# split dataset into training set and test set

from [Link] selection import


train_test split

features train, features test, labels train,


labels test = train_test_split(features, labels,
test _size=0.25)

# import nearest neighbor classifier

from [Link] import


KNeighborsClassifier

classifier = KNeighborsClassifier()

# fit, data

[Link](features
train, labels train)

# get predicted class labels


pred = [Link] (features test)

# get accuracy of model on test set

from [Link] import accuracy score

accuracy = accuracy score(labels test, pred)

print ('Accuracy: {}'.format (accuracy) )


Introduction to Support Vector Machine
Support vector machines also known as support vector
networks is a popular machine learning algorithm used for
classification. The main intuition behind support vector
machines 1s that it tries to locate an optimal hyperplane which
separates data into correct classes making use of only those
data points close to the hyperplane. The data points closest to
the hyperplane are called support vectors.
There may be several hyperplanes that correctly separates
classes but support vector machine algorithm chooses the
hyperplane that has the largest distance (margin) from the
support vectors (data points close to the hyperplane). The
benefit of selecting a hyperplane with the widest margin 1s
because this reduces the chance of mistakenly misclassifying a
data point during test time.

64
Support Vectors

ene
el =
Margin
Width

x,

Support vector machine algorithm ts robust to the presence of


outliers in class distributions and works in a way that ignores
outhers while finding the best hyperplane with a good margin
width.

However, it is not in all cases that a data distribution may be


linearly separable. Support vector machine (SVM) transforms
the features that represents the data from a lower dimension
into a higher dimension by introducing new features which are
based on the original features found in the data. It does this
automatically using a kernel. Some examples of kernels in SVM
are linear kernel, polynomial kernel, Radial Basis Function
(RBF) kernel etc.

Support vector machine has some advantages compared to


other machine learning models. It is particularly potent in high
dimensional spaces and because it only relies on a subset of
data points for classification (the support vectors), it consumes
less memory resulting in efficiency. It is also effective because
65
we can pick different kernels based on the data representation
at hand.

How to create and test the Support Vector Machine


(SVM) classifier

For this section we would use a support vector machine


classifier on the Pima Indian Diabetes Database and compate
its results with the k-nearest neighbor classifier.

Hete is the full code:

# import libraries
import numpy as np

import pandas as pd

import [Link] as plt

# read dataset from csv file

dataset = [Link] _csv('[Link]')

# create features and labels

features = [Link](['Outcome'], axis=1)

labels = dataset['Outcome'"]

# split dataset into training set and test set

from [Link] selection import


train test split

features train, features test, labels train,


labels test = train test _split(features, labels,
test_size=0.25)

66
# import support vector machine classifier

from [Link] import SVC

classifier = SVC()

# fit data
[Link](features train, labels train)

# get predicted class labels


pred = [Link](features
test)

# get accuracy of model on test set

from [Link] import accuracy score

accuracy = accuracy score(labels test, pred)

print ('Accuracy: {}'.format (accuracy) )

We get an accuracy of 0.66 which is worse than 0.74 which we


got for k-nearest neighbor classifier. There are many
hyperparameters that we could try such as changing the type
of kernel used.

from [Link] import SVC

classifier = SVC(kernel='linear')

When we use a linear kernel our accuracy jumps to 0.76. This


is an important lesson in machine learning as often times we
do not know beforehand what the best hyperparametets are.
So we need to experiment with several values before we can
settle on the best performing hyperparameters.

67
sate So ori

eet<a < n
IA te sn |
i a= are 2 em ie,
wool s Ae «4 aii

Cee 1 dary |ore ey Shy? 1H) RS we


fides» - p* Ve ra @ ; T

Shh NS Jae | :
i ane .
ating Dine 6% Kral : a ak “i
ee py ore ge (Ln) et ‘ ;
PI neue emt?dy sh at
SO et eee ae
a teen eis)iteoa) rue of ie 2 rep ae" »

teary —S

eee all me at 7 ; =) 6s _

Nae loa ie WT ie
Bialik
@8 y= .
| nnlindpgenes. G4 Camm
rare tewn) eohiclau@ & & m4
ai - Qe ® : °

Soeneiahe ara to OQ oR enw Gis 4 5


: Si Myod= 2ebing More

aia sth
igs abaea wil wil) ce atyens
591662) &®
aaphlh
dw

ine
7

i A Rete,
Bie Ot Gp hel secs a) 106 PDWer
Fed yetites ‘ae os
EPOiO=
‘ ==

a y
———
as
f= » :
Clustering
Clustering is the most common form of unsupervised learning.
Clustering involves grouping objects or entities into clusters
(groups) based on a similarity metric. What clustering
algorithms aim to achieve is to make all members of a group
as similar as possible but make the cluster dissimilar to other
clusters. At first glance clustering looks a lot like classification
since we ate putting data points into categories, while that may
be the case, the main difference is that in clustering we are
creating categories without the help of a human teacher.
Whereas, in classification, objects were assigned to categories
based on the domain knowledge of a human expert. That is in
classification we had human labelled examples which means
the labels acted as a supervisor teaching the algorithm how to
recognise various categories.

In clustering, the clusters or groups that are discovered are


purely dependent on the data itself. The data distribution is
what drives the kind of clusters that are found by the algorithm.
There are no labels so clustering algorithms are forced to learn
representations 1n an unsupervised manner devoid of direct
human intervention.

Clustering algorithms are divided into two main groups - hard


clustering algorithms and soft clustering algorithms. Hard
clustering algorithms are those clustering algorithms that find
clusters from data such that a data point can only belong to
one cluster and no more. Soft clustering algorithms employ a
technique whereby a data point may belong to more than one

69
cluster, that is the data point is represented across the
distribution of clusters using a probability estimate that assigns
how likely the point belongs to one cluster or the other.

0 peatercnnerinnynininiomnnerenion larger clustersEee


:
]|
D bt @ 8 |

4} " ae 7
| % v %. |
j Ya g ®
©7ob &3 |

2| c ¢ ® Way °
| ie), i
; CE Ty |
j

@eco ®
My ee e, @ 6
4 wo @ % 6 % @ & |
‘ £ & ne oe
| ae) we 3 25, “he 2 |
r- ee
-4|i Fe £ @ se
Che

Rael PERItie ee ah Mere


|

~§ 4 =~? 0 2 4 6

From the data distribution of the image above, we can deduce


that a clustering algorithm has been able to find 5 clusters using
a distance measure such as Euclidean distance. It would be
observed that data points close to cluster boundaries are
equally likely to fall into any neighboring cluster. Some
clustering algorithms are deterministic meaning that they
always produce the same set of clusters regardless of
initialization conditions or how many times they are run. Other
clustering algorithms produce a different cluster collection
everytime they are run and as such it may not be easy to
reproduce results.

70
Introduction to Clustering

The most important input to a clustering algorithm is the


distance measure. This is so because it is used to determine
how similat two ot more points are to each other. It forms the
basis of all clustering algorithms since clustering 1s inherently
about discriminating entities based on similarity.

Another way clustering algorithms ate categorized is using the


relationship structure between clusters. There are two
subetoups - flat clustering and hierarchical clustering
algorithms. In flat clustering, the clusters do not share any
explicit structure so there is no definite way of relating one
cluster to the other. A very popular implementation of a flat
clustering algorithm is K-means algorithm which we would use
as a case study.

Hierarchical clustering algorithms first starts with each data


point belonging to its own cluster, then similar data points are
merged into a bigger cluster and the process continues until all
data points are part of one big cluster. As a result of the process
of finding clusters, there is a clear hierarchical relationship
between discovered clusters.

There ate advantages and disadvantages to the flat and


hierarchical approach. Hierarchical algorithms are usually
deterministic and do not require us to supply the number of
clusters beforehand. However, this leads to computational
inefficiency as we suffer from quadratic cost. The time taken
to discover clusters by an hierarchical clustering algorithm
increases as the size of the data increases.

71
Flat clustering algorithms are intuitive to understand and
feature linear complexity, therefore the time taken to run the
algorithm increases linearly with the number of data points and
because of this flat clustering algorithms scale well to massive
amounts of data. As a rule of thumb, flat clustering algorithms
are generally used for large datasets where a distance metric can
capture similarity while hierarchical algorithms are used for
smaller datasets.

Example of Clustering

We would have a detailed look at an example of a flat clustering


algorithm, K-means. We would also use it on a dataset to see
its performance.
K-means ts an iterative clustering algorithm that seeks to assign
data points to clusters. To run K-means algorithm, we first
need to supply the number of clusters we desire to find. Next,
the algorithm randomly assigns each point to a cluster and
computes the cluster centroids (center of cluster). At this stage
points are reassigned to new clusters based on how close they
are to cluster centroids. We again recompute the cluster
centroids. Finally we repeat the last two steps until no data
points are being reassigned to new clusters. The algorithm has
now converged and we have our final clusters.

Running K-means with Scikit-Learn

For our hands on example we would use K-means algorithm


to find clusters in the Iris dataset. The Iris dataset is a classic in
the machine learning community. It contains 4 attributes (sepal
length, sepal width, petal length, petal width) used to describe
3 species of the Iris plant. The dataset can be found at:

(e)
[Link] /itiscsv downloads
[Link]/1

The first step is to load the data and run the head method on
the dataset to know our features

# import libraries
import numpy as np

import pandas as pd

import [Link] as plt

# read dataset from csv file

dataset = pd.read_csv('[Link]')

# display first five observations

dataset.
head (5)

Id SepalLengthCm SepalWidthCm PetalLengthCm PetalWidthCm Species

G 1 wt a5 14 0.2 Iris-setosa

ee 49 3.0 14 0.2 iris-setosa

ya 47 a2) 3 1.3 0.2 iris-setosa

3 4 46 3.1 1d 0.2 Iris-setosa

45 5.0 3.6 14 0.2 Iris-setosa

We have 4 informative features and 6 columns. Id is an


identifier while Species column contains the label. Since this is
a clustering task, we do not need labels as we would find
clusters in an unsupervised mannetr.

73
xe [Link](['Id', 'Species'], axis=1)

x = [Link] # select values and convert


dataframe to numpy array

The above line of code selects all our features into x dropping
Id and Species.

As was earlier discussed, because K-means 1s a flat clustering


algorithm we need to specify the value of k (number of
clusters) before we run the algorithm. However, we do not
know the optimal value for k, so we use a technique known as
the elbow method. The elbow method plots the percentage of
variance explained as a result of number of clusters. The
optimal value of k from the graph would be the point where
the sum of squared error (SSE) does not improve significantly
with increase in the number of clusters.

Let’s take a at look at these concepts in action

# finding the optimum number of clusters for k-


means classification

from [Link] import KMeans

wess = [] # array to hold sum of squared


distances within clusters

Lormiein erange (i:

kmeans = KMeans(n_ clusters = i, init = 'k-


meanst++', max iter = 300, n_init = 10,
random state = 0)

kmeans.
fit (x)

[Link] (kmeans.inertia_)

74
# plotting the results onto a line graph,
allowing us to observe 'The elbow'

[Link](range(1, 11), wess)

[Link]('The elbow method')

[Link]('Number of clusters')

[Link]('Within Cluster Sum of Squares') #


within cluster sum of squares

plt. show ()

The elbow method

§ Sa
SB

Squares
of
Sum
Cluster
Within

Number of clusters

The k-means algorithm is run for 10 iterations, with n_clusters


ranging from 1 to 10. At each iteration the sum of squared
error (SSE) is recoded. The sum of squared distances within
each cluster configuration is then plotted against the number
of clusters. The “elbow” from the graph is 3 and this is the
optimal value for k.

Ip
Now that we know that the optimal value for k is 3, we create
a K-means object using Scikit-Learn and set the parameter of
n_clusters (number of clusters to generate) to 3.

# creating the kmeans object


kmeans = KMeans(n clusters = 3, init = 'k-
means++', max iter = 300, n_init = 10,
random_state = 0)

Next we use the fit_predict method on our object. This returns


a computation of cluster centers and cluster predictions for
each sample.

y_kmeans = [Link] predict (x)

We then plot the predictions for clusters using a scatter plot of


the first two features.

# visualising the clusters

[Link](x[y_kmeans == 0, 0], xf[y_kmeans ==


0, 1), s = 100, c = 'red', label = 'Iris-
setosa')

[Link](x[y
kmeans == 1, 0], x[y_kmeans ==
ie, so= L007 cl= 'tblue!l, label — Virus
versicolour')

[Link](x[y_ kmeans == 2, 0], x[y_kmeans ==


2 |e st= 1007 c= V"ogreen’, label) =" "iznis=
virginica')

# plotting the centroids of the clusters

76
[Link](kmeans.cluster_centers [:, Oy
[Link] ster
centers [:,1], s = 100, c=
'yellow', label = 'Centroids')

[Link]()

@ Iris-setosa
o ¢ @ lris-versicolour
@ Iris-virginica
Centroids —

45 5.0 5S 6.0 65 70 1.5 6.0

The plot shows 3 clusters - red, blue, green representing types


of Iris plant, setosa, versicolour and virginica respectively. The
yellow point indicates the centroids which is at the center of
each cluster.

Our K-means algorithm was able to find the correct number


of clusters which is 3 because we used the elbow method. It
would be observed that the original dataset had three types
(classes) of Iris plant. Iris setosa, Itis versicolour and Iris
virginica. If this were posed as a classification problem we
would have had 3 classes into which we would have classified
data points. However, because it was posed as a clustering
problem, we were still able to find the optimum number of

al
clusters - 3, which is equal to the number of classes in our
dataset.

What this teaches us is that most classification problems and


datasets can be used for unsupervised learning particularly for
clustering tasks. The main intuition to take out of this is that if
we want to use a classification dataset for clustering, we must
remove labels, that 1s we remove the component of the data
that was annotated by a human to enable supervision. We then
train on the raw dataset to discover inherent patterns contained
in the data distribution.

While we have only touched on a portion of unsupervised


learning, it is important to note that it is a vital branch of
machine learning with lots of real world applications.
Clustering as an example can be used to discover data groups
and get a unique perspective of data before feeding it into
traditional supervised learning algorithms.

78
agi <haph “0
. ae a ae © 1) Coke
om eles os of T : yee oe [Link]>
nits Meno Od ais peta --6
els Thalia at « avon eel
7
oo ee ee
Aapeall jr@peeeses OY ;
yt
-
wit i ONS ae
Ky ee af A
©
& : :
®
- : oe @ i

; - i ae SIG altaT 7

aq , a on -s

» 1 & @aiprieds Gate


or ae, regaeets
7
Sas Hoid Su ow ee Soe = aire “* Wie
a |
Mt silt Lolpeneeiahiied <
‘ a
Introduction to Deep Learning using
TensorFlow
Deep learning 1s a subfield of machine learning that makes use
of algorithms which carry out feature learning to automatically
discover feature representations in data. The “deep” in deep
learning refers to the fact that this kind of learning is done
across several layers, with higher level feature representations
composed of simpler features detected earlier in the chain. The
primary algorithm used in deep learning is a deep neural
network composed of multiple layers. Deep learning is also
known as hierarchical learning because it learns a hierarchy of
features. The complexity of learned features increases as we
move deeper in the network.

Deep learning techniques are loosely inspired by what 1s


currently known about how the brain functions. The main idea
is that the brain is composed of billions of neurons which
interact with each other in some way through the use of
electrical signals from chemical reactions. This interaction
between neurons in conjunction with other organs helps
humans to perform cognitive tasks such as seeing, hearing,
taking decisions etc. What deep learning algorithms do 1s to
arrange a network of neurons in a structure known as an
Artificial Neural Network (ANN) to learn a mapping function
from inputs directly to outputs. The difference between deep
learning and a classical artificial neural network ts that the layers
of the network are several orders of magnitude larger hence it
is commonly called a Deep Neural Network (DNN) of
feedforward neural network.

80
To get a well-grounded understanding, it is important for us to
take a step back and try to understand the concept of a single
neuron.

oy ineNeuron
dendrites
pes i } ¢—
\b=)
Ye a S24 synapses
nucleus —8 aeons ae

First let us develop simple intuitions about a biological neuron.


The image of a biological neuron above shows a single neuron
made up of different parts. The brain consists of billions of
similar neurons connected together to form a network. The
dendrites are the components that carry information signals
from other neurons earlier in the network into a particular
neuron. It is helpful to think of this in the context of machine
learning as features that have so far been learned by other
neurons about our data. The cell body which contains the
nucleus is where calculations that would determine whether we
have identified the presence of a characteristic we are
interested in detecting would take place. Generally, if a neuron
is excited by its chemical composition as a result of inflowing
information, it can decide to send a notification to a connected
neuron in the form of an electrical signal. This electrical signal
is sent through the axon. For our artificial use case, we can

81
think of an artificial neuron firing a signal only when some
condition has been met by its internal calculations. Finally, this
network of neurons learn representations in such a way that
connections between them are either strengthened or
weakened depending on the current task at hand. The
connections between biological neurons are called synapses
and we would see an analogy of synapses in artificial neural
networks known as weights which ate parameters we would
train to undertake a learning problem.

Xi

Threshold
Summer unit

eh OUTUE

w, W, W. w.- Weights of Connection


x,X,x,X,- Inputs —_b - Bias

Let us now translate what we know so far into an artificial


neuron implemented as a computation unit. From the diagram,
we can envisage Xi, Xz, to X, as the features that are passed
into the neuron. A feature represents a dimension of data from
a data point. The combination of features completely describe
that data point as captured by the data. W; to Wz, are the
weights and their job is to tell use how highly we should rank

82
a feature. That is lower values for the weight means that the
connected feature is not as important and higher values signify
greater significance. All inputs to the artificial neural network
are then summed linearly (added side by side). It is at this point
that we determine whether to a send signal to the next neuron
or not using a condition as a threshold. If the result of the
linear calculation is greater than or equal to the threshold value,
we send a signal else we don’t.

From the explanation, it is now plain to see why these


techniques are loosely based on the operation of a biological
neuron. However, it must be noted that deep learning beyond
this pomt does not depend on neuroscience as a complete
understanding of the way the brain functions is not known.

In deep neural networks, the activation criterion (whether we


decide to fite a signal or not) 1s usually replaced with non-linear
activation functions such as sigmoid, tanh or Rectified Linear
Unit (ReLU). The reason a non-linear activation function is
used is so that the artificial neural network can break linearity
in order to enable it learn more complicated representations. If
there were no non-linear activation functions, regardless of the
depth of the network (number of layers), what it would learn
is still a linear function which is severely limiting.

Let us now look at how we can atrange these neurons into an


artificial neural network using the image below to explain the
concepts.

83
Input Layer Hidden Layer Output Layer

Variable - #1 ©
\ : | i “

Variable - #2 ss A
Po Pasar
fo
ony
oN
we, JK — } Output
Gal bad
Variable - #3
we

ay, : ey

Variable - #4 ‘i
ee
An example of a Feed-forward Neural Network with one hidden layer (with 3 neurons )

In the example, Variable 1 through 4 represents an input to the


network in the form of features. An input may be an image,
text or sequential data. The input layer which contains the
variables (features) are connected to neurons in the hidden
layer. The hidden layer is so called because we cannot directly
inspect what goes on inside unlike the input or output layer
which we have direct access to. There are 3 units or neurons in
the hidden layer. ‘The last layer is the output layer. This is where
we get our predictions. In this case it is made up of a single
neuron which suggests that it renders a single output. Such a
neuron may be used to get a single real valued number in a
regression problem or a class in a binary classification problem.
It would be observed that neurons in one layer are connected
to all neurons in the next layer, because of this kind of
configuration, we say that it 1s a fully connected layer. The
strength of connections between neurons are governed by the

84
weights. Finally, the above network is said to be a 2-layer neural
network as the input layer is not counted when describing the
number of layers contained in a network.

There are various architectures of deep neural networks, some


architectures make assumptions about the type of data that the
model expects. Those assumptions usually improve the
performance of such models at the tasks which they were
designed for. Convolutional Neural Networks (CNN) are
designed with the assumption that they would be dealing with
2-dimensional data such as images and ate the de-facto
standard for computer vision tasks like image detection, image
classification, image segmentation, object recognition etc.
Recurrent Neural Networks (RNN) on the other hand are
specially suited for handling sequential and temporal data like
text and time series. Each of these architectures have further
variants that are modifications to suit a particular learning
problem.

Deep Learning Compared to Other Machine


Learning Approaches

The popularity of deep learning is mainly based on the fact that


it has achieved impressive results in diverse fields and is
b eginning to surpass humans in certain tasks such as object
recognition. Compared to other machine learning techniques,
it has been able to achieve state of the art results on many
benchmarks and 1s applicable with slight modifications to a
wide range of learning problems. The reason is that not only
have artificial neural networks been shown to be capable of
learning any function as illustrated by the universal

85
approximation theorem, deep neural networks seem to
improve their performance with more data. They are free from
the plateau effect that traditional machine learning algorithms
suffer from, whereby at some point the performance of the
algorithm does not improve with the availability of additional
data.

Deep learning
®
)&
©
=
Se
fe)
=
a
oO

Amount of data

Deep learning algorithms are seen as being data intensive


because they need enormous amounts of data to achieve high
accuracies and more data always appear to help performance.
Deep learning is quickly becoming the go to solution for many
machine learning problems where vast data 1s available
occasioned by the advent of the internet.

Applications of Deep Learning

86
Deep learning has been applied to solve many problems which
have real world applications and are now being transitioned
into commercial products. In the field of computer vision,
deep learning techniques are used for automatic colorization to
transform old black and white photos, automatic tagging of
friends in photos as seen in social networks and grouping of
photos based on content into folders.

In Natural Language Processing (NLP), these algorithms are


used for speech recognition in digital assistants, smart home
speakers etc. With advances in Natural Language
Understanding (NLU), chatbots are being’ deployed as
customer service agents and machine translation has enabled
real time translations from one language to the other.

Another prominent areas recommender systems, where users


are offered personalized suggestions based on their
preferences and previous spending habits. Simply put, deep
learning algorithms are wildly beneficial and learning them is a
quality investment of time and resoutces.

Python Deep Learning Frameworks

Python has a mature ecosystem with several production ready


deep learning frameworks. Some of them are Pytorch, Chainer,
MXNet Keras, TensorFlow etc. We would concentrate on
using Tensorllow in this book as it is an extremely popular
and supported deep learning framework open sourced by
Google in 2015. TensorFlow uses the concept of a
computation graph to construct a model. Nodes and
operations ate declared on the graph beforehand then at
training time the model is compiled. In the next section we

87
would see how to install TensorFlow and use it to perform
deep learning tasks through a hands on example.

Install TensorFlow

TensorFlow is cross platform and can be installed on various


operating systems. In this section we would see how it can be
installed on the three widely used operating systems - Linux,
macOS and Windows. TensorFlow can be installed using pip
which is the native Python package manager, vittualenyv -
which creates a virtual environment or through a bundled
scientific computation distribution like Anaconda. For the
purpose of this book we would use pip.

Installing TensorFlow on Windows

For TensorFlow to be installed on Windows, you need to first


install Python which comes bundled with pip as the package
manager. Python can be downloaded as an executable at
https: //[Link]/downloads/windows

Once you have it installed, start a terminal and run the


following command:

C:\> pip3 install --upgrade tensorflow

Installing TensorFlow on Linux

For Linux distributions like Ubuntu and its variants, pip 1s


usually already installed. To check which version of pip 1s
installed, from the terminal run

88
$ pip -V

or

$ pip3 -V

This depends on the version of Python you have, pip for


version 2.7 and pip3 for version 3.x

If you do not have pip installed, run the appropriate command


for yyour Python
J version below:

$ sudo apt-get install python-pip python-dev # for Python 2.7


$ sudo apt-get install python3-pip python3-dev # for Python 3.n

It is recommended that your version of pip or pip3 is 8.1 or


greater. Now you can install TensorFlow with the following
command

$ pip install tensorflow # Python 2.7


$ pip3 install tensorflow # Python 3.x

Installing TensorFlow on macOS

To install TensorHlow on macOS, you need to have Python


installed as it is a prerequisite. To check if you have pip
installed run

$ pip -V

ot

$ pip3 -V

89
This depends on the version of Python you have, pip for
version 2.7 and p1p3 for version 3.x

If you do not have pip installed, or you have a version lower


than 8.1, run the commands to install or upgrade:

$ sudo easy_install --upgrade pip

$ sudo easy_install --upgrade six

Now you can install TensorFlow with the following command

$ pip install tensorflow —# Python 2.7

$ pip3 install tensorflow # Python 3.x

How to Create a Neural Network Model

In order to create a neural network in TensorFlow it is


important to understand how TensorFlow works. The
following steps would help us understand what we need to do:

Build a Computation Graph: First we need to define


mathematical operations which would be carried out
on tensors. We can think of tensors as a higher order
matrix. A matrix is an array containing numbers
arranged in rows and columns. Data structures such as
images can be represented as matrices and since deep
learning uses geometric transformations, the input to
any model must be numeric.
I Initialisation of Variables: All previously declared
variables have to be initialised before we can train the
model. In TensorFlow, variables are usually declared

90
using placeholders which register on the computation
graph but the values are actually supplied later.
Creation of Session: After we have described the
computation graph and initialised variables, the next
step is to create a session within which the
computation graph would be executed.
Running of Graph in Session: The complete graph
along with values for placeholders is passed to the
session for execution. This is when the mathematical
operations defined in the graph takes place.
Close Session: When the graph, and related
computations have finished executing, we need to
shutdown or end the session.

gem ( + dl
)
~~ 9%. Operation

AX y aay
Variable Constant

9]
The diagram above shows a simple computation graph for a
function. Using TensorFlow we would describe something
similar that defines a neural network in the next chapter.
How to run the Neural Network using
TensorFlow
For our hands on example, we would do image classification
using the MNIST handwritten digits database which contains
pictures of handwritten digits ranging from 0 to 9 in black and
white. The task is to train a neural network that given an input
digit image, it can predict the class of the number contained
therein.

How to get our data


»

TensorFlow includes several preloaded datasets which we can


use to learn or test out our ideas during experimentation. The
MNIST database is one, of such cleaned up datasets that 1s
simple and easy to understand. Each data point is a black and
white image with only one color channel. Each pixel in the
image denotes the brightness of that point with O indicating
black and 255 white. ‘The numbers range from 0 to 255 for 784
points in a 28 X 28 erid.

Let’s go ahead and load the data from TensorFlow along with
importing other relevant libraries.

# Import MNIST data

from [Link] import


input data

mnist = input _data.read data_sets("/tmp/data/",


one_hot=True)

import numpy as np

import tensorflow as tf
import [Link] as plt

Let us use the matplotlib library to display an image to see what


it looks like by running the following lines of code.

[Link]([Link]([Link][8],
[28, 28]), cmap='gray')

[Link]()

0 5 10 uke 20 BJAt

The displayed image is a handwritten digit of number 9.

How to train and test the data

In order to train an artificial neural network model on our data,


we first need to define the parameters that describe the
computation graph such as number of neurons in each hidden
layer, number of hidden layers, input size, number of output
classes etc. Each image in the dataset is 28 by 28 pixels
therefore, the input shape 1s 784 which is 28 X 28.

94
# Parameters
learning rate = 0.1

num steps = 500

batch size = 128

display step = 100

# Network Parameters
n_hidden 1 = 10 # 1st layer number of neurons

n hidden 2 = 10 # 2nd layer number of neurons


ae ct 9
num input = 784 # MNIST data input (img shape:
28*28)

num classes = 10 # MNIST total classes (0-9


digits)

# tf Graph input

X = tf£[Link]("float", [None, num_input])

Y = t£.placeholder("float", [None, num _classes])

We then declare weights and biases which are trainable


parameters and initialise them randomly to very small values.
The declarations are stored in a Python dictionary.

# Store layers weight & bias


weights = {

Vigil, Us
tf£.Variable(tf.random_normal([num_input,
n hidden 1])),

25
ig Wea &
tf. Variable (tf. random_normal([n_hidden
1,
n_ hidden 2])),

ayia! ¢
tf£.Variable
(tf. random_normal([n_hidden 2,
num_classes]) )

f
biases = {

veal Uae
[Link]
(tf. random_normal([n_hidden_1])),

Np Zi:
tf. Variable(tf. random_normal([n_hidden
2])),

VOUle
[Link]
(tf. random_normal([num_classes]) )

We are would then describe a 3-layer neural network with 10


units in the output for each of the class digits and define the
model by creating a function which forward propagates the
inputs through the layers. Note that we are still describing all
these operations on the computation graph.

# Create model

def neural net (x):

# Hidden fully connected layer with 10


neurons

Jayer 1 = [Link]([Link] (x,


weights['h1i']), biases['bl"'])

# Hidden fully connected layer with 10


neurons

layer 2 = tf£[Link]([Link]
(layer 1,
weights['h2']), biases['b2'])

96
# Output fully connected layer with a neuron
for each class

out_layer = [Link] (layer 2,


weights['out']) + biases['out']

return out layer

Next we call our function, define the loss objective, choose the
optimizer
that would be used to train the model andinitialise
allvariables.

# Construct model ’

logits = neural
_net (X)

# Define loss and optimizer

Loss op =
[Link] mean([Link].softmax_cross_
entropy with_
logits (

logits=logits, labels=yY) )

optimizer =
tf£.[Link]
(learning rate=learning ra
te)

train_op = [Link](loss op)

# Evaluate model (with test logits, for dropout


to be disabled)

correct pred = [Link]([Link](logits, 1),


[Link](Y, 1))

accuracy = tf.reduce_mean([Link] (correct pred,


(one saadleyeyesy))))

# Initialize the variables (i.e. assign their


default value)

oT
init = tf£.global_ variables initializer ()

Finally,
we create a session, supply images in batches to the
model for training and print the loss and accuracy for each
mini-batch.

# Start training
with [Link]() as sess:

# Run the initializer

[Link] (init)

for step in range(1, num_steps+1) :

batchx, batch y =
[Link]
batch (batch_ size)

# Run optimization op (backprop)

[Link](train_op, feed dict={xX:


batchx, Y¥: batch _y})

LE step’ * display step == 0 or step ==

# Calculate batch loss and accuracy

loss, acc = [Link]([loss_op,


accuracy], feed dict={X: batch_x,

¥: batch y})

print("Step " + str(step) + ",


Minibatch Loss= " + \
Wi eat hl rormattloss)l + lly
Training Accuracy= " + \
nosh: format (acc).)

98
print ("Optimization Finished!")

# Calculate accuracy for MNIST test images

print("Testing Accuracy:", \
[Link](accuracy, feed dict={xX:
[Link],

me
[Link]}))

The session was created using with, so it automatically closes


after executing. This is the recommended way of running a
session as we would not need to manually close it. Below is the
output

Step 1, Minibatch Loss= 159.5374, Training Accuracy= 0.156


Step 100, Minibatch Loss= 1.0810, Training Accuracy= 0.773
Step 200, Minibatch Loss= 1.0142, Training Accuracy= 0.797
Step 300, Minibatch Loss= 0.5115, Training Accuracy= 0.844
Step 400, Minibatch Loss= 0.4631, Training Accuracy= 0.891
Step 500, Minibatch Loss= 0.4863, Training Accuracy= 6.867
Optimization Finished!
Testing Accuracy: 0.85

The loss drops to 0.4863 after training for 500 steps and we
achieve an accuracy of 85% on the test set.

Here is the code in full:

# Parameters

learning rate = 0.1

num_steps = 500

batch_size = 128

99
display step = 100

# Network Parameters .

n_ hidden 1 = 10 # 1st layer number of neurons

n_ hidden 2 = 10 # 2nd layer number of neurons

num_input = 784 # MNIST data input (img shape:


28%*28)

num classes = 10 # MNIST total classes (0-9


digits)

# t£ Graph input
X = tf£.placeholder("float", [None, num_input])

Y¥ = tf£.placeholder("float", [None, num_ classes] )

# Store layers weight & bias

weights = {

Weil ys
[Link](tf.random_normal([num_input,
n_ hidden 1])),

Uigia 2
tf£.Variable(tf.random_normal([n_hidden_l,
n_ hidden 2])),

VOU is
[Link](tf.random_normal([n_hidden_ 2,
num_classes]) )

}
biases = {

onlUl
[Link](tf.random_normal([n_hidden_1])),

ep c
[Link](tf.random_normal([n_hidden_2])),

100
Vertes:
[Link](tf.random_normal ([num_classes]) )

# Create model

def neural net(x):

# Hidden fully connected layer with 10


neurons
layer 1 = [Link]([Link]
(x,
weights['h1']), biases['bl1'])

# Hidden fully connected layer with 10


neurons 7

layer 2 = tf£.add(tf£.matmul
(layer 1,
weights['h2']), biases['b2'])

# Output fully connected layer with a neuron


for each class
out layer = [Link] (layer 2,
weights['out']) + biases['out']

return out layer

# Construct model

logits = neural
_net(X)

# Define loss and optimizer


loss op =
tf£.reduce_mean([Link].softmax_cross entropy with_
logits (

logits=logits, labels=yY) )

optimizer =
tf£.[Link] (learning rate=learning ra
te)

train_op = [Link](loss op)

101
# Evaluate model (with test logits, for dropout
to be disabled)

correct pred = [Link]([Link](logits, 1),


teoacgqmax (Ye)

accuracy = tf.reduce_mean([Link](correct
pred,
iene oveleyewesy™))))

# Initialize the variables (i.e. assign their


default value)

init = [Link] variables initializer ()

# Start training
with [Link]() as sess:

# Run the initializer

[Link](init)

for step in range(1, num_steps+1):

batchx, batch_y =
[Link] batch (batch size)

# Run optimization op (backprop)

[Link](train_
op, feed dict={xX:
batchx, Y: batch_y})

if step %* display step = 0 or step ==

# Calculate batch loss and accuracy

loss, acc = [Link]([loss_op,


accuracy], feed _dict={X: batch x,

Y: batch_y})

102
print("Step " + str(step) + ",
Minibatch Loss= " + \
WA cent pa oremeita (LOSS) marten or
Training Accuracy= " + \
"{:.3£}". format (acc))

print ("Optimization Finished!")

# Calculate accuracy for MNIST test images

print ("Testing Accuracy:", \


[Link](accuracy, feed dict={xX:
[Link],

NEE
[Link]}) )

103
104
Case Studies with Real Data
In this chapter we would work with data that can be used for
real world applications. Two case studies would be performed,
the first involves predicting customer churn which means how
likely is a customer to stop patronage to a business and switch
to its competitor. The second study would involve automatic
sentence classification which can be used by reviews sites to
detect users sentiments based on their review.

To enable us develop models quickly and test our hypothesis,


it is reasonable for us to use VensorFlow’s higher level APIs
which are exposed through TFLearn. TFLearn has bundled
components which are similar to Scikit-Learn but for building
deep neural networks. TFLearn is just a convenience wrapper
for TensorFlow’s lower level computation graph components.
As such TensorFlow is a dependency for TFLearn, that is to
say to use or install TFLearn, you first need to have
‘TensorFlow installed.

Since we have Tensorflow already installed from the previous


chapter, we can go ahead to install TFLearn. TFLearn can be
installed across the three major operating systems we covered
in the last chapter by using Python’s native package manager
pip. To install TFLearn we run the following command in a
terminal:

$ pip install tflearn

Bank Churn Modelling

105
Here we are presented with a case whereby a bank wants to use
data collected from its customers over several years to predict
which customers are likely to stop using the bank’s services by
switching to a competing bank. The rewards of such an analysis
to the bank is profound as it can target dissatisfied customers
with incentives which would reduce the churn ratio helping the
bank to grow its customer base and solidify its position.

The dataset contains many informative attributes such as


account balance, number of products subscribed to, credit card
status, estimated customer salary etc. The target variable 1s
whether or not the customer left the bank, so this is a binary
classification task. There are also some categorical features
such as gender and geography which we would need to
transform before feeding them into a neural network.

The churn modelling dataset can be downloaded at:

https: // /[Link]/aakash50897 /churn-


modellingcsv/data

As always we first import all relevant libraries, load the dataset


using Pandas and call the head method on the dataset to see
what ts contained inside.

# import all relevant libraries


import numpy as np

import [Link] as plt

import pandas as pd

import tensorflow as tf

import tflearn

106
# load the dataset

dataset = pd.read_csv('Churn_Modelling.csv')

# get first five rows (observations)

[Link]()

imber Customerld Sumame CreditScore Geography Gender Age Tenure Balance NumOfProducts HasCrCard IsActiveMember EstimatedSalary Exited

1 15634602 Hargrave 619 France Female 42 H (00 1 1 1 10134888 = 1

2 {5647311 Hil 608 = Spain Female 41 1 8380786 1 0 1 {254258 =

3 15619304 Onio 502 France Female 42 8 189660.80 3 1 0 303157 = 1

4 15701354 = Bow 699 France Female 39 1 0.00 2 04 0 9382663 ©0

5 15737888 Mitchel 85 Spain Female 43 2 12551682 1 t 1 73084100

The dataset contains 14 columns, the first 3 columns are


uninformative namely RowNumber, CustomerId and Surname. ‘hose
three columns can be seen as identifiers as they do not provide
any information which would give insights to whether a
customer would stay or leave. They would be removed before
we perform analysis. The last column - Exited is the class label
which our model would learn to predict.

The next step having gotten an overview of the dataset is to


split the columns into features and labels. We do this using
Pandas slicing operation which selects information form
specified indexes. In our case, features start from the 3rd
column and ends in the 12th column. Remember that array
indexing starts at 0 not 1.

X = [Link][:, 3:13].values

y = [Link][:, 13].values

107
Since we have categorical features 1n the dataset (Geography and
Gender), we have to convert them into a form that a deep
learning algorithm can process. We do that using a one-hot
representation. One-hot representation creates a sparse matrix
with zeros in all positions and a | at the position representing
the category under evaluation. We use Scikit-Learn’s
pteprocessing model to first create a label encoder, then create
a one-hot representation from it.

# encoding categorical data

from [Link] import LabelEncoder,


OneHotEncoder

labelencoder
X 1 = LabelEncoder
()

x[:, 1] = labelencoder
x [Link] transform(xX[:,
tj)
labelencoder
X 2 = LabelEncoder()

X[:, 2] = labelencoder
X 2.fit_transform(X[:,
2])
onehotencoder =
OneHotEncoder (categorical features = [1])

X = [Link]
transform (X) . toarray ()

eX sy ee]

We split our data into training and test set. One would be used
to train the model while the other would be use to test
performance.

# spliting the dataset into the training set and


test set

from [Link] selection import


train test split

108
X_ train, X_test, y_ train, y_test
train_test_split(X, y, test_size = ORrz a
random state = 0)

y_train = [Link](y train, (-1, 1)) # reshape


y_train to [None, i]

y_test = [Link](y test, (-1, 1)) # reshape


y_test to [None, 1]

The features as currently contained in X are not in the same


scale so we apply standard scaling which makes all features to
have a mean of 0 and a standard deviation of 1.
2

# feature scaling

from [Link] import StandardScaler

sc = Standardscaler()

x train = [Link] transform


(xX. train)

X_test = [Link](X
test)

We start creating a deep neural network by describing it using


TY earn APL

# build the neural network

net = [Link] data(shape=[None, 11])

net [Link] connected (net, 6 /

activation='relu')

net [Link](net, 0.5)

net = [Link] connected (net, 6 La

activation='relu')

net = [Link](net, 0.5)

net = [Link] connected (net, al v

activation='tanh')

109
net = [Link]
(net)

The network has 11 input features and there are 3 fully


connected layers. We also use dropout as the regularizer in
order to prevent the model from overfitting. Next we define
the model using DNN from TFLearn.

# define model

model = [Link] (net)

# we start training by applying gradient descent


algorithm

[Link](X_ train, y train, n_epoch=10,


batch_size=16, validation_set=(X test, y test),

show_metric=True,
run_id="dense model")

Training Step: 4999 |total loss: nan |time: 2.9385


|Adam |epoch: 610 |loss: nan - binary acc: 0.7647 -- iter: 7984/8000
Training Step: 5600 |total Loss: nan |tine: 4.013s
|Adam |epoch: O16 |loss: nan - binary
acc: 8.7507 |valloss: nan - val acc: @.7885 -- iter: 8080/8000

We train the model for 10 epochs with a batch size of 16. The
model achieves an accuracy of 0.7885 on the test set which we
used to validate the performance of the model.

Sentiment Analysis

For this real world use case we tackle a problem from the field
of Natural Language Processing (NLP). The task is to classify
movie reviews into classes expressing positive sentiment about
a movie or negative sentiment. ‘l’o perform a task like this, the
model must be able to understand natural language, that 1s it

110
must know the meaning of an entire sentence as expressed by
its class prediction. Recurrent Neural Networks (RNNs) are
usually well suited for tasks involving sequential data like
sentences however, we would apply a _ 1-dimensional
Convolutional Neural Network (CNN) model to this task as it
is easier to train and produces comparable results.

The dataset we would use is the IMDB sentiment database


which contains 25,000 movie reviews in the training set and
25,000 reviews in the test set. TFLearn bundles this dataset
alongside others so we would access it from the datasets
>
module.

First we import the IMDB sentiment dataset module and other


relevant components from TFLearn such as convolutional
layers, fully connected layers, data utilities etc.

import tensorflow as tf

import tflearn

from [Link] import input data,


dropout, fully connected

from [Link] import conv_ld,


global_max pool

from [Link]
ops import merge

from [Link] import regression

from tflearn.data_utils import to categorical,


pad_sequences

from [Link] import imdb

The next step is to actually load the dataset into the train and
test splits

dd
# load IMDB dataset

train, test, = imdb.load_data(path='[Link]',


n_words=10000,

valid _portion=0.1)

trainxX, trainY = train

testxX, testY = test

The next phase involves preprocessing the data where we pad


sequences which means we set a maximum sentence length and
for sentences less than the maximum sentence length we add
zeros to them. The reason is to make sure that all sentences are
of the same length before they are passed to the neural network
model. The labels in the train and test sets are also converted
to categorical values.

# data preprocessing
# sequence padding
trainX = pad_sequences(trainX, maxlen=100,
value=0.)

testX = pad _sequences(testX, maxlen=100,


value=0.)

# converting labels to binary vectors

trainY = to _categorical(trainY, nb _classes=2)

testY = to_categorical(testY, nb_classes=2)

As we saw in the previous example, the next step is to describe


a 1-dimensional Convolutional Neural Network model using
the building blocks provided to us by TF'Learn.
# building the convolutional network

network = input data (shape=[None, LOORF


name='input')

network = [Link] (network,


input_dim=10000, output_dim=128)

branchl = conv_ld(network, 128, 3,


padding='valid', activation='relu',
regularizer="L2")

branch2 = conv_ld(network, 128, 4,


padding='valid', activation='relu',
regularizer="L2") 5

branch3 = conv_ld(network, 128, 5,


padding='valid', activation='relu',
regularizer="L2")

network = merge([branchl1, branch2, branch3],


mode='concat', axis=1)

network = [Link]
dims (network, 2)

network = global_max_pool (network)

network = dropout(network, 0.5)

network = fully connected(network, 2,


activation='softmax')

network = regression(network, optimizer='adam',


learning rate=0.001,

loss='categorical crossentropy', name='target')

The network contains 3 1-dimensional convolutional layers, a


global max pooling layer used to reduce the dimension of
convolutions, a dropout layer used for regularization and a
fully connected layer. Finally, we declare the model and call fit
method on it to begin training.
# training
model = [Link] (network,
tensorboard verbose=0) ~

[Link](trainxX, trainY, n_epoch=5,


shuffle=True, validation set=(testX, testY),
show_metric=True, batch _size=32)

Training Step: 3519 |total loss: 6.10591 |time: 339. 604s


|Adan |epoch: 865|loss: 6.10591-acc: 0.9809 -- iter: 22496/22500
Training Step: 3520 ne loss: 0.10458 |time: 348, 865s
|Adan |epoch: 605 | loss; 6.10458-acc: 0.9026
|valloss:0.55262 - val ace: 0.8040 -- iter: 22500/22500

The trained model achieves an accuracy of 0.80 on the test set


which 1s to say it correctly classified the sentiment expressed 1n
80% of sentences.

Here is the code used for training the model in full:

# import tflearn, layers and data utilties

import tensorflow as tf

import tflearn

from [Link] import input data,


dropout, fully connected

from [Link] import conv_ld,


global_max_pool

from [Link]
ops import merge

from [Link] import regression

from tflearn.data_utils import to_categorical,


pad_sequences

from [Link] import imdb

# load IMDB dataset

114
train, test; — = imdb.load_data(path='[Link]',
n_words=10000,

valid portion=0.1)

trainxX, trainY = train

testx, testY = test

# data preprocessing
# sequence padding
trainX = pad_ sequences (trainx, maxlen=100,
value=0.) :

testX = pad_sequences(testX, maxlen=100,


value=0.)

# converting labels
to binary vectors

trainY = to categorical (trainY, nb_classes=2)

testY = to categorical (testY, nb classes=2)

# building the convolutional network

network = input_data(shape=[None, 100],


name='input')

network = [Link] (network,


input_dim=10000, output_dim=128)

branchl = conv_ld(network, 128, 3,


padding='valid', activation='relu',
regularizer="L2")

branch2 = conv_ld(network, 128, 4,


padding='valid', activation='relu',
regularizer="L2")

branch3 = conv_ld(network, 128, 5,


padding='valid', activation='relu'
regularizer="L2")

network = merge([branchl, branch2, branch3],


mode='concat', axis=1)

A
network = tf£.expand_ dims (network, 2)

network = global_max_pool (network)

network = dropout(network, 0.5)

network = fully connected(network, 2,


activation='softmax')

network = regression(network, optimizer='adam',


learning rate=0.001,

loss='categorical crossentropy', name='target')

# training
model = [Link] (network,
tensorboard_ verbose=0)

[Link](trainx, trainY, n_epoch=5,


shuffle=True, validation _set=(testx, testY),
show metric=True, batch _size=32)

116
ie Gee wae,
ipev-ine l fsa Oe | imate

Sy eh aied) paquae ©
$s Lei YSine 4
is cer Cam! Was.

¥eini—a4 Ide ced) aretesryen © ss inten


shh tenga Sate

oa | fa rd nf pankeegesel’ “anes.

ae 7 : = 7

& a >i Sats »

¢ ort we ® dart AG bie feahce *

—SHeatwy Pipseiuhoet.

Qrwaliies jut wr tecmrha


di ij4i> |letieom
{an ~~ spool TRinda _
ts ¥ ie = 19a

: iy a? es mr Ot? tween
Conclusion

There are a lot more teal world applications of deep learning


in consumer products today than at any point in history. It is
generally said that 1f you can get large amounts of data and
enormous computation power to process that data, then deep
learning models could help you provide business value
especially in tasks where humans are experts and the training
data is properly annotated.

Dip
Thank you!
Thank you for buying this book! It is intended to help you
understanding machine learning using Python. If you enjoyed
this book and felt that it added value to your life, we ask that
you please take the time to review It.

Your honest feedback would be greatly appreciated. It really


does make a difference.

a
a
AL SCIENCES
We are a very small publishing company and our
sutvival depends on your reviews.
Please, take a minute to write us your review.

120
Sources & References
Software, libraries, & programming language

Python (https:
Anaconda (https: Lb
Virtualenv (https:
Numpy ( www. [Link]
Pandas (https:/
Matplotlib ([Link]
Scikit-learn (http: [Link] fi/
TensorFlow (ht [Link]/
TFLearn ([Link]

Datasets

Kagele ( www.! [Link] /datase ts)


Boston Housing Dataset
([Link] [Link]/p/rdataset/sourc
e/ file /master/csv/ ‘/[Link])
e@ Pima Indians Diabetes Database
[Link] ‘ucim 5ima-indians-
diabetes-database/ data
e@ Iris Dataset

[Link] / 1)
e@ Bank Churn Modelling
[Link]/aakash50897
modellingcsv /data)

Online books, tutorials, & other references

ea
Coursera Deep Learning Specialization
(https: [Link], /specializations /deep-
learning)
[Link] - Deep Learning for Coders
([Link]
Overfitting
(https: [Link]/wiki/Overfitting)
A Neural Network Program
(https: //[Link]
TensorFlow Examples
(https: //[Link] ‘aymericdamien /TensorFlow-
Examples)
TFLearn Examples
/master/exa
mples)
Machine Learning Crash Course by Google
(https: //[Link]
Choosing the Right Estimator ([Link]
[Link]/stable/tutorial/machine learning ma
ex. html)
Cross-validation: evaluating estimator performance
(http: //sctkit-
[Link]/stable/modules/cross_ validation. html )

122
n> a nee airtel
——s 7ear
roe
perf nbarrod. ee a
Apatite Tye) -
2 ” : 7 \yeeto4s- Pf a

‘— i me 28). stah-
uni? mh ae, hy -?

ep blysati er :
eames.=" We. ©
: newsedia ass
or
=, >~i52 A < ' =) t= ”

ry « 4 : mye lk” i z
, ~

hay db eee 7Tt \ ’ mid a? wf, 6

‘ = «
v 7 cs ia a » er bi
Thank you!
Thank you for buying this book! It is intended to help you
understanding machine learning using Python. If you enjoyed
this book and felt that it added value to your life, we ask that
you please take the time to review it.

Your honest feedback would be greatly appreciated. It really


does make a difference.

a
ae
Al SCIENCES
We are avery small publishing company and our
survival depends on your reviews.
Please, take a minute to write us your review.

[Link]/dp/BO7F193447

124
oe :

Al SCIENCES
NNN
33852192R00076
een
Middletown, DE
18 January 2019
Wi
Python Machine
Learning
from Scratch

If you are looking for a practical book to help you understand Machine
Learning step by step by using Python, then this is a good book for you.
Is this book for me?

e Anyone curious about machine learning but with zero programming


knowledge
* People who want to demystify machine learning (it’s not magic and
it's probably not the end of the world)
e« Technical people who want to quickly gain knowledge in machine
learning

To better understand Machine Learning requires a bit of knowledge


about Statistics, linear algebra, programming and computer science.
We'll explain the most important concepts as we go along. We'll
also dive into Python programming so you'll be familiar with the
behind-the-scenes of mac

oe
SBN 9781725929982

Zov-zN
9°781725°9 AU
UEIUA
CU
UOSOS
QEQUT
ATH
HAUS
A

You might also like