0% found this document useful (0 votes)
18 views156 pages

Machine Learning NETWORK DeepLearning

The document discusses various aspects of machine learning, including supervised, unsupervised, and reinforcement learning algorithms, along with their applications across different industries. It highlights the importance of analytics in predicting trends and making informed decisions, and provides examples of linear regression in predicting house prices. Additionally, it covers the ETL process in machine learning for data transformation and analysis.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
18 views156 pages

Machine Learning NETWORK DeepLearning

The document discusses various aspects of machine learning, including supervised, unsupervised, and reinforcement learning algorithms, along with their applications across different industries. It highlights the importance of analytics in predicting trends and making informed decisions, and provides examples of linear regression in predicting house prices. Additionally, it covers the ETL process in machine learning for data transformation and analysis.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

SUPERVISED LEARNING

MACHINE LEARNING
DEEP LEARNING Algorithms

Narration By

Dr. B. SWAMINATHAN
Dr. B. Swaminathan 2
Deep Learning Applications Used Across Industries
1. Virtual Assistants 2. Chat-bots 3. Healthcare
4. Entertainment 5. Composing Music 6. Image Coloring
7. Robotics 8. Image Captioning 9. Advertising
10. Self Driving Cars 11. Visual Recognition 12. Fraud Detection
13. Personalization's 14. Deep Dreaming 15. Pixel Restoration
16. Automatic Game Playing
17. Language Translations
18. Automatic Handwriting Generation
19. Demographic and Election Predictions
20. Natural Language Processing
21. News Aggregation and Fake News Detection
22. Adding Sounds to Silent Movies
23. Detecting Developmental Delay in Children
24. Colourization of Black and White images
25. Automatic Machine Translation
Dr. B. Swaminathan 3
Analytics(Predictions)
Descriptive Analysis: Which tells us what has already happened (The process of using
current and historical data to identify trends and relationships.)
• Traffic and Engagement Reports,
• Financial Statement Analysis : holistic view of a company’s financial health.; balance
sheet, income statement, cash flow statement,.
Vertical analysis, Horizontal analysis, ratio analysis,
• Demand Trends: NetFlix’s (use case)
• Aggregated Survey Results
• Progress to Goals(Key Performance Analysis)
Predictive Analysis: What could happen
• Finance: Forecasting Future Cash Flow
• Entertainment & Hospitality: Determining Staffing Needs
• Marketing: Behavioral Targeting
• Manufacturing: Preventing Malfunction
• Health Care: Early Detection of Allergic Reactions
Prescriptive Analysis: What should Happen in the Future “What should we do next?”
• .Venture Capital: Investment Decision
• Sales: Lead Scoring; Page views, Email interactions, Site searches, Content
engagement,
• Content Curation: Algorithmic Recommendations(TikTok’s )
• Banking: Fraud Detection
• Product Management: Development and Improvement
Dr. B. Swaminathan 4
Analytics(Predictions) Cont…
Diagnostic Analysis: Why did this happen?”
Cognitive Analysis:
SWOT: Strength Weakness ,Opportunities, and Threads

Dr. B. Swaminathan 5
6
7
What is Intelligence ?

•Intelligence:
•“the capacity to learn and solve problems” in particular,
- the ability to solve novel problems
- the ability to act rationally
- the ability to act like humans

8
9
10
11
12
AI Applications
• Autonomous Planning
& Scheduling:
– Autonomous rovers.

01/04/2021 13
AI Applications
• Autonomous Planning & Scheduling:
– Telescope scheduling

01/04/2021 14
AI Applications
• Autonomous Planning & Scheduling:
– Analysis of data:

01/04/2021 15
AI Applications
• Medicine:
– Image guided surgery

01/04/2021 16
AI Applications
• Medicine:
– Image analysis and enhancement

01/04/2021 17
AI Applications

• Transportation:
– Autonomous vehicle
control:

01/04/2021 18
AI Applications
• Transportation:
– Pedestrian detection:

01/04/2021 19
AI Applications

Games:

01/04/2021 20
AI Applications
• Games:

01/04/2021 21
AI Applications
• Robotic toys:

01/04/2021 22
23
24
Was

25
Was

is

26
27
is

28
Systems:
Other application areas:
Computer Games ● Bioinformatics:
Navigation Systems ○ Gene expression data analysis
○ Prediction of protein structure
Smart phone services ● Text classification, document sorting:
○ Web pages, e-mails
Intelligent email ○ Articles in the news
Search engines ● Video, image classification
● Music composition, picture drawing
Recommender systems ● Natural Language Processing
Self-driving cars ● Perception

29
Introduction to Artificial Intelligence
Ability of a machine to imitate
Artificial intelligent human behavior
Intelligence
System is ability to perform Automatically
Machine Learn and improve the performance
Learning
Machine uses algorithm and
uses the neural net to train the model
Deep Learning

Dr. B. Swaminathan
30
Machine Learning
Supervised Learning Algorithm
It’s Produce the o/p based on previous information(data)
Dataset have attributes (label) to train algorithms.
O/P is classify data or predict
Types
Regression
Classification

Unsupervised Learning Algorithm


Unlabeled and Uncategorized
Similarities
Types
Clustering
Association

Reinforcement Learning Algorithm


Rewards
Dr. B. Swaminathan 31
Supervised Learning Algorithm
Classification vs. Regression
Regression
• Linear Regression
• Regression Trees
• Non-Linear Regression
• Bayesian Linear Regression
• Polynomial Regression
Classification
• Random Forest
• Decision Trees
• Logistic Regression
• Support vector Machines
Classification produces discrete category predictions. Regression produces continuous number predictions.
Dr. B. Swaminathan 32
Unsupervised Learning Algorithm
• Clustering
– K-means clustering
– KNN (k-nearest neighbors)
– Hierarchal clustering
– Anomaly detection
• Association
– Apriori algorithm
– Neural Networks
– Principle Component Analysis
– Independent Component Analysis
– Singular value decomposition
Dr. B. Swaminathan 33
Reinforcement Learning Algorithm
Learning Models,
– Markov Decision Process (parameters)
• Set of actions- A
• Set of states -S
• Reward- R
• Policy- n
• Value- V
– Q learning ---value-based method of supplying
information to inform which action an agent should take
• Value-Based,
• Policy-based
– Deterministic,
– Stochastic
Dr. B. Swaminathan 34
Machine Learning Techniques
Linear Regression
Machine Learning:
Supervised Learning(Linear Regression)
House Price in Square Feet
100000 (y) (x)

𝒙−𝑿 ഥ
𝒚−𝒀 ഥ )(𝒚 − 𝒀
(𝒙 − 𝑿 ഥ) ഥ
𝒙−𝑿 𝟐 ഥ
𝒚−𝒀 𝟐

245 1400 -315.00 -41.50 13072.50 99225 1722.25


312 1600 -115.00 25.50 -2932.50 13225 650.25
279 1700 -15.00 -7.50 112.50 225 56.25
308 1875 160.00 21.50 3440.00 25600 462.25
199 1100 -615.00 -87.50 53812.50 378225 7656.25
219 1550 -165.00 -67.50 11137.50 27225 4556.25
405 2350 635.00 118.50 75247.50 403225 14042.25
324 2450 735.00 37.50 27562.50 540225 1406.25
319 1425 -290.00 32.50 -9425.00 84100 1056.25
255 1700 -15.00 -31.50 472.50 225 992.25
ഥ = 286.50
𝒀 ഥ = 1715.00
𝑿 172500.00 1571500.00 32600.50

𝑆𝑥
𝒚=𝒂+𝒃𝒙 𝑏= 𝑟
𝑆𝑦
; 𝑎 = 𝑌ത − 𝑏 𝑋ത


σ((𝑥 − 𝑋)(𝑦 ത
− 𝑌)) σ 𝑥 − 𝑋ത 2 σ 𝑦 − 𝑌ത 2
𝑟= ; 𝑆𝑥 = ; 𝑆𝑦 =
σ 𝑥 − 𝑋ത 2 σ 𝑦 − 𝑌ത 2 (𝑛 − 1) (𝑛 − 1)
Machine Learning:
Supervised Learning(Linear Regression)

HousePrice vs SquareFeet
3000
house price = 98.24833 + 0.10977 (square feet)
2500
HousePrice(Y)

2000
1500
1000
500
0
0.00 100.00 200.00 300.00 400.00 500.00
Series1 Linear (Series1) SquareFeet(X)

Predict the price for a house with 2000 square feet:


𝒚=𝒂+𝒃𝒙
House Price = 98.25 +0.1098 (2000) = 317;
The predicted price for a house with 2000 square feet is 317.85(100,00s) = 3178500
Mean of x 17150/10
Mean of y 2865/10
r 172500/sqrt(1571500*32600.50) 0.762113713
Sx sqrt(1571500/(10-1)) 417.8649436
Sy sqrt(32600.50/(10-1)) 60.18536182

b 0.762114713 * (60.18536182/417.86) 0.10976918


a 286.5 - (0.10976*1715) 98.2616
• File -> Options
Square Feet (X)
3000

2500

2000

1500 Square Feet (X)


Linear (Square Feet (X))
1000

500

0
0.00 50.00 100.00 150.00 200.00 250.00 300.00 350.00 400.00 450.00
Linear Regression Example
Using Excel

Tools (Select Data)


Data Analysis
Regression
Linear Regression Example Excel
Output
Regression Statistics
Multiple R 0.76211
The regression equation is:
R Square 0.58082

Adjusted R Square 0.52842


house price = 98.24833 + 0.10977 (square feet)
Standard Error 41.33032
Observations 10

ANOVA
df SS MS F Significance F

Regression 1 18934.9348 18934.9348 11.0848 0.01039


Residual 8 13665.5652 1708.1957
Total 9 32600.5000

Coefficients Standard Error t Stat P-value Lower 95% Upper 95%

Intercept 98.24833 58.03348 1.69296 0.12892 -35.57720 232.07386

Square Feet 0.10977 0.03297 3.32938 0.01039 0.03374 0.18580


Linear Regression Example
Graphical Representation

• House price model: scatter plot and regression line


450
400
House Price ($1000s)

350
300 Slope
250
200
= 0.10977
150
100
50
Intercept 0
0 500 1000 1500 2000 2500 3000
= 98.248 Square Feet

house price = 98.24833 + 0.10977 (square feet)


Example Auto insurance
Experience x Premium y xy x2 y2
•3. Regression line, we calculate a and b as follows:
5 64 320 25 4096
2 87 174 4 7569
12 50 600 144 2500
4. Regression line ŷ = a + bx is
9 71 639 81 5041
15 44 660 225 1936 5. Scatter diagram
6 56 336 36 3136
25 42 1050 625 1764
16 60 960 256 3600
Σx = 90 Σy = 474 Σxy = 4739 Σx2= 1396 Σy2 = 29,642

1.
6. The values of r and r2 are computed
2
• standard deviation of errors is

• To construct a 90% confidence interval for B, first we calculate the standard


deviation of b:

• For a 90% confidence level, the area in each tail of the t distribution is

• The degrees of freedom are


Dr. B. Swaminathan 49
Dr. B. Swaminathan 50
Dr. B. Swaminathan 51
Dr. B. Swaminathan 52
Dr. B. Swaminathan 53
Dr. B. Swaminathan 54
Dr. B. Swaminathan 55
Data Extract, transform, load (ETL)
Process of copying data from one or more sources into a destination files
which represents the data differently than the sources.

ETL Process: The typical Machine Learning ETL process is one of successive data
operations:-

Extract Functions [Link]

•Data Extract functions can include:


•base Queries
•File Reading and/or Data Selection
Transform Functions
•Cleaning
•Normalization
•Regularization
Load Functions
•Load functions can include:
•Data base Load
Dr. B. Swaminathan 56
•File Write
Machine Learning Basics
• Artificial Intelligence is a scientific field concerned with the
development of algorithms that allow computers to learn
without being explicitly programmed
• Machine Learning is a branch of Artificial Intelligence, which
focuses on methods that learn from data and make
predictions on unseen data

Machine Learning
Labeled Data algorithm

Training
Prediction

Learned Predictio
Labeled Data model n

Dr. B. Swaminathan 57
Anatomy of a deep neural network
• Layers
• Input data and targets
• Loss function
• Optimizer

Dr. B. Swaminathan 58
Input data and targets

• The network maps the input data X


to predictions Y′
• During training, the predictions Y′
are compared to true targets Y
using the loss function
cat

dog

Dr. B. Swaminathan 59
Deep learning Lasagn Keras TF torch.n Gluo
e Estimator n n
frameworks Thean TensorFlo CNTK PyTorc MXNe Caffe
o w h t
CUDA, cuDNN
HIP, MIOpen MKL, MKL-
• Keras is a high-level DNN
neural networks API
GPUs CPUs
• we will use TensorFlow
as the compute backend
• included in TensorFlow 2 as [Link]
• [Link] , [Link]
• PyTorch is:
• a GPU-based tensor library
• an efficient library for dynamic neural networks
• [Link]

Dr. B. Swaminathan 60
Natural Neural Network

Dendrites : These receive information or signals from other neurons that get
connected to it.
Cell Body : Information processing happens in a cell body. These take in all the
information coming from the different dendrites and process that information.
Axon : It sends the output signal to another neuron for the flow of
information. Here, each of the flanges connects to the dendrite or the hairs on
the next one. Dr. B. Swaminathan 61
Artificial Neural Network

ANN replicates a biological neuron.


Input to a neuron - input layer
Neuron - hidden layer
Output to the next neuron - output layer
Dr. B. Swaminathan 62
Biological Neural Network. Artificial Neural Network

Dendrites (designed to receive communications from other cells) from Biological Neural
Network represent inputs in Artificial Neural Networks, cell nucleus represents Nodes, synapse
represents Weights, and Axon represents Output.
Biological Neural Network Artificial Neural Network

Dendrites Inputs
Cell nucleus Nodes
Synapse Weights
Axon Dr. B. Swaminathan Output 63
Defining Neural Networks

Dr. B. Swaminathan 64
Neural Network Types

Standard NN Convolution NN Recurrent NN

Dr. B. Swaminathan 65
Neural Network Representation
Feed forward Neural Network – Artificial Neuron
Radial basis function Neural Network
Self Organizing Neural Network
Recurrent Neural Network
Convolutional Neural Network
Modular Neural Network

Dr. B. Swaminathan 66
ML vs. Deep Learning

• Conventional machine learning methods rely on human-designed feature


representations
– ML becomes just optimizing weights to best make a final prediction

Dr. B. Swaminathan 67
ML vs. Deep Learning

• Deep learning (DL) is a machine learning subfield that uses multiple layers for
learning data representations
– DL is exceptionally effective at learning patterns

Dr. B. Swaminathan 68
ML vs. Deep Learning

• DL applies a multi-layer process for learning rich hierarchical features (i.e., data
representations)
– Input image pixels → Edges → Textures → Parts → Objects

Low-Level Mid-Level High-Level Trainable


Output
Features Features Features Classifier

Dr. B. Swaminathan 69
Elements of Neural Networks

z = a1w1 + a2 w2 +  + aK wK + b
a1 w1
a2 w2
z  (z )
+ a

wK output

aK weights
Activation
function
input
b
Dr. B. Swaminathan 70

bias
Elements of Neural Networks
• A NN with one hidden layer and one output layer

Weights Biases

Activation functions
4 + 2 = 6 neurons (not counting inputs)
[3 × 4] + [4 × 2] = 20 weights
4 + 2 = 6 biases
26 learnable parameters
Dr. B. Swaminathan 71
Types of Activation functions:
1. Step function,
2. Sign function,
3. Sigmoid function,.

Dr. B. Swaminathan 72
Neural Networks (Activation Function)

Linear Activation Function: Y= MX+ C Sigmoid function: A = 1/(1+e-x)


Tanh Function: F(x)= tanh(x) = 2/(1+e-2X) – 1
(OR)
Tanh (x) = 2 * sigmoid(2x) – 1

Rectified Linear Leaky ReLU Softmax Function


Unit Activation Dr. B. Swaminathan 73
Dr. B. Swaminathan 74
Softmax Layer

• In multi-class classification tasks, the output layer is typically a softmax layer


▪ I.e., it employs a softmax activation function
▪ If a layer with a sigmoid activation function is used as the output layer instead, the
predictions by the NN may not be easy to interpret
o Note that an output layer with sigmoid activations can still be used for binary classification

A Layer with Sigmoid Activations

z1
3

0.95
y1 =  z1 ( )
z2
1

0.73
y2 =  z 2 ( )
z3
-3

0.05
y3 =  z 3 ( )

Slide credit: Hung-yi Lee – Deep Learning Tutorial


Softmax Layer

• The softmax layer applies softmax activations to output a


probability value in the range [0, 1] Probability
▪ The values z inputted to the softmax layer are referred to as ▪ 0 < 𝑦𝑖 <
logits
A Softmax Layer ▪ σ𝑖 𝑦𝑖 =
3 0.88 3

e
20
z1 e e z1
 y1 = e z1
zj

j =1
0.12 3
z2 1
e e z2 2.7  y2 = e z2 e
zj

j =1
0.05 ≈0 3
z3 -3 
e
e
z3 zj
e y3 = e z 3
3 j =1

+ e
zj

j =1

Slide credit: Hung-yi Lee – Deep Learning Tutorial


Activation: Sigmoid

• Sigmoid function σ: takes a real-valued number and “squashes” it into the range
between 0 and 1
▪ The output can be interpreted as the firing rate of a biological neuron
o Not firing = 0; Fully firing = 1
▪ When the neuron’s activation are 0 or 1, sigmoid neurons saturate
o Gradients at these regions are almost zero (almost no signal will flow)
▪ Sigmoid activations are less common in modern NNs

𝑓 𝑥

Slide credit: Ismini Lourentzou – Introduction to Deep Learning


Activation: Tanh

• Tanh function: takes a real-valued number and “squashes” it into range between -
1 and 1
▪ Like sigmoid, tanh neurons saturate
▪ Unlike sigmoid, the output is zero-centered
o It is therefore preferred than sigmoid
▪ Tanh is a scaled sigmoid: tanh(𝑥) = 2 ∙ 𝜎(2𝑥) − 1

𝑓 𝑥 ℝ𝑛 → −1,1
Activation: ReLU

• ReLU (Rectified Linear Unit): takes a real-valued number and thresholds it at


zero
𝑓 𝑥 = max(0, 𝑥)
ℝ𝑛 → ℝ𝑛+
▪ Most modern deep NNs use ReLU
activations
▪ ReLU is fast to compute
o Compared to sigmoid, tanh 𝑓 𝑥
o Simply threshold a matrix at zero
▪ Accelerates the convergence of gradient
descent
o Due to linear, non-saturating form
▪ Prevents the gradient vanishing
problem 𝑥
Activation: Leaky ReLU

• The problem of ReLU activations: they can “die”


▪ ReLU could cause weights to update in a way that the gradients can become zero and
the neuron will not activate again on any data
▪ E.g., when a large learning rate is used

• Leaky ReLU activation function is a variant of ReLU


▪ Instead of the function being 0 when 𝑥 < 0, a leaky ReLU has a small negative slope
(e.g., α = 0.01, or similar)
▪ This resolves the dying ReLU
problem
▪ Most current works still use ReLU
o With a proper setting of the learning 𝛼𝑥 for 𝑥 < 0
𝑓 𝑥 =ቊ
rate, the problem of dying ReLU can 𝑥 for 𝑥 ≫ 0
be avoided
Activation: Linear Function

• Linear function means that the output signal is proportional to the input signal to
the neuron

▪ If the value of the constant c is 1, it is


also called identity activation
function ℝ𝑛 → ℝ𝑛
▪ This activation type is used in
regression problems 𝑓 𝑥 = 𝑐𝑥
o E.g., the last layer can have linear
activation function, in order to output
a real number (and not a class
membership)
Dr. B. Swaminathan 82
Perceptron Models
1. Single-layer Perceptron Model
A single-layered perceptron model consists feed-forward network and also
includes a threshold transfer function inside the model.

Single-layer perceptron model is to analyze the linearly separable objects with


binary outcomes.

Single-layer perceptron can learn only linearly separable patterns.

2. Multi-layer Perceptron model


A multi-layer perceptron model also has the same model structure but has a
greater number of hidden layers.

The multi-layer perceptron model is also known as the Back propagation


algorithm, which executes in two stages as follows:
1. Forward Stage: Activation functions start from the input layer in the forward
stage and terminate on the output layer.
2. Backward Stage: In the backward stage, weight and bias values are modified
as per the model's requirement.
In this stage, the error between actual output and demanded originated
backward on the output layer and ended on the input layer
Dr. B. Swaminathan 83
Multilayer Perceptrons (MLPs)
class of feed-forward neural networks with multiple layers of perceptrons.
• Deep NNs have many hidden layers
– Fully-connected (dense) layers (a.k.a. Multi-Layer Perceptron or MLP)
– Each neuron is connected to all neurons in the succeeding layer

Input Layer 1 Layer 2 Layer L Output


x1 …… y1
x2 …… y2

……
……

……

……

……
xN …… yM

Input Layer Output Layer


Hidden Layers

Dr. B. Swaminathan 84
Five recognized types of neural networks.
Single-layer feed-forward network.
Multilayer feed-forward network.
Model structure is same but has a greater number of hidden layers.
Multi-layer model is also known as the Back propagation algorithm,
It executes in two stages
1. Feed Forward propagation
• A feed forward neural network nodes never form a cycle.
• Neural Network has an input layer, hidden layers, and an output
layer.
2. Backward propagation
• Task is to classify data set.
• Computes the gradient of the loss function for a single weight by
the chain rule.
Single node with its own feedback.
Single-layer recurrent network.
Multilayer recurrent network.

Dr. B. Swaminathan 85
Single-layer feed-forward network.
Input Output
Only Input and output layers.
X1
w11
O1 No computation is performed in Input layer.
Different weights are applied to input nodes &
w12
Cumulative effect per node is taken. ,
X2
w21
O2 Neurons collectively give the output layer to
W31
compute the output signals.
X3 O3 In Output layer only calculations are ther
Wn1

Xn On
Wnn

Dr. B. Swaminathan 86
Sample Calculation
Given:
x1 = 2, x2 =3;
w1 = 0, w2 = 1;
bias = 0;
Activation = Sigmoid;

Activation Function
h1 = x1 * w1 + x2 * w2 ⇒ 2*0 + 3 * 1 = 3 = 3
h2 = x1 * w1 + x2 * w2 ⇒ 2*0 + 3 * 1 = 3 = 3
X = (h1 * w1 + h2 * w2) + b ⇒ 3*0 + 3*1 + 0 = 30

Dr. B. Swaminathan 87
Sample Calculation 2

• A simple network, toy example

1 ∙ 1 + −1 ∙ −2 + 1 = 4
4 0.98 Sigmoid Function
1
1
-2 1
 (z ) =
1 1 + e−z
 (z )
-1 -2 0.12
-1
1 z
0
• A simple network, toy example (cont’d)
▪ For an input vector [1 −1]𝑇 , the output is [0.62 0.83]𝑇

1 4 0.98 2 0.86 3 0.62


1
-2 -1 -1
1 0 -
2
-1 -2 0.12 -2 0.11 -1 0.83
-1
1 -1 4
0 0 2

𝑓: 𝑅2 → 𝑅2 1 0.62
𝑓 =
−1 0.83
Example
Single-layer recurrent network:

1) Network needs to recognize the sentence:


What is the time?

2) Each word comes in as a pattern of sound


Sentence gets sampled into discrete sound waves. Consider first word(“What”).

Dr. B. Swaminathan 90
3) Waveform is split based on every letter. Now we will split the sound wave for the letter
W into smaller segments.
Amplitude varies in the sound wave, Analyze the letter 'W'

4) Random weights get assigned to each interconnection between


the input and hidden layers.

Dr. B. Swaminathan 91
5. Weights get multiplied with the inputs, and a bias is added to form the transfer function

Transfer Function transfer the input to output signal(Weighted Sum with bias)

Binary form or a continuous value


Activation function works on Threshold value (once the value crossed signal is
triggered)
The purpose of the activation function is to introduce non-linearity into the output
of a neuron.

Different Activation function are


Dr. B. Swaminathan 92
Linear, Sigmoid, Rectified Linear Unit Activation, …
Algorithms used in Deep Learning
1. Convolution Neural Networks (CNNs)-> image processing and object detection
• Convolution Layer that has several filters to perform the convolution
operation.
• ReLU layer to perform operations on elements. The output is a rectified
feature map.
• Pooling Layer
• The rectified feature map next feeds into a pooling layer. Pooling is a down-
sampling operation that reduces the dimensions of the feature map.
• The pooling layer then converts the resulting two-dimensional arrays from the
pooled feature map into a single, long, continuous, linear vector by flattening it.

Dr. B. Swaminathan 93
2. Long Short Term Memory Networks (LSTMs)
A type of Recurrent Neural Network (RNN)
That can learn and memorize long-term dependencies.
How LSTMs Work?
First, they forget irrelevant parts of the previous state
Next, they selectively update the cell-state values
Finally, the output of certain parts of the cell state

Dr. B. Swaminathan 94
LSTM building block

Dr. B. Swaminathan 95
Memory Pipe
Forget Gate

New Memory Valve Two signs are the forget valve and the
new memory valve

Dr. B. Swaminathan 96
Generate the output for this LSTM unit forget gate (valve) that shuts the old memory:

new memory valve and the new memory

Dr. B. Swaminathan 97
Dr. B. Swaminathan 98
Dr. B. Swaminathan 99
Dr. B. Swaminathan 100
3. Recurrent Neural Networks (RNNs)
The output from the LSTM becomes an input to the current phase and can
memorize previous inputs due to its internal memory. RNNs are commonly used
for image captioning, time-series analysis, natural-language
processing, handwriting recognition, and machine translation.

How RNNs work?


The output at time t-1 feeds into the input at time t.
Similarly, the output at time t feeds into the input at
time t+1.
RNNs can process inputs of any length.
The computation accounts for historical information,
and the model size does not increase with the input
size.
Google’s auto completing feature works:

Dr. B. Swaminathan 101


4. Generative Adversarial Networks (GANs)
Create new data instances that resemble the training data.
GAN has two components:
A generator, which learns to generate fake data, and
A discriminator, which learns from that false information.
How GANs work?
• The discriminator learns to distinguish between the generator’s fake data and the real
sample data.
• During the initial training, the generator produces fake data, and the discriminator
quickly learns to tell that it's false.
• The GAN sends the results to the generator and the discriminator to update the
model.

Dr. B. Swaminathan 102


5. Radial Basis Function Networks (RBFNs)
How RBFNs Work?
Perform classification by measuring the input's similarity to examples from the
training set.
An input vector that feeds to the input layer. They have a layer of RBF neurons.
The function finds the weighted sum of the inputs, and the output layer has one
node per category or class of data.
The neurons in the hidden layer contain the Gaussian transfer functions, which
have outputs that are inversely proportional to the distance from the neuron's
center.
The network's output is a linear combination of the input’s radial-basis functions
and the neuron’s parameters

Dr. B. Swaminathan 103


6. Multilayer Perceptrons (MLPs)
class of feed-forward neural networks with multiple layers of perceptrons.

Dr. B. Swaminathan 104


7. Self Organizing Maps (SOMs)

Dr. B. Swaminathan 105


8. Deep Belief Networks (DBNs)

Dr. B. Swaminathan 106


9. Restricted Boltzmann Machines( RBMs)

Dr. B. Swaminathan 107


10. Autoencoders

Dr. B. Swaminathan 108


Cost and Loss Function: is Quantifier for discrepancy between predicted values and actual
ground-truth values in a given dataset.
Cost Function: Mathematical function that gives total cost to produce a certain number of
units.

Dr. B. Swaminathan 109


• Cost function measures the model’s error on
a group of objects, whereas the loss function
deals with a single data instance.

Dr. B. Swaminathan 110


Loss Function

Dr. B. Swaminathan 111


Dr. B. Swaminathan 112
Automatic AI Image generation using [Link]
Type a few words to generate original images with our AI!
It text type text and generate the IMAGE

Dr. B. Swaminathan 113


NumPy: Multi-dimensional array and matrix processing
Scikit-learn: Supports most of the classic supervised and unsupervised learning algorithm
Pandas: preparing high-level data sets for machine learning and training.
It relies on two types of data structures, one-dimensional (series) and two-dimensional (DataFr

TensorFlow: machine learning and deep learning models are easily developed and evaluated
Seaborn: ML projects because it can generate plots of learning data.
Theano: numerical computation and specifically for machine learning.
Optimize and evaluate mathematical models and matrix calculations.
used by machine learning and deep learning developers
Keras: developing the neural networks for ML models.
PyTorch: natural language processing or computer vision. For Large dense Data.
Matplotlib: Data Visulization tools

Dr. B. Swaminathan 114


Datasets
Load and return the iris dataset
load_iris(*[, return_X_y, as_frame])
(classification).
Load and return the diabetes dataset
load_diabetes(*[, return_X_y, as_frame, scaled])
(regression).
Load and return the digits dataset
load_digits(*[, n_class, return_X_y, as_frame])
(classification).
Load and return the physical exercise
load_linnerud(*[, return_X_y, as_frame])
Linnerud dataset.
Load and return the wine dataset
load_wine(*[, return_X_y, as_frame])
(classification).

Load and return the breast cancer


load_breast_cancer(*[, return_X_y, as_frame])
wisconsin dataset (classification).

Dr. B. Swaminathan 115


Linear Regression Models # Make predictions using the testing set
import [Link] as plt diabetes_y_pred = [Link](diabetes_X_test)
import numpy as np # The coefficients
from sklearn import datasets, linear_model print("Coefficients: \n", regr.coef_)
from [Link] import
mean_squared_error, r2_score # The mean squared error
print("Mean squared error: %.2f" %
# Load the diabetes dataset mean_squared_error(diabetes_y_test,
diabetes_X, diabetes_y = diabetes_y_pred))
datasets.load_diabetes(return_X_y=True)
# Use only one feature # The coefficient of determination: 1 is perfect
diabetes_X = diabetes_X[:, [Link], 2] prediction
print("Coefficient of determination: %.2f" %
# Split the data into training/testing sets r2_score(diabetes_y_test, diabetes_y_pred))
diabetes_X_train = diabetes_X[:-20]
diabetes_X_test = diabetes_X[-20:] # Plot outputs
[Link](diabetes_X_test, diabetes_y_test,
# Split the targets into training/testing sets color="black")
diabetes_y_train = diabetes_y[:-20] [Link](diabetes_X_test, diabetes_y_pred,
diabetes_y_test = diabetes_y[-20:] color="blue", linewidth=3)
# Create linear regression object [Link](())
regr = linear_model.LinearRegression() [Link](())
[Link]()
# Train the model using the training sets Dr. B. Swaminathan 116
[Link](diabetes_X_train, diabetes_y_train)
Dr. B. Swaminathan 117
Dr. B. Swaminathan 118
Dr. B. Swaminathan 119
Dr. B. Swaminathan 120
Dr. B. Swaminathan 121
Dr. B. Swaminathan 122
Dr. B. Swaminathan 123
Dr. B. Swaminathan 124
Dr. B. Swaminathan 125
Dr. B. Swaminathan 126
Dr. B. Swaminathan 127
Dr. B. Swaminathan 128
Dr. B. Swaminathan 129
Dr. B. Swaminathan 130
Dr. B. Swaminathan 131
Dr. B. Swaminathan 132
Dr. B. Swaminathan 133
Dr. B. Swaminathan 134
Dr. B. Swaminathan 135
Dr. B. Swaminathan 136
Dr. B. Swaminathan 137
Dr. B. Swaminathan 138
Dr. B. Swaminathan 139
Dr. B. Swaminathan 140
Dr. B. Swaminathan 141
Dr. B. Swaminathan 142
Dr. B. Swaminathan 143
Dr. B. Swaminathan 144
Dr. B. Swaminathan 145
Dr. B. Swaminathan 146
Dr. B. Swaminathan 147
Dr. B. Swaminathan 148
Dr. B. Swaminathan 149
Dr. B. Swaminathan 150
Dr. B. Swaminathan 151
Dr. B. Swaminathan 152
Dr. B. Swaminathan 153
Dr. B. Swaminathan 154
Dr. B. Swaminathan 155
Dr. B. Swaminathan 156

You might also like