ADVANCED MACHINE LEARNING AND DEEP LEARNING
Syllabus
Convolutional Neural Networks: The operation, Pooling, Convolution and Pooling as an
Infinitely strong prior, Variants of the basic functions, efficient algorithms, Random or
Unsupervised Features, Neuroscientific Basis for Convolutional Networks.
Books
1. Artificial Intelligence Illuminated - Ben Coppin
2. Deep Learning - Ian Goodfellow, Yoshua Bengio, Aaron Courville
3. Fundamentals of Deep Learning – Nikhil Budama
4. Neural Networks and Deep Learning – Charu Aggarwal
5. Hands-on Deep Learning Algorithms with Python – Sudharsan Ravichandran
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Introduction to Machine Learning
Modulue-3
Convolutional Neural Networks
Dr. Veerabhadrappa S T
Associate Professor
Department of Electronics & Communication,
JSS Academy of Technical Education, Bengaluru
veerabhadrappast@[Link]
Convolutional Neural Networks
➢ A Convolutional Neural Network (CNN) is a type of Deep Learning neural network architecture
commonly used in Computer Vision. Computer vision is a field of Artificial Intelligence that
enables a computer to understand and interpret the image or visual data
➢ Convolutional Neural Network (CNN) is the extended version of artificial neural networks
(ANN) which is predominantly used to extract the feature from the grid-like matrix dataset.
➢ These networks preserve the spatial structure of the problem and were developed for object
recognition tasks such as handwritten digit recognition.
➢ They are popular because people are achieving state-of-the-art results on difficult computer
vision and natural language processing tasks.
➢ Given a dataset of gray scale images with the standardized size of 32 × 32 pixels each, a
traditional feedforward neural network would require 1,024 input weights (plus one bias).
➢ This is fair enough, but the flattening of the image matrix of pixels to a long vector of pixel
values looses all of the spatial structure in the image.
➢ Unless all of the images are perfectly resized, the neural network will have great difficulty with
the problem.
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Convolutional Neural Networks
➢ Convolutional Neural Networks expect and preserve the spatial relationship
between pixels by learning internal feature representations using small squares
of input data.
➢ Feature are learned and used across the whole image, allowing for the objects in
the images to be shifted or translated in the scene and still detectable
by the network.
➢ It is this reason why the network is so useful for object recognition in
photographs, picking out digits, faces, objects and so on with varying orientation.
In summary, below are some of the benefits of using convolutional
neural networks:
➢ They use fewer parameters (weights) to learn than a fully connected network.
➢ They are designed to be invariant to object position and distortion in the scene.
➢ They automatically learn and generalize features from the input domain.
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Convolutional Neural Networks
There are three types of layers in a Convolutional Neural Network
1. Convolutional Layers.
2. Pooling Layers.
3. Fully-Connected Layers.
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Convolutional Neural Networks
Convolutional layers are comprised of filters and feature maps.
• Filters
✓ The filters are essentially the neurons of the layer. They have both weighted
inputs and generate an output value like a neuron.
✓ The input size is a fixed square called a patch or a receptive field. If the
convolutional layer is an input layer, then the input patch will be pixel values.
✓ If they deeper in the network architecture, then the convolutional layer will take
input from a feature map from the previous layer.
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Convolutional Neural Networks
Feature Maps
✓ The feature map is the output of one filter applied to the previous layer. A given filter is
drawn across the entire previous layer, moved one pixel at a time.
✓ Each position results in an activation of the neuron and the output is collected in the
feature map.
✓ You can see that if the receptive field is moved one pixel from activation to activation,
then the field will overlap with the previous activation by (field width - 1) input values.
✓ The distance that filter is moved across the input from the previous layer each
activation is referred to as the stride.
✓ If the size of the previous layer is not cleanly divisible by the size of the filters receptive
field and the size of the stride then it is possible for the receptive field to attempt to
read off the edge of the input feature map.
✓ In this case, techniques like zero padding can be used to invent mock inputs with zero
values for the receptive field to read.
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Convolutional Neural Networks
Pooling Layers
✓ The pooling layers down-sample the previous layers feature map.
✓ Pooling layers follow a sequence of one or more convolutional layers and are
intended to consolidate the features learned and expressed in the previous layers
feature map.
✓ As such, pooling may be consider a technique to compress or generalize feature
representations and generally reduce the overfitting of the training data by the
model.
✓ They too have a receptive field, often much smaller than the convolutional layer.
✓ Also, the stride or number of inputs that the receptive field is moved for each
activation is often equal to the size of the receptive field to avoid any overlap.
✓ Pooling layers are often very simple, taking the average or the maximum of the
input value in order to create its own feature map. Convolutional Neural Network
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Convolutional Neural Networks
Fully Connected Layers
✓ Fully connected layers are the normal flat feedforward neural network layer.
✓ These layers may have a nonlinear activation function or a softmax activation in
order to output probabilities of class predictions.
✓ Fully connected layers are used at the end of the network after feature extraction
and consolidation has been performed by the convolutional and pooling layers.
✓ They are used to create final nonlinear combinations of features and for making
predictions by the network.
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Convolutional Neural Networks
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Convolutional Neural Networks- Operation
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Convolutional Neural Networks- Operation
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Convolutional Neural Networks- Operation
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Convolutional Neural Networks- Operation
Basic convolution slides
the filter on the image
one pixel at a time
• Stride = 1
• Can define a different
stride
• Hyperparameter
• Stride reduces the
number of
multiplications
• Subsamples the image
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Convolutional Neural Networks- Operation
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Convolutional Neural Networks- Operation
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Convolutional Neural Networks- Operation
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Convolutional Neural Networks- Operation
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Convolutional Neural Networks- Operation
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Convolutional Neural Networks- Operation
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Pooling
➢ Pooling is a fundamental operation in Convolutional Neural Networks (CNNs) that
plays a crucial role in down sampling feature maps while retaining important
information.
➢ It reduces the spatial dimensions of feature maps while retaining essential
information.
➢ It helps in controlling the model’s complexity, reducing overfitting, and improving
computational efficiency by reducing the number of parameters and computation
required in subsequent layers.
➢ The pooling operation involves sliding a two-dimensional filter over each channel of
feature map and summarising the features lying within the region covered by the
filter.
➢ Types of Pooling Layers:
➢ Max Pooling
➢ Average Pooling
➢ Global Pooling
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Pooling
Max pooling are often sharper than those derived from average pooling.
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Pooling
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Pooling
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Infinitely strong prior
Prior: probability distribution over the parameters of the model that encode our
believe about what models are reasonable.
•A weak prior: high entropy. Eg: Gaussian with high variance. Such prior allows the
data to move more or less freely.
•A strong prior: low entropy. Eg: Gaussian with low variance. Such piror plays a more
active role in determine where the parameters end up.
Over all, infinitely strong layer:
•Conv layer: the function the layer should learn contains only local interaction and is
equivariant to translation
•Pooling layer: each unit should be invariant to small translation
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Infinitely strong prior
➢ Convolution and Pooling acts as a Infinitely strong prior Infinitely
strong prior places zero probability on some parameters and
says that these parameter values are completely forbidden
➢ In another word, these parameters don’t need to be learned.
➢ Like any strong priors Convolution and pooling can cause
underfitting.
➢We should only compare convolutional models to other
convolutional models in benchmarks of statistical learning
performance.
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Infinitely strong prior
➢ We can imagine a convolutional net as being similar to a fully connected net,
but with an Infinitely strong prior over its weights.
➢ This Infinitely strong prior says that the weights for one hidden unit must be
identical to the weights of its neighbor, but shifted in space.
➢ The prior also says that the weights must be zero, except for in the small,
spatially contiguous receptive field assigned to that hidden
unit.
➢ Overall, we can think of the use of convolution as introducing an infinitely
strong prior probability distribution over the parameters of a layer
➢ Likewise, the use of pooling is an Infinitely strong prior
that each unit should be invariant to small translations.
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Variants of the Basic Convolution
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Variants of the Basic Convolution
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Variants of the Basic Convolution
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Variants of the Basic Convolution
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Variants of the Basic Convolution
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Variants of the Basic Convolution
Comparison of local connections, convolution, and full
connections.
(Top)A locally connected layer with a patch size of two
pixels. Each edge is labeled with a unique letter to show that
each edge is associated with its own weight parameter.
(Center)A convolutional layer with a kernel width of two
pixels. This model has exactly the same connectivity as the
locally connected layer. The difference lies not in which units
interact with each other, but in how the parameters are
shared. The locally connected layer has no parameter
sharing. The convolutional layer uses the same two weights
repeatedly across the entire input, as indicated by the
repetition of the letters labeling each edge.
(Bottom)A fully connected layer resembles a locally
connected layer in the sense that each edge has its own
parameter (there are too many to label explicitly with letters
in this diagram). However, it does not have the restricted
connectivity of the locally connected layer.
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Variants of the Basic Convolution
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Variants of the Basic Convolution
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Variants of the Basic Convolution
A convolutional network with the first two output
channels connected to only the first two input
channels, and the second two output channels
connected to only the second two input channels.
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Variants of the Basic Convolution
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Variants of the Basic Convolution
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Variants of the Basic Convolution
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Variants of the Basic Convolution
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
The Neuroscientific Basis for Convolutional Networks
• In this simplified view, we focus on a part of the brain called V1, also known
as the primary visual cortex.
• V1 is the first area of the brain that begins to perform significantly advanced
processing of visual input.
• In this cartoon view, images are formed by light arriving in the eye and
stimulating the retina, the light-sensitive tissue in the back of the eye.
• The neurons in the retina perform some simple preprocessing of the image but
do not substantially alter the way it is represented.
• The image then passes through the optic nerve and a brain region
called the lateral geniculate nucleus.
• The main role, as far as we are concerned here, of both of these anatomical
regions is primarily just to carry the signal from the eye to V1, which is located at
the back of the head.
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
The Neuroscientific Basis for Convolutional Networks
A convolutional network layer is designed to capture three properties of V1:
1. V1 is arranged in a spatial map. It actually has a two-dimensional structure mirroring the
structure of the image in the retina. For example, light arriving at the lower half of the retina affects
only the corresponding half of V1. Convolutional networks capture this property by having their
features defined in terms of two dimensional maps.
2. V1 contains many simple cells. A simple cell’s activity can to some extent be characterized by a
linear function of the image in a small, spatially localized receptive field. The detector units of a
convolutional network are designed to emulate these properties of simple cells.
3. V1 also contains many complex cells. These cells respond to features that are similar to those
detected by simple cells, but complex cells are invariant to small shifts in the position of the feature.
This inspires the pooling units of convolutional networks. Complex cells are also invariant to some
changes in lighting that cannot be captured simply by pooling over spatial locations.
These invariances have inspired some of the cross-channel pooling strategies in convolutional
networks, such as maxout units
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
The Neuroscientific Basis for Convolutional Networks
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
The Neuroscientific Basis for Convolutional Networks
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru
The Neuroscientific Basis for Convolutional Networks
Veerabhadrappa S T, Department of Electronics & Communication, JSS Academy of Technical Education, Bengaluru