Unit 1 : Revisiting AI Project Cycle &
Ethical Frameworks for AI
1.1 : AI Project Cycle
The AI Project Cycle provides us with an appropriate framework which can lead us towards the goal
(AI Project)
● Problem Scoping Define the project goal by clearly stating the problem you wish to solve.
This involves identifying and analyzing the various parameters that affect the problem to
ensure a clear understanding of the objective.
● Data Acquisition Gather the foundation of your project by collecting large quantities of data
from reliable and authentic sources. This data helps clarify the specific parameters identified
during the scoping phase.
● Data Exploration Convert the acquired data into visual representations (such as graphs,
flow charts, or maps). This visualization allows you to interpret patterns, trends, and
relationships within the data that will influence the model selection.
● Modelling Research, select, and test various AI models to find the most suitable output.
Once the most efficient model is identified, it becomes the base of the project, around which
the final algorithm is developed.
● Evaluation Test the developed model using newly fetched (unseen) data. The results from
this stage are used to assess the model's reliability and implement necessary improvements.
● Deployment Integrate the evaluated model into a real-world environment. This final stage
ensures the solution operates successfully, delivering tangible value and impact to users and
stakeholders.
1.2 : Introduction to AI Domains
AI becomes intelligent according to the training it gets, for training the machine is fed datasets. With
respect to the type of data fed in the AI model, AI models can be broadly categorized into three
domains :
Statistical Data Computer Vision Natural Language Processing
Statistical Data
It is related to data systems and processes, in which the system collects numerous data, maintains
data sets and derives meaning/sense out of them. The information extracted through statistical data
can be used to make a decision about it.
Examples :
1. Price Comparison Websites : PriceGrabber, PriceRunner, Junglee, Shopzilla, DealTime are
some examples of price comparison websites.
Computer Vision (CV)
It depicts the capability of a machine to get and analyse visual information and afterwards predict
some decisions about it. The entire process involves image acquiring, screening, analysing,
identifying and extracting information .The main objective of this domain of AI is to teach machines to
collect information from pixels.
Examples :
1. Agricultural Monitoring : Crop monitoring, Pest Detection, and yield estimation with the help
of drone with cameras which capture aerial images
2. Surveillance Systems : monitor public spaces, buildings and borders
Natural Language Processing (NLP)
It deals with the interaction between computers and humans using natural language. It attempts to
extract information from the spoken and written word using algorithms.
The ultimate objective of NLP is to read, decipher, understand, and make sense of human languages
in a valuable manner
Examples :
1. Email Filters
2. Machine Translation : These systems analyze the structure and semantics of sentences in the
source language and generate equivalent translations in the target language (Google
translate and Microsoft Translator)
1.3 : Ethical Frameworks for AI
Frameworks
● Frameworks are a set of steps that help us in solving problems.
● It provides a step-by-step guide for solving problems in an organized manner.
● It ensures that all relevant factors and considerations are taken into account.
● They serve as a common language for communication and collaboration, facilitating the
sharing of best practices and promoting consistency in problem- solving methodologies
Ethical Frameworks
Ethics are a set of values or morals which help us separate right from wrong.
● Ethical frameworks are frameworks which help us ensure that the choices we make do not
cause unintended harm.
● It provides a systematic approach to navigating complex moral dilemmas by considering
various ethical principles and perspectives.
● By using ethical frameworks, individuals and organisations can make well informed decisions
and promote positive outcomes for all stakeholders involved
Why do we need Ethical Frameworks for AI?
Ethical frameworks ensure that AI makes morally acceptable choices. If we use ethical frameworks
while building our AI solutions, we can avoid unintended outcomes (Caused by bias)
Types of Ethical Frameworks
1. Sector Based Frameworks :
● These are frameworks tailored to specific sectors or industries.
● Bio ethics - Common sector based framework which focuses on ethical
considerations in healthcare
○ It addresses issues such as patient privacy, data security, and the ethical
use of AI in medical decision-making.
● These frameworks may also apply to domains such as finance, education,
transportation, agriculture, governance, and law enforcement
2. Value - Based Frameworks :
● Value-based frameworks focus on fundamental ethical principles and values guiding
decision- making.
● It reflects the different moral philosophies that inform ethical reasoning.
● Value-based frameworks are concerned with assessing the moral worth of actions
and guiding ethical behaviour
I. Rights Based : Prioritizes the protection of human rights and dignity, valuing human
life over other considerations. It emphasizes the importance of respecting individual
autonomy, dignity, and freedoms. In the context of AI, this could involve ensuring
that AI systems do not violate human rights or discriminate against certain groups
II. Utility - Based : Evaluates actions based on the principle of maximizing utility or
overall good, aiming to achieve outcomes that offer the greatest benefit and
minimize harm. It seeks to maximize overall utility or benefit for the greatest number
of people. In AI, this might involve weighing the potential benefits of AI applications
against the risks they pose to society, such as job displacement or privacy concerns
III. Virtue - based : This framework focuses on the character and intentions of the
individuals involved in decision-making. It asks whether the actions of individuals or
organizations align with virtuous principles such as honesty, compassion, and
integrity. In the context of AI, virtue ethics could involve considering whether
developers, users, and regulators uphold ethical values throughout the AI lifecycle.
These classifications provide a structured approach for addressing ethical concerns in AI
development and deployment, ensuring that considerations relevant to specific sectors and
fundamental ethical values are adequately addressed.
Bioethics
Bioethics is an ethical framework used in healthcare and life sciences.
Principles of bioethics :
● Respect for Autonomy
● Do not harm
● Ensure maximum benefit for all
● Give justice
Framework Definition What It Is Questions Proper Definition
Rights-Based The rights-based Focuses on Does this action A framework that prioritizes the
framework focuses ensuring no violate anyone's protection of human rights and
on protecting the person’s rights fundamental dignity, emphasizing respect for
rights, dignity, and are violated and human rights, individual autonomy and freedom
well-being of all all individuals privacy, or dignity? regardless of the outcome.
individuals. are treated fairly. ● The Individual (Protecting the
person's "shield")
Utility-Based The utility-based Evaluates Do the total A framework that evaluates actions
framework evaluates actions by their benefits of this based on maximizing overall utility or
actions based on consequences technology good, aiming to achieve the greatest
overall benefit and to create the outweigh the benefit for the greatest number of
harm for the largest greatest benefit risks/harms to people while minimizing harm.
number of people. for the largest society? ● The Outcome (The Math of
number. Pros vs. Cons)
Virtue-Based The virtue ethics Judges actions Are the developers A framework that focuses on the
framework looks at based on values acting with moral character and intentions of the
moral character, like honesty, integrity, honesty, individuals (or developers) involved,
intentions, and integrity, and responsibility? asking if they are upholding virtues
responsible responsibility, like honesty and integrity.
behaviour. and care. ● The Creator (The
Intent/Character)
Applying in Case studies :
● Rights Based : Look for victims. If even one person is treated unfairly, discriminated against,
or has their privacy stolen, this framework rejects the action.
● Utility Based : Weigh the scale. List the Pros (Benefits) and Cons (Risks). If the Benefits are
heavier, a Utilitarian accepts it. If the risks are heavier, they reject it.
● Virtue Based : Check the intent. Did the company hide data? Did they cut corners? If they
were greedy or dishonest, they failed this framework.
Unit 2 : Advanced Concepts of
Modeling in AI
AI : Artificial Intelligence,refers to any technique that enables computers to mimic human intelligence.
An artificially intelligent machine works on algorithms and data fed to it and gives the desired output.
ML : Machine Learning, enables machines to improve at tasks with experience. The machine learns
from the new data fed to it while testing and uses it for the next iteration. It also takes into account
the times when it went wrong and considers the exceptions too.
DL : Deep Learning, enables software to train itself to perform tasks with vast amounts of [Link]
to the availability of a huge set of data, it is able to train itself with the help of multiple machine
learning algorithms working altogether to perform a specific task.
2.1 Revisiting AI, ML, DL
Machine Learning (ML)
The machine learns from its mistakes and takes them into consideration in the next execution. This
enables machines to improve at tasks with experience. It improvises using its own experiences.
Examples of ML :
1. Object Classification : Identifies and labels the objects present within an image or data
point. It determines the category it belongs to
2. Anomaly Detection : Anomaly detection helps us find the unexpected things hiding in our
data (Blood Pressure)
Deep Learning (DL)
These machines are intelligent enough to develop algorithms for themselves (it trains itself with vast
amounts of data)
Input is given to an ANN, and after processing, the output is generated by the DL block.
Examples of DL:
1. Object Identification : Object classification in deep learning tackles the task of identifying
and labeling objects within an image. It essentially uses powerful algorithms to figure out
what's in a picture and categorize those things.
2. Digit Recognition : Digit recognition in deep learning tackles the challenge of training
computers to identify handwritten digits (0-9) within images.
Common terminologies used with data
What Is Data?
● Data is information in any form
What are data features?
● Columns of the tables are called features
● Some features are special, they are called labels
What are Labels?
● Data Labeling is the process of attaching meaning to data
● Data can be of two types - Labelled and Unlabeled
What do you mean by a training data set?
● The training data set is a collection of examples given to the model to analyze and learn. ( a
set of labeled data is used to train the AI model )
What do you mean by a testing data set?
● The testing data set is used to test the accuracy of the model. (Test is performed without
labeled data and then verify results with labels)
Aspect AI (Artificial Intelligence) ML (Machine Learning) DL (Deep Learning)
Definition Mimics human intelligence using Learns from data to improve over time Learns from large data using neural networks
algorithms
Learning Type Rule-based or data-driven Supervised, unsupervised, reinforcement Mostly supervised with deep neural networks
Input Type Rules + data Labeled or structured data Raw data (images, pixels, etc.)
Output Human-like decisions and actions Predictions based on patterns Complex tasks like object/digit recognition
Structure Algorithm-based Input → ML Model → Output Input → Artificial Neural Network → Output
Complexity Varies (low to high) Moderate High (multi-layered)
Examples Chatbots, game bots, voice Spam detection, product Facial recognition, CAT/DOG ID, handwritten
assistants recommendation, fruit classification digit recognition
2.2 Modeling
AI Modelling refers to developing algorithms, also called models which can be trained to get
intelligent outputs. That is, writing codes to make a machine artificially intelligent.
Rule Based Approach
Rule based approach refers to the AI modelling where the relationship or patterns in data are defined
by the developer. The machine follows the rules or instructions mentioned by the developer, and it
performs its task accordingly. Rule based chatbots are commonly used in customer service
● A limitation of rule-based models is that learning is static.
● Once trained, the machine cannot adapt to changes in the data or rules.
● It does not learn from feedback or new inputs.
● This limitation led to the rise of machine learning, where the system adapts, learns from
feedback, and updates itself based on new data.
Learning Based Approach
A learning-based approach is a method where a computer learns how to do something by looking at
examples or getting feedback, similar to how we learn from [Link] of being explicitly
programmed for a task, the computer learns to perform it by analysing data and finding patterns or
rules on its own
● A learning-based AI model can identify patterns in unlabeled data
● It learns from data by itself and adapts algorithms according to changes in the dataset.
● Unlike rule-based models, it modifies itself with new data and exceptions.
● if the model is trained on a certain type of data, it develops an algorithm around it. As the
data changes over time, the model can adjust itself accordingly to handle new situations and
exceptions.
● Example : A learning based spam email filter
Categories of Machine Learning based Models
Supervised learning :
In a supervised learning model, the dataset which is
fed to the machine is labelled. In other words, the
dataset is known to the person
Supervised Learning is when you make the machine
learn by teaching or training the machine using labeled
data.
The model learns from the training data and then
applies the same knowledge to test data.
Example : Math teacher, coin currency, Social Media
Platforms
Unsupervised Learning :
An unsupervised learning model works on an
unlabelled [Link] data which is fed into the
machine is random and there is a possibility that the
person who is training the model does not have any information regarding it.
Unsupervised Learning is a type of learning without any guidance
Example : Street dogs and A child learning to swim, Grocery shopping, OTT platform
recommendations based in watch history,Detect and flag suspicious or fraudulent bank transactions
● The unsupervised learning models are used to identify relationships, patterns and trends out
of the data which is fed into it. It helps the user in understanding what the data is about and
what are the major features identified by the machine in it.
● Reinforcement Learning :
This learning approach enables the computer to make a series of decisions that maximize a reward metric
for the task without human intervention and without being explicitly programmed to achieve the task.
● Reinforcement learning is a type of learning in which a machine learns to perform a task through a
repeated trial-and-error method
Why is it different?
● Supervised/Unsupervised Learning needs clear data and predefined solutions.
● RL is suited for complex, dynamic, and unknown environments.
● It works well when:
○ You don’t have enough data.
○ The environment is unpredictable or constantly changing
● RL learns through interaction and adaptation, not from pre-existing knowledge.
Examples : Parking a car, Humanoids walking
Subcategories of Supervised Learning Model
Classification Model :
Data is classified according to the labels. This
model works on discrete dataset which means
data need no be continuos
Examples : Grading system students are
classified, Weather, classification of emails
Continuous data → Regression models (predict values).
Non-continuous data → Classification models (predict
categories).
Regression Model : These models work on continuous data. Regression algorithms predict a continuous
value based on the input [Link] values as Temperature, Price, Income, Age, etc.
Examples : Prediction of salary, Prediction of the price of the house, Used Car price Prediction
Subcategories of Unsupervised Learning Model :
Clustering Model :
Clustering is a process of dividing the data points into different groups or clusters based on their similarity
between them.
● Classification uses predefined classes in which objects are assigned whereas clustering finds
similarities between objects and places them in the same cluster and it differentiates them from
objects in other clusters
Example : Music recommendation on spotify and other OTT platforms
Association :
Association Rule is an unsupervised learning method that is used to find interesting relationships between
variables from the database.
Example : Customer recommendation, Association rule
Aspect Supervised Learning Unsupervised Learning
Data Type Uses labelled data Uses unlabelled data
Purpose Predicts outcomes based on known patterns (e.g., Finds unknown patterns or structures in data (e.g.,
price prediction using past data) clustering observations)
Use Case Example Real-world predictions like item pricing Making sense of messy data from experiments
Computational Need Lower – due to clean, structured input Higher – due to messy, unstructured data
Sub Categories of Deep Learning
These machines are intelligent enough to develop algorithms for themselves
Artificial Neural Network (ANN) : Artificial Neural networks are modelled on the human brain and
nervous system. They are able to automatically extract features without input from the programmer.
Every neural network node is essentially a machine learning algorithm. It is useful when solving
problems for which the data set is very large
Convolutional Neural Network (CNN) : Convolutional Neural Network is a Deep Learning algorithm
which can take in an input image, assign importance (learnable weights and biases) to various
aspects/objects in the image and be able to differentiate one from the other
2.3 Artificial Neural Networks
What is a Neural Network?
Neural networks are loosely modelled after how neurons in the human brain behave.
● The key advantage of neural networks is that they are able to extract data features
automatically without needing the input of the programmer.
● A neural network is essentially a system of organizing machine learning algorithms to
perform certain tasks.
● It is a fast and efficient way to solve problems for which the dataset is very large, such as in
images.
How Neural Networks Work :
● It is divided into multiple layers, and each layer is further divided into several blocks called nodes.
Each node has its own task to accomplish which is then passed to the next layer
● Neural Network consists of an input layer, hidden layer which performs computation using
weights and biases on each node and finally, information is passed through these layers to reach
the output layer
● Input layer : First layer, its job is to acquire data and feed it to the neural network. No processing
occurs in the input layer
● Hidden layers : Layers in which the whole processing occurs. These layers are not visible to the
user. Each node of the layer has its own ML algorithm which executes on the data received from
the input layer
● The hidden layer performs computation by means of weights and biases Information passes from
one layer to the other after the value found from this calculation passes through a selected
activation function.
● The process of finding the right output begins with trial and error until the network finally learns.
● With each try, the weights are adjusted based on the error found between the desired output and
the network output.
Input and Output layer is meant for user-interface, There can be multiple hidden layers in a neural network
system and their number depends upon the complexity of the function for which the network has been
configured. Same with the number of nodes
Examples of neural networks : Facial recognition, customer support, chatbot, vegetable price prediction,
etc
How does AI make a Decision?
Step 1: Identify inputs
● List the factors given in the question (e.g., jacket, umbrella, sunny, forecast).
● Convert them to numbers (Yes = 1, No = 0).
Step 2: Assign weights and bias
● Assign bigger weight to more important factors.
● Bias is a fixed number that acts like a threshold.
Note: Weights can be any numbers; they don’t need to add up to anything.
Step 3: Multiply and add
● Multiply each input by its weight.
● Add them all to get the total weighted sum.
Step 4: Apply bias
● Net = (Weighted sum – Bias).
● This is the formula you should write.
Step 5: Check decision rule
● If Net ≥ 0 → Output = 1 (YES). If Net < 0 → Output = 0 (NO).
Step 6: Interpret in words
● Write one sentence: “Since net = X, the decision is YES/NO (e.g., I will go to the park).”
Unit 3 : Evaluating Models
3.1 Importance of Model Evaluation
What is evaluation?
● Model evaluation is the process of using different evaluation metrics to understand a
machine learning model’s performance. An AI model gets better with constructive feedback
Need of model evaluation
It helps you understand its strengths, weaknesses, and suitability for the task at hand. This feedback
loop is essential for building trustworthy and reliable AI systems.
Example : Report card
(You learn → test → assess result → thrive for better results
Training w train data → test model w test data → evaluate → Fine tuning for better performance)
3.2 Splitting the training set data for Evaluation
Train-test split
● The train-test split is a technique for evaluating the performance of a machine learning
algorithm
● It can be used for any supervised learning algorithm
● The procedure involves taking a dataset and dividing it into two subsets: The training dataset
and the testing dataset
● The train-test procedure is appropriate when there is a sufficiently large dataset available
Need of Train-test split
● The train dataset is used to make the model learn
● The input elements of the test dataset are provided to the trained model. The model makes
predictions, and the predicted values are compared to the expected values
● The objective is to estimate the performance of the machine learning model on new data:
data not used to train the model
Overfitting : Overfitting is a modeling error that occurs when a machine learning model learns the
training data too well, including its noise and random fluctuations, instead of capturing the general
patterns. As a result, the model performs very well on the training data but fails to generalize to new,
unseen data, leading to poor performance on test or real-world datasets.
3.3 Accuracy and Error
Accuracy
● Accuracy is an evaluation metric that allows you to measure the total number of predictions a
model gets right.
● The accuracy of the model and performance of the model is directly proportional
Error
● Error can be described as an action that is inaccurate or wrong.
● In Machine Learning, the error is used to see how accurately our model can predict data it
uses to learn new, unseen data.
Error refers to the difference between a model's prediction and the actual outcome. It quantifies how
often the model makes mistakes.
Calculating the accuracy of the AI Model
● Error : Actual - Predicted
● Error rate : error/actual
● Accuracy : 1-Error rate
● Accuracy % : accuracy * 100%
● Accuracy of AI model : mean accuracy
Predicted Actual Error Abs (Actual - Predicted) Error Rate (Error / Accuracy (1 - Accuracy% (Accuracy * 100)%
Actual) Error rate)
Key Points :
● Here the goal is to minimize error and maximize accuracy.
● Real-world data can be messy, and even the best models make mistakes.
● Sometimes, focusing solely on accuracy might not be ideal. For instance, in medical
diagnosis, a model with slightly lower accuracy but a strong focus on avoiding incorrectly
identifying a healthy person as sick might be preferable.
3.4 Evaluation metrics for Classification
What is Classification?
Classification usually refers to a problem where a specific type of class label is the result to be
predicted from the given input field of data
Classification Metrics :
● Confusion Matrix
● Classification accuracy
● Precision
● Recall
● F1 Score
Confusion matrix
The confusion matrix is a handy presentation of the accuracy of a model with two or more classes
The table presents the actual values on the y-axis
and predicted values on the x-axis
True Positive (TP)
Outcome of the model correctly predicting the
positive class (Prediction and reality match)
True Negative (TN)
Outcome of the model correctly predicting the
negative class (Prediction and reality match)
False Positive (FP)
Outcome of the model is wrongly predicting the
negative class as positive class (Prediction was yes,
Reality was no)
False Negative (FN)
Outcome of the model is wrongly predicting the positive class as the negative class (Prediction was
no, Reality was yes)
Accuracy
Classification accuracy is the number of correct predictions made as a ratio of all predictions made
𝐶𝑜𝑟𝑟𝑒𝑐𝑡 𝑝𝑟𝑒𝑑𝑖𝑐𝑡𝑖𝑜𝑛𝑠 𝑇𝑃 + 𝑇𝑁
𝐶𝑙𝑎𝑠𝑠𝑖𝑓𝑖𝑐𝑎𝑡𝑖𝑜𝑛 𝐴𝑐𝑐𝑢𝑟𝑎𝑐𝑦 = 𝑇𝑜𝑡𝑎𝑙 𝑃𝑟𝑒𝑑𝑖𝑐𝑡𝑖𝑜𝑛
= 𝑇𝑃 + 𝑇𝑁 + 𝐹𝑃 + 𝐹𝑁
● It is only suitable when there are an equal number of observations in each class, i.e., a
balanced dataset (which is rarely the case), and that all predictions and prediction errors are
equally important, which is often not the case.
in cases of unbalanced data, we should use other metrics such as Precision, Recall or F1 score
Precision
Precision is the ratio of the total number of correctly classified positive examples and the total
number of predicted positive examples.
𝐶𝑜𝑟𝑟𝑒𝑐𝑡 𝑝𝑜𝑠𝑖𝑡𝑖𝑣𝑒 𝑝𝑟𝑒𝑑𝑖𝑐𝑡𝑖𝑜𝑛𝑠 𝑇𝑃
𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛 = 𝑇𝑜𝑡𝑎𝑙 𝑝𝑜𝑠𝑖𝑡𝑖𝑣𝑒 𝑝𝑟𝑒𝑑𝑖𝑐𝑡𝑖𝑜𝑛
= 𝑇𝑃 + 𝐹𝑃
● The metrics precision is generally used for unbalanced datasets when dealing with the False
Positives becomes important, and the model needs to reduce the FPs as much as possible.
Recall
The recall is the measure of our model correctly identifying True Positives
𝐶𝑜𝑟𝑟𝑒𝑐𝑡 𝑝𝑜𝑠𝑖𝑡𝑖𝑣𝑒 𝑝𝑟𝑒𝑑𝑖𝑐𝑡𝑖𝑜𝑛𝑠 𝑇𝑃
𝑅𝑒𝑐𝑎𝑙𝑙 = 𝑇𝑜𝑡𝑎𝑙 𝑎𝑐𝑡𝑢𝑎𝑙 𝑝𝑜𝑠𝑖𝑡𝑖𝑣𝑒 𝑣𝑎𝑙𝑢𝑒𝑠
= 𝑇𝑃 + 𝐹𝑁
The metrics Recall is generally used for unbalanced dataset when dealing with the False Negatives
becomes important and the model needs to reduce the FNs as much as possible.
F1 Score
F1-Score provides a way to combine both precisions and recall into a single measure that captures
both properties
In those use cases, where the dataset is unbalanced, and we are unable to decide whether FP is
more important or FN, we should use the F1 score as the suitable metric.
2 × 𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛 × 𝑅𝑒𝑐𝑎𝑙𝑙
𝐹1 𝑆𝑐𝑜𝑟𝑒 = 𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛 + 𝑅𝑒𝑐𝑎𝑙𝑙
3.5 Ethical concerns around model evaluation
Unit 5 : Computer Vision
5.1 Introduction
The Computer Vision domain of Artificial Intelligence, enables machines to see through images or
visual data, process and analyse them on the basis of algorithms and methods to analyse actual
phenomena with images
Computer vision is the process of extraction of information from images, text, videos, etc.
A system that can process, analyze and make sense of visual data in the same way as humans do.
Computer Vision Image processing
Computer vision deals with extracting information from the Image processing is mainly focused on processing the raw
input images or videos to infer meaningful information and input images to enhance them or preparing them to do
understanding them to predict the visual input other tasks
Computer Vision is a superset of Image Processing. Image Processing is a subset of Computer Vision.
Examples - Object detection, Handwriting recognition, etc. Examples - Resizing image, Correcting brightness,
Changing tones, etc.
Applications of Computer Vision
1. Facial Recognition : Home security, Schools attendance systems
2. Face Filters : Through the camera, the machine or the algorithm is able to identify the facial
dynamics of the person and applies the facial filter selected.
3. Google's Search by Image : This uses Computer Vision as it compares different features of
the input image to the database of images and gives us the search result while at the same
time analysing various features of the image.
4. Computer Vision in Retail : Retailers can use Computer Vision techniques to track customers'
movements through stores, analyse navigational routes and detect walking patterns.
Inventory Management is another such application. Through security camera image analysis,
a Computer Vision algorithm can generate a very accurate estimate of the items available in
the store
5. Self - Driving Cars
6. Medical Imaging
7. Google Translate App : Uses optical character recognition to see the image and augmented
reality to overlay accurate translation
5.2 Computer Vision Tasks
The various applications of Computer Vision are based on a certain number of tasks that are
performed to get certain information from the input image which can be directly used for prediction
or forms the base for further analysis
Classification : The image Classification problem is the task of
assigning an input image one label from a fixed set of categories.
Classification + Localisation : This is the task that involves both processes of identifying what
object is present in the image and at the same time identifying at what location that object is present
in that image.
Object Detection : Object detection is the process of finding instances (finds and labels) of
real-world objects such as faces, bicycles, and buildings in images or videos. Object detection
algorithms typically use extracted features and learning algorithms to recognize instances of an
object category (Vehicle parking systems and image retrieval)
Instance Segmentation : Instance Segmentation means finding each object in a picture, naming it,
and coloring every pixel that belongs to it.
Basics of Pixels
The word "pixel" means a picture element. They are the smallest unit of information that make up a
picture
Resolution
● The number of pixels in an image is sometimes called the resolution. When the term is used
to describe pixel count, one convention is to express resolution as the width by the height
● Another convention is to express the number of pixels as a single number (Area)
Pixel Value
Each of the pixels that represent an image stored inside a computer has a pixel value that describes
how bright that pixel is, and/or what colour it should be.
The most common pixel format is the byte image, where this number is stored as an 8-bit integer
giving a range of possible values from O to 255.(Typically, zero is to be taken as no colour or black
and 255 is taken to be full colour or white.)
● In computer systems, computer data is in the form of ones and zeros, which we call the
binary system. Each bit in a computer system can have either a zero or a one. Since each
pixel uses 1 byte of an image, which is equivalent to 8 bits of data. Since each bit can have
two possible values which tell us that the 8 bits can have 255 possibilities of values that
starts from 0 and ends at 255
Each pixel = 1 byte = 8 bits = Binary system hence 2 values = 2^8 which 256 values (0-255)
Greyscale Images
Grayscale images are images that have a range of shades of gray without apparent colour.
Intermediate shades of gray are represented by equal brightness levels of the three primary colours
● A grayscale has each pixel of size 1 byte having a single plane of 2d array of pixels. The size
of a grayscale image is defined as the Height x Width of that image.
A grayscale image is an image where each pixel represents only the intensity of light, not color. This
means: There are no Red, Green, or Blue [Link] image is stored as a 2D array
(Height x Width). There is one single plane (channel) in memory.
RGB Images
Every RGB image is stored in the form of three different channels called the R channel, G channel,
and the B channel.
● Each plane separately has many pixels with each pixel value varying from O to 255. All the
three planes when combined form a colour image. This means that in an RGB image, each
pixel has a set of three different values which together give colour to that particular pixel.
● Pixel = [R, G, B] → e.g. [125, 200, 90]
● An RGB image is stored in the computer’s memory as a 3D array of numbers
→ The first dimension is the height (number of rows of pixels),
→ The second dimension is the width (number of columns),
→ The third dimension has 3 elements per pixel — one each for Red, Green, and Blue.
● Stored as: 3 separate 2D arrays → [3][Height][Width]
● When you separate an RGB image into its Red, Green, and Blue channels, each channel is
just a 2D grid of numbers from 0 to 255. That’s exactly what a grayscale image is too:
A 2D array where each pixel has just one intensity value (0 = black, 255 = white)
So when we look at just one channel (like the Red channel), we are looking at a grayscale
image — not because the image itself has turned grayscale, but because: its just the intensity
Conclusion : An RGB image is stored as a 3D array in memory, with dimensions [Height × Width ×
3].Each pixel has 3 values—one for Red, one for Green, and one for Blue—each ranging from 0 to
255 and stored using 1 [Link] three channels together define the full color of every pixel.
Unit 6 : Natural Language Processing
A natural language is a human language, such as French, Spanish, English, Japanese, etc.
Features of Natural Languages
● They are governed by set rules that include syntax, lexicon, and semantics.
● All natural languages are redundant, i.e., the information can be conveyed in multiple ways.
● All natural languages change over time.
Thus, in natural language, it is important to understand that a word can have multiple meanings and
the meanings fit into the statement according to the context of it.
Computer Language
Computer languages are languages used to interact with a computer, such as Python, C++, Java,
HTML, etc.
Computers require a specific set of instructions to understand human input called programs. To talk
to a computer, we convert natural language into a language that the computer understands, This is
done by NLP
Why is NLP important?
Computers can only process electronic signals in the form of binary language. Natural Language
Processing facilitates this conversion to digital form from the natural form.[ Thus, the whole purpose
of NLP is to make communication between computer systems and humans possible ]. This includes
creating different tools and techniques that facilitate better communication of intent and context.
6.2 Applications of Natural Language Processing
1. Voice Assistants : Voice assistants take our natural speech, process it, and give us an
output. They use NLP to understand natural language and execute tasks effectively
2. Autogenerated captions : Captions are generated by turning natural speech into text in
real-time. It is a valuable feature for enhancing the accessibility of video content.
3. Language Translation : It incorporates the generation of translation from another language.
[This involves the conversion of text or speech from one language to another, facilitating
cross-linguistic communication and fostering global connectivity.]
4. Sentiment Analysis : Sentiment Analysis is a tool to express an opinion, whether the
underlying sentiment is positive, negative, or neutral. Customer sentiment analysis helps in
the automatic detection of emotions when customers interact with the products or services
5. Text Classification : Text classification is a tool which classifies a sentence or document
category-wise. This process classifies the raw texts into predefined groups or categories.
6. Keyword Extraction : Keyword extraction is a tool that automatically extracts the most
used, important words and expressions from a text. It can give valuable insights into
people’s opinions (or customer service) about any business on social media
6.3 Stages of NLP
1. Lexical Analysis : Lexicon stands for a collection of the various words and phrases used in
a language. NLP starts with identifying the structure of input words. It is the process of
dividing a large chunk of words into structural paragraphs, sentences, and words
2. Syntactic Analysis/Parsing : It is the process of checking the grammar of sentences and
phrases. It forms a relationship among words and eliminates logically incorrect sentences
3. Semantic Analysis : The input text is now checked for meaning, and every word and phrase
is checked for meaningfulness
4. Discourse Integration : It is the process of forming the story of the sentence (relationship
b/w its preceding and succeeding sentences)
5. Pragmatic Analysis : Pragmatic analysis is the process of interpreting the intended meaning
of a sentence based on real-world context, speaker intention, social norms, and common
sense knowledge.
Stage What it Does Easy Way to Remember
1. Lexical Analysis Breaks text into words, phrases, and sentences using "Split into words and check structure."
vocabulary rules (lexicon).
2. Syntactic Analysis Checks grammar and sentence structure. Removes "Is it grammatically correct?"
incorrect ones.
3. Semantic Analysis Finds literal meaning of words and phrases in context. "What do the words mean?"
4. Discourse Integration Connects current sentence to previous and next ones "How does this fit in the conversation?"
for context.
5. Pragmatic Analysis Understands intended meaning using real-world "What does the speaker really want to say?"
knowledge and tone.
6.4 Chatbots
A chatbot is a computer program that's designed to simulate human conversation through voice
commands or text chats or both. It can learn over time how to best interact with humans.
Examples : Eliza, Mitsuku, Cleverbot and Singtel
Script Bot Smart Bot
Easy to make Flexible and Powerful
Work around script which is programmed in them Work on bigger databases and other resources
directly
They are free and easy to integrate to a messaging Learn more with data
platform
No or little language processing skills Coding is required to take this up on board
Limited functionality Wide functionality
6.5 Text Processing
Computers understand only numerical data, so human language must be converted to numbers. The
first step is Text Normalisation. Text Normalisation cleans and simplifies the text to reduce its
complexity.
Text Normalisation
Sentence Segmentation : the whole corpus is divided into sentences, each sentence is taken as a
different data
Tokenization : After segmenting the sentences, each sentence is then further divided into tokens.
Tokens is a term used for any word or number or special character occurring in a sentence
Removing Stop words, Special Characters and Numbers : The tokens which are not necessary
are removed from the token list. Stop words are the words which occur very frequently but don't add
any value.
Converting Text to a Common Case : Converting it into a similar case. This ensures that the case
sensitivity of the machine does not consider the same words as different just because of different
cases.
Stemming : The remaining words are reduced to their root words. (affixes are removed). Stemming
does not take into account whether the stemmed word is meaningful or not. It just removes the
affixes hence it is faster.
Lemmatization : Stemming and lemmatization both aim to remove affixes from words. Their goal is
similar, but they differ in output quality. Lemmatization returns a meaningful word (called a lemma)
after removal. Because of this, lemmatization is more accurate but also slower than stemming.
Bag of Words
Bag of Words (BoW) is an NLP model used to extract features from text. It helps in preparing text
data for machine learning algorithms. BoW counts the occurrences (frequency) of each word in the
text. It then builds a vocabulary for the entire text corpus.
1. Text Processing
2. Create a dictionary
3. Create document vectors
4. Create document vectors for all documents
TDIDF : Term Frequency and Inverse Document Frequency
TFIDF helps is identify the value of each word
Term Frequency : Is the frequency of a word in one document
Inverse Document Frequency : Document Frequency is the number of documents in which the
word occurs irrespective of how many times it has occurred in those documents.
Talking about inverse document frequency, we need to put the document frequency in the
denominator while the total number of documents is the numerator
𝑇𝑜𝑡𝑎𝑙 𝑛𝑜. 𝑜𝑓 𝑑𝑜𝑐𝑢𝑚𝑒𝑛𝑡𝑠
𝐼𝐷𝐹 = 𝐷𝑜𝑐𝑢𝑚𝑒𝑛𝑡 𝐹𝑟𝑒𝑞𝑢𝑒𝑛𝑐𝑦
𝑇𝐹𝐼𝐷𝐹(𝑊) = 𝑇𝐹(𝑊) × 𝑙𝑜𝑔(𝐼𝐷𝐹(𝑊))
Now the words have been converted into numbers
● For a word to have a high TFIDF value, the word needs to have a high term frequency but
less document frequency which shows that the word is important for one document but is
not a common word for all documents.
● These values help the computer understand which words are to be considered while
processing the natural language. The higher the value, the more important the word is for a
given corpus.
● Words that occur in all the documents with high term frequencies have the lowest values and
are considered to be the stop words
Applications of TFIDF