Introduction to Machine Learning Concepts
Introduction to Machine Learning Concepts
Module -1
Introduction & Understanding Data – 1
Machine Learning.
In the real world, we are surrounded by humans who can learn everything from their experiences
with their learning capability, and we have computers or machines which work on our instructions.
But can a machine also learn from experiences or past data like a human does? So here comes the
role of Machine Learning.
programmed, aids in making predictions or decisions with the assistance of sample historical data,
or training data. . A machine can learn if it can gain more data to improve its performance.
“Machine learning is the field of study that gives the computers ability to learn without being
explicitly programmed.”
Machine learning has become so popular because of three reasons:
1. High volume of available data to manage:
Big companies such as Facebook, Twitter, and YouTube generate huge amount of data that grows at
a phenomenal rate. It is estimated that the data approximately gets doubled every year.
2. Second reason is that the cost of storage has reduced. The hardware cost has also dropped.
Therefore, it is easier now to capture, process, store, distribute, and transmit the digital information.
3. Third reason for popularity of machine learning is the availability of complex algorithms now.
Especially with the advent of deep learning, many algorithms are available for machine learning.
Before starting the machine learning journey, let us establish these terms - data, information,
knowledge, intelligence, and wisdom. It makes sense to include intelligence as the unit of analysis
in the Data Information Knowledge Wisdom hierarchy since intelligence has inseparable
relationships with knowledge and wisdom. Hence the DIKW acronym to DIKIW. A knowledge
pyramid as shown in Fig
Data refers to raw, unprocessed facts and figures without context. It is the foundation for all
subsequent layers but holds limited value in isolation.
The Processed & Analyzed Data is Information.
Information is organized, structured, and contextualized data. Information is useful for answering
basic questions like "who," "what," "where," and "when."
Ex. Sales data can be analysed to extract information like which is the fast selling producct.
Condensed information is Knowledge.
Knowledge is the result of analyzing and interpreting information to uncover patterns, trends, and
relationships. It provides an understanding of "how" and "why" certain phenomena occur.
Ex. The histirical patterns and future trends obtained in the above sales data can be called
knowledge.
Unless Knowledge is not extracted , data is no [Link] is useful only when it is put into
action. Applied knowlede is Intelligence
The ultimate objective of knowledge pyramid is wisdom that represents the maturity of mind that is
so far, exhibited only by humans.
Wisdom is the ability to make well-informed decisions and take effective action based on
understanding of the underlying knowledge.
How does Machine Learning work
As humans take decisions based on an experience, computers make models based on extracted
patterns in the input data and then use these data-filled models for prediction and to take decisions.
For computers, the learnt model is equivalent to human experience
Figure 1.2: (a) A Learning System for Humans (b) A Learning System for Machine Learning
A machine learning system builds prediction models, learns from previous data, and predicts the
output of new data whenever it receives it. The amount of data helps to build a better model that
accurately predicts the output, which in turn affects the accuracy of the predicted output.
Let we have a complex problem in which we need to make predictions. Instead of writing code, we
just need to feed the data to generic algorithms, which build the logic based on the data and predict
the output. Our perspective on the issue has changed as a result of machine learning.
The learning program summarizes the raw data in a model. Formally stated, a model is an explicit
description of patterns within the data in the form of:
1. Mathematical equation
2. Relational diagrams like trees/graphs
3. Logical if/else rules, or
4. Groupings called clusters
The Machine Learning algorithm's operation is depicted in the following block diagram:
Data Analytics Another branch of data science is data analytics. It aims to extract useful
knowledge from crude data. There are different types of analytics(Descriptive, Diagnostic,
Predictive & Perspective. Ex- Predictive data analytics is used for making predictions.) Machine
learning is closely related to this branch of analytics and shares almost all algorithms.
Pattern Recognition It is an engineering field. It uses machine learning algorithms to extract the
features for pattern analysis and pattern classification. One can view pattern recognition as a
specific application of machine learning
• Machine learning is much similar to data mining as it also deals with the huge amount of
the data.
Need for ML
The demand for machine learning is steadily rising. Because it is able to perform tasks that are too
complex for a person to directly implement, machine learning is required. Humans are constrained
by our inability to manually access vast amounts of data; as a result, we require computer systems,
which is where machine learning comes in to simplify our lives.
By providing them with a large amount of data and allowing them to automatically explore the data,
build models, and predict the required output, we can train machine learning algorithms. The cost
function can be used to determine the amount of data and the machine learning algorithm's
performance. We can save both time and money by using machine learning.
Presently, AI is utilized in self-driving vehicles, digital misrepresentation identification, face
acknowledgment, and companion idea by Facebook, and so on. Different top organizations, for
example, Netflix and Amazon have constructed AI models that are utilizing an immense measure of
information to examine the client interest and suggest item likewise.
Following are some key points which show the importance of Machine Learning:
• Rapid increment in the production of data
• Solving complex problems, which are difficult for a human
• Decision making in various sector including finance
• Finding hidden patterns and extracting useful information from data.
>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>
Labelled and Unlabelled Data
Data is a raw fact. It is represented in the form of a table. Data also can be referred to as a data
point, sample, or an example. Each row of the table represents a data point. Features are attributes
or characteristics of an object. Normally, the columns of the table are attributes. Out of all attributes,
one attribute is important and is called a label. Label is the feature that we aim to predict. Thus,
there are two types of data – labelled and unlabelled.
Labelled Data To illustrate labelled data, let us take one example dataset called Iris flower dataset
[Link]. Length of Width of Length of Width of Class
Petal Petal Sepal Sepal
1. 5.5 4.2 1.4 0.2 Setosa
2. 7 3.2 4.7 1.4 Versicolor
3. 7.3 2.9 6.3 1.8 Virginica
The dataset has 50 samples of Iris – with four attributes, length and width of sepals and petals. The
target variable is called class. There are three classes – Iris setosa, Iris virginica, and Iris versicolor.
The partial data of Iris dataset is shown in Table
A dataset need not be always numbers. It can be images or video frames. Deep neural networks
can handle images with labels.
The following Figure deep neural network takes images of dogs and cats with labels for
classification.
Dog
Cat
3. Reinforcement learning
1) Supervised Learning
In supervised learning, labeled sample data are provided to the machine learning system for
training, and the system then predicts the output based on the training data. The system uses
labeled data to build a model that understands the datasets and learns about each one. After the
training and processing are done, we test the model with sample data to see if it can accurately
predict the output.
The mapping of the input data to the output data is the objective of supervised learning. The
managed learning depends on oversight, and it is equivalent to when an understudy learns things in
the management of the educator. Spam filtering is an example of supervised learning.
Supervised learning can be grouped further in two categories of algorithms:
• Classification
• Regression
Classification
Classification is a supervised learning method. The input attributes of the classification algorithms
are called independent variables. The target attribute is called label or dependent variable.
The relationship between the input and target variable is represented in the form of a structure
which is called a classification model. So, the focus of classification is to predict the ‘label’ that is
in a discrete form (a value from the set of finite values). An example is shown in Figure where a
classification algorithm takes a set of labelled data images such as dogs and cats to construct a
model that can later be used to classify an unknown test image data.
In classification, learning takes place in two stages. During the first stage, called training stage,
the learning algorithm takes a labelled dataset and starts learning. After the training set, samples
are processed and the model is generated. In the second stage, the constructed model is tested with
test or unknown sample and assigned a label. This is the classification process.
Similarly, in the case of Iris dataset, if the test is given as (6.3, 2.9, 5.6, 1.8, ?), the classification
will generate the label for this. This is called classification. One of the examples of classification is
Image recognition, which includes classification of diseases like cancer, classification of plants, etc.
Classification models can also be classified as generative models and discriminative models.
Generative models deal with the process of data generation and its distribution. Probabilistic models
are examples of generative models. Discriminative models do not care about the generation of data.
Instead, they simply concentrate on classifying the given data.
Some of the key algorithms of classification are:
• Decision Tree
• Random Forest
• Support Vector Machines
• Naïve Bayes
• Artificial Neural Network and Deep Learning networks like CNN
Regression models
Regression models, unlike classification algorithms, predict continuous variables like price.
In other words, it is a number. A fitted regression model is shown in Figure 1.8 for a dataset that
represent weeks input x and product sales y.
2) Unsupervised Learning
Unsupervised learning is a learning method in which a machine learns without any supervision.
The training is provided to the machine with the set of data that has not been labeled, classified, or
categorized, and the algorithm needs to act on that data without any supervision. The goal of
unsupervised learning is to restructure the input data into new features or a group of objects with
similar patterns
In unsupervised learning, we don't have a predetermined result. The machine tries to find useful
insights from the huge amount of data. It can be further classifieds into two categories of
algorithms:
• Clustering
• Association Clustering
Cluster analysis is an example of unsupervised learning. It aims to group objects into disjoint
clusters or groups. Cluster analysis clusters objects based on its attributes. All the data objects of the
partitions are similar in some aspect and vary from the data objects in the other partitions
significantly.
Some of the examples of clustering processes are — segmentation of a region of interest in an
image, detection of abnormal growth in a medical image, and determining clusters of signatures in a
gene database.
An example of clustering scheme is shown in Figure 1.9 where the clustering algorithm takes a set
of dogs and cats images and groups it as two clusters-dogs and cats. It can be observed that the
samples belonging to a cluster are similar and samples are different radically across clusters.
Example of Clustering Scheme. some of the key clustering algorithms are:
• k-means algorithm
• Hierarchical algorithms
Dimensionality reduction
Dimensionality reduction algorithms are examples of unsupervised algorithms. It takes a higher di-
mension data as input and outputs the data in lower dimension by taking advantage of the variance
of the data. It is a task of reducing the dataset with few features without losing the generality
Differences between Supervised and Unsupervised Learning
Sl. No. Supervised Learning Unsupervised Learning
1. There is a supervisor component No supervisor component
2. Uses Labelled data Uses Unlabelled data
3. Assigns categories or labels Performs grouping process such that
similar objects will be in one cluster
2) Semi-supervised Learning
There are circumstances where the dataset has a huge collection of unlabelled data and some
labelled data. Labelling is a costly process and difficult to perform by the humans. Semi-supervised
algorithms use unlabelled data by assigning a pseudo-label. Then, the labelled and pseudo-labelled
dataset can be combined.
3) Reinforcement Learning
Reinforcement learning mimics human beings. Like human beings use ears and eyes to perceive the
world and take actions, reinforcement learning allows the agent to interact with the environment to
get rewards. The agent can be human, animal, robot, or any independent program. The rewards
enable the agent to gain experience. The agent aims to maximize the reward.
The reward can be positive or negative (Punishment). When the rewards are more, the behaviorgets
reinforced and learning becomes possible.
The robotic dog, which automatically learns the movement of his arms, is an example of
Reinforcement learning.
Here the experience helps in constructing a model. It can be said in summary, compared to
supervised learning, there is no supervisor or labelled dataset. Many sequential decisions need to be
taken to reach the final decision. Therefore, reinforcement algorithms are reward-based, goal-
oriented algorithms.
Is a model for this test data be multiplication? That is, y = x 1 × x 2 . Well! It is true! But, this is
equally true that y may be y = x 1 ÷ x 2 , or y = x 1x2 . So, there are three functions that fit the data
2 Huge data – This is a primary requirement of machine learning. Availability of a quality data is a
challenge. A quality data means it should be large and should not have data problems such as
missing data or incorrect data.
3. High computation power – With the availability of Big Data, the computational resource requi-
rement has also increased. Systems with Graphics Processing Unit (GPU) or even Tensor
Processing Unit (TPU) are required to execute machine learning algorithms. Also, machine learning
tasks have become complex and hence time complexity has increased, and that can be solved only
with high computing power.
4. Complexity of the algorithms – The selection of algorithms, describing the algorithms,
application of algorithms to solve machine learning task, and comparison of algorithms have
become necessary for machine learning or data scientists now. Algorithms have become a big topic
of discussion and it is a challenge for machine learning professionals to design, select, and evaluate
optimal algorithms.
5. #Bias/Variance – Variance is the error of the model. This leads to a problem called bia s /
variance tradeoff. A model that fits the training data correctly but fails for test data, in general lacks
generalization, is called overfitting. The reverse problem is called underfitting where the model fails
for training data but has good generalization. Overfitting and underfitting are great challenges for
machine learning algorithms.
1. Understanding the business – This step involves understanding the objectives and requ-
irements of the business organization. Generally, a single data mining algorithm is enough for
giving the solution. This step also involves the formulation of the problem statement for the data
mining process.
2. Understanding the data – It involves the steps like data collection, study of the characteristics of
the data, formulation of hypothesis, and matching of patterns to the selected hypothesis.
3. Preparation of data – This step involves producing the final dataset by cleaning the raw data
and preparation of data for the data mining process. The missing values may cause problems during
both training and testing phases. Missing data forces classifiers to produce inaccurate results.
This is a perennial problem for the classification models. Hence, suitable strategies should be
adopted to handle the missing data.
4. Modelling – This step plays a role in the application of data mining algorithm for the data to
obtain a model or pattern.
5. Evaluate – This step involves the evaluation of the data mining results using statistical analysis
and visualization methods. The performance of the classifier is determined by evaluating the
accuracy of the classifier. The process of classification is a fuzzy issue.
For example, classification of emails requires extensive domain knowledge and requires domain
experts. Hence, performance of the classifier is very crucial.
6. Deployment – This step involves the deployment of results of the data mining algorithm to
improve the existing process or for a new situation.
Key Terms:
Machine Learning – A branch of AI that concerns about machines to learn automatically without
being explicitly programmed.
Data – A raw fact.
Model – An explicit description of patterns in a data.
Experience – A collection of knowledge and heuristics in humans and historical training data in case
of machines.
Predictive Modelling – A technique of developing models and making a prediction of unseen data.
Deep Learning – A branch of machine learning that deals with constructing models using neural
networks.
Data Science – A field of study that encompasses capturing of data to its analysis covering all stages
of data management.
Data Analytics – A field of study that deals with analysis of data.
Big Data – A study of data that has characteristics of volume, variety, and velocity.
Statistics – A branch of mathematics that deals with learning from data using statistical methods.
Hypothesis – An initial assumption of an experiment.
Learning – Adapting to the environment that happens because of interaction of an agent with the
environment.
Label – A target attribute.
Labelled Data – A data that is associated with a label.
Unlabelled Data – A data without labels.
Supervised Learning – A type of machine learning that uses labelled data and learns with the help of
a supervisor or teacher component.
Classification Program – A supervisory learning method that takes an unknown input and assigns a
label for it. In simple words, finds the category of class of the input attributes.
Regression Analysis – A supervisory method that predicts the continuous variables based on the
input variables.
Unsupervised Learning – A type of machine leaning that uses unlabelled data and groups the
attributes to clusters using a trial and error approach.
Cluster Analysis – A type of unsupervised approach that groups the objects based on attributes so
that similar objects or data points form a cluster.
Semi-supervised Learning – A type of machine learning that uses limited labelled and large
unlabelleddata. It first labels unlabelled data using labelled data and combines it for learning
purposes.
Reinforcement Learning – A type of machine learning that uses agents and environment interaction
for creating labelled data for learning.
Well-posed Problem – A problem that has well-defined specifications. Otherwise, the problem is
called ill-posed.
Bias/Variance – The inability of the machine learning algorithm to predict correctly due to lack of
generalization is called bias. Variance is the error of the model for training data. This leads to
problems called overfitting and underfitting.
Model Deployment – A method of deploying machine learning algorithms to improve the existing
business processes for a new situation.
Chapter 2
Understanding Data
2.1 WHAT IS DATA?
All facts are data. In computer systems, bits encode facts present in numbers, text, images, audio,
and video. Data can be directly human interpretable (such as numbers or texts) or diffused data such
as images or video that can be interpreted only by a computer.
Data is available in different data sources like flat files, databases, or data warehouses. It can either
be an operational data or a non-operational data. Operational data is the one that is encountered in
normal business procedures and processes. For example, daily sales data is operational data, on the
other hand, non-operational data is the kind of data that is used for decision making.
Data by itself is meaningless. It has to be processed to generate any information. A string of bytes is
meaningless. Only when a label is attached like height of students of a class, the data becomes
meaningful.
Processed data is called information that includes patterns, associations, or relationships among
data. For example, sales data can be analyzed to extract information like which product was sold
larger in the last quarter of the year.
Elements of Big Data
Data whose volume is less and can be stored and processed by a small-scale computer is called
‘small data’. These data are collected from several sources, and integrated and processed by a small-
scale computer. Big data, on the other hand, is a larger data whose volume is much larger than
‘small data’ and is characterized as follows:
1. Volume – Since there is a reduction in the cost of storing devices, there has been a tremendous
growth of data. Small traditional data is measured in terms of gigabytes (GB) and terabytes (TB),
but Big Data is
measured in terms of petabytes (PB) and exabytes (EB). One exabyte is 1 million terabytes.
2. Velocity – The fast arrival speed of data and its increase in data volume is noted as velocity. The
availability of IoT devices and Internet power ensures that the data is arriving at a faster rate.
Velocity helps to understand the relative growth of big data and its accessibility by users, systems
and applications.
3. Variety – The variety of Big Data includes:
• Form – There are many forms of data. Data types range from text, graph, audio, video, to maps.
There can be composite data too, where one media can have many other sources of data, for
example, a video can have an audio song.
• Function – These are data from various sources like human conversations, transaction records, and
old archive data.
• Source of data – This is the third aspect of variety. There are many sources of data. Broadly, the
data source can be classified as open/public data, social media data and multimodal data. Some of
the other forms of Vs that are often quoted in the literature as characteristics of Big data are:
4. Veracity of data – Veracity of data deals with aspects like conformity to the facts, truthfulness,
believablity, and confidence in data. There may be many sources of error such as technical errors,
typographical errors, and human errors. So, veracity is one of the most important aspects of data.
5. Validity – Validity is the accuracy of the data for taking decisions or for any other goals that are
needed by the given problem.
6. Value – Value is the characteristic of big data that indicates the value of the information that is
extracted from the data and its influence on the decisions that are taken based on it.
Thus, these 6 Vs are helpful to characterize the big data. The data quality of the numeric attributes
is determined by factors like precision, bias, and accuracy.
2.1.1 Types of Data
In Big Data, there are three kinds of data. They are structured data, unstructured data, and
semistructured data.
Structured Data
In structured data, data is stored in an organized manner such as a database where it is available in
the form of a table. The data can also be retrieved in an organized manner using tools like SQL. The
structured data frequently encountered in machine learning are listed below:
Record Data A dataset is a collection of measurements taken from a process. We have a collection
of objects in a dataset and each object has a set of measurements. The measurements can be
arranged in the form of a matrix. Rows in the matrix represent an object and can be called as
entities, cases, or records. The columns of the dataset are called attributes, features, or fields. The
table is filled with observed data. Also, it is better to note the general jargons that are associated
with the dataset. Label is the term that is used to describe the individual observations.
Data Matrix It is a variation of the record type because it consists of numeric attributes. The
standard matrix operations can be applied on these data. The data is thought of as points or vectors
in the multidimensional space where every attribute is a dimension describing the object.
Graph Data It involves the relationships among objects. For example, a web page can refer to
another web page. This can be modeled as a graph. The modes are web pages and the hyperlink is
an edge that connects the nodes.
Ordered Data Ordered data objects involve attributes that have an implicit order among them. The
examples of ordered data are:
Temporal data – It is the data whose attributes are associated with time. For example, the customer
purchasing patterns during festival time is sequential data. Time series data is a special type of
sequence data where the data is a series of measurements over time.
Sequence data – It is like sequential data but does not have time stamps. This data involves the
sequence of words or letters. For example, DNA data is a sequence of four characters – A T G C.
Spatial data – It has attributes such as positions or areas. For example, maps are spatial data
where the points are related by location.
Unstructured Data
Unstructured data includes video, image, and audio. It also includes textual documents, programs,
and blog data. It is estimated that 80% of the data are unstructured data.
Semi-Structured Data
Semi-structured data are partially structured and partially unstructured. These include data like
XML/JSON data, RSS feeds, and hierarchical data.
2.1.2 Data Storage and Representation
Once the dataset is assembled, it must be stored in a structure that is suitable for data analysis. The
goal of data storage management is to make data available for analysis. There are different
approaches to organize and manage data in storage files and systems from flat file to data
warehouses. Some of them are listed below:
Flat Files These are the simplest and most commonly available data source. It is also the cheapest
way of organizing the data. These flat files are the files where data is stored in plain ASCII or
EBCDIC format. Minor changes of data in flat files affect the results of the data mining algorithms.
Hence, flat file is suitable only for storing small dataset and not desirable if the dataset becomes
larger.
Some of the popular spreadsheet formats are listed below:
• CSV files – CSV stands for comma-separated value files where the values are separated by
commas. These are used by spreadsheet and database applications. The first row may have
attributes and the rest of the rows represent the data.
• TSV files – TSV stands for Tab separated values files where values are separated by Tab. Both
CSV and TSV files are generic in nature and can be shared. There are many tools like Google
Sheets and Microsoft Excel to process these files.
Database System It normally consists of database files and a database management system
(DBMS). Database files contain original data and metadata. DBMS aims to manage data and
improve operator performance by including various tools like database administrator, query
processing, and transaction manager. A relational database consists of sets of tables. The tables have
rows and columns. The columns represent the attributes and rows represent tuples. A tuple
corresponds to either an object or a relationship between objects. A user can access and manipulate
the data in the database using SQL.
Different types of databases are listed below:
1 A transactional database is a collection of transactional records. Each record is a transaction. A
transaction may have a time stamp, identifier and a set of items, which may have links to other
tables. Normally, transaction databases are created for performing associational analysis that
indicates the correlation among the items.
2. Time-series database stores time related information like log files where data is associated with a
time stamp. This data represents the sequences of data, which represent values or events obtained
over a period (for example, hourly, weekly or yearly) or repeated time span. Observing sales of
product continuously may yield a time-series data.
3. Spatial databases contain spatial information in a raster or vector format. Raster formats are
either bitmaps or pixel maps. For example, images can be stored as a raster data. On the other hand,
the vector format can be used to store maps as maps use basic geometric primitives like points,
lines, polygons and so forth.
World Wide Web (WWW) It provides a diverse, worldwide online information source.
The objective of data mining algorithms is to mine interesting patterns of information present in
WWW.
XML (eXtensible Markup Language) It is both human and machine interpretable data format that
can be used to represent data that needs to be shared across the platforms.
Data Stream It is dynamic data, which flows in and out of the observing environment. Typical
characteristics of data stream are huge volume of data, dynamic, fixed order movement, and
realtime constraints.
RSS (Really Simple Syndication) It is a format for sharing instant feeds across services.
JSON (JavaScript Object Notation) It is another useful data interchange format that is often used
for many machine learning algorithms.
2.2 BIG DATA ANALYTICS AND TYPES OF ANALYTICS
The primary aim of data analysis is to assist business organizations to take decisions. For example,
a business organization may want to know which is the fastest selling product, in order for them to
market activities. Data analysis is an activity that takes the data and generates useful information
and insights for assisting the organizations.
Data analysis and data analytics are terms that are used interchangeably to refer to the same
concept. However, there is a subtle difference. Data analytics is a general term and data analysis is a
part of it. Data analytics refers to the process of data collection, preprocessing and analysis. It deals
with the complete cycle of data management. Data analysis is just analysis and is a part of data
analytics. It takes historical data and does the analysis. Data analytics, instead, concentrates more on
future and helps in prediction.
There are four types of data analytics:
1. Descriptive analytics
2. Diagnostic analytics
3. Predictive analytics
4. Prescriptive analytics
Descriptive Analytics : It is about describing the main features of the data. After data collection is
done, descriptive analytics deals with the collected data and quantifies it. It is often stated that
analytics is essentially statistics. There are two aspects of statistics – Descriptive and Inference.
Descriptive analytics only focuses on the description part of the data and not the inference part.
Diagnostic Analytics :It deals with the question – ‘Why?’. This is also known as causal analysis, as
it aims to find out the cause and effect of the events. For example, if a product is not selling,
diagnostic analytics aims to find out the reason. There may be multiple reasons and associated
effects are analyzed as part of it.
Predictive Analytics :It deals with the future. It deals with the question – ‘What will happen in
future given this data?’. This involves the application of algorithms to identify the patterns to
predict the future.
Prescriptive Analytics :It is about the finding the best course of action for the business
organizations. Prescriptive analytics goes beyond prediction and helps in decision making by giving
a set of actions. It helps the organizations to planbetter for the future and to mitigate the risks that
are involved.
2.3 BIG DATA ANALYSIS FRAMEWORK
For performing data analytics, many frameworks are proposed. All proposed analytics frameworks
have some common factors. Big data framework is a layered architecture. Such an architecture has
many advantages such as genericness. A 4-layer architecture has the following layers:
1. Date connection layer
2. Data management layer
3. Data analytics later
4. Presentation layer
Data Connection Layer It has data ingestion mechanisms and data connectors. Data ingestion means
taking raw data and importing it into appropriate data structures. It performs the tasks of ETL
process. By ETL, it means extract, transform and load operations.
Data Management Layer It performs preprocessing of data. The purpose of this layer is to allow
parallel execution of queries, and read, write and data management tasks. There may be many
schemes that can be implemented by this layer such as data-in-place, where the data is not moved at
all, or constructing data repositories such as data warehouses and pull data on-demand mechanisms.
Data Analytic Layer It has many functionalities such as statistical tests, machine learning algorithms
to understand, and construction of machine learning models. This layer implements many model
validation mechanisms too. The processing is done as shown in Box 2.1.
Presentation Layer It has mechanisms such as dashboards, and applications that display the results
of analytical engines and machine learning algorithms.
Thus, the Big Data processing cycle involves data management that consists of the following steps.
1. Data collection
2. Data preprocessing
3. Applications of machine learning algorithm
4. Interpretation of results and visualization of machine learning algorithm
This is an iterative process and is carried out on a permanent basis to ensure that data is suitable for
data mining.
Application and interpretation of machine learning algorithms constitute the basis for the rest of the
book. So, primarily, data collection and data preprocessing are covered as part of this chapter.
2.3.1 Data Collection
The first task of gathering datasets are the collection of data. It is often estimated that most of the
time is spent
for collection of good quality data. A good quality data yields a better result. It is often difficult to
characterize a ‘Good data’. ‘Good data’ is one that has the following properties:
1. Timeliness – The data should be relevant and not stale or obsolete data.
2. Relevancy – The data should be relevant and ready for the machine learning or data mining
algorithms. All the necessary information should be available and there should be no bias in the
data.
3. Knowledge about the data – The data should be understandable and interpretable, and should be
self- sufficient for the required application as desired by the domain knowledge engineer.
Broadly, the data source can be classified as open/public data, social media data and multimodal
data.
1. Open or public data source – It is a data source that does not have any stringent copyright rules or
restrictions. Its data can be primarily used for many purposes. Government census data are good
examples of open data:
• Digital libraries that have huge amount of text data as well as document images • Scientific
domains with a huge collection of experimental data like genomic data and biological data
• Healthcare systems that use extensive databases like patient databases, health insurance data,
doctors’ information, and bioinformatics information
2. Social media – It is the data that is generated by various social media platforms like Twitter,
Facebook, YouTube, and Instagram. An enormous amount of data is generated by these platforms.
data – It includes data that involves many modes such as text, video, audio and mixed types. Some
of them are listed below:
• Image archives contain larger image databases along with numeric and text data
• The World Wide Web (WWW) has huge amount of data that is distributed on the Internet. These
data are heterogeneous in nature.
2.3.2 Data Preprocessing
In real world, the available data is ’dirty’. By this word ’dirty’, it means:
• Incomplete data
• Inaccurate data
• Outlier data
• Data with missing values
• Data with inconsistent values
• Duplicate data
Data preprocessing improves the quality of the data mining techniques. The raw data must be
preprocessed to give accurate results. The process of detection and removal of errors in data is
called data cleaning. Data wrangling means making the data processable for machine learning
algorithms. Some of the data errors include human errors such as typographical errors or incorrect
measurement and structural errors like improper data formats. Data errors can also arise from
omission and duplication of attributes. Noise is a random component and involves distortion of a
value or introduction of spurious objects. Often, the noise is used if the data is a spatial or temporal
component. Certain deterministic distortions in the form of a streak are known as artifacts.
Consider, for example, the following patient Table 2.1. The ‘bad’ or ‘dirty’ data can be observed in
this table.
It can be observed
that data like Salary = ’ ’ is incomplete data. The DoB of patients, John, Andre, and Raju, is the
missing data. The age of David is recorded as ‘5’ but his DoB indicates it is 10/10/1980. This is
called inconsistent data.
Inconsistent data occurs due to problems in conversions, inconsistent formats, and difference in
units. Salary for John is -1500. It cannot be less than ‘0’. It is an instance of noisy data. Outliers are
data that exhibit the characteristics that are different from other data and have very unusual values.
The age of Raju cannot be 136. It might be a typographical error. It is often required to distinguish
between noise and outlier data.
Outliers may be legitimate data and sometimes are of interest to the data mining algorithms. These
errors often come during data collection stage. These must be removed so that machine learning
algorithms yield better results as the quality of results is determined by the quality of input data.
This removal process is called data cleaning.
Missing Data Analysis
The primary data cleaning process is missing data analysis. Data cleaning routines attempt to fill up
the missing values, smoothen the noise while identifying the outliers and correct the inconsistencies
of the data. This enables data mining to avoid overfitting of the models.
The procedures that are given below can solve the problem of missing data:
1. Ignore the tuple – A tuple with missing data, especially the class label, is ignored. This method is
not effective when the percentage of the missing values increases.
2. Fill in the values manually – Here, the domain expert can analyse the data tables and carry out
the analysis and fill in the values manually. But, this is time consuming and may not be feasible for
larger sets.
3. A global constant can be used to fill in the missing attributes. The missing values may be
’Unknown’ or be ’Infinity’. But, some data mining results may give spurious results by analysing
these labels.
4. The attribute value may be filled by the attribute value. Say, the average income can replace a
missing value.
5. Use the attribute mean for all samples belonging to the same class. Here, the average value
replaces the missing values of all tuples that fall in this group.
6. Use the most possible value to fill in the missing value. The most probable value can be obtained
from other methods like classification and decision tree prediction.
Some of these methods introduce bias in the data. The filled value may not be correct and could be
just an estimated value. Hence, the difference between the estimated and the original value is called
an error or bias.
Removal of Noisy or Outlier Data
Noise is a random error or variance in a measured value. It can be removed by using binning, which
is a method where the given data values are sorted and distributed into equal frequency bins. The
bins are also called as buckets. The binning method then uses the neighbor values to smooth the
noisy data.
Some of the techniques commonly used are ‘smoothing by means’ where the mean of the bin
removes the values of the bins, ‘smoothing by bin medians’ where the bin median replaces the bin
values, and ‘smoothing by bin boundaries’ where the bin value is replaced by the closest bin
boundary. The maximum and minimumvalues are called bin boundaries. Binning methods may be
used as a discretization technique. Example 2.1 illustrates this principle.
Example 2.1: Consider the following set: S = {12, 14, 19, 22, 24, 26, 28, 31, 34}. Apply various
binning techniques and show the result.
Solution: By equal-frequency bin method, the data should be distributed across bins. Let us assume
the bins of size 3, then the above data is distributed across the bins as shown below:
Bin 1 : 12 , 14, 19
Bin 2 : 22, 24, 26
Bin 3 : 28, 31, 32
By smoothing bins method, the bins are replaced by the bin means. This method results in:
Bin 1 : 15, 15, 15
Bin 2 : 24, 24, 24
Bin 3 : 30.3, 30.3, 30.3
Using smoothing by bin boundaries method, the bins' values would be like:
Bin 1 : 12, 12, 19
Bin 2 : 22, 22, 26
Bin 3 : 28, 32, 32
As per the method, the minimum and maximum values of the bin are determined, and it serves as
bin boundary and does not change. Rest of the values are transformed to the nearest value. It can be
observed in Bin 1, the middle value 14 is compared with the boundary values 12 and 19 and
changed to the closest value, that is 12. This process is repeated for all bins.
Data Integration and Data Transformations
Data integration involves routines that merge data from multiple sources into a single data source.
So, this may lead to redundant data. The main goal of data integration is to detect and remove
redundancies that arise fromintegration. Data transformation routines perform operations like
normalization to improve the performance of the data mining algorithms. It is necessary to
transform data so that it can be processed. This can be considered as a preliminary stage of data
conditioning. Normalization is one such technique. In normalization, the attribute values are scaled
to fit in a range (say 0-1) to improve the performance of the data mining algorithm. Often, in neural
networks, these techniques are used. Some of the normalization procedures used are:
1. Min-Max
2. z-Score
Min-Max Procedure It is a normalization technique where each variable V is normalized by its
difference with the minimum value divided by the range to a new range, say 0–1. Often, neural
networks require this kind of normalization. The formula to implement this normalization is given
as:
Here max-min is the range. Min and max are the minimum and maximum of the given data, new
max and new min are the minimum and maximum of the target range, say 0 and 1.
Consider the set: V = {88, 90, 92, 94}. Apply Min-Max procedure and map the marks to a new
range 0–1.
Solution: The minimum of the list V is 88 and maximum is 94. The new min and new max are 0 and
1, respectively. The mapping can be done using Eq. (2.1) as:
So, it can be observed that the marks{88, 90, 92, 94}are mapped to the new range{0, 0.33, 0.66, 1}.
Thus, the Min-Max normalization range is between 0 and 1.
z-Score Normalization This procedure works by taking the difference between the field value and
mean value, and by scaling this difference by standard deviation of the attribute.
Here, s is the standard deviation of the list V and m is the mean of the list V.
Example 2.3: Consider the mark list V = {10, 20, 30}, convert the marks to z-score.
Solution: The mean and Sample Standard deviation (s) values of the list V are 20 and 10, respec-
tively. So the
z-scores of these marks are calculated using Eq. (2.2) as:
Hence, the z-score of the marks 10, 20, 30 are -1, 0 and 1, respectively.
Data Reduction
Data reduction reduces data size but produces the same results. There are different ways in which
data reduction can be carried out such as data aggregation, feature selection, and dimensionality
reduction.
operations like (=, ≠) are meaningful for these data. For example, the patient ID can be checked for
equality and nothing else.
•Ordinal Data – It provides enough information and has natural order. For example, Fever = {Low,
Medium, High} is an ordinal data. Certainly, low is less than medium and medium is less than high,
irrespective of the value. Any transformation can be applied to these data to get a new value.
Numeric or Qualitative Data It can be divided into two categories. They are interval type and ratio
type.
•Interval Data – Interval data is a numeric data for which the differences between values are
meaningful. For example, there is a difference between 30 degree and 40 degree. Only the
permissible
operations are + and -.
•Ratio Data – For ratio data, both differences and ratio are meaningful. The difference between the
ratio and interval data is the position of zero in the scale. For example, take the Centigrade-
Fahrenheit conversion. The zeroes of both scales do not match. Hence, these are interval data.
Another way of classifying the data is to classify it as:
1. Discrete value data
[Link] data
Discrete Data This kind of data is recorded as integers. For example, the responses of the survey
can be discrete data. Employee identification number such as 10001 is discrete data.
Continuous Data It can be fitted into a range and includes decimal point. For example, age is a
continuous data. Though age appears to be discrete data, one may be 12.5 years old and it makes
sense. Patient height and weight are all continuous data.
Third way of classifying the data is based on the number of variables used in the dataset. Based on
that, the data can be classified as univariate data, bivariate data, and multivariate data. This is shown
in Figure 2.2.
description involves finding the frequency distributions, central tendency measures, dispersion or
variation, and shape of the data.
2.5.1 Data Visualization
Data visualization is a graphical representation of data and information. With the help of data
visualization, we can see how the data looks like and what kind of correlation is held by the
attributes of the data. It is the fastest way to see if the features correspond to the output.
Let us consider some forms of graphs
Bar Chart A Bar chart (or Bar graph) is used to display the frequency distribution for variables.
Bar charts are used to illustrate discrete data. The charts can also help to explain the counts of
nominal data. It also helps in comparing the frequency of different groups.
The bar chart for students' marks {45, 60, 60, 80, 85} with Student ID = {1, 2, 3, 4, 5} is shown
below in Figure 2.3.
Pie Chart These are equally helpful in illustrating the univariate data. The percentage frequency
distribution of students' marks {22, 22, 40, 40, 70, 70, 70, 85, 90, 90} is below in Figure 2.4.
It can be observed that the number of students with 22 marks are 2. The total number of students are
10. So, 2/10 × 100 = 20% space in a pie of 100% is allotted for marks 22
Histogram It plays an important role in data mining for showing frequency distributions.
The histogram for students’ marks {45, 60, 60, 80, 85} in the group range of 0-25, 26-50, 51-75, 76-
100 is given below in Figure 2.5. One can visually inspect from Figure 2.5 that the number of
students in the range 76-100 is 2
Histogram conveys useful information like nature of data and its mode. Mode indicates the peak of
dataset. In other words, histograms can be used as charts to show frequency, skewness present in the
data, and shape.
These are similar to bar charts. They are less clustered as compared to bar charts, as they illustrate
the bars only with single points. The dot plot of English marks for five students with ID as {1, 2, 3,
4, 5} and marks {45, 60, 60, 80, 85} is given in Figure 2.6. The advantage is that by visual
inspection one can find out who got more marks.
at certain values, normally in the central location. It is called measure of central tendency (or
averages). Popular measures are mean, median and mode.
1. Mean – Arithmetic average (or mean) is a measure of central tendency that represents the ‘center’
of the dataset. Mathematically, the average of all the values in the sample (population) is denoted as
x. Let x1, x2, ... , xN be a set of ‘N’ values or observations, then the arithmetic mean is given as:
For example, the mean of the three numbers 10, 20, and 30 is 20
•Weighted mean – Unlike arithmetic mean that gives the weightage of all items equally, weighted
mean gives different importance to all items as the item importance varies.
Hence, different weightage can be given to items. In case of frequency distribution, mid values of
the range are taken for computation. This is illustrated in the following computation.
In weighted mean, the mean is computed by adding the product of proportion and group mean. It is
mostly used when the sample sizes are unequal.
•Geometric mean – Let x1, x2, ... , xN be a set of ‘N’ values or observations. Geometric mean is the
Nth root of the product of N items. The formula for computing geometric mean is given as follows:
Here, n is the number of items and xi are values. For example, if the values are 6 and 8, the
geometric mean is given as In larger cases, computing geometric mean is difficult. Hence, it is
usually calculated as:
The problem of mean is its extreme sensitiveness to noise. Even small changes in the input affect
the mean drastically. Hence, often the top 2% is chopped off and then the mean is calcu- lated for a
larger dataset.
2. Median – The middle value in the distribution is called median. If the total number of items in the
distribution is odd, then the middle value is called median. A median class is that class where
(N/2)th item is present.
In the continuous case, the median is given by the formula:
Median class is that class where N/2th item is present. Here, i is the class interval of the median
class and L1 is the lower limit of median class, f is the frequency of the median class, and cf is the
cumulative frequency of all classes preceding median.
3. Mode – Mode is the value that occurs more frequently in the dataset. In other words, the value
that has the highest frequency is called mode.
2.5.3 Dispersion
The spreadout of a set of data around the central tendency (mean, median or mode) is called
dispersion. Dispersion is represented by various ways such as range, variance, standard deviation,
and standard error. These are second order measures. The most common measures of the dispersion
data are listed below:
Range Range is the difference between the maximum and minimum of values of the given list of
data.
Standard Deviation The mean does not convey much more than a middle point. For example, the
following datasets {10, 20, 30} and {10, 50, 0} both have a mean of 20. The difference between
these two sets is the spread of data. Standard deviation is the average distance from the mean of the
dataset to each point.
The formula for sample standard deviation is given by:
Here, N is the size of the population, xi is observation or value from the population and m is the
population mean. Often, N – 1 is used instead of N in the denominator of Eq. (2.8).
Quartiles and Inter Quartile Range It is sometimes convenient to subdivide the dataset using
coordinates. Percentiles are about data that are less than the coordinates by some percentage of the
total value. Kth percentile is the property that the k% of the data lies at or below Xi. For example,
median is 50th percentile and can be denoted as Q0.50. The 25th percentile is called first quartile
(Q1) and the 75th percentile is called third quartile (Q3). Another measure that is useful to measure
dispersion is Inter Quartile Range (IQR). The IQR is the difference between Q3 and Q1.
Interquartile percentile = Q3 – Q1 (2.9)
Outliers are normally the values falling apart at least by the amount 1.5 × IQR above the third
quartile or below the first quartile.
Interquartile is defined by Q0.75 – Q0.25. (2.10)
Example 2.4: For patients’ age list {12, 14, 19, 22, 24, 26, 28, 31, 34}, find the IQR.
Solution: The median is in the fifth position. In this case, 24 is the median.
The first quartile is median of the scores below the mean i.e., {12, 14, 19, 22}. Hence, it’s the
median of the list below 24.
In this case, the median is the average of the second and third values, that is, Q0.25 = 16.5.
Similarly, the third quartile is the median of the values above the median, that is {26, 28, 31, 34}.
So, Q0.75 is the average of the seventh and eighth score. In this case, it is 28 + 31/2 = 59/2 = 29.5.
Hence, the
IQR using Eq. (2.10) is:
= Q0.75 – Q0.25
= 29.5-16.5 = 13
Semi-Quartile range SIQR = 1/2X13 =6.5
Five-point Summary and Box Plots The median, quartiles Q1 and Q3, and minimum and
maximum written in the order < Minimum, Q1, Median, Q3, Maximum > is known as five-point
summary.
Example 2.5: Find the 5-point summary of the list {13, 11, 2, 3, 4, 8, 9}.
Solution: The minimum is 2 and the maximum is 13. The Q1, Q2 and Q3 are 3, 8 and 11,
respectively. Hence, 5- point summary is {2, 3, 8, 11, 13}, that is, {minimum, Q1, median, Q3,
maximum}. Box plots are useful for describing 5-point summary. The Box plot for the set is given
in Figure 2.7.
2.5.4 Shape
Skewness and Kurtosis (called moments) indicate the symmetry/asymmetry and peak location of
the dataset.
Skewness
The measures of direction and degree of symmetry are called measures of third order. Ideally,
skewness should be zero as in ideal normal distribution. More often, the given dataset may not have
perfect symmetry (consider the following Figure 2.8).
Generally, for negatively skewed distribution, the median is more than the mean. The relationship
between skew and the relative size of the mean and median can be summarized by a convenient
numerical skew index known as Pearson 2 skewness coefficient.
Also, the following measure is more commonly used to measure skewness. Let X1, X2, ..., XN be a
set of ‘N’ values or observations then the skewness can be given as:
Here, m is the population mean and s is the population standard deviation of the univariate data.
Sometimes, for bias correction instead of N, N - 1 is used.
Kurtosis
Kurtosis also indicates the peaks of data. If the data is high peak, then it indicates higher kurtosis
and vice versa. Kurtosis is measured using the formula given below:
It can be observed that N - 1 is used instead of N in the numerator of Eq. (2.14) for bias correction.
Here, x and s are the mean and standard deviation of the univariate data, respectively. Some of the
other useful measures for finding the shape of the univariate dataset are mean absolute deviation
(MAD) and coefficient of variation (CV).
Mean Absolute Deviation (MAD)
MAD is another dispersion measure and is robust to outliers. Normally, the outlier point is detected
by computing the deviation from median and by dividing it by MAD. Here, the absolute deviation
between the data and mean is taken. Thus, the absolute deviation is given as:
variation is used to compare datasets with different units. CV is the ratio of standard deviation and
mean, and %CV is the percentage of coefficient of variations.
2.5.5 Special Univariate Plots
The ideal way to check the shape of the dataset is a stem and leaf plot. A stem and leaf plot are a
display that help us to know the shape and distribution of the data. In this method, each value is
split into a ’stem’ and a ’leaf’. The last digit is usually the leaf and digits to the left of the leaf
mostly form the stem. For example, marks 45 are divided into stem 4 and leaf 5 in Figure 2.9. The
stem and leaf plot for the English subject marks, say, {45, 60, 60, 80, 85} is given in
Figure 2.9
It can be seen from Figure 2.9 that the first column is stem and the second column is leaf. For the
given English marks, two students with 60 marks are shown in stem and leaf plot as stem6 with 2
leaves with 0. The normal Q-Q plot for marks x = [13 11 2 3 4 8 9] is given below in Figure 2.10.
Chapter – 1
INTRODUCTION TO MACHINE LEARNING
Numerical Problems and Activities
1. Let us assume a regression algorithm generates a model y = 0.54 + 0.66 x for data
pertaining to week and sales of a product. Here, x is the week and y is the product sales.
Find the prediction for the 5th and 8th week.
Solution
The generated model is given as y = 0.54 + 0.66 x. For the fifth week, x=5, the
prediction value of y, product sales, is given as y = 0.54 + 0.665 = 3.84.
The prediction value of y, product sales, for 8th week is given as
y = 0.54 + 0.668 = 5.34
Solution
Models are global and applicable for the entire dataset. Examples of models are
y = 0.54 + 0.66x
y = 0.62 + 0.73x
Patterns are local and reflect the properties of local data. Example, association rules in
the form x → y , showing the associations
bread → butter
milk → coffee
3. Survey and find out at least five latest applications of machine learning.
Solution
Bank credit
Company Churn application
Sentiment analysis
Tourism Scheduling
Planning Algorithms
4. Survey and list out at least five products that use machine learning.
Solution
Alexa
Cortana
Google Translate
Amazon Recommendations
Netflix Recommendations
Data science, machine learning, and data analytics intersect in their use of data to drive decision-making and extract insights. Data science encompasses capturing and analyzing data for various applications. Machine learning, a branch of data science, focuses on learning patterns from data for predictions. Data analytics, another branch, aims to extract useful knowledge from data, often employing descriptive, diagnostic, predictive, or prescriptive analytics. The differentiation primarily lies in their scope and focus, with data science being broader, incorporating machine learning, and focusing on data management .
Machine learning models face challenges such as overfitting to training data, difficulty handling noisy datasets, and the need for vast amounts of labeled data for supervised learning. Evaluating models requires robust measures to ensure that they generalize well to unseen data, often necessitating cross-validation techniques. Additionally, models must adapt to changing data distributions, which can require retraining or adjusting parameters. These limitations underscore the need for carefully designed evaluation frameworks and adaptive learning algorithms .
Central tendency measures such as mean, median, and mode summarize the central point of a dataset, making data analysis simpler and allowing for comparisons. Dispersion measures, which include range, variance, and standard deviation, describe the spread of data around this central point. Together, these measures help in understanding the distribution’s shape and variability, crucial for accurate data analysis and interpretation .
Big data significantly impacts machine learning by providing diverse and vast datasets that enhance model training and accuracy. Fields like language translation and image recognition have benefited from big data, which allows machine learning algorithms to learn from comprehensive and varied data sets. This ability to handle large volumes, varieties, and velocities of data has accelerated advancements in deep learning, furthering capabilities in fields such as natural language processing, autonomous vehicles, and fraud detection .
Data mining plays a role in machine learning by uncovering hidden patterns within data, which can then be used by machine learning algorithms for predictive modeling. While data mining aims to find these hidden patterns, machine learning extends this by creating models that can predict future data based on identified patterns, focusing more on the application of these patterns for predictions .
Predictive modeling in medicine helps in predicting disease outcomes and the effectiveness of treatments using patient history and other data. In multimedia, machine learning enhances face recognition, biometric identification, and multimedia retrieval by constructing predictive models that discern patterns. These sectors utilize machine learning to analyze vast datasets efficiently, resulting in improved decision-making and personalized services .
Machine learning differs from traditional statistical methods in that it focuses on extracting patterns for prediction from data, rather than starting with a predefined hypothesis. Statistics involves setting a hypothesis and using data to verify relationships, often requiring complex mathematical models and assumptions. The objective of machine learning is more oriented towards application and prediction, while statistics emphasize hypothesis testing and understanding relationships within the data .
Humans utilize past experiences to solve new problems by searching for similar past situations and applying the heuristics developed from previous experiences. This process involves collecting data, forming abstract concepts, and generalizing these abstractions into actionable intelligence. Similarly, a machine learning model learns from historical data, forming patterns and abstractions that it can use to make predictions on new data .
Machine learning is a subset of artificial intelligence focused on extracting patterns for prediction from data. While AI encompasses the broader goal of developing intelligent agents capable of performing complex tasks autonomously, machine learning specifically targets learning from examples and data to improve these tasks. AI's objective extends to logic, reasoning, and decision-making processes, whereas machine learning is more about optimization of algorithms for data-driven predictions .
Supervised learning involves training a model on labeled data, using input-output pairs to learn a mapping for future predictions. It relies on guidance from labeled examples to classify or predict outcomes. Unsupervised learning, conversely, deals with unlabeled data and aims to infer the natural structure present within a dataset. It uses techniques like clustering to group data based on similarities without any prior labels, primarily focusing on exploring data patterns .