0% found this document useful (0 votes)
16 views37 pages

Introduction to Machine Learning Concepts

The document provides an overview of Machine Learning, its necessity in business, and its relationship with data science, statistics, and artificial intelligence. It outlines key concepts, types of learning (supervised, unsupervised, semi-supervised, and reinforcement), and challenges faced in the field. Additionally, it explains important terms and definitions related to data, information, knowledge, and intelligence within the context of machine learning applications.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
16 views37 pages

Introduction to Machine Learning Concepts

The document provides an overview of Machine Learning, its necessity in business, and its relationship with data science, statistics, and artificial intelligence. It outlines key concepts, types of learning (supervised, unsupervised, semi-supervised, and reinforcement), and challenges faced in the field. Additionally, it explains important terms and definitions related to data, information, knowledge, and intelligence within the context of machine learning applications.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Module 1

Syllabus:

Introduction: Need for Machine Learning, Machine Learning Explained, Machine Learning
in Relation to other Fields, Types of Machine Learning, Challenges of Machine Learning,
Machine Learning Process, Machine Learning Applications.

Understanding Data – 1: Introduction, Big Data Analysis Framework, Descriptive


Statistics, Univariate Data Analysis and Visualization.

Need of Machine Learning:

 Business organizations use huge amount of data for their daily activities.
 The full potential of this data was not utilized due to two reasons.
 1. Data being scattered across different archive systems and organizations were not
able to integrate these sources fully.
 2. Lack of awareness about software tools that could help to unearth the useful
infom1ation from data.
 Machine learning has become so popular because of three reasons:
1. High volume of available data to manage: Big companies such as Facebook,
Twitter, and YouTube generate huge amount of data that grows at a
phenomenal rate. It is estimated that the data approximately gets doubled
every year.
2. Second reason is that the cost of storage has reduced. The hardware cost
has also dropped. Therefore, it is easier now to capture, process, store,
distribute, and transmit the digital infom1ation.
3. Third reason for popularity of machine learning is the availability of complex
algorithms now. Especially with the advent of deep learning, many
algorithms are available for machine learning.

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 1


Important Terms in Machine Learning:

 Data: All facts are data, Data can be numbers or text that can be processed by a
computer.
 Information: Processed data is called information; this includes patterns,
associations, or relationships among data. For example, sales data can be analysed to
extract information like which is the fast selling product.
 Knowledge: Condensed information is called knowledge. Unless knowledge is
extracted, data is of no use. knowledge is not useful unless it is put into action.
 Intelligence: Intelligence is the applied knowledge for actions, An actionable fom1 of
knowledge is called intelligence.
 Wisdom: The ultimate objective of knowledge pyramid is wisdom that represents the
maturity of mind that is, so far, exhibited only by humans.

Machine Learning Explained

 Machine learning is an important sub-branch of Artificial Intelligence (AI).


 Arthur Samuel: He stated that “Machine learning is the field of study that gives the
computer’s ability to learn without being explicitly programmed."
 In conventional programming, after understanding the problem, a detailed design of
the program such as a flowchart or an algorithm needs to be created and converted
into programs using a suitable programming language.

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 2


 This approach could be difficult for many real-world problems such as puzzles,
games, and complex image recognition applications.
 Artificial intelligence aims to understand these problems and develop general purpose
rules manually. Then, these rules are fom1Ulated into logic and implemented in a
program to create intelligent systems.
 This idea of developing intelligent systems by using logic and reasoning by
converting an expert's knowledge into a set of rules and programs is called an expert
system
 Example: An expert system called MYCIN was designed for medical diagnosis after
converting the expert knowledge of many doctors into a system, this approach did not
progress much as programs lacked real intelligence, it was impractical in many
domains as programs still depended on human expertise and hence did not truly
exhibit intelligence.
 The aim of machine learning is to learn a model or set of rules from the given dataset
automatically so that it can predict the unknown data correctly.
 As humans take decisions based on an experience, computers make models based on
extracted patterns in the input data and then use these data-filled models for prediction
and to take decisions.

 Tom Mitchell's definition of machine learning states that, "A computer program is
said to learn from experience E, with respect to task T and some performance measure
P, if its performance on T measured by P improves with experience E." The important
components of this definition are experience E, task T, and performance measure P.
 Example: Task T could be detecting an object in an image. The machine can gain
the knowledge of object using training dataset of thousands of images. This is
called experience E. the focus is to use this experience E for the task of object

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 3


detection T. The ability of the system to detect the object is measured by
performance measures like precision and recall.
 experience is gathered by these steps:
1. Collection of data
2. Once data is gathered, abstract concepts are fom1ed out of that
data. Abstraction is used to generate concepts,
3. Generalization converts the abstraction into an actionable form of
intelligence.
4. Heuristics are educated guesses for all tasks. For example, if one runs
or encounters a danger, it is the resultant of human experience or his
heuristics fom1ation.

Machine Learning In Relation To Other Fields:

Machine Learning and Artificial Intelligence

 The aim of AI is to develop intelligent agents.


 Initially, the idea of AI was ambitious, that is, to develop intelligent systems like
human beings.
 Machine learning is the sub branch of AI, whose aim is to extract the patterns for
prediction. It is a broad field that includes learning from examples and other areas like
reinforcement learning.

 Deep learning is a sub branch of machine learning. In deep learning, the models are
constructed using neural network technology. Neural networks are based on the
human neuron models. Many neurons fom1 a network connected with the activation
functions that trigger further neurons to perfom1 tasks.

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 4


Machine Learning, Data Science, Data Mining, Data Analytics

 Data science deals with gathering of data for analysis, it includes the following

Big Data:

 Big data is a field of data science that deals with data's following characteristics:
1. Volume: Huge amount of data is generated by big companies like Facebook,
Twitter, and YouTube.
2. Variety: Data is available in variety of forms like images, videos, and in
different fom1ats.
3. Velocity: It refers to the speed at which the data is generated and processed.
 Big data is used by many machine learning algorithms for applications such as
language Translation and image recognition.

Data Mining:

 Like while mining the earth one gets into precious resources
 It is often believed that unearthing of the data produces hidden information that
otherwise would have eluded the attention
 Data mining and machine learning are same. There is no difference between these
fields except that data mining aims to extract the hidden patterns that are present in
the data, whereas, machine learning aims to use it for prediction.

Data Analytics:

 It aims to extract useful knowledge from crude data.

Pattern Recognition

 It uses machine learning algorithms to extract the features for pattern analysis and
pattern classification.

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 5


Machine Learning and Statistics:

 Statistics is a branch of mathematics that has a solid theoretical foundation regarding


statistical learning.
 The difference between statistics and ML is that statistical methods look for regularity
in data called patterns
 Statistics sets a hypothesis and perfom1s experiments to verify and validate the
hypothesis in order to find relationships among data.
 Statistics requires knowledge of the statistical procedures and the guidance of a good
statistician.
 It is mathematics intensive and models are often complicated equations and involve
many assumptions.
 Statistical methods are coherent and rigorous.
 It has strong theoretical foundations and interpretations that require a strong statistical
knowledge.
 Machine learning, comparatively, has less assumptions and requires less statistical
knowledge.

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 6


Types of Machine Learning:

 Before discussing the types of learning, it is necessary to discuss about data.


 Labelled Data: data that has both input features and corresponding output labels.
 Unlabelled Data: data that has input features (X) but no corresponding output labels
(Y). This means that the data points do not have predefined categories, classifications,
or target values.

Supervised Learning:

 Supervised algorithms use labelled dataset.


 There is a supervisor or teacher component in supervised learning.

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 7


 A supervisor provides labelled data so that the model is constructed and generates test
data.
 In supervised learning algorithms, learning takes place in two stages.
 First stage, the teacher communicates the infom1ation to the student that the student is
supposed to master.
 The student receives the infom1ation and understands it.
 During this stage, the teacher has no knowledge of whether the infom1ation is grasped
by the student.
 In second stage, The teacher then asks the student a set of questions to find out how
much information has been grasped by the student. Based on these questions, the
student is tested, and the teacher informs the student about his assessment. This kind
of learning is typically called supervised learning.
 Supervised learning has two methods:
1. Classification
2. Regression

Classification

 Classification is a supervised learning method.


 The input attributes of the classification algorithms are called independent variables.
 The target attribute is called label or dependent variable.
 The relationship between the input and target variable is represented in the form of a
structure which is called a classification model.
 The focus of classification is to predict the 'label' that is in a discrete form (a value
from the set of finite values).
 Example: classification algorithm takes a set of labelled data images such as dogs and
cats to construct a model that can later be used to classify an unknown test image
data.

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 8


 In classification, learning takes place in two stages.
 During the first stage, called training stage, the learning algorithm takes a labelled
dataset and starts learning.
 After the training set, samples are processed and the model is generated.
 In the second stage, the constructed model is tested with test or unknown sample and
assigned a label. This is the classification process.
 Some of the key algorithms of classification are:
1. Decision Tree
2. Random Forest
3. Support Vector Machines
4. Naive Bayes
5. Artificial Neural Network and Deep Learning networks like CNN

Regression Model:

 They predict continuous variables like price.


 In other words, it is a number.
 A fitted regression model is shown in Figure 1.8 for a dataset that represent weeks
input x and product sales y.

 The regression model takes input x and generates a model in the fom1 of a fitted line
of the form y =f(x).
 Here, x is the independent variable that may be one or more attributes and y is the
dependent variable.

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 9


 In Figure 1.8, linear regression takes the training set and tries to fit it with a line -
product sales= 0.66 x Week+ 0.54. Here, 0.66 and 0.54 are all regression coefficients
that are learnt from data.
 The advantage of this model is that prediction for product sales (y) can be made for
unknown week data (x).
 For example, the prediction for unknown eighth week can be made by substituting x
as 8 in that regression formula to get y.
 One of the most important regression algorithms is linear regression.

What is the difference between classification and regression models? The main difference is
that regression models predict continuous variables such as product price, while classification
concentrates on assigning labels such as class.

Unsupervised Learning:

 As the name suggests, there are no supervisor or teacher components.


 In the absence of a supervisor or teacher, self-instruction is the most common kind of
learning process.
 This process of self-instruction is based on the concept of trial and error
 Here, the program is supplied with objects, but no labels are defined. The algorithm
itself observes the examples and recognizes patterns based on the principles of
grouping.
 Grouping is done in ways that similar objects form the same group.

Ouster analysis and Dimensional reduction algorithms are examples of unsupervised


algorithms.

Cluster Analysis:

 Cluster analysis is an example of unsupervised learning.


 It aims to group objects into disjoint clusters or groups.
 It aims to group objects into disjoint clusters or groups.
 Cluster analysis clusters objects based on its attributes.
 Some of the examples of clustering processes are - segmentation of a region of
interest in an image, detection of abnom1al growth in a medical image, and
detem1ining clusters of signatures in a gene database.

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 10


 Example: Algorithm takes a set of dogs and cats images and groups it as two
clusters-dogs and cats. It can be observed that the samples belonging to a cluster are
similar and samples are different radically across clusters.

 Some of the key clustering algorithms are:


1. k-means algorithm
2. Hierarchical algorithms

Dimensionality Reduction:

 Dimensionality reduction algorithms are examples of unsupervised algorithms. It


takes a higher dimension data as input and outputs the data in lower dimension by
taking advantage of the variance of the data.
 It is a task of reducing the dataset with few features without losing the generality.

Difference Between Supervised and Un supervised Learning

Semi-Supervised Learning:

 There are circumstances where the dataset has a huge collection of unlabelled data
and some labelled data.
 Labelling is a costly process and difficult to perfom1by the humans.
 Semi-supervised algorithms use unlabelled data by assigning a pseudo-label.
 Then, the labelled and pseudo-labelled dataset can be combined.

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 11


Reinforcement Learning:

 Reinforcement learning allows the agent to interact with the environment to get
rewards.
 The agent can be human, animal, robot, or any independent program.
 The rewards enable the agent to gain experience.
 The agent aims to maximize the reward.
 The reward can be positive or negative (Punishment). When the rewards are more, the
behaviour gets reinforced and learning becomes possible.
 Example : Grid Game
Consider the following example of a Grid game as shown in Figure

 In this grid game, the gray tile indicates the danger, black is a block, and the tile with
diagonal lines is the goal. The aim is to start, say from bottom-left grid, using the
actions left, right, top and bottom to reach the goal state.
 To solve this sort of problem, The agent interacts with the environment to get
experience.
 In the above case, the agent tries to create a model by simulating many paths and
finding rewarding paths.
 This experience helps in constructing a model.

There is no supervisor or labelled dataset. Many sequential decisions need to be taken to


reach the final decision.

Therefore, reinforcement algorithms are reward-based, goal-oriented algorithms.

Challenges Of Machine Learning:

1. Problems -Machine learning can deal with the 'well-posed' problems where
specifications are complete and available. Computers cannot solve 'ill-posed'
problems.

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 12


 From the above table
 That is, y =x1 x x2. Well! It is true! But, this is equally true that y may bey=
x1+ x2, or y =xt2. So, there are three functions that fit the data.
 This means that the problem is ill-posed.

2. Huge data - This is a primary requirement of machine learning. Availability of a


quality data is a challenge. A quality data means it should be large and should not
have data problems such as missing data or incorrect data.
3. High computation power - With the availability of Big Data, the computational
resource requirement has also increased. Systems with Graphics Processing Unit
(GPU) or even Tensor Processing Unit (TPU) are required to execute machine
learning algorithms. Also, machine learning tasks have become complex and hence
time complexity has increased, and that can be solved only with high computing
power.
4. Complexity of the algorithms - The selection of algorithms, describing the
algorithms, application of algorithms to solve machine learning task, and comparison
of algorithms have become necessary for machine learning or data scientists now.
Algorithms have become a big topic of discussion and it is a challenge for machine
teaching professionals to design, select, and evaluate optimal algorithms.
5. Bias/Variance - Variance is the error of the model. This leads to a problem called
bias/ variance trade off. A model that fits the training data correctly but fails for test
data, in general lacks generalization, is called over fitting. The reverse problem is
called underfitting where the model fails for training data but has good generalization.
Overfitting and underfitting are great challenges for machine learning algorithms.

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 13


Machine Learning Process:

1. Understanding the business - This step involves understanding the objectives and
requirements of the business organization. Generally, a single data mining algorithm
is enough for giving the solution. This step also involves the formulation of the
problem statement for the data mining process.
2. Understanding the data - It involves the steps like data collection, study of the
characteristics of the data, fom1ulation of hypothesis, and matching of patterns to the
selected hypothesis.
3. Preparation of data - This step involves producing the final dataset by cleaning the
raw data and preparation of data for the data mining process. The missing values may
cause problems during both training and testing phases. Missing data forces classifiers
to produce inaccurate results. This is a perennial problem for the classification
models. Hence, suitable strategies should be adopted to handle the missing data.
4. Modelling - This step plays a role in the application of data mining algorithm for the
data to obtain a model or pattern.
5. Evaluate - This step involves the evaluation of the data mining results using
statistical analysis and visualization methods. The performance of the classifier is
detem1ined by evaluating the accuracy of the classifier. The process of classification
is a fuzzy issue. For example, classification of emails requires extensive domain

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 14


knowledge and requires domain experts. Hence, performance of the classifier is very
crucial.
6. Deployment - This step involves the deployment of results of the data mining
algorithm to improve the existing process or for a new situation.

Machine Learning Applications

1. Sentiment analysis - This is an application of natural language processing (NLP)


where the words of documents are converted to sentiments like happy, sad, and angry
which are captured by emoticons effectively. For movie reviews or product reviews,
five stars or one star are automatically attached using sentiment analysis programs.
2. Recommendation systems - These are systems that make personalized purchases
possible. For example, Amazon recommends users to find related books or books
bought by people who have the same taste like you, and Netflix suggests shows or
related movies of your taste. The recommendation systems are based on machine
learning.
3. Voice assistants -Products like Amazon Alexa, Microsoft Cortana, Apple Siri, and
Google Assistant are all examples of voice assistants. They take speech commands
and perfom1 tasks. These chatbots are the result of machine learning technologies.
4. Technologies like Google Maps and those used by Uber are all examples of machine
learning which offer to locate and navigate shortest paths to reduce time.

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 15


What is Data:

 All facts are data.


 Data by itself is meaningless. It has to be processed to generate any information.
 A string of bytes is meaningless. Only when a label is attached like height of students
of a class, the data becomes meaningful.
 Processed data is called information that includes patterns, associations, or
relationships among data.
 For example, sales data can be analysed to extract information like which product
was sold larger in the last quarter of the year.

Elements of Big Data:

Big data, is a larger data whose volume is much larger than 'small data' and is characterized
as follows:

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 16


1. Volume - Since there is a reduction in the cost of storing devices, there has been a
tremendous growth of data. Small traditional data is measured in terms of gigabytes
(GB) and terabytes (TB), but Big Data is measured in terms of petabytes (PB) and
exabytes (EB). One exabyte is 1 million terabytes.
2. Velocity- The fast arrival speed of data and its increase in data volume is noted as
velocity. The availability of IoT devices and Internet power ensures that the data is
arriving at a faster rate. Velocity helps to understand the relative growth of big data
and its accessibility by users, systems and applications.
3. Variety - The variety of Big Data includes:
 Form - There are many forms of data. Data types range from text, graph,
audio, video, to maps. There can be composite data too, where one media can
have many other sources of data, for example, a video can have an audio song.
 Function - These are data from various sources like human conversations,
transaction records, and old archive data.
 Source of data - This is the third aspect of variety. There are many sources of
data. Broadly, the data source can be classified as open/public data, social
media data and multimodal data. These are discussed in Section 2.3.1 of this
chapter.
4. Veracity of data - Veracity of data deals with aspects like conformity to the facts,
truth- fulness, believability, and confidence in data. There may be many sources of
error such as technical errors, typographical errors, and human errors. So, veracity is
one of the most important aspects of data.
5. Validity- Validity is the accuracy of the data for taking decisions or for any other
goals that are needed by the given problem.
6. Value - Value is the characteristic of big data that indicates the value of the
information that is extracted from the data and its influence on the decisions that are
taken based on it.

Types of Data:

In Big Data, there are three kinds of data.

1. Structured data
2. Unstructured data
3. Semi-structured data.

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 17


Structured Data

In structured data, data is stored in an organized manner such as a database where it is


available in the form of a table.

The structured data frequently encountered in machine learning are listed below:
1. Record Data:
 A dataset is a collection of measurements taken from a process.
 Collection of objects in a dataset and each object has a set of
measurements.
 The measurements can be arranged in the form of a matrix.
 Rows in the matrix represent an object and can be called as entities,
cases, or records.
 The columns of the dataset are called attributes, features, or fields.
2. Data Matrix:
 It is a variation of the record type because it consists of numeric
attributes.
 The standard matrix operations can be applied on these data.
 The data is thought of as points or vectors in the multidimensional space
where every attribute is a dimension describing the object.
3. Graph Data:
 It involves the relationships among objects.
 For example, a web page can refer to another web page.
 This can be modeled as a graph.
 The modes are web pages and the hyperlink is an edge that connects the
nodes.
4. Ordered Data:

Ordered data objects involve attributes that have an implicit order among them.
The examples of ordered data are:
 Temporal data - It is the data whose attributes are associated with time. For
example, the customer purchasing patterns during festival time is sequential
data. Time series data is a special type of sequence data where the data is a
series of measurements over time.

 Sequence data - It is like sequential data but does not have time stamps. This
data involves the sequence of words or letters. For example, DNA data is a

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 18


sequence of four characters A T G C.
 Spatial data - It has attributes such as positions or areas. For example, maps
are spatial data where the points are related by location.

Unstructured Data:

 Unstructured data includes video, image, and audio.


 It also includes textual documents, programs, and blog data.

Semi Structured Data

 Semi-structured data are partially structured and partially unstructured.


 These include data like XMLflSON data, RSS feeds, and hierarchical data.

Data Storage and Representation:

Once the dataset is assembled, it must be stored in a structure that is suitable for data
analysis. The goal of data storage management is to make data available for analysis

There are different approaches to organize and manage data are:

 Database System: It normally consists of database files and a database management


system (DBMS). Database files contain original data and metadata. A relational
database consists of sets of tables. The tables have rows and columns. The columns
represent the attributes and rows represent tuples. A tuple corresponds to either an
object or a relationship between objects. A user can access and manipulate the data in
the database using SQL.

Different types of databases are listed below:


1. A transactional database is a collection of transactional records. Each record
is a transaction. A transaction may have a time stamp, identifier and a set of
items, which may have links to other tables.
2. Time-series database stores time related information like log files where data
is associated with a time stamp. This data represents the sequences of data,
which represent values or events obtained over a period (for example, hourly,
weekly or yearly) or repeated time span.
3. Spatial databases contain spatial information in a raster or vector format.
 World Wide Web (WWW) It provides a diverse worldwide online information
source. The objective of data mining algorithms is to mine teresting patterns of
information present in WWW.

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 19


 XML (eXtenslble Markup Language) It is both human and machine interpretable
data format that can beused torepresent data that needs to be shared across the
platforms.

 Data Stream It is dynamic data, which flows in and out of the observing
environment. Typical characteristics of data stream are huge volume of data, dynamic,
fixed order movement, and real-time constraints.

 RSS (Really Simple Syndication) It is a format for sharing instant feeds across
services.

 JSON (JavaScript Object Notation) It is another useful data interchange format that
is often used for many machine learning algorithms.

BIG DATA ANALYTICS AND TYPES OF ANALYTICS

 The primary aim of data analysis is to assist business organizations to take decisions.

 Data analysis is an activity that takes the data and generates useful information and
insights for assisting the organizations.

 Data analytics is a general term and data analysis is a part of it Data analytics
refers to the process of data collection, preprocessing and analysis.

 Data analysis is just analysis and is a part of data analytics. It takes historical data and
does the analysis. Data analytics, instead, concentrates more on future and helps in
prediction.

 There are four types of data analytics:


1. Descriptive analytics
2. Diagnostic analytics
3. Predictive analytics
4. Prescriptive analytics

Descriptive Analytics:

 It is about describing the main features of the data.


 After data collection is done, descriptive analytics deals with the collected data and
quantifies it.
 There are two aspects of statistics - Descriptive and Inference. Descriptive analytics
only focuses on the description part of the data and not the inference part.

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 20


Diagnostic Analytics:

 It deals with the question - 'Why?'.


 It aims to find out the cause and effect of the events.
 For example, if a product is not selling, diagnostic analytics aims to find out the
reason.

Predictive Analytics

 It deals with the future. It deals with the question - 'What will happen in future given
this data?'.
 This involves the application of algorithms to identify the patterns to predict the
future.

Prescriptive Analytics

 It is about the finding the best course of action for the business organizations.
 Prescriptive analytics goes beyond prediction and helps in decision making by giving
a set of actions.
 It helps the organizations to plan better for the future and to mitigate the risks that are
involved.

BIG DATA ANALYSIS FRAMEWORK

Big data framework is a layered architecture. . A 4-layer architecture has the following
layers:

1. Date connection layer


2. Data management layer
3. Data analytics later
4. Presentation layer

Data Connection Layer:

 It has data ingestion mechanisms and data connectors.


 Data ingestion means taking raw data and importing it into appropriate data
structures.
 It performs the tasks of ETL process. By ETL, it means extract, transform and load
operations.

Data Management Layer

 It performs preprocessing of data.


 The purpose of this layer is to allow parallel execution of queries, and read, write and
data management tasks.

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 21


Data Analytic Layer

 It has many functionalities such as statistical tests, machine learning algorithms to


understand, and construction of machine learning models.
 This layer implements many model validation mechanisms too.

Presentation Layer

 It has mechanisms such as dashboards, and applications that display the results of
analytical engines and machine learning algorithms.

The Big Data processing cycle involves data management that consists of the following
steps.

1. Data collection
2. Data preprocessing
3. Applications of machine learning algorithm
4. Interpretation of results and visualization of machine learning algorithm

Data Collection

 The first task of gathering datasets are the collection of data.


 Most of the time is spent for collection of good quality data.
 A good quality data yields a better result.
 'Good data' is one that has the following properties:
1. Timeliness - The data should be relevant and not stale or obsolete data.
2. Relevancy- The data should be relevant and ready for the machine learning or
data mining algorithms. All the necessary information should be available and
there should be no bias in the data.
3. Knowledge_about the data - The data should be understandable and
interpretable, and should be self-sufficient for the required application as
desired by the domain knowledge engineer.
 The data source can be classified as open/public data, social media data and
[Link] data.
1. Open or public data source: It is a data source that does not have any
stringent copyright rules or restrictions.
•Digital libraries that have huge amount of text data as well as document
images
•Scientific domains with a huge collection of experimental data like genomic
data and biological data
•Healthcare systems that use extensive databases like patient databases, health
insurance data, doctors' information, and bioinformatics information
2. Social media - It is the data that is generated by various social media
platforms like Twitter, Facebook, YouTube, and Instagram. An enormous
amount of data is generated by these platforms.

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 22


3. [Link] data - It includes data that involves many modes such as text,
video, audio and mixed types, ex: WWW.

Data Preprocessing

In real world, the available data is 'dirty'. By this word 'dirty', it means:

 Incomplete data
 Inaccurate data
 Outlier data
 Data with missing values
 Data with inconsistent values
 Duplicate data

Data preprocessing improves the quality of the data mining techniques. The raw data must
be preprocessed to give accurate results. The process of detection and removal of errors in
data is called data cleaning.

Data wrangling means making the data processable for machine learning algorithms.

Some of the data errors include human errors such as typographical errors or incorrect
measurement and structural errors like improper data formats. Data errors can also arise from
omission and duplication of attributes.

Noise is a random component and involves distortion of a value or introduction of spurious


objects. Often, the noise is used if the data is a spatial or temporal component.

 In the above table It can be observed that data like Salary = ' ' is incomplete data.
 The DOB of patients, John, Andre, and Raju, is the missing data.
 The age of David is recorded as '5' but his DOB indicates it is 10/10/1980. This is
called inconsistent data.
 Salary for John is -1500. It cannot be less than 'O'.
 The age of Raju cannot be 136.

Outliers may be legitimate data and sometimes are of interest to the data mining algorithms.
These errors often come during data collection stage. These must be removed so that machine

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 23


learning algorithms yield better results as the quality of results is determined by the quality of
input data. This removal process is called data cleaning.

Missing Data Analysis

The primary data cleaning process is missing data analysis. Data cleaning routines attempt to
fill up the missing values, smoothen the noise while identifying the outliers and correct the
inconsistencies of the data. This enables data mining to avoid overfitting of the models.

The procedures that are given below can solve the problem of missing data:

1. Ignore the tuple - A tuple with missing data, especially the class label, is ignored.
This method is not effective when the percentage of the missing values increases.
2. Fill in the values manually - Here, the domain expert can analyze the data tables and
carry out the analysis and fill in the values manually. But, this is time consuming and
may not be feasible for larger sets.
3. A global constant can be used to fill in the missing attributes. The missing values
may be 'Unknown' or be 'Infinity'. But, some data mining results may give spurious
results by analyzing these labels.
4. The attribute value may be filled by the attribute value. Say, the average income can
replace a missing value.
5. Use the attribute mean for all samples belonging to the same class. Here, the average
value replaces the missing values of all tuples that fall in this group.
6. Use the most possible value to fill in the missing value. The most probable value
can be obtained from other methods like classification and decision tree prediction

Removal of Noisy or Outlier Data

 Noise is a random error or variance in a measured value.


 It can be removed by using binning, which is a method where the given data values
are sorted and distributed into equal frequency bins.
 The bins are also called as buckets. The binning method then uses the neighbor
values to smooth the noisy data.
 Some of the techniques commonly used are 'smoothing by means' where the mean
of the bin removes the values of the bins.
 'Smoothing by bin medians' where the bin median replaces the bin values, and
'smoothing by bin boundaries' where the bin value is replaced by the closest bin
boundary.
 The maximum and minimum values are called bin boundaries.

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 24


Data Integration and Data Transformations

 Data integration involves routines that merge data from multiple sources into a single
data source.
 The main goal of data integration is to detect and remove redundancies that arise from
integration.
 Data transformation routines perform operations like normalization to improve the
performance of the data mining algorithms.
 It is necessary to transform data so that it can be processed.
 Normalization is one such technique. In normalization, the attribute values are scaled
to fit in a range (say 0-1) to improve the performance of the data mining algorithm.
Some of the normalization procedures used are:
1. Min-Max
2. z-Score

Min-Max Procedure

It is a normalization technique where each variable is normalized by its difference with the
minimum value divided by the range to a new range, say 0-1. Often, neural networks require
this kind of normalization. The formula to implement this normalization is given as:

Min and max are the minimum and maximum of the given data, new max and new min are
the minimum and maximum of the target range, say 0 and 1.

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 25


z-Score Normalization

This procedure works by taking the difference between the field value and mean value, and
by scaling this difference by standard deviation of the attribute.

Here, a is the standard deviation of the list V and µ is the mean of the list V.

z-scores are used to detect outlier detection. If the data value z-score function is either less
than -3 or greater than +3, then it is possibly an outlier.

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 26


Data Reduction

Data reduction reduces data size but produces the same results. There are different ways in
which data reduction can be carried out such as data aggregation, feature selection, and
dimensionality reduction.

DESCRIPTIVE STATISTICS

 Descriptive statistics is a branch of statistics that does dataset summarization.


 Data visualization is a branch of study that is useful for investigating the given data.
Mainly, the plots are useful to explain and present data to customers.
 Descriptive analytics and data visualization techniques help to understand the
nature of the data, which further helps to determine the kinds of machine learning or
data mining tasks that can be applied to the data. This step is often known as
Exploratory Data Analysis

Dataset and Data Types

 A dataset can be assumed to be a collection of data objects.


 The data objects may be records, points, vectors, patterns, events, cases, samples
or observations.
 These records contain many attributes. An attribute can be defined as the property or
characteristics of an object.

 Every attribute should be associated with a value. This process is called measurement.
The type of attribute determines the data types, often referred to as measurement scale
types.

 Broadly, data can be classified into two types:


1. Categorical or qualitative data
2. Numerical or quantitative data

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 27


Categorical or Qualitative Data

The categorical data can be divided into tw9 types.

 Nominal Data - In Table 2.2, patient ID is nominal data. Nominal data are symbols
and cannot be processed like a number. For example, the average of a patient ID does
not make any statistical sense. Nominal data type provides only information but has
no ordering among data. Only operations like (=, *) are meaningful for these data. For
example, the patient ID can be checked for equality and nothing else.
 Ordinal Data - It provides enough information and has natural order. For example,
Fever= {Low, Medium, High} is an ordinal data. Certainly, low is less than medium
and medium is less than high, irrespective of the value. Any transformation can be
applied to these data to get a new value.

Numeric or Qualitative Data

It can be divided into two categories

 Interval Data - Interval data is a numeric data for which the differences between
values are meaningful. For example, there is a difference between 30 degree and 40
degree. Only the permissible operations are+ and-.
 Ratio Data - For ratio data, both differences and ratio are meaningful. The difference
between the ratio and interval data is the position of zero in the scale. For example,
take the Centigrade-Fahrenheit conversion. The zeroes of both scales do not match.
Hence, these are interval data.

Another way of classifying the data is to classify it as:

1. Discrete value data


2. Continuous data

Discrete Data This kind of data is recorded as integers. For example, the responses of the
survey can be discrete data. Employee identification number such as 10001 is discrete data.

Continuous Data It can be fitted into a range and includes decimal point. For example, age is
a continuous data. Though age appears to be discrete data, one may be 12.5 years old and it
makes sense. Patient height and weight are all continuous data.

Third way of classifying the data is based on the number of variables used in the dataset.

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 28


In case of univariate data, the dataset has only one variable. Bivariate data indicates that the
number of variables used are two and multivariate data uses three or more variables.

UNIVARIATE DATA ANALYSIS AND VISUALIZATION

 As the name indicates, the dataset has only one variable.


 The aim of univariate analysis is to describe data and find patterns.
 Univariate data description involves finding the frequency distributions, central
tendency measures, dispersion or variation, and shape of the data.

Data Visualization
 To understand data, graph visualization is must.
 It helps to present information and data to customers.
 Some of the graphs that are used in univariate data analysis are bar charts, histograms,
frequency polygons and pie charts.

Bar Chart: A Bar chart (or Bar graph) is used to display the frequency distribution for
variables. Bar charts are used to illustrate discrete data. The charts can also help to explain
the counts of nominal data.

Pie Chart: The percentage frequency distribution of students' marks {22, 22, 40, 40, 70, 70,
70, 85, 90, 90)

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 29


It can be observed that the number of students with 22 marks are 2. The total number of
students are 10. So, 2/10 x 100 = 20% space in a pie of 100% is allotted for marks 22

Histogram It plays an important role in data mining for showing frequency distributions. The
histogram for students' marks {45, 60, 60, 80, 85} in the group range of 0-25, 26-50, 51-75,
76-100 is given below in Figure 2.5. One can visually inspect from Figure 2.5 that the
number of students in the range 76-100 is 2.

Histogram conveys useful information like nature of data and its mode. Mode indicates the
peak of dataset.

Dot Plot These are similar to bar charts. they are less clustered as compared to bar charts, as
they illustrate the bars only with single points. The dot plot of English marks for five students
with ID as {l, 2, 3, 4, 5) and marks {45.,60, 60, 80, 85} is given in Figure

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 30


Central Tendency

 One cannot remember all the data.


 Therefore, a condensation or summary of the data is necessary. This makes the data
analysis easy and simple. One such summary is called central tendency.
 Popular measures are mean, median and mode.

1. Mean - Arithmetic average (or mean) is a measure of central tendency that represents
the 'center' of the dataset. . It can be found by adding all the data and dividing the sum
by the number of observations. the average of all the values in the sample
(population) is denoted as x. Let x1, x2, • • • , xN be a set of 'N' values or
observations, then the arithmetic mean is given as:

 Weighted mean - Unlike arithmetic mean that gives the weightage of all
items equally, weighted mean gives different importance to all items as the
item importance varies. Hence, different weightage can be given to items.
 Geometric mean - Let x1, x2, .••, xN be a set of 'N' values or observations.
Geometric mean is the N'th root of the product of N items. The formula for
computing geometric mean is given as follows:

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 31


2. Median - The middle value in the distribution is called median. If the total number of
items in the distribution is odd, then the middle value is called median. If the numbers
are even, then the average value of two items in the center is the median. It can be
observed that the median is the value where x; is divided into two equal halves, with
half of the values being lower than the median and half higher than the median.

3. Mode - Mode is the value that occurs more frequently in the dataset. The procedure
for finding the mode is to calculate the frequencies for all the values in the data, and
mode is the value (or values) with the highest frequency.

Dispersion

The spreadout of a set of data around the central tendency (mean, median or mode) is called
dispersion. Dispersion is represented by various ways such as range, variance, standard
deviation, and standard error.

Range Range is the difference between the maximum and minimum of values of the given
list of data.

Standard Deviation The mean does not convey much more than a middle point. For
example, the following datasets {10, 20, 30} and {10, 50, 0} both have a mean of 20. The
difference between these two sets is the spread of data.

Standard deviation is the average distance from the mean of the dataset to each point.

The formula for sample standard deviation is given by;

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 32


Here, N is the size of the population, x. is observation or value from the population and µ is
the population mean. Often, N - 1 is used instead of N in the denominator

Quartiles and Inter Quartile Range It is sometimes convenient to subdivide the dataset
using coordinates.

For example, median is 50th percentile and can be denoted as Qo so The 25th percentile is
called first quartile (Q1) and the 75th percentile is called third quartile (Q3 ). Another
measure that is useful to measure dispersion is Inter Quartile Range (IQR). The IQR is the
difference between Q3 and Q1.

Five-point Summary and Box Plots The median, quartiles Q1 and Q3 and minimum and
maximum written in the order < Minimum, Q1 , Median, Q3 Maximum > is known as five-
point summary.

Box plots are suitable for continuous variables and a nominal variable. Box plots can be used
to illustrate data distributions and summary of data. It is the popular way for plotting five
number summaries.

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 33


Shape

Skewness and Kurtosis (called moments) indicate the symmetry/asymmetry and peak
location of the dataset.

Skewness

 If the dataset has far higher values, then it is said to be skewed to the right.
 On the other hand, if the dataset has far more low values then it is said to be skewed
towards left
 If the tail is longer on the left-hand side and hump on the right-hand side, it is called
positive skew. Otherwise, it is called negative skew.
 A perfect symmetry means the skewness is zero.
 In the case of skew, the median is greater than the mean. In positive skew, the mean is
greater than the median.

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 34


Kurtosis
 Kurtosis also indicates the peaks of data. If the data is high peak, then it indicates
higher kurtosis and vice versa.
 Kurtosis is the measure of whether the data is heavy tailed or light tailed
 Low kurtosis tends to have light tails.
 The implication is that there is no outlier data. Let x1, x2, • • •, xN be a set of 'N'
values or observations. Then, kurtosis is measured using the formula given below:

Mean Absolute Deviation (MAD)


MAD is another dispersion measure and is robust to outliers. Normally, the outlier point is
detect d by computing the deviation from median and by dividing it by MAD. Here, the
absolute deviation between the data and mean is taken. Thus, the absolute deviation is given
as:

Special Univariate Plots

 The ideal way to check the shape of the dataset is a stem and leaf plot. A stem
and leaf plot are a display that help us to know the shape and distribution of the
data.
 In this method, each value is • split into a 'stem' and a1eaf. The last digit is usually the
leaf and digits to the left of the leaf mostly form the stem.

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 35


It can be seen from Figure that the first column is stem and the second column is leaf. For the
given English marks, two students with 60 marks are shown in stem and leaf plot as stem-6
with 2 leaves with 0.

A Q-Q plot can be used to assess the shape of the dataset. The Q-Q plot is a 2D scatter plot of
an univariate data against theoretical normal distribution data or of two datasets - the quartiles
of the first and second datasets. The normal Q-Q plot for marks x = [1311 2 3 4 8 9] is given
below.

Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 36


Sonia Fathima B.E, [Link] Dept of AIML, AIT CKM Page 37

You might also like