0% found this document useful (0 votes)
18 views140 pages

Data Science Course Overview and Modules

Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPT, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
18 views140 pages

Data Science Course Overview and Modules

Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPT, PDF, TXT or read online on Scribd

SRI RAMAKRISHNA ENGINEERING

COLLEGE
[Educational Service : SNR Sons Charitable Trust]
[Autonomous Institution, Reaccredited by NAAC with ‘A+’ Grade]
[Approved by AICTE and Permanently Affiliated to Anna University, Chennai]
[ISO 9001:2015 Certified and all Eligible Programmes Accredited by NBA]
VATTAMALAIPALAYAM, N.G.G.O. COLONY POST, COIMBATORE – 641 022.

Department of Information Technology

20IT211- Data Science

Presentation by
[Link] Rani, AP([Link])/IT
COURSE OUTCOMES
20IT211- Data Science
Understand the basic concepts of data science and
CO1 PO1,PO2,PO12
data mining

CO2 Identify the techniques to explore and evaluate data PO3,PO5,PO12

Apply various data mining algorithms for real time PO2,PO3,PO5,P


CO3
applications O12

Implement the concepts of clustering and model


CO4 PO3,PO5,PO12
evaluation

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 2


20IT211- Data Science

Module I : INTRODUCTION 9 hours

What is data science – Case for data science – Data


science classification – Data science algorithms – Data
science process – Prior knowledge – Data preprocessing –
Data cleaning – Data integration – Data reduction – Data
transformation and data discretization – Feature selection
– Data sampling – Modeling – Application.

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 3


20IT211- Data Science

Module II : DATA EXPLORATION AND VISUALIZATION


9 hours

Objectives of Data exploration – Datasets – Descriptive


statistics – Data Visualization – Univariate visualization –
Multivariate visualization – Visualizing high dimensional
data – Roadmap for data exploration.

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 4


20IT211- Data Science

Module III : CLASSIFICATION AND ASSOCIATION


ANALYSIS 18 hours

Basic concepts of Classification – Decision tree induction –


Bayes classification methods – Rule based classification –
Techniques to improve classification accuracy – Support vector
machines – Regression methods: Linear regression – Logistic
regression – Association analysis: Frequent Item set mining
methods – Pattern evaluation methods.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 5
20IT211- Data Science

Module IV : CLUSTERING AND MODEL EVALUATION 9


hours

Basic concepts and methods in cluster analysis – Partitioning


methods – Density based methods – Model evaluation:
Confusion matrix – Receiver Operator Characteristics (ROC) and
Area under the Curve (AUC) – Lift curves – Evaluating the
Predictions – Implementation

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 6


TEXTBOOKS
1. Vijay Kotu and Bala Deshpande, “Data Science Concepts and Practice”, 2
Edition, Morgan Kaufmann Publishers, 2019.

2. Jiawei Han, Micheline Kamber and Jian Pei, “Data Mining: Concepts and
Techniques”, 3 Edition, Morgan Kaufmann Publishers, 2012.

3. Cathy O’Neil and Rachel Schutt, “Doing Data Science, Straight Talk From
The Frontline”, O’Reilly, 2016.

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 7


Reference(s)
1. Mohammed J. Zaki and Wagner Miera Jr, “Data Mining and Analysis:
Fundamental Concepts and Algorithms”, Cambridge University Press, 2014.

2. Matt Harrison, “Learning the Pandas Library: Python Tools for Data
Munging, Analysis and Visualization O’Reilly, 2016.

3. Joel Grus, “Data Science from Scratch: First Principles with Python”, O’Reilly
Media, 2015. 4. Wes McKinney, “Python for Data Analysis: Data Wrangling
with Pandas, NumPy, and IPython”, O’Reilly Media, 2012

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 8


WEB REFERENCES
1. [Link]

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 9


Acknowledgement
Resources are taken from the internet and textbooks

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 10


What is Data

Data are plain


• Data is afacts
collection of raw,
unorganised facts and details like
text, observations, figures, symbols
and description of things etc.
• In other words, data does not carry
any specific purpose and has no
significance by itself.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 11
What is Information

Information is the processed, organised and structured data.

It provides context for data and enables decision making.

For example, a single customer’s sale at a restaurant is data – this becomes information
when the business is able to identify the most popular or least popular dish.

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 12


What is Information
Information is defined as classified or organized data that has some meaningful value
for the user.

 Information is also the processed data used to make decisions and take action.

Processed data must meet the following criteria for it to be of any significant use in
decision-making:
Accuracy: The information must be accurate.

Completeness: The information must be complete.

Timeliness: The information must be available when it’s needed.

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 13


Example

 Marks / student’s test score

 Roll numbers Data


Data

 Report card/sheet /average score of the class- Information.


Information
 Other examples for information are pay-slips, schedules, reports, worksheet,

bar charts, invoices and account returns etc

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 14


Computer Data
Computer data is information processed or stored by a
computer.
◦ This information may be in the form of text documents, images, audio
clips, software programs, or other types of data.

Computer data may be processed by the computer’s CPU and


is stored in files and folders on the computer’s hard disk.

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 15


Types of Data
•Quantitative data is numerical and can be counted,

quantified, and mathematically analyzed (e.g., GPAs,

standardized test scores, attendance patterns). Quantitative

data is often considered a highly reliable source of information.

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 16


•Qualitative data is usually non-numerical and used to

provide meaning and understanding. Student narratives

describing their reasons for participating in your program

each month are examples of qualitative data. Qualitative data

is believed to have great validity and depth.

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 17


QUANTITATIVE DATA
Quantitative data is information that you can measure. It’s
numbers –something you can count. Because it’s countable it can
be reliable evidence. Examples include:

◦ How many people took part?

◦ How much did it cost?

◦ How long did it run for?

◦ Average attendance at each programme session?

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 18


QUALITATIVE DATA
Qualitative data is information about qualities, you can’t count it.
That is, it’s information about how people feel about something.
Examples include:

◦ Sharing what people like about a programme.

◦ How they think it could be improved.

◦ What difference it has made to their lives.

◦ Whether they would recommend the programme to others.

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 19


Data

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 20


Data

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 21


19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 22
Example

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 23


Example

Google data centre normally holds petabytes to exabytes(10 18 bytes)of data

Google currently processes over 20 petabytes of data per day

Facebook: 500+ terabytes of data each day

20IT261- R Programming for Data Science 01/19/25 24


19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 24
ARTIFICIAL INTELLIGENCE
John McCarthy was an American computer scientist coined the

term “Artificial Intelligence” in 1956

AI is a branch of computer science dealing with the simulation of

intelligent behavior in computers

The capability of a machine to imitate intelligent human behavior

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 25


Types of AI

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 26


TYPES OF AI
here
We
are

27
SREC
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 27
Machine Learning
Machine learning can either be considered a sub-field or one of the tools of
artificial intelligence, is providing machines with the capability of learning
from experience.

Experience for machines comes in the form of data. Data that is used to
teach machines is called training data.

Machine learning algorithms, also called “learners”, take both the known
input and output (training data) to figure out a model for the program which
converts input to output.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 28
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 29
What is Machine Learning ?
Machine learning is an application of artificial intelligence(AI) that
provides systems the ability to automatically learn and improve from
experience without being programmed.
Definition by Tom Mitchell:
Machine Learning is the study of algorithms that
• improve their performance P
• at some task T
• with experience E.
A well-defined learning task is given by <P , T , E>.

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 30


Defining the Learning Task
T: Playing checkers
P: Percentage of games won against an arbitrary opponent
E: Playing practice games against itself

T: Categorize email messages as spam or legitimate.


P: Percentage of email messages correctly classified.
E: Database of emails, some with human-given labels

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 31


EXAMPLE

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 32


Types of Machine Learning
Supervised learning
Given: training data + desired outputs (labels)
Classification : Output is discrete variable (eg., cat/dog)
Regression : Output is continuous (eg., price/temperature)

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 33


Algorithms Used:
Used
CLASSIFICATION:
[Link] Bayes
[Link](Support Vector Machines)
[Link] Decision Forest
REGRESSION:
[Link] Regression
[Link] Regression

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 34


Unsupervised learning
Given: training data (without desired outputs)
Clustering : It groups the similar instances
Dimensionality Reduction : It reduces the number of input variables in
training data.

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 35


Algorithms Used
CLUSTERING
a.K-means
[Link]
DIMENSIONALITY REDUCTION
[Link]
[Link]

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 36


Differences

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 37


Differences

A B
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 38
SEMI SUPERVISED LEARNING
•Given: training data + a few desired outputs
•Combination of supervised and unsupervised models, it learns from a
dataset that includes both labeled and unlabeled data.(eg.,Speech
analysis).

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 39


Reinforcement learning
•Reinforcement learning is also known as Active learning
•Rewards from sequence of actions
•An agent interacts with an environment and watches the result of the
interaction.

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 40


19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 41
Algorithms Used For
Reinforcement learning
•Markov Decision Process
•Approximate Dynamic Programming
•Brute Force

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 42


What is Data science
Data science is the field of study that combines domain

expertise, programming skills, and knowledge of mathematics

and statistics to extract meaningful insights from data.

Data Science is a blend of various tools, algorithms, and

machine learning principles with the goal to discover hidden

patterns from the raw data

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 43


Data Science
Data science is a collection of techniques used to extract value
from data.

Data science techniques rely on finding useful patterns,


connections, and relationship within data

Data science is also commonly referred to as knowledge


discovery, machine learning, predictive analytics, and data mining

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 44


Data Science Models

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 45


Types of Data
Structured Data
◦ Quantitative data represented in table

◦ Eg for structured data are Excel files or SQL databases

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 46


Unstructured Data

◦ Qualitative data

◦ Not organized in a pre-defined manner

◦ Eg for unstructured data are Text data, image, audio, video

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 47


Types of Data

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 48


19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 49
DATA SOURCES

From
Where data
comes

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 50


DATA SOURCES
Company data Open data
• Collected by companies • Free, open data sources
• Helps them make data-driven decisions • Can be used, shared, and built-on by anyone

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 51


COMPANY DATA
WEB DATA CUSTOMER DATA LOGISTICS DATA

SURVEY
DATA

FINANCIAL TRANSACTIONS

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 52


OPEN DATA
Open data is data that can be freely used, re-used and redistributed by anyone
Public APIs Public APIs

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 53


OTHER DATA TYPES

Image data

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 54


OTHER DATA TYPES

Text data
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 55
OTHER DATA TYPES

Network data

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 56


Data Science impression in the
domain

01/19/25 57
20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT
19/01/25 57
Data Science Jobs Roles
Most prominent Data Scientist job titles are:
◦ Data Scientist

◦ Data Engineer

◦ Data Analyst

◦ Statistician

◦ Data Architect

◦ Data Admin

◦ Business Analyst

◦ Data/Analytics Manager

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 58


Data Analyst
Data Analyst usually explains what is going on by processing history of the data.

Data Analyst Job Description

[Link] reports

[Link] patterns

[Link] with Stakeholders

[Link] data and setting up infrastructure


*Harvard Business Review has declared data science the Hot job of the 21st century, and IBM
predicts demand for data scientists will rise 28% by 2020.

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 59


Data Scientist
Data Scientist not only does the exploratory analysis
to discover insights from it, but also uses various
advanced machine learning algorithms to identify the
occurrence of a particular event in the future.
◦ A Data Scientist will look at the data from many angles,

sometimes angles not known earlier.

20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 19/01/25 60


Data Scientist Job Description
•Pulling, merging and analyzing data
•Looking for patterns or trends
•Using a wide variety of tools like Tableau, Python, Hive, R,
Impala, PySpark, Excel, Hadoop, etc to develop and test new
algorithms
•Trying to simplify data problems and developing predictive
models
•Building data visualizations
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 61
Skill required for data analyst
Sound knowledge of Excel, SQL, R, and Python.

Communication and Data visualization skills.

In-depth knowledge of Data wrangling skills.

Mathematics and Statistical skills.

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 62


Data Science Process

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 63


Data Science Process
1. Discovery:

Discovery step involves acquiring data from all the identified internal & external sources which helps you to
answer the business question

2. Preparation:

Data can have lots of inconsistencies like missing value, blank columns, incorrect data format which needs to be
cleaned. Need to process, explore, and condition data before modeling.

3. Model Planning

Need to determine the method and technique to draw the relation between input variables. Planning for a model
is performed by using different statistical formulas and visualization tools

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 64


Data Science Process
4. Model Building:

In this step, the actual model building process starts. Here, Data scientist distributes datasets for training and testing. Techniques
like association, classification, and clustering are applied to the training data set. The model once prepared is tested against the
"testing" dataset.

5. Operationalize:

In this stage, deliver the final baselined model with reports, code, and technical documents. Model is deployed into a real-time
production environment after thorough testing.

6. Communicate Results

In this stage, the key findings are communicated to all stakeholders. This helps you to decide if the results of the project are a
success or a failure based on the inputs from the model.

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 65


Data Science Process *
The standard data science process involves

(1) understanding the problem,

(2) preparing the data samples,

(3) developing the model,

(4) applying the model on a dataset to see how the model may work in
the real world,

(5) deploying and maintaining the models *-Content from Book

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 66


19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 67
Prior Knowledge
The prior knowledge step in the data science process helps to

define what problem is being solved, how it fits in the business

context, and what data is needed in order to solve the problem

Prior knowledge refers to information that is already known

about a subject

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 68


Prior Knowledge

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 69


Example for Data
Identifiers

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 70


Data Science Classification: Data
Science Tasks

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 71


19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 72
What is Data Mining?
•Data mining is the process of discovering interesting patterns and
knowledge from large amounts of data.

•The data sources can include databases, data warehouses, the Web, other
information repositories, or data that are streamed into the system
dynamically.

•The mining of gold from rocks or sand, we say gold mining instead of rock
or sand mining. Likewise many people treat data mining as a synonym for
another popularly used term, knowledge discovery from data, or KDD

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 73


Alternate terms for data mining
•Knowledge mining from data
•Knowledge extraction
•Data / pattern analysis
•Data Archaeology
•Data dredging

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 74


Steps Involved in KDD
Process

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 75


Data Pre-Processing

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 76


Data Quality: Why Preprocess the
Data?
Preprocessing of data is mainly to check the data quality. The quality can be
checked by the following
Accuracy: To check whether the data entered is correct or not.
Completeness: To check whether the data is available or not recorded.
Consistency: To check whether the same data is kept in all the places that do or
do not match.
Timeliness: The data should be updated correctly.
Believability: The data should be trustable.
Interpretability: The understandability of the data.

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 77


Major Tasks in Data
Preprocessing

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 78


Major Tasks in Data
Preprocessing
outliers=exceptions!
Data cleaning
◦ Fill in missing values, smooth noisy data, identify or remove outliers, and resolve
inconsistencies

Data integration
◦ Integration of multiple databases, data cubes, or files

Data transformation
◦ Normalization and aggregation

Data reduction
◦ Obtains reduced representation in volume but produces the same or similar analytical
results

Data discretization
◦ Part of data reduction but with particular importance, especially for numerical data
Data Cleaning
Data in the Real World Is Dirty: Lots of potentially incorrect
data, e.g., instrument faulty, human or computer error,
transmission error
◦ incomplete: lacking attribute values, lacking certain
attributes of interest, or containing only aggregate data
◦ e.g., Occupation=“ ” (missing data)
◦ noisy: containing noise, errors, or outliers
◦ e.g., Salary=“−10” (an error)

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 80


◦ inconsistent: containing discrepancies in codes or
names, e.g.,
◦ Age=“42”, Birthday=“03/07/2010”
◦ Was rating “1, 2, 3”, now rating “A, B, C”
◦ discrepancy between duplicate records
◦ Intentional (e.g., disguised missing data)
◦ Jan. 1 as everyone’s birthday?

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 81


Incomplete (Missing) Data
Data is not always available
◦ E.g., many tuples have no recorded value for several attributes, such as
customer income in sales data
Missing data may be due to
◦ equipment malfunction
◦ inconsistent with other recorded data and thus deleted
◦ data not entered due to misunderstanding
◦ certain data may not be considered important at the time of entry
◦ not register history or changes of the data

Missing data may need to be inferred

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 82


How to Handle Missing Data?
Ignore the tuple

Fill in the missing value manually

Fill in it automatically with


◦ a global constant: e.g., “unknown”, a new class?!
◦ the attribute mean
◦ the attribute mean for all samples belonging to the same class: smarter
◦ the most probable value

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 83


How to Handle Missing
Data?

Age Income Religion Gender


23 24,200 Muslim M
39 ? Christian F
45 45,390 ? F

Fill missing values using aggregate functions (e.g., average) or probabilistic


estimates on global value distribution
E.g., put the average income here, or put the most probable income based on
the fact that the person is 39 years old
E.g., put the most frequent religion here
Noisy Data
Noise: random error or variance in a measured variable

Incorrect attribute values may be due to


◦ faulty data collection instruments

◦ data entry problems

◦ data transmission problems

◦ technology limitation

◦ inconsistency in the naming convention

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 85


How to Handle Noisy Data?
Binning
◦ first sort data and partition into (equal-frequency) bins
◦ then one can smooth by bin means, smooth by bin median, smooth by bin
boundaries, etc.
Regression
◦ smooth by fitting the data into regression functions
Clustering
◦ detect and remove outliers
Combined computer and human inspection
◦ detect suspicious values and check by humans (e.g., deal with possible
outliers)

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 86


Binning Example

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 87


Data Cleaning as a Process
Data discrepancy detection
◦ Use metadata (e.g., domain, range, dependency,
distribution)
◦ Check field overloading
◦ Check uniqueness rule, consecutive rule and null rule
◦ Use commercial tools
◦ Data scrubbing: use simple domain knowledge (e.g.,
postal code, spell-check) to detect errors and make
corrections
◦ Data auditing: by analyzing data to discover rules and
relationship to detect violate

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 88


Data Cleaning as a Process
Data migration and integration
◦ Data migration tools: allow transformations to be
specified. E.g:- male into “M” etc
◦ ETL (Extraction/Transformation/Loading) tools: allow
users to specify transformations through a graphical
user interface

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 89


Data Integration
Data Integration—the merging of data from multiple datastores.

Data Integration is a data preprocessing technique that combines data

from multiple heterogeneous data sources into a coherent data store

and provides a unified view of the data.

These sources may include multiple data cubes, databases, or flat files.

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 90


Issues in Data Integration
Schema integration and object matching

- entity identification problem (e.g:- cus_no and cust_id ?)

Resolve: Metadata- name, meaning, data type, and range of


values permitted for the attribute, and null rules for handling
blank, zero, or null values

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 91


Issues in Data Integration

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 92


Issues in Data Integration
Redundancy and Correlation Analysis
Redundancy is another important issue in data integration.

An attribute (such as annual revenue, for instance) may be


redundant if it can be “derived” from another attribute or set of
attributes.

Inconsistencies in attribute or dimension naming can also


cause redundancies in the resulting data set.

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 93


Issues in Data Integration

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 94


Correlation Analysis (Nominal
Data)
2
(Observed  Expected )
 
2 The larger the Χ2 value, the more likely the variables are related

Expected

Play Not play Sum


chess chess (row)
Like science fiction 250(90) 200(360) 450
Not like science 50(210) 1000(840) 1050 * From book
fiction
Sum(col.) 300 1200 1500
Expected=count(playchess)x count(Like science fiction)
n
(250  90) 2 (50  210) 2 (200  360) 2 (1000  840) 2
Expected=300x450  
2
   507.93
90 210 360 840
1500
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 95
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 96
Correlation Coefficient for Numeric
Data
 
n n
(ai  A)(bi  B ) (ai bi )  n A B
For numeric attributes, we can rA, B  i 1
 i 1
(n  1) A B (n  1) A B
evaluate the correlation between A B
two attributes, A and B, by
computing the correlation
coefficient

Also known as Pearson’s product


moment coefficient, named after
its inventer, Karl Pearson).

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 97


Covariance of Numeric Data
correlation and
covariance are two
similar measures for
assessing how much two
attributes change
together

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 98


Covariance of Numeric Data
For two attributes A and B that tend to change together, if A is larger than A (the
expected value of A), then B is likely to be larger than B (the expected value of B).
Therefore, the covariance between A and B is positive.

On the other hand, if one of the attributes tends to be above its expected value
when the other attribute is below its expected value, then the covariance of A and B
is negative.

If A and B are independent (i.e., they do not have correlation), then E(A . B)= E(A).E(B).
Therefore, the covariance is . However, the
converse is not true. Some pairs of random variables (attributes) may have a covariance of 0
but are not independent.

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 99


Covariance of Numeric Data

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 100


Example
Suppose two stocks A and B have the following values in one week: (2, 5),
(3, 8), (5, 10), (4, 11), (6, 14).

Question: If the stocks are affected by the same industry trends, will their
prices rise or fall together?
◦ E(A) = (2 + 3 + 5 + 4 + 6)/ 5 = 20/5 = 4

◦ E(B) = (5 + 8 + 10 + 11 + 14) /5 = 48/5 = 9.6

◦ Cov(A,B) = (2×5+3×8+5×10+4×11+6×14)/5 − 4 × 9.6 = 4

Thus, A and B rise together since Cov(A, B) > 0.

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 101


Issues in Data Integration
Tuple Duplication:

In addition to detecting redundancies between attributes, duplication


should also be detected at the tuple level (e.g., where there are two or
more identical tuples for a given unique data entry case).

The use of denormalized tables is another source of data redundancy.

Inconsistencies often arise between various duplicates, due to inaccurate


data entry or updating some but not all data occurrences.

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 102


Issues in Data Integration
Data Value Conflict Detection and Resolution

Data integration also involves the detection and resolution of data value conflicts. For
example, for the same real-world entity, attribute values from different sources may differ.

This may be due to differences in representation, scaling, or encoding. For instance, a weight
attribute may be stored in metric units in one system and British imperial units in another.

For a hotel chain, the price of rooms in different cities may involve not only different
currencies but also different services (e.g., free breakfast) and taxes

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 103


Data Transformation and Data
Discretization

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 104


Data Transformation and Data
Discretization

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 105


Data Transformation

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 106


Normalization
Min-max normalization: to [new_minA, new_maxA]
v  minA
v'  (new _ maxA  new _ minA)  new _ minA
maxA  minA

Ex. Let income range $12,000 to $98,000 normalized to [0.0, 1.0]. Then $73,000 is mapped to

73,600  12,000
(1.0  0)  0 0.716
98,000  12,000

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 107


Normalization
Z-score normalization (μ: mean, σ: standard deviation):

v  A
v' 
 A

Ex. Let μ = 54,000, σ = 16,000. Then

73,600  54,000
a value of $73,600 for income 1.225
16,000

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 108


Normalization
Normalization by decimal scaling

v
v'  j Where j is the smallest integer such that Max(|ν’|) < 1
10

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 109


19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 110
Data Reduction
Data reduction techniques can be applied to obtain a reduced
representation of the data set that is much smaller in volume, yet
closely maintains the integrity of the original data

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 111


Data Reduction
Data Reduction Methods

•Data cube aggregation

•Dimensionality reduction: e.g., remove unimportant attributes

•Data compression

•Numerosity reduction: e.g., fit data into models

•Discretization and concept hierarchy generation

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 112


Data Reduction Strategies
•Dimensionality reduction is the process of reducing the number of
random variables or attributes

•Numerosity reduction techniques replace the original data volume by


alternative, smaller forms of data representation

•Data Compression, transformations are applied so as to obtain a


reduced or “compressed” representation of the original data.

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 113


Data Reduction Strategies

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 114


Principal Components Analysis
Principal components analysis (PCA; also called the Karhunen-Loeve, or K-
L, method)

Basic procedure :

1. The input data are normalized so that each attribute falls within the same range. This step
helps ensure that attributes with large domains will not dominate attributes with smaller
domains.

2. PCA computes k orthonormal vectors that provide a basis for the normalized input data.
These are unit vectors that each point in a direction perpendicular to the others. These vectors
are referred to as the principal components. The input data are a linear combination of the
principal components.

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 115


Principal Components Analysis
3. The principal components are sorted in order of decreasing
“significance” or strength.

4. Because the components are sorted in decreasing order of


“significance,” the data size can be reduced by eliminating the weaker
components, that is, those with low variance. Using the strongest
principal components, it should be possible to reconstruct a good
approximation of the original data.

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 116


Steps for PCA Algorithm
Standardize the data

Calculate the covariance matrix

Calculate the eigenvectors and eigenvalues

Choose the principal components

Transform the data- The final step is to transform the original data
into the lower-dimensional space defined by the principal
components

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 117


Wavelet Transforms

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 118


Attribute Sub-set Selection
The large data set has many attributes, some of which are irrelevant
to data mining or some are redundant.

The core attribute subset selection reduces the data volume and
dimensionality.

The attribute subset selection reduces the volume of data by


eliminating redundant and irrelevant attributes.

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 119


How can we find a ‘good’ subset of the original
attributes

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 120


Stepwise forward selection

Stepwise backward elimination

Combination of forward selection and backward elimination

Decision tree induction

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 121


Numerosity Reduction
Numerosity Reduction is a data reduction technique which
replaces the original data by smaller form of data representation.

2 techniques for numerosity reduction- Parametric and Non-


Parametric methods.

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 122


Parametric Methods

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 123


Parametric Methods

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 124


Non-Parametric
Histogram

A histogram is a ‘graph’ that represents frequency distribution which


describes how often a value appears in the data.

Histogram uses the binning method and to represent data distribution


of an attribute.

It uses disjoint subset which we call as bin or buckets.

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 125


Histogram Example
Data for AllElectronics data set,
which contains prices for
regularly sold items.

1, 1, 5, 5, 5, 5, 5, 8, 8, 10, 10,
10, 10, 12, 14, 14, 14, 15, 15,
15, 15, 15, 15, 18, 18, 18, 18,
18, 18, 18, 18, 20, 20, 20, 20,
20, 20, 20, 21, 21, 21, 21, 25,
25, 25, 25, 25, 28, 28, 30, 30,
30.

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 126


Clustering

Clustering techniques groups the


similar objects from the data in
such a way that the objects in a
cluster are similar to each other
but they are dissimilar to objects
in another cluster.

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 127


Sampling

Sampling can be used for data reduction because it allows a large data set to be
represented by a much smaller random data sample (or subset).

Methods
Simple Random Sample Without Replacement of sizes
Simple Random Sample with Replacement of sizes
Cluster Sample
Stratified Sample

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 128


19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 129
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 130
Data Cube Aggregation

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 131


Data Compression

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 132


hierarchy
generation
Discretization, where the raw values of a numeric attribute
(e.g., age) are replaced by interval labels (e.g., 0–10, 11–20,
etc.) or conceptual labels (e.g., youth, adult, senior)

Concept hierarchy generation for nominal data, where


attributes such as street can be generalized to higher-level
concepts, like city or country.

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 133


MODELING
A model is the abstract representation of the data and the
relationships in a given dataset

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 134


Modeling: Training and Testing
Datasets
The modeling step creates a representative model inferred from the data.

The dataset used to create the model, with known attributes and target, is
called the training dataset.

The validity of the created model will also need to be checked with another
known dataset called the test dataset or validation dataset.

A standard rule of thumb is two-thirds of the data are to be used as training


and one-third as a test dataset.

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 135


Modeling: Learning Algorithms
Classification task many algorithms can be chosen from:
decision trees, rule induction, neural networks, Bayesian
models, k-NN, etc.

Decision tree techniques- Classification And Regression Tree


(CART), CHi-squared Automatic Interaction Detector (CHAID)

Regression

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 136


Modeling: Evaluation of the
Model
Accuracy

Prediction error

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 137


Modeling: Ensemble Modeling

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 138


Bagging
Boosting

19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 139


Thank You!!!
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 140

You might also like