SRI RAMAKRISHNA ENGINEERING
COLLEGE
[Educational Service : SNR Sons Charitable Trust]
[Autonomous Institution, Reaccredited by NAAC with ‘A+’ Grade]
[Approved by AICTE and Permanently Affiliated to Anna University, Chennai]
[ISO 9001:2015 Certified and all Eligible Programmes Accredited by NBA]
VATTAMALAIPALAYAM, N.G.G.O. COLONY POST, COIMBATORE – 641 022.
Department of Information Technology
20IT211- Data Science
Presentation by
[Link] Rani, AP([Link])/IT
COURSE OUTCOMES
20IT211- Data Science
Understand the basic concepts of data science and
CO1 PO1,PO2,PO12
data mining
CO2 Identify the techniques to explore and evaluate data PO3,PO5,PO12
Apply various data mining algorithms for real time PO2,PO3,PO5,P
CO3
applications O12
Implement the concepts of clustering and model
CO4 PO3,PO5,PO12
evaluation
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 2
20IT211- Data Science
Module I : INTRODUCTION 9 hours
What is data science – Case for data science – Data
science classification – Data science algorithms – Data
science process – Prior knowledge – Data preprocessing –
Data cleaning – Data integration – Data reduction – Data
transformation and data discretization – Feature selection
– Data sampling – Modeling – Application.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 3
20IT211- Data Science
Module II : DATA EXPLORATION AND VISUALIZATION
9 hours
Objectives of Data exploration – Datasets – Descriptive
statistics – Data Visualization – Univariate visualization –
Multivariate visualization – Visualizing high dimensional
data – Roadmap for data exploration.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 4
20IT211- Data Science
Module III : CLASSIFICATION AND ASSOCIATION
ANALYSIS 18 hours
Basic concepts of Classification – Decision tree induction –
Bayes classification methods – Rule based classification –
Techniques to improve classification accuracy – Support vector
machines – Regression methods: Linear regression – Logistic
regression – Association analysis: Frequent Item set mining
methods – Pattern evaluation methods.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 5
20IT211- Data Science
Module IV : CLUSTERING AND MODEL EVALUATION 9
hours
Basic concepts and methods in cluster analysis – Partitioning
methods – Density based methods – Model evaluation:
Confusion matrix – Receiver Operator Characteristics (ROC) and
Area under the Curve (AUC) – Lift curves – Evaluating the
Predictions – Implementation
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 6
TEXTBOOKS
1. Vijay Kotu and Bala Deshpande, “Data Science Concepts and Practice”, 2
Edition, Morgan Kaufmann Publishers, 2019.
2. Jiawei Han, Micheline Kamber and Jian Pei, “Data Mining: Concepts and
Techniques”, 3 Edition, Morgan Kaufmann Publishers, 2012.
3. Cathy O’Neil and Rachel Schutt, “Doing Data Science, Straight Talk From
The Frontline”, O’Reilly, 2016.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 7
Reference(s)
1. Mohammed J. Zaki and Wagner Miera Jr, “Data Mining and Analysis:
Fundamental Concepts and Algorithms”, Cambridge University Press, 2014.
2. Matt Harrison, “Learning the Pandas Library: Python Tools for Data
Munging, Analysis and Visualization O’Reilly, 2016.
3. Joel Grus, “Data Science from Scratch: First Principles with Python”, O’Reilly
Media, 2015. 4. Wes McKinney, “Python for Data Analysis: Data Wrangling
with Pandas, NumPy, and IPython”, O’Reilly Media, 2012
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 8
WEB REFERENCES
1. [Link]
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 9
Acknowledgement
Resources are taken from the internet and textbooks
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 10
What is Data
Data are plain
• Data is afacts
collection of raw,
unorganised facts and details like
text, observations, figures, symbols
and description of things etc.
• In other words, data does not carry
any specific purpose and has no
significance by itself.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 11
What is Information
Information is the processed, organised and structured data.
It provides context for data and enables decision making.
For example, a single customer’s sale at a restaurant is data – this becomes information
when the business is able to identify the most popular or least popular dish.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 12
What is Information
Information is defined as classified or organized data that has some meaningful value
for the user.
Information is also the processed data used to make decisions and take action.
Processed data must meet the following criteria for it to be of any significant use in
decision-making:
Accuracy: The information must be accurate.
Completeness: The information must be complete.
Timeliness: The information must be available when it’s needed.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 13
Example
Marks / student’s test score
Roll numbers Data
Data
Report card/sheet /average score of the class- Information.
Information
Other examples for information are pay-slips, schedules, reports, worksheet,
bar charts, invoices and account returns etc
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 14
Computer Data
Computer data is information processed or stored by a
computer.
◦ This information may be in the form of text documents, images, audio
clips, software programs, or other types of data.
Computer data may be processed by the computer’s CPU and
is stored in files and folders on the computer’s hard disk.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 15
Types of Data
•Quantitative data is numerical and can be counted,
quantified, and mathematically analyzed (e.g., GPAs,
standardized test scores, attendance patterns). Quantitative
data is often considered a highly reliable source of information.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 16
•Qualitative data is usually non-numerical and used to
provide meaning and understanding. Student narratives
describing their reasons for participating in your program
each month are examples of qualitative data. Qualitative data
is believed to have great validity and depth.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 17
QUANTITATIVE DATA
Quantitative data is information that you can measure. It’s
numbers –something you can count. Because it’s countable it can
be reliable evidence. Examples include:
◦ How many people took part?
◦ How much did it cost?
◦ How long did it run for?
◦ Average attendance at each programme session?
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 18
QUALITATIVE DATA
Qualitative data is information about qualities, you can’t count it.
That is, it’s information about how people feel about something.
Examples include:
◦ Sharing what people like about a programme.
◦ How they think it could be improved.
◦ What difference it has made to their lives.
◦ Whether they would recommend the programme to others.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 19
Data
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 20
Data
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 21
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 22
Example
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 23
Example
Google data centre normally holds petabytes to exabytes(10 18 bytes)of data
Google currently processes over 20 petabytes of data per day
Facebook: 500+ terabytes of data each day
20IT261- R Programming for Data Science 01/19/25 24
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 24
ARTIFICIAL INTELLIGENCE
John McCarthy was an American computer scientist coined the
term “Artificial Intelligence” in 1956
AI is a branch of computer science dealing with the simulation of
intelligent behavior in computers
The capability of a machine to imitate intelligent human behavior
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 25
Types of AI
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 26
TYPES OF AI
here
We
are
27
SREC
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 27
Machine Learning
Machine learning can either be considered a sub-field or one of the tools of
artificial intelligence, is providing machines with the capability of learning
from experience.
Experience for machines comes in the form of data. Data that is used to
teach machines is called training data.
Machine learning algorithms, also called “learners”, take both the known
input and output (training data) to figure out a model for the program which
converts input to output.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 28
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 29
What is Machine Learning ?
Machine learning is an application of artificial intelligence(AI) that
provides systems the ability to automatically learn and improve from
experience without being programmed.
Definition by Tom Mitchell:
Machine Learning is the study of algorithms that
• improve their performance P
• at some task T
• with experience E.
A well-defined learning task is given by <P , T , E>.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 30
Defining the Learning Task
T: Playing checkers
P: Percentage of games won against an arbitrary opponent
E: Playing practice games against itself
T: Categorize email messages as spam or legitimate.
P: Percentage of email messages correctly classified.
E: Database of emails, some with human-given labels
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 31
EXAMPLE
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 32
Types of Machine Learning
Supervised learning
Given: training data + desired outputs (labels)
Classification : Output is discrete variable (eg., cat/dog)
Regression : Output is continuous (eg., price/temperature)
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 33
Algorithms Used:
Used
CLASSIFICATION:
[Link] Bayes
[Link](Support Vector Machines)
[Link] Decision Forest
REGRESSION:
[Link] Regression
[Link] Regression
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 34
Unsupervised learning
Given: training data (without desired outputs)
Clustering : It groups the similar instances
Dimensionality Reduction : It reduces the number of input variables in
training data.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 35
Algorithms Used
CLUSTERING
a.K-means
[Link]
DIMENSIONALITY REDUCTION
[Link]
[Link]
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 36
Differences
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 37
Differences
A B
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 38
SEMI SUPERVISED LEARNING
•Given: training data + a few desired outputs
•Combination of supervised and unsupervised models, it learns from a
dataset that includes both labeled and unlabeled data.(eg.,Speech
analysis).
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 39
Reinforcement learning
•Reinforcement learning is also known as Active learning
•Rewards from sequence of actions
•An agent interacts with an environment and watches the result of the
interaction.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 40
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 41
Algorithms Used For
Reinforcement learning
•Markov Decision Process
•Approximate Dynamic Programming
•Brute Force
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 42
What is Data science
Data science is the field of study that combines domain
expertise, programming skills, and knowledge of mathematics
and statistics to extract meaningful insights from data.
Data Science is a blend of various tools, algorithms, and
machine learning principles with the goal to discover hidden
patterns from the raw data
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 43
Data Science
Data science is a collection of techniques used to extract value
from data.
Data science techniques rely on finding useful patterns,
connections, and relationship within data
Data science is also commonly referred to as knowledge
discovery, machine learning, predictive analytics, and data mining
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 44
Data Science Models
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 45
Types of Data
Structured Data
◦ Quantitative data represented in table
◦ Eg for structured data are Excel files or SQL databases
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 46
Unstructured Data
◦ Qualitative data
◦ Not organized in a pre-defined manner
◦ Eg for unstructured data are Text data, image, audio, video
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 47
Types of Data
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 48
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 49
DATA SOURCES
From
Where data
comes
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 50
DATA SOURCES
Company data Open data
• Collected by companies • Free, open data sources
• Helps them make data-driven decisions • Can be used, shared, and built-on by anyone
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 51
COMPANY DATA
WEB DATA CUSTOMER DATA LOGISTICS DATA
SURVEY
DATA
FINANCIAL TRANSACTIONS
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 52
OPEN DATA
Open data is data that can be freely used, re-used and redistributed by anyone
Public APIs Public APIs
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 53
OTHER DATA TYPES
Image data
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 54
OTHER DATA TYPES
Text data
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 55
OTHER DATA TYPES
Network data
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 56
Data Science impression in the
domain
01/19/25 57
20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT
19/01/25 57
Data Science Jobs Roles
Most prominent Data Scientist job titles are:
◦ Data Scientist
◦ Data Engineer
◦ Data Analyst
◦ Statistician
◦ Data Architect
◦ Data Admin
◦ Business Analyst
◦ Data/Analytics Manager
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 58
Data Analyst
Data Analyst usually explains what is going on by processing history of the data.
Data Analyst Job Description
[Link] reports
[Link] patterns
[Link] with Stakeholders
[Link] data and setting up infrastructure
*Harvard Business Review has declared data science the Hot job of the 21st century, and IBM
predicts demand for data scientists will rise 28% by 2020.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 59
Data Scientist
Data Scientist not only does the exploratory analysis
to discover insights from it, but also uses various
advanced machine learning algorithms to identify the
occurrence of a particular event in the future.
◦ A Data Scientist will look at the data from many angles,
sometimes angles not known earlier.
20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 19/01/25 60
Data Scientist Job Description
•Pulling, merging and analyzing data
•Looking for patterns or trends
•Using a wide variety of tools like Tableau, Python, Hive, R,
Impala, PySpark, Excel, Hadoop, etc to develop and test new
algorithms
•Trying to simplify data problems and developing predictive
models
•Building data visualizations
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 61
Skill required for data analyst
Sound knowledge of Excel, SQL, R, and Python.
Communication and Data visualization skills.
In-depth knowledge of Data wrangling skills.
Mathematics and Statistical skills.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 62
Data Science Process
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 63
Data Science Process
1. Discovery:
Discovery step involves acquiring data from all the identified internal & external sources which helps you to
answer the business question
2. Preparation:
Data can have lots of inconsistencies like missing value, blank columns, incorrect data format which needs to be
cleaned. Need to process, explore, and condition data before modeling.
3. Model Planning
Need to determine the method and technique to draw the relation between input variables. Planning for a model
is performed by using different statistical formulas and visualization tools
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 64
Data Science Process
4. Model Building:
In this step, the actual model building process starts. Here, Data scientist distributes datasets for training and testing. Techniques
like association, classification, and clustering are applied to the training data set. The model once prepared is tested against the
"testing" dataset.
5. Operationalize:
In this stage, deliver the final baselined model with reports, code, and technical documents. Model is deployed into a real-time
production environment after thorough testing.
6. Communicate Results
In this stage, the key findings are communicated to all stakeholders. This helps you to decide if the results of the project are a
success or a failure based on the inputs from the model.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 65
Data Science Process *
The standard data science process involves
(1) understanding the problem,
(2) preparing the data samples,
(3) developing the model,
(4) applying the model on a dataset to see how the model may work in
the real world,
(5) deploying and maintaining the models *-Content from Book
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 66
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 67
Prior Knowledge
The prior knowledge step in the data science process helps to
define what problem is being solved, how it fits in the business
context, and what data is needed in order to solve the problem
Prior knowledge refers to information that is already known
about a subject
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 68
Prior Knowledge
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 69
Example for Data
Identifiers
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 70
Data Science Classification: Data
Science Tasks
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 71
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 72
What is Data Mining?
•Data mining is the process of discovering interesting patterns and
knowledge from large amounts of data.
•The data sources can include databases, data warehouses, the Web, other
information repositories, or data that are streamed into the system
dynamically.
•The mining of gold from rocks or sand, we say gold mining instead of rock
or sand mining. Likewise many people treat data mining as a synonym for
another popularly used term, knowledge discovery from data, or KDD
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 73
Alternate terms for data mining
•Knowledge mining from data
•Knowledge extraction
•Data / pattern analysis
•Data Archaeology
•Data dredging
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 74
Steps Involved in KDD
Process
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 75
Data Pre-Processing
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 76
Data Quality: Why Preprocess the
Data?
Preprocessing of data is mainly to check the data quality. The quality can be
checked by the following
Accuracy: To check whether the data entered is correct or not.
Completeness: To check whether the data is available or not recorded.
Consistency: To check whether the same data is kept in all the places that do or
do not match.
Timeliness: The data should be updated correctly.
Believability: The data should be trustable.
Interpretability: The understandability of the data.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 77
Major Tasks in Data
Preprocessing
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 78
Major Tasks in Data
Preprocessing
outliers=exceptions!
Data cleaning
◦ Fill in missing values, smooth noisy data, identify or remove outliers, and resolve
inconsistencies
Data integration
◦ Integration of multiple databases, data cubes, or files
Data transformation
◦ Normalization and aggregation
Data reduction
◦ Obtains reduced representation in volume but produces the same or similar analytical
results
Data discretization
◦ Part of data reduction but with particular importance, especially for numerical data
Data Cleaning
Data in the Real World Is Dirty: Lots of potentially incorrect
data, e.g., instrument faulty, human or computer error,
transmission error
◦ incomplete: lacking attribute values, lacking certain
attributes of interest, or containing only aggregate data
◦ e.g., Occupation=“ ” (missing data)
◦ noisy: containing noise, errors, or outliers
◦ e.g., Salary=“−10” (an error)
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 80
◦ inconsistent: containing discrepancies in codes or
names, e.g.,
◦ Age=“42”, Birthday=“03/07/2010”
◦ Was rating “1, 2, 3”, now rating “A, B, C”
◦ discrepancy between duplicate records
◦ Intentional (e.g., disguised missing data)
◦ Jan. 1 as everyone’s birthday?
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 81
Incomplete (Missing) Data
Data is not always available
◦ E.g., many tuples have no recorded value for several attributes, such as
customer income in sales data
Missing data may be due to
◦ equipment malfunction
◦ inconsistent with other recorded data and thus deleted
◦ data not entered due to misunderstanding
◦ certain data may not be considered important at the time of entry
◦ not register history or changes of the data
Missing data may need to be inferred
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 82
How to Handle Missing Data?
Ignore the tuple
Fill in the missing value manually
Fill in it automatically with
◦ a global constant: e.g., “unknown”, a new class?!
◦ the attribute mean
◦ the attribute mean for all samples belonging to the same class: smarter
◦ the most probable value
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 83
How to Handle Missing
Data?
Age Income Religion Gender
23 24,200 Muslim M
39 ? Christian F
45 45,390 ? F
Fill missing values using aggregate functions (e.g., average) or probabilistic
estimates on global value distribution
E.g., put the average income here, or put the most probable income based on
the fact that the person is 39 years old
E.g., put the most frequent religion here
Noisy Data
Noise: random error or variance in a measured variable
Incorrect attribute values may be due to
◦ faulty data collection instruments
◦ data entry problems
◦ data transmission problems
◦ technology limitation
◦ inconsistency in the naming convention
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 85
How to Handle Noisy Data?
Binning
◦ first sort data and partition into (equal-frequency) bins
◦ then one can smooth by bin means, smooth by bin median, smooth by bin
boundaries, etc.
Regression
◦ smooth by fitting the data into regression functions
Clustering
◦ detect and remove outliers
Combined computer and human inspection
◦ detect suspicious values and check by humans (e.g., deal with possible
outliers)
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 86
Binning Example
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 87
Data Cleaning as a Process
Data discrepancy detection
◦ Use metadata (e.g., domain, range, dependency,
distribution)
◦ Check field overloading
◦ Check uniqueness rule, consecutive rule and null rule
◦ Use commercial tools
◦ Data scrubbing: use simple domain knowledge (e.g.,
postal code, spell-check) to detect errors and make
corrections
◦ Data auditing: by analyzing data to discover rules and
relationship to detect violate
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 88
Data Cleaning as a Process
Data migration and integration
◦ Data migration tools: allow transformations to be
specified. E.g:- male into “M” etc
◦ ETL (Extraction/Transformation/Loading) tools: allow
users to specify transformations through a graphical
user interface
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 89
Data Integration
Data Integration—the merging of data from multiple datastores.
Data Integration is a data preprocessing technique that combines data
from multiple heterogeneous data sources into a coherent data store
and provides a unified view of the data.
These sources may include multiple data cubes, databases, or flat files.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 90
Issues in Data Integration
Schema integration and object matching
- entity identification problem (e.g:- cus_no and cust_id ?)
Resolve: Metadata- name, meaning, data type, and range of
values permitted for the attribute, and null rules for handling
blank, zero, or null values
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 91
Issues in Data Integration
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 92
Issues in Data Integration
Redundancy and Correlation Analysis
Redundancy is another important issue in data integration.
An attribute (such as annual revenue, for instance) may be
redundant if it can be “derived” from another attribute or set of
attributes.
Inconsistencies in attribute or dimension naming can also
cause redundancies in the resulting data set.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 93
Issues in Data Integration
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 94
Correlation Analysis (Nominal
Data)
2
(Observed Expected )
2 The larger the Χ2 value, the more likely the variables are related
Expected
Play Not play Sum
chess chess (row)
Like science fiction 250(90) 200(360) 450
Not like science 50(210) 1000(840) 1050 * From book
fiction
Sum(col.) 300 1200 1500
Expected=count(playchess)x count(Like science fiction)
n
(250 90) 2 (50 210) 2 (200 360) 2 (1000 840) 2
Expected=300x450
2
507.93
90 210 360 840
1500
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 95
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 96
Correlation Coefficient for Numeric
Data
n n
(ai A)(bi B ) (ai bi ) n A B
For numeric attributes, we can rA, B i 1
i 1
(n 1) A B (n 1) A B
evaluate the correlation between A B
two attributes, A and B, by
computing the correlation
coefficient
Also known as Pearson’s product
moment coefficient, named after
its inventer, Karl Pearson).
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 97
Covariance of Numeric Data
correlation and
covariance are two
similar measures for
assessing how much two
attributes change
together
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 98
Covariance of Numeric Data
For two attributes A and B that tend to change together, if A is larger than A (the
expected value of A), then B is likely to be larger than B (the expected value of B).
Therefore, the covariance between A and B is positive.
On the other hand, if one of the attributes tends to be above its expected value
when the other attribute is below its expected value, then the covariance of A and B
is negative.
If A and B are independent (i.e., they do not have correlation), then E(A . B)= E(A).E(B).
Therefore, the covariance is . However, the
converse is not true. Some pairs of random variables (attributes) may have a covariance of 0
but are not independent.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 99
Covariance of Numeric Data
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 100
Example
Suppose two stocks A and B have the following values in one week: (2, 5),
(3, 8), (5, 10), (4, 11), (6, 14).
Question: If the stocks are affected by the same industry trends, will their
prices rise or fall together?
◦ E(A) = (2 + 3 + 5 + 4 + 6)/ 5 = 20/5 = 4
◦ E(B) = (5 + 8 + 10 + 11 + 14) /5 = 48/5 = 9.6
◦ Cov(A,B) = (2×5+3×8+5×10+4×11+6×14)/5 − 4 × 9.6 = 4
Thus, A and B rise together since Cov(A, B) > 0.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 101
Issues in Data Integration
Tuple Duplication:
In addition to detecting redundancies between attributes, duplication
should also be detected at the tuple level (e.g., where there are two or
more identical tuples for a given unique data entry case).
The use of denormalized tables is another source of data redundancy.
Inconsistencies often arise between various duplicates, due to inaccurate
data entry or updating some but not all data occurrences.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 102
Issues in Data Integration
Data Value Conflict Detection and Resolution
Data integration also involves the detection and resolution of data value conflicts. For
example, for the same real-world entity, attribute values from different sources may differ.
This may be due to differences in representation, scaling, or encoding. For instance, a weight
attribute may be stored in metric units in one system and British imperial units in another.
For a hotel chain, the price of rooms in different cities may involve not only different
currencies but also different services (e.g., free breakfast) and taxes
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 103
Data Transformation and Data
Discretization
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 104
Data Transformation and Data
Discretization
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 105
Data Transformation
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 106
Normalization
Min-max normalization: to [new_minA, new_maxA]
v minA
v' (new _ maxA new _ minA) new _ minA
maxA minA
Ex. Let income range $12,000 to $98,000 normalized to [0.0, 1.0]. Then $73,000 is mapped to
73,600 12,000
(1.0 0) 0 0.716
98,000 12,000
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 107
Normalization
Z-score normalization (μ: mean, σ: standard deviation):
v A
v'
A
Ex. Let μ = 54,000, σ = 16,000. Then
73,600 54,000
a value of $73,600 for income 1.225
16,000
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 108
Normalization
Normalization by decimal scaling
v
v' j Where j is the smallest integer such that Max(|ν’|) < 1
10
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 109
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 110
Data Reduction
Data reduction techniques can be applied to obtain a reduced
representation of the data set that is much smaller in volume, yet
closely maintains the integrity of the original data
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 111
Data Reduction
Data Reduction Methods
•Data cube aggregation
•Dimensionality reduction: e.g., remove unimportant attributes
•Data compression
•Numerosity reduction: e.g., fit data into models
•Discretization and concept hierarchy generation
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 112
Data Reduction Strategies
•Dimensionality reduction is the process of reducing the number of
random variables or attributes
•Numerosity reduction techniques replace the original data volume by
alternative, smaller forms of data representation
•Data Compression, transformations are applied so as to obtain a
reduced or “compressed” representation of the original data.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 113
Data Reduction Strategies
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 114
Principal Components Analysis
Principal components analysis (PCA; also called the Karhunen-Loeve, or K-
L, method)
Basic procedure :
1. The input data are normalized so that each attribute falls within the same range. This step
helps ensure that attributes with large domains will not dominate attributes with smaller
domains.
2. PCA computes k orthonormal vectors that provide a basis for the normalized input data.
These are unit vectors that each point in a direction perpendicular to the others. These vectors
are referred to as the principal components. The input data are a linear combination of the
principal components.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 115
Principal Components Analysis
3. The principal components are sorted in order of decreasing
“significance” or strength.
4. Because the components are sorted in decreasing order of
“significance,” the data size can be reduced by eliminating the weaker
components, that is, those with low variance. Using the strongest
principal components, it should be possible to reconstruct a good
approximation of the original data.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 116
Steps for PCA Algorithm
Standardize the data
Calculate the covariance matrix
Calculate the eigenvectors and eigenvalues
Choose the principal components
Transform the data- The final step is to transform the original data
into the lower-dimensional space defined by the principal
components
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 117
Wavelet Transforms
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 118
Attribute Sub-set Selection
The large data set has many attributes, some of which are irrelevant
to data mining or some are redundant.
The core attribute subset selection reduces the data volume and
dimensionality.
The attribute subset selection reduces the volume of data by
eliminating redundant and irrelevant attributes.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 119
How can we find a ‘good’ subset of the original
attributes
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 120
Stepwise forward selection
Stepwise backward elimination
Combination of forward selection and backward elimination
Decision tree induction
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 121
Numerosity Reduction
Numerosity Reduction is a data reduction technique which
replaces the original data by smaller form of data representation.
2 techniques for numerosity reduction- Parametric and Non-
Parametric methods.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 122
Parametric Methods
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 123
Parametric Methods
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 124
Non-Parametric
Histogram
A histogram is a ‘graph’ that represents frequency distribution which
describes how often a value appears in the data.
Histogram uses the binning method and to represent data distribution
of an attribute.
It uses disjoint subset which we call as bin or buckets.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 125
Histogram Example
Data for AllElectronics data set,
which contains prices for
regularly sold items.
1, 1, 5, 5, 5, 5, 5, 8, 8, 10, 10,
10, 10, 12, 14, 14, 14, 15, 15,
15, 15, 15, 15, 18, 18, 18, 18,
18, 18, 18, 18, 20, 20, 20, 20,
20, 20, 20, 21, 21, 21, 21, 25,
25, 25, 25, 25, 28, 28, 30, 30,
30.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 126
Clustering
Clustering techniques groups the
similar objects from the data in
such a way that the objects in a
cluster are similar to each other
but they are dissimilar to objects
in another cluster.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 127
Sampling
Sampling can be used for data reduction because it allows a large data set to be
represented by a much smaller random data sample (or subset).
Methods
Simple Random Sample Without Replacement of sizes
Simple Random Sample with Replacement of sizes
Cluster Sample
Stratified Sample
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 128
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 129
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 130
Data Cube Aggregation
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 131
Data Compression
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 132
hierarchy
generation
Discretization, where the raw values of a numeric attribute
(e.g., age) are replaced by interval labels (e.g., 0–10, 11–20,
etc.) or conceptual labels (e.g., youth, adult, senior)
Concept hierarchy generation for nominal data, where
attributes such as street can be generalized to higher-level
concepts, like city or country.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 133
MODELING
A model is the abstract representation of the data and the
relationships in a given dataset
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 134
Modeling: Training and Testing
Datasets
The modeling step creates a representative model inferred from the data.
The dataset used to create the model, with known attributes and target, is
called the training dataset.
The validity of the created model will also need to be checked with another
known dataset called the test dataset or validation dataset.
A standard rule of thumb is two-thirds of the data are to be used as training
and one-third as a test dataset.
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 135
Modeling: Learning Algorithms
Classification task many algorithms can be chosen from:
decision trees, rule induction, neural networks, Bayesian
models, k-NN, etc.
Decision tree techniques- Classification And Regression Tree
(CART), CHi-squared Automatic Interaction Detector (CHAID)
Regression
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 136
Modeling: Evaluation of the
Model
Accuracy
Prediction error
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 137
Modeling: Ensemble Modeling
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 138
Bagging
Boosting
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 139
Thank You!!!
19/01/25 20IT211- DATA SCIENCE - [Link] RANI, AP([Link])/IT 140