0% found this document useful (0 votes)
6 views18 pages

Chapter 4

Chapter 4 focuses on data preprocessing, a crucial step in data mining that transforms raw data into a usable format by addressing issues like incompleteness and inconsistency. It outlines four main phases of data preprocessing: data cleaning, data integration, data transformation, and data reduction. The chapter emphasizes the importance of these processes in ensuring the quality of data for effective data mining results.

Uploaded by

biwam75142
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
6 views18 pages

Chapter 4

Chapter 4 focuses on data preprocessing, a crucial step in data mining that transforms raw data into a usable format by addressing issues like incompleteness and inconsistency. It outlines four main phases of data preprocessing: data cleaning, data integration, data transformation, and data reduction. The chapter emphasizes the importance of these processes in ensuring the quality of data for effective data mining results.

Uploaded by

biwam75142
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Chapter 4

Data
Preprocessing
CHAPTER OBJECTIVES

1. To understand the need for data preprocessing.


2. To identify different phases of data preprocessing such as data
cleaning, data integration,
3. Data transformation and data reduction
Need for Data Processing
• Data preprocessing is a data mining technique that involves
transformation of raw data into an understandable format, because real
world data can often be incomplete, inconsistent or even erroneous in
nature.
• Data preprocessing resolves such issues. Data preprocessing ensures
that further data mining process are free from errors. It is a prerequisite
preparation for data mining, it prepares raw data for the core processes.

These slides are designed to complement 'Data Mining and Data Warehousing' by Parteek Bhatia, published by Cambridge University Press.
Data Pre-processing example

These slides are designed to complement 'Data Mining and Data Warehousing' by Parteek Bhatia, published by Cambridge University Press.
Data Pre-processing Methods
Raw data is highly vulnerable
to missing values, noise and
inconsistency and the quality
of data affects the data
mining results.

The various stages in which


data preprocessing is
performed:

1. Data Cleaning
2. Data Integration
3. Data Transformation
4. Data Reduction
These slides are designed to complement 'Data Mining and Data Warehousing' by Parteek Bhatia, published by Cambridge University Press.
Data Pre-processing Methods
Data Cleaning
 Raw data or noisy data goes
through the process of
cleansing first. In Data
cleansing missing values are
filled, noisy data is
smoothened, inconsistencies
are resolved, outliers are
identified and removed in
order to clean the data.

These slides are designed to complement 'Data Mining and Data Warehousing' by Parteek Bhatia, published by Cambridge University Press.
Data Pre-processing Methods
Operations Performed
during Data Cleaning
 Handling of Missing Values
 Handling of Noisy Data
 Handling of Inconsistent
data

These slides are designed to complement 'Data Mining and Data Warehousing' by Parteek Bhatia, published by Cambridge University Press.
Data Pre-processing Methods
Data Integration
 One of the most necessary
steps taken during the data
analysis is Data Integration.
Data integration is a method
which combines data from
plethora of sources (such as
multiple databases, flat files
or data cubes) into a unified
data store.

These slides are designed to complement 'Data Mining and Data Warehousing' by Parteek Bhatia, published by Cambridge University Press.
Data Pre-processing Methods
Data Transformation
 Sometimes, the value of one
attribute may be small as
compared to other attributes,
then in this scenario, that attribute
will not have much influence on
mining of information, since the
values of this attribute were
smaller than other attributes and
the variation within the attribute
will also be small.
 Normalization and Standardization
are most popular and widely used
Data Transformation methods.
These slides are designed to complement 'Data Mining and Data Warehousing' by Parteek Bhatia, published by Cambridge University Press.
Data Pre-processing Methods
Data Reduction
 It is often seen that, when the complex
data analysis and mining processes are
carried out over humongous data sets,
they cost a very long time, henceforth
making the whole data mining or
analysis process unsuitable (or not
feasible). Data reduction techniques
come for the rescue in such scenarios.
Using data reduction techniques
dataset could be represented in a
reduced manner without actually
compromising the integrity of original
data.

These slides are designed to complement 'Data Mining and Data Warehousing' by Parteek Bhatia, published by Cambridge University Press.
Book Details
Table of Contents Published by
1. Beginning with machine learning Cambridge University, Press
2. Introduction to data mining (UK).
3. Beginning with Weka and R language Recommended as Six Best
New Data Warehousing
4. Data pre-processing
Books to read in 2020 and
5. Classification 43 Best Data Mining Books
6. Implementing classification in Weka and R of all time by
7. Cluster analysis [Link]
8. Implementing clustering with Weka and R
9. Association mining
10. Implementing association mining with Weka and R
11. Web mining and search engine
12. Operational data store and data warehouse
13. Data warehouse schema
14. Online analytical processing
15. Big data and NoSQL
Order Your Copy Today
 [Link]: Data Mining and Data Warehousing: Principles
and Practical Techniques eBook : Bhatia, Parteek: Kindle Store

[Link]

[Link]
Other Books from Parteek Bhatia:
Machine Learning with Python
Machine learning python
principles and practical
techniques | Pattern
recognition and machine
learning | Cambridge
University Press
Other Books from Parteek Bhatia:
Simplified Approach to DBMS
Table of Contents 19. Enterprise Database Products
1. Fundamentals of Database Management System 20. Beginning with SQL
2. The architecture of Database Management 21. Invoking SQL* Plus
System 22. Performing basic SQL
3. Data Models operations
4. Relational Database Management System 23. Basic SELECT statement
5. Relational Algebra and Calculus 24. Inbuilt functions
6. Entity-Relationship Model 25. Grouping of data
7. Conversion of ER Diagrams to Tables 26. Joining of tables
8. Normalization for refinement of data 27. Sub Queries
9. Physical Database design 28. Managing Tables
10. Transaction Management 29. Database objects, DCL and
11. Concurrency Control TCL statements
12. Security and Integrity of data 30. Pl/SQL Fundamentals and its
13. Recovery of data statements
14. Distributed Database 31. Error Handling
15. Object-Oriented Databases and expert system 32. Cursor Management
16. DBTG Model 33. Subprograms and packages
17. Data Warehouse and Data mining 34. Database triggers
18. No-SQL Database 35. Advanced Topics
Other Books from Parteek Bhatia

For more information visit: [Link]


Online Courses on Udemy by Parteek Bhatia
Parteek Bhatia’s YouTube Channel

Premier resource for over 450 in-depth videos covering Machine


Learning, Data Mining, DBMS, SQL, PL/SQL, Python, and Big Data. Our
channel simplifies complex topics without compromising on depth,
catering to both beginners and seasoned professionals. Explore tutorials,
step-by-step guides, and practical tips to enhance your tech skills.
Whether you're preparing for exams, advancing your career, or exploring
new technologies, our content is designed to empower your learning
journey. Subscribe and join our community of tech enthusiasts to stay
updated with educational and engaging content!
Thanks
Happy Learning.

You might also like