[Link] Asst.
prof KITS
DATA WARE HOUSING AND DATA MINING
UNIT-1:
Introduction: Fundamentals of Data Mining? Data Mining Functionalities ? Classification of
Data Mining system? Data mining task primitives ? Integration of data mining system with a
database or a data warehouse system ? Major issues in data mining ? CRISP model ?
Data preprocessing: Need for preprocessing the data? Data cleaning , Data Integration, Data
Transformation , Data Reduction.
INTRODUCTION: Data mining is nothing but discovery of knowledge data from large
database. Generally the term mining refers to mining of gold from rocks or sand is called
gold mining.
Why Data Mining?
• The major reason that data mining has attracted a great deal of attention in the
information industry in recent years is due to the wide availability of huge amounts of
data and need for turning such data into useful information and knowledge.
• The information and knowledge gained can be used for applications ranging from
business management, production control, and market analysis, to engineering design
and science exploration.
• Data mining can be viewed as a result of the natural evolution of information
technology. It means, providing a path to extract the required data of an industry from
warehousing machine. This is the witness of developing knowledge of an industry.
• It includes data collection, database creation, data management (i.e data storage and
retrieval, and database transaction processing) and data analysis and
understanding(involving data warehousing and data mining).
Evolution of data mining and data warehousing:
In the development of data mining, we should know the evolution of database. This
includes,
Data collection and Database creation: In the 1960’s, database and information technology
began with file processing system. It is powerful database system. But it is providing
inconsistency of data. It means, a user needs to maintain duplicate data of an industry.
Database Management System: In b/w 1970 – 1980, the progress of database is
Hierarchical and network database systems were developed.
Relational database systems were developed
Data modeling tools were developed in early 1980s (such as E-R model etc.
Indexing and data organization techniques were developed. ( such as B+ tree, hashing etc).
Query languages were developed. (such as SQL, PL/SQL)
User interfaces, forms and reports, query processing.
On-line transaction processing (OLTP)
Advanced Database Systems: In mid 1980s to till date,
Advanced data models were developed. (such as extended relational, object-
Department of CAI 1
[Link] [Link] KITS
oriented,object-relational, spatial, temporal, multimedia .
Data Warehousing and Data mining: In late 1980 to till date
Developed Data warehouse and OLAP technology
Data mining and knowledge discovery were introduced.
Web-based Databases Systems: In 1990 – ll date
XML based database systems and web mining were developed.
New Genera on of Integrated Informa on Systems: From 2000 onwards developed an
integratedinforma on system.
What is Data Mining: The term Data Mining refers to extracting or “mining” knowledge
from large amounts of data. The term mining is actually a misnomer (i.e. unstructured data). For
example, mining of gold from rocks or sand is referred to as gold mining.
Data mining steps in the knowledge discovery process:
Data mining is the process of discovering meaningful new trends bystoring the large amount of
data in repository of database. It also uses pattern recognition techniques as well as statistical
techniques.
(KDD): The Data mining is a step in the Knowledge Discovery in Databases (KDD). It has
different stages, such as
Data Cleaning: It is the process of removing noise and inconsistent data.
Data Integrating: It is the process of combining data from multiple sources.
Data Selection: It is the process of retrieving relevant data from database.
Data Transformation: In this process, data are transformed or consolidated into forms or
reportsby performing summary or aggregation operation.
Data Mining: It is an essential process to extracting data from raw data by using
intelligentmethods.
Pattern Evaluation: to identify the discovered data is in the knowledge based on some
interestingness measures.(i.e identify the mined data is in the required format or not.).
Knowledge presentation: Visualization and knowledge representation techniques are
used to present the mined data to the user.
Department of CAI 2
ti
ti
ti
ti
[Link] [Link] KITS
Architecture of Data Mining System:
The architecture of data mining is the process of discovering the interesting knowledge from large
amounts of data stored either in databases or in the data warehouse or information
repositories. It has various stages to extract the data into user view from unstructured sources.
Database, data warehouse, or information repositories:
This is single or set of databases, data warehouses, spreadsheets, or other kinds of
information repositories. In this step, the Data Cleaning and Data Integration techniques may be
performed
Database or data warehouse server: The database or data warehouse server warehouse
server is responsible for fetching the relevant data based on the user’s data mining request.
Knowledge base: This is the domain knowledge that is used to guide the search or evaluate the
interestingness of resulting Database patterns.
Data mining engine: This is essential to the data mining system and ideally consists of a
set of functional modules for tasks such as characterization, association, classification,
cluster analysis, and evolution and deviation analysis.
pattern evaluation module: This step providing measures, constraints(rules), methods etc to
filterout the discovered patterns or data. This is most useful for efficient data mining.
Graphical user interface: This step provides the communication b/w user and data mining
system. It allows the user to interact with the system by specifying a data mining query or task.
Department of CAI 3
[Link] [Link] KITS
What Kind of Data Can Be Mined?
Data mining can be applied to any kind of information repositories such as Databases data,
data warehouse, transactional data bases, advanced systems, flat files and the World Wide
Web. Advanced databases systems include object-oriented, object-relation databases, time
series databases, text databases and multimedia databases.
Databases Data: A database system is also called a database management system
(DBMS). It consists of a collection of interrelated data, known as a database, and set of
software programs to manage and access the data. The software programs provide
mechanisms for defining database structures and data storage. These also provide data
consistency and security, concurrency,shared or distributed data access etc.
relational database: is a collection of tables, each of which is assigned a unique name. Each
table consists of a set of attributes (columns or fields) and a set of tuples (records or rows).
Each tuple is identified by a unique key and is described by a set of attribute values. For this, ER
models are constructed for relational databases.
Data Warehouses: A data warehouse is a repository of information collected from
multiple sources, stored under a schema and resides at a single site. The data warehouses are
constructed by a process of data cleaning, data transformation, data integration, data loading
and periodic data refreshing.
Transactional Databases: A transactional database consists of a file where each record
represents a transaction. A transaction includes a unique transaction such as data of the
transaction, the customer id number, the ID number of the sales person and so on.
AllElectronics transactions can be stored in a table with one record per transaction. This is
shown in fig.
Transaction_id List of items Transaction dates
T100 I1, I3, I8, I16 18-12-2018
T200 I2, I8 18-12-2018
1.2 WHAT KINDS OF PATTERNS CAN BE MINED? (OR) DATA
MINING FUNCTIONALITIES:
Data mining functionalities are used to specify the kind of patterns to be found in data mining
[Link] mining tasks are classified into two categories descriptive and predictive.
Descriptive mining tasks characterize the general properties of the data in the database.
Predictive mining tasks perform inference on the current data in order to make predictions.
There are five data mining functionalities
1. Concept/Class Description: Data can be associated with classes or concepts.
Descriptions of a individual classes or a concepts in summarized, concise and precise
terms called class or concept descriptions. These descriptions can be divided into
Department of CAI 4
[Link] [Link] KITS
1. Data Characterization 2. Data Discrimination.
Data Characterization:
• It is summarization of the general characteristics of class or cocept.
• The data corresponding to the user specified class are collected by a database query.
The output of data characterization can be presented in various forms like pie charts, bar charts
curves, multidimensional cubes, multidimensional tables etc. The resulting descriptions can be
presented as generalized relations are called characteristic rules.
Data Discriminations: Comparison of two target class data objects from one or set of
contrasting (distinct) classes. The target and contrasting classes can be specified by the user,
and the corresponding data objects are retrieved through database queries.
For example, comparison of products whose sales increased by 10% in the last year
with those whose sales decreased by 30% during the same period. This is called data
discrimination.
2. Mining Frequent Patterns, Associations and Correlations:
Frequent Patterns: Frequent patterns are those patterns that occur frequently in transactional
data. Here isthe list of kind of frequent patterns −
Frequent Item Set − It refers to a set of items that frequently appear together, for example, milk
and bread.
Frequent Subsequence − A sequence of patterns that occur frequently such aspurchasing a
camera is followed by memory card.
Frequent Sub Structure − Substructure refers to different structural forms, such asgraphs, trees,
or lattices, which may be combined with item-sets or subsequences.
Association Analysis: “What is association analysis ?”
Association analysis is the discovery of association rules. it is a way of identifying the relation
between various items. It is used for transaction data analysis.
For example, In AllElectronics relational database, data mining system may find association
rules like
buys(X, “computer”) ==> buys(X, “so ware”)
Here, who buys “computer”, they buys “software”.
Correlation analysis:
Correlation analysis is a mathematical technique . it shows how strongly pair of attributes are
related together.
Example: Tall people tend to have more weight.
3. Classi ication and Regression for predictive analysis:
Classification: It is the process of finding a set of models that describes and distinguishes
data classes or concepts.
• The derived model may be represented in various forms such as classification (IF-
THEN) rules, decision trees, mathematical formulae or neural networks.
Department of CAI 5
f
ft
[Link] [Link] KITS
• A decision tree is a flow-chart like tree structure. The decision trees can easily
• converted to classification rule. The neural networks are used for classification to
provide connection b/w computers.
Regression for Prediction: It is a statistical method. It is used to predict numerical missing or
unavailable data values rather than class labels. Prediction refers to both data value prediction and
class label prediction. The predicted values are numerical data and are often referred to as
prediction.
4. Cluster Analysis: (“What is cluster analysis?”)
Clustering is a method of grouping data into different groups, it grouping the data based on
the one principle that is maximizing the intra class similarity and minimizing the inter class
similarity. The objectives of clustering are
• To uncover natural groupings
• To initiate hypothesis about the data
• To find consistent and valid organization
of data.
For example, Cluster analysis can be performed
on AllElectronics customers. It means, to
identify homogeneous (same group) customers.
By this cluster may represent target groups for
marketing to increase the sales.
5. Outlier Analysis − Outliers may be defined as the data objects that do not comply with
the general behavior or model of the data available.
Most data mining methods discard outliers as noise or exceptions. Finding such type of
applications are fraud detection is referred as outlier mining.
For example, Outlier analysis may uncover usage of credit cards by detecting
purchases of large amount of products when comparing with regular purchase of large
product customers.
1.3 WHICH TECHNOLOGIES ARE USED? (OR)
CLASSIFICATION OF DATA MINING SYSTEMS:
Data mining is classified with many techniques. Such as statistics, machine learning, pattern
recognition, database and data warehouse systems, information retrieval, visualization,
algorithms, high performance computing, and many application domains (Shown in Figure).Data
mining system can be categorized according to various criteria.
Because of diversity of disciplines contributing to data mining ,data mining research is
Department of CAI 6
[Link] [Link] KITS
expected to generate large variety of data mining systems. data mining systems can be categorized
according to various criteria ,as fallows.
Classification according to the kinds of databases mined: a data mining system can
be classified according to the kinds of database mined. database systems can be classified
according to different criteria such as models or types of data.
For instance, if classifying according to the data models, we may have relational, transactional
Object relational or data warehouse mining system.
Classification according to the kinds of knowledge mined: a data mining system can
be classified according to the kinds of knowledge they mined .that is, based on the mining
functionalities such as characterization , discrimination association and correlation analysis,
classification ,prediction, clustering outlier analysis and evolution analysis .
Classification according to the kinds of techniques utilized:
1. Statistics:
• It uses the mathematical analysis to express representations, model and summarize
empirical data or real world observations.
• Statistical analysis involves the collection of methods, applicable to large amount ofdata to
conclude and report the trend.
2. Machine learning
• Arthur Samuel defined machine learning as a field of study that gives computers the
ability to learn without being programmed.
• When the new data is entered in the computer, algorithms help the data to grow orchange
due to machine learning.
• In machine learning, an algorithm is constructed to predict the data from the available
database (Predictive analysis).
• It is related to computational statistics.
The four types of machine learning are:
Supervised learning
Department of CAI 7
[Link] [Link] KITS
• It is based on the classification.
• It is also called as inductive learning. In this method, the desired outputs are included in
the training dataset.
Unsupervised learning:
Unsupervised learning is based on clustering. Clusters are formed on the basis of similarity
measures and desired outputs are not included in the training dataset.
Semi-supervised learning:
Semi-supervised learning includes some desired outputs to the training dataset to generate the
appropriate functions. This method generally avoids the large number of labeled examples (i.e.
desired outputs).
Active learning:
Active learning is a powerful approach in analyzing the data efficiently.
The algorithm is designed in such a way that, the desired output should be decided by the
algorithm itself (the user plays important role in this type)
3. Visualization:
It is the process of extracting and visualizing the data in a very clear and
understandable way without any form of reading or writing by displaying the results
in the form of pie charts, bar graphs, statistical representation and through graphical
forms as well.
4. Database systems and data warehouse:
• Databases are used for the purpose of recording the data as well as data warehousing.
• Online Transactional Processing (OLTP) uses databases for day to day transaction
purpose.
• Data warehouses are used to store historical data which helps to take strategically decision
for business.
• It is used for online analytical processing (OALP), which helps to analyze the data.
5. Information retrieval
Information deals with uncertain representations of the semantics of objects (text, images).
• For example: Finding relevant information from a large document.
•
Classification according to the application adopted: a data mining system can be
classified according to the applications they adapt .
Financial Data Analysis:
• The financial data in banking and financial industry is generally reliable and of high
quality which facilitates systematic data analysis and data mining. Like,
• Loan payment prediction and customer credit policy analysis.
• Detection of money laundering and other financial crimes.
Retail Industry:
• Data Mining has its great application in Retail Industry because it collects large amount of
data from on sales, customer purchasing history, goods transportation,consumption and
services. It is natural that the quantity of data collected will continue to expand rapidly
because of the increasing ease, availability and popularity of the web.
Telecommunication Industry:
• Today the telecommunication industry is one of the most emerging industries providing
Department of CAI 8
[Link] [Link] KITS
various services such as fax, pager, cellular phone, internet messenger, images, e- mail,
web data transmission, etc. Due to the development of new computer andcommunication
technologies, the telecommunication industry is rapidly expanding. This is the reason
why data mining is become very important to help and understand the business.
Biological Data Analysis:
• In recent times, we have seen a tremendous growth in the field of biology such as
genomics, proteomics, functional Genomics and biomedical research. Biological data
mining is a very important part of Bioinformatics.
• Other Scientific Applications
• The applications discussed above tend to handle relatively small and homogeneous data
sets for which the statistical techniques are appropriate. Huge amount of data have been
collected from scientific domains such as geosciences, astronomy, etc.
• A large amount of data sets is being generated because of the fast numerical simulations in
various fields such as climate and ecosystem modeling, chemical engineering, fluid
dynamics, etc.
1.4 DATA MINING TASK PRIMITIVES:
A data mining task can be specified in the form of a data mining query, which is input to the data
mining system. A data mining query is defined in terms of data mining task primitives. These
primitives allow the user to interactively communicate with the data mining system during discovery
in order to direct the mining process, or examine the findings from different angles or depths. The
data mining primitives specify the following, as illustrated in Figure
Department of CAI 9
[Link] [Link] KITS
➢ The set of task-relevant data to be mined: This specifies the portions of the database or
the set of data in which the user is interested. This includes the database attributes or data
warehouse dimensions of interest (referred to as the relevant attributes or dimensions).
➢ The kind of knowledge to be mined: This specifies the data mining functions to be
performed, such as characterization, discrimination, association or correlation analysis,
classification, prediction, clustering, outlier analysis, or evolution analysis.
➢ The background knowledge: to be used in the discovery process: This knowledge about
the domain to be mined is useful for guiding the knowledge discovery process and for
evaluating the patterns found. Concept hierarchies are a popular form of background
knowledge, which allow data to be mined at multiple levels of abstraction. An example of
a concept hierarchyfor the attribute (or dimension) age is shown in Figure 1.14. User
beliefs regarding relationships in the data are another form of background knowledge.
➢ The interestingness measures and thresholds for pattern evaluation: They may be
used to guide the mining process or, after discovery, to evaluate the discovered patterns.
Different kinds of knowledge may have different interestingness measures. For example,
interestingness measures for association rules include support and confidence. Rules
whose support and confidence values are below user-specified thresholds are considered
uninteresting.
➢ The expected representation for visualizing the discovered patterns: This refers to the
form in which discovered patterns are to be displayed, which may include rules, tables,
Department of CAI 10
[Link] [Link] KITS
charts, graphs, decision tree, and cubes.
A data mining query language can be designed to incorporate these primitives, allowing
users to flexibly interact with data mining systems. Having a data mining query language
provides a foundation on which user-friendly graphical interfaces can be built.
1.5 INTEGRATION OG DATA MINING SYSTEM WITH A
DATABASE OR DATA WAREHOUSE SYSTEM
If a data mining system works as a stand- alone system or is embedded in an application program,
there are no DB or DW system with which it has no communicate. This simple scheme is called no
coupling where the main focus of the data mining design rest on developing effective and efficient
algorithms for mining.
When data mining system works in n environment that requires it to communicate with other
information system components, such as DB or DW system, possible integration schemes include no
coupling, loose coupling, semi tight coupling , tight coupling.
No Coupling: No coupling means that a DM system will not utilize any function of a DB or DW
system. It may fetch data from a particular source (such as a file system), process data using some
data mining algorithms, and then store the mining results in another file. First, a DB system
provides a great deal of flexibility and efficiency at storing, organizing accessing, and processing
data.
Without using a DB/DW system, a DM system may spend a substantial amount of time
finding, collecting, cleaning, and transforming data. In DB and/or DW systems, data tend to be
well organized, indexed, cleaned, integrated, or consolidated, so that finding the task-relevant,
high quality data becomes an easy task.
Moreover, most data have been or will be stored in DB/DW systems. Without any coupling
of such systems, a DM system will need to use other tools to extract data, making it difficult to
integrate such a system into an information processing environment. Thus, no coupling
represents a poor design.
Loose Coupling: Loose coupling means that a DM system will use some facilities of a DB or
DW system, fetching data from a data repository managed by these systems, performing data
mining, and then storing the mining results either in a file or in a designated place in a database or
data warehouse.
Loose coupling is better than no coupling because it can fetch any portion of data stored in
databases or data warehouses by using query processing, indexing, and other system facilities. It
incurs some advantages of the flexibility, efficiency, and other features provided by such systems.
However, many loosely coupled mining systems are main memory-based. Because mining
does not explore data structures and query optimization methods provided by DB or DW systems,
it is difficult for loose coupling to achieve high scalability and good performance with large data
sets.
Department of CAI 11
[Link] [Link] KITS
Semi tight Coupling: Semi tight coupling means that besides linking a DM system to a DB/DW
system, efficient implementations of a few essential data mining primitives (identified by the
analysis of frequently encountered data mining functions) can be provided in the DB/DW system.
These primitives can include sorting, indexing, aggregation, histogram analysis, and pre
computation of some essential statistical measures, such as sum, count, max, min, standard
deviation, and so on. Moreover, some frequently used intermediate mining results can be pre
computed and stored in the DB/DW system. this design will enhance the performance of a DM
system.
Tight Coupling: Tight coupling means that a DM system is smoothly integrated into the DB/DW
system. The data mining subsystem is treated as one functional component of an information
system.
Data mining queries and functions are optimized based on query analysis ,data structures .
With further technology advances ,DM,DB, and DW systems will evolve and integrate
together as one information system with multiple functionalities. This will provide a uniform
information processing environment.
This approach is highly desirable because it facilitates efficient implementation of data mining
functions high system performance and an integrated information processing environment.
With this analysis , it is easy to see that data mining system should be coupled with a DB/DW
system. Loose coupling though not efficient ,is better than no coupling because it uses both data
and system facilities of DB system tight coupling is highly desirable but its implementation needs
more research .Semi tight coupling is a compromise between loose and tight coupling.
1.6 MAJOR ISSUES IN DATA MINING
Data mining is a dynamic and fast-expanding field with great strengths. The major issuescan
divided into three groups:
1. Mining methodologies and user interaction issues
2. Performance issues
3. Issues related to the diversity of database types
1. Mining methodologies and user interaction issues: these reflects the kind of knowledge
mined .
It refers to the following kinds of issues −
➢ Mining different kinds of knowledge in databases − Different users may be
interested in different kinds of knowledge. Therefore it is necessary for data
mining to cover a broad range of knowledge discovery task. Including data
characterization ,data discrimination association and correlation analysis and
classification prediction ,clustering ,outlier analysis.
➢ Interactive mining of knowledge at multiple levels of abstraction − The data
Department of CAI 12
[Link] [Link] KITS
mining process needs to be interactive because it allows users to focus the search
for patterns, providing and refining data mining requests based on the returned
results.
➢ Incorporation of background knowledge − To guide discovery process and to express the
discovered patterns, the background knowledge can be used. Backgroundknowledge may be
used to express the discovered patterns not only in concise terms but at multiple levels of
abstraction.
➢ Data mining query languages and ad hoc data mining − Data Mining Query language that
allows the user to describe ad hoc mining tasks, should be integrated with a data warehouse
query language and optimized for efficient and flexible data mining.
➢ Presentation and visualization of data mining results − Once the patterns are
discovered it needs to be expressed in high level languages, and visual representations.
These representations should be easily understandable. This requires the system to adopt
expressive knowledge representation techniques such as trees ,tables, rules , graphs, charts
etc.
➢ Handling noisy or incomplete data − the data cleaning methods are required to handle
the noise and incomplete objects while mining the data regularities. If the datacleaning
methods are not there then the accuracy of the discovered patterns will be poor.
➢ Pattern evaluation − the interestingness measure problem - the patterns discovered
should be interesting because either they represent common knowledge or lack novelty.
the interestingness measure to guide the discovery process and reduce the search space.
2. Performance issues: this include efficiency ,scalability, and parallelization of data mining
algorithms.
➢ Efficiency and scalability of data mining algorithms − In order to effectively extract the
information from huge amount of data in databases, data mining algorithm must be
efficient and scalable.
➢ Parallel, distributed, and incremental mining algorithms − The factors such as huge size of
databases, wide distribution of data, and complexity of data mining methods motivate the
development of parallel and distributed data mining algorithms. These algorithms divide
the data into partitions which is further processed in a parallel fashion. Then the results
from the partitions is merged. The incremental algorithms, update databases without
mining the data again from scratch.
3. Issues related to the diversity of database types:
➢ Handling of relational and complex types of data − The database may contain complex
data objects, multimedia data objects, spatial data, temporal data etc. It is not possible for
one system to mine all these kind of data.
➢ Mining information from heterogeneous databases and global information systems −
The data is available at different data sources on LAN or WAN. Thesedata source may
Department of CAI 13
[Link] [Link] KITS
be structured, semi structured or unstructured. Therefore mining the knowledge from them
adds challenges to data mining.
The above issues are considered as major requirements and challenges for the further evolution of
data mining research and development
DATA PREPROCESSING
Preprocessing:
Real-world databases are highly susceptible to noisy, missing, and inconsistent data due to
their typically huge size (often several gigabytes or more) and their likely origin from multiple,
heterogeneous sources. Low-quality data will lead to low-quality mining results, so we prefer a
preprocessing concepts.
Data Preprocessing Techniques
1. Data cleaning: can be applied to remove noise and correct inconsistencies in the data.
2. Data integration: merges data from multiple sources into coherent data store, such
asa data warehouse.
3. Data reduction can reduce the data size by aggregating, eliminating redundant features,
or clustering, for instance. These techniques are not mutually exclusive; they may work
together.
4. Data transformations, such as normalization, may be applied.
Need for preprocessing
➢ Incomplete, noisy and inconsistent data are common place properties of large real world
databases and data warehouses.
➢ Incomplete data can occur for a number of reasons:
• Attributes of interest may not always be available
• Relevant data may not be recorded due to misunderstanding, or because of equipment
malfunctions.
• Data that were inconsistent with other recorded data may have been deleted.
• Missing data, particularly for tuples with missing values for some attributes, mayneed to
be inferred.
• The data collection instruments used may be faulty.
• There may have been human or computer errors occurring at data entry.
• Errors in data transmission can also occur.
• There may be technology limitations, such as limited buffer size for coordinating
synchronized data transfer and consumption.
• Data cleaning routines work to ―cleanǁ the data by filling in missing values,
smoothing noisy data, identifying or removing outliers, and resolving inconsistencies.
• Data integration is the process of integrating multiple databases cubes or files. Yet some
Department of CAI 14
[Link] [Link] KITS
attributes representing a given may have different names in different databases, causing
inconsistencies and redundancies.
• Data transformation is a kind of operations, such as normalization and aggregation, are
additional data preprocessing procedures that would contribute toward the successof the
mining process.
• Data reduction obtains a reduced representation of data set that is much smaller in volume,
yet produces the same(or almost the same) analytical results.
Fig: forms of data preprocessing.
1. DATA CLEANING:
Real-world data tend to be incomplete, noisy, and inconsistent. Data cleaning (or data
cleansing) routines attempt to fill in missing values, smooth out noise while identifying outliers
and correct inconsistencies in the data.
Missing Values
Many tuples have no recorded value for several attributes, such as customer [Link] we
can fill the missing values for this attributes.
Department of CAI 15
[Link] [Link] KITS
The following methods are useful for performing missing values over several attributes:
a. Ignore the tuple: This is usually done when the class label missing (assuming the mining task
involves classification). This method is not very effective, unless thetuple contains several
attributes with missing values. It is especially poor when the percentage of the missing values per
attribute varies considerably.
b. Fill in the missing values manually: This approach is time –consuming and may not be
feasible given a large data set with many missing values.
c. Use a global constant to fill in the missing value: Replace all missing attribute valueby the
same constant, such as a label like ―unknownǁ or -∞.
d. Use the attribute mean to fill in the missing value: For example, suppose that the average
income of customers is $56,000. Use this value to replace the missing value for income.
e. Use the most probable value to fill in the missing value: This may be determined with
regression, inference-based tools using a Bayesian formalism or decision tree induction. For
example, using the other customer attributes in the sets decision tree is constructed to predict the
missing value for income.
Noisy Data:
Noise is a random error or variance in a measured variable. Noise is removed using data
smoothing techniques.
Binning: Binning methods smooth a sorted data value by consulting its ―neighborhood,ǁ that is
the value around it. The sorted values are distributed into a number of ―bucketsǁ or ―bins―.
Because binning methods consult the neighborhood of values, they perform local smoothing.
Sorted data for price (in dollars): 3,7,14,19,23,24,31,33,38.
Example 1: Partition into (equal-frequency) bins:
Bin 1: 3,7,14
Bin 2: 19,23,24
Bin 3: 31,33,38
In the above method the data for price are first sorted and then
partitioned into equal-frequency bins of size 3.
Smoothing by means:
Bin 1: 8,8,8
Bin 2: 22,22,22
Bin 3: 34,34,34
In smoothing by bin means method, each value in a bin is replaced by the
mean value ofthebin. For example, the mean of the values 3,7&14 in bin 1
is 8[(3+7+14)/3].
Smoothing by bin boundaries:
Bin 1: 3,3,14
Department of CAI 16
[Link] [Link] KITS
Bin 2: 19,24,24
Bin 3: 31,31,38
In smoothing by bin boundaries, the maximum & minimum values in give
bin or identify asthe bin boundaries. Each bin value is then replaced by the
closest boundary value.
In general, the large the width, the greater the effect of the smoothing.
Alternatively, bins may be equal-width, where the interval range of values in
each bin is constant .
Department of CAI 17
[Link] [Link] KITS
Regression: Data can be smoothed by fitting the data to function, such as with regression. Linear
regression involves finding the ―bestǁ line to fit two attributes (or variables), so that one attribute
can be used to predict the other. Multiple linear regressions is an extension of linear regression,
where more than two attributes are involved and the data are fit to a multidimensional surface.
Clustering: Outliers may be detected by clustering, where similar values are organized into
groups, or ―clusters.ǁ Intuitively, values that fall outside of the set of clusters may be
considered outliers.
[Link] Integration
Data mining often requires data integration - the merging of data from stores into a
coherent data store, as in data warehousing. These sources may include multiple data bases,
datacubes, or flat files.
Issues in Data Integration:
1. Entity identification problem
2. Redundancy and correlation analysis
3. Tuple duplication
4. Data value conflict detection and resolution
1. Entity identification problem : schema integration and object matching are very important
issues in data integration
Schema integration : mismatch in attribute names.
Example: cust_id, customer_id, cust_no, etc
Handling blank, zero, null values.
Object matching: mismatch in structure of the data
Example : discount issues.
Currency type.
2. Redundancy and correlation analysis: Redundancy is another important issue an attribute
(such as annual revenue, for instance) may be redundant if it can be ―derived from another
attribute are set of attributes. Inconsistencies in attribute of dimension naming can also cause
redundancies in the resulting data set. Some redundancies can be detected by correlation analysis
and covariance analysis.
Example : DOB ,AGE.
quarter sales, year sales
NAME DOB AGE
Department of CAI 18
[Link] [Link] KITS
correlation analysis: given two attributes ,such analysis can measure how strongly one attribute
implies the other, based on the available data.
3. Tuple duplication : the use of demoralized tables (often done to improve performance by
avoiding joins) is another source of data redundancy . Inconsistencies often arise between various
duplicates, due to in accurate data entry or updating some but not all data occurrences.
Example :
NAME DOB BRANCH OOCUPATION ADDRESS
A 25 Tpg govt tpg
B 30 Tnk govt rjy
A 25 Tpg private tpg
C 30 Tnk private rjy
4. Data value conflict detection and resolution: A fourth important issue in data
integration is the data value conflict detection and resolution . For example, for the
same real–world entity, attribute value from different sources may differ. This may be
due to difference in representation, scaling, or encoding.
For example, a weight attribute may be stored in metric units in one system
andBritish imperial units in another.
School curriculum (grading system).
Attribute may also different on the abstraction level, where an attribute in one system is recorded
at ,a lower level abstraction than the “same” attribute in another.
Example: Monthly total sales in store.
Monthly total sales from all stores in that region.
Department of CAI 19
[Link] [Link] KITS
3. Data Transformation
The measurement unit used can affect the data analysis. For example, changing measurement
units from meters to inches for height, or from kilograms to pounds for weight, may lead to very
different results.
For distance-based methods, normalization helps prevent attributes with initially large ranges
(e.g., income) from outweighing attributes with initially smaller ranges (e.g., binary attributes). It
is also useful when given no prior knowledge of the data.
There are many methods for data normalization. We study min-max normalization, z- score
normalization, and normalization by decimal scaling. For our discussion, let A be a numeric
attribute with n observed values, v1, v2, …., vn.
a. Min-max normalization performs a linear transformation on the original data. Suppose that
minAand maxAare the minimum and maximum values of an attribute, [Link]- maxnormalization
maps a value, vi, of A to vi’in the range [new_minA,new_maxA]by computing
Min-max normalization preserves the relationships among the original data values. Itwill
encounter an ―out-of-boundsǁ error if a future input case for normalization fallsoutside of the
original data range for A.
Example:-Min-max normalization. Suppose that the minimum and maximum values
fortheattribute income are $12,000 and $98,000, respectively. We would like to map income to
the range [0.0, 1.0]. By min-max normalization, a value of $73,600 for income istransformed to
b. Z-Score Normalization
The values for an attribute, A, are normalized based on the mean (i.e., average) and standard
deviation of A. A value, vi, of A is normalized to vi’ by computing
Where and σA are the mean and standard deviation, respectively, of attribute A. Example z-
score normalization. Suppose that the mean and standard deviation of the values for the attribute
income are $54,000 and $16,000, respectively. With z-score normalization, avalue of $73,600 for
income is transformed to
c. Normalization by Decimal Scaling:
Normalization by decimal scaling normalizes by moving the decimal point of values of attribute
A. The number of decimal points moved depends on the maximum absolute value of
A. A value, vi, of A is normalized to vi’ by computing
Department of CAI 20
𝐴
[Link] [Link] KITS
where j is the smallest integer such that max(|vi’)< 1.
Example :Decimal scaling. Suppose that the recorded values of A range from -986 to
917. The maximum absolute value of A is 986. To normalize by decimal scaling, we
Therefore divide each value by 1000 (i.e., j = 3) so that -986 normalizes to -0.986
and917normalizes to 0.917.
4. Data Reduction:
Obtain a reduced representation of the data set that is much smaller in volume but yet
produces the same (or almost the same) analytical results.
Why data reduction? — A database/data warehouse may store terabytes of data.
Complex data analysis may take a very long time to run on the complete data set.
Data reduction strategies
[Link] cube aggregation
[Link] Subset Selection
[Link] reduction — e.g., fit data into models
[Link] reduction - Data Compression
Data cube aggregation:
For example, the data consists of All Electronics sales per quarter for the years 2014
to [Link] are, however, interested in the annual sales, rather than the total per
quarter. Thus, the data can be aggregated so that the resulting data summarize the
total sales per year insteadof per quarter.
Department of CAI 21
[Link] [Link] KITS
Year/ 201 201 201 201 Yea Sal
Quarter 4 5 6 7 r es
Quarter 1 200 210 320 230 201 164
4 0
Quarter 2 400 440 480 420 201 171
5 0
Quarter 3 480 480 540 460 201 202
6 0
Quarter 4 560 580 680 640 201 175
7 0
Attribute Subset Selection
Attribute subset selection reduces the data set size by removing irrelevant or
redundant attributes (or dimensions). The goal of attribute subset selection is to find
a minimum set of attributes. It reduces the number of attributes appearing in the
discovered patterns, helping to make the patterns easier to understand.
For n attributes, there are 2n possible subsets. An exhaustive search for the optimal
subset of attributes can be prohibitively expensive, especially as n and the number
of data classes increase. Therefore, heuristic methods that explore a reduced search
space are commonly used for attribute subset selection. These methods are typically
greedy in that, while searching to attribute space, they always make what looks to
be the best choice at that time. Their strategy to make a locally optimal choice in the
hope that this will lead to a globally optimal solution. Many other attributes
evaluation measure can be used, such as the information gain measure used in
building decision trees for classification.
Department of CAI 22
[Link] [Link] KITS
Techniques for heuristic methods of attribute sub set selection
➢ Stepwise forward selection
➢ Stepwise backward elimination
➢ Combination of forward selection and backward elimination
➢ Decision tree induction
1. Stepwise forward selection: The procedure starts with an empty set of attributes as the
reduced set. The best of original attributes is determined and added to the reduced set. At each
subsequent iteration or step, the best of the remaining original attributes is added to the set.
2. Stepwise backward elimination: The procedure starts with full set of attributes. At each
step, it removes the worst attribute remaining in the set.
3. Combination of forward selection and backward elimination: The stepwise forward
selection and backward elimination methods can be combined so that, at each step, the
procedure selects the best attribute and removes the worst from among the remaining
attributes
4. Decision tree induction: Decision tree induction constructs a flowchart like structure where
each internal node denotes a test on an attribute, each branch corresponds to an outcome of
the test, and each leaf node denotes a class prediction. At each node, the algorithm choices
the ―bestǁ attribute to partition the data into individual classes. A tree is constructed from
the given data. All attributes that do not appear in the tree are assumed to beirrelevant. The set
of attributes appearing in the tree from the reduced subset of attributes. Threshold measure is
used as stopping criteria.
Numerosity Reduction:
Numerosity reduction is used to reduce the data volume by choosing alternative,
smallerforms of the data representation
Techniques for Numerosity reduction:
• Parametric - In this model only the data parameters need to be stored, instead of theactual
data. (e.g.,) Log-linear models, Regression
• Nonparametric – This method stores reduced representations of data includehistograms,
clustering, and sampling.
Parametric model
1. Regression
Linear regression
• In linear regression, the data are model to fit a straight line. For example, a random
variable, Y called a response variable), can be modeled as a linear function of another
random variable, X called a predictor variable), with theequation Y=αX+β
• Where the variance of Y is assumed to be constant. The coefficients, α and β (called
regression coefficients), specify the slope of the line and the Y- intercept, respectively.
Multiple- linear regression
• Multiple linear regression is an extension of (simple) linear regression, allowing a
response variable Y, to be modeled as a linear function of two or more predictor variables.
2. Log-Linear Models
Department of CAI 23
[Link] [Link] KITS
Log-Linear Models can be used to estimate the probability of each point in a multidimensional
space for a set of discretized attributes, based on a smaller subset of dimensional combinations
Nonparametric Model
1. Histograms
A histogram for an attribute A partitions the data distribution of A into disjointsubsets, or
buckets. If each bucket represents only a single attribute-value/frequency pair, the buckets are
called singleton buckets.
Ex: The following data are bast of prices of commonly sold items at All Electronics. The numbers
have been sorted:
1,1,5,5,5,5,5,8,8,10,10,10,10,12,14,14,14,15,15,15,15,15,18,18,18,18,18,18,18,18,20,20,20,2
0,20,20,21,21,21,21,21,25,25,25,25,25,28,28,30,30,30
There are several partitioning rules including the following:
• Equal-width: The width of each bucket range is uniform.
• (Equal-frequency (or equi-depth): the frequency of each bucket is constant
2. Clustering
Clustering technique consider data tuples as objects. They partition the objects into groups or
Department of CAI 24
[Link] [Link] KITS
clusters, so that objects within a cluster are similar to one another and dissimilar to objects in
other clusters. Similarity is defined in terms of how close the objects are in space, based on a
distance function. The quality of a cluster may be represented by its diameter, the maximum
distance between any two objects in the cluster. Centroid distance is an alternative measure of
cluster quality and is defined as the average distance of each cluster object from the cluster
centroid.
3. Sampling:
Sampling can be used as a data reduction technique because it allows a large data setto be
represented by a much smaller random sample (or subset) of the data. Suppose that a large data
set D, contains N tuples, then the possible samples are Simple Random sample without
Replacement (SRS WOR) of size n: This is created by drawing „n‟ of the „N‟ tuples from D
(n<N), where the probability of drawing any tuple in D is 1/N, i.e., all tuples are equally likely to
be sampled.
Dimensionality Reduction:
In dimensionality reduction, data encoding or transformations are applied so as to obtained
reduced or ―compressedǁ representation of the oriental data.
Dimension Reduction Types
• Lossless - If the original data can be reconstructed from the compressed data without any
loss of information
• Lossy - If the original data can be reconstructed from the compressed data with loss of
information, then the data reduction is called lossy.
Effective methods in lossy dimensional reduction
a) Wavelet transforms
b) Principal components analysis.
1. Wavelet transforms:
The discrete wavelet transform (DWT) is a linear signal processing technique that, when
applied to a data vector, transforms it to a numerically different vector, of wavelet coefficients.
The two vectors are of the same length. When applying this technique to data reduction, we
consider each tuple as an n-dimensional data vector, that is, X=(x1,x2,…………,xn),
depicting n measurements made on the tuple from n database attributes.
2. Principal components analysis
Suppose that the data to be reduced, which Karhunen-Loeve, K-L, method consists of tuples
or data vectors describe by n attributes or dimensions. Principal components analysis, or PCA
(also called the Karhunen-Loeve, or K-L, method), searches for k n-dimensional orthogonal
vectors that can best be used to represent the data where k<=n. PCA combines the essence of
attributes by creating an alternative, smaller set of variables. The initial data can then be
projected onto this smaller set.
Department of CAI 25