Adama Science & Technology
University
CoEEC:CS
Data Mining(CoSc4152)
Introduction
1
Lecture Outline
Why Data Mining?
What Is Data Mining?
Data Can Be Mined
Patterns that Can Be Mined
Technologies Used
Applications Area of Data Mining
Major Issues in Data Mining
Summary
2
Why Data Mining?
The Explosive Growth of Data: from terabytes to petabytes
Data collection and data availability
Automated data collection tools, database systems, Web,
computerized society
Major sources of abundant data
Business: Web, e-commerce, transactions, stocks, …
Science: Remote sensing, bioinformatics, scientific
simulation, …
Society and everyone: news, digital cameras, YouTube
We are drowning in data, but starving for knowledge!
“Necessity is the mother of invention”—Data mining—
Automated analysis of massive data sets
3
What Is Data Mining?
Data mining (knowledge discovery from data)
Extraction of interesting (non-trivial, implicit, previously
unknown and potentially useful) patterns or knowledge from
huge amount of data
Data mining: a misnomer?
Alternative names
Knowledge discovery (mining) in databases (KDD), knowledge
extraction, data/pattern analysis, data archeology, data
dredging, information harvesting, business intelligence, etc.
Watch out: Is everything “data mining”?
Simple search and query processing
(Deductive) expert systems
4
What Is Data Mining?
Data mining is the science of discovering
structure and making predictions in large or
complex data sets
5
Knowledge Discovery (KDD) Process
This is a view from typical database
systems and data warehousing
Pattern Evaluation
communities
Data mining plays an essential role in
the knowledge discovery process
Data Mining
Task-relevant Data
Data Selection
Warehouse
Data Cleaning
Data Integration
Databases
6
Data Mining in Business Intelligence
Increasing potential
to support
business decisions End User
Decisio
n
Making
Data Presentation Business
Analyst
Visualization Techniques
Data Mining Data
Information Discovery Analyst
Data Exploration
Statistical Summary, Querying, and Reporting
Data Preprocessing/Integration, Data Warehouses
DBA
Data Sources
Paper, Files, Web documents, Scientific experiments, Database Systems
7
KDD Process: A Typical View from ML
and Statistics
Input Data Data Pre- Data Post-
Processing Mining Processin
g
Data integration Pattern discovery Pattern evaluation
Normalization Association & Pattern selection
correlation
Feature selection Classification Pattern
interpretation
Dimension reduction Clustering
Pattern visualization
Outlier analysis
…………
This is a view from typical machine learning and statistics communities
8
Why Data Mining?—Potential Applications
Data analysis and decision support
Market analysis and management
Target marketing, customer relationship management (CRM), market
basket analysis, cross selling, market segmentation
Risk analysis and management
Forecasting, customer retention, improved underwriting, quality
control, competitive analysis
Fraud detection and detection of unusual patterns (outliers)
Other Applications
Text mining (news group, email, documents) and Web mining
Intelligent query answering
DNA and bio-data analysis
9
Why Data Mining?—Potential Applications
Market Analysis and Management(1)
Where does the data come from?
Credit card transactions, loyalty cards, discount coupons, customer complaint calls, plus
(public) lifestyle studies
Target marketing
Find clusters of “model” customers who share the same characteristics: interest, income
level, spending habits, etc.
Determine customer purchasing patterns over time
Cross-market analysis
Associations/co-relations between product sales, & prediction based on such association
10
Why Data Mining?—Potential Applications
Market Analysis and Management(2)
Customer profiling
What types of customers buy what products (clustering or classification)
Customer requirement analysis
identifying the best products for different customers
predict what factors will attract new customers
Provision of summary information
multidimensional summary reports
statistical summary information (data central tendency and variation)
11
Why Data Mining?—Potential Applications
Corporate Analysis & Risk Management
Finance planning and asset evaluation
cash flow analysis and prediction
contingent claim analysis to evaluate assets
cross-sectional and time series analysis (financial-ratio, trend analysis,
etc.)
Resource planning
summarize and compare the resources and spending
Competition
monitor competitors and market directions
group customers into classes and a class-based pricing procedure
set pricing strategy in a highly competitive market
12
Why Data Mining?—Potential Applications
Approaches: Clustering & model construction for frauds, outlier analysis, based on historical data
Applications: Health care, retail, credit card service, telecomm.
Auto insurance: detect a group of people who stage accidents to collect insurance
Money laundering: suspicious monetary transactions
Medical insurance
Professional patients, ring of doctors, and ring of references
Unnecessary or correlated screening tests
Detecting inappropriate medical treatment
Telecommunications: phone-call fraud
Phone call model: destination of the call, duration, time of day
or week. Analyze patterns that deviate from an expected
13
Why Data Mining?—Potential Applications
Financial Data Analysis
Financial data
complete
reliable
high quality
Loan payment prediction and customer credit policy analysis
Factors influencing loan payment performance
loan-to-value ratio
term of the loan
debt ratio (total monthly debt/total monthly income)
payment-to-income ratio
income level
education level
residence region
credit history
Analysis may find that
payment-income ratio is a dominant factor while
education level and debt ratio are not
14
Why Data Mining?—Potential Applications
Data Mining for the Retail Industry
Multidimensional analysis of sales, customers, products, time and region
OLAP cubes
Effectiveness of sales campaigns
Advertisements, coupons, discounts, bonuses
promote products and attract customers
can help improve profits
Compare amount of sales and number of transactions
during the sales period versus before or after the sales campaign
Association analysis
which items are likely to be purchased together with the items on sale
Customer retention Analysis of Customer loyalty
sequences of purchases of particular customers
goods purchased at different periods by the same customers can be grouped into sequences
changes in customer consumption or loyalty
suggests adjustments on the pricing and variety of goods
to retain old customers and attract new customers
Purchase recommendation and cross-reference of items
associations from sales records
a customer who buy a PC is likely to buy a printer
purchase recommendations
15
Why Data Mining?—Potential Applications
Data Mining for the Telecommunication Industry
Telecommunication data are multidimensional
calling-time duration
location of caller location of called
type of call
used to identify and compare
data traffic system workload
resource usage user group behaviour
profit
fraudulent pattern analysis and identification of unusual patterns
to achieve customer loyalty
characteristics of customers affecting line usage
16
Multi-Dimensional View of Data
Mining
Data to be mined
Database data (extended-relational, object-oriented, heterogeneous, legacy),
data warehouse, transactional data, stream, spatiotemporal, time-series,
sequence, text and web, multi-media, graphs & social and information
networks
Knowledge to be mined (or: Data mining functions)
Characterization, discrimination, association, classification, clustering,
trend/deviation, outlier analysis, etc.
Descriptive vs. predictive data mining
Multiple/integrated functions and mining at multiple levels
Techniques utilized
Data-intensive, data warehouse (OLAP), machine learning, statistics, pattern
recognition, visualization, high-performance, etc.
Applications adapted
Retail, telecommunication, banking, fraud analysis, bio-data mining, stock
market analysis, text mining, Web mining, etc. 17
Data Mining: On What Kinds of
Data?
Database-oriented data sets and applications
Relational database, data warehouse, transactional database
Object-relational databases, Heterogeneous databases and legacy
databases
Advanced data sets and advanced applications
Data streams and sensor data
Time-series data, temporal data, sequence data (incl. bio-sequences)
Structure data, graphs, social networks and information networks
Multimedia database
Text databases
The World-Wide Web
18
Data Mining Function/Tasks
Characterization
Discrimination
Association and Correlation Analysis
Classification
Cluster Analysis
Outlier Analysis
19
Data Mining Functionalities / Tasks (1)
Characterization & Discrimination
Class/Concept Description
Data characterization is a summarization of general features of objects in
a target class
The data relevant to a target class are retrieved by a database query and run through a
summarization module to extract the essence of the data at different levels of
abstraction
example: characterize the video store’s customers who regularly rent more than 30
movies a year
Data discrimination is used to compare of the general features of objects
between two classes
comparison relevant features of objects between a target class and a contrasting class
example: compare the general characteristics of the customers who rented more than
30 movies in the last year with those who rented less than 5
The techniques used for data discrimination are very similar to the techniques used for
data characterization with the exception that data discrimination results include
comparative measures
20
Data Mining Function: (1)
Association and Correlation Analysis
Frequent patterns (or frequent itemsets)
What items are frequently purchased together in your
Shoa supermarket?
Association, correlation vs. causality
A typical association rule
Milk Diapper [0.5%, 75%] (support, confidence)
Are strongly associated items also strongly correlated?
How to mine such patterns and rules efficiently in large
datasets?
How to use such patterns for classification, clustering, and
other applications?
21
Data Mining Function: (3)
Classification
Classification and label prediction
Construct models (functions) based on some training examples
Describe and distinguish classes or concepts for future prediction
E.g., classify countries based on (climate), or classify cars
based on (gas mileage)
Predict some unknown class labels
Typical methods
Decision trees, naïve Bayesian classification, support vector
machines, neural networks, rule-based classification, pattern-
based classification, logistic regression, …
Typical applications:
Credit card fraud detection, direct marketing, classifying stars,
diseases, web-pages, …
22
Data Mining Functionalities / Tasks (4)
Prediction
Same as classification and estimation
except classify according to some predicted or estimated future
value
In prediction, historical data is used to build a (predictive) model that
explains the current observed behavior
Model can then be applied to new instances to predict future
behavior or forecast the future value of some missing attribute
Examples
predicting the size of balance transfer if the prospect accepts the
offer
predicting the load on a Web server in a particular time period
Suitable data mining tools
Association rule discovery
Decision Trees
23
Figure: (a) Affinity Grouping, (b) Decision Trees, (c)
Neural Networks
24
Data Mining Function: (4) Cluster
Analysis
Unsupervised learning (i.e., Class label is unknown)
Group data to form new categories (i.e., clusters), e.g., cluster
houses to find distribution patterns
Principle: Maximizing intra-class similarity & minimizing
interclass similarity
Many methods and applications
25
Data Mining Function: (5) Outlier
Analysis
Outlier analysis
Outlier: A data object that does not comply with the general
behavior of the data
Noise or exception? ― One person’s garbage could be
another person’s treasure
Methods: by product of clustering or regression analysis, …
Useful in fraud detection, rare events analysis
26
Can We Find All and Only Interesting Patterns?(1)
A data mining operation may generate thousands of
patterns
not all of them are interesting
e.g., the association rule {bread, milk} ==> {butter}
Possible Approaches
human-centered - use subjective measures of interestingness
and observation to find interesting patterns
query-based - use a “knowledge query” mechanism to query
the results after the mining operations
focused mining - establish constraints during mining to narrow
down result sets
But, how do we “measure” interestingness?
27
Can We Find All and Only Interesting Patterns?(2)
Objective vs. subjective interestingness measures
Objective: based on statistics and structures of patterns,
e.g., support, confidence, lift, etc.
Subjective: based on user’s beliefs (or domain-specific
beliefs) related to the data, e.g., “unexpectedness”,
“novelty”, etc.
Interestingness measures: A pattern is interesting if it is
easily understood by humans
valid on new or test data with some degree of certainty
potentially useful / actionable
novel, or validates some hypothesis that a user seeks to
confirm
28
Can We Find All and Only Interesting Patterns?(3)
Find all the interesting patterns: Completeness
Can a data mining system find all the interesting patterns?
Search for only interesting patterns: Optimization
Can a data mining system find only the interesting
patterns?
Approaches
First find all the patterns and then filter out the uninteresting ones
Generate only the interesting patterns -- mining query optimization
29
Structure and Network Analysis
Graph mining
Finding frequent subgraphs (e.g., chemical compounds), trees (XML),
substructures (web fragments)
Information network analysis
Social networks: actors (objects, nodes) and relationships (edges)
e.g., author networks in CS, terrorist networks
Multiple heterogeneous networks
A person could be multiple information networks: friends, family,
classmates, …
Links carry a lot of semantic information: Link mining
Web mining
Web is a big information network: from PageRank to Google
Analysis of Web information networks
Web community discovery, opinion mining, usage mining, …
30
Data Mining: Classification
Schemes
General functionality
Descriptive data mining
Predictive data mining
Different views, different classifications
Kinds of databases to be mined
Kinds of knowledge to be discovered
Kinds of techniques utilized
Kinds of applications adapted
31
Data Mining: Confluence of Multiple
Disciplines
Machine Pattern Statistics
Learning Recognition
Applications Data Mining Visualization
Algorithm Database High-Performance
Technology Computing
32
Why Confluence of Multiple
Disciplines?
Tremendous amount of data
Algorithms must be scalable to handle big data
High-dimensionality of data
Micro-array may have tens of thousands of dimensions
High complexity of data
Data streams and sensor data
Time-series data, temporal data, sequence data
Structure data, graphs, social and information networks
Spatial, spatiotemporal, multimedia, text and Web data
Software programs, scientific simulations
New and sophisticated applications
33
Different Issues in Data
Mining
Security and social issues:
Private and sensitive data is gathered and mined without
individual’s knowledge and/or consent
New implicit knowledge is disclosed (confidentiality, integrity)
Appropriate use and distribution of discovered knowledge
Is there a need for privacy and DM policies? What are
implications for the Web and e-commerce?
User Interface Issues:
Data visualization
Understandability and interpretation of results
Information representation and rendering
Interactivity
Manipulation of mined knowledge
Focus and refine mining tasks and/or results
Cont...
Mining methodology issues:
Mining different kinds of knowledge in databases
Interactive mining of knowledge at multiple levels of
abstraction
Data mining query languages and ad-hoc data mining
Expression and visualization of data mining results
Handling noise and incomplete data
Pattern evaluation: the interestingness problem
Performance Issues:
Efficiency and scalability of data mining algorithms (need
linear algorithms)
Sampling
Parallel and distributed methods
Cont…
Data source issues:
Diversity of data types
Handling complex types of data
Mining information from heterogeneous
databases and global information systems
(e.g., WWW)
Is it possible to expect a DM system to
perform well on all kinds of data? (do we
need distinct algorithms for distinct data
sources?)
typical
Data Mining System
Graphical user interface
Pattern evaluation
Data mining engine
Database or data Knowledge-
Data warehouse server base
cleaning & Filtering
data Data
integratio Database Warehouse
s
n
37
Summary
Data mining: Discovering interesting patterns and knowledge from massive
amount of data
A natural evolution of science and information technology, in great
demand, with wide applications
A KDD process includes data cleaning, data integration, data selection,
transformation, data mining, pattern evaluation, and knowledge
presentation
Mining can be performed in a variety of data
Data mining functionalities: characterization, discrimination, association,
classification, clustering, trend and outlier analysis, etc.
Data mining technologies and applications
Major issues in data mining
38
ASK YOUR DOUBTS