0% found this document useful (0 votes)
5 views39 pages

Introduction to Data Mining Concepts

The document provides an introduction to data mining, outlining its significance due to the exponential growth of data and the need for knowledge extraction. It covers the definition of data mining, its applications in various fields such as business intelligence, finance, and healthcare, and discusses the knowledge discovery process. Additionally, it highlights different data mining functionalities including classification, clustering, and outlier analysis, along with the challenges of identifying interesting patterns.

Uploaded by

debrituandarsa
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPT, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
5 views39 pages

Introduction to Data Mining Concepts

The document provides an introduction to data mining, outlining its significance due to the exponential growth of data and the need for knowledge extraction. It covers the definition of data mining, its applications in various fields such as business intelligence, finance, and healthcare, and discusses the knowledge discovery process. Additionally, it highlights different data mining functionalities including classification, clustering, and outlier analysis, along with the challenges of identifying interesting patterns.

Uploaded by

debrituandarsa
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPT, PDF, TXT or read online on Scribd

Adama Science & Technology

University
CoEEC:CS
Data Mining(CoSc4152)
Introduction

1
Lecture Outline
 Why Data Mining?
 What Is Data Mining?
 Data Can Be Mined
 Patterns that Can Be Mined
 Technologies Used
 Applications Area of Data Mining
 Major Issues in Data Mining
 Summary

2
Why Data Mining?
 The Explosive Growth of Data: from terabytes to petabytes
 Data collection and data availability


Automated data collection tools, database systems, Web,
computerized society
 Major sources of abundant data


Business: Web, e-commerce, transactions, stocks, …

Science: Remote sensing, bioinformatics, scientific
simulation, …

Society and everyone: news, digital cameras, YouTube
 We are drowning in data, but starving for knowledge!
 “Necessity is the mother of invention”—Data mining—
Automated analysis of massive data sets
3
What Is Data Mining?

 Data mining (knowledge discovery from data)


 Extraction of interesting (non-trivial, implicit, previously
unknown and potentially useful) patterns or knowledge from
huge amount of data
 Data mining: a misnomer?
 Alternative names
 Knowledge discovery (mining) in databases (KDD), knowledge
extraction, data/pattern analysis, data archeology, data
dredging, information harvesting, business intelligence, etc.
 Watch out: Is everything “data mining”?
 Simple search and query processing
 (Deductive) expert systems
4
What Is Data Mining?
 Data mining is the science of discovering
structure and making predictions in large or
complex data sets

5
Knowledge Discovery (KDD) Process
 This is a view from typical database
systems and data warehousing
Pattern Evaluation
communities
 Data mining plays an essential role in
the knowledge discovery process
Data Mining

Task-relevant Data

Data Selection
Warehouse
Data Cleaning

Data Integration

Databases
6
Data Mining in Business Intelligence

Increasing potential
to support
business decisions End User
Decisio
n
Making
Data Presentation Business
Analyst
Visualization Techniques
Data Mining Data
Information Discovery Analyst

Data Exploration
Statistical Summary, Querying, and Reporting

Data Preprocessing/Integration, Data Warehouses


DBA
Data Sources
Paper, Files, Web documents, Scientific experiments, Database Systems
7
KDD Process: A Typical View from ML
and Statistics

Input Data Data Pre- Data Post-


Processing Mining Processin
g

Data integration Pattern discovery Pattern evaluation


Normalization Association & Pattern selection
correlation
Feature selection Classification Pattern
interpretation
Dimension reduction Clustering
Pattern visualization
Outlier analysis
…………

 This is a view from typical machine learning and statistics communities

8
Why Data Mining?—Potential Applications

 Data analysis and decision support


 Market analysis and management

Target marketing, customer relationship management (CRM), market
basket analysis, cross selling, market segmentation
 Risk analysis and management

Forecasting, customer retention, improved underwriting, quality
control, competitive analysis
 Fraud detection and detection of unusual patterns (outliers)
 Other Applications
 Text mining (news group, email, documents) and Web mining
 Intelligent query answering
 DNA and bio-data analysis
9
Why Data Mining?—Potential Applications

 Market Analysis and Management(1)


 Where does the data come from?
 Credit card transactions, loyalty cards, discount coupons, customer complaint calls, plus
(public) lifestyle studies

 Target marketing
 Find clusters of “model” customers who share the same characteristics: interest, income
level, spending habits, etc.
 Determine customer purchasing patterns over time

 Cross-market analysis
 Associations/co-relations between product sales, & prediction based on such association

10
Why Data Mining?—Potential Applications
 Market Analysis and Management(2)
 Customer profiling
 What types of customers buy what products (clustering or classification)
 Customer requirement analysis
 identifying the best products for different customers
 predict what factors will attract new customers
 Provision of summary information
 multidimensional summary reports
 statistical summary information (data central tendency and variation)

11
Why Data Mining?—Potential Applications
 Corporate Analysis & Risk Management
 Finance planning and asset evaluation

cash flow analysis and prediction

contingent claim analysis to evaluate assets

cross-sectional and time series analysis (financial-ratio, trend analysis,
etc.)
 Resource planning

summarize and compare the resources and spending
 Competition

monitor competitors and market directions

group customers into classes and a class-based pricing procedure

set pricing strategy in a highly competitive market

12
Why Data Mining?—Potential Applications

 Approaches: Clustering & model construction for frauds, outlier analysis, based on historical data
 Applications: Health care, retail, credit card service, telecomm.
 Auto insurance: detect a group of people who stage accidents to collect insurance
 Money laundering: suspicious monetary transactions
 Medical insurance

Professional patients, ring of doctors, and ring of references

Unnecessary or correlated screening tests
 Detecting inappropriate medical treatment
 Telecommunications: phone-call fraud

Phone call model: destination of the call, duration, time of day
or week. Analyze patterns that deviate from an expected

13
Why Data Mining?—Potential Applications
 Financial Data Analysis
 Financial data

complete

reliable

high quality
 Loan payment prediction and customer credit policy analysis
 Factors influencing loan payment performance

loan-to-value ratio

term of the loan

debt ratio (total monthly debt/total monthly income)

payment-to-income ratio

income level

education level

residence region

credit history
 Analysis may find that

payment-income ratio is a dominant factor while

education level and debt ratio are not
14
Why Data Mining?—Potential Applications
 Data Mining for the Retail Industry
 Multidimensional analysis of sales, customers, products, time and region

OLAP cubes
 Effectiveness of sales campaigns

Advertisements, coupons, discounts, bonuses

promote products and attract customers

can help improve profits

Compare amount of sales and number of transactions

during the sales period versus before or after the sales campaign

Association analysis

which items are likely to be purchased together with the items on sale
 Customer retention Analysis of Customer loyalty

sequences of purchases of particular customers

goods purchased at different periods by the same customers can be grouped into sequences

changes in customer consumption or loyalty

suggests adjustments on the pricing and variety of goods

to retain old customers and attract new customers
 Purchase recommendation and cross-reference of items

associations from sales records

a customer who buy a PC is likely to buy a printer

purchase recommendations

15
Why Data Mining?—Potential Applications

 Data Mining for the Telecommunication Industry

 Telecommunication data are multidimensional


 calling-time duration
 location of caller location of called
 type of call

 used to identify and compare


 data traffic system workload
 resource usage user group behaviour
 profit

 fraudulent pattern analysis and identification of unusual patterns


 to achieve customer loyalty
 characteristics of customers affecting line usage

16
Multi-Dimensional View of Data
Mining
 Data to be mined
 Database data (extended-relational, object-oriented, heterogeneous, legacy),

data warehouse, transactional data, stream, spatiotemporal, time-series,


sequence, text and web, multi-media, graphs & social and information
networks
 Knowledge to be mined (or: Data mining functions)
 Characterization, discrimination, association, classification, clustering,

trend/deviation, outlier analysis, etc.


 Descriptive vs. predictive data mining

 Multiple/integrated functions and mining at multiple levels

 Techniques utilized
 Data-intensive, data warehouse (OLAP), machine learning, statistics, pattern

recognition, visualization, high-performance, etc.


 Applications adapted
 Retail, telecommunication, banking, fraud analysis, bio-data mining, stock

market analysis, text mining, Web mining, etc. 17


Data Mining: On What Kinds of
Data?
 Database-oriented data sets and applications
 Relational database, data warehouse, transactional database
 Object-relational databases, Heterogeneous databases and legacy
databases
 Advanced data sets and advanced applications
 Data streams and sensor data
 Time-series data, temporal data, sequence data (incl. bio-sequences)
 Structure data, graphs, social networks and information networks
 Multimedia database
 Text databases
 The World-Wide Web
18
Data Mining Function/Tasks
 Characterization
 Discrimination
 Association and Correlation Analysis
 Classification
 Cluster Analysis
 Outlier Analysis

19
Data Mining Functionalities / Tasks (1)
Characterization & Discrimination
 Class/Concept Description
 Data characterization is a summarization of general features of objects in
a target class

The data relevant to a target class are retrieved by a database query and run through a
summarization module to extract the essence of the data at different levels of
abstraction

example: characterize the video store’s customers who regularly rent more than 30
movies a year

 Data discrimination is used to compare of the general features of objects


between two classes

comparison relevant features of objects between a target class and a contrasting class

example: compare the general characteristics of the customers who rented more than
30 movies in the last year with those who rented less than 5

The techniques used for data discrimination are very similar to the techniques used for
data characterization with the exception that data discrimination results include
comparative measures

20
Data Mining Function: (1)
Association and Correlation Analysis
 Frequent patterns (or frequent itemsets)
 What items are frequently purchased together in your
Shoa supermarket?
 Association, correlation vs. causality
 A typical association rule

Milk  Diapper [0.5%, 75%] (support, confidence)
 Are strongly associated items also strongly correlated?
 How to mine such patterns and rules efficiently in large
datasets?
 How to use such patterns for classification, clustering, and
other applications?
21
Data Mining Function: (3)
Classification
 Classification and label prediction
 Construct models (functions) based on some training examples
 Describe and distinguish classes or concepts for future prediction

E.g., classify countries based on (climate), or classify cars
based on (gas mileage)
 Predict some unknown class labels
 Typical methods
 Decision trees, naïve Bayesian classification, support vector
machines, neural networks, rule-based classification, pattern-
based classification, logistic regression, …
 Typical applications:
 Credit card fraud detection, direct marketing, classifying stars,
diseases, web-pages, …
22
Data Mining Functionalities / Tasks (4)
Prediction
 Same as classification and estimation

except classify according to some predicted or estimated future
value
 In prediction, historical data is used to build a (predictive) model that
explains the current observed behavior

Model can then be applied to new instances to predict future
behavior or forecast the future value of some missing attribute
 Examples

predicting the size of balance transfer if the prospect accepts the
offer

predicting the load on a Web server in a particular time period
 Suitable data mining tools

Association rule discovery

Decision Trees

23
Figure: (a) Affinity Grouping, (b) Decision Trees, (c)
Neural Networks
24
Data Mining Function: (4) Cluster
Analysis
 Unsupervised learning (i.e., Class label is unknown)
 Group data to form new categories (i.e., clusters), e.g., cluster
houses to find distribution patterns
 Principle: Maximizing intra-class similarity & minimizing
interclass similarity
 Many methods and applications

25
Data Mining Function: (5) Outlier
Analysis
 Outlier analysis
 Outlier: A data object that does not comply with the general
behavior of the data
 Noise or exception? ― One person’s garbage could be
another person’s treasure
 Methods: by product of clustering or regression analysis, …
 Useful in fraud detection, rare events analysis

26
Can We Find All and Only Interesting Patterns?(1)

 A data mining operation may generate thousands of


patterns

not all of them are interesting

e.g., the association rule {bread, milk} ==> {butter}
 Possible Approaches

human-centered - use subjective measures of interestingness
and observation to find interesting patterns

query-based - use a “knowledge query” mechanism to query
the results after the mining operations

focused mining - establish constraints during mining to narrow
down result sets
 But, how do we “measure” interestingness?

27
Can We Find All and Only Interesting Patterns?(2)

 Objective vs. subjective interestingness measures



Objective: based on statistics and structures of patterns,
e.g., support, confidence, lift, etc.

Subjective: based on user’s beliefs (or domain-specific
beliefs) related to the data, e.g., “unexpectedness”,
“novelty”, etc.
 Interestingness measures: A pattern is interesting if it is

easily understood by humans

valid on new or test data with some degree of certainty

potentially useful / actionable

novel, or validates some hypothesis that a user seeks to
confirm

28
Can We Find All and Only Interesting Patterns?(3)

 Find all the interesting patterns: Completeness



Can a data mining system find all the interesting patterns?
 Search for only interesting patterns: Optimization

Can a data mining system find only the interesting
patterns?

Approaches

First find all the patterns and then filter out the uninteresting ones

Generate only the interesting patterns -- mining query optimization

29
Structure and Network Analysis
 Graph mining
 Finding frequent subgraphs (e.g., chemical compounds), trees (XML),

substructures (web fragments)


 Information network analysis
 Social networks: actors (objects, nodes) and relationships (edges)


e.g., author networks in CS, terrorist networks
 Multiple heterogeneous networks


A person could be multiple information networks: friends, family,
classmates, …
 Links carry a lot of semantic information: Link mining

 Web mining
 Web is a big information network: from PageRank to Google

 Analysis of Web information networks


Web community discovery, opinion mining, usage mining, …

30
Data Mining: Classification
Schemes
 General functionality
 Descriptive data mining
 Predictive data mining
 Different views, different classifications
 Kinds of databases to be mined
 Kinds of knowledge to be discovered
 Kinds of techniques utilized
 Kinds of applications adapted

31
Data Mining: Confluence of Multiple
Disciplines

Machine Pattern Statistics


Learning Recognition

Applications Data Mining Visualization

Algorithm Database High-Performance


Technology Computing

32
Why Confluence of Multiple
Disciplines?
 Tremendous amount of data
 Algorithms must be scalable to handle big data
 High-dimensionality of data
 Micro-array may have tens of thousands of dimensions
 High complexity of data
 Data streams and sensor data
 Time-series data, temporal data, sequence data
 Structure data, graphs, social and information networks
 Spatial, spatiotemporal, multimedia, text and Web data
 Software programs, scientific simulations
 New and sophisticated applications

33
Different Issues in Data
Mining
 Security and social issues:

Private and sensitive data is gathered and mined without
individual’s knowledge and/or consent

New implicit knowledge is disclosed (confidentiality, integrity)

Appropriate use and distribution of discovered knowledge

Is there a need for privacy and DM policies? What are
implications for the Web and e-commerce?
 User Interface Issues:
 Data visualization

Understandability and interpretation of results

Information representation and rendering
 Interactivity

Manipulation of mined knowledge

Focus and refine mining tasks and/or results
Cont...
 Mining methodology issues:

Mining different kinds of knowledge in databases

Interactive mining of knowledge at multiple levels of
abstraction

Data mining query languages and ad-hoc data mining

Expression and visualization of data mining results

Handling noise and incomplete data

Pattern evaluation: the interestingness problem
 Performance Issues:
 Efficiency and scalability of data mining algorithms (need
linear algorithms)
 Sampling
 Parallel and distributed methods
Cont…

Data source issues:



Diversity of data types

Handling complex types of data

Mining information from heterogeneous
databases and global information systems
(e.g., WWW)

Is it possible to expect a DM system to
perform well on all kinds of data? (do we
need distinct algorithms for distinct data
sources?)
typical
Data Mining System

Graphical user interface

Pattern evaluation

Data mining engine

Database or data Knowledge-


Data warehouse server base
cleaning & Filtering
data Data
integratio Database Warehouse
s
n
37
Summary
 Data mining: Discovering interesting patterns and knowledge from massive
amount of data
 A natural evolution of science and information technology, in great
demand, with wide applications
 A KDD process includes data cleaning, data integration, data selection,
transformation, data mining, pattern evaluation, and knowledge
presentation
 Mining can be performed in a variety of data
 Data mining functionalities: characterization, discrimination, association,
classification, clustering, trend and outlier analysis, etc.
 Data mining technologies and applications
 Major issues in data mining

38
ASK YOUR DOUBTS

You might also like