Data Mining –
Introduction
APTAK@[Link]
Content
• What is Data Mining?
• Origins and development of Data Mining
• Types of data and problems with data
• Methods and techniques of Data Mining
• Examples of applications
What is Data Mining?
Exploring large data sets to discover links, patterns and trends unknown before which can be
used to take better decisions.
Extracting potentially useful information from the data, unknown before Interdisciplinary field
which combines the techniques of machine learning, pattern recognizing, statistics, databases,
and visualizations to obtain information from large databases.
There are many terms (names) connected with this issue:
• Data exploration
• Revealing knowledge from database
• Data exploitation
However, the most frequently used term in literature, on conferences, and among practitioners
is Data Mining.
What is Data Mining ? – definition
Many Definitions
◦ Non-trivial extraction of implicit, previously unknown and potentially useful information
from data
◦ Exploration & analysis, by automatic or semi-automatic means, of large quantities of data
in order to discover meaningful patterns
4
Origins of Data Mining
Draws ideas from machine learning/AI, pattern recognition, statistics, and
database systems
Traditional techniques may be unsuitable due to data that is
◦ Large-scale
◦ High dimensional
◦ Heterogeneous
◦ Complex
◦ Distributed
A key component of the emerging field of data science and data-driven discovery
5
Data Mining development
➢Development of Data Mining methods is connected with the development of computer
techniques, statistical procedures, economic procedures, organizational and management
methods and inference in uncertainty conditions.
➢The development of those methods caused that the starting point in the decision making
process more and more frequently bases on data net on the research hypothesis or theoretical
methods
➢The factors which support the development of Data Mining methods are as follows
➢Huge increase in the data sets volumes,
➢Storing data in warehouses so that all the enterprise have access to utilized database – Access to data
from internet and intranet
➢Development of software packages to explore data,
➢Rapid increase in the calculation powers (capabilities) of computers as well as memory storage,
➢Competition pressure to increase market share in global economy.
Types of DATA
Historically majority of data was collected (for scientific purposes), currently majority are
business operational data, data are opportunistic:
Problems with data
Currently the lack of data and information is not a problem. The problem is insufficient number
of analysts qualified to process the data into knowledge. The knowledge of specialists from
many domains is needed.
• Problem (domain expert) – understands the specific business or scientific problem,
terminology, strong and weak points of existing solutions,
• Data (data expert) – understands data, their structure, size, and format,
• Analytical methods (analytical expert) – understands the possibilities and limitations of
methods which can be applied to the problem.
In general, data users are not scientists. They use the system to take important business
decisions.
Data exploration – main tasks…
• Description
• Estimation Tid Refund
Data
Marital Taxable
Cheat
Status Income
1 Yes Single 125K No
• Prediction 2
3
No
No
Married
Single
100K
70K
No
No
4 Yes Married 120K No
5 No Divorced 95K Yes
• Classification 6
7
No
Yes
Married 60K
Divorced 220K
No
No
8 No Single 85K Yes
9 No Married 75K No
• Clustering 10
11
No
No
Single
Married
90K
60K
Yes
No
12 Yes Divorced 220K No
13 No Single 85K Yes
• Association and sequences 10
14
15
No
No
Married
Single
75K
90K
No
Yes
Milk
Methods and techniques of Data Mining
Data exploration covers many technics from
different disciplines:
• Database technology
• Data visualization
• Statistics
• Information search
• Machine learning
• Image and sound processing
• Pattern recognition
• Spatial analysis
• Neuron networks
Predictive Modeling: Classification
Find a model for class attribute as a function of the values of other attributes
Model for predicting credit
Class Employed
# years at
worthiness
Level of Credit Yes
Tid Employed present No
Education Worthy
address
1 Yes Graduate 5 Yes
2 Yes High School 2 No No Education
3 No Undergrad 1 No
{ High school,
4 Yes High School 10 Yes Graduate
Undergrad }
… … … … …
10
Number of Number of
years years
> 3 yr < 3 yr > 7 yrs < 7 yrs
Yes No Yes No
Classification Example
# years at
Level of Credit
Tid Employed present
Education Worthy
address
1 Yes Undergrad 7 ?
# years at 2 No Graduate 3 ?
Level of Credit
Tid Employed present 3 Yes High School 2 ?
Education Worthy
address
… … … … …
1 Yes Graduate 5 Yes 10
2 Yes High School 2 No
3 No Undergrad 1 No
4 Yes High School 10 Yes
… … … … … Test
10
Set
Training
Learn
Model
Set Classifier
Examples of Classification Task
➢Classifying credit card transactions as legitimate or fraudulent
➢Classifying land covers (water bodies, urban areas, forests, etc.) using satellite data
➢Categorizing news stories as finance, weather, entertainment, sports, etc
➢Identifying intruders in the cyberspace
➢Predicting tumor cells as benign or malignant
➢Classifying secondary structures of protein as alpha-helix, beta-sheet, or random coil
13
Classification: Fraud detection
Financial transactions are used in money laundering, embezzlement and other financial frauds,
smuggling of goods, relations in organized crime or planning terrorism. High value financial
transactions are monitored (personal details, account balance, transaction amount). Sometimes
in order to hide large flows many transactions are made on different accounts. A central
monitoring system will then be following the number of accounts opened by a given person at
small intervals etc.
Goal: Predict fraudulent cases in credit card transactions.
Approach:
➢ Use credit card transactions and the information on its account-holder as attributes.
➢ When does a customer buy, what does he buy, how often he pays on time, etc
➢ Label past transactions as fraud or fair transactions. This forms the class attribute.
➢ Learn a model for the class of the transactions.
➢ Use this model to detect fraud by observing credit card transactions on an account.
It must be remembered that a number of illegal transactions is small in comparison with all high
value transactions, therefore the Data Mining model built must be sensitive.
09/09/2020 INTRODUCTION TO DATA MINING, 2ND EDITION TAN, STEINBACH, KARPATNE, KUMAR 14
Classification: Churn
Many companies adjust their products to the requirements of majority of clients. An important factor
seems to be the client loyalty, i.e. relying on one seller or service provider in consumer behaviours.
The clients were divided into four classes based on their loyalty levels (number of purchase
transactions, volumes, product diversification)
C.a. 10% of the most loyal clients made frequent purchases, high value and within a given product
category. The most numerous group (third, slightly below 40%) made their purchases several times a
year and for low value. The least loyal group did shopping 1-2 times a year and focused on products
which were not available anywhere else.
Churn prediction for telephone customers
◦ Goal: To predict whether a customer is likely to be lost to a competitor.
◦ Approach:
◦ Use detailed record of transactions with each of the past and present customers, to find attributes.
◦ How often the customer calls, where he calls, what time-of-the day he calls most, his
financial status, marital status, etc.
◦ Label the customers as loyal or disloyal.
◦ Find a model for loyalty.
Classification – the improvement in the
quality of telecommunication services
• In every second millions and millions voice connection operations are conducted in the
telecommunication network. Although the quality of the telecommunication equipment is high, many
errors appear during data transmission. All operations are monitored and recorded in the database,
including the errors, thanks to which it is possible to find the reason of the error. The application of
Data Mining techniques will not only make it possible to find the reasons but also allow for predicting
errors in the future.
• However, the users are located in different places in the world and the information about errors are
recorded in different databases. In order to make it possible to apply Data Mining techniques it is
necessary to transfer data to the central warehouse. The raw data concerning connections are
processed to the form of aggregated information on each connection. In this way we obtain historical
data, allowing for prediction.
• The dataset can be divided into the learning set (containing data from the previous week) and the
test set (containing data from the next week). The model obtained in this way will be able to predict
emergency situations one week ahead. The prediction models of neural networks or decision trees
can be applied. • It appeared that the model can be built based on 10%, 20%, 33%, 50%, 67%, 100%
number of cases and the error values in the model can be observed. It appeared that c.a. 33% of data
provided satisfactory results and c.a. 10,000 cases were enough to build a correct model.
Classification – predicting bankruptcies
of companies
➢The economic crisis in East Asia triggered incredible number of company bankruptcies in this
region and worldwide. The data consisted of two groups: Korean companies which went into
bankruptcy in a stable period of 1991-1995 and Korean companies which bankrupted in the
economic crisis of 1997-1998.
➢Based on the literature, researchers identified 40 financial factors including the ratios of growth,
profitability, debt, activity and efficiency.
➢Decision tree models were adopted separately for data in stable conditions and data in crisis
conditions. Based on the models, set of rules were generated. Examples of rules: – If the capital
profitability ratio is bigger than 19.65, predict lack of bankruptcy as precise as 86% – If the relation
of the cash flow ratio to the total assets is equal -5.65 or less, predict bankruptcy as precise as 84%
➢The researchers consulted with the experts in finance in order to interpret the data. To make
sure that the model can be generalized for the set of all Korean companies, a control sample was
selected consisting of non-bankrupt companies and the attributes of both samples were
compared.
Regression
Predict a value of a given continuous valued variable based on the values of
other variables, assuming a linear or nonlinear model of dependency.
Extensively studied in statistics, neural network fields.
Examples:
◦ Predicting sales amounts of new product based on advetising expenditure.
◦ Predicting wind velocities as a function of temperature, humidity, air
pressure, etc.
◦ Time series prediction of stock market indices.
09/09/2020 INTRODUCTION TO DATA MINING, 2ND EDITION TAN, STEINBACH, KARPATNE, KUMAR 18
Clustering
Finding groups of objects such that the objects in a group will be similar (or related) to one another
and different from (or unrelated to) the objects in other groups
Inter-cluster
Intra-cluster distances are
distances are maximized
minimized
Clustering - application
Understanding
◦ Custom profiling for targeted marketing
◦ Group related documents for browsing
◦ Group genes and proteins that have similar functionality
◦ Group stocks with similar price fluctuations
Summarization
◦ Reduce the size of large data sets
Clustering: Market segmentation
◦ Goal: subdivide a market into distinct subsets of customers where any subset may
conceivably be selected as a market target to be reached with a distinct marketing mix.
◦ Approach:
◦ Collect different attributes of customers based on their geographical and lifestyle related
information.
◦ Find clusters of similar customers.
◦ Measure the clustering quality by observing buying patterns of customers in same cluster vs.
those from different clusters.
Clustering: Document clustering
◦ Goal: To find groups of documents that are similar to each other based on the important
terms appearing in them.
◦ Approach: To identify frequently occurring terms in each document. Form a similarity
measure based on the frequencies of different terms. Use it to cluster.
Enron email dataset
Clustering – banking and finance 1/2
• The bank has essential data warehouse where the aggregated data from specific branches are
introduced. Not only standard services for clients are included in the scope of bank activity such as:
granting credits, maintaining accounts and deposits, but also investment management and insurance.
• The bank has decided to withdraw with the mass marketing policies and move to the more specific
activity dedicated to particular clients. Actions were taken in order to implement Data Mining
techniques.
• The implementation was held regularly in 4 stages:
1. Specifying the requirement regarding the effects of Data Mining model.
2. Building the model.
3. Testing the model.
4. Checking correctness of the model.
Clustering – banking and finance 2/2
• Data Mining was directed to select bank’s clients to specific group because of:
◦ 1. Maximum and minimum account balance.
◦ 2. overdraft.
◦ 3. Frequency of cash withdrawal.
• Neural network classifiers and k-means nearest neighbours were applied. Finally the k-means
nearest neighbours was selected, because it allowed for detailed specification of the similarity of
relevant clients.
• Several hundred of groups (clusters) were received sharing similar requirement and needs.
This number of clusters was too big, to create a separate marketing policy for each group, so the
algorithm was applied where 14 groups (clusters) were created only.
Association Rule Discovery: Definition
➢Given a set of records each of which contain some number of items from a given collection.
➢Produce dependency rules which will predict occurrence of an item based on occurrences of
other items.
TID Items
1 Bread, Coke, Milk
Rules Discovered:
2 Beer, Bread {Milk} --> {Coke}
3 Beer, Coke, Diaper, Milk {Diaper, Milk} --> {Beer}
4 Beer, Bread, Diaper, Milk
5 Coke, Diaper, Milk
Association Analysis: Applications
Market-basket analysis
◦ Rules are used for sales promotion, shelf management, and inventory management
Telecommunication alarm diagnosis
◦ Rules are used to find combination of alarms that occur together frequently in the same time
period
Medical Informatics
◦ Rules are used to find combination of patient symptoms and test results associated with
certain diseases
Motivating Challenges
➢Scalability
➢High Dimensionality
➢Heterogeneous and Complex Data
➢Data Ownership and Distribution
➢Non-traditional Analysis
Summary
Many people think that Data Mining means magical finding of the hidden information without
formulating a problem and without considering the content of data. But it is not like that Data
exploration is a field which must be learned, it is not a ready-made product, i.e. software
delivering the solution to problems without the need of human cooperation and supervision. A
broader term is applicable – searching for knowledge in databases (1989)
• KDD (Knowledge discovery in databases) is an multidisciplinary field connected with finding
patterns from large data sets
• Pattern recognition
• Machine learning
• Neurocomputing – connected with neural networks
• Sometimes KDD is identified with DM. More specifically, DM is one of the stages in the process
of searching for knowledge in databases.