0% found this document useful (0 votes)
0 views8 pages

CDSModule 2 Notes

Data mining is the process of extracting useful information from large datasets to identify patterns and trends that aid in data-driven decision-making. It is categorized into predictive and descriptive analyses, with various techniques such as classification, regression, clustering, and association rule mining. The document also discusses challenges, advantages, disadvantages, and applications of data mining across different sectors including healthcare, finance, and education.

Uploaded by

tanmayibv2006
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
0 views8 pages

CDSModule 2 Notes

Data mining is the process of extracting useful information from large datasets to identify patterns and trends that aid in data-driven decision-making. It is categorized into predictive and descriptive analyses, with various techniques such as classification, regression, clustering, and association rule mining. The document also discusses challenges, advantages, disadvantages, and applications of data mining across different sectors including healthcare, finance, and education.

Uploaded by

tanmayibv2006
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Global Academy of Technology, Bengaluru

Computer Science and Engineering


Course: Computational Data Science
Module 2: Data Mining
2.1 What is Data Mining
 The process of extracting information to identify patterns, trends, and useful data that would
allow the business to take the data-driven decision from huge sets of data is called Data
Mining.
 In other words, it is the process of investigating hidden patterns of information from various
perspectives to generate revenue and achieve cost-cutting.
 Data mining is the act of automatically searching for large stores of information to find
trends and patterns that go beyond simple analysis procedures.
 Data mining utilizes complex mathematical algorithms for data segments and evaluates the
probability of future events. Data Mining is also called Knowledge Discovery of Data
(KDD).

2.2 Types of Data Mining


Data mining can be broadly categorized into two types –
1. Predictive Data Mining Analysis
2. Descriptive Data Mining Analysis.

Predictive Data Mining Analysis

 Predictive data mining analysis is one of the types of data mining that involves using
historical data to make predictions about future events or trends. It involves building models
using various statistical and machine learning algorithms to forecast future outcomes based
on patterns found in past data.
 The models generated from predictive data mining can be used to predict various scenarios,
such as predicting customer behavior, identifying potential fraud, forecasting sales, and
predicting the likelihood of a disease occurring in a particular population.
 Predictive types of data mining typically involve a large amount of data preparation and
preprocessing, as well as selecting appropriate algorithms to build models that can
accurately predict future outcomes.
 Predictive Data-Mining can also be further divided into four types that are listed below:

1. Classification Analysis
2. Regression Analysis
3. Time Serious Analysis
4. Prediction Analysis

Descriptive Data Mining Analysis.

 Descriptive data mining analysis is one of the types of data mining that focuses on exploring
and understanding the underlying patterns and relationships within a dataset. Unlike
predictive data mining, descriptive data mining is not concerned with making predictions
about future events but rather with summarizing and visualizing the data to gain insights into
its structure and characteristics.
 Descriptive types of data mining techniques are often used for exploratory data analysis, to
discover patterns and relationships in data, and to gain a deeper understanding of the data's
underlying distribution.
 The Descriptive Data-Mining Tasks can also be further divided into four types that are as
follows:

1. Clustering Analysis
2. Summarization Analysis
3. Association Rules Analysis
4. Sequence Discovery Analysis

2.3 Challenges of Implementation in Data Mining

Figure 2.1: Challenges in Data Mining


1. Incomplete and noisy data:
 The data in the real-world is heterogeneous, incomplete and noisy.
 Data in huge quantities will usually be inaccurate or unreliable.
 These problems may occur due to data measuring instrument or because of human errors.
Example: Suppose a retail chain collects phone numbers of customers who spend more than $ 500,
and the accounting employees put the information into their system. The person may make a digit
mistake when entering the phone number, which results in incorrect data. Even some customers
may not be willing to disclose their phone numbers, which results in incomplete data. The data
could get changed due to human or system error.
All these consequences (noisy and incomplete data) make data mining challenging.

2. Data Distribution:
 Real-worlds data is usually stored on various platforms in a distributed computing
environment.
 It might be in a database, individual systems, or even on the internet.
 Practically, it is a quite tough task to make all the data to a centralized data repository
mainly due to organizational and technical concerns.
For example, various regional offices may have their servers to store their data. It is not feasible to
store, all the data from all the offices on a central server.

3. Complex Data:
 Real-world data is heterogeneous, and it could be multimedia data, including audio and
video, images, complex data, spatial data, time series, and so on.
 Managing these various types of data and extracting useful information is a tough task.
 Most of the time, new technologies, new tools, and methodologies would have to be refined
to obtain specific information.

4. Performance:
 The data mining system's performance relies primarily on the efficiency of algorithms and
techniques used.
 If the designed algorithm and techniques are not up to the mark, then the efficiency of the
data mining process will be affected adversely.

5. Data Privacy and Security:


 Data mining usually leads to serious issues in terms of data security, governance, and
privacy.
 For example, if a retailer analyzes the details of the purchased items, then it reveals data
about buying habits and preferences of the customers without their permission.

6. Data Visualization:
 In data mining, data visualization is the primary method that shows the output to the user in
a presentable way.
 The extracted data should convey the exact meaning of what it intends to express.
 But many times, representing the information to the end-user in a precise and easy way is
difficult.

2.4 Advantages and Disadvantages


Advantages:
o The Data Mining technique enables organizations to obtain knowledge-based data.
o Data mining enables organizations to make lucrative modifications in operation and
production.
o Compared with other statistical data applications, data mining is a cost-efficient.
o Data Mining helps the decision-making process of an organization.
o It Facilitates the automated discovery of hidden patterns as well as the prediction of trends
and behaviors.
o It can be induced in the new system as well as the existing platforms.
o It is a quick process that makes it easy for new users to analyze enormous amounts of data in
a short time.

Disadvantages:
o There is a probability that the organizations may sell useful data of customers to other
organizations for money. As per the report, American Express has sold credit card purchases
of their customers to other organizations.
o Many data mining analytics software is difficult to operate and needs advance training to
work on.
o Different data mining instruments operate in distinct ways due to the different algorithms
used in their design. Therefore, the selection of the right data mining tools is a very
challenging task.
o The data mining techniques are not precise, so that it may lead to severe consequences in
certain conditions.

2.5 Applications of Data Mining


Data Mining is primarily used by organizations with intense consumer demands- Retail,
Communication, Financial, marketing company, determine price, consumer preferences, product
positioning, and impact on sales, customer satisfaction, and corporate profits.
Figure 2.2: Applications of Data Mining
Data Mining in Healthcare:
 Data Mining uses data and analytics for better insights and to identify best practices that will
enhance health care services and reduce costs.
 Analysts use data mining approaches such as Machine learning, multi-dimensional database,
Data visualization, soft computing, and statistics.
 Data Mining can be used to forecast patients in each category. The procedures ensure that
the patients get intensive care at the right place and at the right time.
 Data mining also enables healthcare insurers to recognize fraud and abuse.

Data Mining in Market Basket Analysis:


 Market basket analysis is a modeling method based on a hypothesis.
 If you buy a specific group of products, then you are more likely to buy another group of
products.
 This technique may enable the retailer to understand the purchase behavior of a buyer.
 This data may assist the retailer in understanding the requirements of the buyer and altering
the store's layout accordingly.
 Analytical comparison of results between various stores, between customers in different
demographic groups can also be done.

Data mining in Education:


 Education data mining is a newly emerging field, concerned with developing techniques that
explore knowledge from the data generated from educational Environments
 EDM objectives are recognized as affirming student's future learning behavior, studying the
impact of educational support, and promoting learning science.
 An organization can use data mining to make precise decisions and also to predict the results
of the student. With the results, the institution can concentrate on what to teach and how to
teach.
Data Mining in Manufacturing Engineering:
 Knowledge is the best asset possessed by a manufacturing company.
 Data mining tools can be beneficial to find patterns in a complex manufacturing process.
 Data mining can be used in system-level designing to obtain the relationships between
product architecture, product portfolio, and data needs of the customers.
 It can also be used to forecast the product development period, cost, and expectations among
the other tasks.

Data Mining in CRM (Customer Relationship Management):


 Customer Relationship Management (CRM) is all about obtaining and holding Customers,
also enhancing customer loyalty and implementing customer-oriented strategies.
 To get a decent relationship with the customer, a business organization needs to collect data
and analyze the data.
 With data mining technologies, the collected data can be used for analytics.

Data Mining in Fraud detection:


 Billions of dollars are lost to the action of frauds.
 Traditional methods of fraud detection are a little bit time consuming and sophisticated.
 Data mining provides meaningful patterns and turning data into information.
 An ideal fraud detection system should protect the data of all the users.
 Supervised methods consist of a collection of sample records, and these records are
classified as fraudulent or non-fraudulent.
 A model is constructed using this data, and the technique is made to identify whether the
document is fraudulent or not.

Data Mining in Lie Detection:


 Apprehending a criminal is not a big deal, but bringing out the truth from him is a very
challenging task.
 Law enforcement may use data mining techniques to investigate offenses, monitor suspected
terrorist communications, etc.
 This technique includes text mining also, and it seeks meaningful patterns in data, which is
usually unstructured text.
 The information collected from the previous investigations is compared, and a model for lie
detection is constructed.

Data Mining in Financial Banking:


 The Digitalization of the banking system is supposed to generate an enormous amount of
data with every new transaction.
 The data mining technique can help bankers by solving business-related problems in
banking and finance by identifying trends, casualties, and correlations in business
information and market costs.
 These patterns are not instantly evident to managers or executives because the data volume
is too large or are produced too rapidly on the screen by experts.
 The manager may find these data for better targeting, acquiring, retaining, segmenting, and
maintain a profitable customer.

2.6 Overview of Basic Data Mining Tasks


Basic Data Mining Tasks are
 Classification
 Regression
 Time Series Analysis
 Prediction
 Clustering
 Sequence Discovery

Classification:
This type of data mining technique is generally used in fetching or retrieving important and relevant
information about the data & metadata. It is also even used to categories the different types of data
format into different classes.

In the classification analysis, you have to apply or implement the algorithms to decide in which way
the new data should be categorized or classified. A classic example of classification analysis would
be Outlook email. In Outlook, they use certain algorithms to characterize an email is legitimate or
spam.

This technique is usually very helpful for retailers who can use it to study the buying habits of their
different customers. Retailers can also study the past sales data and then lookout (or search) for
products that customers usually buy together. After which, they can put those products nearby of
each other in their retail stores to help customers save their time and as well as to increase their
sales.

Regression:
Regression is a statistical data mining technique used to analyze the relationship between a
dependent variable and one or more independent variables. The goal is to build a model that can
predict the dependent variable value based on the independent variables' values. Regression is used
in data mining to identify patterns, trends, and relationships between variables and to make
predictions and forecasts.

Time Series Analysis:


Time series analysis is one of the types of data mining techniques used to analyze sequential data,
such as time-series data. The goal is to identify patterns, trends, and relationships in the data over
time. Time series analysis is used in data mining for various applications such as stock market
forecasting, weather prediction, and trend analysis.

Prediction Analysis:
This technique is generally used to predict the relationship that exists between both the independent
and dependent variables as well as the independent variables alone. It can also use to predict profit
that can be achieved in future depending on the sale. Let us imagine that profit and sale are
dependent and independent variables, respectively. Now, on the basis of what the past sales data
says, we can make a profit prediction of the future using a regression curve.

Clustering:
Clustering is a data mining technique used to group similar objects based on their characteristics.
The goal is to identify natural groupings or clusters in the data. Clustering is used in data mining for
various purposes, such as customer segmentation, image segmentation, and anomaly detection.

Association rule mining:


Association rule mining is a data mining technique used to discover relationships between variables
in large datasets. The goal is to identify rules that describe the relationships between variables.
Association rule mining is used in data mining for various purposes, such as market basket analysis,
product recommendation, and website navigation analysis.

Summarization:
Summarization is one of the types of data mining techniques used to summarize the characteristics
of a dataset into a more compact and understandable form. The goal is to provide an overview of the
dataset and highlight the most important aspects. Summarization is used in data mining for various
purposes, such as data visualization, report generation, and decision-making.

Sequential Pattern:
Sequential pattern mining is a data mining technique used to discover patterns in sequential
data, such as time series or transactional data. The goal is to identify frequent patterns or
sequences of events that occur in the data. Sequential pattern mining is used in data mining
for various applications such as web log analysis, customer behavior analysis, and DNA
sequence analysis.

You might also like