0% found this document useful (0 votes)
3 views144 pages

Module 1

Data mining is the process of using computers to analyze large data sets for patterns and trends, leading to business insights and predictions. The data mining process includes phases such as business understanding, data understanding, data preparation, modeling, evaluation, and deployment. Data warehousing supports decision-making by organizing and integrating data from various sources, allowing for complex analysis and strategic insights.

Uploaded by

danmathews575
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views144 pages

Module 1

Data mining is the process of using computers to analyze large data sets for patterns and trends, leading to business insights and predictions. The data mining process includes phases such as business understanding, data understanding, data preparation, modeling, evaluation, and deployment. Data warehousing supports decision-making by organizing and integrating data from various sources, allowing for complex analysis and strategic insights.

Uploaded by

danmathews575
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

CST466 -DATA MINING

DATA MINING
• Data mining is most commonly defined as the process of using computers
and automation to search large sets of data for patterns and trends,
turning those findings into business insights and predictions.

• Data mining goes beyond the search process, as it uses data to evaluate
future probabilities and develop actionable analyses.

• Data mining is the process of finding patterns in data.


PHASES OF DATA MINING
• Business Understanding
• To get started, first ask these questions: What is our objective? What
problem are we trying to solve? What data do we need to solve it?

• Without a clear understanding of the proper data to mine, the project can
produce errors, inaccurate results, or results that don’t answer the correct
questions.
• Data Understanding
• Once the overall objective is determined, proper data needs to be collected.

• The data must be relevant to subject matter and usually comes from a
variety of sources such as sales records, customer surveys, and geolocation
data.

• This phase’s goal is to ensure the data correctly encompasses all necessary
data sets to address the objective.
• Data Preparation
• The most time-consuming phase, the preparation phase, consists of three
steps: extraction, transformation, and loading — also referred to as ETL.
• First, data is extracted from various sources and deposited into a staging
area.
• Next, during the transformation step: the data is cleaned, null sets are
populated, duplicative data is removed, errors are resolved, and all data is
allocated into tables.
• In the final step, loading, the formated data is loaded into the database
for use.
• Modeling
• Data modeling addresses the relevant data set and considers the best
statistical and mathematical approach to answering the objective
question(s).

• There are a variety of modeling techniques available, such as


classification, clustering, and regression analysis .

• It’s also common to use different models on the same data to address
specific objectives.
• Evaluation
• After the models are built and tested, it’s time to evaluate their efficiency in
answering the question identified during the business understanding
phase. This is a human-driven phase, as the individual running the project
must determine whether the model output sufficiently meets their
objectives. If not, a different model can be created, or different data can be
prepared.
• Deployment
• Once the data mining model is seemed accurate and successful in
answering the objective question, it’s time to put it to use. Deployment can
occur in the form of a visual presentation or a report sharing insights. It also
can lead to action such as generating a new sales strategy or implementing
risk-reduction measures.
Most Common Types of Data Mining
• Classification Analysis
• With this technique, data points are assigned to groups, or classes, based
on a specific question or problem to address.
• Association Rule Learning
• This function seeks to uncover the relationships between data points; it is
used to determine whether a specific action or variable has any traits that
can be linked to other actions (e.g., business travelers’ room choices and
dining habits). A hotelier might use association rule insights to offer room
upgrades or food and beverage promotions to attract additional business
travelers.
• Anomaly or Outlier Detection
• In addition to searching for patterns, data mining seeks to uncover unusual
data within a set. Anomaly detection is the process of finding data that
doesn’t conform to the pattern.
• Clustering Analysis
• Clustering looks for similarities within a data set, separating data points
that share common traits into subsets.
• Regression Analysis
• Regression analysis is about understanding which factors within a data set
are most important, which can be ignored, and how these factors interact.
With this technique, data miners are able to validate theories such as
“when a lot of snow is predicted, more bread and milk will be sold before
the snow.”
What is Data Warehouse

• Data warehousing provides architectures and tools for business


executives to systematically organize, understand, and use their data
to make strategic decision.

• A data warehouse refers to a data repository that is maintained


separately from an organization’s operational databases.

• Data warehouse systems allow for integration of a variety of


application systems. They support information processing by
providing a solid platform of consolidated historic data for analysis.
A data warehouse is a subject-oriented, integrated, time-variant, and
nonvolatile collection of data in support of management’s decision
making process.

Subject-oriented: A data warehouse is organized around major subjects


such as customer, supplier, product, and sales. Rather than concentrating
on the day-to-day operations and transaction processing of an
organization, a data warehouse focuses on the modeling and analysis of
data for decision makers.
• Integrated: A data warehouse is usually constructed by integrating multiple
heterogeneous sources, such as relational databases, flat files, and online
transaction records. Data cleaning and data integration techniques are
applied to ensure consistency in naming conventions, encoding structures,
attribute measures, and so on.

• Time-variant: Data are stored to provide information from an historic


perspective (e.g., the past 5–10 years). Every key structure in the data
warehouse contains, either implicitly or explicitly, a time element.

• Nonvolatile: A data warehouse is always a physically separate store of data


transformed from the application data found in the operational environment.
Due to this separation, a data warehouse does not require transaction
processing, recovery, and concurrency control mechanisms. It usually requires
only two operations in data accessing: initial loading of data and access of
data.
• A data warehouse is a semantically consistent data store that serves as a
physical implementation of a decision support data model.

• It stores the information an enterprise needs to make strategic decisions.

• The construction of a data warehouse requires data cleaning, data


integration, and data consolidation.
• Many organizations use this information to support business decision-making
activities, including
(1) increasing customer focus, which includes the analysis of customer buying
patterns (such as buying preference, buying time, budget cycles, and appetites
for spending);
(2) repositioning products and managing product portfolios by comparing the
performance of sales by quarter, by year, and by geographic regions in order to
fine-tune production strategies;
(3) analyzing operations and looking for sources of profit; and
(4) managing customer relationships, making environmental corrections, and
managing the cost of corporate assets.
• Rather than using a query-driven approach, data warehousing
employs an update driven approach in which information from
multiple, heterogeneous sources is integrated in advance and stored
in a warehouse for direct querying and analysis. Unlike online
transaction processing databases, data warehouses do not contain
the most current information.
Need for Data Warehousing
• For a company to be successful, they must make good decisions. For making
good decisions require all relevant data to ranking into consideration.

• The best source for this data is a well designed Data Warehouse

• Performing OLAP in Operational DB is time consuming

• Operational DB is designed for known tasks & workloads like hashing,


indexing, searching for particular records etc.

• DWH queries are more complex, involves the computation of large data
groups summarized level
Advantages & Disadvantages of DWH
• ADVANTAGES
• Increase productivity and decrease computing costs
• It is helpful to predict future trends
• Helps in decision making
• Able to combine data from different sources into one place
• DISADVANTAGES
• Extracting, cleaning and loading data is time consuming

• Privacy Issues: collects information about people that are using some market-
based techniques and information technology

• Provide training to end users: require a very skilled specialist person to prepare
the data and understand the output

• Security Issues: Especially DWH is web accessible, some of this data which is
very critical might be hacked by hackers

• Misuse of information: Some can misuse this information to harm others in


their own way.
Multidimensional Data Model
• Data warehouses and OLAP tools are based on a multidimensional data
model.
• This model views data in the form of a data cube
• The actual physical storage of such data may differ from this logical
representation.
• A data cube allows data to be modeled and viewed in multiple dimensions.
• It is defined by dimensions and facts.
• Dimensions are the perspectives or entities with respect to which an
organization wants to keep records
• Dimension table: A table associated with each dimension, which further
describes the dimension.
• Table represents a 2-D table
• It shows employment in California by sex ,by year and by profession.
• The rows and columns represent more than one dimension
• The rows in table 2.1 represent the two dimensions
Sex and Year which are in arbitrary order.
• Represent sex first after that year.
• The columns do not represent 2 distinct dimensions, but it represent some
sort of taxonomy of a dimension
• summary information is maintained which is the main theme of the table.
Here the summary function is sum
DATA CUBE
• A Popular conceptual model for DATA WAREHOUSE is the
Multidimensional view of data.
• This model views data in the form of a data cube(or Hypercube)
• It has 3 dimensions namely sex, profession and year
• Each dimension can be sub divided in to subdimensions.
• In Multidimensional data model there is a set of numeric measures,
That are the main theme or subject of the analysis.
From the above example the numeric measure is Employment
• We can have more than one numeric measure.
• Some examples of numeric measures are sales,budget,revenue etc..
• All the dimensions together are assumed to uniquely determine the
measure.
• Each dimension in turn is described as a set of attributes.
• Dimensions are the perspectives or entities used to keep record in an
organization.
• Each dimension can be described by a set of attributes.
Dimension modelling
• Dimension modelling is a special technique for structuring data
around business concepts.
• Unlike ER (Entity Relationship) Modelling dimension modelling
structures the numeric measures and the dimensions .
• The apex cuboid, or 0-D cuboid, refers to the case where the group-by
is empty.
• The base cuboid is the least generalized (most specific) of the
cuboids.
• The apex cuboid is the most generalized (least specific) of the
cuboids, and is often denoted as all.
Multidimensional data model components
• Summary measures eg: employment,sales
• Summary function eg: sum
• Dimension eg:sex,year,profession,state
• Dimension hierarchy eg:professional class  profession
Summary measures
• Which is the main theme of the analysis in a multidimensional model.
• A measure value is computed by aggregating the data..
• Measure can be categorized in to 3 groups.
• [Link]-measure can be simply the aggregation of the measures
of all partitions. Example:count,sum,min,max
• [Link]-it can be computed by an algebraic function with some set
of [Link] : average is obtained by sum/count.
• [Link]-there does not exist an algebraic function that can be used
to compute the aggregate [Link]:median,mode,most frequent.
Multidimensional data model-Warehouse
schema
• Schema is a logical description of the entire database.
• It includes the name and description of records of all record types
including all associated data-items and aggregates.
• Much like a database, a data warehouse also requires to maintain a
schema.
• A database uses relational model, while a data warehouse uses Star,
Snowflake, and Fact Constellation schema.
• Types
• [Link] schema
• [Link] schema
• [Link] constellation schema
Star Schema
• Each dimension in a star schema is represented with only one-dimension table.
• This dimension table contains the set of attributes.
• The following diagram shows the sales data of a company with respect to the four
dimensions, namely time, item, branch, and location.

•There is a fact table at the center. It contains the keys to each of four dimensions.
•The fact table also contains the attributes, namely dollars sold and units sold.
• A star schema is the elementary form of a dimensional model, in
which data are organized into facts and dimensions.
• A fact is an event that is counted or measured, such as a sale or log
in.
• A dimension includes reference data about the fact, such as date,
item, or customer.
• It is known as star schema because the entity-relationship diagram of
this schemas simulates a star, with points, diverge from a central
table. The center of the schema consists of a large fact table, and the
points of the star are the dimension tables.
Fact Tables
• A table in a star schema which contains facts and connected to
dimensions.
• A fact table has two types of columns: those that include fact and
those that are foreign keys to the dimension table.
Dimension Table
• Dimension table consist of columns that correspond to the attributes
of the dimension. If a dimension has not got hierarchies and levels, it
is called a flat dimension or list.
• Fact tables store data about sales while dimension tables data about
the geographic region (markets, cities), clients, products, times,
channels.
Advantages
• Easy to understand
• Easy to define hierarchies
• Reduces the number of physical joins
• Requires low maintenance
• Very simple meta data
SNOWFLAKE SCHEMA
• A schema is known as a snowflake if one or more dimension tables do
not connect directly to the fact table but must join through other
dimension tables.
• The snowflake schema is an expansion of the star schema where each
point of the star explodes into more points. It is called snowflake
schema because the diagram of snowflake schema resembles a
snowflake.
• Snowflaking is a method of normalizing the dimension tables in a
STAR schemas. When we normalize all the dimension tables entirely,
the resultant structure resembles a snowflake with the fact table in
the middle.
• Some dimension tables in the Snowflake schema are normalized.

• The normalization splits up the data into additional tables.

• Unlike Star schema, the dimensions table in a snowflake schema are


normalized. For example, the item dimension table in star schema is
normalized and split into two dimension tables, namely item and
supplier table.
• Now the item dimension table contains the attributes item_key,
item_name, type, brand, and supplier-key.

• The supplier key is linked to the supplier dimension table. The


supplier dimension table contains the attributes supplier_key and
supplier_type
• The dimension tables can be normalized to create snowflake
schemas.
• A snowflake schema consists of a single fact table and multiple
dimension tables.
• Like the Star Schema, each tuple of the fact table consists of a
(foreign) key pointing to each of the dimension tables that provide its
multidimensional coordinates.
• Dimension Tables in a star schema are denormalized, while those in a
snowflake schema are normalized.
FACT CONSTELLATION
• A Fact constellation means two or more fact tables sharing one or more
dimensions. It is also called Galaxy schema.
• There may be a need to have more than one Fact Table and these are
called Fact constellations.
• A Fact Constellation is a kind of schema where we have more than one Fact
Table sharing among them some Dimension Tables. It is also called Galaxy
Schema.
• For example, let us assume that Deccan Electronics would like to have
another Fact Table for supply and delivery. It may contain five dimensions,
or keys: time, item, delivery-agent, origin, destination along with the
numeric measure as the number of units supplied and the cost of delivery.
It can be seen that both Fact Tables can share the same item-Dimension
Table as well as time-Dimension Table.
OLAP Operations

• Roll-up
• Drill-down
• Slice and dice
• Pivot (rotate)
• Advantages of ROLAP –
• ROLAP is used for handle the large amount of data.
• ROLAP tools don’t use pre-calculated data cubes.
• Data can be stored efficiently.

• (2) A Multidimensional OLAP (MOLAP) model, i.e., a particular


purpose server that directly implements multidimensional
information and operations.
• Data is pre computed pre summarized and stored.
• HOLAP-HYBRID ONLINE ANALYTICAL PROCESSING
• Combination of rolap and molap
• Holap allows storing part of data in MOLAP store,and another part of
data in ROLAP store.
• Get quicker response for each queries.
Data Warehouse Architecture
• There are three approaches to creating a data warehouse layer:
Single-tier, two-tier, and three-tier.
• 3-tier architecture.
• Tier I is essentially the warehouse server,
• Tier 2 is the OLAP-engine for analytical processing,
• Tier 3 is a client containing reporting tools, visualization tools, data
mining tools, querying tools, etc.
• There is also the backend process which is concerned with extracting
data from multiple operational databases and from external sources;
with cleaning, transforming and integrating this data for loading into
the data warehouse server; and periodically refreshing the
warehouse.
Stages of Data Mining
Data warehousing to Datamining
• Data warehouse refers to the process of compiling and organizing data
into one common database, whereas data mining refers to the process
of extracting useful data from the databases.

• The data mining process depends on the data compiled in the data
warehousing phase to recognize meaningful patterns.
Applications of Data Warehousing
• Retail and E-trade: Retailers use data warehousing to research sales tendencies,
tune stock stages, and optimize supply chain control. It allows for knowing
customer behavior, permitting personalized marketing campaigns and product
predictions.
• Finance and Banking: Data warehousing helps risk evaluation, fraud detection, and
compliance reporting in the financial quarter. It allows for evaluating transactional
data, customer profiles, and market trends to make knowledgeable selections.
• Healthcare: Healthcare professionals use data warehousing to manipulate affected
patient data and tune medical histories. Data evaluation helps predict outbreaks,
optimize treatment effectiveness, and improve patient care.
• Manufacturing: Data warehousing helps track production processes, manage
inventory, and optimize supply chains. It enables friendly manipulation using
reading sensor statistics from production equipment.
Applications of Data Mining
• Marketing and Customer Analysis: Data mining analyzes customer behavior,
possibilities, and purchase history. This information enables growing focused
advertising campaigns, enhancing client retention, and growing sales.
• Fraud Detection: In finance and banking, data mining detects uncommon
transaction patterns, figuring out fraud cases. It helps in early detection and
prevention of fraudulent activities.
• Healthcare and Medical Research: Data mining aids in studying medical data,
patient histories, and medical data to detect disease outbreaks, verify remedy
effectiveness, and improve patient care.
• Retail and Inventory Management: Retailers utilize data mining to forecast
demand, optimize inventory, and identify developments in income. This
results in efficient supply chain control and reduced operational fees.
Datamining Issues
• Data mining is not an easy task, as the algorithms used
can get very complex and data is not always available
at one place. It needs to be integrated from various
heterogeneous data sources.

• Mining Methodology and User Interaction


• Performance Issues
• Diverse Data Types Issues
• Mining Methodology and User Interaction Issues
• Mining different kinds of knowledge in databases − Different users may be interested in different
kinds of knowledge. Therefore it is necessary for data mining to cover a broad range of knowledge
discovery task.
• Interactive mining of knowledge at multiple levels of abstraction − The data mining process
needs to be interactive because it allows users to focus the search for patterns, providing and
refining data mining requests based on the returned results.
• Incorporation of background knowledge − To guide discovery process and to express the
discovered patterns, the background knowledge can be used. Background knowledge may be
used to express the discovered patterns not only in concise terms but at multiple levels of
abstraction.
• Data mining query languages and ad hoc data mining − Data Mining Query language that allows
the user to describe ad hoc mining tasks, should be integrated with a data warehouse query
language and optimized for efficient and flexible data mining.
• Presentation and visualization of data mining results − Once the patterns are discovered it needs
to be expressed in high level languages, and visual representations. These representations should
be easily understandable.
• Handling noisy or incomplete data − The data cleaning methods are required to handle the noise
and incomplete objects while mining the data regularities. If the data cleaning methods are not
there then the accuracy of the discovered patterns will be poor.
• Pattern evaluation − The patterns discovered should be interesting because either they represent
common knowledge or lack novelty.
Performance Issues

• Efficiency and scalability of data mining algorithms − In order to


effectively extract the information from huge amount of data in databases,
data mining algorithm must be efficient and scalable.

• Parallel, distributed, and incremental mining algorithms − The factors


such as huge size of databases, wide distribution of data, and complexity of
data mining methods motivate the development of parallel and distributed
data mining algorithms. These algorithms divide the data into partitions
which is further processed in a parallel fashion. Then the results from the
partitions is merged. The incremental algorithms, update databases
without mining the data again from scratch.
Diverse Data Types Issues

• Handling of relational and complex types of data − The database


may contain complex data objects, multimedia data objects, spatial
data, temporal data etc. It is not possible for one system to mine all
these kind of data.

• Mining information from heterogeneous databases and global


information systems − The data is available at different data sources
on LAN or WAN. These data source may be structured, semi
structured or unstructured. Therefore mining the knowledge from
them adds challenges to data mining.
• Goal identification: Develop and understand the application
domain and the relevant prior knowledge and identify the
KDD process's goal from the customer perspective.
• Creating a target data set: Selecting the data set or focusing
on a set of variables or data samples on which the discovery
was made.
• Data cleaning and pre processing: Basic operations include
removing noise if appropriate, collecting the necessary
information to model or account for noise, deciding on
strategies for handling missing data fields, and accounting for
time sequence information and known changes.
• Data reduction and projection: Finding useful features to
represent the data depending on the purpose of the task. The
effective number of variables under consideration may be
reduced through dimensionality reduction methods or
conversion, or invariant representations for the data can be
found.
• Matching process objectives: KDD with step 1 a method of mining
particular. For example, summarization, classification, regression,
clustering, and others.
• Modeling and exploratory analysis and hypothesis selection:
Choosing the algorithms or data mining and selecting the method or
methods to search for data patterns. This process includes deciding
which model and parameters may be appropriate (e.g., definite data
models are different models on the real vector) and the matching of
data mining methods, particularly with the general approach of the
KDD process (for example, the end-user might be more interested in
understanding the model in its predictive capabilities).
• Data Mining: The search for patterns of interest in a particular
representational form or a set of these representations, including
classification rules or trees, regression, and clustering. The user can
significantly aid the data mining method to carry out the preceding steps
properly.
• Presentation and evaluation: Interpreting mined patterns, possibly
returning to some of the steps between steps 1 and 7 for additional
iterations. This step may also involve the visualization of the extracted
patterns and models or visualization of the data given the models drawn.
• Taking action on the discovered knowledge: Using the knowledge directly,
incorporating the knowledge in another system for further action, or
simply documenting and reporting to stakeholders. This process also
includes checking and resolving potential conflicts with previously believed
knowledge (or extracted).
• Data Mining is only a step within the overall KDD process.
There are two major Data Mining goals .
• verification of discovery.
• Verification verifies the user's hypothesis about data, while discovery
automatically finds interesting patterns.
• There are four major data mining tasks: clustering,
classification, regression, and association (summarization).
• Clustering is identifying similar groups from unstructured data.
Classification is learning rules that can be applied to new data.
• Regression is finding functions with minimal error to model
data. And the association looks for relationships between
variables. Then, the specific data mining algorithm needs to be
selected. Different algorithms like linear regression, logistic
regression, decision trees, and Naive Bayes can be selected
depending on the goal. Then patterns of interest in one or
more symbolic forms are searched.
• Finally, models are evaluated either using predictive accuracy
or understandability
Why do we need Data Mining?

• The volume of information is increasing every day that we can handle


from business transactions, scientific data, sensor data, Pictures,
videos, etc.
• So, we need a system that will be capable of extracting the essence of
information available and that can automatically generate reports,
views, or summaries of data for better decision-making.
Why is Data Mining used in business?

• Data mining is used in business to make better managerial decisions


by:
• Automatic summarization of data.
• Discovering patterns in raw data.
• Extracting the essence of information stored.
Difference between KDD and Data Mining

• Although the two terms KDD and Data Mining are heavily used
interchangeably, they refer to two related yet slightly different
concepts.
• KDD is the overall process of extracting knowledge from data, while
Data Mining is a step inside the KDD process, which deals with
identifying patterns in data.
• And Data Mining is only the application of a specific algorithm based
on the overall goal of the KDD process.
• KDD is an iterative process where evaluation measures can be
enhanced, mining can be refined, and new data can be integrated and
transformed to get different and more appropriate results.
Architecture of typical data mining system
• Data mining is a significant method where previously unknown and
potentially useful information is extracted from the vast amount of
data.
• The data mining process involves several components, and these
components constitute a data mining system architecture.
• The significant components of data mining systems are a data source,
data mining engine, data warehouse server, the pattern evaluation
module, graphical user interface, and knowledge base.
Data Source:

• The actual source of data is the Database, data warehouse, World


Wide Web (WWW), text files, and other documents.
• We need a huge amount of historical data for data mining to be
successful.
• Organizations typically store data in databases or data warehouses.
• Data warehouses may comprise one or more databases, text files
spreadsheets, or other repositories of data. Sometimes, even plain
text files or spreadsheets may contain information.
• Another primary source of data is the World Wide Web or the
internet.
Different processes:

• Before passing the data to the database or data warehouse server,


the data must be cleaned, integrated, and selected.
• As the information comes from various sources and in different
formats, it can't be used directly for the data mining procedure
because the data may not be complete and accurate.
• So, the first data requires to be cleaned and unified. More
information than needed will be collected from various data sources,
and only the data of interest will have to be selected and passed to
the server.
• Several methods may be performed on the data as part of selection,
integration, and cleaning.
Database or Data Warehouse Server:
• The database or data warehouse server consists of
the original data that is ready to be processed. Hence,
the server is cause for retrieving the relevant data
that is based on data mining as per user request.
Data Mining Engine:

• The data mining engine is a major component of any data mining


system. It contains several modules for operating data mining tasks,
including association, characterization, classification, clustering,
prediction, time-series analysis, etc.
• In other words, we can say data mining is the root of our data mining
architecture. It comprises instruments and software used to obtain
insights and knowledge from data collected from various data sources
and stored within the data warehouse.
Pattern Evaluation Module:

• The Pattern evaluation module is primarily responsible for the


measure of investigation of the pattern by using a threshold value. It
collaborates with the data mining engine to focus the search on
exciting patterns.
• This segment commonly employs stake measures that cooperate with
the data mining modules to focus the search towards fascinating
patterns.
• It might utilize a stake threshold to filter out discovered patterns. On
the other hand, the pattern evaluation module might be coordinated
with the mining module, depending on the implementation of the
data mining techniques used.
Graphical User Interface:

• The graphical user interface (GUI) module communicates between


the data mining system and the user. This module helps the user to
easily and efficiently use the system without knowing the complexity
of the process. This module cooperates with the data mining system
when the user specifies a query or a task and displays the results.
Knowledge Base:

• The knowledge base is helpful in the entire process of data mining. It


might be helpful to guide the search or evaluate the stake of the
result patterns. The knowledge base may even contain user views and
data from user experiences that might be helpful in the data mining
process. The data mining engine may receive inputs from the
knowledge base to make the result more accurate and reliable. The
pattern assessment module regularly interacts with the knowledge
base to get inputs, and also update it.
Tasks and Functionalities of Data Mining

• Data mining tasks are designed to be semi-automatic or fully


automatic and on large data sets to uncover patterns such as groups
or clusters, unusual or over the top data called anomaly detection
and dependencies such as association and sequential pattern.
• Once patterns are uncovered, they can be thought of as a summary of
the input data, and further analysis may be carried out using Machine
Learning and Predictive analytics.
Data mining activities can be divided into two
categories:
• Descriptive Data Mining: It includes certain knowledge to understand
what is happening within the data without a previous idea. The
common data features are highlighted in the data set. For example,
count, average etc.
• Predictive Data Mining: It helps developers to provide unlabeled
definitions of attributes. With previously available or historical data,
data mining can be used to make predictions about critical business
metrics based on data's linearity. For example, predicting the volume
of business next quarter based on performance in the previous
quarters over several years or judging from the findings of a patient's
medical examinations that is he suffering from any particular disease.
Data mining Functionalities
1. Class/Concept Descriptions
• A class or concept implies there is a data set or set of features that
define the class or a concept.
• A class can be a category of items on a shop floor, and a concept
could be the abstract idea on which data may be categorized like
products to be put on clearance sale and non-sale products.
• There are two concepts here, one that helps with grouping and the
other that helps in differentiating.

• Data Characterization: This refers to the summary of general


characteristics or features of the class, resulting in specific rules that
define a target class. A data analysis technique called Attribute-
oriented Induction is employed on the data set for achieving
characterization.
• Data Discrimination: Discrimination is used to separate distinct data
sets based on the disparity in attribute values. It compares features
of a class with features of one or more contrasting classes. bar
charts, curves and pie charts.
2. Mining Frequent Patterns

• One of the functions of data mining is finding data patterns. Frequent


patterns are things that are discovered to be most common in data.
Various types of frequency can be found in the dataset.

• Frequent item set: This term refers to a group of items that are commonly
found together, such as milk and sugar.
• Frequent substructure: It refers to the various types of data structures that
can be combined with an item set or subsequences, such as trees and
graphs.
• Frequent Subsequence: A regular pattern series, such as buying a phone
followed by a cover.
3. Association Analysis

• It analyses the set of items that generally occur together in a


transactional dataset. It is also known as Market Basket Analysis for
its wide use in retail sales. Two parameters are used for determining
the association rules:
• It provides which identifies the common item set in the database.
• Confidence is the conditional probability that an item occurs when
another item occurs in a transaction.
4. Classification

• Classification is a data mining technique that categorizes items in a


collection based on some predefined properties. It uses methods like
if-then, decision trees or neural networks to predict a class or
essentially classify a collection of items. A training set containing
items whose properties are known is used to train the system to
predict the category of items from an unknown collection of items.
5. Prediction

• It defines predict some unavailable data values or spending trends. An


object can be anticipated based on the attribute values of the object and
attribute values of the classes. It can be a prediction of missing numerical
values or increase or decrease trends in time-related information. There
are primarily two types of predictions in data mining: numeric and class
predictions.
• Numeric predictions are made by creating a linear regression model that is
based on historical data. Prediction of numeric values helps businesses
ramp up for a future event that might impact the business positively or
negatively.
• Class predictions are used to fill in missing class information for products
using a training data set where the class for products is known.
6. Cluster Analysis

• In image processing, pattern recognition and bioinformatics,


clustering is a popular data mining functionality. It is similar to
classification, but the classes are not predefined. Data attributes
represent the classes. Similar data are grouped together, with the
difference being that a class label is not known. Clustering algorithms
group data based on similar features and dissimilarities.
7. Outlier Analysis
• Outlier analysis is important to understand the quality of data. If
there are too many outliers, you cannot trust the data or draw
patterns. An outlier analysis determines if there is something out of
turn in the data and whether it indicates a situation that a business
needs to consider and take measures to mitigate. An outlier analysis
of the data that cannot be grouped into any classes by the algorithms
is pulled up.
8. Evolution and Deviation Analysis

• Evolution Analysis pertains to the study of data sets that change over
time. Evolution analysis models are designed to capture evolutionary
trends in data helping to characterize, classify, cluster or discriminate
time-related data.
9. Correlation Analysis

• Correlation is a mathematical technique for determining whether and


how strongly two attributes is related to one another. It refers to the
various types of data structures, such as trees and graphs, that can be
combined with an item set or subsequence. It determines how well
two numerically measured continuous variables are linked.
Researchers can use this type of analysis to see if there are any
possible correlations between variables in their study.

PREVIOUS YEAR QUESTIONS
Q.3
Q.4

Q.5

Q.6
Q.7

Q.8
Q.9

Q.10

Q.11
Q.12

Q.13

You might also like