Predictive Analysis
Predictive Analysis
Unit:1
Introduction
Slide - 1
Unit:1
What is Analytics?
Slide - 2
Analytics
Slide - 3
1
2/10/2026
Why is Analytics?
Slide - 4
Business Analytics
Business) Analytics is the use of:
• data,
• information technology,
• statistical analysis,
• quantitative methods, and
• mathematical or computer-based models
to help managers gain improved insight about their business
operations and make better, fact-based decisions.
Slide - 5
Business Analytics
Business) Analytics is :
• “a process of transforming data into actions through
analysis and insights in the context of organizational
decision making and problem solving.”
Slide - 6
2
2/10/2026
Business Analytics
Business) Analytics is :
• “a process by which business use statistical methods
and technologies for analyzing historical data in order to
gain insight and improve strategic decision-making”
• Is the combination of skills technologies and practices
• Used to examine an organization’s data and performance as
a way to gain insights and make data driven decisions in
future
Slide - 7
Uses/ Applications
• To start a business
• Business growth
• Analyse data to get desired best information
• Create business strategies (Prescriptive analysis)
• Stay updated to business trends using business tools
• Helps to analyse mistakes
Slide - 8
Examples of Applications (1 of 2)
• Pricing
– setting prices for consumer and industrial goods,
government contracts, and maintenance contracts
• Customer segmentation
– identifying and targeting key customer groups in retail,
insurance, and credit card industries
• Merchandising
– determining brands to buy, quantities, and allocations
• Location
– finding the best location for bank branches and ATMs, or
where to service industrial equipment Slide - 9
3
2/10/2026
Examples of Applications (2 of 2)
Impacts of Analytics
• Benefits
– …reduced costs, better risk management, faster decisions,
better productivity and enhanced bottom-line performance
such as profitability and customer satisfaction.
• Challenges
– …lack of understanding of how to use analytics, competing
business priorities, insufficient analytical skills, difficulty in
getting good data and sharing information, and not
understanding the benefits versus perceived costs of
analytics studies.
Slide - 11
Impacts of Analytics
• Changes
– Business analytics is changing how managers make
decisions.
– To thrive in today’s business world, organizations must:
– continually innovate to differentiate themselves from
competitors,
– seek ways to grow revenue and market share,
– reduce costs,
– retain existing customers and acquire new ones,
– and become faster and leaner.
Slide - 12
4
2/10/2026
• Analytic Foundations
– Business Intelligence (BI)
– facilitated the collection, management, analysis, and
reporting of data.
– A term coined by IBM researcher Hans Perter Luhn in
1958
– Information Systems (IS)
– Statistics
– Operations Research/Management Science (OR/M
S)
– Was born from for military operations prior to and during
World War II. Later on these mathematical tools and
techniques applied successfully to problems in business and
Slide - 13
industry.
5
2/10/2026
Slide - 16
Slide - 17
Slide - 18
6
2/10/2026
Slide - 20
Slide - 21
7
2/10/2026
Slide - 22
Slide - 23
Slide - 24
8
2/10/2026
Slide - 25
Slide - 27
9
2/10/2026
Types of Analytics
Slide - 28
Slide - 29
Slide - 30
10
2/10/2026
11
2/10/2026
Slide - 34
Types of Analytics
Slide - 35
Slide - 36
12
2/10/2026
SWOC Analysis
SWOC Analysis
• Strengths-
• What are the core competencies of the business?
• Differentiating factors in business
• Order Winners and Order Qualifiers
• Weakness-
• What are the areas of improvement of business?
• Identification of customer pain points and turning them
into opportunities
Slide - 38
SWOC Analysis
• Opportunities-
• What are the areas to be focused on to increase revenue?
• Study of new market trends
• Environment scanning
• Challenges-
• What are the environmental factors that can hamper
business?
• Risk analysis (Consumer choice, competitive pricing,
chaning technology)
Slide - 39
13
2/10/2026
Reports in Analytics
Slide - 40
Reports in Analytics
Reports in Analytics
Slide - 42
14
2/10/2026
Slide - 43
Slide - 45
15
2/10/2026
Slide - 46
Slide - 47
• 1. Quantitative data
• It answers key questions such as “how many, “how much” and
“how often”.
• Quantitative data can be expressed as a number or can be
quantified. Simply put, it can be measured by numerical variables.
Tangibles
• Quantitative data are easily amenable to statistical manipulation
and can be represented by a wide variety of statistical types of
graphs and charts such as line, bar graph, scatter plot, and etc.
• Examples of quantitative data:
• Scores on tests and exams e.g. 85, 67, 90 and etc.
• The weight of a person or a subject.
• Your shoe size.
• The temperature in a room. Slide - 48
16
2/10/2026
• 2. Qualitative data
• Qualitative data can’t be expressed as a number and can’t be
measured. Qualitative data consist of words, pictures, and
symbols, not numbers. Intangibles
• Qualitative data is also called categorical data because the
information can be sorted by category, not by number.
• Qualitative data can answer questions such as “how this has
happened” or and “why this has happened”.
• Examples of qualitative data:
• Colors e.g. the color of the sea
• Your favorite holiday destination such as Hawaii, New Zealand and etc.
Slide - 50
• 3. Nominal data
• Nominal data is used just for labelling variables, without any
type of quantitative value. The name ‘nominal’ comes from the
Latin word “nomen” which means ‘name’.
• The nominal data just name a thing without applying it to order.
Actually, the nominal data could just be called “labels.”
• Examples of nominal data:
• Gender (Women, Men)
Slide - 51
17
2/10/2026
• 4. Ordinal data
• Shows where a number is in order. This is the crucial difference from
nominal types of data. Ordinal data is data which is placed into some
kind of order by their position on a scale.
• Ordinal data may indicate superiority. However, you cannot do
arithmetic with ordinal numbers because they only show sequence.
• Ordinal variables are considered as “in between” qualitative and
quantitative variables. In other words, the ordinal data is qualitative
data for which the values are ordered.
• Examples of ordinal data:
• The first, second and third person in a competition.
• When a company asks a customer to rate the sales experience on a scale of 1-10.
Slide - 53
• 5. Discrete data
• Discrete data is a count that involves only integers. The discrete
values cannot be subdivided into parts. For example, the number of
children in a class is discrete data. You can count whole individuals.
You can’t count 1.5 kids.
• Discrete data can take only certain values. The data variables cannot
be divided into smaller parts. It has a limited number of possible
values e.g. days of the month.
• Examples of discrete data:
• The number of students in a class.
18
2/10/2026
• [Link] data
• Continuous data is information that could be meaningfully divided into
finer levels. It can be measured on a scale or continuum and can
have almost any numeric value.
• The continuous variables can take any value between two numbers.
For example, between 50 and 72 inches, there are literally millions of
possible heights: 52.04762 inches, 69.948376 inches and etc.
• Examples of continuous data:
• The amount of time required to complete a project.
Slide - 55
Slide - 56
Organization/Sources of Data
• Data organization is the practice of categorizing and classifying
data to make it more usable. Similar to a file folder, where we
keep important documents, you’ll need to arrange your data in
the most logical and orderly fashion, so you — and anyone else
who accesses it — can easily find what they’re looking for.
• It is also needed that unauthorized person should not have
access to our data.
Slide - 57
19
2/10/2026
Organization/Sources of Data
DATA IS BEING COLLECTED
• Data includes information produced by humans and devices.
• Device-driven data is largely clean and organized,
• But of far greater interest is human-driven data that exist in
various formats and need more exquisite tools for proper
processing and management.
Slide - 58
Organization/Sources of Data
The data collection is focused on the following types of data:
• Network data. This type of data is gathered on all kinds of
networks, including social media, information and technological
networks, the Internet and mobile networks, etc.
• Real-time data. They are produced on online streaming media,
such as YouTube, Twitch, Skype, or Netflix.
• Transactional data. They are gathered when a user makes an
online purchase (information on the product, time of purchase,
payment methods, etc.)
• Geographic data. Location data of everything, humans,
vehicles, building, natural reserves, and other objects are
continuously supplied with satellites.
Slide - 59
Organization/Sources of Data
• Natural language data. These data are gathered mostly from
voice searches that can be made on different devices accessing
the Internet.
• Time series data. This type of data is related to the observation
of trends and phenomena taking place at this very moment and
over a period of time, for instance, global temperatures, mortality
rates, pollution levels, etc.
• Linked data. They are based on HTTP, RDF, SPARQL, and
URIs web technologies and meant to enable semantic
connections between various databases so that computers could
read and perform semantic queries correctly.
Slide - 60
20
2/10/2026
Organization/Sources of Data
Explanation
• RDF. The Resource Description Framework (RDF) is a
general framework for representing interconnected data on the
web. RDF statements are used for describing and exchanging
metadata, which enables standardized exchange of data
based on relationships
• HTTP. The Hypertext Transfer Protocol (HTTP) is the
foundation of the World Wide Web, and is used to load
webpages using hypertext links. HTTP is an application layer
protocol designed to transfer information between networked
devices and runs on top of other layers of the network protocol
stack
Slide - 61
Organization/Sources of Data
The data collection is focused on the following types of data:
• SPARQL. SPARQL Protocol and RDF Query Language,
enables users to query information from databases or any data
source that can be mapped to RDF. The SPARQL standard is
designed and endorsed by the W3C (World Wide Web
Consortium) and helps users and developers focus on what they
would like to know instead of how a database is organized.
• URIs. A Uniform Resource Identifier, is a unique sequence of
characters that identifies an abstract or physical resource, such
as resources on a webpage, mail address, phone number,
books, real-world objects such as people and places, concepts.
Slide - 62
Organization/Sources of Data
The data collection is focused on the following types of data:
• Asking for it. the majority of firms prefer asking users directly to
share their personal information. They give these data when
creating website accounts or buying online. The minimum
information to be collected includes a username and an email
address, but some profiles require more details.
• Cookies and Web Beacons. Cookies and web beacons are two
widely used methods to gather the data on users, namely, what
web pages they visit and when. They provide basic statistics
about how a website is used. Cookies and web beacons in no
way compromise your privacy but just serve to personalize your
experience with one or another web source.
Slide - 63
21
2/10/2026
Organization/Sources of Data
The data collection is focused on the following types of data:
• Email tracking. Email trackers are meant to give more
information on the user actions in the mailbox. In particular, an
email tracker allows detecting when an email was opened. Both
Google and Yahoo use this method to learn their users’
behavioural patterns and provide personalized advertising.
Slide - 64
Slide - 65
22
2/10/2026
Slide - 69
23
2/10/2026
Slide - 72
24
2/10/2026
Slide - 73
Slide - 74
Slide - 75
25
2/10/2026
Slide - 76
Slide - 77
Slide - 78
26
2/10/2026
Slide - 79
Slide - 80
Slide - 81
27
2/10/2026
Slide - 82
Big Data
Slide - 84
28
2/10/2026
Slide - 85
Slide - 86
Slide - 87
29
2/10/2026
Slide - 88
Slide - 89
30
2/10/2026
Slide - 91
Slide - 92
Slide - 93
31
2/10/2026
Slide - 94
Slide - 95
32
2/10/2026
Slide - 99
33
2/10/2026
Slide -
100
Slide -
107
Slide -
108
34
2/10/2026
[Link]
Slide -
109
Slide -
110
Slide -
111
35
2/10/2026
Slide -
112
Slide -
113
Slide -
114
36
2/10/2026
Slide -
115
37
2/10/2026
Slide -
118
Slide -
120
38
2/10/2026
Slide -
121
1. Recognizing a problem
2. Defining the problem
3. Structuring the problem
4. Analyzing the problem
5. Interpreting results and making a decision
6. Implementing the solution
Slide -
122
Recognizing a Problem
Problems exist when there is a gap between what
is happening and what we think should be
happening.
• For example, costs are too high compared with
competitors.
Slide -
123
39
2/10/2026
Slide -
124
Slide -
125
Slide -
126
40
2/10/2026
Slide -
127
Slide -
128
Slide -
129
41
2/10/2026
Slide -
130
Impact Cycle
Slide -
131
Slide -
132
42
2/10/2026
Slide -
133
Slide -
134
How it works?
• Data Collection: Gathers large volumes of
historical and real-time data.
• Pattern Identification: Employs statistical
models, machine learning (like decision trees,
neural networks), and AI to find relationships and
trends.
• Forecasting: Applies these patterns to predict
future events or probabilities, such as customer
actions, sales, or potential risks
Slide -
135
43
2/10/2026
Key Applications
• Gathers Marketing: Personalizing offers,
predicting customer lifetime value.
• Finance: Detecting credit card fraud, assessing
loan risk, forecasting revenue.
• Operations: Optimizing staffing, predicting
maintenance needs.
• Healthcare: Predicting patient outcomes or
disease outbreaks
Slide -
136
Benefits
• Enables proactive, data-driven decisions.
• Identifies risks and opportunities before they fully
develop.
• Improves efficiency, customer satisfaction, and
profitability.
Slide -
137
Techniques
• In general, there are two types of predictive
analytics models:
• classification and
• regression models and forecasting models.
Slide -
138
44
2/10/2026
Techniques
• Classification models attempt to put data objects
(such as customers or potential outcomes) into one
category or another. For instance, if a retailer has a
lot of data on different types of customers, they
may try to predict what types of customers will be
receptive to market emails.
• Regression models & Forecasting Models try to
predict continuous data, such as how much
revenue that customer will generate during their
relationship with the company.
Slide -
139
Techniques
• Predictive analytics tends to be performed with
three main types of techniques:
• Regression analysis & Forecasting
• Decision tree
• Neural Network
Slide -
140
Slide -
141
45
2/10/2026
Slide -
143
Slide -
144
46
2/10/2026
Slide -
145
Slide -
146
Slide -
147
47
2/10/2026
Techniques
• Decision trees
• Decision trees are classification models that place data into
different categories based on distinct variables.
• The method is best used when trying to understand an
individual's decisions.
• The model looks like a tree, with each branch representing
a potential choice, with the leaf of the branch representing
the result of the decision.
• Decision trees are typically easy to understand and work
well when a dataset has several missing variables.
Slide -
148
Slide -
150
48
2/10/2026
Slide -
151
Slide -
152
49
2/10/2026
Slide -
154
• Advantages
• Easy to Understand: Decision Trees are visual which makes it
easy to follow the decision-making process.
• Versatility: Can be used for both classification and regression
problems.
• No Need for Feature Scaling: Unlike many machine learning
models, it don’t require us to scale or normalize our data.
• Handles Non-linear Relationships: It capture complex, non-
linear relationships between features and outcomes effectively.
• Interpretability: The tree structure is easy to interpret helps in
allowing users to understand the reasoning behind each decision.
• Handles Missing Data: It can handle missing values by using
strategies like assigning the most common value or ignoring
missing data during splits.
Slide -
155
• Disadvantages
• Overfitting: They can overfit the training data if they are too deep
which means they memorize the data instead of learning general
patterns. This leads to poor performance on unseen data.
• Instability: It can be unstable which means that small changes in
the data may lead to significant differences in the tree structure and
predictions.
• Bias towards Features with Many Categories: It can become
biased toward features with many distinct values which focuses too
much on them and potentially missing other important features
which can reduce prediction accuracy.
• Difficulty in Capturing Complex Interactions: Decision Trees may
struggle to capture complex interactions between features which
helps in making them less effective for certain types of data.
• Computationally Expensive for Large Datasets: For large
datasets, building and pruning a Decision Tree can be
computationally intensive, especially as the tree depth increases.
Slide -
156
50
2/10/2026
• Applications
• Loan Approval in Banking: Banks use Decision Trees to assess
whether a loan application should be approved. The decision is
based on factors like credit score, income, employment status and
loan history. This helps predict approval or rejection helps in
enabling quick and reliable decisions.
• Medical Diagnosis: In healthcare they assist in diagnosing
diseases. For example, they can predict whether a patient has
diabetes based on clinical data like glucose levels, BMI and blood
pressure. This helps classify patients into diabetic or non-diabetic
categories, supporting early diagnosis and treatment.
• Predicting Exam Results in Education: Educational institutions
use to predict whether a student will pass or fail based on factors
like attendance, study time and past grades. This helps teachers
identify at-risk students and offer targeted support.
Slide -
157
• Applications …
• Customer Churn Prediction: Companies use Decision Trees to
predict whether a customer will leave or stay based on behavior
patterns, purchase history, and interactions. This allows businesses
to take proactive steps to retain customers.
• Fraud Detection: In finance, Decision Trees are used to detect
fraudulent activities, such as credit card fraud. By analyzing past
transaction data and patterns, Decision Trees can identify
suspicious activities and flag them for further investigation.
Slide -
158
Slide -
159
51
2/10/2026
We choose the product with the highest EV. The option with
the lower EV is shown with two lines cutting across it. We’ve
“rolled back” the tree from 6 uncertain outcomes to the one
branch at the decision point with the highest expected value.
We’re pruning off poor uses of capital that lead to undesirable
outcomes.
Slide -
160
Neural networks
• Neural networks are machine learning methods that are
useful in predictive analytics when modeling very complex
relationships.
• Essentially, they are powerhouse pattern recognition
engines.
• Neural networks are best used to determine nonlinear
relationships in datasets, especially when no known
mathematical formula exists to analyze the data. Neural
networks can be used to validate the results of decision
trees and regression models.
Slide -
161
Neural networks
• Neural networks are machine learning models that mimic
the complex functions of the human brain. These models
consist of interconnected nodes or neurons that process
data, learn patterns and enable tasks such as pattern
recognition and decision-making.
Slide -
162
52
2/10/2026
Neural networks
• Neural networks are capable of
learning and identifying patterns
directly from data without pre-defined
rules. These networks are built from
several key components:
• Neurons: The basic units that receive inputs, each
neuron is governed by a threshold and an activation
function.
Neural networks
• Process
• Input Computation: Data is fed into the network.
• Output Generation: Based on the current parameters, the
network generates an output.
• Iterative Refinement: The network refines its output by
adjusting weights and biases, gradually improving its
performance on diverse tasks.
Slide -
164
Neural networks
• Structure
53
2/10/2026
Slide -
166
Applications
• Pattern Recognition: Identifying complex patterns in
customer behavior
• Anomaly Detection: Finding unusual patterns that might
indicate fraud or errors
• Predictive Modeling: Forecasting future trends with high
accuracy
• Natural Language Processing: Understanding and
processing human language
Slide -
167
Challenges
• Resource intensive
• Results are often hard to interpret
Slide -
168
54
2/10/2026
Data Cleaning/Cleansing
Slide -
169
Key Aspects
• Handling Missing Data: Imputing missing values, or
removing rows/columns with substantial missing data.
• Removing Duplicates: Identifying and eliminating
redundant or identical records.
• Standardization/Formatting: Fixing structural errors, such
as inconsistent date formats, typos, or inconsistent naming
conventions.
• Outlier Handling: Removing or correcting data points that
are drastically different from others, which may skew
analysis
Slide -
171
55
2/10/2026
Slide -
172
Slide -
173
Slide -
174
56
2/10/2026
Slide -
177
57
2/10/2026
Slide -
178
Slide -
179
Slide -
180
58