0% found this document useful (0 votes)
16 views58 pages

Predictive Analysis

The document provides an overview of business analytics, defining it as the use of data, statistical analysis, and technology to enhance decision-making and operational insights. It discusses various applications of analytics in business, such as pricing, customer segmentation, and supply chain design, while also highlighting the benefits and challenges associated with its implementation. Additionally, it distinguishes between business analysis and business analytics, outlining their scopes and similarities.

Uploaded by

mahajanmegha1305
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
16 views58 pages

Predictive Analysis

The document provides an overview of business analytics, defining it as the use of data, statistical analysis, and technology to enhance decision-making and operational insights. It discusses various applications of analytics in business, such as pricing, customer segmentation, and supply chain design, while also highlighting the benefits and challenges associated with its implementation. Additionally, it distinguishes between business analysis and business analytics, outlining their scopes and similarities.

Uploaded by

mahajanmegha1305
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

2/10/2026

Unit:1

Introduction

Slide - 1

Unit:1

What is Analytics?

Slide - 2

Analytics

Slide - 3

1
2/10/2026

Why is Analytics?

Slide - 4

Business Analytics
Business) Analytics is the use of:
• data,
• information technology,
• statistical analysis,
• quantitative methods, and
• mathematical or computer-based models
to help managers gain improved insight about their business
operations and make better, fact-based decisions.

Slide - 5

Business Analytics
Business) Analytics is :
• “a process of transforming data into actions through
analysis and insights in the context of organizational
decision making and problem solving.”

Slide - 6

2
2/10/2026

Business Analytics
Business) Analytics is :
• “a process by which business use statistical methods
and technologies for analyzing historical data in order to
gain insight and improve strategic decision-making”
• Is the combination of skills technologies and practices
• Used to examine an organization’s data and performance as
a way to gain insights and make data driven decisions in
future

Slide - 7

Uses/ Applications

• To start a business
• Business growth
• Analyse data to get desired best information
• Create business strategies (Prescriptive analysis)
• Stay updated to business trends using business tools
• Helps to analyse mistakes

Slide - 8

Examples of Applications (1 of 2)
• Pricing
– setting prices for consumer and industrial goods,
government contracts, and maintenance contracts
• Customer segmentation
– identifying and targeting key customer groups in retail,
insurance, and credit card industries
• Merchandising
– determining brands to buy, quantities, and allocations
• Location
– finding the best location for bank branches and ATMs, or
where to service industrial equipment Slide - 9

3
2/10/2026

Examples of Applications (2 of 2)

• Supply Chain Design


– determining the best sourcing and transportation
options and finding the best delivery routes
• Staffing
– ensuring appropriate staffing levels and capabilities,
and hiring the right people
• Health care
– scheduling operating rooms to improve utilization,
improving patient flow and waiting times, purchasing
supplies, and predicting health risk factors
Slide - 10

Impacts of Analytics

• Benefits
– …reduced costs, better risk management, faster decisions,
better productivity and enhanced bottom-line performance
such as profitability and customer satisfaction.
• Challenges
– …lack of understanding of how to use analytics, competing
business priorities, insufficient analytical skills, difficulty in
getting good data and sharing information, and not
understanding the benefits versus perceived costs of
analytics studies.

Slide - 11

Impacts of Analytics

• Changes
– Business analytics is changing how managers make
decisions.
– To thrive in today’s business world, organizations must:
– continually innovate to differentiate themselves from
competitors,
– seek ways to grow revenue and market share,
– reduce costs,
– retain existing customers and acquire new ones,
– and become faster and leaner.

Slide - 12

4
2/10/2026

Evolution of Business Analytics (1 of 5)

• Analytic Foundations
– Business Intelligence (BI)
– facilitated the collection, management, analysis, and
reporting of data.
– A term coined by IBM researcher Hans Perter Luhn in
1958
– Information Systems (IS)
– Statistics
– Operations Research/Management Science (OR/M
S)
– Was born from for military operations prior to and during
World War II. Later on these mathematical tools and
techniques applied successfully to problems in business and
Slide - 13
industry.

Evolution of Business Analytics (2 of 5)

• Modern Business Analytic


– Data mining
– databases using a variety of statistical and analytical
tools. Uses standard statistical tools as well as more
advanced ones
– Simulation and risk analysis
– relies on spreadsheet models and statistical
analysis to examine the impacts of uncertainty in
the estimates and their potential interaction with
one another on the output variable of interest. Allow
to manipulate data to perform what-if analysis—how
specific combinations of inputs that reflect key
assumptions will affect model outputs.
Slide - 14

Evolution of Business Analytics (3 of 5)

• Modern Business Analytic


– Decision Support Systems (DSS) includes 3
components
– 1. Data management. The data management component
includes databases for storing data and allows the user to
input, retrieve, update, and manipulate data.
– 2. Model management. The model management
component consists of various statistical tools and
management science models and allows the user to
easily build, manipulate, analyze, and solve models.
– 3. Communication system. The communication system
component provides the interface necessary for the user
to interact with the data and model management
components. Slide - 15

5
2/10/2026

Evolution of Business Analytics (4 of 5)

• Modern Business Analytic


– Decision Support Systems (DSS) includes 3
components
– DSSs have been used for many applications, including
pension fund management, portfolio management, work-
shift scheduling, global manufacturing and facility
location, advertising-budget allocation, media planning,
distribution planning, airline operations planning,
inventory control, library management, classroom
assignment, nurse scheduling, blood distribution, water
pollution control, ski-area design, police-beat design, and
energy planning.

Slide - 16

Evolution of Business Analytics (5 of 5)

• Modern Business Analytic


– Visualization
– Visualizing data and results of analyses provide a way of
easily communicating data at all levels of a business and
can reveal surprising patterns and relationships.
– The Cincinnati Zoo, has used this on an iPad to display hourly,
daily, and monthly reports of attendance, food and retail location
revenues and sales, and other metrics for prediction and
marketing strategies.
– UPS uses telematics to capture vehicle data and display them to
help make decisions to improve efficiency and performance.

Slide - 17

Business Analysis vs Business Analytics


• Business analytics:
• is data and reporting—examining past business
performance and forecasting future business performance.
• Business analysis:
• focuses on functions and processes—determining business
requirements and suggesting solutions.

Slide - 18

6
2/10/2026

Business Analysis vs Business Analytics


• Business analysis: Scope
• Business analysis is the practice of assisting firms in resolving
their technical difficulties by understanding, defining, and
solving those issues.
• Scope:
– Company analysis- Requirement ascertain, strategic
decisions. Initiatives
– Requirement planning and management
– Requirement elicitation- collecting needs from relevant
members of the project team.
– Requirement analysis and documentation
– Requirement communication
– Solution assessment and validation
Slide - 19

Business Analysis vs Business Analytics


• Business analytics: Scope
• It is a process of collecting, evaluating, and drawing valuable
outcomes from the enormous amount of data available.
• Scope:
– Finance
– Marketing
– HR
– CRM
– Manufacturing
– Banking and Credit Cards

Slide - 20

Business Analysis vs Business Analytics

Slide - 21

7
2/10/2026

Business Analysis vs Business Analytics


• Similarities
• Examine and enhance businesses
• Determine solutions to issues
• Establish things based on the requirements

Slide - 22

Business Analysis vs Business Analytics


• Business analysis is a practice of identifying business
requirements and figuring out solutions to specific business
problems.
• This has a heavy overlap with the analysis of business needs to
function normally and to enhance how they function.
• Sometimes, the solutions include a system’s development
feature. It can also incorporate business change, process
enhancement or strategic planning, and policy improvement.

Slide - 23

Business Analysis vs Business Analytics


• Business analytics is all about the group of tools, techniques,
and skills that help the investigation of previous business
performance.
• It also aids to gain insights into future performance. In general,
business analytics aims mostly at data and statistical analysis.

Slide - 24

8
2/10/2026

A Visual Perspective of Business


Analytics
Focuses on better
understanding characteristics
and patterns
Visualizing data and results of among variables in large
analyses provide a way of databases using a variety of
easily communicating data at statistical and analytical tools..
all levels of a business and
can reveal surprising Facilitates the
patterns and relationships. collection,
management,
analysis, and
reporting of data,

Relies on spreadsheet models


and statistical analysis used to assess the sensitivity
to examine of optimization models to
the impacts of uncertainty in changes in data inputs and
the estimates and their provide better insight
potential interaction with for making good decisions.
one another on the output
variable of interest.

Slide - 25

Software Support and Spreadsheet


Technology
• Commercial software
– IBM Cognos Express
– SAS Analytics
– Tableau
• Spreadsheets
– Widely used
– Effective for manipulating data and developing and
solving models
– Support powerful commercial add-ons
– Facilitate communication of results
Slide - 26

Some illustrative applications


• Analysing supply chains (Hewlett-Packard)
• Inventory optimization(Procter & Gamble)
• Selection of internal projects (Lockheed Martin)
• Performance measurement and evaluation(American
Red Cross)

Slide - 27

9
2/10/2026

Types of Analytics

Slide - 28

Descriptive, Predictive, Prescriptive


and Diagnostic Analytics
What Happened?

• Descriptive analytics: the use of data to understand


past and current business performance and make
informed decisions.
• Takes a critical look at your current state of business
• Helps to know business's strength and weaknesses
• Based upon these insights, organization can develop
strategies for business improvement.
• Data mining and data mining techniques are
employed

Slide - 29

Descriptive, Predictive, Prescriptive


and Diagnostic Analytics
• Descriptive analytics Examples:
• Business reports of revenue and expenses, cash flow, accounts
receivable and accounts payable, inventory and production.
• Financial metric and other business KPIs. These include metrics
that assess the health and value of a business, such as the price to
earnings ratio, current ratio and return on invested capital.
• Social media engagement: Descriptive analytics generates metrics
that help determine the return on social media initiatives, such as
growth in followers, engagement rates and revenue attributable to
specific social media platforms.
• Surveys: Descriptive analytics produces summaries of internal and
external survey results, such as a net promoter score.

Slide - 30

10
2/10/2026

Descriptive, Predictive, Prescriptive


and Diagnostic Analytics
What might happen in the future?”
• Predictive analytics: predict the future by examining
historical data, detecting patterns or relationships in these
data, and then extrapolating these relationships forward
in time.
• Gives a look at what is going to happen
• Gives detailed reports
• Examples
• Finance: Forecasting Future Cash Flow
• Entertainment & Hospitality: Determining Staffing Needs
• Marketing: Behavioral Targeting
• Manufacturing: Preventing Malfunction
• Health Care: Early Detection of Allergic Reactions Slide - 31

Descriptive, Predictive, Prescriptive


and Diagnostic Analytics “What
next?”
should we do

• Prescriptive analytics: identify the best alternatives to


minimize or maximize some objective
• Helps to make models to make accurate predictions
• Make real time changes that will give the best possible
outcomes.
• Provides recommendations based upon expected results
• Recommendation engine
• Examples:
• Venture Capital: Investment Decisions
• Sales: Lead Scoring
• Content Curation: Algorithmic Recommendations
Slide - 32
• Banking: Fraud Detection

Descriptive, Predictive, Prescriptive


and Diagnostic Analytics“Why did this happen?”
• Diagnostic analytics: Why and How?
• Give more critical look into the past and present of
business and
• Helps to understand why these issues have occurred and
reasons behind
• Tells why such things happened and
• Based upon findings. Organization cav strategize to
improve business
• Techniques used – Data mining, Data discovery, drill
down and correlations.
Slide - 33

11
2/10/2026

Descriptive, Predictive, Prescriptive


and Diagnostic Analytics
• Diagnostic analytics Examples:
• Examining Market Demand
• Correlation vs. Causation
• Explaining Customer Behavior
• Identifying Technology Issues
• Improving Company Culture

Slide - 34

Types of Analytics

Slide - 35

Descriptive, Predictive, Prescriptive


and Diagnostic Analytics

Slide - 36

12
2/10/2026

SWOC Analysis

• An effective business today is always built on a solid rock


foundation of introspection and analysis.
• Abide by rule of “survival of fittest.”
• Companies that can identify and adapt new opportunities
succeed.
• SWOC analysis combines data from internal and external
factors for better strategic decision making.
• Allows to identify Strength and Weakness existing in
corporate structure, translate them into Opportunities
while keeping an eye out for Challenges to business.
Slide - 37

SWOC Analysis
• Strengths-
• What are the core competencies of the business?
• Differentiating factors in business
• Order Winners and Order Qualifiers
• Weakness-
• What are the areas of improvement of business?
• Identification of customer pain points and turning them
into opportunities

Slide - 38

SWOC Analysis
• Opportunities-
• What are the areas to be focused on to increase revenue?
• Study of new market trends
• Environment scanning
• Challenges-
• What are the environmental factors that can hamper
business?
• Risk analysis (Consumer choice, competitive pricing,
chaning technology)
Slide - 39

13
2/10/2026

Reports in Analytics

• Once data is collected it is organized using graphs,


charts and tables. Analytics is the process of organizing
this data and analyzing it in order to gain valuable
insights of business.
• Approach is called Data Visualization

Slide - 40

Reports in Analytics

• Five key elements of custom reports are:


• User- A segmentation level option. Broadest level; each
individual person is a unique user.
• Sessions- Segmentation level options. Most users make
multiple visits, which is known as a “session”.
• Hits- Segmentation level option. Within each session
there are hits
• Dimensions- Every report comprises of dimensions and
metrics. Dimension describe characteristics (geographic
location or browser.
• Metrics- are quantitative measurements ( sessions or
conversion rate Slide - 41

Reports in Analytics

• Three different types of reports are:


• Explorer- This is a basic report. Includes line graph, data
table
• Flat Table- Most common type of custom report. It is
essentially sortable data table.
• Map overlay- Simply a global map with colours to
indicate engagements, traffic etc)

Slide - 42

14
2/10/2026

Example : Retail Markdown Decisions

• Most department stores clear seasonal inventory by reducing


prices.
• Key question: When to reduce the price and by how much to
maximize revenue?

Slide - 43

Example : Retail Markdown Decisions

• For example, suppose that a store has 100 garments of a


certain style that go on sale from April 1 and wants to sell all of
them by the end of June. (Seasonal Sale)
• Over each week of the 12-week selling season, they can make
a decision to discount the price.
• They face two decisions:
– When to reduce the price?
– By how much?
• This results in 24 decisions to make. For a major national
chain that may carry thousands of products, this can easily
result in millions of decisions that store managers have to
make. Slide - 44

Example : Retail Markdown Decisions

• Potential applications of analytics in retail:


– Descriptive analytics: examine historical data for similar
products (prices, units sold, advertising …)
– Predictive analytics: predict sales based on price
– Prescriptive analytics: find the best sets of pricing and
advertising to maximize sales revenue

Slide - 45

15
2/10/2026

Data for Business Analytics

• Data: numbers or textual data that are collected


through some type of measurement process
• Information: result of analyzing data; that is,
extracting meaning from data to support
evaluation and decision making

Slide - 46

Types of Data for Business Analytics

Slide - 47

• Qualitative vs. Quantitative Data

• 1. Quantitative data
• It answers key questions such as “how many, “how much” and
“how often”.
• Quantitative data can be expressed as a number or can be
quantified. Simply put, it can be measured by numerical variables.
Tangibles
• Quantitative data are easily amenable to statistical manipulation
and can be represented by a wide variety of statistical types of
graphs and charts such as line, bar graph, scatter plot, and etc.
• Examples of quantitative data:
• Scores on tests and exams e.g. 85, 67, 90 and etc.
• The weight of a person or a subject.
• Your shoe size.
• The temperature in a room. Slide - 48

16
2/10/2026

• Qualitative vs. Quantitative Data

• 2. Qualitative data
• Qualitative data can’t be expressed as a number and can’t be
measured. Qualitative data consist of words, pictures, and
symbols, not numbers. Intangibles
• Qualitative data is also called categorical data because the
information can be sorted by category, not by number.
• Qualitative data can answer questions such as “how this has
happened” or and “why this has happened”.
• Examples of qualitative data:
• Colors e.g. the color of the sea

• Your favorite holiday destination such as Hawaii, New Zealand and etc.

• Names as Ram, Shyam, Sita, Gita.

• Ethnicity such as American Indian, Asian, etc. Slide - 49

• Qualitative vs. Quantitative Data

Slide - 50

• Nominal vs. Ordinal Data

• 3. Nominal data
• Nominal data is used just for labelling variables, without any
type of quantitative value. The name ‘nominal’ comes from the
Latin word “nomen” which means ‘name’.
• The nominal data just name a thing without applying it to order.
Actually, the nominal data could just be called “labels.”
• Examples of nominal data:
• Gender (Women, Men)

• Hair color (Blonde, Brown, Brunette, Red, etc.)

• Marital status (Married, Single, Widowed)

• Ethnicity (European,Latin, Asian)

Slide - 51

17
2/10/2026

• Nominal vs. Ordinal Data

• 4. Ordinal data
• Shows where a number is in order. This is the crucial difference from
nominal types of data. Ordinal data is data which is placed into some
kind of order by their position on a scale.
• Ordinal data may indicate superiority. However, you cannot do
arithmetic with ordinal numbers because they only show sequence.
• Ordinal variables are considered as “in between” qualitative and
quantitative variables. In other words, the ordinal data is qualitative
data for which the values are ordered.
• Examples of ordinal data:
• The first, second and third person in a competition.

• Letter grades: A, B, C, and etc.

• When a company asks a customer to rate the sales experience on a scale of 1-10.

• Economic status: low, medium and high. Slide - 52

• Nominal vs. Ordinal Data

Slide - 53

• Discrete vs. Continuous Data


• In statistics, marketing research, and data science, many decisions
depend on whether the basic data is discrete or continuous.

• 5. Discrete data
• Discrete data is a count that involves only integers. The discrete
values cannot be subdivided into parts. For example, the number of
children in a class is discrete data. You can count whole individuals.
You can’t count 1.5 kids.
• Discrete data can take only certain values. The data variables cannot
be divided into smaller parts. It has a limited number of possible
values e.g. days of the month.
• Examples of discrete data:
• The number of students in a class.

• The number of workers in a company.

• The number of test questions you answered correctly


Slide - 54

18
2/10/2026

• Discrete vs. Continuous Data

• [Link] data
• Continuous data is information that could be meaningfully divided into
finer levels. It can be measured on a scale or continuum and can
have almost any numeric value.
• The continuous variables can take any value between two numbers.
For example, between 50 and 72 inches, there are literally millions of
possible heights: 52.04762 inches, 69.948376 inches and etc.
• Examples of continuous data:
• The amount of time required to complete a project.

• The height of children.

• The square footage of a two-bedroom house.

• The speed of cars.

Slide - 55

• Discrete vs. Continuous Data

Slide - 56

Organization/Sources of Data
• Data organization is the practice of categorizing and classifying
data to make it more usable. Similar to a file folder, where we
keep important documents, you’ll need to arrange your data in
the most logical and orderly fashion, so you — and anyone else
who accesses it — can easily find what they’re looking for.
• It is also needed that unauthorized person should not have
access to our data.

Slide - 57

19
2/10/2026

Organization/Sources of Data
DATA IS BEING COLLECTED
• Data includes information produced by humans and devices.
• Device-driven data is largely clean and organized,
• But of far greater interest is human-driven data that exist in
various formats and need more exquisite tools for proper
processing and management.

Slide - 58

Organization/Sources of Data
The data collection is focused on the following types of data:
• Network data. This type of data is gathered on all kinds of
networks, including social media, information and technological
networks, the Internet and mobile networks, etc.
• Real-time data. They are produced on online streaming media,
such as YouTube, Twitch, Skype, or Netflix.
• Transactional data. They are gathered when a user makes an
online purchase (information on the product, time of purchase,
payment methods, etc.)
• Geographic data. Location data of everything, humans,
vehicles, building, natural reserves, and other objects are
continuously supplied with satellites.

Slide - 59

Organization/Sources of Data
• Natural language data. These data are gathered mostly from
voice searches that can be made on different devices accessing
the Internet.
• Time series data. This type of data is related to the observation
of trends and phenomena taking place at this very moment and
over a period of time, for instance, global temperatures, mortality
rates, pollution levels, etc.
• Linked data. They are based on HTTP, RDF, SPARQL, and
URIs web technologies and meant to enable semantic
connections between various databases so that computers could
read and perform semantic queries correctly.

Slide - 60

20
2/10/2026

Organization/Sources of Data
Explanation
• RDF. The Resource Description Framework (RDF) is a
general framework for representing interconnected data on the
web. RDF statements are used for describing and exchanging
metadata, which enables standardized exchange of data
based on relationships
• HTTP. The Hypertext Transfer Protocol (HTTP) is the
foundation of the World Wide Web, and is used to load
webpages using hypertext links. HTTP is an application layer
protocol designed to transfer information between networked
devices and runs on top of other layers of the network protocol
stack

Slide - 61

Organization/Sources of Data
The data collection is focused on the following types of data:
• SPARQL. SPARQL Protocol and RDF Query Language,
enables users to query information from databases or any data
source that can be mapped to RDF. The SPARQL standard is
designed and endorsed by the W3C (World Wide Web
Consortium) and helps users and developers focus on what they
would like to know instead of how a database is organized.
• URIs. A Uniform Resource Identifier, is a unique sequence of
characters that identifies an abstract or physical resource, such
as resources on a webpage, mail address, phone number,
books, real-world objects such as people and places, concepts.

Slide - 62

Organization/Sources of Data
The data collection is focused on the following types of data:
• Asking for it. the majority of firms prefer asking users directly to
share their personal information. They give these data when
creating website accounts or buying online. The minimum
information to be collected includes a username and an email
address, but some profiles require more details.
• Cookies and Web Beacons. Cookies and web beacons are two
widely used methods to gather the data on users, namely, what
web pages they visit and when. They provide basic statistics
about how a website is used. Cookies and web beacons in no
way compromise your privacy but just serve to personalize your
experience with one or another web source.

Slide - 63

21
2/10/2026

Organization/Sources of Data
The data collection is focused on the following types of data:
• Email tracking. Email trackers are meant to give more
information on the user actions in the mailbox. In particular, an
email tracker allows detecting when an email was opened. Both
Google and Yahoo use this method to learn their users’
behavioural patterns and provide personalized advertising.

Slide - 64

Data For Analytics

Slide - 65

Examples of Data Sources and Uses


• Annual reports
• Accounting audits
• Financial profitability analysis
• Economic trends
• Marketing research
• Operations management performance
• Human resource measurements
• Web behavior
– page views, visitor’s country, time of view, length of time, origin and
destination paths, products they searched for and viewed, products
purchased, what reviews they read, and many others
Slide - 66

22
2/10/2026

Data Quality Management


Data quality is defined as: “The degree to which data meets a
company’s expectations of accuracy, validity, completeness,
and consistency”
• By tracking data quality, a business can pinpoint potential issues
harming quality, and ensure that shared data is fit to be used for
a given purpose.
• When collected data fails to meet the company expectations of
accuracy, validity, completeness, and consistency, it can have
massive negative impacts on customer service, employee
productivity, and key strategies.
• Quality data is key to making accurate, informed decisions.
While all data has some level of “quality,” a variety of
characteristics and factors determines the degree of data quality
(high-quality versus low-quality). Different data quality
characteristics will be more important to various users Slide - 67

Data Quality Management


• Six popular data quality dimensions are:
• Completeness: Completeness is defined as a measure of the
percentage of data that is missing within a dataset.
• Timeliness: Timeliness measures how up-to-date or antiquated
the data is at any given moment.
• Validity: Validity refers to information that fails to follow specific
company formats, rules, or processes.
• Integrity: Integrity of data refers to the level at which the
information is reliable and trustworthy.
• Uniqueness: Uniqueness is a data quality characteristic most
often associated with customer profiles.
• Consistency: It ensures that the source of the information
collection is capturing the correct data based on the unique
objectives of the department or company. Slide - 68

Dealing with Missing or Incomplete


Data
The concept of missing data is implied in the
name: its data that is not captured for a variable for
the observation in question.
Missing data reduces the statistical power of the
analysis, which can distort the validity of the
results.

Slide - 69

23
2/10/2026

Dealing with Missing or Incomplete


Data
Imputation vs Removing Data
• The imputation method develops reasonable guesses
for missing data. It’s most useful when the percentage of
missing data is low. If the portion of missing data is too
high, the results lack natural variation that could result in
an effective model.
• The other option is to remove/delete data. When dealing
with data that is missing at random, related data can be
deleted to reduce bias. Removing data may not be the
best option if there are not enough observations to result
in a reliable analysis. In some situations, observation of
specific events or factors may be required.
Slide - 70

Dealing with Missing or Incomplete


Data
Reasons for missing Data
• Missing at Random(MAR): means the data is missing relative
to the observed data. It is not related to the specific missing
values. The data is not missing across all observations but only
within sub-samples of the data. It is not known if the data should
be there; instead, it is missing given the observed data. The
missing data can be predicted based on the complete observed
data.
• Missing Completely at Random (MCAR) : In the MCAR
situation, the data is missing across all observations regardless
of the expected value or other variables. Data scientists can
compare two sets of data, one with missing observations and
one without. Using statistical tests( [Link] t-test), if there is no
difference between the two data sets, the data is characterized
Slide - 71
as MCAR.

Dealing with Missing or Incomplete


Data
Reasons for missing Data
• Missing at Not at Random(MNAR): The MNAR category
applies when the missing data has a structure to it. In other
words, there appear to be reasons the data is missing. In a
survey, perhaps a specific group of people – say women ages
45 to 55 – did not answer a question.
• Deletion:
• Data is missed because of deletion :
• List wise
• Pair wise
• Dropping Variables

Slide - 72

24
2/10/2026

Dealing with Missing or Incomplete


Data
• List wise:
• All data for an observation that has one or more missing values
are deleted.
• The analysis is run only on observations that have a complete set
of data.
• If the data set is small, it may be the most efficient method to
eliminate those cases from the analysis.
• However, in most cases, the data are not missing completely at
random (MCAR). Deleting the instances with missing
observations can result in biased parameters and estimates and
reduce the statistical power of the analysis.

Slide - 73

Dealing with Missing or Incomplete


Data
• Pair wise:
• Pair wise deletion assumes data are missing completely at
random (MCAR), but all the cases with data, even those with
missing data, are used in the analysis.
• Pairwise deletion allows data scientists to use more of the data.
However, the resulting statistics may vary because they are
based on different data sets. The results may be impossible to
duplicate with a complete set of data.
• Dropping Variables
• If data is missing for more than 60% of the observations, it may
be wise to discard it if the variable is insignificant.

Slide - 74

Dealing with Missing or Incomplete


Data
Imputation
• When data is missing, it may make sense to delete data,.
However, that may not be the most effective option.
• For example, if too much information is discarded, it may not be
possible to complete a reliable analysis.
• Or there may be insufficient data to generate a reliable prediction
for observations that have missing data.
• Instead of deletion, data scientists have multiple solutions to
impute the value of missing data. Depending why the data are
missing, imputation methods can deliver reasonably reliable
results.
• Discussed in next slides

Slide - 75

25
2/10/2026

Dealing with Missing or Incomplete


Data
Mean, Median and Mode
• This is one of the most common methods of imputing values
when dealing with missing data.
• In cases where there are a small number of missing
observations, data scientists can calculate the mean or median
of the existing observations.
• However, when there are many missing variables, mean or
median results can result in a loss of variation in the data.
• This method does not use time-series characteristics or depend
on the relationship between the variables.

Slide - 76

Dealing with Missing or Incomplete


Data
Time-Series Specific Methods
• There are four types of time-series data:
• No trend or seasonality.
• Trend, but no seasonality.
• Seasonality, but no trend.
• Both trend and seasonality
• The time series methods of imputation assume the adjacent
observations will be like the missing data.
• These methods work well when that assumption is valid.
However, these methods won’t always produce reasonable
results, particularly in the case of strong seasonality.

Slide - 77

Dealing with Missing or Incomplete


Data
Last Observation Carried Forward (LOCF) & Next
Observation Carried Backward (NOCB)
• These options are used to analyze longitudinal repeated
measures data, in which follow-up observations may be missing.
In this method, every missing value is replaced with the last
observed value.
• Longitudinal data track the same instance at different points
along a timeline.
• This method is easy to understand and implement.
• However, this method may introduce bias when data has a
visible trend. It assumes the value is unchanged by the missing
data.

Slide - 78

26
2/10/2026

Dealing with Missing or Incomplete


Data
Linear Interpolation
• Linear interpolation is often used to approximate a value of some
function by using two known values of that function at other
points.
• This formula can also be understood as a weighted average.
The weights are inversely related to the distance from the end
points to the unknown point. The closer point has more influence
than the farther point.
• When dealing with missing data, one should use this method in
a time series that exhibits a trend line, but it’s not appropriate for
seasonal data.

Slide - 79

Dealing with Missing or Incomplete


Data
Seasonal Adjustment with Linear Interpolation
• When dealing with data that exhibits both trend and seasonality
characteristics, use seasonal adjustment with linear
interpolation.
• First perform the seasonal adjustment by computing a centered
moving average or taking the average of multiple averages –
say, two one-year averages – that are offset by one period
relative to another.

Slide - 80

Dealing with Missing or Incomplete


Data
Multiple Imputations
• Multiple imputations is considered a good approach for data sets
with a large amount of missing data.
• Instead of substituting a single value for each missing data
point, the missing values are exchanged for values that
encompass the natural variability and uncertainty of the right
values.
• Using the imputed data, the process is repeated to make
multiple imputed data sets.
• Each set is then analyzed using the standard analytical
procedures, and the multiple analysis results are combined to
produce an overall result.

Slide - 81

27
2/10/2026

Dealing with Missing or Incomplete


Data
K Nearest Neighbours (KNN)
• In this method, data scientists choose a distance measure for k
neighbours, and the average is used to impute an estimate.
• The data scientist must select the number of nearest neighbours
and the distance metric. KNN can identify the most frequent
value among the neighbours and the mean among the nearest
neighbours.

Slide - 82

Big Data

• Big data refers to massive amounts of business data


(volume) from a wide variety of sources (variety), much of
which is available in real time (velocity), and much of
which is uncertain or unpredictable (veracity).
• “The effective use of big data has the potential to
transform economies, delivering a new wave of
productivity growth and consumer surplus. Using big data
will become a key basis of competition for existing
companies, and will create new competitors who are able
to attract employees that have the critical skills for a big
data world.” - McKinsey Global Institute, 2011
Slide - 83

Data Reliability and Validity

• Reliability - data are accurate and consistent.


• Validity - data correctly measures what it is
supposed to measure.

Slide - 84

28
2/10/2026

Data Reliability and Validity Examples


(1 of 3)

A tire pressure gage that consistently reads


several pounds of pressure below the true value is
not reliable, although it is valid because it does
measure tire pressure.

Slide - 85

Data Reliability and Validity Examples


(2 of 3)

The number of calls to a customer service desk


might be counted correctly each day (and thus is a
reliable measure) but not valid if it is used to
assess customer dissatisfaction, as many calls
may be simple queries.

Slide - 86

Data Reliability and Validity Examples


(3 of 3)

A survey question that asks a customer to rate the


quality of the food in a restaurant may be neither
reliable (because different customers may have
conflicting perceptions) nor valid (if the intent is to
measure customer satisfaction, as satisfaction
generally includes other elements of service
besides food).

Slide - 87

29
2/10/2026

What is big Data?

Slide - 88

Big Data Management


Big Data consists of huge amounts of information
that cannot be stored or processed using
traditional data storage mechanisms or processing
techniques. It generally consists of three different
variations.
I. Structured
II. Semi Structured
III. Unstructured

Slide - 89

Big Data Management …….


I. Structured data
• (As its name suggests) has a well-defined
structure and follows a consistent order.
• This kind of information is designed so that it
can be easily accessed and used by a person
or computer.
• Structured data is usually stored in the well-
defined rows and columns of a table (such as a
spreadsheet) and databases — particularly
relational database management systems, or
RDBMS. Slide - 90

30
2/10/2026

Big Data Management …….


II. Semi-structured data
• Exhibits a few of the same properties as
structured data.
• But for the most part, this kind of information
has no definite structure and cannot conform to
the formal rules of data models such as an
RDBMS.

Slide - 91

Big Data Management …….


III. Unstructured data
• Possesses no consistent structure across its
various forms and does not obey conventional
data models’ formal structural rules.
• In very few instances, it may have information
related to date and time.

Slide - 92

Big Data Management : Characteristics

Slide - 93

31
2/10/2026

Big Data Management : Characteristics

Slide - 94

Big Data Management : Characteristics


I. Volume
• This trait refers to the immense amounts of
information generated every second via social
media, cell phones, cars, transactions,
connected sensors, images, video, and text.
• In petabytes (1000 TB), terabytes (1000GB),
or even zettabytes ( 1 billion TB), these
volumes can only be managed by big data
technologies.

Slide - 95

Big Data Management : Characteristics


II. Variety
• To the existing landscape of transactional and
demographic data such as phone numbers and
addresses, information in the form of
photographs, audio streams, video, and a host of
other formats now contributes to a multiplicity of
data types — about 80% of which are completely
unstructured.
• Having different attributes
• Collected at various social media platforms
• Ca be digital, audio or video
Slide - 96

32
2/10/2026

Big Data Management : Characteristics


III. Velocity
• Information is streaming into data repositories at a
prodigious rate, and this characteristic alludes to
the speed of data accumulation.
• It also refers to the speed with which big data can
be processed and analyzed to extract the insights
and patterns it contains. These days, that speed
is often real-time.
• Also refers to data flow direction
• Stock Exchange collects 1 terabyte of data in a single
trading session, and having current data and real-time
rules for trades and predictive modeling are important for
managing stock portfolios.
Slide - 97

Big Data Management : Characteristics


IV. Veracity
• This is the degree of reliability and truth that big
data has to offer in terms of its relevance,
cleanliness, and accuracy.
• Concerned with uncertainty in data. Missing values
V. Value
• Since the primary aim of big data gathering and
analysis is to discover insights that can inform
decision-making and other processes, this
characteristic explores the benefit or otherwise that
information and analytics can ultimately produce.
Slide - 98

Big Data Management : Characteristics

Slide - 99

33
2/10/2026

Big Data Management : Opportunities,


Challenges & Technologies
• Although big data represents opportunities, it also presents
challenges in terms of data storage and processing, security
and available analytical talent.
• The four Vs indicate that big data creates challenges in terms
of how these complex data can be captured, stored,
processed, secured, and then analyzed.
• Traditional databases typically assume that data fit into nice
rows and columns, but that is not always the case with
• big data. Also, the sheer volume (the first V) often means
that it is not possible to store all of the data on a single
computer.

Slide -
100

Big Data Management : Flow Chart

Slide -
107

Big Data Management :Video

Slide -
108

34
2/10/2026

Big Data Management :Video (Online)

[Link]

Slide -
109

Understanding Different Types of Data

Slide -
110

Understanding Different Types of Data

Slide -
111

35
2/10/2026

Understanding Different Types of Data

Slide -
112

Understanding Different Types of Data

Slide -
113

Understanding Different Types of Data

Slide -
114

36
2/10/2026

Understanding Different Types of Data

Slide -
115

Big Data Management : Live Examples


• Kroger Understands Its Customers
• Kroger is the largest retail grocery chain in the United States.
It sends over 11 million pieces of direct mail to its customers
each quarter. The quarterly mailers each contain 12 coupons
that are tailored to each household based on several years of
shopping data obtained through its customer loyalty card
program.
• By collecting and analyzing consumer behavior at the
individual household level, and better matching its coupon
offers to shopper interests, Kroger has been able to realize a
far higher redemption rate on its coupons.
• In the six-week period following distribution of the mailers,
over 70% of households redeem at least one coupon, leading
to an estimated coupon revenue of $10 billion for Kroger. Slide -
116

Big Data Management : Live Examples


• MagicBand at Disney
• The Walt Disney Company offers a wristband to visitors to its
Orlando, Florida, Disney World theme park.
• Known as the MagicBand, the wristband contains technology
that can transmit more than 40 feet and can be used to track
each visitor’s location in the park in real time. The band can link
to information that allows Disney to better serve its visitors.
• Prior to the trip to Disney World, a visitor might be asked to fill
out a survey on their birth date and favorite rides, characters,
and restaurant table type and location. This information, linked to
the MagicBand, can allow Disney employees using smartphones
to greet you by name as you arrive, offer you products they know
you prefer, wish you a happy birthday, have your favorite
characters show up as you wait in line or have lunch at your
favorite table. Slide -
117

37
2/10/2026

Big Data Management : Live Examples


• The MagicBand can be linked to your credit card, so there is no
need to carry cash or a credit card. And during your visit, your
movement throughout the park can be tracked and the data can
be analyzed to better serve you during your next visit to the park.

Slide -
118

Big Data Management : Live Examples


• Coca-Cola Freestyle Gives Consumers Their Own
Personal Soft Drinks
• Coca-Cola, the largest beverage company in the world, sells its
products in over 200 countries. One of Coca-Cola’s innovations
is Coca-Cola Freestyle, a touch screen self-service soda
dispenser that can be found in numerous restaurants such as
Subway and Burger King.
• Coca-Cola Freestyle allows the customer to create their own
customized drink by combining existing flavors. For example, a
customer might combine Sprite with Hi-C Orange to create a
flavor that is not available in stores.
• Coca-Cola Freestyle is not only an innovation that better serves
customers through mass customization—the machines are also
a goldmine of customer preference data. The Freestyle
machines collect data on mixes that consumers around the world
Slide -
choose. 119

Big Data Management : Live Examples


• These data are analyzed to better understand customer
preferences and how they differ by country and region. Whereas
in the past, marketing analysts would need to rely on expensive
focus groups and multiple rounds of market testing to help
develop new products, Freestyle machines directly provide data
on consumer preferences.
• For example, a new bottled product, Cherry Sprite, was
launched when Freestyle data indicated that demand for that
combination of flavors would be strong in the United States.

Slide -
120

38
2/10/2026

Process of Business Analytics

Slide -
121

Problem Solving with Analytics

1. Recognizing a problem
2. Defining the problem
3. Structuring the problem
4. Analyzing the problem
5. Interpreting results and making a decision
6. Implementing the solution

Slide -
122

Recognizing a Problem
Problems exist when there is a gap between what
is happening and what we think should be
happening.
• For example, costs are too high compared with
competitors.

Slide -
123

39
2/10/2026

Defining the Problem

• Clearly defining the problem is not a trivial task.


• Complexity increases when the following occur:
– large number of courses of action
– the problem belongs to a group and not an individual
– competing objectives
– external groups are affected
– problem owner and problem solver are not the same
person
– time limitations exist

Slide -
124

Structuring the Problem

• Stating goals and objectives


• Characterizing the possible decisions
• Identifying any constraints or restrictions

Slide -
125

Analyzing the Problem

• Analytics plays a major role.


• Analysis involves some sort of experimentation
or solution process, such as evaluating different
scenarios, analyzing risks associated with
various decision alternatives, finding a solution
that meets certain goals, or determining an
optimal solution.

Slide -
126

40
2/10/2026

Interpreting Results and Making a


Decision

• Models cannot capture every detail of the real


problem.
• Managers must understand the limitations of
models and their underlying assumptions and
often incorporate judgment into making a
decision.

Slide -
127

Implementing the Solution

• Translate the results of the model back to the


real world.
• Requires providing adequate resources,
motivating employees, eliminating resistance to
change, modifying organizational policies, and
developing trust.

Slide -
128

What is Metrics in Analytics?

Slide -
129

41
2/10/2026

Indicative Metrics of Various segments of


Analytics

Slide -
130

Impact Cycle

Slide -
131

Challenges with Analytics

Slide -
132

42
2/10/2026

What is Predictive Analytics?

Slide -
133

• Predictive analytics uses historical and current


data, along with statistical algorithms, machine
learning, and AI, to forecast future outcomes,
trends, and behaviors, helping organizations
move beyond what happened to what will likely
happen next for better decision-making, risk
mitigation, and opportunity identification.
• It finds hidden patterns in data to predict things
like customer churn, sales, or fraud, enabling
proactive strategies in marketing, finance,
operations, and more, with accuracy depending
on data quality

Slide -
134

How it works?
• Data Collection: Gathers large volumes of
historical and real-time data.
• Pattern Identification: Employs statistical
models, machine learning (like decision trees,
neural networks), and AI to find relationships and
trends.
• Forecasting: Applies these patterns to predict
future events or probabilities, such as customer
actions, sales, or potential risks

Slide -
135

43
2/10/2026

Key Applications
• Gathers Marketing: Personalizing offers,
predicting customer lifetime value.
• Finance: Detecting credit card fraud, assessing
loan risk, forecasting revenue.
• Operations: Optimizing staffing, predicting
maintenance needs.
• Healthcare: Predicting patient outcomes or
disease outbreaks

Slide -
136

Benefits
• Enables proactive, data-driven decisions.
• Identifies risks and opportunities before they fully
develop.
• Improves efficiency, customer satisfaction, and
profitability.

Slide -
137

Techniques
• In general, there are two types of predictive
analytics models:
• classification and
• regression models and forecasting models.

Slide -
138

44
2/10/2026

Techniques
• Classification models attempt to put data objects
(such as customers or potential outcomes) into one
category or another. For instance, if a retailer has a
lot of data on different types of customers, they
may try to predict what types of customers will be
receptive to market emails.
• Regression models & Forecasting Models try to
predict continuous data, such as how much
revenue that customer will generate during their
relationship with the company.
Slide -
139

Techniques
• Predictive analytics tends to be performed with
three main types of techniques:
• Regression analysis & Forecasting
• Decision tree
• Neural Network

Slide -
140

Regression analysis & Forecasting


• Regression is a statistical analysis technique that estimates
relationships between variables. Regression is useful to
determine patterns in large datasets to determine the
correlation between inputs.
• It is best employed on continuous data that follows a known
distribution. Regression is often used to determine how one
or more independent variables affects another, such as how
a price increase will affect the sale of a product.

Slide -
141

45
2/10/2026

Regression analysis & Forecasting


• Meaning
• The literal meaning of the word 'regression' is 'stepping back
towards the average'.
• British biometrician Sir Francis Galton (1822-1911) studies the
heights of many persons and concluded that the offspring of
abnormally tall or short parents tend to regress to the average
population height.
• In analytics, regression analysis is concerned with the
measure of average relationship between variables.
• Maily the derivation of appropriate functional relationships
between variables is dealt with.
• Regression explains the nature of relationship between
variables. Slide -
142

Regression analysis & Forecasting


• Uses
i) Helps in establishing relationship between dependent variable
and independent variables. The independent variables may be
more than one. Such relationships are very useful in further
studies of the variables, under consideration.
(ii) Used for prediction. Once a relation is established between
dependent variable and independent variables, the value of
dependent variable can be predicted for given values of the
independent variables. This is very useful for predicting sale,
profit, investment, income, population, etc.

Slide -
143

Regression analysis & Forecasting


• Uses …
(iii) Specially used in Economics for estimating demand function,
production function, consumption function, supply function, etc.
A very important branch of Economics, called Econometrics, is
based on the techniques of regression analysis.
(iv) The coefficient of correlation between two variables can be
found easily by using the regression lines between the
variables.

Slide -
144

46
2/10/2026

Regression analysis & Forecasting


• Types
• Simple regression: Only two variables are under consideration,
then the regression is called simple regression. For example, the
study of regression between 'income' and 'expenditure’ for a group
of family would be termed as simple regression.
• Multiple regression: There are more than two variables under
consideration The regression is called partial regression if there
are more than two variables under consideration and relation
between only two variables is established after excluding the effect
of other variables.
• The simple regression is called linear regression if the point on
the scatter diagram of variables lies almost along a line otherwise
it is termed as non-linear regression or curvilinear regression

Slide -
145

Regression analysis & Forecasting


• Linear regression

Slide -
146

Regression analysis & Forecasting


• Multiple Linear regression

Slide -
147

47
2/10/2026

Techniques
• Decision trees
• Decision trees are classification models that place data into
different categories based on distinct variables.
• The method is best used when trying to understand an
individual's decisions.
• The model looks like a tree, with each branch representing
a potential choice, with the leaf of the branch representing
the result of the decision.
• Decision trees are typically easy to understand and work
well when a dataset has several missing variables.
Slide -
148

• Decision analysis is helpful for financial comparison of


alternatives under conditions of risk and uncertainty.
• A decision tree is a schematic model of the sequence of
steps in a problem and the conditions and consequences of
each step.
• They are used in situations involving multiphase decisions
and interdependent decisions that must be made in
sequence.
• Decision tree analysis provides:
I. A way of structuring complex multi decisions
II. A direct way of dealing with uncertain event
III. An objective way of determining the relative value of each
decision alternative Slide -
149

• The concept of expected value (EV) used gives only


relative measure of value and not absolute measures.
• Components of Decision tree
• Root Node: Starting point representing the whole dataset.
• Branches: Lines connecting nodes showing the flow from
one decision to another.
• Internal Nodes: Points where decisions are made based
on data features.
• Leaf Nodes: End points of the tree where the final decision
or prediction is made.

Slide -
150

48
2/10/2026

Slide -
151

• Main two types of Decision tree based upon target


variable
• Classification Trees: Used for predicting categorical
outcomes like spam or not spam. These trees split the data
based on features to classify data into predefined
categories.
• Regression Trees: Used for predicting continuous
outcomes like predicting house prices. Instead of assigning
categories, it provides numerical predictions based on the
input features.

Slide -
152

• How Decision Tree works?


1. Start with the Root Node: It begins with a main question at the
root node which is derived from the dataset’s features.
2. Ask Yes/No Questions: From the root, the tree asks a series of
yes/no questions to split the data into subsets based on specific
attributes.
3. Branching Based on Answers: Each question leads to different
branches:
• If the answer is yes, the tree follows one path.
• If the answer is no, the tree follows another path.

4. Continue Splitting: This branching continues through further


decisions helps in reducing the data down step-by-step.
5. Reach the Leaf Node: The process ends when there are no more
useful questions to ask leading to the leaf node where the final
decision or prediction is made Slide -
153

49
2/10/2026

• How Decision Tree works: Example


Imagine we need to decide whether to drink coffee based on the
time of day and how tired we feel. The tree first checks the time:
1. In the morning: It asks “Tired?”
• If yes, the tree suggests drinking coffee.
• If no, it says no coffee is needed.
2. In the afternoon: It asks again “Tired?”
• If yes, it suggests drinking coffee.
• If no, no coffee is needed.

Slide -
154

• Advantages
• Easy to Understand: Decision Trees are visual which makes it
easy to follow the decision-making process.
• Versatility: Can be used for both classification and regression
problems.
• No Need for Feature Scaling: Unlike many machine learning
models, it don’t require us to scale or normalize our data.
• Handles Non-linear Relationships: It capture complex, non-
linear relationships between features and outcomes effectively.
• Interpretability: The tree structure is easy to interpret helps in
allowing users to understand the reasoning behind each decision.
• Handles Missing Data: It can handle missing values by using
strategies like assigning the most common value or ignoring
missing data during splits.
Slide -
155

• Disadvantages
• Overfitting: They can overfit the training data if they are too deep
which means they memorize the data instead of learning general
patterns. This leads to poor performance on unseen data.
• Instability: It can be unstable which means that small changes in
the data may lead to significant differences in the tree structure and
predictions.
• Bias towards Features with Many Categories: It can become
biased toward features with many distinct values which focuses too
much on them and potentially missing other important features
which can reduce prediction accuracy.
• Difficulty in Capturing Complex Interactions: Decision Trees may
struggle to capture complex interactions between features which
helps in making them less effective for certain types of data.
• Computationally Expensive for Large Datasets: For large
datasets, building and pruning a Decision Tree can be
computationally intensive, especially as the tree depth increases.
Slide -
156

50
2/10/2026

• Applications
• Loan Approval in Banking: Banks use Decision Trees to assess
whether a loan application should be approved. The decision is
based on factors like credit score, income, employment status and
loan history. This helps predict approval or rejection helps in
enabling quick and reliable decisions.
• Medical Diagnosis: In healthcare they assist in diagnosing
diseases. For example, they can predict whether a patient has
diabetes based on clinical data like glucose levels, BMI and blood
pressure. This helps classify patients into diabetic or non-diabetic
categories, supporting early diagnosis and treatment.
• Predicting Exam Results in Education: Educational institutions
use to predict whether a student will pass or fail based on factors
like attendance, study time and past grades. This helps teachers
identify at-risk students and offer targeted support.

Slide -
157

• Applications …
• Customer Churn Prediction: Companies use Decision Trees to
predict whether a customer will leave or stay based on behavior
patterns, purchase history, and interactions. This allows businesses
to take proactive steps to retain customers.
• Fraud Detection: In finance, Decision Trees are used to detect
fraudulent activities, such as credit card fraud. By analyzing past
transaction data and patterns, Decision Trees can identify
suspicious activities and flag them for further investigation.

Slide -
158

Slide -
159

51
2/10/2026

We choose the product with the highest EV. The option with
the lower EV is shown with two lines cutting across it. We’ve
“rolled back” the tree from 6 uncertain outcomes to the one
branch at the decision point with the highest expected value.
We’re pruning off poor uses of capital that lead to undesirable
outcomes.

Slide -
160

Neural networks
• Neural networks are machine learning methods that are
useful in predictive analytics when modeling very complex
relationships.
• Essentially, they are powerhouse pattern recognition
engines.
• Neural networks are best used to determine nonlinear
relationships in datasets, especially when no known
mathematical formula exists to analyze the data. Neural
networks can be used to validate the results of decision
trees and regression models.

Slide -
161

Neural networks
• Neural networks are machine learning models that mimic
the complex functions of the human brain. These models
consist of interconnected nodes or neurons that process
data, learn patterns and enable tasks such as pattern
recognition and decision-making.

Slide -
162

52
2/10/2026

Neural networks
• Neural networks are capable of
learning and identifying patterns
directly from data without pre-defined
rules. These networks are built from
several key components:
• Neurons: The basic units that receive inputs, each
neuron is governed by a threshold and an activation
function.

• Connections: Links between neurons that carry


information, regulated by weights and biases.

• Weights and Biases: These parameters determine


the strength and influence of connections.

• Propagation Functions: Mechanisms that help


process and transfer data across layers of neurons.

• Learning Rule: The method that adjusts weights


Slide -
and biases over time to improve accuracy. 163

Neural networks
• Process
• Input Computation: Data is fed into the network.
• Output Generation: Based on the current parameters, the
network generates an output.
• Iterative Refinement: The network refines its output by
adjusting weights and biases, gradually improving its
performance on diverse tasks.

Slide -
164

Neural networks
• Structure

• Three layers of neural network algorithm:


• The Input Layer: This enters past data values into the next
layer.
• The Hidden Layer: This is a key component of a neural
network. It has complex functions that create predictors. A set
of nodes in the hidden layer, called neurons, represents
mathematical functions that modify the input data.
• The Output Layer: Here, the predictions made in the hidden
layer are collected to produce the final layer, which is the
model's prediction. Slide -
165

53
2/10/2026

How Neural networks predicts?


• Each neuron takes into consideration a set of input values.
Each of them gets linked to a "weight", which is a numerical
value that can be derived using either supervised or
unsupervised training, such as data clustering, and a value
called "bias".
• The network chooses from the answer produced by a
neuron based on its weight and bias.

Neuron is the most fundamental


unit of processing. Also called
perceptron

Slide -
166

Applications
• Pattern Recognition: Identifying complex patterns in
customer behavior
• Anomaly Detection: Finding unusual patterns that might
indicate fraud or errors
• Predictive Modeling: Forecasting future trends with high
accuracy
• Natural Language Processing: Understanding and
processing human language

Slide -
167

Challenges
• Resource intensive
• Results are often hard to interpret

Slide -
168

54
2/10/2026

Data Cleaning/Cleansing

Slide -
169

What is data cleaning?


• Data cleaning is the essential process of identifying and
correcting, removing, or modifying inaccurate, incomplete,
improperly formatted, or duplicate data within a dataset to
ensure high-quality, reliable analysis.
• It increases dataset accuracy and consistency, which
improves machine learning model performance and
decision-making
• When combining multiple data sources, there are many
opportunities for data to be duplicated or mislabeled.
• If data is incorrect, outcomes and algorithms are unreliable,
even though they may look correct.
• There is no one absolute way to prescribe the exact steps in
the data cleaning process because the processes will vary
Slide -
from dataset to dataset. 170

Key Aspects
• Handling Missing Data: Imputing missing values, or
removing rows/columns with substantial missing data.
• Removing Duplicates: Identifying and eliminating
redundant or identical records.
• Standardization/Formatting: Fixing structural errors, such
as inconsistent date formats, typos, or inconsistent naming
conventions.
• Outlier Handling: Removing or correcting data points that
are drastically different from others, which may skew
analysis

Slide -
171

55
2/10/2026

Why do we need clean data?


• It ensures accuracy
• It improves efficiency
• It builds trust
• It brings about compliance

Slide -
172

Difference between Data Cleaning &


Data Transformation
• Data cleaning is the process that removes data that does
not belong in your dataset.
• Data transformation is the process of converting data from
one format or structure into another.
• Transformation processes can also be referred to as data
wrangling, or data munging, transforming and mapping data
from one "raw" data form into another format for
warehousing and analyzing.

Slide -
173

Data Cleaning Cycle

Slide -
174

56
2/10/2026

Data Cleaning steps


• Profiling & Assessment: Inspect the data to understand its
current state and identify issues like missing values or
strange patterns.
• Removing Duplicates: Eliminate identical or near-identical
records that can inflate metrics and skew results.
• Fixing Structural Errors: Correct typos, inconsistent
capitalization (e.g., "Iron" vs. "iron"), and mislabeled
categories (e.g., "N/A" vs. "Not Applicable").
• Standardizing Formats: Ensure dates, currencies, and
units are uniform across the dataset (e.g., converting all
dates to YYYY-MM-DD).
• Handling Missing Data: Address gaps by either removing
incomplete rows, imputing values using averages or
predictive regression, or flagging them as "Unknown". Slide -
175

Data Cleaning steps


• Filtering Outliners: Often, there will be one-off
observations where, at a glance, they do not appear to fit
within the data you are analyzing. If you have a legitimate
reason to remove an outlier, like improper data-entry, doing
so will help the performance of the data you are working
with.
• Validate & QA: At the end of the data cleaning process,
you should be able to answer following questions as a part
of basic validation:
– Does the data make sense?
– Does the data follow the appropriate rules for its field?
– Does it prove or disprove your working theory, or bring any insight
to light?
– Can you find trends in the data to help you form your next theory?
– If not, is that because of a data quality issue?
Slide -
176

Tools and Software


• Spreadsheets: Microsoft Excel is a standard for small
datasets, offering built-in functions like TRIM, CLEAN, and
"Remove Duplicates".
• Programming: Python (Pandas/NumPy) and R are the go-
to for large-scale automation and complex statistical
cleaning.
• Specialized Software: Tools like OpenRefine (open-
source), Tableau Prep, Trifacta, and Alteryx provide visual,
no-code interfaces for cleaning complex data.
• Cloud & AI: Modern platforms like [Link] and Tamr
use AI to automate deduplication and format normalization
at an enterprise scale.

Slide -
177

57
2/10/2026

5 Characteristics of Quality Data


• Validity. The degree to which your data conforms to
defined business rules or constraints.
• Accuracy. Ensure your data is close to the true values.
• Completeness. The degree to which all required data is
known.
• Consistency. Ensure your data is consistent within the
same dataset and/or across multiple data sets.
• Uniformity. The degree to which the data is specified using
the same unit of measure.

Slide -
178

Benefits of Data Cleaning


• Removal of errors when multiple sources of data are at
play.
• Fewer errors make for happier clients and less-frustrated
employees.
• Ability to map the different functions and what your data is
intended to do.
• Monitoring errors and better reporting to see where errors
are coming from, making it easier to fix incorrect or corrupt
data for future applications.
• Using tools for data cleaning will make for more efficient
business practices and quicker decision-making.

Slide -
179

Limitations of Data Cleaning


• Time-consuming: It is very time-consuming task specially
for large and complex datasets.
• Error-prone: It can result in loss of important information.
• Cost and resource-intensive: It is resource-intensive
process that requires significant time, effort and expertise. It
can also require the use of specialized software tools.
• Overfitting: Data cleaning can contribute to overfitting by
removing too much data.

Slide -
180

58

You might also like