MB25202 - BUSINESS ANALYTICS
Module 1: Foundations of Business Analytics 10
Business Analytics – Definitions – Importance – Key terminologies – Business Analytics
Process – Relationship between Business Analytics and Organizational Decision Making
– Business Analytics for Competitive Advantage – Strategic role of analytics.
Module 2: Data and Resource Management in Analytics 10
Managing Analytics Infrastructure – Human Resources and Data Requirements –
Aligning Organizational Structure with BA – Role of Technology – Information Policy –
Data Governance – Data Quality Management – Change Management in BA.
Module 3: Descriptive Analytics and Visualization 10
Introduction to Descriptive Analytics – Visualizing and Exploring Data – Descriptive
Statistics – Sampling Techniques – Estimation Methods – Probability Distributions used
in Descriptive Analytics – Use of dashboards and visualization tools.
Module 4: Predictive Analytics and Data Mining 10
Introduction to Predictive Analytics – Predictive Models – Data-driven vs. Logic-driven
Models – Predictive Modeling Procedures – Supervised vs. Unsupervised Learning –
Data Mining Techniques – Model Validation – Application in Business cases.
Module 5: Prescriptive Analytics and Optimization 10
Overview of Prescriptive Analytics – Prescriptive Modeling Techniques – Non-linear
Optimization – Simulation – Scenario Planning – Demonstrating Business Performance
Improvement – Role of decision support systems.
Module 6: Applications and Emerging Trends in BA 10
Industry applications in Marketing, HR, Finance, and Operations – Real-time analytics –
AI & ML integration – Cloud-based BA tools – Business Intelligence vs. BA – Ethics in
Analytics – Case Studies on contemporary analytics usage.
Dr. G. Sindhu, Professor, PPG Business School
UNIT I INTRODUCTION TO BUSINESS ANALYTICS
1.1BUSINESS ANALYTICS
Business Analytics is the process by which businesses use statistical methods and
technologies for analyzing historical data in order to gain new insight and improve
strategic decision-making.
Business analytics, a data management solution and business intelligence subset,
refers to the use of methodologies such as data mining, predictive analytics, and
statistical analysis in order to analyze and transform data into useful information,
identify and anticipate trends and outcomes, and ultimately make smarter,
data-driven business decisions.
COMPONENTS OF A BUSINESS ANALYTICS
a. Data Aggregation: Data is collected to one single, central location from where
sorting can begin. Inaccurate and incomplete data is removed and only usable
data is left behind. Even duplicate data is checked for and removed completely.
b. Data Mining: Data mining for business analytics sorts through large datasets to
identify trends and establish relationships
c. Association and Sequence Identification: These components are a pattern of
consumer behaviour. In the association part of the behaviour, consumers buy
products that are associated with each other like toothpaste and toothbrush or
shampoo and conditioner. In the sequence identification, the sequence on buying
is extracted like a sequence of an airline ticket, airport cab ride, and hotel room.
d. Text Mining: Explores and organizes large, unstructured text datasets for the
purpose of qualitative and quantitative analysis
e. Forecasting: Analyzes historical data from a specific period in order to make
informed estimates that are predictive in determining future events or behaviors
f. Predictive Analytics: Predictive business analytics uses a variety of statistical
techniques to create predictive models, which extract information from datasets,
identify patterns, and provide a predictive score for an array of organizational
outcomes.
Dr. G. Sindhu, Professor, PPG Business School
g. Optimization: Once trends have been identified and predictions have been made,
businesses can engage simulation techniques to test out best-case scenarios.
Businesses can really make great use of analytics and optimize their operations.
They can anticipate surges in demand and step production to maintain supply.
h. Data Visualization: Provides visual representations such as charts and graphs for
easy and quick data analysis.
1.2IMPORTANCE OF BA
a. Enhance Customer Experience
b. Make Informed Decisions
c. Reduce Employee Turnover
d. Improve Efficiency
e. Identify Frauds
f. Cut Manufacturing Costs
g. Make The Most Of Your Investment
h. Improved Advertising
i. Better Product Management
j. Tackle Problems
1.3TERMINOLOGIES
Business analytics begins with a data set (a simple collection of data or a data file) or
commonly with a database (a collection of data files that contain information on
people, locations, and so on). As databases grow, they need to be stored somewhere.
Technologies such as computer clouds and data warehousing store data. Database
storage areas have become so large that a new term was devised to describe them.
Big data describes the collection of data sets that are so large and complex that
software systems are hardly able to process them. Little data describes the smaller
data segments or files that help individual businesses keep track of customers.
Three terms in business literature are often related to one another: analytics, business
analytics, and business intelligence. Analytics can be defined as a process that
involves the use of statistical techniques, information system software, and
Dr. G. Sindhu, Professor, PPG Business School
operations research methodologies to explore, visualize, discover, and communicate
patterns or trends in data. Simply, analytics converts data into useful information.
Analytics is an older term commonly applied to all disciplines, not just business. A
typical example of the use of analytics is the weather measurements collected and
converted into statistics, which in turn predict weather patterns.
1.4PROCESS OF DATA ANALYTICS
Step 1: Identifying the Problem
The first step of the process is identifying the business problem. The problem could
be an actual crisis; it could be something related to recognizing business needs or
optimizing current processes. This is a crucial stage in Business Analytics as it is
important to clearly understand what the expected outcome should be. When the
desired outcome is determined, it is further broken down into smaller goals. Then,
business stakeholders decide the relevant data required to solve the problem. Some
important questions must be answered in this stage, such as: What kind of data is
available? Is there sufficient data? And so on.
Step 2: Exploring Data
Once the problem statement is defined, the next step is to gather data and, more
importantly, cleanse the data—most organizations would have plenty of data, but
not all data points would be accurate or useful. Organizations collect huge amounts
of data through different methods, but at times, junk data or empty data points
would be present in the dataset. These faulty pieces of data can hamper the analysis.
Hence, it is very important to clean the data that has to be analyzed. To do this, we
need computations for the missing data, remove outliers, and find new variables as a
combination of other variables. Plot time series graphs as they generally indicate
patterns and outliers. It is very important to remove outliers as they can have a
heavy impact on the accuracy of the model that you create. Moreover, cleaning the
data helps you get a better sense of the dataset.
Step 3: Analysis
Once the data is ready, the next thing to do is analyze it. Now to execute the same,
there are various kinds of statistical methods (such as hypothesis testing, correlation,
etc.) involved to find out the insights that you are looking for. You can use all of the
Dr. G. Sindhu, Professor, PPG Business School
methods for which you have the data. The prime way of analyzing is pivoting
around the target variable, so we need to take into account whatever factors that
affect the target variable. In addition to that, a lot of assumptions are also considered
to find out what the outcomes can be. Generally, at this step, the data is sliced, and
the comparisons are made. Through these methods, you are looking to get actionable
insights.
Step 4: Prediction and Optimization
Gone are the days when analytics was used to react. In today’s era, Business
Analytics is all about being proactive. Prediction techniques, such as neural networks
or decision trees are used to model the data. These prediction techniques will help to
find out hidden insights and relationships between variables, which will further help
to uncover patterns on the most important metrics. By principle, a lot of models are
used simultaneously, and the models with the most accuracy are chosen. In this
stage, a lot of conditions are also checked as parameters, and answers to a lot of
‘what if…?’ questions are provided.
Step 5: Making a Decision and Evaluating the Outcome
From the insights of the model built on target variables, a viable plan of action will
be established in this step to meet the organization’s goals and expectations. The said
plan of action is then put to work, and the waiting period begins. The actual
outcomes of the predictions are to be analyzed and find out how successful it is.
Measure and evaluate the outcomes.
Step 6: Optimizing and Updating
Dr. G. Sindhu, Professor, PPG Business School
Post the implementation of the solution, the outcomes are measured as mentioned
above. The methods through which the plan of action can be optimized, then those
can be implemented. This step is crucial for any analytics in the future because of
having an ever-improving database. Through this database, you can get closer and
closer to maximum optimization. Evaluate the ROI.
1.3.1 Data Analytics Types
Type of
Definition Purpose Methodologies
Analytics
To identify
Descriptive
possible trends in
statistics,
large data sets or
including
databases. The
measures of
The application of purpose is to get
central tendency,
simple statistical a rough picture of
measures of
techniques that what
dispersion charts,
Descriptive describe what is generally the data
graphs, sorting
contained in a looks like and
methods,
data set or what criteria
frequency
database. might have
distributions,
potential for
probability
identifying trends
distributions, and
or future business
sampling methods.
behavior.
Diagnostic
Diagnostic The purpose of
analytics is usually
Analytics, which diagnostic
performed using
helps you analytics is to
Diagnostic such techniques as
understand determine the
Analytics data discovery,
why something root cause of an
drill-down, data
happened in the occurrence or
mining, and
past trend.
correlations
An application of
Statistical methods
advanced
like multiple
statistical,
To build regression and
information
predictive ANOVA.
software, or
models designed Information system
Predictive operations
to identify and methods like data
research methods
predict future mining and sorting.
to identify
trends. Operations research
predictive
methods like
variables and
forecasting models.
build predictive
Dr. G. Sindhu, Professor, PPG Business School
models to identify
trends and
relationships not
easily observed in
a descriptive
analysis.
An application of
decision science, To allocate
Operations
management resources
research
science, and optimally to
methodologies
operations take advantage of
Prescriptive like linear
research predicted trends
programming and
methodologies to or
decision
make best use of future
theory.
allocable opportunities.
resources
1. Descriptive Analytics describes the happenings over time, such as whether the number
of views increased or decreased and whether the current month’s sales are better than
the last one.
2. Diagnostic Analytics focuses on the reason for the occurrence of any event. It requires
hypothesizing and involves a much diverse dataset. It examines data to answer
questions, such as “Did the weather impact the selling of soft drinks?” or “Did the new
ad strategy affect sales?”
3. Predictive Analytics focuses on the events that are expected to occur in the immediate
future. Predictive analytics tries to find answers to questions like, what happened to the
sales in the last hot summer season?
4. Prescriptive Analytics indicates a plan of action. If the chance of a hot summer
calculated as the average of the five weather models is above 58%, an evening shift can
be added to the brewery, and an additional tank can be rented to maximize the
production.
1.5 RELATIONSHIP WITH ORGANIZATIONAL DECISION MAKING
The BA process can solve problems and identify opportunities to improve business
performance. In the process, organizations may also determine strategies to guide
operations and help achieve competitive advantages. Typically, solving problems and
identifying strategic opportunities to follow are organization decision-making tasks.
Dr. G. Sindhu, Professor, PPG Business School
Identifying opportunities can be viewed as a problem of strategy choice requiring a
solution. It should come as no surprise that the BA process closely parallels classic
organization decision-making processes. As depicted in Fig, the business analytic
process has an inherent relationship to the steps in typical organization decision-making
processes.
In the BA process, the first step is to recognize that databases may contain information
that could both solve problems and find opportunities to improve business performance.
Then in Step 2 of the Organization decision making process, an exploration of the
problem to determine its size, impact, and other factors is undertaken to diagnose what
the problem is. Likewise, the BA descriptive analytic analysis explores factors that might
prove useful in solving problems and offering opportunities. The Organization decision
making process problem statement step is similarly structured to the BA predictive
analysis to find strategies, paths, or trends that clearly define a problem or opportunity
for an organization to solve problems. Finally, the ODMP’s last steps of strategy
selection and implementation involve the same kinds of tasks that the BA process
requires in the final prescriptive step (make an optimal selection of resource allocations
that can be implemented for the betterment of the organization).
Dr. G. Sindhu, Professor, PPG Business School
The decision-making foundation that has served ODMP for many decades parallels the
BA process. The same logic serves both processes and supports organization
decision-making skills and capacities.
1.6 BA FOR COMPETITIVE ADVANTAGE
Competitive advantage refers to an organization’s ability to create greater value than its
competitors by:
● Offering lower costs
● Delivering differentiated products or services
● Responding faster to customer and market changes
● Improving efficiency and innovation
BA enables firms to move from experience-based decisions to data-driven strategies.
Business Analytics Creates Competitive Advantage through
a. Superior Decision Making
b. Cost Optimization and Efficiency
c. Customer Insights and Differentiation
d. Predictive and Prescriptive Advantage
e. Speed, Agility, and Innovation
Dr. G. Sindhu, Professor, PPG Business School
Types of Competitive Advantage
Competitive Advantage Description
From a marketing standpoint, offer products or services
Price Leadership at the lowest cost to customers in the industry, while
making acceptable profit for the company.
To ensure the firm's resource usage in a way that seeks
Sustainability balance to hurt neither the environment nor the bottom
line of a firm's profitability.
Improve the internal business operations and activities
over competitors, lessening the cost to the customer.
Operations Efficiency
That reduced cost, if passed on to customers, can
provide a lower price advantage based on efficiency.
Make customer transactions easier or more pleasurable
than with other firms. This improves the service
Service Effectiveness characteristics of the firm while lowering the time it
takes to get services to the customer, thus enhancing
customer value.
Introduce completely new or notably better products or
services with the intention of disrupting competitors'
Innovation
businesses by obsoleting the current market entries with
break-through product offerings.
Provide customers with a variety of products, services,
Product Differentiation or features that competitors are not yet offering or are
unable to offer.
1.7 STRATEGIC ROLE OF ANALYTICS
The strategic role of analytics refers to how organizations use data analysis not just for
reporting or operational efficiency, but as a core driver of strategy, competitive
advantage, and long-term decision-making. The strategic role of analytics lies in
transforming data into actionable insights that shape organizational strategy, enhance
competitiveness, and ensure sustainable growth.
1. Analytics as a Strategic Asset
● Data is treated as an organizational asset, similar to capital or technology.
● Firms leverage analytics to identify opportunities, reduce risks, and create
value.
● Strategic analytics supports top-level management decisions, not just routine
operations.
Dr. G. Sindhu, Professor, PPG Business School
2. Enabling Evidence-Based Decision Making
● Moves decision-making from intuition-based to fact-based.
● Uses historical data, real-time data, and predictive models to support:
● Market entry decisions
● Pricing strategies
● Investment and expansion plans
3. Competitive Advantage Creation
● Analytics helps organizations:
o Understand customer behavior better than competitors
o Optimize supply chains and operations
o Innovate products and services
● Companies like Amazon, Netflix, and Google use analytics as a core
competitive differentiator.
4. Supporting Strategic Planning
● Analytics aids in:
o Environmental scanning (PESTLE, industry trends)
o Scenario analysis and forecasting
o SWOT analysis supported by data
● Predictive and prescriptive analytics help in anticipating future outcomes and
choosing optimal strategies.
5. Customer-Centric Strategy
● Analytics enables:
o Customer segmentation and targeting
o Personalization of products and services
o Measurement of customer lifetime value and retention
● Helps align business strategy with changing customer expectations,
especially in digital markets.
6. Performance Measurement and Control
● Strategic analytics supports:
o KPIs and Balanced Scorecards
o Monitoring strategic initiatives
o Identifying gaps between planned and actual performance
Dr. G. Sindhu, Professor, PPG Business School
● Ensures continuous alignment between strategy and execution.
7. Risk Management and Strategic Resilience
● Analytics helps identify:
o Financial, operational, and market risks
o Fraud and compliance issues
● Scenario modeling helps organizations prepare for uncertainty and
disruptions.
8. Innovation and Digital Transformation
● Analytics drives innovation by:
o Identifying unmet customer needs
o Supporting data-driven product development
o Enabling AI, machine learning, and automation initiatives
● Plays a key role in digital and Industry 4.0 strategies.
9. Strategic Alignment Across Functions
● Integrates data across marketing, HR, finance, and operations.
● Ensures all departments work toward common strategic goals.
● For example, HR analytics aligns workforce planning with business strategy.
10.Role of Leadership and Analytics Culture
● Strategic value of analytics depends on:
o Top management support
o Data-driven culture
o Skilled analytics teams and governance
Dr. G. Sindhu, Professor, PPG Business School
UNIT II
MANAGING RESOURCES FOR BUSINESS ANALYTICS
Managing Analytics Infrastructure – Human Resources and Data Requirements –
Aligning Organizational Structure with BA – Role of Technology – Information Policy –
Data Governance – Data Quality Management – Change Management in BA.
2.1 MANAGING ANALYTICS INFRASTRUCTURE
Analytics Infrastructure refers to the integrated set of technologies, processes, data
resources, and governance mechanisms that enable organizations to collect, store,
process, analyze, and visualize data effectively. Managing this infrastructure is critical
for transforming data into actionable insights and supporting data-driven
decision-making.
Components of Analytics Infrastructure
a) Data Sources
● Internal data: ERP, CRM, HR systems, transaction databases
● External data: Social media, market data, IoT devices, third-party data
b) Data Storage Systems
● Data warehouses
● Data lakes
● Cloud storage platforms
c) Data Processing & Integration
● ETL / ELT tools
● Real-time and batch processing frameworks
d) Analytics & Modeling Tools
● Descriptive, predictive, and prescriptive analytics tools
● Statistical, machine learning, and AI platforms
e) Visualization & Reporting
● Dashboards and BI tools
● Self-service analytics
2.2 HUMAN RESOURCES AND DATA REQUIREMENTS
Dr. G. Sindhu, Professor, PPG Business School
Human Resources (HR) plays a crucial role by providing workforce-related data that
supports strategic, tactical, and operational decisions. HR data enables organizations to
analyze employee performance, talent needs, productivity, and retention, thereby
improving organizational effectiveness. Human Resources (HR) data refers to
information collected and used to manage employees effectively and support strategic
workforce decisions.
2.2.1 ROLE OF HUMAN RESOURCES IN BUSINESS ANALYTICS
● Acts as a key data provider for people-related analytics
● Supports workforce planning and talent management
● Enables data-driven HR decisions instead of intuition-based practices
● Aligns HR strategy with business objectives
2.2.2 SOURCES OF HR DATA
● HRIS / HCM systems
● Payroll and attendance systems
● Performance management systems
● Learning Management Systems (LMS)
● Employee surveys and exit interviews
2.2.3 HR DATA REQUIREMENTS IN BUSINESS ANALYTICS
a. Employee Demographic Data
● Age, gender, education, experience
● Job role, department, location
● Useful for diversity analysis and workforce segmentation
b. Recruitment & Selection Data
● Source of hire, time-to-hire, cost-per-hire
● Candidate assessment scores
● Helps in predicting hiring success and recruitment efficiency
c. Training & Development Data
● Training hours, programs attended
Dr. G. Sindhu, Professor, PPG Business School
● Skill assessments, certifications
● Used to evaluate training effectiveness and skill gaps
d. Performance Management Data
● Appraisal scores, KPIs, productivity metrics
● Goal achievement and competency ratings
● Supports performance forecasting and reward decisions
e. Compensation & Benefits Data
● Salary structure, incentives, bonuses
● Benefits utilization
● Helps in pay equity and compensation optimization
f. Attendance & Workforce Productivity Data
● Absenteeism, overtime, shift data
● Utilization rates
● Used for capacity planning and productivity analysis
g. Employee Engagement & Retention Data
● Survey results, feedback scores
● Attrition rate, reasons for exit
● Enables attrition prediction and retention strategies
2.3 ORGANIZATION STRUCTURES ALIGNING BUSINESS ANALYTICS
According to Isson and Harriottto successfully implement business analytics (BA)
within organizations, the BA in whatever organizational form it takes must be fully
integrated throughout a firm. This requires BA resources to be aligned in a way that
permits a view of customer information within and across all departments, access to
customer information from multiple sources (internal and external to the
organization), access to historical analytics from a central repository, and making
technology resources align to be accountable for analytic success.
Organization Structures
Most organizations are hierarchical, with senior managers making the strategic
planning decisions, middle-level managers making tactical planning decisions, and
Dr. G. Sindhu, Professor, PPG Business School
lower-level managers making operational planning decisions. Within the hierarchy,
other organizational structures exist to support the development and existence of
groupings of resources like those needed for BA. These additional structures include
programs, projects, and teams. A program in this context is the process that seeks to
create an outcome and usually involves managing several related projects with the
intention of improving organizational performance. A program can also be a large
project. A project tends to deliver outcomes and can be defined as having temporary
rather than permanent social systems within or across organizations to accomplish
particular and clearly defined tasks, usually under time constraints. Projects are often
composed of teams. A team consists of a group of people with skills to achieve a
common purpose.
Teams are especially appropriate for conducting complex tasks that have many
interdependent sub tasks. The relationship of programs, projects, and teams with a
business hierarchy is presented in Figure.
Figure: Hierarchal relationships program, project, and team planning
In summary, one way to look at the alignment of BA resources is to view it as a
progression of assigned planning tasks from a BA program, to BA projects, and
eventually to BA teams for implementation, this hierarchical relationship is a way to
examine how firms align planning and decision-making workload to fit strategic
needs and requirements.
Dr. G. Sindhu, Professor, PPG Business School
BA organization structures usually begin with an initiative that recognizes the need
to use and develop some kind of program in analytics. Fortunately, most firms today
recognize this need. The question then becomes how to match the firm’s needs
within the organization to achieve its strategic, tactical, and operations objectives
within resource limitations. Planning the BA resource allocation within the
organizational structure of a firm is a starting place for the alignment of BA to best
serve a firm’s needs.
Even larger investments in BA resources might be required by firms that decide to
establish a whole BA department containing all the BA resources for a particular
organization. Although some firms create BA departments, the departments don’t
have to be large. Whatever the organization structure that is used, the role of BA is a
staff (not line management) role in their advisory and consulting mission for the firm.
In general, there are different ways to structure an organization to align its BA
resources to serve strategic plans. In organizations where functional departments are
structured on a strict hierarchy, separate BA departments or teams have to be
allocated to each functional area, as presented in this functional organization
structure may have the benefit of stricter functional control by the VPs of an
organization and greater efficiency in focusing on just the analytics within each
specialized area. On the other hand, this structure does not promote the
cross-department access that is suggested as a critical success factor for the
implementation of a BA program.
Dr. G. Sindhu, Professor, PPG Business School
Figure: Functional organization structure with BA
The needs of each firm for BA sometimes dictate positioning BA within existing
organization functional areas. Clearly, many alternative structures can house a BA
grouping. For example, because BA provides information to users, BA could be
included in the functional area of management information systems, with the chief
information officer (CIO) acting as both the director of information systems (which
includes database management) and the leader of the BA grouping.
An alternative organizational structure commonly found in large organizations aligns
resources by project or product and is called a matrix organization. As illustrated, this
structure allows the VPs some indirect control over their related specialists, which
would include the BA specialists but also allows direct control by the project or product
manager. This, similar to the functional organizational structure, does not promote the
cross-department access suggested for a successful implementation of a BA program.
Dr. G. Sindhu, Professor, PPG Business School
Figure: Matrix Organization Structure
This centralized BA organization structure minimizes investment costs by avoiding
duplications found in both the functional and the matrix styles of organization
structures. At the same time, it maximizes information flow between and across
functional areas in the organization. This is a logical structure for a BA group in its
advisory role to the organization
Figure: Centralized BA department, project, or team organization structure
Dr. G. Sindhu, Professor, PPG Business School
Given the advocacy and logic recommending a centralized BA grouping, there are
reasons for all BA groupings to be centralized. These reasons help explain why BA
initiatives that seek to integrate and align BA resources in to any type of BA group
within the organization sometimes fail. The listing in Table is not exhaustive, but it
provides some of the important issues to consider in the process of structuring a BA
group.
Table: Reasons for BA Initiative and Organization Failure
2.4 ROLE OF TECHNOLOGY
Technology is the backbone of Business Analytics, enabling organizations to collect,
process, analyze, and interpret large volumes of data to support evidence-based
decision-making. In MBA curricula (Anna University–aligned), the role of technology
can be explained across the analytics lifecycle—from data acquisition to strategic
action.
a. Data Collection & Integration
Dr. G. Sindhu, Professor, PPG Business School
● Enterprise systems (ERP, CRM, SCM) and digital channels (web, mobile, IoT)
generate structured and unstructured data.
● ETL tools and APIs integrate data from multiple sources into centralized
repositories.
● Ensures timely, comprehensive, and reliable datasets for analysis.
b. Data Storage & Management
● Data Warehouses support historical, structured data for reporting.
● Big Data platforms handle volume, velocity, and variety.
● Cloud storage provides scalability, flexibility, and cost efficiency.
c. Data Processing & Analytics
● Descriptive & Diagnostic analytics: SQL, OLAP, reporting tools.
● Predictive analytics: statistical models, machine learning algorithms.
● Prescriptive analytics: optimization and simulation tools for decision
recommendations.
● High-performance computing enables real-time and near–real-time analytics.
d. Visualization & Decision Support
● BI & visualization tools (Tableau, Power BI, Qlik) convert complex data into
intuitive dashboards.
● Supports management reporting, KPI tracking, and scenario analysis.
● Enhances data-driven culture across functional areas.
e. Automation & Artificial Intelligence
● AI/ML automates pattern recognition, forecasting, and anomaly detection.
● Robotic Process Automation (RPA) streamlines data preparation and reporting.
● Enables faster insights with reduced manual intervention.
f. Data Governance, Security & Quality
● Data governance frameworks define ownership, standards, and compliance.
● Security technologies ensure confidentiality, integrity, and availability of data.
● Data quality tools improve accuracy, consistency, and completeness—critical for
trustworthy analytics.
g. Strategic & Competitive Advantage
● Technology-driven analytics supports strategic planning, customer
personalization, risk management, and operational efficiency.
Dr. G. Sindhu, Professor, PPG Business School
● Organizations leveraging advanced analytics gain sustainable competitive
advantage through superior insights and agility.
2.5 INFORMATION POLICY
Information Policy in Business Analytics refers to the set of rules, guidelines, and
standards that govern how data is collected, stored, accessed, shared, protected, and
used within an organization to support ethical, legal, and effective analytics-driven
decision-making.
2.5.1 OBJECTIVES OF INFORMATION POLICY
● Ensure data accuracy, consistency, and reliability
● Protect confidential and sensitive information
● Enable controlled data sharing for analytics
● Ensure legal and regulatory compliance
● Support strategic and operational decision-making
2.6 ROLE OF INFORMATION POLICY IN BUSINESS ANALYTICS
a. Enables trustworthy analytics and insights
b. Reduces risk of data misuse and compliance violations
c. Supports ethical AI and responsible analytics
d. Improves decision quality and organizational transparency
e. Facilitates cross-functional data sharing
Information Policy: Defines rules and principles for data usage
Data Governance: Implements and enforces those rules through structures and processes
Information Policy is a critical foundation of Business Analytics. It ensures that analytics
activities are secure, ethical, high-quality, and compliant, enabling organizations to
extract maximum value from data while minimizing risk
Dr. G. Sindhu, Professor, PPG Business School
2.6 DATA GOVERNANCE
Data Governance in Business Analytics refers to the framework of policies, standards,
roles, responsibilities, and processes that ensure data used for analytics is accurate,
consistent, secure, and ethically managed. It provides control and accountability over
data assets across the organization, enabling reliable and trustworthy analytics.
Importance of Data Governance
● Ensures high-quality and reliable data for analytics
● Builds trust in analytical insights
● Supports regulatory and legal compliance
● Reduces data-related risks
● Enables organization-wide data-driven decision-making
Key Elements of Data Governance
a. Data Ownership
b. Data Quality Management
c. Data Policies and Standards
d. Data Security and Privacy
e. Metadata Management
Framework of Data Governance in BA
a. Governance committees or councils
b. Defined roles and responsibilities
c. Data standards and quality rules
d. Monitoring and audit mechanisms
e. Continuous improvement processes
Role of Data Governance in Analytics Lifecycle
● Data Collection: Ensures valid and authorized data sources
● Data Preparation: Maintains consistency and quality
● Data Analysis: Provides reliable inputs for models
● Reporting & Decision Making: Builds confidence in insights
Dr. G. Sindhu, Professor, PPG Business School
Challenges in Implementing Data Governance
● Organizational resistance and lack of awareness
● Complexity of managing large and diverse data sources
● Balancing control with flexibility
● High implementation and maintenance effort
2.7 DATA QUALITY MANAGEMENT
Data Quality Management refers to the systematic process of ensuring that data used
in business operations and analytics is accurate, complete, consistent, reliable, and
timely. High-quality data is essential for meaningful analysis, effective
decision-making, and organizational success.
2.7.1 DIMENSIONS OF DATA QUALITY
a. Accuracy
● Data correctly represents real-world values
● Example: Correct employee salary or customer address
b. Completeness
● All required data fields are filled
● Example: No missing values in critical records
c. Consistency
● Data is uniform across different systems
● Example: Same customer ID showing identical details in CRM and ERP
d. Timeliness
● Data is up to date and available when needed
● Example: Real-time sales data for decision-making
e. Validity
● Data follows defined rules and formats
● Example: Date fields following a standard format
f. Uniqueness
● No duplicate records exist
● Example: One unique record per employee or customer
Dr. G. Sindhu, Professor, PPG Business School
2.7.2 DATA QUALITY MANAGEMENT PROCESS
a. Data Profiling
● Examining data to identify errors, duplicates, and inconsistencies
b. Data Cleansing
● Correcting, standardizing, and removing inaccurate data
c. Data Validation
● Applying rules and constraints to ensure data correctness
d. Data Monitoring
● Continuous tracking of data quality metrics
e. Data Improvement
● Root cause analysis and process improvements
2.8 CHANGE MANAGEMENT IN BA
Change Management in Business Analytics refers to the structured approach used to
prepare, support, and guide people and organizations in adopting analytics-driven
processes, technologies, and decision-making practices. Since BA initiatives often alter
workflows, roles, and culture, effective change management is essential for successful
analytics implementation.
2.8.1 NEED FOR CHANGE MANAGEMENT IN BA
● Shift from intuition-based to data-driven decision-making
● Adoption of new analytics tools and technologies
● Redefinition of roles, skills, and responsibilities
● Resistance due to fear of transparency and job displacement
2.8.2 KEY ELEMENTS OF CHANGE MANAGEMENT IN BA
a. Leadership Commitment
● Top management sponsorship is critical
● Leaders must promote analytics vision and value
● Encourages organizational buy-in
b. Clear Analytics Vision & Strategy
● Define objectives of BA initiatives
Dr. G. Sindhu, Professor, PPG Business School
● Align analytics goals with business strategy
● Communicate expected benefits clearly
c. Stakeholder Involvement
● Engage users from HR, Finance, Marketing, Operations
● Encourage collaboration between business and analytics teams
● Reduces resistance and increases acceptance
d. Skill Development & Training
● Training in data literacy, BI tools, and analytics techniques
● Upskilling employees for analytics roles
● Builds confidence and competence
e. Process & Culture Change
● Embed analytics into daily workflows
● Promote a data-driven culture
● Reward evidence-based decision-making
f. Communication & Feedback
● Continuous communication on progress and outcomes
● Address concerns and misconceptions
● Use feedback to refine analytics solution
Dr. G. Sindhu, Professor, PPG Business School
UNIT III - DESCRIPTIVE ANALYTICS
Introduction to Descriptive analytics - Visualizing and Exploring Data - Descriptive
Statistics - Sampling Techniques and Estimation methods - Probability Distribution for
Descriptive Analytics - Use of dashboards and visualization tools
3. 1 INTRODUCTION TO DESCRIPTIVE ANALYTICS
Descriptive analytics is the foundational stage of data analytics that focuses on
summarizing and interpreting historical data to understand what has happened in
the past. It transforms raw data into meaningful information through aggregation,
visualization, and basic statistical techniques, enabling organizations to gain insights
into patterns, trends, and performance. The primary objective of descriptive analytics
is to answer the question: “What happened?” By analyzing past data,
decision-makers can assess outcomes, identify strengths and weaknesses, and
establish a factual basis for further analysis.
Descriptive analytics helps organizations understand past performance, monitor key
metrics, and communicate insights clearly to stakeholders. It serves as the base for
advanced analytics such as diagnostic, predictive, and prescriptive analytics.
Key Features of Descriptive Analytics
● Uses historical data
● Involves data summarization and reporting
● Employs tables, charts, graphs, and dashboards
● Relies on basic statistical measures such as mean, median, mode, percentages,
and frequency distributions
Dr. G. Sindhu, Professor, PPG Business School
3.2 VISUALIZING AND EXPLORING DATA
There is no single best way to explore a data set, but some way of
conceptualizing what the data set looks like is needed for this step of the BA process.
Charting is often employed to visualize what the data might reveal.
When determining the software options to generate charts in SPSS or Excel,
consider that each software can draft a variety of charts for the selected variables in
the data sets. Using the data in the above Table, charts can be created for the
illustrative sales data sets. Some of these charts are discussed in Table as a set of
exploratory tools that are helpful in understanding the informational value of data
sets. The chart to select depends on the objectives set for the chart.
Dr. G. Sindhu, Professor, PPG Business School
The charts presented in Table reveal interesting facts. The area chart is able to
clearly contrast the magnitude of the values in the two variable data sets (Sales 1 and
Sales 4). The column chart is useful in revealing the almost perfect linear trend in the
Sales 3 data, whereas the scatter chart reveals an almost perfect nonlinear function in
Sales 4 data. Additionally, the cluttered pie chart with 20 different percentages
illustrates that all charts can or should be used in some situations. The best practices
suggest charting should be viewed as an exploratory activity of BA. BA analysts
should run a variety of charts and see which ones reveal interesting and useful
information. Those charts can be further refined to drill down to more detailed
information and more appropriate charts related to the objectives of the BA initiative.
Dr. G. Sindhu, Professor, PPG Business School
Of course, a cursory review of the Sales 4 data in Figure makes the concave
appearance of the data in the scatter chart in Table unnecessary. But most BA problems
involve big data—so large as to make it impossible to just view it and make judgment
calls on structure or appearance. This is why descriptive statistics can be employed to
view the data in a parameter- based way in the hopes of better understanding the
information that the data has to reveal.
3.3 DESCRIPTIVE STATISTICS
When selecting the option of descriptive statistics in SPSS or Excel, a number of useful
statistics are automatically computed for the variables in the data sets. Some of these
descriptive statistics are discussed as exploratory tools that are helpful in
understanding the informational value of data sets.
Dr. G. Sindhu, Professor, PPG Business School
The main purpose of descriptive statistics is to describe what the data shows, making
complex data easier to interpret and communicate.
Components of Descriptive Statistics
1. Measures of Central Tendency
These measures indicate the average or central value of a dataset.
● Mean – Arithmetic average of all observations
● Median – Middle value when data is arranged in order
● Mode – Most frequently occurring value
2. Measures of Dispersion
These measures show the spread or variability of data.
● Range – Difference between the highest and lowest values
● Variance – Average of squared deviations from the mean
● Standard Deviation – Square root of variance, showing data spread around
the mean
3. Measures of Shape
These describe the distribution pattern of data.
● Skewness – Degree of asymmetry
● Kurtosis – Degree of peakedness or flatnes
Dr. G. Sindhu, Professor, PPG Business School
3.4 SAMPLING TECHNIQUES
Sampling is the process of selecting a subset (sample) from a population to
draw conclusions about the entire population. It saves time, cost, and effort while
maintaining acceptable accuracy.
SAMPLING METHODS
Sampling is an important strategy of handling large data. If data files are too
big to be run by software or just too large to work with, the number of items in the
data file can be sampled to provide a new data file that seeks to accurately represent
the population from which it comes. In sampling data, there are three components that
should be recognized: a population, a sample, and a sample element (the items that
make up the sample). A firm’s collection of customer service performance documents
for one year could be designated as a population of customer service performance for
that year.
Population: The population is the entire group that you want to draw conclusions
about.
Sample: The sample is the specific group of individuals that you will collect data from.
Sampling frame: The sampling frame is the actual list of individuals that the sample
will be drawn from. Ideally, it should include the entire target population
Sample Size: A sample size is a part of the population chosen for a survey or
experiment.
Dr. G. Sindhu, Professor, PPG Business School
Sampling Techniques
A. Probability Sampling Methods
Probability sampling is based on the fact that every member of a population has a
known and equal chance of being selected.
1. Simple random sampling
In this case, each individual is chosen entirely by chance and each member of
the population has an equal chance, or probability, of being selected. For
example, if you have a sampling frame of 1000 individuals, labelled 0 to 999,
use groups of three digits from the random number table to pick your sample.
So, if the first three numbers from the random number table were 094, select the
individual labelled “94”, and so on. Simple random sampling allows the
sampling error to be calculated and reduces selection bias. A specific advantage
is that it is the most straightforward method of probability sampling.
2. Systematic sampling
Individuals are selected at regular intervals from the sampling frame. The
intervals are chosen to ensure an adequate sample size. If you need a sample
size n from a population of size x, you should select every x/nth individual for
Dr. G. Sindhu, Professor, PPG Business School
the sample. For example, if you wanted a sample size of 100 from a population
of 1000, select every 1000/100 = 10th member of the sampling frame.
3. Stratified sampling: In this method, the population is first divided into
subgroups (or strata) who all share a similar characteristic. It is used when we
might reasonably expect the measurement of interest to vary between the
different subgroups, and we want to ensure representation from all the
subgroups. For example, in a study of stroke outcomes, we may stratify the
population by sex, to ensure equal representation of men and women.
4. Clustered sampling: In a clustered sample, subgroups of the population are
used as the sampling unit, rather than individuals. The population is divided
into subgroups, known as clusters which are randomly selected to be included
in the study. Clusters are usually already defined, for example individual
general practices or towns could be identified as clusters.
B. Non-Probability Sampling Methods
Non-probability sampling is a sampling method in which not all members of the
population have an equal chance of participating in the study.
1. Convenience sampling: Convenience sampling is perhaps the easiest method of
sampling, because participants are selected based on availability and
willingness to take part. Useful results can be obtained, but the results are
prone to significant bias, because those who volunteer to take part may be
different from those who choose not to (volunteer bias), and the sample may
not be representative of other characteristics, such as age or sex. Note:
volunteer bias is a risk of all non-probability sampling methods.
2. Quota sampling: This method of sampling is often used by market researchers.
Interviewers are given a quota of subjects of a specified type to attempt to
recruit. For example, an interviewer might be told to go out and select 20 adult
men, 20 adult women, 10 teenage girls and 10 teenage boys so that they could
interview them about their television viewing. Ideally the quotas chosen would
proportionally represent the characteristics of the underlying population.
Dr. G. Sindhu, Professor, PPG Business School
3. Judgement or Purposive Sampling: Also known as selective, or subjective,
sampling, this technique relies on the judgement of the researcher when
choosing who to ask to participate. Researchers may implicitly thus choose a
“representative” sample to suit their needs, or specifically approach individuals
with certain characteristics. This approach is often used by the media when
canvassing the public for opinions and in qualitative research. Judgement
sampling has the advantage of being time-and cost-effective.
4. Snowball sampling: This method is commonly used in social sciences when
investigating hard-to-reach groups. Existing subjects are asked to nominate
further subjects known to them, so the sample increases in size like a rolling
snowball. For example, when carrying out a survey of risk behaviours amongst
intravenous drug users, participants may be asked to refer other users to be
interviewed.
3.5 ESTIMATION METHODS
Estimation is a statistical technique used to approximate unknown population
parameters (mean, proportion, variance, etc.) based on sample data.
3.5.1 TYPES OF ESTIMATION METHODS
Estimation methods are broadly classified into:
a. Point Estimation
Point estimation is a statistical method in which a single numerical value
(point) is used to estimate an unknown population parameter based on sample
data. The value obtained from the sample is called a point estimator, and the
calculated value is called a point estimate.
Common Point Estimators
Population Parameter Point Estimator Symbol
Population Mean Sample Mean x̄
Population Proportion Sample Proportion p
Dr. G. Sindhu, Professor, PPG Business School
Population Variance Sample Variance s²
Population Standard Deviation Sample SD s
Examples of Point Estimation
1. Mean Estimation
If the average salary of a sample of employees is ₹40,000, then ₹40,000 is the
point estimate of the population mean salary.
2. Proportion Estimation
If 60 out of 100 customers are satisfied, then
p=60/100 =0.6
Here, 0.6 is the point estimate of the population proportion.
3. Variance Estimation
Sample variance (s²) is used to estimate population variance (σ²).
b. Interval Estimation
Interval estimation is a statistical method in which an unknown
population parameter is estimated using a range of values (interval) rather
than a single value. This interval is called a confidence interval, and it is
associated with a confidence level (such as 90%, 95%, or 99%).
Confidence Interval: It consists of Lower limit, Upper limit, Confidence
level
Types of Interval Estimation
a. Interval Estimation of Population Mean (σ Known) – Z Method
Used when:
● Population standard deviation (σ) is known
● Sample size is large (n ≥ 30)
Dr. G. Sindhu, Professor, PPG Business School
b. Interval Estimation of Population Mean (σ Unknown) – t Method
Used when:
Population standard deviation is unknown
Sample size is small (n < 30)
Where s, is sample standard deviation and t depends on degrees of
freedom (n − 1).
c. Interval Estimation of Population Proportion
Used to estimate population proportion (P).
Where: p = sample proportion, q = 1 – p
3.6 PROBABILITY DISTRIBUTION FOR DESCRIPTIVE ANALYTICS
A probability distribution describes how the values of a random variable are
distributed and the likelihood of occurrence of each value. In descriptive
analytics, probability distributions help summarize data behavior, identify
patterns, and understand uncertainty by explaining how often values occur and
how data is spread.
Probability distribution answers: “How are data values distributed?”
Types of Probability Distributions Used in Descriptive Analytics
1. Discrete Probability Distribution
a. Binomial Distribution
The Binomial Distribution is a discrete probability distribution that
describes the number of successes in a fixed number of independent trials,
where each trial has only two possible outcomes: success or failure. It is
Dr. G. Sindhu, Professor, PPG Business School
widely used in descriptive analytics to analyze situations involving
yes/no, pass/fail, or defective/non-defective outcomes.
Explanation: A random variable X follows a binomial distribution if it
represents the number of successes in ‘n’ independent trials, with a constant
probability of success ‘p’ in each trial.
Where,
● x = number of successes (0, 1, 2, …, n)
● n = total number of trials
● p = probability of success
● 1−p = probability of failure
Example
If a product has a 10% defect rate and 5 items are inspected, the
probability that exactly 2 items are defective is:
b. Poisson Distribution
The Poisson distribution is a discrete probability distribution used to describe
the number of times an event occurs within a fixed interval of time, space, or
area, when these events occur randomly and independently at a constant
average rate. It is widely used in descriptive analytics to model count-based
data and understand event frequency.
Explanation:
A random variable X follows a Poisson distribution if it represents the
number of occurrences of an event in a given interval and the events occur
independently with a known average rate (λ).
Dr. G. Sindhu, Professor, PPG Business School
Where,
● x = number of occurrences (0, 1, 2, …)
● λ = average number of occurrences in the interval
● e = Euler’s constant (≈ 2.718)
Example:
If a call center receives an average of 4 calls per minute, the probability of
receiving exactly 2 calls in a minute is:
2. Continuous Probability Distribution
a. Normal Distribution
The Normal Distribution, also known as the Gaussian Distribution, is a
continuous probability distribution that is widely used in descriptive
analytics to represent real-world data that clusters around a central
value. It is characterized by a symmetrical, bell-shaped curve, making
it one of the most important distributions in statistics and business
analytics.
Key Characteristics
● Bell-shaped and symmetrical curve
● Mean = Median = Mode
● Defined by two parameters: μ (mean) and σ (standard deviation)
● Total area under the curve equals 1
● Curve extends infinitely in both directions
Dr. G. Sindhu, Professor, PPG Business School
b. Uniform Distribution
A Uniform Distribution is a probability distribution in which all
possible outcomes are equally likely to occur. It represents complete
fairness or randomness within a specified range.
Key Characteristics
● Equal likelihood: Every value has the same probability
● Constant PDF: Height of the graph remains the same
● Mean: (a + b) / 2
● Variance: (b - a)2/12
● Symmetrical distribution
c. Exponential Distribution
The Exponential Distribution is a continuous probability distribution
used to model the time until an event occurs, assuming the event
happens independently and at a constant average rate. It is commonly
Dr. G. Sindhu, Professor, PPG Business School
used to study waiting time, failure time, or arrival time in descriptive
and business analytics.
Key Characteristics
● Continuous distribution
● Right-skewed curve
● Defined only for non-negative values
3.7 USE OF DASHBOARDS AND VISUALIZATION TOOLS
Dashboards and visualization tools are business analytics instruments used to
convert raw data into visual formats such as charts, graphs, tables, and KPIs. They
help managers and decision-makers understand trends, patterns, and performance
quickly and effectively.
3.7.1 DASHBOARD
A dashboard is a single-screen visual display that presents key performance
indicators (KPIs) and critical metrics required to monitor business performance.
Types of Dashboards
1. Strategic Dashboards
o Used by top management
o Focus on long-term goals and performance
o Example: Revenue growth, market share
Dr. G. Sindhu, Professor, PPG Business School
2. Operational Dashboards
o Used by middle-level managers
o Monitor day-to-day operations
o Example: Daily sales, inventory levels
3. Analytical Dashboards
o Used by analysts
o Allow drill-down and interactive analysis
o Example: Trend analysis, forecasting
3.7.2 VISUALIZATION TOOLS
Visualization tools are software applications that graphically represent data to
improve understanding and insight.
Common Visualization Tools
● Tableau
● Microsoft Power BI
● Google Data Studio (Looker Studio)
● Qlik Sense
● Excel (Advanced Charts & Pivot Tables)
Uses of Dashboards and Visualization Tools
1. Simplifies Complex Data
● Converts large datasets into easy-to-understand visuals
● Reduces cognitive load on decision-makers
2. Improves Decision-Making
● Enables faster and data-driven decisions
● Highlights trends, outliers, and exceptions
3. Real-Time Performance Monitoring
● Tracks KPIs in real time
● Helps organizations respond quickly to issues
4. Enhances Communication
Dr. G. Sindhu, Professor, PPG Business School
● Visuals communicate insights better than tables
● Useful for presentations to stakeholders
5. Identifies Trends and Patterns
● Helps detect seasonal patterns, growth trends, and risks
● Supports forecasting and planning
6. Increases Operational Efficiency
● Helps identify bottlenecks and inefficiencies
● Supports process improvement initiatives
Dr. G. Sindhu, Professor, PPG Business School
UNIT IV
PREDICTIVE ANALYTICS AND DATA MINING
Introduction to Predictive analytics – Predictive Models - Data Driven vs Logic Driven
Models - Predictive Modeling and procedure – Supervised vs. Unsupervised Learning
– Data Mining Techniques – Model Validation – Application in Business cases
4.1INTRODUCTION TO PREDICTIVE ANALYTICS
An application of advanced statistical information software, or operations
research methods to identify predictive variables and build predictive models to
identify trends and relationships not readily observed in a descriptive analysis.
Ex: Multiple regressions are used to show the relationship (or lack of relationship)
between age, weight and exercise on diet food sales. Knowing that the relationships
exist helps explain why one set of independent variables influences dependent
variables such as business performance.
Picture a situation in which big data files are available from a firm’s sales and
customer information (responses to differing types of advertisements, customer
surveys on product quality, customer surveys on supply chain performance, sale
prices, and so on). Assume also that a previous descriptive analytic analysis suggests
there is a relationship between certain customer variables, but there is a need to
precisely establish a quantitative relationship between sales and customer behavior.
Satisfying this need requires exploration into the big data to first establish whether a
measurable, quantitative relationship does in fact exist and then develop a statistically
valid model in which to predict future events. This is what the predictive analytics step
in BA seeks to achieve.
Whatever methodology is used, the identification of future trends or forecasts is
the principle output of the predictive analytics step in the BA process.
4.2 PREDICTIVE MODELING
Dr. G. Sindhu, Professor, PPG Business School
Predictive modeling means developing models that can be used to
forecast or predict future events. In business analytics, models can be developed
based on logic or data.
TYPES OF PREDICTIVE MODELING
1. Logic-Driven Models
A logic-driven model is one based on experience, knowledge, and logical
relationships of variables and constants connected to the desired business
performance outcome situation. The question here is how to put variables and
constants together to create a model that can predict the future. Doing this
requires business experience. Model building requires an understanding of
business systems and the relationships of variables and constants that seek to
generate a desirable business performance outcome. To help conceptualize the
relationships inherent in a business system, diagramming methods can be
helpful. For example, the cause-and-effect diagram is a visual aid diagram that
permits a user to hypothesize relationships between potential causes of an
outcome.
Dr. G. Sindhu, Professor, PPG Business School
This diagram lists potential causes in terms of human, technology, policy, and
process resources in an effort to establish some basic relationships that impact
business performance. The diagram is used by tracing contributing and relational
factors from the desired business performance goal back to possible causes, thus
allowing the user to better picture sources of potential causes that could affect the
performance. This diagram is sometimes referred to as a fishbone diagram because
of its appearance.
Another useful diagram to conceptualize potential relationships with business
performance variables is called the influence diagram. According to Evans,
influence diagrams can be useful to conceptualize the relationships of variables
in the development of models. It maps the relationship of variables and a
constant to the desired business performance outcome of profit. From such a
diagram, it is easy to convert the information into a quantitative model with
constants and variables that define profit in this situation:
Profit = Revenue − Cost, or
Profit = (Unit Price × Quantity Sold) − [(Fixed Cost) + (Variable Cost × Quantity
Sold)], or
P = (UP × QS) − [FC + (VC × QS)]
Dr. G. Sindhu, Professor, PPG Business School
The relationships in this simple example are based on fundamental business
knowledge. Consider, however, how complex cost functions might become
without some idea of how they are mapped together. It is necessary to be
knowledgeable about the business systems being modeled in order to capture the
relevant business behavior. Cause-and-effect diagrams and influence diagrams
provide tools to conceptualize relationships, variables, and constants, but it often
takes many other methodologies to explore and develop predictive models.
2. Data-Driven Models
Logic-driven modeling is often used as a first step to establish relationships
through data-driven models (using data collected from many sources to
quantitatively establish model relationships).
a. Sampling and Estimation: Generate statistical confidence intervals to define
limitations and boundaries on future forecasts for other forecasting models
b. Regression Analysis: Creates a predictive equation useful for forecasting time
series forecasts. It helps to weed out predictive variables in forecasting
models that add little to predicting values. It also generates a trend line for
forecasting.
c. Correlation Analysis: The analysis assesses the relationships between
variables. It also helps to weed out predictive variables in forecasting models
that add little to predicting values.
d. Probability Distribution: This analysis estimates the trend behavior that
follows certain types of probability distributions. It conducts statistical tests to
confirm significance of variables.
e. Predictive Modeling & Analysis: Predictive modeling fits linear and
nonlinear models to data to use the models for forecasting.
f. Forecasting Models: Forecasting models are one of the many tools businesses use
to predict outcomes regarding sales, supply and demand, consumer behavior and
more. Various models like time series model, econometric models, smoothing
models can be used to forecast values.
g. Simulation: It projects the future behavior in variables by simulating the past
behavior found in probability distributions.
Dr. G. Sindhu, Professor, PPG Business School
3. Other Models
a. Discriminant analysis is similar to a multiple regression model except
that it permits continuous independent variables and a categorical
dependent variable. The analysis generates a regression function whereby
values of the independent variables can be incorporated to generate a
predicted value for the dependent variable.
b. Logistic regression is also like multiple regression. Like discriminant
analysis, its dependent variable can be categorical. The independent
variables, though, in logistic regression can be either continuous or
categorical. For example, in predicting potential outsource providers, a
firm might use a logistic regression, in which the dependent variable
would be to either classify an outsource provider as rejected (represented
by the value of the dependent variable being zero) or classify the
outsource provider as acceptable (represented by the value of one for the
dependent variable).
c. Hierarchical clustering is a methodology that establishes a hierarchy of
clusters that can be grouped by the hierarchy. Two strategies are suggested
for this methodology: agglomerative and divisive. The agglomerative
strategy is a bottom-up approach, where one starts with each item in the
data and begins to group them. The divisive strategy is a top-down
approach, where one starts with all the items in one group and divides the
group into clusters. One method commonly used is to employ a Euclidean
distance formula that looks at the square root of the sum of distances
between two variables, their differences squared. Basically, the formula
seeks to match up variable candidates that have the least squared error
differences. (In other words, they’re closer together.)
Dr. G. Sindhu, Professor, PPG Business School
d. K-mean clustering is a classification methodology that permits a set of
data to be reclassified into K groups, where K can be set as the number of
groups desired. The algorithmic process identifies initial candidates for
the K groups and then interactively searches other candidates in the data
set to be averaged into a mean value that represents a particular K group.
The process of selection is based on maximizing the distance from the
initial K candidates selected in the initial run through the list. Each run or
iteration through the data set allows the software to select further
candidates for each group. The K-mean clustering process provides a
quick way to classify data into differentiated groups. To illustrate this
process, use the sales data in Figure and assume these are sales from
individual customers. Suppose a company wants to classify the sales
customers into high and low sales groups.
Dr. G. Sindhu, Professor, PPG Business School
Figure: Sales data for cluster classification problem
The SPSS K-Mean cluster software can be found in Analyze > Classify >
K-Means Cluster Analysis. Any integer value can designate the K number of
clusters desired. In this problem set, K=2. The SPSS printout of this classification
process is shown in Table. The solution is referred to as a Quick Cluster because
it initially selects the first two high and low values. The Initial Cluster Centers
table listed the initial high (20167) and a low (12369) value from the data set as
the clustering process begins. As it turns out, the software divided the customers
into nine high sales customers with a group mean sales of 18,309 and eleven low
sales customers with a group mean sales of 14,503.
Dr. G. Sindhu, Professor, PPG Business School
Table: SPSS K-Mean Cluster Solution
Consider how large big data sets can be. Then realize this kind of
classification capability can be a useful tool for identifying and predicting sales
based on the mean values.
There are so many BA methodologies that no single section, chapter, or
even book can explain or contain them all. The analytic treatment and computer
usage in this chapter have been focused mainly on conceptual use.
4.3 PREDICTIVE MODELLING PROCEDURES
Predictive modelling is a statistical and analytical process that uses historical
data, statistical techniques, and machine learning algorithms to predict future
outcomes or trends. It is a core component of Business Analytics and Decision
Science.
Predictive Modelling Procedures
1. Problem Definition
● Clearly define the objective of prediction
● Identify the dependent (target) variable
Dr. G. Sindhu, Professor, PPG Business School
● Decide how predictions will be used
● Example: Predict employee attrition, sales demand, or customer churn.
2. Data Collection
● Gather relevant historical and current data
● Sources include databases, surveys, ERP systems, CRM, HRMS
● Key requirement: Data should be reliable, relevant, and sufficient.
3. Data Preparation (Data Pre-processing)
● Handling missing values
● Removing outliers
● Data normalization and transformation
● Encoding categorical variables
● Data integration
4. Exploratory Data Analysis (EDA)
● Understand data patterns and relationships
● Use descriptive statistics and visualization
● Identify correlations and trends
● Tools used: Charts, scatter plots, correlation matrices
5. Feature Selection and Engineering
● Select relevant independent variables
● Create new meaningful features
● Remove redundant or irrelevant variables
● Purpose: Improve model performance and interpretability.
6. Model Selection
● Choose appropriate predictive techniques based on data type and
objective.
● Common Predictive Models:
a. Linear Regression
b. Logistic Regression
c. Decision Trees
Dr. G. Sindhu, Professor, PPG Business School
d. Random Forest
e. Time Series Models
f. Neural Networks
7. Model Building (Training)
● Split data into training and testing sets
● Fit the model using training data
● Estimate model parameters
8. Model Validation and Evaluation
● Evaluate accuracy and reliability using test data.
● Common Evaluation Metrics:
a. Accuracy
b. Precision and Recall
c. RMSE (Root Mean Square Error)
d. R² value
e. Confusion Matrix
9. Model Deployment
● Implement the model in real business environment
● Integrate with dashboards or decision systems
● Use predictions for planning and strategy
10.Monitoring and Model Updating
● Continuously monitor model performance
● Update model with new data
● Rebuild if accuracy declines
Dr. G. Sindhu, Professor, PPG Business School
4.4 SUPERVISED VS UNSUPERVISED LEARNING
Supervised and unsupervised learning are two major categories of machine learning
techniques used in predictive analytics and business analytics to discover patterns
and make decisions from data.
4.4.1 SUPERVISED LEARNING
Supervised learning is a machine learning approach where the model is trained using
labeled data, i.e., data that contains both input variables (features) and known output
(target variable).
Types of Supervised Learning
1. Regression – Predicts continuous values
Example: Sales forecasting, salary prediction
2. Classification – Predicts categorical outcomes
Example: Customer churn (Yes/No), employee attrition
Common Supervised Learning Techniques
● Linear Regression
● Logistic Regression
● Decision Trees
● Random Forest
● Support Vector Machines (SVM)
● Neural Networks
4.4.2 UNSUPERVISED LEARNING
Unsupervised learning is a machine learning approach where the model works with
unlabeled data, i.e., no predefined output variable is provided. The goal is to identify
hidden patterns or structures in the data.
Dr. G. Sindhu, Professor, PPG Business School
Common Unsupervised Learning Techniques
● K-Means Clustering
● Hierarchical Clustering
● DBSCAN
● Principal Component Analysis (PCA)
● Association Rule Mining (Apriori)
4.4.3 SUPERVISED VS UNSUPERVISED LEARNING
Supervised Learning Supervised Learning Supervised Learning
Data Type Labeled data Unlabeled data
Target Variable Present Absent
Objective Prediction Pattern discovery
Output Known Unknown
Accuracy Measurement Yes No direct measure
Complexity Lower Higher
Examples Regression, Classification Clustering, PCA
Limitations:
a. Supervised Learning
● Requires large labeled datasets
● Time-consuming data preparation
● Risk of overfitting
b. Unsupervised Learning
● Difficult to validate results
● Interpretation can be complex
● Results may vary based on parameters
Dr. G. Sindhu, Professor, PPG Business School
4.5DATA MINING
Data mining is the process of sorting through large data sets to identify patterns and
relationships that can help solve business problems through data analysis.
4.5.4 DATA MINING PROCESS
The data mining process can be broken down into these four primary stages:
● Data gathering: Relevant data for an analytics application is identified and
collected. The data may be located in different source systems, a data warehouse
or a data lake, an increasingly common repository in big data environments that
contain a mix of structured and unstructured data.
● Data preparation: This stage includes a set of steps to get the data ready to be
mined. It starts with data exploration, profiling and pre-processing, followed by
data cleansing work to fix errors and other data quality issues. Data
transformation is also done to make data sets consistent for a particular
application.
● Mining the data: Once the data is prepared, a data scientist chooses the
appropriate data mining technique and then implements one or more algorithms
to do the mining. In machine learning applications, the algorithms typically must
be trained on sample data sets to look for the information being sought before
they're run against the full set of data.
● Data analysis and interpretation: The data mining results are used to create
analytical models that can help drive decision-making and other business actions.
The data scientist or another member of a data science team also must
communicate the findings to business executives and users, often through data
visualization and the use of data storytelling techniques.
Dr. G. Sindhu, Professor, PPG Business School
4.5.2 DATA MINING TECHNIQUES
a. Association: It is used to find a correlation between two or more items by
identifying the hidden pattern in the data set and hence also called relation
analysis.
b. Classification: This data mining method is used to distinguish the items in the
data sets into classes or groups. It helps to predict the behaviour of entities
within the group accurately.
c. Clustering Analysis: Clustering is almost similar to classification, but depends
on the similarities of data items. Different groups have dissimilar or unrelated
objects. It is also called data segmentation as it partitions into groups based on
similarities.
d. Prediction: Prediction is used to predict the future based on the past and
present trends or data set. Prediction is mostly used to combine other mining
methods such as classification, pattern matching, trend analysis, and relation.
e. Sequential Patterns or Pattern Tracking: It is used to identify patterns that
frequently occur over a certain period of time.
Dr. G. Sindhu, Professor, PPG Business School
f. Decision Trees: Decision tree is a tree structure (as its name suggests), where
a. Each internal node represents a test on the attribute.
b. Branch denotes the result of the test.
c. Terminal nodes hold the class label.
d. The topmost node is the root node which has a simple question that has two
or more answers. Accordingly, the tree grows, and a flow chart like structure
is generated.
Dr. G. Sindhu, Professor, PPG Business School
g. Outlier Analysis or Anomaly Analysis: It identifies the data items that do not
comply with the expected pattern or expected behavior. These unexpected data
items are considered as outliers or noise.
h. Neural Network: It is based on biological neural networks. It is a collection of
neurons like processing units with weighted connections between them. They
are used to model the relationship between inputs and outputs.
4.6 MODEL VALIDATION TECHNIQUES
Model Validation is the process of evaluating how well a predictive model
performs on new data. It ensures that the model is accurate, reliable, and
generalizable, and not just fitting the training data.
Dr. G. Sindhu, Professor, PPG Business School
Objectives of Model Validation
● Check prediction accuracy
● Detect overfitting or underfitting
● Ensure model stability and reliability
● Select the best-performing model
4.6.1 VALIDATION TECHNIQUES
Validation is the process of assessing how well the mining models perform against real
data. Data validation is the practice of checking the integrity, accuracy and structure of
data before it is used for a business operation.
a. Hold-out
The hold-out method involves splitting the data into multiple parts and using one
part for training the model and the rest for validating and testing it.
Model evaluation using the hold-out method entails splitting the dataset into
training and test datasets, evaluating model performance, and determining the
most optimal model. This diagram illustrates the hold-out method for model
evaluation.
Dr. G. Sindhu, Professor, PPG Business School
Hold-out method for model evaluation
There are two parts to the dataset in the diagram above. One split is held aside as a
training set. Another set is held back for testing or evaluation of the model. The
percentage of the split is determined based on the amount of training data
available. A typical split of 70–30% is used in which 70% of the dataset is used for
training and 30% is used for testing the model.
The objective of this technique is to select the best model based on its accuracy on
the testing dataset and compare it with other models. There is, however, the
possibility that the model can be well fitted to the test data using this technique. In
other words, models are trained to improve model accuracy on test datasets based
on the assumption that the test dataset represents the population. As a result, the
test error becomes an optimistic estimation of the generalization error. Obviously,
this is not what we want. Since the final model is trained to fit well (or overfit) the
test data, it won’t generalize well to unknowns or future datasets.
b. K-fold cross-validation
K-fold cross-validation is a technique for evaluating predictive models. The dataset
is divided into k subsets or folds. The model is trained and evaluated k times,
using a different fold as the validation set each time. Performance metrics from
each fold are averaged to estimate the model’s generalization performance. This
Dr. G. Sindhu, Professor, PPG Business School
method aids in model assessment, selection, and hyper parameter tuning,
providing a more reliable measure of a model’s effectiveness.
Lets have a generalized k value, if K=5, it means, the given dataset has to split into
5 folds. During each run, one fold is considered for testing and the rest will be for
training and moving on with iterations.
c. LOOCV
The Leave-One-Out Cross-Validation, or LOOCV, procedure is used to
estimate the performance of machine learning algorithms when they are
used to make predictions on data not used to train the model.
It is a computationally expensive procedure to perform, although it results
in a reliable and unbiased estimate of model performance. Although simple
to use and no configuration to specify, there are times when the procedure
should not be used, such as when you have a very large dataset or a
computationally expensive model to evaluate.
d. Bootstraping
Bootstrapping is a resampling technique used in predictive analytics and
statistics where multiple samples are drawn from the original dataset with
replacement. It helps estimate the accuracy, variability, and stability of a
predictive model, especially when the dataset is small.
Dr. G. Sindhu, Professor, PPG Business School
Process:
● Start with an original dataset of size n
● Draw a bootstrap sample of size n with replacement
● Some observations may appear multiple times, while others may not
appear
● Train the model on the bootstrap sample
● Test the model on the out-of-bag (OOB) data
● Repeat the process many times (e.g., 500 or 1000 samples)
● Average the performance results
4.7 APPLICATION IN BUSINESS CASES
a. Marketing Analytics
● Customer Segmentation: Predictive models identify profitable customer
segments for targeted campaigns.
● Churn Prediction: Anticipates which customers may leave so retention
strategies can be applied.
● Campaign Optimization: Estimates customer response to marketing
campaigns, improving ROI.
Example: Telecom companies predict churn to offer personalized retention
plans.
b. Finance and Banking
● Credit Risk Scoring: Predicts likelihood of loan defaults to reduce financial risk.
● Fraud Detection: Detects unusual patterns in transactions to prevent fraud.
● Portfolio Management: Forecasts asset performance and manages investment
risk.
Example: Banks use predictive models to approve credit applications with
higher accuracy.
c. Human Resource Management
Dr. G. Sindhu, Professor, PPG Business School
● Employee Attrition Prediction: Identifies employees at risk of leaving to
implement retention strategies.
● Performance Forecasting: Predicts employee performance for promotions,
training, or appraisals.
Example: Organizations use predictive analytics to plan talent retention
programs.
d. Sales and Demand Forecasting
● Sales Prediction: Forecasts future sales based on historical trends.
● Inventory Optimization: Ensures the right stock levels, reducing
understocking or overstocking.
● Pricing Strategy: Determines optimal pricing to maximize revenue.
Example: Retailers predict seasonal demand to manage inventory efficiently.
e. Operations and Supply Chain
● Predictive Maintenance: Forecasts equipment failures to reduce downtime.
● Supply Chain Optimization: Anticipates delays and bottlenecks, improving
efficiency.
Example: Manufacturing firms predict machine breakdowns to schedule
maintenance proactively.
f. Healthcare and Insurance
● Risk Prediction: Estimates patient risk for diseases or insurance claims.
● Treatment Effectiveness: Predicts outcomes for personalized treatment plans.
Example: Health insurers predict high-risk claims to set premiums or design
preventive plans.
Dr. G. Sindhu, Professor, PPG Business School
UNIT V
PRESCRITIVE ANALYTICS
Introduction to Prescriptive analytics - Prescriptive Modeling - Non-Linear
Optimization - Demonstrating Business Performance Improvement
5.1 INTRODUCTION TO PRESCRIPTIVE ANALYTICS
Prescriptive analytics is an advanced analytics technique that recommends
optimal actions based on predictive insights, business constraints, and
organizational objectives to achieve the best possible outcomes.
After undertaking the descriptive and predictive analytics steps in the BA
process, one should be positioned to undertake the final step: prescriptive
analytics analysis. The prior analysis should provide a forecast or prediction of
what future trends in the business may hold. For example, there may be
significant statistical measures of increased (or decreased) sales, profitability
trends accurately measured in dollars for new market opportunities, or measured
cost savings from a future joint venture.
If a firm knows where the future lies by forecasting trends, it can best plan
to take advantage of possible opportunities that the trends may offer. Step 3 of
the BA process, prescriptive analytics, involves the application of decision
science, management science, or operations research methodologies to make best
use of allocable resources. These are mathematically based methodologies and
algorithms designed to take variables and other parameters into a quantitative
framework and generate an optimal or near- optimal solution to complex
problems. These methodologies can be used to optimally allocate a firm’s limited
resources to take best advantage of the opportunities it has found in the
predicted future trends. Limits on human, technology, and financial resources
prevent any firm from going after all the opportunities. Using prescriptive
Dr. G. Sindhu, Professor, PPG Business School
analytics allows the firm to allocate limited resources to optimally or
near-optimally achieve the objectives as fully as possible.
Figure: Prescriptive analytic methodologies
5.2 PRESCRIPTIVE MODELING
Prescriptive Modeling is a type of advanced analytics that goes beyond
descriptive and predictive analytics. It recommends the best course of action to
achieve a desired outcome, based on predictions and business objectives.
Prescriptive analytics model businesses while taking into account all inputs,
processes and outputs. Models are calibrated and validated to ensure they
accurately reflect business processes. Prescriptive analytics recommend the best
way forward with actionable information to maximize overall returns and
profitability.
Dr. G. Sindhu, Professor, PPG Business School
5.2.1 TYPES OF PRESCRIPTIVE MODELS
● Optimization Models: Linear programming, integer programming, goal
programming. Example: Minimizing production cost or maximizing profit
a. Linear Programming (LP)
b. Integer Programming (IP)
c. Non-linear Programming
d. Goal Programming
● Simulation Models: Used to study the behavior of systems under uncertainty by
experimenting with different scenarios. Example: Inventory demand forecasting
under uncertain conditions
a. Monte Carlo Simulation
b. What-if Analysis
● Decision Analysis Models: Used to choose the best alternative under risk or
uncertainty. Example: Choosing between investment alternatives.
a. Decision Trees
b. Payoff Tables
c. Expected Monetary Value (EMV)
● Heuristic and Metaheuristic Models: Used for complex problems where
exact solutions are difficult. Solve complex problems where exact solutions are
computationally expensive.
a. Genetic Algorithms
b. Simulated Annealing
c. Tabu Search
Dr. G. Sindhu, Professor, PPG Business School
5.2.2 STEPS INVOLVED IN BUILDING A PRESCRIPTIVE MODEL
Prescriptive modeling provides optimal decision recommendations by using
mathematical and analytical techniques. The systematic steps involved are:
1. Problem Identification and Definition
● Clearly define the business problem, objectives, decision scope, and
constraints. Example: Minimizing production cost or maximizing profit.
2. Data Collection and Preparation
● Gather relevant internal and external data. Clean, organize, and validate the
data for accuracy and completeness.
3. Model Formulation
● Translate the problem into a mathematical or logical model by:
● Defining decision variables
● Stating the objective function
● Identifying constraints
4. Selection of Appropriate Technique
Choose suitable tools such as:
● Linear Programming
● Simulation
● Decision Trees
● Optimization Algorithms
5. Model Solution
● Apply computational methods or software (e.g., Excel Solver, SPSS, R,
Python) to obtain the optimal solution.
6. Model Validation
● Test the model using historical or sample data to ensure reliability and
accuracy of results.
7. Sensitivity and Scenario Analysis
● Analyze how changes in inputs affect outputs to assess risk and robustness
of decisions.
8. Implementation
Dr. G. Sindhu, Professor, PPG Business School
● Apply the recommended solution in real business operations.
9. Monitoring and Review
● Continuously track performance and update the model when business
conditions change.
5.2.3 APPLICATION OF PRESCRIPTIVE ANALYSIS
1. Financial Management
Banks are heavily relying on business intelligence solutions for financial
management today, in order to achieve substantial benefits of deploying currently
available predictive and prescriptive analytics tools such as machine learning
applications and big data analytics. Various business intelligence solutions are
integrated with these technologies that can enable banking institutions to make
data-driven decisions, especially while devising important business strategies
vis-a-vis financial management.
2. Fraud Prevention
Fraud prevention is one of the most critical and sensitive applications of predictive
and prescriptive analytics tools in the banking sector, mainly due to the rise of
cyber attacks and cyber security threats. Banks that are equipped with advanced
tools with predictive and prescriptive analytics can easily identify such problems.
This can ultimately help them accurately predict the potential threats and prevent
the breach of sensitive data and potential for fraud.
3. Application Screening
Banking institutions and many such financial bodies are often flooded with
innumerable applications. In the modern, technological era, banks are making use
of predictive and prescriptive analytics tools to process and analyze massive
volumes of applications without any delays or errors.
Dr. G. Sindhu, Professor, PPG Business School
4. Better Liquidity/Cash Planning
Banks are using predictive and prescriptive analytics, not only in financial
management applications, but also in ensuring better planning for liquidity or
availability of cash. As optimal management of liquid assets through business
intelligence can lead to more profitable outcomes, predictive and prescriptive
analytics tools are likely to gain more popularity in the finance sector.
5. Customer Acquisition and Retention
Predictive and prescriptive analytics tools make use of state-of-the-art, data-driven
technologies, such as Artificial Intelligence, machine learning, and Big Data, to
identify opportunities for customer acquisition and retention. Banks can leverage
business intelligence to run campaigns for customer retention and new customer
acquisition with more optimized targeting, which can help them easily spot
high-value customer segments that are more likely to respond to these campaigns.
6. Knowing Customer Buying Habits
Banking institutions and financial organizations, especially the ones that operate
independently, usually have to grapple with the problems associated with
introducing the right scheme for their target customers. Now, with the use of
AI-driven predictive and prescriptive analytics tools, the entire banking sector can
easily identify the potential changes in the behaviour of their target customer
group. This can ultimately result in more successful products and schemes, and
thereby, the positive growth of the banking sector.
7. Loan Approval
Banks are getting more sophisticated with the system they follow to evaluate load
applications. Taking into consideration the fact that some applications can be
approved despite them not having a high FICO, banking and financial institutions
are relying on advanced predictive and prescriptive analytics tools to reduce
unnecessary denial of loans, helping non-traditional borrowers get their loans
approved.
Dr. G. Sindhu, Professor, PPG Business School
8. Cross-selling
By analyzing the behaviours and purchasing trends of their existing customers,
banking and financial organizations are trying to spot opportunities where
multiple products and schemes can be pitched. Predictive and prescriptive
analytics tools can help these financial bodies to identify the potential for successful
and efficient cross-selling. This does not only contribute to the profitability of the
organization, but also helps in improving customer relationships, opening better
opportunities in the future.
9. Customer Lifetime Value (LTV)
Banking institutions, today, are struggling to ensure that their customers do not
bounce off to better schemes offered by their competitors, as the intensity of
competition is only increasing with time. With the help of predictive and
prescriptive analytics tools, banking organizations are streamlining their customer
engagement efforts, in order to achieve some wins based on customer lifetime
value.
10.Customer relationship management
This is one of the most practical applications of predictive and prescriptive
analytics tools, as they can precisely notify banks about how their customer
relationship management should be modified or improved. With the help of
various insights such as the list of customers that the banks must focus on with
better customer engagement efforts, how customers have responded to certain
promotions in the past, and how to get better returns on various customer
engagement activities.
11.Marketing: Email Automation
Email automation is a clear-cut example of prescriptive analytics at work.
Marketers use email automation to sort leads into categories based on their
motivations, mindsets, and intentions and deliver email content to them based on
those categories. Any interactions leads have with emails can put them in another
category, resulting in a different set of messages being triggered. Email automation
allows companies to provide personalized messaging at scale and increase the
Dr. G. Sindhu, Professor, PPG Business School
chance of converting a lead into a customer using content that applies to their
motivations and needs.
12.Product Management: Development and Improvement
Prescriptive analytics can also inform product development and improvements.
Product managers can gather user data by surveying customers, running tests with
a product’s beta versions, conducting market research with people who aren’t
current product users, and collecting behavioral data as current users interact. All
this data can be analyzed—either manually or algorithmically—to identify trends,
discover the reasons for those trends, and predict whether the trends are predicted
to recur.
13.Sales: Lead Scoring
Prescriptive analytics plays a prominent role in sales through lead scoring, also
called lead ranking. Lead scoring is the process of assigning a point value to
various actions along the sales funnel, enabling you, or an algorithm, to rank leads
based on how likely they are to convert into customers.
1. Actions you can assign value to include:
2. Page views
3. Email interactions
4. Site searches
5. Content engagement, such as attending webinars, downloading e-books, or
watching videos
14.Supply Chain & Logistics Management
Prescriptive Analytics is crucial for route optimization in the Supply Chain
industry. Logistic companies leverage it to prevent logistical issues like incorrect
shipping locations. They also rely on these analyses for better route planning at
lesser energy consumption while saving time & money.
Dr. G. Sindhu, Professor, PPG Business School
5.3 OPTIMIZATION MODEL
An Optimization Model is a type of prescriptive analytics technique used to
identify the best possible solution for problem under given constraints.
5.3.1 LINEAR OPTIMIZATION MODEL
Bright Furniture Company manufactures Tables and Chairs. Three processes are
required viz. Carpentry, Sizing and Polishing. The available time for carpentry, sizing
and polishing are 95 hrs, 35 hrs and 25 hrs respectively. Time taken for table and chair
are given. The profit for Table and Chair sold is Rs. 35 and Rs.42 respectively.
Determine how many chairs and table are to be manufactured to maximize profit.
LP model of the problem
Maximize Z = 35x1 + 42x2
Subject to constraints,
• 4x1+ 6x2 ≤ 95 (Carpentry)
• 2x1+ 2x2 ≤ 25 (Sizing)
• 2x1+ x2 ≤ 35 (Polishing)
where x1, x2 ≥ 0
Solving using computer (Tora)
Input Screen
Dr. G. Sindhu, Professor, PPG Business School
Enter the data as shown in the figure and enter solve and select Graphical
method and click Output Screen.
Result: Profit Z max = Rs. 700.00
No. of Tables to be produced, X1 = 5
No. of Chairs to be produced, X2 = 12.5 say 13
The same can be solved using Simplex method
Select Solve Problem – Algebraic – Iterations – All slack starting solution
Dr. G. Sindhu, Professor, PPG Business School
Click All-slack solution for Simpex Method,
Results can be read from the optimal table (ie., Iteration 3, last column).
The prescriptive analytics analysis step brings the prior statistical analytic steps
into an applied decision-making process where a potential business performance
improvement is shown to better this organization’s ability to use its resources more
effectively. The management job of monitoring performance and checking to see that
business performance is in fact improved is a needed final step in the BA analysis.
Dr. G. Sindhu, Professor, PPG Business School
Without proof that business performance is improved, it’s unlikely that BA would
continue to be used.
5.3.2 NONLINEAR OPTIMIZATION
Nonlinear optimization, when business performance cost or profit functions
become too complex for simple linear models to be useful, exploration of nonlinear
functions is a standard practice in BA. Although the predictive nature of exploring for
a mathematical expression to denote a trend or establish a forecast falls mainly in the
predictive analytics step of BA, the use of the nonlinear function to optimize a decision
can fall in the prescriptive analytics step.
As mentioned previously, there are many mathematical programing nonlinear
methodologies and solution procedures designed to generate optimal business
performance solutions. Most of them require careful estimation of parameters that may
or may not be accurate, particularly given the precision required of a solution that can
be so precariously dependent upon parameter accuracy. This precision is further
complicated in BA by the large data files that should be factored into the
model-building effort.
To overcome these limitations and be more inclusive in the use of large data,
regression software can be applied. As illustrated, Curve Fitting software can be used
to generate predictive analytic models that can also be utilized to aid in making
prescriptive analytic decisions.
For purposes of illustration, SPSS’s Curve Fitting software will be used.
Suppose that a resource allocation decision is being faced whereby one must decide
how many computer servers a service facility should purchase to optimize the firm’s
costs of running the facility. The firm’s predictive analytics effort has shown a growth
Dr. G. Sindhu, Professor, PPG Business School
trend. A new facility is called for if costs can be minimized. The firm has a history of
setting up large and small service facilities and has collected the 20 data points.
Whether there are 20 or 20,000 items in the data file, this SPSS function fits the
data based on regression mathematics to a nonlinear line that best minimizes the
distance from the data items to the line. The software then converts the line into a
mathematical expression useful for forecasting.
Figure: Data and SPSS Curve Fitting function selection window
In this server problem, the basic data has a u-shaped function, as presented. This
is a classic shape for most cost functions in business. In this problem, it represents the
balancing of having too few servers (resulting in a costly loss of customer business
through dissatisfaction and complaints with the service) or too many servers
(excessive waste in investment costs as a result of underutilized servers). Although this
Dr. G. Sindhu, Professor, PPG Business School
is an overly simplified example with little and nicely ordered data for clarity purposes,
in big data situations, cost functions are considerably less obvious.
Figure: Server problem basic data cost function
The first step in using the curve-fitting methodology is to generate the best- fitting
curve to the data. By selecting all the SPSS models, the software applies each point of
data using the regression process of minimizing distance from a line. The result is a
series of regression models and statistics, including ANOVA and other testing
statistics. It is known from the previous illustration of regression that the adjusted
R-Square statistic can reveal the best estimated relationship between the independent
(number of servers) and dependent (total cost) variables. These statistics are presented
in Table. The best adjusted R-Square value (the largest) occurs with the quadratic
model, followed by the cubic model. The more detailed supporting statistics for both
of these models are presented in Table. The graph for all the SPSS curve-fitting models
appears in Figure
Dr. G. Sindhu, Professor, PPG Business School
Table: Adjusted R-Square Values of All SPSS Models
Dr. G. Sindhu, Professor, PPG Business School
Table: Quadratic and Cubic Model SPSS Statistics
Dr. G. Sindhu, Professor, PPG Business School
Figure: Graph of all SPSS curve-fitting models
From Table, the resulting two statistically significant curve-fitted models follow:
Yp = 35417.772 − 5589.432 X + 268.445 X2 [Quadratic model]
Yp = 36133.696 − 5954.738 X + 310.895 X2 − 1.347 X3 [Cubic model]
where:
Yp = the forecasted or predicted total cost, and
X = can be the number of computer servers.
For purposes of illustration, we will use the quadratic model. In the next
step of using the curve-fitting models, one can either use calculus to derive the
cost minimizing value for X (number of servers) or perform a deterministic
simulation where values of X are substituted into the model to compute and
predict the total cost (Yp).
Dr. G. Sindhu, Professor, PPG Business School
As a simpler solution method to finding the optimal number of servers,
simulation can be used. Representing a deterministic simulation, the resulting costs of
servers can be computed using the quadratic model, as presented in Figure. These
values were computed by plugging the number of server values (1 to 20) into the Yp
quadratic function one at a time to generate the predicted values for each of the server
possibilities. Note that the lowest value in these predicted values occurs with the
acquisition of 10 servers at $6357.952, and the next lowest is at 11 servers at $6415.865.
In the actual data in Figure, the minimum total cost point occurs at 9 servers at $4533,
whereas the next lowest total cost is $4678 occurring at 10 servers. The differences are
due to the estimation process of curve fitting. Note in Figure that the curve that is
fitted does not touch the lowest 5 cost values. Like regression in general, it is an
estimation process, and although the ANOVA statistics in the quadratic model
demonstrate a strong relationship with the actual values, there is some error. This
process provides a near-optimal solution but does not guarantee one.
Figure: Predicted total cost in server problem for each server alternative
Like all regression models, curve fitting is an estimation process and has risks,
but the supporting statistics, like ANOVA, provide some degree of confidence in the
resulting solution.
Dr. G. Sindhu, Professor, PPG Business School
Finally, it must be mentioned that many other nonlinear optimization
methodologies exist. Some, like quadratic programming, are considered constrained
optimization models (like LP). Other methodologies, like the use of calculus in this
chapter, are useful in solving for optimal solutions in unconstrained problem settings.
Simulation
A simulation is the imitation of the operation of a real-world process or system
over time. Simulations require the use of models; the model represents the key
characteristics or behaviors of the selected system or process, whereas the simulation
represents the evolution of the model over time.
5.4DEMONSTRATING BUSINESS PERFORMANCE IMPROVEMENT
Business performance improvement refers to the systematic process of enhancing
an organization’s efficiency, effectiveness, and profitability through data-driven
decision-making, process optimization, and strategic initiatives. Example: A retail
company implemented inventory optimization:
● Stock-out rate reduced from 15% to 5%
● Holding cost reduced by 20%
● Customer satisfaction increased by 12%
This demonstrates measurable business performance improvement.
5.4.1 KEY AREAS OF IMPROVEMENT
1. Operational Efficiency – Reducing costs, improving productivity
2. Financial Performance – Increasing revenue, profitability, ROI
3. Customer Satisfaction – Improving service quality and retention
4. Employee Performance – Enhancing skills and engagement
5. Process Effectiveness – Streamlining workflows and reducing waste
Dr. G. Sindhu, Professor, PPG Business School
Methods to Demonstrate Performance Improvement
a. Key Performance Indicators (KPIs) : Sales growth, profit margin, customer
retention rate, productivity ratio
b. Benchmarking: Comparing performance with industry standards or best
competitors
c. Before-and-After Analysis: Measuring performance metrics before and after
improvement initiatives
d. Dashboards and Scorecards: Using Balanced Scorecard and business dashboards
for visual monitoring
e. Data Analytics and Reporting: Using descriptive, predictive, and prescriptive
analytics to drive decisions.
Dr. G. Sindhu, Professor, PPG Business School
UNIT VI
APPLICATIONS AND EMERGING TRENDS IN BA
Industry applications in Marketing, HR, Finance, and Operations – Real-time
analytics – AI & ML integration – Cloud-based BA tools – Business Intelligence
vs. BA – Ethics in Analytics – Case Studies on contemporary analytics usage.
6.1 INDUSTRY APPLICATIONS IN MARKETING, HR, FINANCE, AND
OPERATIONS
Business Analytics refers to the use of data, statistical tools, models, and
technologies to analyze business performance and support fact-based
decision-making.
6.1.1 APPLICATIONS OF BA IN MARKETING
a. Customer Segmentation – Grouping customers based on behavior and
demographics
b. Market Basket Analysis – Identifying product combinations for
cross-selling
c. Customer Lifetime Value (CLV) – Estimating long-term profitability
d. Pricing Analytics – Optimizing prices using demand data
e. Campaign Effectiveness Analysis – Measuring ROI of advertisements
f. Churn Prediction – Identifying customers likely to leave
6.1.2 APPLICATIONS OF BA IN HUMAN RESOURCES (HR)
a. Talent Acquisition Analytics – Predicting hiring success
b. Employee Performance Analysis – Identifying high performers
c. Attrition Prediction – Reducing employee turnover
d. Training Effectiveness Measurement – Evaluating skill development
programs
e. Workforce Planning – Forecasting manpower requirements
Dr. G. Sindhu, Professor, PPG Business School
f. Compensation Analytics – Designing fair and competitive pay structures
6.1.3. APPLICATIONS OF BA IN FINANCE
a. Financial Forecasting & Budgeting: Predict future financial performance
and allocate resources effectively.
b. Credit Risk Analysis: Assess applicants’ income, credit history, and
repayment behavior before approving personal or business loans.
c. Fraud Detection: Identifying unusual patterns or transactions that indicate
fraudulent activities.
d. Portfolio Optimization: Allocating investments across assets in a way that
maximizes returns while minimizing risk.
e. Cash Flow Management: To forecast, monitor, and optimize cash inflows
and outflows to ensure adequate liquidity and financial stability.
f. Cost Optimization: To identify, control, and reduce unnecessary expenses
while maintaining or improving operational efficiency and profitability.
6.1.4 APPLICATIONS OF BA IN OPERATIONS
a. Demand Forecasting: Predicting future customer demand for products or
services to support production and inventory planning.
b. Inventory Optimization: Inventory optimization is making sure a business
keeps the right amount of stock i.e enough to meet customer demand
without having too much or too little.
c. Production Planning: Planning and scheduling production to make the
right products on time, using resources efficiently.
d. Supply Chain Optimization: Deliver products faster, reduce costs, and meet
customer demand efficiently.
e. Quality Control Analytics: Using data to monitor and improve product
quality, detect defects, and ensure standards are consistently met.
f. Maintenance Scheduling (Predictive Maintenance): Planning and
organizing equipment maintenance to prevent breakdowns and ensure
smooth, uninterrupted operations
Dr. G. Sindhu, Professor, PPG Business School
6.2REAL-TIME ANALYTICS
Real-time analytics is the process of analyzing data immediately as it is
generated, allowing businesses to make instant, informed decisions. Unlike
traditional analytics, which works on historical data, real-time analytics works
on live data streams from operations, transactions, or customer interactions.
6.2.1 APPLICATIONS IN BUSINESS ANALYTICS
● Retail: Detecting sales trends or stock shortages instantly.
● Finance: Monitoring transactions for fraud in real-time.
● Manufacturing: Predictive maintenance alerts for machines.
● Customer Service: Providing instant support using live data on queries and
complaints
6.2.3 BENEFITS
● Faster decision-making and response to market changes.
● Improved operational efficiency.
● Enhanced customer experience.
● Reduced risks and losses due to immediate detection of issues.
6.3AI & ML INTEGRATION
AI (Artificial Intelligence) and ML (Machine Learning) integration in business
analytics involves using intelligent algorithms and models to analyze data,
make predictions, automate decisions, and uncover patterns that are not easily
visible to humans.
Dr. G. Sindhu, Professor, PPG Business School
6.3.1 KEY FEATURES
a. Predictive Analytics: ML models forecast trends, demand, or customer
behavior.
b. Automation: AI systems can automate repetitive tasks like data processing
or report generation.
c. Pattern Recognition: AI/ML can detect anomalies, fraud, or hidden insights
in large datasets.
d. Continuous Learning: ML algorithms improve over time as they process
more data.
6.3.2 APPLICATIONS IN BUSINESS ANALYTICS
● Sales & Marketing: Predicting customer preferences, churn, or targeted
offers.
● Finance: Detecting fraudulent transactions or credit risk assessment.
● Operations: Optimizing supply chain and inventory using predictive
models.
● Customer Service: AI chatbots and recommendation systems for real-time
assistance.
6.4CLOUD-BASED BA TOOLS
Cloud-based BA tools are software applications hosted on the cloud that allow
businesses to collect, store, analyze, and visualize data over the internet
without needing on-premise infrastructure.
Key Features:
a. Accessibility: Data and analytics can be accessed anytime, anywhere.
b. Scalability: Easily handles growing volumes of data without additional
hardware.
c. Collaboration: Teams can work together on dashboards and reports in
real-time.
d. Cost-Effective: Reduces IT infrastructure and maintenance costs.
Dr. G. Sindhu, Professor, PPG Business School
6.5BUSINESS INTELLIGENCE VS. BA
Business Analytics: BA uses statistical, predictive, and AI-based methods to analyze
data and forecast future outcomes. It answers “Why did it happen?” and “What will
happen next?”
Business Intelligence: BI focuses on analyzing past and present data to generate
reports, dashboards, and summaries that support operational decision-making. It
answers “What happened?” and “What is happening?”
Aspect Business Intelligence (BI) Business Analytics (BA)
Focus Past & present data Future predictions
Methods Reports, dashboards Statistics, ML, AI
Decision Type Descriptive Predictive & prescriptive
6.6ETHICS IN ANALYTICS
Ethics in analytics refers to the responsible collection, use, and protection of
data to ensure fairness, privacy, transparency, and accuracy. Ethics in analytics
refers to the responsible use of data to ensure privacy, fairness, transparency,
and accuracy in decision-making. It involves protecting personal information,
avoiding bias in models, ensuring data security, and using insights honestly
and responsibly.
Ethics in Analytics
a. Data Privacy: Personal and sensitive information must be collected lawfully
and used only for legitimate purposes, with proper consent.
b. Data Security: Organizations must protect data from unauthorized access,
breaches, and misuse through strong security measures.
c. Fairness & Bias Prevention: Analytical models should be free from
discrimination and bias that could unfairly affect individuals or groups.
d. Transparency & Explainability: The methods, data sources, and logic
behind analytical models should be understandable and clearly
communicated.
Dr. G. Sindhu, Professor, PPG Business School
e. Accuracy & Integrity: Data must be reliable, complete, and used honestly to
avoid misleading outcomes and decisions.
f. Accountability: Organizations must take responsibility for how data and
analytics outcomes are used in business decisions.
Importance of Ethics in Analytics
● Builds customer and stakeholder trust
● Ensures legal and regulatory compliance
● Prevents misuse of data and discrimination
● Improves the credibility of business decisions
6.7CASE STUDIES ON CONTEMPORARY ANALYTICS USAGE
Example: 1
Retail Sales Forecasting (Predictive Analytics)
Case:
A supermarket chain faces frequent stock outs during weekends and excess
inventory on weekdays.
Case Explanation:
The supermarket chain notices two major problems:
● Weekends: Products run out quickly (stock outs), causing loss of sales and
unhappy customers.
● Weekdays: Excess stock remains unsold, increasing storage costs and
wastage, especially for perishable items.
This happens because the store is not accurately predicting daily demand
patterns.
Analytics Used:
Time series forecasting techniques such as moving averages and regression
analysis.
Time series forecasting uses past sales data over time to identify:
● Trends (overall increase/decrease in sales)
● Seasonality (weekend vs weekday demand)
Dr. G. Sindhu, Professor, PPG Business School
● Patterns and fluctuations
Methods like moving averages smooth past data to identify trends, while
regression models predict future demand based on historical sales behavior.
Solution Explanation:
1. The supermarket collects daily sales data from previous months.
2. Forecasting models predict how much of each product will sell on each day.
3. Inventory orders are adjusted based on these predictions:
● More stock for weekends
● Less stock for weekdays
Outcome:
● Reduced stock outs
● Lower inventory holding costs
● Less product wastage
● Higher customer satisfaction and profits
Example 2:
Bank Credit Risk Assessment (Predictive Analytics)
Case:
A bank wants to reduce loan defaults while approving more genuine customers.
Analytics Used:
Classification models (Logistic Regression, Decision Trees).
Solution:
Customer credit history and income data are analyzed to predict default risk, helping the
bank approve safer loans and reduce NPAs.
Dr. G. Sindhu, Professor, PPG Business School