0% found this document useful (0 votes)
19 views46 pages

Business Analytics Course Overview

Uploaded by

amairakathuria10
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
19 views46 pages

Business Analytics Course Overview

Uploaded by

amairakathuria10
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

INTRODUCTION TO BUSINESS ANALYTICS

Dr Rishi Rajan Sahay


SSCBS
Outline:

Overview of the syllabus

Some interesting quotes about ‘Data’

Some facts related to data

An overview of the Business Analytics

Descriptive Analytics
Syllabus
Unit 1: Introduction to Business Analytics and Descriptive Analytics (14 hours)
Introduction to Business Analytics: Role of Analytics for Data Driven Decision Making; Types:
Descriptive Analytics, Predictive Analytics, and Prescriptive Analytics. Introduction to the concepts of
Big Data Analytics, Web and Social Media Analytics. Overview of Machine Learning Algorithms.
Introduction to relevant statistical software packages and carrying out descriptive analysis through it.

Unit 2: Predictive Analytics 1 (9 hours)


Simple Linear Regression: Estimation of Parameters, validation of simple linear regression model,
Coefficient of determination, Significance tests, Residual analysis, Confidence and Prediction
intervals.
Multiple Linear Regression: Interpretation of Partial regression coefficients, working with categorical
variables, Multi-collinearity and VIF, Outlier Analysis, Auto-correlation, transformation of variables,
variable selection in regression model building.
Unit 3: Predictive Analytics 2 (9 hours)
Logistic and Multinomial Regression: Logistic function, Estimation of probability using logistic
regression, Omnibus Test, Wald Test, Hosmer Lemshow Test, Pseudo R Square. Model Performance:
Classification table (sensitivity, specificity, accuracy paradox, precision, F score), Gini coefficient,
ROC, AUC, methods for determining the optimal cutoff probability.

Unit 4: Machine Learning Models (13 hours)


Decision Trees: Introduction, Chi-Square Automatic Interaction Detection, Bonferroni Correction,
Classification and Regression Tree, Gini Impurity Index, Entropy, Cost based splitting Criteria,
Ensemble Methods, Random Forest. Clustering: Introduction, Distance and Dissimilarity measures
used in clustering, Quality and Optimal Number of clusters, Clustering Algorithms, K-Means
clustering, Hierarchical Clustering.
Practical component (30 hours)

Practical Exercises:
1. Prepare and import data (financial data of companies, macroeconomic data,
primary data collected through questionnaires). Calculate and interpret
descriptive statistics on R/Python.
2. Perform simple OLS regression on R/Python and interpret the results obtained.
3. Test the assumptions of OLS (multicollinearity, autocorrelation, normality etc.)
on R/Python.
4. Perform regression analysis with categorical/dummy/qualitative variables on
R/Python.
5. Perform probabilistic regression models (logit and probit) along with validation
tests and classification table on R/Python.
6. Apply and interpret the results of decision trees and clustering models on R and
Python.
Essential/recommended readings
1. Business Analytics: The Science of Data Driven Decision Making, First Edition (2017), U Dinesh
Kumar, Wiley India.

Suggestive readings
1. Introduction to Machine Learning with Python, Andreas C. Mueller and Sarah Guido, O'Reilly
Media, Inc.
2. Data Mining for Business Analytics – Concepts, Techniques, and Applications in Python.
Galit Shmueli, Peter C. Bruce, Peter Gedeck, and Nitin R. Patel. Wiley.
3. Relevant Case Studies from different functional domains of business to be used while covering
the Predictive Analytics and Machine Learning models.

Following Case Studies may be taken up along with the course topics:
■ Merton Truck Company (HBS Case)
■ Supply Chain Optimization at Madurai Aavin Milk Dairy (IIMB Case).
■ Red Brand Canners (Stanford Case); Managing Linen at Apollo Hospitals (IIMB Case).
Note: Examination scheme and mode shall be as prescribed by the Examination Branch, University
of
Some quotes on Data..
In God we trust; all others must bring data
-Edward Deming
Data is the new oil
-Clive Humby (the famous phrase
was later embraced by World Economic
Forum in 2011)

Data is the new oil. We need to find it,


extract it, distribute it and monetize it.
-David Buckingham

The world’s most valuable resource is


no longer oil, but data…-The Economist, May 2017

The data is not only new oil, but also new soil.
-Mukesh Ambani (Hindustan Times Leadership summit, 2017)
Some facts about data:

1 gigabyte (GB) = 1024 megabytes (MB)

1 terabyte (TB) = 1024 gigabytes (GB)

1 petabyte = 1024 terabyte (TB)

1 Exabyte = 1024 petabytes

1 zettabytes = 1024 exabytes

1 yottabyte = 1024 zettabyte


Some interesting facts about data:

-Every 2 days we create as much information as we did from


the beginning of time until 2003;

- Over 90% of all the data in the world was created in the past 2 years;

-The total amount of data being captured and stored by industry


doubles every 1.2 years;

-If you burned all of the data created in just one day onto DVDs, you
could stack them on top of each other and reach the moon – twice.

(As per a report published in 2016)


Contd..

• The Global Datasphere will grow from 33 Zettabytes (ZB) in 2018


to 175 ZB by 2025, a 26% annual compound growth rate (CAGR), as
per IDC’s DataAge 2018 report.

• "The Global DataSphere is expected to more than double in size from


2022 to 2026…”
Business Analytics

It is the process of analyzing data to gain insights and make informed decisions.

Wayne Winston, a prominent scholar and consultant in management science


and prescriptive analytics, defines analytics as simply “using data for better
decision making”

Data analytics converts raw data into actionable insights. It includes a range of
tools, technologies, and processes used to find trends and solve problems by
using data. Data analytics can shape business processes, improve decision-
making, and foster business growth.---aws website
Data Analysis vs Data Analytics
Analysis: the act of studying or examining something in detail, in order to discover or
understand more about it, or your opinion and judgment after doing this.

Analytics: a process in which a computer examines information using mathematical methods in


order to find useful patterns.

Data Analysis: It involves applying various statistical and computational techniques to identify
patterns, trends, correlations, and anomalies within the data. Data analysis is typically focused
on understanding the past or current state of affairs based on historical data.

Data analytics: is a broader concept that encompasses various processes and techniques for
extracting insights and value from data. It includes data analysis as one of its components but
extends beyond it. Data analytics involves a more comprehensive approach that not only
analyzes historical data but also focuses on predicting future outcomes and prescribing actions
to achieve desired goals.
Three main
Analytics types of
analytics:

Descriptive Predictive Prescriptive


Analytics Analytics Analytics
Big Data
Big data refers to extremely large and diverse collections of structured,
unstructured, and semi-structured data that continues to grow exponentially
over time. These datasets are so huge and complex that traditional data
management systems cannot store, process, and analyze them.
---Google cloud website

Big data has following characteristics (Four Vs):

• Volume (large amounts of data)

• Velocity (fast data processing)

• Variety (different types of data)

• Veracity (accuracy in data).


Big Data Analytics
Big data analytics is the process of finding patterns, trends, and
relationships in massive datasets. These complex analytics require
specific tools and technologies, computational power, and data storage
that support the scale.—aws website
Web Analytics
Web Analytics involves the collection, measurement, and analysis of web data to
understand and optimize web usage. It focuses on analyzing user behavior on
websites and mobile apps.

Key Metrics: Common metrics in web analytics include page views, session duration,
bounce rate, conversion rate, and user demographics.

Tools: Popular web analytics tools include Google Analytics, Adobe Analytics, and
Matomo. These tools help track and report on website traffic, user behavior, and
conversion metrics.

Applications: Web Analytics is essential for digital marketing, website optimization,


and user experience enhancement. It helps businesses understand how users
interact with their websites, which content or products are most popular, and how
marketing campaigns perform.
Social Media Analytics
Social Media Analytics involves collecting and analyzing data from social media
platforms to gain insights into user sentiment, engagement, and trends. It helps
understand how brands, products, or services are perceived by the public.

ROI (Return on Investment) and ROE (Return on Engagement) are two important
metrics used in social media analytics to measure the effectiveness and success of
social media activities.

ROI measures the profitability of an investment relative to its cost. In the context
of social media, it quantifies the financial return from social media activities
compared to the investment made in them

ROI=(Net Profit/ Total Investment)​×100


Return on Engagement (ROE)
ROE measures the effectiveness of social media engagement in terms of interactions
like likes, shares, comments, and other forms of user engagement. Unlike ROI, which
focuses on financial returns, ROE focuses on the value derived from user engagement
and interaction.

Importance: ROE is particularly valuable for assessing the success of social media
efforts in building brand awareness, fostering community, and engaging with
audiences. It can also provide insight into customer sentiment and brand perception.
ROE is not as easily quantifiable as ROI and often involves tracking various
engagement metrics. It may include:

• Engagement Rate: The ratio of total engagement (likes, shares,


comments, etc.) to the total number of followers or impressions.

• Conversion Rate: The percentage of engaged users who take a desired


action, such as signing up for a newsletter or making a purchase.

• Sentiment Analysis: Analyzing the sentiment (positive, negative, neutral)


of comments and posts to gauge public perception.
Machine Learning

• Machine learning is a set of algorithms


that have the capability to learn from
the data.

• Machine learning is a set of methods


that can automatically detect patterns
in data.

• These uncovered patterns are then


used to predict future data, or to
perform other kinds of decision-
making under uncertainty.

• The key premise is learning from data!!


Supervised learning
• Machines are trained using
well "labelled" training data,
and on basis of that data,
machines predict the output.
• The labelled data means
some input data is already
tagged with the correct
output.

Unsupervised learning
• label is not present for any
observation

Semi-supervised learning
• falls in between the two

Reinforcement learning
• algorithm (called the agent)
continuously learns from the
environment in an iterative
fashion.
Some common applications of ML:

• Spam filtering
• Image classification
• Medical diagnosis
• forecasting
• Risk Assessment
• Clustering
• Putting stocks in different clusters based on financial ratios
• Customer segmentation- More Hypermarket, Zomato, Paytm, Blinkit,
Google collage theme based photos
• Dimesionality reduction (PCA)
• Recommender system
• Marketing (Pricing, sale prediction)/ Finance (Portfolio optimization, Loan
Application Analysis, fraud detection)
Supervised learning

Supervised learning is a type of machine learning where the model is trained on a


labeled dataset. In this paradigm, the dataset contains input-output pairs, where the
inputs are the features (independent variables) and the outputs are the target labels
(dependent variables). The goal of supervised learning is to learn a mapping from
inputs to outputs that can be generalized to unseen data. The learned model is then
used to make predictions or classify new data points.

• Regression

• Classification-Logistic regression, KNN, Decision Tree, SVM, Random Forest etc.


Unsupervised learning

Unsupervised learning is a type of machine learning where the algorithm learns


patterns from unlabelled data. Unlike supervised learning, there are no
predefined labels or outcomes. The goal is to find hidden structures, patterns, or
features in the data. Here are some key topics and techniques in unsupervised
learning:

• Clustering
• Dimensionality Reduction
• Association Rule Mining
Descriptive Statistics

Central Tendency: Mean, median, Mode

Measure of Dispersion: Range, Interquartile range, Standard deviation,


Variance, Mean Absolute deviation

Other measures of dispersion: Coefficient of variation, Coefficient of


quartile deviation

Shape of data: Skewness and Kurtosis

Five number summary or five point summary


Data Preparation
Before an analysis is performed, data requires to be treated so that it is
ready/suitable for analysis.

• Data description
• Data cleaning
• Missing value treatment
• Outlier treatment
• Variable transformation
• Variable addition/deletion
• Rows addition/deletion
• Data splitting
Central Tendency

• Average

Arithmetic Mean
Geometric Mean
Harmonic Mean
→(Ungrouped data & Grouped data)

• Median, Quartiles
→ (Ungrouped data & Grouped data)

• Mode
→ (Ungrouped data & Grouped data)
Measure of Dispersion

• Range

• Inter Quartile Range (IQR)

• Variance

• Standard deviation

• Coefficient of variation
Measure of Shape

• Skewness
• Kurtosis

Symmetric data: In statistics, symmetric data refers to a dataset where the


values are evenly distributed around the mean value, resulting in a
symmetrical distribution. This means that the data points on the left side of
the distribution mirror those on the right side, creating a shape that looks like
a mirror image.
Skewness- is a measure of the asymmetry of a distribution. In statistics, it is
used to describe how much a distribution deviates from a symmetrical
distribution, such as a normal distribution. Skewness can be positive, negative,
or zero.
Skewed data can have an impact on data analysis and modeling, as it can affect
the accuracy of statistical measures and lead to biased predictions. To address
this, it may be necessary to transform the data or use alternative statistical
measures that are more robust to skewness.
Kurtosis
Kurtosis is a statistical measure that describes the degree of peakedness or
flatness of a probability distribution compared to a normal distribution. A
normal distribution has a kurtosis of 0, and a distribution that is more peaked
than a normal distribution (i.e., has more values in the tails and fewer in the
center) has positive kurtosis, while a distribution that is flatter than a normal
distribution (i.e., has fewer values in the tails and more in the center) has
negative kurtosis.
kurtosis is a useful tool for investors to evaluate the risk and return
characteristics of an investment or portfolio, and to make informed
investment decisions.

In finance, kurtosis can be used to measure the degree of risk associated


with an investment. A distribution with high kurtosis (i.e., a fat-tailed
distribution) indicates that the probability of extreme events (such as large
gains or losses) is higher than in a normal distribution. This means that an
investment with a high kurtosis distribution may be riskier than an
investment with a low kurtosis distribution, even if they have similar means
and standard deviations.

A portfolio with assets that have low or negative kurtosis can be more
attractive for investors seeking a well-diversified portfolio, since it may be
less sensitive to extreme market events.
Empirical Rule
The empirical rule, also known as the 68-95-99.7 rule, is a statistical rule of
thumb that describes the approximate proportion of data that falls within a
certain number of standard deviations from the mean in a normal distribution.
Specifically, the empirical rule states that:
•About 68% of the data falls within one standard deviation of the mean.
•About 95% of the data falls within two standard deviations of the mean.
•About 99.7% of the data falls within three standard deviations of the mean.
Role of Normality
• Many statistical methods require that the numeric
variables we are working with have an approximate
normal distribution.
• For example, t-tests, F-tests, and regression analyses
all require in some sense that the numeric variables
are approximately normally distributed.
How to assess normality?

1. Graphically (Data visualization)

[Link] of the data

3. Goodness of fit test


Graphical approach
Normality of the data can be assessed through various data visualizations.
Some of the important data visualizations useful for assessing normality are:

Histogram

Density plot

Boxplot

Normal Quantile plot


Normal quantile plot
QQ plot
The quantile-quantile (q-q) plot is a graphical technique for determining
if two data sets come from populations with a common distribution.
Summary of the data
Descriptive statistics can be used to get some insight about the symmetricity
of the data. For example:
The location of central value (mean, median, mode)
Skewness
Kurtosis
Goodness of fit test for Normality
Ho: The distribution is normal
HA: The distribution is NOT normal

Common tests of normality include:


Shapiro-Wilk
Kolmogorov-Smirnov
Anderson-Darling
Lillefor’s
Problem: THEY DON’T ALWAYS AGREE!!
What if the data is not normal?

Transform the data


Z-score
A standardized value, commonly called a z-score, provides a relative
measure of the distance an observation is from the mean. It is independent
of the units of measurement.

It can be used to identify outliers or anomalies in the data that may be


indicative of fraudulent activity. For example, let's say we have a dataset of
transaction amounts and we want to identify any potentially fraudulent
transactions. We can calculate the z-score for each transaction using the
mean and standard deviation of the transaction amounts, and then flag any
transactions with a z-score greater than 3 or less than -3 for further
investigation.
z-score = (x-μ)/σ
Measure of association

Covariance
Correlation
Scatter diagram
Coefficient of Correlation

Correlation means relation between two or more variables. Correlation analysis


deals with studying the relationship between two or more variables. The
foundation of correlation analysis was laid by Sir Francis Galton (Millar, 1996) in
1880s. It was further developed by Karl Pearson (Stigler, 1986).
The correlation between two variables is referred to as simple correlation. In case
of more than two variables, we can think of multiple and partial correlation.

Karl Pearson’s coefficient of correlation

𝑁
𝑐𝑜𝑣(𝑥,𝑦) ෌𝑖=1 𝑥𝑖 −𝑥ҧ 𝑦𝑖 −𝑦ത
r= =
𝜎𝑥 .𝜎𝑦
σ𝑁
𝑖=1(𝑥𝑖 −𝑥ҧ)
2 σ𝑁
𝑖=1(𝑦𝑖 −𝑦ҧ)
2
Exploratory data analysis

Exploratory Data Analysis (EDA) is the process of analyzing and visualizing


data to better understand its patterns, trends, and underlying
relationships. EDA relies heavily on data visualization, and summary
statistics
Data Visualization

• Categorical data
• Continuous data
• Scatter diagram

You might also like