PA Notes
PA Notes
Predictive analytics is a branch of advanced data science that uses historical data, statistical
modeling, and machine learning techniques to forecast future events or behaviors. By
identifying patterns in data, it helps organizations anticipate risks and opportunities. Key
applications include forecasting customer needs, managing supply chains, and identifying
financial fraud.
Predictive analytics is the use of statistics and modeling techniques to forecast future outcomes.
Current and historical data patterns are examined and plotted to determine the likelihood that
those patterns will repeat.
Businesses use predictive analytics to fine-tune their operations and decide whether new
products are worth the investment. Investors use predictive analytics to decide where to put
their money. Internet retailers use predictive analytics to fine-tune purchase recommendations
to their users and increase sales.
A common misconception is that predictive analytics and machine learning are the same.
Predictive analytics help us understand possible future occurrences by analyzing the past. At
its core, predictive analytics includes a series of statistical techniques (including machine
learning, predictive modeling, and data mining) and uses statistics (both historical and current)
to estimate, or predict, future outcomes.
Machine learning is a subfield of computer science that means "the programming of a digital
computer to behave in a way which, if done by human beings or animals, would be described
as involving the process of learning."
Predictive analytics is the use of historical data, statistical modeling, and machine learning
techniques to forecast future outcomes, trends, and behaviors. It helps organizations anticipate
potential scenarios—from machine failures to customer actions—to drive proactive, data-
driven decisions.
• Improved customer targeting: Analyzing data can help businesses identify and target
their ideal customers more effectively.
• Increased customer retention: Data can help businesses identify customers at risk of
churning and take steps to prevent them from leaving.
• Reduced fraud: Data can help businesses identify fraudulent transactions and prevent
them from occurring in the first place.
• Optimized operations: Data can help businesses optimize their operations, such as
supply chain and inventory management.
• Increased revenue: Data can help businesses increase revenue by identifying new
opportunities, such as upselling and cross-selling to existing customers.
Descriptive Analytics
Descriptive analytics is the process of analyzing historical data to understand what has
happened in the past. It focuses on summarizing and interpreting data to provide insights into
past performance and trends. Descriptive analytics answers the question, "What
happened?"
• Data Visualization: Using charts, graphs, and dashboards to represent data visually.
• Business Reporting: Generating regular reports on sales, revenue, and other key
performance indicators (KPIs).
Predictive Analytics
Predictive analytics uses historical data and statistical algorithms to forecast future events. It
aims to predict what is likely to happen based on past trends and patterns. Predictive analytics
answers the question, "What could happen?"
• Time Series Analysis: Analyzing data points collected or recorded at specific time
intervals.
• Machine Learning: Using algorithms to learn from data and make predictions.
• Python: For machine learning and data analysis libraries like scikit-learn and
TensorFlow.
• Risk Management: Predicting potential risks and their impact on business operations.
Prescriptive Analytics
Prescriptive analytics goes beyond predicting future outcomes by recommending actions to
achieve desired results. It combines data, algorithms, and business rules to suggest the best
course of action. Prescriptive analytics answers the question, "What should we do?"
• Machine Learning: Using algorithms to learn from data and make recommendations.
Key Differences Between Descriptive, Predictive and Prescriptive data analytics model
While descriptive, predictive, and prescriptive analytics are interconnected, they serve different
purposes and provide different insights:
Tools Used - data mining, Tools Used - machine Tools Used - heuristics,
data aggregation learning, statistical model optimization
Example:
Example:
Ecommerce businesses
Example : Identifying techniques to
which uses the customer's
Annual Revenue Report optimize the patient care in
browsing history to
the healthcare
recommend products
Descriptive, Predictive and Prescriptive data analytics are important types of analytics where
Descriptive analytics is used to summarize the data, Predictive analytics is used to make future
predictions based on the past data and Prescriptive analytics is used to identify the possible
future outcomes and show the best option.
The predictive analytics process is a systematic sequence of steps used to extract insights from
data and build models that can predict future outcomes. It ensures that predictions are accurate,
relevant, and useful for decision-making.
CRISP-DM (Cross-Industry Standard Process for Data Mining) is a standard framework used
to plan and execute data mining and predictive analytics projects. It provides a structured
approach to convert raw data into useful insights and predictions.
CRISP-DM was originally developed as a data mining process model, but its steps are general
and flexible, which makes it equally suitable for predictive analytics projects.
This is the first and most important step where the problem is clearly defined. Organizations
identify their objectives, such as predicting customer churn or forecasting sales. It ensures that
the analytics work is aligned with business goals.
In this stage, relevant data is gathered from various sources such as databases, customer
records, or transaction systems. The data is then explored to understand its structure, quality,
and relevance to the problem.
3. Data Preparation
This step involves cleaning and transforming the data. Missing values are handled, errors are
corrected, and useful variables (features) are selected. Proper data preparation improves the
accuracy of predictive models.
4. Model Building
In this stage, statistical and machine learning techniques such as regression, decision trees, or
classification models are applied. The goal is to develop a model that can identify patterns and
make predictions.
5. Model Evaluation
The developed model is tested to check its accuracy and reliability. Techniques like validation
and testing are used to ensure the model performs well on new, unseen data and avoids issues
like overfitting.
6. Deployment
Once the model is validated, it is implemented in a real business environment. The model is
used to make predictions that support decision-making in day-to-day operations.
After deployment, the model’s performance is continuously monitored. Since data and business
conditions change over time, the model may need updates or retraining to remain accurate.
Predictive analytics promotes data-driven decision-making (DDD), where decisions are based
on empirical evidence rather than intuition. According to Provost and Fawcett, even small
improvements in predictive accuracy can lead to better business outcomes. This reduces bias
and enhances objectivity in managerial decisions.
By analyzing historical data and identifying patterns, predictive models estimate future
outcomes with higher accuracy. This reduces uncertainty in decision-making and allows
managers to make informed choices with greater confidence.
Predictive analytics allows organizations to anticipate future events and take action in advance.
Instead of reacting to problems after they occur, businesses can prevent them. For example,
predicting customer churn helps companies retain customers before they leave.
Predictive analytics supports long-term planning by forecasting trends such as demand, market
growth, and customer behavior. It helps managers develop effective strategies and allocate
resources efficiently.
Predictive analytics plays a crucial role in identifying potential risks such as fraud, credit
defaults, and operational failures. By quantifying risk probabilities, it helps organizations take
preventive measures and minimize losses (Abbott emphasizes risk-based decision-making).
Organizations can better understand customer preferences and behavior through predictive
models. This leads to more personalized marketing strategies, improved customer satisfaction,
and stronger customer relationships.
Predictive analytics supports automated decision systems, where decisions are made instantly
based on model outputs. This is widely used in credit scoring, fraud detection, and
recommendation systems, improving efficiency and consistency.
Organizations that effectively use predictive analytics gain a competitive edge by making
smarter and faster decisions. They can identify opportunities earlier and respond to market
changes more effectively than competitors.
Better decisions lead to improved operational efficiency, cost reduction, and revenue growth.
Predictive analytics directly contributes to enhanced organizational performance and
profitability.
APPLICATIONS IN BUSINESS CASE STUDIES FROM VARIOUS INDUSTRIES
(E.G., FINANCE, MARKETING, OPERATIONS)
Predictive analytics is widely used across industries to improve decision-making, optimize
operations, reduce risks, and increase profitability. By analyzing historical and real-time data,
organizations can predict future trends and take proactive actions.
1. Finance Industry
Application Areas
• Credit scoring
• Fraud detection
• Risk management
• Investment forecasting
2. Marketing Industry
Application Areas
• Customer segmentation
• Personalized marketing
• Customer churn prediction
• Product recommendation
Recommendation System by Amazon
Amazon uses predictive analytics to provide personalized product recommendations to
customers. The company analyzes browsing history, previous purchases, search patterns,
ratings, and customer preferences to predict products that customers are likely to buy.
Its recommendation engine generates suggestions such as “Frequently Bought Together” and
“Customers Also Bought.” These predictions significantly influence customer purchasing
decisions.
Impact
• Increased sales and cross-selling
• Improved customer experience
• Higher customer engagement and retention
4. Healthcare Industry
Application Areas
• Disease prediction
• Patient risk analysis
• Hospital resource planning
Data Collection and Preparation: Data Sources and Collection: Types of data (structured
vs. unstructured)/ Data collection methods and tools.
Data Cleaning and Preparation: Handling missing data. Data transformation and
normalization. Data Preparation Using Excel or Python/R for data cleaning and preparation.
Data is the raw form of information, a collection of facts, figures, symbols or observations that
represent details about events, objects or phenomena. By itself, data may appear meaningless,
but when organized, processed and interpreted, it transforms into valuable insights that support
decision-making, problem-solving and innovation.
• Data refers to raw facts, figures, or information that can be processed and analysed to
extract meaningful insights.
• In data science and computing, data is categorised into different types based on its
structure and nature.
• Understanding its type helps in selecting appropriate analysis and processing methods.
Data collection is the systematic process of acquiring, collecting, extracting, and storing a
voluminous amount of data from various sources to analyze, research, or make data-driven
decisions. It involves identifying the required data type—structured (organized in
rows/columns) or unstructured (raw formats like text/video)—and utilizing methods like
surveys, interviews, or automated tools (APIs, IoT sensors) to gather data for analysis in
relational databases or data lakes.
TYPES OF DATA
Data can be categorized in different ways depending on how it is collected, stored and
represented. Broadly, it falls into the following:
Quantitative Data
Quantitative data is information that can be measured, counted and expressed in numerical
form. It provides objective values that can be analyzed statistically to identify patterns, trends
and relationships.
• Can be divided into: Discrete data (Whole numbers) and Continuous data (Values on a
scale).
Example: Age of people, number of customers visiting a store, temperature readings, sales
revenue.
Qualitative Data
Data exists in many different forms and sizes, but most can be presented as structured or
unstructured. This table has collected the main differences between these two data types.
Structured data is highly organized and exists in tabular format with interconnected rows and
columns, so you can easily search for specific details and single out the relationships between
its pieces. It doesn’t normally require much storage space. To manipulate it, there is a special
language called SQL, which stands for Structured Query Language and was developed back in
the 1970s by IBM.
Structured data examples. Most of us are familiar with structured data — Google Sheets and
Microsoft Office Excel files are the first things that spring to mind. This data can comprise
both textual elements and numbers, such as employee names, contacts, ZIP codes, addresses,
credit card numbers, etc.
The typical structured data example: an Excel spreadsheet that contains information
about customers and purchases.
Pretty much everyone has dealt with booking a ticket via one of the airline reservation systems.
They operate structured data: passenger names, location names, flight numbers, number of
passengers, etc. Once you have entered information into the appropriate field, the application
saves the data and allocates it to the appropriate tables in the database.
This information will be stored, read, changed, or deleted when needed. And it’s easy to
analyze. For example, just retrieve specific columns and rows of the corresponding tables to
compare the prices or the number of tickets purchased on different dates.
Unstructured data is schemaless, meaning it has no pre-defined structure and is stored in its
native format. This includes imagery, text documents, and video and audio files. It’s easy to
capture, provides many insights, and can be used in various ways, given that you have enough
storage to keep it and advanced technologies for analysis.
Unstructured data examples. Unstructured data includes a wide array of forms, such as
email, text files, social media posts, video, images, audio, sensor data, and so on.
For instance, a travel agency posts new travel tours on social media and wants to know the
audience's reactions. Each post contains metadata descriptions or attributes like shares or
hashtags that can be quantified and structured. However, the post itself (pictures plus text) and
comments belong to the category of unstructured data. Collecting valuable insights will take
advanced techniques like sentiment analysis.
Semi-Structured Data
Semi-structured data combines aspects of structured and unstructured data. It does not reside
in traditional tables but still contains tags or markers that provide a loose structure.
• Easier to analyze than unstructured data, but less rigid than structured data.
The actual data is then further divided mainly into two types known as:
• Primary data
• Secondary data
Primary data
The data which is Raw, original, and extracted directly from the official sources is known as
primary data. This type of data is collected directly by performing techniques such as
questionnaires, interviews, and surveys. The data collected must be according to the demand
and requirements of the target audience on which analysis is performed otherwise it would be
a burden in the data processing.
1. Interview method:
The data collected during this process is through interviewing the target audience by a person
called interviewer and the person who answers the interview is known as the interviewee. Some
basic business or product related questions are asked and noted down in the form of notes,
audio, or video and this data is stored for processing. These can be both structured and
unstructured like personal interviews or formal interviews through telephone, face to face,
email, etc.
2. Survey method:
The survey method is the process of research where a list of relevant questions are asked and
answers are noted down in the form of text, audio, or video. The survey method can be obtained
in both online and offline mode like through website forms and email. Then that survey answers
are stored for analyzing data. Examples are online surveys or surveys through social media
polls.
3. Observation method:
The observation method is a method of data collection in which the researcher keenly observes
the behavior and practices of the target audience using some data collecting tool and stores the
observed data in the form of text, audio, video, or any raw formats. In this method, the data is
collected directly by posting a few questions on the participants. For example, observing a
group of customers and their behavior towards the products. The data obtained will be sent for
processing.
4. Experimental method:
The experimental method is the process of collecting data through performing experiments,
research, and investigation. The most frequently used experiment methods are CRD, RBD,
LSD, FD.
• LSD - Latin Square Design is an experimental design that is similar to CRD and RBD
blocks but contains rows and columns. It is an arrangement of NxN squares with an
equal amount of rows and columns which contain letters that occurs only once in a row.
Hence the differences can be easily found with fewer errors in the experiment. Sudoku
puzzle is an example of a Latin square design.
• FD - Factorial design is an experimental design where each experiment has two factors
each with possible values and on performing trail other combinational factors are
derived.
5. Focus Groups:
Small group discussions moderated to gain insights into specific topics or products.
6. Case Studies:
Secondary data
Secondary data is the data which has already been collected and reused again for some valid
purpose. This type of data is previously recorded from primary data and it has two types of
sources named internal source and external source.
1. Internal source:
These types of data can easily be found within the organization such as market record, a sales
record, transactions, customer data, accounting resources, etc. The cost and time
consumption is less in obtaining internal sources.
2. External source:
The data which can’t be found at internal organizations and can be gained through external
third party resources is external source data. The cost and time consumption is more because
this contains a huge amount of data. Examples of external sources are Government
publications, news publications, Registrar General of India, planning commission,
international labor bureau, syndicate services, and other non-governmental publications.
3. Literature Review:
Analyzing existing studies, journals, and articles that can be internal or external.
3 Other sources:
• Sensors data: With the advancement of IoT devices, the sensors of these devices
collect data which can be used for sensor data analytics to track the performance and
usage of products.
• Satellites data: Satellites collect a lot of images and data in terabytes on daily basis
through surveillance cameras which can be used to collect useful information.
• Web traffic: Due to fast and cheap internet facilities many formats of data which is
uploaded by users on different platforms can be predicted and collected with their
permission for data analysis. The search engines also provide their data through
keywords and queries searched mostly.
• Recording Devices: Cameras, audio recorders, and video tools for qualitative studies.
• Digital Tools/Software: Programming languages (Python, R), SQL, and data analysis
software (SPSS, Tableau).
DATA CLEANING AND PREPARATION
Data Preparation is the third stage of the predictive modeling process, intended to convert data
identified for modeling into a form that is better for the predictive modeling algorithms. Each
data set can provide different challenges to data preparation, especially with data cleansing.
The key steps in data preparation related to the columns in the data are variable cleaning,
variable selection, and feature creation. Data preparation steps related to the rows in the data
are record selection, sampling, and feature creation (again).
Variable Cleaning
Variable cleaning refers to fixing problems with values of variables themselves, including
incorrect or miscoded values, outliers, and missing values.
Incorrect Values: Incorrect values are problematic because predictive modeling algorithms
assume that every value in each column is completely correct. If any values are coded
incorrectly, the only mechanism the algorithms have to overcome these errors is to overwhelm
the errors with correctly coded values, thus making the incorrect values insignificant.
Outliers: Outliers are unusual values that are separated from the main body of the distribution,
typically as measured by standard deviations from the mean or by the IQR. Whether or not you
remove or mitigate the influence of outliers is a critical decision in the modeling process.
If the outliers to be cleaned are examples of values that are correctly coded, there are four
typical approaches to handling them:
Multidimensional Outliers: Nearly all outlier detection algorithms in modeling software refer
to outliers in single variables. Outliers can also be multidimensional, but these are much more
difficult to identify. To mitigate the effects of multidimensional outliers, the modeler could
apply the same approaches described for single-variable outliers.
Missing Values: Missing values are arguably the most problematic of all data problems.
Missing values are typically coded in data with a null value or as an empty cell, although many
more representations can exist in data. Table 4-2 shows typical missing values you may
encounter in data.
Missing value correction is perhaps the most time-consuming of all the variable cleaning steps
needed in data preparation. Whenever possible, imputing missing values is the most desirable
action. Missing value imputation means changing values of missing data to a value that
represents a plausible or expected value in the variable if it were actually known. These are the
most commonly used methods for fixing problems associated with missing values.
Listwise and Column Deletion: The simplest method of handling missing values is listwise
deletion, meaning one removes any record with any missing values, leaving only records with
fully populated values for every variable to be used in the analysis. An alternative to listwise
deletion is column deletion: removing any variable that has any missing values at all, leaving
only variables that are fully populated. both listwise deletion and column deletion are practiced,
especially when the number of missing values is particularly large or when the timeline to
complete the modeling is particularly short.
Imputation with a Constant: This option is almost always available in predictive analytics
software. For categorical variables, this can be as simple as filling missing values with a “U”
or another appropriate string to indicate missing. For continuous variables, this is most often a
0. Sometimes, imputing with 0 causes significant problems, for example, age.
Mean and Median Imputation for Continuous Variables: The next level of sophistication
is imputing with a constant value that is not predefined by the modeler. Mean imputation is by
far the most common method, but in some circumstances, if the mean and median are different
from one another, imputing with the median may be better because the median will represent
better the most typical value of the variable. However, median imputation can be more
computationally expensive, especially if the number of records in the data is large.
Imputing with Distributions: When large percentages of values are missing, the summary
statistics are affected by mean imputation. An alternative to this is, rather than imputing with a
con stant value, to impute randomly from a known distribution. For the variable AGE, if instead
of imputing with the mean (61.6), you impute using a random number generated from a normal
distribution with mean 61.6 and standard deviation 16.6 (see Table 4-4). These imputed values
will retain the same shape as the original shape of AGE.
Random Imputation from Own Distributions: A similar procedure is called “Hot Deck”
imputation, where the “deck” referred originally to the days when the Census Bureau had cards
for individuals. If someone did not respond, one could take another card from the deck at
random and use that information as a substitute. In predictive modeling terms, this is random
imputation, but instead of using a random number generator to pick a number, a random actual
value of the variable of the non-missing values is selected. Predictive modeling software rarely
provides this option, but it is easy to do.
Imputing Missing Values from a Model: The model approach to missing value imputation
begins with changing the role of the input variable with missing values to now be a target
variable. The inputs to the new model are other input variables that may predict this new target
variable well. The training data should be large enough.
Dummy Variables Indicating Missing Values: Capturing the existence of missing data can
be done with a dummy variable coded as 1 when the variable is missing and 0 when it is
populated.
Imputation for Categorical Variables: When the data is categorical, imputation is not always
necessary; the missing value can be imputed with a value that represents missing so that no cell
contains a null any longer.
Data transformation and normalization are important steps in data preparation. They help
convert raw data into a suitable format and scale, improving the accuracy and performance
of predictive models.
DATA TRANSFORMATION
Meaning
Data transformation is the process of changing the format, structure, or values of data to
make it suitable for analysis.
1. Scaling
Scaling adjusts the range of data so that all variables are comparable. It prevents variables with
large values from dominating the model.
Example: Converting income from ₹10,000–₹1,00,000 into a smaller range like 0–1.
2. Encoding
Encoding converts categorical (text) data into numerical form so that it can be used in models.
Most machine learning algorithms require numeric input.
3. Aggregation
Aggregation combines multiple data points into a summarized form. It helps in identifying
patterns and reducing data complexity.
4. Feature Construction
Feature construction involves creating new variables from existing data to improve model
performance. It helps capture hidden relationships.
5. Log Transformation
Log transformation reduces skewness in data by compressing large values. It helps in making
data more normally distributed.
Example: Applying log to income data where a few values are extremely high.
DATA NORMALIZATION
Meaning
Normalization is a technique used to rescale data into a common range, usually between 0
and 1.
Techniques of Normalization
1. Min-Max Normalization
This method rescales data to a fixed range (0 to 1) while preserving relationships between
values.
𝑥 − 𝑥𝑚𝑖𝑛
𝑥′ =
𝑥𝑚𝑎𝑥 − 𝑥𝑚𝑖𝑛
This method transforms data so that it has a mean of 0 and a standard deviation of 1. It is useful
when data has different units.
𝑥−𝜇
𝑧=
𝜎
3. Decimal Scaling
This method moves the decimal point based on the maximum value in the dataset.
Key Points
Data preparation is the process of converting raw data into a clean and usable format for
analysis and predictive modeling. Tools such as Excel, Python, and R are widely used for
cleaning, transforming, and preparing data because they provide powerful functions for
handling large datasets efficiently.
Microsoft Excel is one of the most commonly used tools for basic data cleaning and preparation
because it is simple and user-friendly.
Excel helps identify and handle missing values using filters, conditional formatting, and
formulas. Missing values can be replaced using mean, median, or manually entered values.
Example: Replacing blank sales entries with the average sales value using the AVERAGE()
function.
Excel provides the “Remove Duplicates” feature to identify and delete repeated records. This
improves data accuracy and prevents biased analysis.
Excel allows users to sort and filter data for easier analysis and organization. Data can be
arranged in ascending or descending order.
Excel functions such as CONCATENATE(), LEFT(), RIGHT(), and TEXT() help transform
data into the required format.
Example: Splitting full names into first name and last name columns.
e. Normalization and Calculations
Excel formulas can be used for scaling and normalization of data. Pivot tables and charts also
support summarization and visualization.
Python is widely used in predictive analytics because of libraries such as Pandas, NumPy, and
Scikit-learn.
Python provides functions like fillna() and dropna() in Pandas to handle missing values
efficiently.
Example: Replacing missing customer income values with the mean income.
[Link]([Link]())
b. Removing Duplicates
df.drop_duplicates()
c. Data Transformation
Python supports encoding, aggregation, and feature engineering for predictive models.
pd.get_dummies(df['Gender'])
d. Data Normalization
Libraries like Matplotlib and Seaborn help visualize patterns and detect outliers.
R is a statistical programming language widely used for data analysis and predictive modeling.
Functions like [Link]() and [Link]() are used to identify and remove missing values.
[Link](data)
b. Data Transformation in R
Packages like dplyr and tidyr help in reshaping and transforming datasets.
c. Data Normalization in R
d. Data Visualization in R
STATISTICAL CONCEPTS
PROBABILITY DISTRIBUTIONS
Probability distribution yields the possible outcomes for any random event. It is also defined
based on the underlying sample space as a set of possible outcomes of any random experiment.
A probability distribution is a statistical function that describes all the possible values and
likelihoods that a random variable can take within a given range. This range will be
bounded between the minimum and maximum possible values, but precisely where the possible
value is likely to be plotted on the probability distribution depends on a number of factors.
These factors include the distribution's mean (average), standard deviation, skewness,
and kurtosis. In simple terms, it tells us the pattern of probabilities for all possible outcomes.
A discrete distribution is used when the random variable can take countable values (finite or
countably infinite). Each possible value has a specific probability assigned to it. The sum of all
probabilities equals 1.
Example
Rolling a die:
• Possible values: 1, 2, 3, 4, 5, 6
• Each has probability = 1/6
Other examples: number of customers, number of defects
2. Continuous Probability Distribution
A continuous distribution is used when the variable can take any value within a range.
Probabilities are represented using a probability density function (PDF). The probability of
a single value is zero; instead, we calculate probability over intervals.
Example
• Height of people (e.g., between 150 cm and 180 cm)
• Temperature values
Important Concepts
The binomial distribution is a discrete distribution with a finite number of possibilities. When
observing a series of what are known as Bernoulli trials, the binomial distribution emerges. A
Bernoulli trial is a scientific experiment with only two outcomes: success or failure.
Consider a random experiment in which you toss a biased coin six times with a 0.4 chance of
getting head. If 'getting a head' is considered a ‘success’, the binomial distribution will show
the probability of r successes for each value of r.
The binomial random variable represents the number of successes (r) in n consecutive
independent Bernoulli trials.
The Bernoulli distribution is a variant of the Binomial distribution in which only one
experiment is conducted, resulting in a single observation. As a result, the Bernoulli distribution
describes events that have exactly two outcomes.
A Poisson distribution is a probability distribution used in statistics to show how many times
an event is likely to happen over a given period of time. To put it another way, it's a count
distribution. Poisson distributions are frequently used to comprehend independent events at a
constant rate over a given time interval. Siméon Denis Poisson, a French mathematician, was
the inspiration for the name.
Normal distribution, also called Gaussian distribution, the most common distribution
function for independent, randomly generated variables. Its familiar bell-shaped curve
is ubiquitous in statistical reports, from survey analysis and quality control to resource
allocation.
The graph of the normal distribution is characterized by two parameters: the mean, or average,
which is the maximum of the graph and about which the graph is always symmetric; and
the standard deviation, which determines the amount of dispersion away from the mean.
The Standard Normal Model: A standard normal model is a normal distribution with a mean
of 0 and a standard deviation of 1.
MEANING OF HYPOTHESES
A hypothesis is an assumption that is made based on some evidence. This is the initial point of
any investigation that translates the research questions into predictions. It includes components
like variables, population and the relation between the variables. A research hypothesis is a
hypothesis that is used to test the relationship between two or more variables.
TYPES OF HYPOTHESES
ii. Complex Hypothesis: It shows the relationship between two or more dependent
variables and two or more independent variables. Eating more vegetables and fruits
leads to weight loss, glowing skin, reduces the risk of many diseases such as heart
disease.
vi. Associative and Causal Hypothesis: Associative hypothesis occurs when there is a
change in one variable resulting in a change in the other variable. Whereas, causal
hypothesis proposes a cause and effect interaction between two or more variables.
Examples of Hypothesis
• All lilies have the same number of petals is an example of a null hypothesis.
• If a person gets 7 hours of sleep, then he will feel less fatigue than if he sleeps less.
CHARACTERISTICS OF HYPOTHESES
i. Hypothesis should be clear and precise. If the hypothesis is not clear and precise, the
inferences drawn on its basis cannot be taken as reliable.
v. Hypothesis should be stated as far as possible in most simple terms so that the same
is easily understandable by all concerned. But one must remember that simplicity of
hypothesis has nothing to do with its significance.
vi. Hypothesis should be consistent with most known facts i.e., it must be consistent
with a substantial body of established facts. In other words, it should be one which
judges accept as being the most likely.
vii. Hypothesis should be amenable to testing within a reasonable time. One should not
use even an excellent hypothesis, if the same cannot be tested in reasonable time for
one cannot spend a life-time collecting data to test it.
viii. Hypothesis must explain the facts that gave rise to the need for explanation. This
means that by using the hypothesis plus other known and accepted generalizations, one
should be able to deduce the original problem condition. Thus hypothesis must actually
explain what it claims to explain; it should have empirical reference.
FORMULATION OF HYPOTHESES
Formulating a hypothesis can take place at the very beginning of a research project, or after a
bit of research has already been done. Sometimes a researcher knows right from the start which
variables she is interested in studying, and she may already have a hunch about their
relationships. Other times, a researcher may have an interest in a particular topic, trend, or
phenomenon, but he may not know enough about it to identify variables or formulate a
hypothesis.
Whenever a hypothesis is formulated, the most important thing is to be precise about what one's
variables are, what the nature of the relationship between them might be, and how one can go
about conducting a study of them.
PROCEDURE FOR TESTING HYPOTHESES
ERRORS IN HYPOTHESES
There are two types of errors in hypothesis testing, Type I and Type II errors. Both types of
error relate to incorrect conclusions about the null hypothesis.
Type I Error: We may reject H0 when H0 is true. Type I error means rejection of hypothesis
which should have been accepted. Type I error is denoted by α (alpha) known as α error, also
called the level of significance of test
Type II Error: We may accept H0 when in fact H0 is not true. Type II error means accepting
the hypothesis which should have been rejected. Type II error is denoted by β (beta) known as
β error.
Hypothesis testing determines the validity of the assumption (technically described as null
hypothesis) with a view to choose between two conflicting hypotheses about the value of a
population parameter. Hypothesis testing helps to decide on the basis of a sample data, whether
a hypothesis about the population is likely to be true or false.
Parametric tests usually assume certain properties of the parent population from which we draw
samples. Assumptions like observations come from a normal population, sample size is large,
assumptions about the population parameters like mean, variance, etc., must hold good before
parametric tests can be used. parametric tests require measurement equivalent to at least an
interval scale.
There are situations when the researcher cannot or does not want to make such assumptions.
In such situations we use statistical methods for testing hypotheses which are called non-
parametric tests because such tests do not depend on any assumption about the parameters of
the parent population. Non-parametric tests assume only nominal or ordinal data.
Parameters of
Parametric Nonparametric
Comparison
Type of It is used on data that follows a It is used on data that follows any
distribution normal distribution. arbitrary distribution.
Meaning
Regression analysis is a statistical technique used to study the relationship between a dependent
variable and one or more independent variables. It is mainly used for prediction and estimation
of continuous values.
Example
A company wants to predict sales (Y) based on advertising expenditure (X). Regression helps
estimate how much sales will increase when advertising increases.
Key Components
• Dependent Variable (Y): The variable we want to predict or explain. Example: Sales,
profit, demand
• Independent Variable (X): The variable(s) used to predict Y. Example: Price,
advertising, income
Uses of Regression
• Sales forecasting
• Demand prediction
• Financial analysis
• Marketing effectiveness
Advantages
• Simple and easy to interpret
• Useful for prediction
• Helps understand relationships
Limitations
• Assumes linear relationship
• Sensitive to outliers
• May not capture complex patterns
BUILDING STATISTICAL MODELS
Building a regression model involves developing a mathematical equation that best fits the
data.
Meaning
Simple linear regression uses one independent variable to predict the dependent variable.
Model Equation
Y = a + bX
Explanation
The model tries to fit a straight line that minimizes the error between actual and predicted
values.
Example
Meaning
Multiple regression uses two or more independent variables to predict the dependent
variable.
Model Equation
𝑌 = 𝑎 + 𝑏1 𝑋1 + 𝑏2 𝑋2 + ⋯ + 𝑏𝑛 𝑋𝑛
Explanation
Each independent variable contributes to predicting Y. This model is more realistic because
business problems usually depend on multiple factors.
Example
MODEL ASSUMPTIONS
Model assumptions are the conditions that must be satisfied for a regression model to produce
valid, reliable, and unbiased results. If these assumptions are violated, the model’s predictions
and inferences may become inaccurate.
1. Linearity
This assumption states that there must be a linear relationship between the dependent variable
and independent variables. It means that a change in X should result in a proportional change
in Y. If the relationship is not linear, the model may not fit the data properly.
Example: Sales increasing steadily with advertising expenditure indicates linearity, whereas a
curved relationship violates this assumption.
2. Independence of Errors
The residuals (errors) should be independent of each other, meaning the error in one
observation should not influence another. This is especially important in time-series data where
values may be correlated over time.
Example: Daily sales errors should not depend on the previous day’s errors; otherwise,
autocorrelation exists.
Example: If prediction errors are small for low sales but large for high sales, the assumption
is violated.
4. Normality of Errors
Residuals should be normally distributed, especially for hypothesis testing and confidence
interval estimation. While regression can still work without perfect normality, significant
deviations may affect statistical conclusions.
5. No Multicollinearity
Independent variables should not be highly correlated with each other, as this makes it difficult
to isolate their individual effects on the dependent variable. High multicollinearity can lead to
unstable coefficient estimates.
Example: If price and discount are strongly related, it becomes difficult to determine their
separate impact on sales.
6. No Significant Outliers
The dataset should not contain extreme values that disproportionately influence the regression
results. Outliers can distort the regression line and reduce model accuracy.
Example: A one-time unusually high sales value may skew the results and affect predictions.
MODEL DIAGNOSTICS
Model diagnostics refers to the set of techniques used to evaluate the performance, validity,
and reliability of a regression model. It helps in checking whether the model assumptions are
satisfied and whether the model is suitable for prediction.
1. R² (Coefficient of Determination)
R² measures the proportion of variation in the dependent variable explained by the independent
variables. Its value ranges from 0 to 1, where a higher value indicates a better fit of the model.
However, a very high R² does not always mean the model is perfect, as it may also indicate
overfitting.
2. Adjusted R²
Adjusted R² modifies the R² value by considering the number of independent variables in the
model. It penalizes unnecessary variables that do not improve the model significantly. This
makes it more reliable for multiple regression models.
Example: If adding a new variable does not improve the model, Adjusted R² may decrease.
3. Residual Analysis
Residuals are the differences between actual and predicted values. Analyzing residuals helps
check whether assumptions like linearity, independence, and constant variance are satisfied.
Ideally, residuals should be randomly distributed without any clear pattern.
The p-value helps determine whether an independent variable has a significant effect on the
dependent variable. A p-value less than 0.05 generally indicates that the variable is statistically
significant.
The F-test evaluates whether the regression model as a whole is statistically significant. It
checks if at least one independent variable affects the dependent variable.
VIF is used to detect multicollinearity among independent variables. A high VIF value indicates
that variables are highly correlated, which can distort regression results.
Example: VIF > 10 suggests serious multicollinearity.
Outliers are extreme values that differ significantly from other observations. Influential points
can strongly affect the regression line. Identifying and handling them is important for
improving model accuracy.
Example: A very high sales value due to a one-time event may distort the model.
Data mining is the process of discovering useful patterns, relationships, and insights from large
volumes of data using statistical, machine learning, and analytical techniques.
In simple terms, It is the process of turning raw data into meaningful information for decision-
making.
Data mining goes beyond simple data analysis by identifying hidden patterns and trends that
are not easily visible. It plays a key role in predictive analytics by helping organizations
understand past behavior and predict future outcomes. According to Provost & Fawcett, data
mining is a core part of data-driven decision-making (DDD).
Examples
• Identifying customer buying patterns in retail
• Detecting fraudulent transactions in banking
• Predicting customer churn in telecom
CRISP-DM (Cross-Industry Standard Process for Data Mining) is the most widely used
framework for data mining projects. It provides a structured and iterative approach.
1. Business Understanding
This is the first step where the problem is defined from a business perspective. Objectives,
goals, and success criteria are clearly identified. It ensures that the data mining project is
aligned with organizational needs.
2. Data Understanding
In this stage, data is collected and explored to understand its structure, quality, and relevance.
Initial analysis is performed to identify patterns and issues such as missing values.
3. Data Preparation
This step involves cleaning, transforming, and organizing data into a usable format. It includes
handling missing data, normalization, and feature selection. It is often the most time-consuming
stage.
Example: Removing duplicate records and converting categorical data into numerical form.
4. Modeling
In this stage, various data mining techniques such as classification, regression, or clustering are
applied. Multiple models may be tested to find the best-performing one.
Example: Building a model to predict whether a customer will churn.
5. Evaluation
The model is evaluated to check whether it meets business objectives and performs accurately.
It ensures that the model is valid and reliable before implementation.
6. Deployment
The final model is implemented in a real-world environment. The results are used for decision-
making, and the model may be integrated into business systems.
Example: Using the model to identify customers likely to leave and targeting them with offers.
MODULE – 4
Regression models are statistical and machine learning techniques used to study the
relationship between dependent and independent variables. They help predict continuous
numerical outcomes such as sales, profit, demand, stock prices, customer spending, and
production levels. Advanced regression techniques improve prediction accuracy and handle
complex business problems more effectively than simple linear regression.
Advanced regression techniques are used when traditional regression models are unable to
handle non-linear relationships, multicollinearity, or large numbers of variables efficiently.
POLYNOMIAL REGRESSION
Polynomial regression is an advanced form of linear regression used when the relationship
between variables is non-linear or curved rather than straight.
In many real-world business situations, the relationship between independent and dependent
variables does not follow a straight line. Polynomial regression solves this problem by adding
higher-degree terms such as 𝑋 2 , 𝑋 3 , etc., to the regression equation.
This allows the model to fit curved patterns in the data and improve prediction accuracy.
Polynomial regression is commonly used in sales forecasting, trend analysis, and production
optimization.
Equation
𝑌 = 𝑎 + 𝑏1 𝑋 + 𝑏2 𝑋 2 + 𝑏3 𝑋 3 + ⋯ + 𝑏𝑛 𝑋 𝑛
Where:
• 𝑌= Dependent variable
• 𝑋= Independent variable
• 𝑎= Intercept
• 𝑏= Regression coefficients
Example
A company may observe that increasing advertising expenditure initially increases sales
rapidly, but after a certain point, sales growth slows down. This curved relationship can be
modeled effectively using polynomial regression.
Characteristics
• Polynomial Regression models are usually fit with the method of least squares. The
least square method minimizes the variance of the coefficients, under the Gauss
Markov Theorem.
• The errors are independent, normally distributed with mean zero and a constant
variance (OLS).
• We make our model and find out that it performs very badly,
• We observe between the actual value and the best fit line, which we predicted and it
seems that the actual value has some kind of curve in the graph and our line is no where
near to cutting the mean of the points.
• This where polynomial Regression comes to the play, it predicts the best fit line that
follows the pattern(curve) of the data, as shown in the pic below:
• Polynomial Regression does not require the relationship between the independent and
dependent variables to be linear in the data set, This is also one of the main difference
between the Linear and Polynomial Regression.
• Polynomial Regression is generally used when the points in the data are not captured
by the Linear Regression Model and the Linear Regression fails in describing the best
result clearly.
As we increase the degree in the model, it tends to increase the performance of the model.
However, increasing the degrees of the model also increases the risk of over-fitting and under-
fitting the data.
In order to find the right degree for the model to prevent over-fitting or under-fitting, we can
use:
1. Forward Selection: This method increases the degree until it is significant enough to
define the best possible model.
2. Backward Selection: This method decreases the degree until it is significant enough
to define the best possible model.
If you know what Linear Regression is then you will probably understand the maths behind the
polynomial regression too. Linear Regression is basically the first degree Polynomial.
Advantages
• Complex interpretation
1. Lasso Regression (L1 Regularization): Lasso stands for Least Absolute Shrinkage
and Selection Operator. In lasso regression, the penalty term added to the objective
function is the absolute value of the coefficients' sum. This leads to some coefficients
becoming exactly zero, effectively performing variable selection by excluding certain
features from the model.
2. Ridge Regression (L2 Regularization): Ridge regression adds a penalty term to the
objective function that is the squared sum of the coefficients. This penalizes large
coefficients but does not force them to be exactly zero, allowing all features to be
included in the model. Ridge regression is particularly useful when dealing with
multicollinearity (high correlation among predictor variables).
Ridge Regression is a version of linear regression that includes a penalty to prevent the model
from overfitting, especially when there are many predictors or not enough data.
The standard loss function (mean squared error) is modified to include a regularization term:
𝑛
Loss = MSE + 𝜆 ∑𝑖=1 𝑤𝑖2
Here, λ is the regularization parameter that controls the strength of the penalty, and wi are
the coefficients.
Example
• House size
• Number of rooms
• Property value
may be strongly correlated. Ridge regression helps stabilize the model and improve forecasting
accuracy.
Advantages
• Reduces overfitting
Limitations
Lasso regression, also known as L1 regularization, is a linear regression technique that adds
a penalty to the loss function to prevent overfitting. This penalty is based on the absolute
values of the coefficients.
Lasso regression is a version of linear regression including a penalty equal to the absolute value
of the coefficient magnitude. By encouraging sparsity, this L1 regularization term reduces
overfitting and helps some coefficients to be absolutely zero, hence facilitating feature
selection.
The standard loss function (mean squared error) is modified to include a regularization term:
Here, λ is the regularization parameter that controls the strength of the penalty, and wi are
the coefficients.
Example
A telecom company predicting customer churn may use variables such as:
• Customer age
• Call duration
• Internet usage
• Subscription type
Lasso regression may automatically eliminate variables that contribute very little to churn
prediction.
Advantages
• Improves interpretability
Limitations
The key differences between ridge and lasso regression are discussed below:
Characteristic Ridge Regression Lasso Regression
Use when you have many Use when you believe only some
predictors, all contributing to predictors are truly important (e.g.,
Example Use
the outcome (e.g., predicting genetic studies where only a few
Case
house prices where all features genes out of thousands are
like size, location, etc., matter) relevant).
Model evaluation metrics help measure the performance and accuracy of regression models.
These metrics compare predicted values with actual values to determine how well the model
performs.
The objective of Linear Regression is to find a line that minimizes the prediction error of all
the data points.
The essential step in any machine learning model is to evaluate the accuracy of the model.
The Mean Squared Error, Mean absolute error, Root Mean Squared Error, and R-Squared or
Coefficient of determination metrics are used to evaluate the performance of the model in
regression analysis.
1. R² (Coefficient of Determination)
R² measures the proportion of variation in the dependent variable explained by the regression
model.
R² = 1 → Perfect prediction
A higher R² value indicates a better model fit because more variation in the dependent variable
is explained by the independent variables.
Example
If:
R² = 0.85
then 85% of variation in sales is explained by variables such as advertising expenditure and
pricing.
RMSE squares prediction errors before averaging them. Because errors are squared, larger
errors receive greater importance.
Formula
∑(𝑦𝑖 − 𝑦̂𝑖 )2
𝑅𝑀𝑆𝐸 = √
𝑛
Where:
• 𝑦𝑖 = Actual value
• 𝑛= Number of observations
Example
If predicted sales are very different from actual sales, RMSE increases significantly.
Advantages
Limitations
• Sensitive to outliers
MAE measures the average absolute difference between predicted and actual values.
Unlike RMSE, MAE does not square errors. All prediction errors are treated equally.
MAE is easier to understand because the error is expressed in the same units as the original
data.
Formula
∑ ∣ 𝑦𝑖 − 𝑦̂𝑖 ∣
𝑀𝐴𝐸 =
𝑛
Where:
• 𝑦𝑖 = Actual value
• 𝑦̂𝑖 = Predicted value
• 𝑛= Number of observations
Example
• MAE = 500
Advantages
• Easy to interpret
Limitations
Classification models are machine learning and statistical techniques used to predict categorical
outcomes or classes. Unlike regression models, which predict continuous numerical values,
classification models predict categories such as:
• Yes / No
• Fraud / Not Fraud
• Pass / Fail
• Disease / No Disease
• Customer Churn / No Churn
Classification models are widely used in predictive analytics for fraud detection, customer
segmentation, medical diagnosis, spam filtering, and recommendation systems.
LOGISTIC REGRESSION
Logistic regression is a supervised classification algorithm used to predict the probability of a
categorical outcome, especially binary outcomes.
Although the name contains “regression,” logistic regression is mainly used for classification
problems. It predicts probabilities between 0 and 1 using a mathematical function called the
sigmoid function.
The output is usually:
• 1 → Positive class
• 0 → Negative class
The model estimates the likelihood that a particular event will occur.
Logistic regression is suitable when:
• The dependent variable is categorical
• The problem involves binary classification
Sigmoid Function
The sigmoid curve converts predicted values into probabilities.
1
𝑃(𝑌 = 1) =
1 + 𝑒 −𝑧
Where:
• 𝑃(𝑌 = 1)= Probability of positive outcome
• 𝑒= Exponential constant
• 𝑧= Linear combination of variables
2. Multinomial Logistic Regression: Used when the dependent variable has more than two
categories without any order.
Examples
• Product preference: Brand A, Brand B, Brand C
• Transportation choice: Bus, Train, Car
Healthcare
Hospitals use logistic regression for:
• Disease diagnosis
• Patient risk prediction
Example
Predicting whether a patient has diabetes.
Marketing
Businesses use logistic regression for:
• Customer churn prediction
• Purchase behavior analysis
Example
Predicting whether a customer will buy a product.
Human Resource Management
Organizations use logistic regression for:
• Employee attrition prediction
• Recruitment analysis
Advantages
• Simple and easy to interpret
• Produces probability outputs
• Works well for binary classification
• Computationally efficient
Limitations
• Assumes linear relationship in log odds
• Less effective for highly complex datasets
• Sensitive to outliers
DECISION TREES
Decision trees are one of the most widely used supervised machine learning algorithms for
classification and prediction problems. A decision tree is a tree-structured model that divides
data into branches based on conditions or decision rules. It helps organizations make
predictions and decisions in a simple, visual, and interpretable manner.
Decision trees are commonly used in:
• Loan approval prediction
• Customer segmentation
• Fraud detection
• Medical diagnosis
• Employee attrition analysis
• Marketing analytics
The model predicts outcomes by asking a sequence of questions and splitting data into smaller
subsets until a final decision is reached.
A decision tree is a predictive model that represents decisions and their possible outcomes in
the form of a tree-like structure.
The model starts with a root node and repeatedly splits data into branches based on conditions.
Each branch represents a possible outcome, and the final nodes represent predictions or
classifications.
Pruning in tree structures, particularly in the context of decision trees, is a technique used to
prevent overfitting and improve the generalization ability of the model. Overfitting occurs
when a model learns the training data too well, capturing noise and irrelevant patterns that do
not generalize well to unseen data.
Pruning involves selectively removing parts of the tree that do not contribute significantly to
its predictive accuracy or performance on unseen data. There are mainly two types of pruning
techniques:
1. Pre-Pruning: This involves setting constraints on the growth of the tree during the
construction phase. Instead of growing the tree to its maximum depth or until all leaves
are pure (contain only one class), pre-pruning techniques stop the tree growth based on
conditions such as the maximum depth of the tree, minimum number of samples
required to split a node, or minimum improvement in impurity measure (e.g., Gini
impurity, entropy) required for a split.
2. Post-Pruning (or Cost-Complexity Pruning): Post-pruning, also known as cost-
complexity pruning, involves growing the tree to its maximum size and then iteratively
removing nodes or branches that have the least impact on the tree's performance. This
is typically done by calculating a pruning parameter (often based on impurity measures
or error rates) for each subtree and removing the subtree with the smallest increase in
error rate or impurity when pruned. The process continues until further pruning does
not improve the model's performance or until a predefined stopping criterion is met.
RANDOM FORESTS
Leo Breiman developed an extension of decision trees called random forests. There is publicly
available software for this method.
Random forest is an ensemble learning method that combines multiple decision trees to
improve prediction accuracy.
Instead of depending on one decision tree, random forest creates several trees using different
subsets of data and variables. Each tree gives a prediction, and the final prediction is based on
majority voting.
For classification tasks, the output of the random forest is the class selected by most trees. For
regression tasks, the mean or average prediction of the individual trees is returned (is typically
the average prediction of all individual trees in the ensemble.). Random decision forests correct
for decision trees' habit of overfitting to their training set and improves model stability.
Example
An e-commerce company predicts whether a customer will purchase a product based on:
• Browsing history
• Purchase behavior
• Product preferences
Multiple decision trees make predictions, and the final output is determined through majority
voting.
Advantages
• High prediction accuracy
• Reduces overfitting
• Handles large datasets efficiently
• Works well with missing values
Limitations
• Computationally expensive
• Difficult to interpret
• Requires more memory and processing power
MODEL EVALUATION METRICS (ACCURACY, PRECISION, RECALL, F1
SCORE)
Model evaluation metrics are used to measure how well a classification model performs. After
building a classification model such as logistic regression, decision tree, or random forest, it is
important to evaluate whether the model is making correct predictions.
These metrics compare the predicted values with the actual values and help determine the
accuracy and reliability of the model.
Confusion Matrix
Predicted Class
Positive Negative
Actual Positive TP FN
Actual Negative FP TN
Where:
• TP (True Positive) → Model correctly predicts positive class
• TN (True Negative) → Model correctly predicts negative class
• FP (False Positive) → Model incorrectly predicts positive class
• FN (False Negative) → Model incorrectly predicts negative class
1. Accuracy
Accuracy measures the overall correctness of the classification model. It shows the proportion
of total predictions that were classified correctly.
Formula
𝑇𝑃 + 𝑇𝑁
𝐴𝑐𝑐𝑢𝑟𝑎𝑐𝑦 =
𝑇𝑃 + 𝑇𝑁 + 𝐹𝑃 + 𝐹𝑁
Accuracy calculates how many predictions made by the model are correct out of all predictions.
A higher accuracy value indicates better model performance. However, accuracy may become
misleading when the dataset is imbalanced. For example, if 95% of observations belong to one
class, the model may achieve high accuracy simply by predicting the majority class.
Example
Suppose a model predicts:
90 correct predictions out of 100 total predictions
Then:
Accuracy = 90%
Advantages
• Easy to understand
• Useful for balanced datasets
• Measures overall model performance
Limitations
• Misleading for imbalanced datasets
• Does not distinguish between types of errors
2. Precision
Precision measures how many predicted positive observations are actually positive.
Formula
𝑇𝑃
𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛 =
𝑇𝑃 + 𝐹𝑃
Precision focuses on the quality of positive predictions.
A high precision value means:
• Few false positives
• More reliable positive predictions
Precision becomes important when false positive errors are costly.
Example
In email spam detection:
• High precision means emails classified as spam are truly spam.
In fraud detection:
• High precision ensures that flagged transactions are actually fraudulent.
Advantages
• Reduces false alarms
• Important in fraud detection and spam filtering
Limitations
• Does not consider false negatives
3. Recall
Recall measures how many actual positive cases are correctly identified by the model.
Formula
𝑇𝑃
𝑅𝑒𝑐𝑎𝑙𝑙 =
𝑇𝑃 + 𝐹𝑁
Recall focuses on detecting all positive cases.
A high recall value means:
• Fewer false negatives
• Most positive cases are identified successfully
Recall is important when missing positive cases is risky.
Example
In disease diagnosis:
• High recall ensures that most patients with the disease are detected.
In fraud detection:
• High recall helps identify most fraudulent transactions.
Advantages
• Reduces missed positive cases
• Useful in healthcare and security applications
Limitations
• High recall may increase false positives
4. F1 Score
Example
In medical diagnosis:
• F1 Score helps balance:
o Correct disease detection
o Avoiding unnecessary alarms
Advantages
• Balances precision and recall
• Effective for imbalanced datasets
Limitations
• Slightly difficult to interpret
• Does not consider true negatives directly
Comparison of Evaluation Metrics
Time Series Analysis: Components of time series data. Meaning of Stationarity &
differencing, ARIMA models.
Basics of Big Data: Meaning, Characteristic’s (5V’s) and Sources of Big Data. Handling
Large Data Sets - Techniques and Challenges. Introduction to Hadoop and Spark.
Time Series Analysis is a way of studying the characteristics of the response variable
concerning time as the independent variable. To estimate the target variable in predicting or
forecasting, use the time variable as the reference point. TSA represents a series of time-based
orders, it would be Years, Months, Weeks, Days, Horus, Minutes, and Seconds. It is an
observation from the sequence of discrete time of successive intervals. Since TSA involves
producing the set of information in a particular sequence, this makes it distinct from spatial and
other analyses.
Time series analysis is a statistical technique used to analyze data collected over a period of
time at regular intervals. It helps identify patterns, trends, and future movements in data. Time
series analysis is widely used in predictive analytics for forecasting sales, stock prices, weather
conditions, demand, and economic indicators.
Time series data refers to observations recorded sequentially over time, such as daily, monthly,
quarterly, or yearly data.
• Trend: The long-term, sustained direction of the data. It represents continuous, gradual
increases or decreases over an extended period (e.g., an overall increase in global
population or a gradual decline in traditional cable subscriptions).
• Seasonality: Predictable, repeating patterns tied to specific, fixed timeframes. These
fluctuations occur at regular intervals (daily, weekly, monthly, or quarterly) and are
usually calendar-related (e.g., a spike in retail sales every December or increased ice
cream purchases every summer).
• Cyclical Variations: Long-term fluctuations that follow an up-and-down pattern,
typically spanning multiple years. Unlike seasonality, cycles are not fixed in duration
and are usually linked to broader macroeconomic factors or business cycles (e.g.,
economic recessions and expansions).
• Irregular Variations (Noise): Random, unpredictable short-term fluctuations that do
not fit a discernible pattern. These are unexplainable shifts caused by one-time events
(e.g., a sudden drop in tourism due to an unexpected weather event or a spike in a
specific stock price due to a sudden news announcement).
Stationarity
A time series has stationarity if a shift in time doesn’t cause a change in the shape of the
distribution. Basic properties of the distribution like the mean, variance and covariance are
constant over time.
Types of Stationary
• Strict stationarity means that the joint distribution of any moments of any degree
(e.g. expected values, variances, third order and higher moments) within the process
is never dependent on time. This definition is in practice too strict to be used for any
real-life model.
• First-order stationarity series have means that never changes with time. Any other
statistics (like variance) can change.
• Second-order stationarity (also called weak stationarity) time series have a constant
mean, variance and an autocovariance that doesn’t change with time. Other statistics in
the system are free to change over time. This constrained version of strict stationarity
is very common.
It can be difficult to tell if a model is stationary or not. Unlike the obvious example showing
seasonality above, you usually can’t tell by looking at a graph. If you aren’t sure about the
stationarity of a model, a hypothesis test can help.
• Unit root tests (e.g. Augmented Dickey-Fuller (ADF) test or Zivot-Andrews test),
• The Priestley-Subba Rao (PSR) Test or Wavelet-Based Test, which are less common
tests based on spectrum analysis.
Importance of Stationarity
Most forecasting methods assume that a distribution has stationarity. For example,
autocovariance and autocorrelations rely on the assumption of stationarity. An absence of
stationarity can cause unexpected or bizarre behaviors, like t-ratios not following a t-
distribution or high r-squared values assigned to variables that aren’t correlated at all.
Transforming Models
Most real-life data sets just aren’t stationary. To quote Thomson (1994):
“Experience with real-world data, however, soon convinces one that both stationarity and
Gaussianity are fairy tales invented for the amusement of undergraduates.”
To put it another way, if you’ve got a real-life data set (and not a theoretical one from a class),
you’re going to need to make it stationary in order to get any useful predictions from it. A
model can sometimes be stationarized through mathematical transformation (usually
performed by software) which makes the model relatively easy to predict; It will have the same
statistical properties at a later date. The mathematical transformations are then reversed so that
the new model predicts the behavior of the original time series model. Transformations might
include:
• Difference the data: differenced data has one less point than the original data. For
example, given a series Zt you can create a new series Yi = Zi – Zi – 1.
• Fit a curve to the data, then model the residuals from that curve.
• Take the logarithm or square root (usually works for data with non-constant variance).
Some models can’t be transformed in this way — like models with seasonality. These can
sometimes be broken down into smaller pieces (a process called stratification) and individually
transformed. Another way to deal with seasonality is to subtract the mean value of the periodic
function from the data.
Differencing
Differencing is where your data has one less data point than the original data set; You’re
subtracting (or moving) a point—a “difference”. For example, given a series Zt you can create
a new series Yi = Zi – Zi – 1. As well as its general use in transformations, differencing is
widely used in time series analysis.
“A series with no deterministic component which has a stationary,
invertible ARMA representation after differencing d times is said to be integrated of order d”
In simple terms, Differencing is a technique used to make non-stationary data stationary.
It removes trends and seasonality by calculating the difference between consecutive
observations.
The first difference is calculated as:
𝑌𝑡′ = 𝑌𝑡 − 𝑌𝑡−1
Where:
• 𝑌𝑡 = current value
• 𝑌𝑡−1 = previous value
′
𝑌𝑡′′ = 𝑌𝑡′ − 𝑌𝑡−1
or directly,
ARIMA is one of the most widely used statistical models for time series forecasting. It helps
predict future values based on past observations. ARIMA models are especially useful when
data shows patterns such as trends over time.
Components of ARIMA
The autoregressive part means that the current value depends on previous values of the same
series.
If previous observations influence current observations, the series has autoregressive behavior.
For example:
• Today's sales may depend on yesterday’s sales.
• Today's stock price may depend on past stock prices.
Mathematical Representation:
For AR(1):
𝑌𝑡 = 𝑐 + 𝜙𝑌𝑡−1 + 𝑒𝑡
Where:
• 𝑌𝑡 = Current value
• 𝑌𝑡−1 = Previous value
• 𝑝ℎ𝑖= AR coefficient
• 𝑒𝑡 = Error term
Example
If a retail store had high sales yesterday, it is likely to have relatively high sales today as well.
Value of d Meaning
d=0 Already stationary
d=1 First differencing applied
d=2 Second differencing applied
Usually, d = 1 is sufficient.
Method Purpose
ACF Helps identify MA(q)
PACF Helps identify AR(p)
5. Build ARIMA Model: The appropriate ARIMA(p,d,q) model is selected.
6. Evaluate the Model: Metrics used are RMSE, MAE, AIC, BIC
7. Forecast Future Values: The model predicts future observations.
Advantages of ARIMA
Limitations of ARIMA
MEANING
Big Data refers to extremely large and complex datasets that cannot be efficiently processed,
stored, or analyzed using traditional data processing tools and database systems. Big Data
includes structured, semi-structured, and unstructured data generated from various digital
sources such as social media, sensors, websites, and business transactions.
Big Data analytics helps organizations extract meaningful insights, identify patterns, improve
decision-making, and support predictive analytics.
CHARACTERISTIC’S (5V’S)
Big data is defined by its five foundational characteristics. Together, the "5 Vs" explain how
organizations capture, manage, and transform massive, complex, and fast-moving information
into actionable business insights.
• Volume: Refers to the sheer scale and massive amount of data generated every second.
It includes terabytes and petabytes of information that cannot be processed by
traditional databases.
• Velocity: The rapid speed at which data is created, gathered, and processed. Real-time
data processing is critical for handling continuous, fast-flowing information like
financial transactions or social media streams.
• Variety: Represents the different types and formats of data. It includes traditional
structured data (e.g., SQL tables), semi-structured data (e.g., JSON logs), and
unstructured data (e.g., audio, video, and emails).
• Veracity: Measures the accuracy, reliability, and trustworthiness of the data. Dealing
with veracity involves cleaning out noise, biases, and duplicates to ensure the data is of
high quality.
• Value: The ultimate goal of big data. It refers to an organization's ability to extract
meaningful, profitable, and strategic insights from the vast amounts of raw data.
V Meaning Example
Social media platforms generate massive amounts of data continuously. Platforms like
Facebook, Instagram, Twitter, LinkedIn, and TikTok record posts, likes, shares, comments,
videos, and user profiles.
Example: Over 500 million tweets are posted daily. Businesses analyze this to understand
trends, customer behavior, and preferences. Social media is one of the most prominent sources
of big data today.
Connected devices like smartwatches, fitness trackers, smart TVs, home assistants, and
connected cars generate constant streams of data. Sensors track locations, activities, and device
performance.
Example: A smart fridge monitors food inventory and usage patterns. With billions of
connected devices worldwide, IoT is among the world’s biggest sources of big data due to its
real-time updates.
3. Healthcare Systems
Hospitals, clinics, and wearable devices generate terabytes of medical data every day. This
includes patient records, diagnostic images, and device readings.
Example: MRI scans and ECG readings are recorded for millions of patients. Healthcare relies
on these sources of big data for disease prediction, personalized treatment, and research.
4. Financial Transactions
Every online purchase, card swipe, or stock trade produces high-frequency data. Banks and
financial institutions track this information for security and analysis.
Example: Visa, Mastercard, and Paytm record millions of transactions daily. Financial data is
one of the world’s biggest sources of big data because of continuous, high-speed activity.
5. E-Commerce Platforms
Online marketplaces track clicks, searches, reviews, and purchases. This helps businesses offer
personalized recommendations.
Example: Amazon monitors which products users view and their browsing time. E-commerce
platforms remain key sources of big data for understanding customer behavior.
6. Telecommunication Networks
Telecom providers gather call records, SMS logs, and internet usage data.
Example: Providers track network usage to identify dropped calls and optimize coverage.
Telecom data is an important source of big data for infrastructure planning.
Governments generate huge datasets, including census information, tax records, and vehicle
registrations.
Example: India’s Aadhaar system records biometric and demographic details for over a billion
people. Such records are critical sources of big data for policy-making and planning.
8. Education Systems
Student performance, online course engagement, and learning platform activity generate data
continuously.
Example: MOOCs record user progress and participation. Education data is a key source of
big data for improving teaching and learning experiences.
Retail stores collect information from barcode scans, loyalty programs, and customer footfall.
Example: Walmart processes over 1 million transactions every hour. Retail data is an important
source of big data for inventory management and sales prediction.
Search engines record billions of queries daily, reflecting user interests and trends.
Example: Google processes over 3.5 billion searches every day. Search data is one of the
world’s biggest sources of big data due to its volume, speed, and global coverage.
Data from GPS tracking, ride-hailing apps, and airline bookings provide insights for route
optimization.
Example: Uber collects ride and location data from millions of users daily. Transport data is a
vital source of big data for operational efficiency.
Example: Netflix uses user data to recommend shows. Media and entertainment contribute as
significant sources of big data.
Example: NASA and ISRO track global weather patterns. These measurements are key
sources of big data for forecasting and disaster planning.
Example: Automotive factories track assembly line operations to reduce errors. Industrial data
is a crucial source of big data for predictive maintenance.
Emails and messaging apps create enormous volumes of text, attachments, and usage records.
Example: Over 300 billion emails are sent daily worldwide. Communication data is a
prominent source of big data for analysis and automation.
Handling large datasets is one of the major challenges in Big Data analytics because traditional
systems are often unable to process massive volumes of structured and unstructured data
efficiently. To manage and analyze Big Data effectively, organizations use several advanced
techniques and technologies.
1. Distributed Computing
Distributed computing is a technique in which large datasets are divided and processed across
multiple computers or servers instead of relying on a single machine. Each computer in the
network performs a portion of the processing task, and the results are combined later. This
technique improves processing speed, scalability, and reliability. Distributed computing is
widely used in Big Data frameworks such as Hadoop. For example, e-commerce companies
process millions of customer transactions simultaneously using distributed systems.
2. Parallel Processing
Parallel processing refers to the execution of multiple computational tasks at the same time.
Instead of processing data sequentially, large datasets are divided into smaller parts and
processed simultaneously by multiple processors. This significantly reduces computation time
and improves efficiency. Parallel processing is commonly used in scientific research, banking
systems, and real-time analytics applications where fast processing is essential.
3. Cloud Computing
Cloud computing provides scalable storage and processing capabilities through internet-based
services. Organizations can store and process large datasets using cloud platforms such as
AWS, Microsoft Azure, and Google Cloud without investing heavily in physical infrastructure.
Cloud computing offers flexibility, scalability, cost efficiency, and remote accessibility.
Businesses use cloud computing to handle increasing data volumes and perform Big Data
analytics efficiently.
4. Data Compression
Data compression is a technique used to reduce the storage space required for large datasets.
Compression algorithms minimize file size while preserving important information. This helps
organizations save storage costs and improve data transfer speed. Data compression is
especially useful for multimedia files such as videos, images, and audio data generated by
streaming platforms and social media applications.
5. NoSQL Databases
NoSQL databases are designed to handle unstructured and semi-structured data efficiently.
Unlike traditional relational databases, NoSQL databases can store large volumes of diverse
data formats such as text, images, videos, and sensor data. Databases such as MongoDB,
Cassandra, and HBase are commonly used in Big Data environments because they provide
high scalability and flexibility.
6. Data Partitioning
Data partitioning involves dividing a large dataset into smaller, manageable segments called
partitions. Each partition can be processed independently, improving system performance and
reducing processing time. Partitioning also enhances scalability and simplifies data
management. Large organizations often partition customer or transaction data based on region,
time, or category.
7. In-Memory Processing
In-memory processing stores data in the system’s main memory (RAM) instead of reading it
repeatedly from disk storage. This significantly increases processing speed and supports real-
time analytics. Technologies such as Apache Spark use in-memory processing to perform faster
computations compared to traditional disk-based systems.
8. Data Sampling
Data sampling is a technique where a smaller representative subset of a large dataset is selected
for analysis. Instead of processing the entire dataset, analysts use samples to identify patterns
and trends more quickly. Sampling reduces computation time and is useful in predictive
analytics and statistical analysis when processing the full dataset is not feasible.
9. Batch Processing
Batch processing involves processing large volumes of data in groups or batches at scheduled
intervals instead of processing data continuously in real time. This technique is useful for
handling repetitive and large-scale processing tasks such as payroll systems, banking
transactions, and report generation.
Stream processing is used for analyzing continuously generated real-time data. Instead of
storing data first and analyzing later, stream processing analyzes data instantly as it is
generated. This technique is widely used in stock market analysis, fraud detection, social media
monitoring, and IoT applications where immediate insights are required.
CHALLENGES IN HANDLING LARGE DATA SETS
Handling large datasets is a major challenge in Big Data analytics because the size, speed, and
complexity of data exceed the capabilities of traditional data processing systems. Organizations
face several technical, operational, and security-related difficulties while storing, processing,
and analyzing Big Data effectively.
One of the biggest challenges is storing massive amounts of data generated every second from
multiple sources such as social media, IoT devices, banking systems, and e-commerce
platforms. Traditional storage systems are often insufficient to handle such huge volumes of
data. Organizations require scalable and distributed storage infrastructure, which can be
expensive and complex to manage.
Big Data is generated at very high speed, and organizations often require real-time or near real-
time analysis for decision-making. Processing such large datasets quickly is challenging
because it demands high computational power and advanced processing frameworks. Delays
in processing may reduce the usefulness of insights, especially in applications like fraud
detection and stock market analysis.
Big Data frequently contains sensitive information such as financial records, customer details,
healthcare information, and business transactions. Protecting this data from cyberattacks,
unauthorized access, and data breaches is a major challenge. Organizations must implement
strong security measures, encryption techniques, and privacy policies to ensure data protection
and regulatory compliance.
5. Scalability Challenges
As organizations continue generating more data, systems must be capable of scaling efficiently
to handle increasing workloads. Expanding storage capacity and processing infrastructure
without affecting performance is a difficult task. Poor scalability may lead to slower processing
and system failures.
Big Data comes from multiple sources and exists in different formats such as text, images,
videos, audio files, sensor data, and databases. Integrating these diverse types of data into a
unified system for analysis is a complex process. Managing structured, semi-structured, and
unstructured data together requires specialized technologies and tools.
Managing Big Data requires advanced hardware, distributed computing systems, cloud
platforms, and skilled professionals. Setting up and maintaining such infrastructure can be
expensive, especially for small and medium-sized organizations. Continuous upgrades and
maintenance also increase operational costs.
Many businesses require instant processing and analysis of continuously generated streaming
data. Handling real-time data efficiently is challenging because it requires low-latency systems,
fast networks, and powerful processing frameworks. Applications such as online
recommendations, traffic monitoring, and financial trading depend heavily on real-time
analytics.
Big Data technologies such as Hadoop, Spark, machine learning, and cloud computing require
specialized technical knowledge. Many organizations face difficulty finding skilled data
scientists, analysts, and Big Data engineers who can effectively manage and analyze large
datasets.
Organizations must follow legal and regulatory requirements related to data usage, privacy, and
storage. Ensuring compliance with regulations while handling huge datasets is challenging.
Improper data governance can result in legal penalties, reputational damage, and loss of
customer trust.
INTRODUCTION TO HADOOP
Hadoop is an open-source framework developed for storing and processing extremely large
datasets across multiple computers in a distributed environment. It was designed to handle Big
Data efficiently when traditional database systems became insufficient for managing massive
volumes of structured and unstructured data.
Hadoop allows organizations to divide large datasets into smaller parts and process them
simultaneously using clusters of computers. This distributed approach improves processing
speed, scalability, and fault tolerance. Hadoop is widely used in industries such as banking,
healthcare, retail, telecommunications, and social media analytics.
One of the major advantages of Hadoop is its ability to process huge amounts of data at low
cost using commodity hardware instead of expensive high-end systems. Hadoop also supports
scalability, meaning additional systems can be added easily as data volume increases.
Working of Hadoop
The working of Hadoop involves storing data in HDFS and processing it using MapReduce
across multiple machines.
Large Data
↓
Stored in HDFS
↓
Processed using MapReduce
↓
Results Generated
For example, a bank analyzing millions of customer transactions for fraud detection can
distribute the processing tasks across several computers using Hadoop.
Features of Hadoop
Hadoop provides several important features that make it suitable for Big Data analytics:
• Scalability – New systems can be added easily to handle increasing data volumes.
• Fault Tolerance – Data replication ensures reliability even if systems fail.
• Distributed Processing – Data is processed across multiple machines simultaneously.
• Cost Efficiency – Uses low-cost commodity hardware.
• Flexibility – Can process structured, semi-structured, and unstructured data.
Advantages of Hadoop
Hadoop enables organizations to process huge datasets efficiently and economically. It supports
parallel processing, high scalability, and distributed storage. Hadoop is also highly reliable
because data is replicated across multiple systems. It is widely used for data mining, predictive
analytics, recommendation systems, and customer behavior analysis.
Limitations of Hadoop
Although Hadoop is powerful, it has some limitations. Hadoop MapReduce processing can be
slower because it relies heavily on disk storage. It is also complex to configure and manage.
Real-time processing is difficult in Hadoop, making it less suitable for applications requiring
instant analytics.
INTRODUCTION TO SPARK
Apache Spark is an open-source Big Data processing framework designed for high-speed and
real-time data analytics. Spark was developed to overcome some limitations of Hadoop
MapReduce, especially processing speed.
Unlike Hadoop, Spark performs in-memory processing, meaning data is processed directly in
RAM instead of repeatedly reading from disk storage. This makes Spark significantly faster
than Hadoop MapReduce for many analytics tasks.
Spark is widely used in:
• Real-time analytics
• Machine learning
• Fraud detection
• Recommendation systems
• Streaming data analysis
Components of Spark
Spark consists of several important modules that support advanced analytics and data
processing.
1. Spark Core
Spark Core is the main processing engine responsible for task scheduling, memory
management, and distributed processing.
2. Spark SQL
Spark SQL is used for processing structured and semi-structured data using SQL queries.
3. MLlib
MLlib is Spark’s machine learning library used for predictive analytics, classification,
clustering, and regression tasks.
4. Spark Streaming
Spark Streaming processes real-time streaming data such as stock market transactions, sensor
data, and social media feeds.
5. GraphX
GraphX is used for graph processing and network analysis.
Features of Spark
Spark provides several advanced features that make it highly popular in Big Data analytics.
• High-Speed Processing – Faster than Hadoop MapReduce because of in-memory
computation.
• Real-Time Analytics – Supports streaming and real-time data processing.
• Machine Learning Support – Includes built-in ML libraries.
• Ease of Use – Supports programming languages such as Python, Java, Scala, and R.
• Scalability – Handles large-scale distributed data processing efficiently.
Advantages of Spark
Spark offers much faster processing compared to Hadoop MapReduce and is highly suitable
for real-time analytics applications. It supports advanced machine learning and streaming
analytics, making it ideal for predictive analytics and AI applications. Spark is also easier to
use because it supports multiple programming languages.
Limitations of Spark
Spark requires large memory resources because it performs in-memory processing. Managing
very large datasets entirely in memory can be expensive. Spark can also be more complex when
handling extremely large distributed systems.
Difference Between Hadoop and Spark
Predictive Analytics Tools: Basic features and comparison between Python, R, SAS, SPSS,
Tableau, Power BI, Google Big Query, Azure ML Studio.
Ethical issues in predictive analytics: Data privacy and security. Bias & Fairness in
predictive models
Predictive analytics tools are software platforms and programming environments used for data
analysis, machine learning, forecasting, statistical modeling, visualization, and business
intelligence. These tools help organizations identify patterns from historical data and make
accurate future predictions for better decision-making.
Python
1. Simple and Easy Syntax: Python uses simple English-like commands, making it easier
to learn and write programs compared to many programming languages.
2. Large Collection of Libraries: Python provides powerful libraries for analytics and
machine learning, including:
• Pandas → Data manipulation
• NumPy → Numerical computing
• Matplotlib → Data visualization
• Scikit-learn → Machine learning
• TensorFlow → Deep learning
3. Machine Learning and AI Support: Python is highly suitable for advanced analytics,
artificial intelligence, and neural networks.
4. Big Data Integration: Python integrates easily with Hadoop, Spark & Cloud
platforms. This makes it useful for handling massive datasets.
5. Cross-Platform Compatibility: Python works on Windows, Linux & macOS
6. Automation Support: Python can automate repetitive analytical tasks such as Data
extraction, Report generation & Forecasting
Applications
• Fraud detection
• Customer segmentation
• Sales forecasting
• Recommendation systems
• Time series analysis
Advantages
• Open-source and free
• Flexible and scalable
• Strong community support
• Suitable for advanced predictive analytics
Limitations
• Requires programming knowledge
• Slightly slower execution speed compared to compiled languages
R
Major Features of R
Advantages
• Excellent statistical capabilities
• Strong data visualization support
• Free and flexible
Limitations
• Difficult for beginners
• Slower with extremely large datasets
SAS
SAS (Statistical Analysis System) is a commercial software suite used for advanced analytics,
predictive modeling, business intelligence, and data management. It is widely used by large
organizations because of its reliability, scalability, and security.
Major Features of SA
Applications
• Risk management
• Fraud analytics
• Healthcare analytics
• Customer behavior analysis
Advantages
• Highly reliable
• Strong technical support
• Excellent for enterprise analytics
Limitations
• Expensive licensing cost
• Requires specialized training
SPSS
SPSS (Statistical Package for the Social Sciences) is a statistical software developed mainly
for data analysis, survey analysis, and research applications. It is widely used in academic,
marketing, and social science research.
SPSS is known for its user-friendly graphical interface and minimal coding requirements.
1. Easy-to-Use Interface: Users can perform analysis using menus and dialog boxes
instead of programming.
2. Statistical Analysis Tools: Supports Regression analysis, Correlation analysis,
Hypothesis testing & ANOVA
3. Data Management: Allows Data cleaning, Data transformation & Missing value
handling
4. Report Generation: SPSS generates tables, charts, and statistical reports
automatically.
5. Survey Data Analysis: Highly useful for questionnaire and survey analysis.
Applications
• Market research
• Academic research
• Employee surveys
• Consumer behavior studies
Advantages
• Beginner-friendly
• Minimal programming needed
• Good for statistical analysis
Limitations
• Limited advanced AI capabilities
• Expensive commercial software
Tableau
Tableau is a powerful business intelligence and data visualization tool used to create interactive
dashboards, charts, and reports.
It helps organizations convert raw data into understandable visual insights for better decision-
making.
1. Interactive Dashboards: Users can create dynamic dashboards with filters and drill-
down analysis.
2. Drag-and-Drop Interface: No advanced programming knowledge is required.
3. Real-Time Analytics: Supports real-time data updates and monitoring.
4. Multiple Data Source Integration: Connects with Excel, SQL databases, Cloud
platforms & Hadoop
5. Advanced Visualization: Creates Heat maps, Geographic maps, Trend charts & KPI
dashboards
Applications
• Sales analysis
• Performance tracking
• Business reporting
• Data storytelling
Advantages
• Excellent visualization quality
• Easy to use
• Fast dashboard development
Limitations
• Limited advanced predictive modeling
• Expensive enterprise version
Power BI
Power BI integrates strongly with Microsoft products such as Excel, Azure, and SQL Server.
Applications
• Financial reporting
• Sales analytics
• KPI monitoring
• Business intelligence
Advantages
• User-friendly
• Affordable
• Excellent Microsoft integration
Limitations
• Less flexible than Python for advanced analytics
• Performance limitations with very large datasets
Google BigQuery
1. Massive Data Processing: Can process terabytes and petabytes of data quickly.
2. Serverless Architecture: No infrastructure management required.
3. High-Speed SQL Queries: Supports extremely fast analytical queries.
4. Cloud Scalability: Automatically scales based on workload.
5. Integration with Google Cloud: Works with: Google Cloud Storage, AI tools &
Machine learning services
Applications
• Big Data analytics
• Real-time business intelligence
• Customer analytics
• Predictive modeling
Advantages
• Very fast processing
• Scalable cloud infrastructure
• Easy Big Data management
Limitations
• Cloud dependency
• Cost may increase with large query usage
Azure ML Studio
Applications
• Predictive analytics
• AI model deployment
• Classification and forecasting
• Business intelligence
Advantages
• Easy model deployment
• Scalable cloud platform
• Supports automated machine learning
Limitations
• Requires Azure knowledge
• Internet dependency for cloud access
Google Azure ML
Feature Python R SAS SPSS Tableau Power BI
BigQuery Studio
Cloud
Statistical Commercial Data Business
Programming Statistical Cloud Data Machine
Type of Tool Programming Analytics Visualization Intelligence
Language Software Warehouse Learning
Language Software Tool Tool
Platform
Python
SAS Tableau
Developed By Software R Foundation IBM Microsoft Google Microsoft
Institute Software
Foundation
Machine
Statistical Visualization Machine
Primary Learning & Statistical Enterprise Business Big Data
& Survey & Learning &
Purpose Data Analysis Analytics Intelligence Analytics
Analysis Dashboards AI
Analytics
Ease of Moderate to
Moderate Moderate Easy Easy Easy Moderate Moderate
Learning Difficult
Programming SQL
Yes Yes Limited Very Little No Minimal Minimal
Required Knowledge
Pay-as- Subscription
Cost Free Free Expensive Expensive Expensive Affordable
you-use Based
Statistical
Very Good Excellent Excellent Excellent Limited Moderate Moderate Good
Analysis
Machine
Basic AI
Learning Excellent Very Good Good Limited Limited Good Excellent
Features
Support
Data
Good Excellent Good Moderate Excellent Excellent Moderate Moderate
Visualization
Big Data
Excellent Moderate Very Good Limited Moderate Moderate Excellent Excellent
Handling
Cloud
Good Moderate Good Limited Good Excellent Excellent Excellent
Integration
Real-Time
Good Limited Moderate Limited Good Good Excellent Excellent
Analytics
User Coding- Coding- GUI + Drag-and- Drag-and- Web Drag-and-
GUI-Based
Interface Based Based Coding Drop Drop Interface Drop
Scalability High Moderate High Moderate Moderate High Very High Very High
Large-
Advanced
Statistical Enterprise Academic Dashboards Business Scale Predictive
Best For Analytics &
Research Analytics Research & Reports Intelligence Cloud Modeling
AI
Analytics
Industries IT, Finance, Research, Banking, Academics, Business Corporate Big Data AI & Cloud
Using It Healthcare Education Insurance Surveys Analytics Reporting Companies Analytics
Ease of Massive
Flexibility & Statistical Enterprise Interactive Microsoft Cloud ML
Key Strength Statistical Data
AI Support Modeling Reliability Visualization Integration Deployment
Analysis Processing
Limited Limited
Main Requires Difficult for Limited AI Cloud
High Cost Predictive Advanced Query Cost
Limitation Coding Skills Beginners Features Dependency
Modeling Analytics
Predictive analytics helps organizations forecast future outcomes using historical data,
statistical models, and machine learning techniques. Although predictive analytics provides
many benefits in decision-making, it also creates several ethical concerns related to privacy,
security, fairness, transparency, and responsible use of data.
Ethical issues arise when predictive models misuse personal information, produce biased
decisions, or negatively affect individuals and society. Therefore, organizations must ensure
that predictive analytics systems are accurate, transparent, secure, and fair.
Data privacy refers to the protection of personal and sensitive information from unauthorized
access, misuse, or disclosure. Data security refers to the methods and technologies used to
protect data from theft, cyberattacks, and data breaches.
Predictive analytics systems often use large amounts of customer, employee, medical, financial,
and behavioral data. If this data is not handled properly, it may violate privacy rights and create
ethical and legal problems.
• Encryption: Encryption converts data into coded form so unauthorized users cannot
read it.
• Access Control: Only authorized users should be allowed to access sensitive datasets.
• Authentication Systems: Passwords, OTPs, and biometric verification help secure
systems.
• Firewalls and Antivirus Software: Protect systems from external cyber threats.
• Data Backup: Regular backups help recover data during system failures or attacks.
Organizations should:
• Obtain user consent before collecting data
• Use data only for intended purposes
• Protect sensitive information using security measures
• Follow data protection laws and regulations
• Minimize unnecessary data collection
A healthcare company using patient medical records for predictive analytics without patient
permission may violate ethical and legal standards.
Bias in predictive analytics occurs when a predictive model produces unfair or discriminatory
results toward certain individuals or groups.
Fairness means ensuring that predictive models make objective and unbiased decisions for all
users.
Bias can occur because predictive models learn patterns from historical data. If historical data
contains discrimination or imbalance, the model may continue those unfair patterns.
1. Biased Training Data: If historical data contains unfair patterns, the model learns and
repeats them.
Example: A hiring dataset dominated by male employees may cause the model to
prefer male candidates.
2. Incomplete Data: Missing or unrepresentative data can create inaccurate predictions
for certain groups.
Example: A medical prediction model trained mainly on adults may not perform well
for children.
3. Human Bias: Bias may enter the model through human decisions during:
• Data collection
• Feature selection
• Model design
4. Algorithmic Bias: Some algorithms may unintentionally favor one category over
another.
Types of Bias in Predictive Analytics
Organizations can reduce bias and improve fairness using several methods:
1. Using Diverse and Balanced Data: Training data should represent all groups fairly.
2. Regular Model Testing: Models should be tested continuously for discrimination and
unfair outcomes.
3. Transparent Algorithms: Organizations should explain how predictions are made.
4. Human Oversight: Human experts should review important predictive decisions.
5. Ethical AI Practices: Organizations should adopt responsible AI policies and fairness
standards.
An AI recruitment system trained using past hiring data may reject female applicants because
historical hiring patterns favored male candidates.
Importance of Ethics in Predictive Analytics
Ethical predictive analytics helps organizations:
• Build customer trust
• Improve transparency
• Ensure fairness
• Protect sensitive information
• Comply with legal regulations
Ethical practices also improve the reliability and social acceptance of predictive analytics
systems.