0% found this document useful (0 votes)
3 views120 pages

PA Notes

Predictive analytics utilizes historical data and statistical modeling to forecast future outcomes, significantly enhancing decision-making across various industries. It differs from descriptive and prescriptive analytics by focusing on predicting potential future events and trends, thus enabling proactive strategies. The process involves several steps, including business understanding, data collection, model building, and deployment, ultimately leading to improved operational efficiency and competitive advantage.

Uploaded by

bhavansiddesh
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views120 pages

PA Notes

Predictive analytics utilizes historical data and statistical modeling to forecast future outcomes, significantly enhancing decision-making across various industries. It differs from descriptive and prescriptive analytics by focusing on predicting potential future events and trends, thus enabling proactive strategies. The process involves several steps, including business understanding, data collection, model building, and deployment, ultimately leading to improved operational efficiency and competitive advantage.

Uploaded by

bhavansiddesh
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

MODULE – 1

Introduction to Predictive Analytics Definition and significance. Predictive vs. Descriptive


vs. Prescriptive Analytics. Overview of the predictive analytics process. Applications in
Business Case studies from various industries (e.g., finance, marketing, operations)
Discussion on the impact of predictive analytics on decision-making.

INTRODUCTION TO PREDICTIVE ANALYTICS

Predictive analytics is a branch of advanced data science that uses historical data, statistical
modeling, and machine learning techniques to forecast future events or behaviors. By
identifying patterns in data, it helps organizations anticipate risks and opportunities. Key
applications include forecasting customer needs, managing supply chains, and identifying
financial fraud.

Predictive analytics is the use of statistics and modeling techniques to forecast future outcomes.
Current and historical data patterns are examined and plotted to determine the likelihood that
those patterns will repeat.

Businesses use predictive analytics to fine-tune their operations and decide whether new
products are worth the investment. Investors use predictive analytics to decide where to put
their money. Internet retailers use predictive analytics to fine-tune purchase recommendations
to their users and increase sales.

A common misconception is that predictive analytics and machine learning are the same.
Predictive analytics help us understand possible future occurrences by analyzing the past. At
its core, predictive analytics includes a series of statistical techniques (including machine
learning, predictive modeling, and data mining) and uses statistics (both historical and current)
to estimate, or predict, future outcomes.

Thus, machine learning is a tool used in predictive analysis.

Machine learning is a subfield of computer science that means "the programming of a digital
computer to behave in a way which, if done by human beings or animals, would be described
as involving the process of learning."

DEFINITION OF PREDICTIVE ANALYTICS

Predictive analytics is the use of historical data, statistical modeling, and machine learning
techniques to forecast future outcomes, trends, and behaviors. It helps organizations anticipate
potential scenarios—from machine failures to customer actions—to drive proactive, data-
driven decisions.

SIGNIFICANCE OF PREDICTIVE ANALYTICS

1. Improves Decision-Making: Predictive analytics helps organizations make better


decisions by using data instead of guesswork. It analyzes past patterns to predict future
outcomes, allowing managers to choose actions that are more likely to succeed. This
reduces uncertainty and improves the quality and accuracy of decisions.
2. Helps in Forecasting Future Trends: Predictive analytics is useful in forecasting
future trends such as sales, customer demand, and market changes. By studying
historical data, businesses can anticipate what might happen next and plan accordingly,
which supports effective strategic planning and resource allocation.
3. Identifies Risks and Opportunities: It helps organizations identify potential risks like
customer churn, fraud, or financial losses before they occur. At the same time, it
highlights new opportunities such as untapped markets or profitable customer
segments, enabling businesses to take timely action.
4. Enhances Customer Experience: Predictive analytics improves customer experience
by understanding customer behavior and preferences. Businesses can offer personalized
products, services, and recommendations, which increases customer satisfaction,
loyalty, and retention.
5. Increases Operational Efficiency: By predicting demand and optimizing processes,
predictive analytics helps organizations use their resources efficiently. It reduces
wastage, improves inventory management, and streamlines operations, leading to lower
costs and better productivity.
6. Provides Competitive Advantage: Organizations that use predictive analytics can
respond quickly to changes and make smarter decisions compared to competitors. This
gives them an advantage in the market, helping them stay ahead in a competitive
business environment.
7. Supports Proactive Decision-Making: Predictive analytics allows businesses to act in
advance rather than reacting after problems occur. For example, companies can predict
equipment failure or customer dissatisfaction and take preventive measures, reducing
losses and improving performance.
8. Improves Financial Performance: Accurate predictions help organizations increase
revenue and reduce unnecessary expenses. By optimizing pricing, marketing strategies,
and operations, predictive analytics contributes to better profitability and overall
financial performance.
9. Enables Automation of Decisions: Predictive analytics supports automated decision-
making systems where decisions are made quickly without human intervention. This is
especially useful in areas like credit scoring, fraud detection, and online
recommendations, saving time and increasing efficiency.
10. Scalable and Applicable Across Industries: Predictive analytics can be applied across
various industries such as marketing, finance, healthcare, and operations. Its flexibility
and scalability make it a valuable tool for organizations of all sizes to improve
performance and decision-making.

Some benefits of predictive analytics to solve several kinds of business problems:

• Improved customer targeting: Analyzing data can help businesses identify and target
their ideal customers more effectively.

• Increased customer retention: Data can help businesses identify customers at risk of
churning and take steps to prevent them from leaving.

• Reduced fraud: Data can help businesses identify fraudulent transactions and prevent
them from occurring in the first place.

• Optimized operations: Data can help businesses optimize their operations, such as
supply chain and inventory management.

• Improved decision-making: Data can help businesses make better decisions by


providing actionable insights into their customers, operations, and the market.

• Increased revenue: Data can help businesses increase revenue by identifying new
opportunities, such as upselling and cross-selling to existing customers.

• Gained competitive advantage: Data can help businesses gain a competitive


advantage by providing them with metrics and insights their competitors do not have.

PREDICTIVE VS. DESCRIPTIVE VS. PRESCRIPTIVE ANALYTICS.

Descriptive Analytics
Descriptive analytics is the process of analyzing historical data to understand what has
happened in the past. It focuses on summarizing and interpreting data to provide insights into
past performance and trends. Descriptive analytics answers the question, "What
happened?"

Techniques and Tools for Descriptive Analytics

Descriptive analytics employs various techniques and tools, including:

• Data Aggregation: Combining data from multiple sources to provide a comprehensive


view.

• Data Mining: Extracting patterns and relationships from large datasets.

• Data Visualization: Using charts, graphs, and dashboards to represent data visually.

• Statistical Analysis: Applying statistical methods to summarize and describe data.

Common tools used in descriptive analytics include:

• Excel: For basic data analysis and visualization.

• Tableau: For advanced data visualization and dashboard creation.

• Power BI: For interactive data visualization and business intelligence.

• SQL: For querying and managing databases.

Applications of Descriptive Analytics

Descriptive analytics is widely used across various industries for:

• Business Reporting: Generating regular reports on sales, revenue, and other key
performance indicators (KPIs).

• Customer Segmentation: Analyzing customer data to identify different segments and


their characteristics.

• Market Analysis: Understanding market trends and consumer behavior.

• Operational Efficiency: Monitoring and improving business processes.

Predictive Analytics
Predictive analytics uses historical data and statistical algorithms to forecast future events. It
aims to predict what is likely to happen based on past trends and patterns. Predictive analytics
answers the question, "What could happen?"

Techniques and Tools for Predictive Analytics

Predictive analytics involves several techniques and tools, including:

• Regression Analysis: Modeling the relationship between dependent and independent


variables.

• Time Series Analysis: Analyzing data points collected or recorded at specific time
intervals.

• Machine Learning: Using algorithms to learn from data and make predictions.

• Classification and Clustering: Grouping data into categories or clusters based on


similarities.

Common tools used in predictive analytics include:

• R: For statistical computing and graphics.

• Python: For machine learning and data analysis libraries like scikit-learn and
TensorFlow.

• SAS: For advanced analytics, business intelligence, and data management.

• IBM SPSS: For statistical analysis and predictive modeling.

Applications of Predictive Analytics

Predictive analytics is applied in various fields, such as:

• Risk Management: Predicting potential risks and their impact on business operations.

• Customer Retention: Identifying customers at risk of churning and developing


retention strategies.

• Sales Forecasting: Estimating future sales based on historical data.

• Healthcare: Predicting disease outbreaks and patient outcomes.

Prescriptive Analytics
Prescriptive analytics goes beyond predicting future outcomes by recommending actions to
achieve desired results. It combines data, algorithms, and business rules to suggest the best
course of action. Prescriptive analytics answers the question, "What should we do?"

Techniques and Tools for Prescriptive Analytics

Prescriptive analytics utilizes various techniques and tools, including:

• Optimization: Finding the best solution from a set of feasible options.

• Simulation: Modeling complex systems to evaluate different scenarios.

• Decision Analysis: Assessing and comparing different decision options.

• Machine Learning: Using algorithms to learn from data and make recommendations.

Common tools used in prescriptive analytics include:

• Gurobi: For mathematical optimization.

• IBM ILOG CPLEX: For optimization and decision support.

• AnyLogic: For simulation modeling.

• MATLAB: For numerical computing and optimization.

Applications of Prescriptive Analytics

Prescriptive analytics is used in various industries for:

• Supply Chain Optimization: Improving inventory management and logistics.

• Revenue Management: Setting optimal pricing strategies.

• Healthcare: Recommending personalized treatment plans.

• Finance: Optimizing investment portfolios and risk management strategies.

Key Differences Between Descriptive, Predictive and Prescriptive data analytics model

While descriptive, predictive, and prescriptive analytics are interconnected, they serve different
purposes and provide different insights:

• Descriptive Analytics: Focuses on understanding past events and trends. It provides a


summary of historical data and helps identify patterns and relationships.
• Predictive Analytics: Uses historical data to forecast future events. It helps anticipate
potential outcomes and trends, enabling proactive decision-making.

• Prescriptive Analytics: Recommends actions to achieve desired outcomes. It combines


data, algorithms, and business rules to suggest the best course of action.

The key differences can be summarized as follows:

Descriptive Predictive Prescriptive

Predictive data analytical


Descriptive data analytical Prescriptive data analytical
model use statistical
model use data aggregation model use optimization and
models and forecast
and data mining to provide simulation algorithms to
techniques to understand
insight into the past. advice on possible outcomes.
the future.

It focuses on - What has It focuses on - What could It focuses on - What should


happened in the past? happen in the future? we do?

This analysis showcases


This is the analysis of the
This analysis is used to feasible solutions to a
past or historical data used
determine the future problem and the impact of
to understand the trends and
trends. considering a solution on the
estimate metrics over time.
future trend.

It is used when the user It is used when the user


It is used when the user have
want to summarize results want to make an educated
to make complex or time
for all part or a part of guess at the likely
sensitive decisions.
business. outcomes.

Tools Used - data mining, Tools Used - machine Tools Used - heuristics,
data aggregation learning, statistical model optimization

Use reactive approach. Use proactive approach. Use proactive approach.

Example:
Example:
Ecommerce businesses
Example : Identifying techniques to
which uses the customer's
Annual Revenue Report optimize the patient care in
browsing history to
the healthcare
recommend products
Descriptive, Predictive and Prescriptive data analytics are important types of analytics where
Descriptive analytics is used to summarize the data, Predictive analytics is used to make future
predictions based on the past data and Prescriptive analytics is used to identify the possible
future outcomes and show the best option.

OVERVIEW OF THE PREDICTIVE ANALYTICS PROCESS

The predictive analytics process is a systematic sequence of steps used to extract insights from
data and build models that can predict future outcomes. It ensures that predictions are accurate,
relevant, and useful for decision-making.

CRISP-DM (Cross-Industry Standard Process for Data Mining) is a standard framework used
to plan and execute data mining and predictive analytics projects. It provides a structured
approach to convert raw data into useful insights and predictions.

CRISP-DM was originally developed as a data mining process model, but its steps are general
and flexible, which makes it equally suitable for predictive analytics projects.

Both fields follow a similar objective:

• Data mining: discovering patterns in data

• Predictive analytics: using those patterns to predict future outcomes

Since predictive analytics is essentially an advanced application of data mining techniques,


CRISP-DM fits both.
1. Business Understanding

This is the first and most important step where the problem is clearly defined. Organizations
identify their objectives, such as predicting customer churn or forecasting sales. It ensures that
the analytics work is aligned with business goals.

2. Data Collection (Data Understanding)

In this stage, relevant data is gathered from various sources such as databases, customer
records, or transaction systems. The data is then explored to understand its structure, quality,
and relevance to the problem.

3. Data Preparation

This step involves cleaning and transforming the data. Missing values are handled, errors are
corrected, and useful variables (features) are selected. Proper data preparation improves the
accuracy of predictive models.

4. Model Building

In this stage, statistical and machine learning techniques such as regression, decision trees, or
classification models are applied. The goal is to develop a model that can identify patterns and
make predictions.

5. Model Evaluation
The developed model is tested to check its accuracy and reliability. Techniques like validation
and testing are used to ensure the model performs well on new, unseen data and avoids issues
like overfitting.

6. Deployment

Once the model is validated, it is implemented in a real business environment. The model is
used to make predictions that support decision-making in day-to-day operations.

7. Monitoring and Maintenance

After deployment, the model’s performance is continuously monitored. Since data and business
conditions change over time, the model may need updates or retraining to remain accurate.

DISCUSSION ON THE IMPACT OF PREDICTIVE ANALYTICS ON DECISION-


MAKING

Predictive analytics has significantly transformed organizational decision-making by enabling


a data-driven, forward-looking approach. Instead of relying on intuition or past experience
alone, organizations now use data models to make more accurate and timely decisions.

1. Shift from Intuition to Data-Driven Decisions

Predictive analytics promotes data-driven decision-making (DDD), where decisions are based
on empirical evidence rather than intuition. According to Provost and Fawcett, even small
improvements in predictive accuracy can lead to better business outcomes. This reduces bias
and enhances objectivity in managerial decisions.

2. Improved Accuracy and Reduced Uncertainty

By analyzing historical data and identifying patterns, predictive models estimate future
outcomes with higher accuracy. This reduces uncertainty in decision-making and allows
managers to make informed choices with greater confidence.

3. Enables Proactive Decision-Making

Predictive analytics allows organizations to anticipate future events and take action in advance.
Instead of reacting to problems after they occur, businesses can prevent them. For example,
predicting customer churn helps companies retain customers before they leave.

4. Faster and Real-Time Decisions


With advanced analytics tools, organizations can process large volumes of data quickly and
generate real-time predictions. This enables faster decision-making, which is critical in
dynamic environments such as finance, marketing, and operations.

5. Enhances Strategic Planning

Predictive analytics supports long-term planning by forecasting trends such as demand, market
growth, and customer behavior. It helps managers develop effective strategies and allocate
resources efficiently.

6. Supports Risk Management

Predictive analytics plays a crucial role in identifying potential risks such as fraud, credit
defaults, and operational failures. By quantifying risk probabilities, it helps organizations take
preventive measures and minimize losses (Abbott emphasizes risk-based decision-making).

7. Improves Customer-Centric Decisions

Organizations can better understand customer preferences and behavior through predictive
models. This leads to more personalized marketing strategies, improved customer satisfaction,
and stronger customer relationships.

8. Enables Automation of Decisions

Predictive analytics supports automated decision systems, where decisions are made instantly
based on model outputs. This is widely used in credit scoring, fraud detection, and
recommendation systems, improving efficiency and consistency.

9. Increases Competitive Advantage

Organizations that effectively use predictive analytics gain a competitive edge by making
smarter and faster decisions. They can identify opportunities earlier and respond to market
changes more effectively than competitors.

10. Improves Overall Business Performance

Better decisions lead to improved operational efficiency, cost reduction, and revenue growth.
Predictive analytics directly contributes to enhanced organizational performance and
profitability.
APPLICATIONS IN BUSINESS CASE STUDIES FROM VARIOUS INDUSTRIES
(E.G., FINANCE, MARKETING, OPERATIONS)
Predictive analytics is widely used across industries to improve decision-making, optimize
operations, reduce risks, and increase profitability. By analyzing historical and real-time data,
organizations can predict future trends and take proactive actions.
1. Finance Industry
Application Areas
• Credit scoring
• Fraud detection
• Risk management
• Investment forecasting

Fraud Detection by PayPal


PayPal uses predictive analytics and machine learning models to detect fraudulent transactions
in real time. The company processes millions of transactions daily and analyzes customer
behavior, transaction history, device information, and spending patterns to identify suspicious
activities.
For example, if a user suddenly makes unusually large transactions from a different location or
device, the system flags the transaction for verification. Predictive models continuously learn
from past fraud cases to improve detection accuracy.
Impact
• Reduced financial fraud losses
• Faster fraud detection
• Improved customer trust and security

2. Marketing Industry
Application Areas
• Customer segmentation
• Personalized marketing
• Customer churn prediction
• Product recommendation
Recommendation System by Amazon
Amazon uses predictive analytics to provide personalized product recommendations to
customers. The company analyzes browsing history, previous purchases, search patterns,
ratings, and customer preferences to predict products that customers are likely to buy.
Its recommendation engine generates suggestions such as “Frequently Bought Together” and
“Customers Also Bought.” These predictions significantly influence customer purchasing
decisions.
Impact
• Increased sales and cross-selling
• Improved customer experience
• Higher customer engagement and retention

3. Operations and Supply Chain


Application Areas
• Demand forecasting
• Inventory management
• Production planning
• Predictive maintenance

Demand Forecasting by Walmart


Walmart uses predictive analytics for inventory management and demand forecasting. The
company analyzes historical sales data, seasonal trends, weather conditions, and local events
to predict future demand for products.
For example, before hurricanes or storms, Walmart’s predictive models identified increased
demand for products such as bottled water and emergency supplies. This allowed stores to stock
products in advance.
Impact
• Reduced stock shortages
• Improved supply chain efficiency
• Better inventory optimization

4. Healthcare Industry
Application Areas
• Disease prediction
• Patient risk analysis
• Hospital resource planning

Disease Prediction by Mayo Clinic


Mayo Clinic uses predictive analytics to identify patients at risk of chronic diseases and
hospital readmissions. Patient records, medical history, laboratory reports, and lifestyle data
are analyzed to predict health risks.
For example, predictive models help doctors identify patients with a high probability of heart
disease or diabetes so that preventive treatment can be provided early.
Impact
• Early disease detection
• Improved patient care
• Reduced hospital readmission rates

5. Logistics and Transportation


Application Areas
• Route optimization
• Delivery forecasting
• Fleet maintenance

Route Optimization by UPS


UPS uses predictive analytics through its ORION (On-Road Integrated Optimization and
Navigation) system to optimize delivery routes. The system analyzes traffic conditions,
delivery schedules, weather, and fuel usage to predict the most efficient routes.
The predictive model helps drivers avoid unnecessary travel and delays, improving delivery
efficiency.
Impact
• Reduced fuel consumption
• Faster deliveries
• Lower operational costs
MODULE – 2

Data Collection and Preparation: Data Sources and Collection: Types of data (structured
vs. unstructured)/ Data collection methods and tools.
Data Cleaning and Preparation: Handling missing data. Data transformation and
normalization. Data Preparation Using Excel or Python/R for data cleaning and preparation.

MEANING OF DATA & DATA COLLECTION

Data is the raw form of information, a collection of facts, figures, symbols or observations that
represent details about events, objects or phenomena. By itself, data may appear meaningless,
but when organized, processed and interpreted, it transforms into valuable insights that support
decision-making, problem-solving and innovation.

• Data refers to raw facts, figures, or information that can be processed and analysed to
extract meaningful insights.

• In data science and computing, data is categorised into different types based on its
structure and nature.

• Understanding its type helps in selecting appropriate analysis and processing methods.

Data collection is the systematic process of acquiring, collecting, extracting, and storing a
voluminous amount of data from various sources to analyze, research, or make data-driven
decisions. It involves identifying the required data type—structured (organized in
rows/columns) or unstructured (raw formats like text/video)—and utilizing methods like
surveys, interviews, or automated tools (APIs, IoT sensors) to gather data for analysis in
relational databases or data lakes.

TYPES OF DATA

Data can be categorized in different ways depending on how it is collected, stored and
represented. Broadly, it falls into the following:
Quantitative Data

Quantitative data is information that can be measured, counted and expressed in numerical
form. It provides objective values that can be analyzed statistically to identify patterns, trends
and relationships.

• Represents numbers and measurable values.

• Can be divided into: Discrete data (Whole numbers) and Continuous data (Values on a
scale).

• Widely used in research, finance, engineering and business analytics.

Example: Age of people, number of customers visiting a store, temperature readings, sales
revenue.

Qualitative Data

Qualitative data is descriptive, non-numeric information that explains qualities, characteristics


or categories rather than quantities. It helps understand opinions, experiences and meanings
behind behaviors.

• Focuses on qualities, attributes and categories rather than numbers.

• Often collected through surveys, interviews or observations.

• Useful for understanding opinions, motivations and behaviors.

Example: Customer feedback (“satisfied”, “unsatisfied”), product colors, interview transcripts,


social media comments.
Structured vs unstructured data in a nutshell

Data exists in many different forms and sizes, but most can be presented as structured or
unstructured. This table has collected the main differences between these two data types.

What is structured data?

Structured data is highly organized and exists in tabular format with interconnected rows and
columns, so you can easily search for specific details and single out the relationships between
its pieces. It doesn’t normally require much storage space. To manipulate it, there is a special
language called SQL, which stands for Structured Query Language and was developed back in
the 1970s by IBM.

Structured data examples. Most of us are familiar with structured data — Google Sheets and
Microsoft Office Excel files are the first things that spring to mind. This data can comprise
both textual elements and numbers, such as employee names, contacts, ZIP codes, addresses,
credit card numbers, etc.

The typical structured data example: an Excel spreadsheet that contains information
about customers and purchases.

Pretty much everyone has dealt with booking a ticket via one of the airline reservation systems.
They operate structured data: passenger names, location names, flight numbers, number of
passengers, etc. Once you have entered information into the appropriate field, the application
saves the data and allocates it to the appropriate tables in the database.
This information will be stored, read, changed, or deleted when needed. And it’s easy to
analyze. For example, just retrieve specific columns and rows of the corresponding tables to
compare the prices or the number of tickets purchased on different dates.

What is unstructured data?

Unstructured data is schemaless, meaning it has no pre-defined structure and is stored in its
native format. This includes imagery, text documents, and video and audio files. It’s easy to
capture, provides many insights, and can be used in various ways, given that you have enough
storage to keep it and advanced technologies for analysis.

Unstructured data examples. Unstructured data includes a wide array of forms, such as
email, text files, social media posts, video, images, audio, sensor data, and so on.

The travel agency Facebook post: an example of unstructured data.

For instance, a travel agency posts new travel tours on social media and wants to know the
audience's reactions. Each post contains metadata descriptions or attributes like shares or
hashtags that can be quantified and structured. However, the post itself (pictures plus text) and
comments belong to the category of unstructured data. Collecting valuable insights will take
advanced techniques like sentiment analysis.

Semi-Structured Data

Semi-structured data combines aspects of structured and unstructured data. It does not reside
in traditional tables but still contains tags or markers that provide a loose structure.

• Provides a balance between flexibility and structure.

• Easier to analyze than unstructured data, but less rigid than structured data.

• Often used in web applications, IoT devices and log systems.

Example: JSON files, XML documents, NoSQL databases, sensor logs.


DATA SOURCES/ DATA COLLECTION METHODS AND TOOLS

The actual data is then further divided mainly into two types known as:

• Primary data

• Secondary data

Primary data

The data which is Raw, original, and extracted directly from the official sources is known as
primary data. This type of data is collected directly by performing techniques such as
questionnaires, interviews, and surveys. The data collected must be according to the demand
and requirements of the target audience on which analysis is performed otherwise it would be
a burden in the data processing.

Few methods of collecting primary data:

1. Interview method:

The data collected during this process is through interviewing the target audience by a person
called interviewer and the person who answers the interview is known as the interviewee. Some
basic business or product related questions are asked and noted down in the form of notes,
audio, or video and this data is stored for processing. These can be both structured and
unstructured like personal interviews or formal interviews through telephone, face to face,
email, etc.

2. Survey method:

The survey method is the process of research where a list of relevant questions are asked and
answers are noted down in the form of text, audio, or video. The survey method can be obtained
in both online and offline mode like through website forms and email. Then that survey answers
are stored for analyzing data. Examples are online surveys or surveys through social media
polls.

3. Observation method:

The observation method is a method of data collection in which the researcher keenly observes
the behavior and practices of the target audience using some data collecting tool and stores the
observed data in the form of text, audio, video, or any raw formats. In this method, the data is
collected directly by posting a few questions on the participants. For example, observing a
group of customers and their behavior towards the products. The data obtained will be sent for
processing.

4. Experimental method:

The experimental method is the process of collecting data through performing experiments,
research, and investigation. The most frequently used experiment methods are CRD, RBD,
LSD, FD.

• CRD - Completely Randomized design is a simple experimental design used in data


analytics which is based on randomization and replication. It is mostly used for
comparing the experiments.
• RBD - Randomized Block Design is an experimental design in which the experiment
is divided into small units called blocks. Random experiments are performed on each
of the blocks and results are drawn using a technique known as analysis of variance
(ANOVA). RBD was originated from the agriculture sector.

• LSD - Latin Square Design is an experimental design that is similar to CRD and RBD
blocks but contains rows and columns. It is an arrangement of NxN squares with an
equal amount of rows and columns which contain letters that occurs only once in a row.
Hence the differences can be easily found with fewer errors in the experiment. Sudoku
puzzle is an example of a Latin square design.

• FD - Factorial design is an experimental design where each experiment has two factors
each with possible values and on performing trail other combinational factors are
derived.

5. Focus Groups:

Small group discussions moderated to gain insights into specific topics or products.

6. Case Studies:

In-depth, detailed examination of a particular case (person, group, event).

Secondary data

Secondary data is the data which has already been collected and reused again for some valid
purpose. This type of data is previously recorded from primary data and it has two types of
sources named internal source and external source.

1. Internal source:

These types of data can easily be found within the organization such as market record, a sales
record, transactions, customer data, accounting resources, etc. The cost and time
consumption is less in obtaining internal sources.

2. External source:

The data which can’t be found at internal organizations and can be gained through external
third party resources is external source data. The cost and time consumption is more because
this contains a huge amount of data. Examples of external sources are Government
publications, news publications, Registrar General of India, planning commission,
international labor bureau, syndicate services, and other non-governmental publications.

3. Literature Review:

Analyzing existing studies, journals, and articles that can be internal or external.

3 Other sources:

• Sensors data: With the advancement of IoT devices, the sensors of these devices
collect data which can be used for sensor data analytics to track the performance and
usage of products.

• Satellites data: Satellites collect a lot of images and data in terabytes on daily basis
through surveillance cameras which can be used to collect useful information.

• Web traffic: Due to fast and cheap internet facilities many formats of data which is
uploaded by users on different platforms can be predicted and collected with their
permission for data analysis. The search engines also provide their data through
keywords and queries searched mostly.

Data Collection Tools (Instruments Used)

• Questionnaires/Survey Forms: Physical or digital instruments (e.g., Google Forms)


containing survey questions.

• Checklists: Lists of items or behaviors to be observed for quick verification.

• Rating Scales: Tools for capturing the level or intensity of opinions.

• Recording Devices: Cameras, audio recorders, and video tools for qualitative studies.

• Digital Tools/Software: Programming languages (Python, R), SQL, and data analysis
software (SPSS, Tableau).
DATA CLEANING AND PREPARATION

Data Preparation is the third stage of the predictive modeling process, intended to convert data
identified for modeling into a form that is better for the predictive modeling algorithms. Each
data set can provide different challenges to data preparation, especially with data cleansing.

The key steps in data preparation related to the columns in the data are variable cleaning,
variable selection, and feature creation. Data preparation steps related to the rows in the data
are record selection, sampling, and feature creation (again).

Variable Cleaning

Variable cleaning refers to fixing problems with values of variables themselves, including
incorrect or miscoded values, outliers, and missing values.

Incorrect Values: Incorrect values are problematic because predictive modeling algorithms
assume that every value in each column is completely correct. If any values are coded
incorrectly, the only mechanism the algorithms have to overcome these errors is to overwhelm
the errors with correctly coded values, thus making the incorrect values insignificant.

Consistency in Data Formats: A second problem to be addressed with variable cleaning is


inconsistency in data. The format of the variable values within a single column must be
consistent throughout the column. Most often, mismatches in variable value types occur when
data from multiple sources are combined into a single table.

Outliers: Outliers are unusual values that are separated from the main body of the distribution,
typically as measured by standard deviations from the mean or by the IQR. Whether or not you
remove or mitigate the influence of outliers is a critical decision in the modeling process.

If the outliers to be cleaned are examples of values that are correctly coded, there are four
typical approaches to handling them:

• Remove the outliers from the modeling data


• Separate the outliers and create separate models just for outliers
• Transform the outliers so that they are no longer outliers
• Bin the data. Some outliers are too extreme for transformations to capture them well;
they are so far from the main density of the data that they remain outliers even after
transformations have been applied. An alternative to numeric transformations is to
convert the numeric variable to categorical through binning; this approach
communicates to the modeling algorithms that the actual value of the outliers are not
important
• Leave the outliers in the data without modification. Here, the modeler could decide to
use only algorithms unaffected by outliers, such as decision trees

Multidimensional Outliers: Nearly all outlier detection algorithms in modeling software refer
to outliers in single variables. Outliers can also be multidimensional, but these are much more
difficult to identify. To mitigate the effects of multidimensional outliers, the modeler could
apply the same approaches described for single-variable outliers.

Missing Values: Missing values are arguably the most problematic of all data problems.
Missing values are typically coded in data with a null value or as an empty cell, although many
more representations can exist in data. Table 4-2 shows typical missing values you may
encounter in data.

Fixing Missing Data/ Handling Missing Data

Missing value correction is perhaps the most time-consuming of all the variable cleaning steps
needed in data preparation. Whenever possible, imputing missing values is the most desirable
action. Missing value imputation means changing values of missing data to a value that
represents a plausible or expected value in the variable if it were actually known. These are the
most commonly used methods for fixing problems associated with missing values.
Listwise and Column Deletion: The simplest method of handling missing values is listwise
deletion, meaning one removes any record with any missing values, leaving only records with
fully populated values for every variable to be used in the analysis. An alternative to listwise
deletion is column deletion: removing any variable that has any missing values at all, leaving
only variables that are fully populated. both listwise deletion and column deletion are practiced,
especially when the number of missing values is particularly large or when the timeline to
complete the modeling is particularly short.

Imputation with a Constant: This option is almost always available in predictive analytics
software. For categorical variables, this can be as simple as filling missing values with a “U”
or another appropriate string to indicate missing. For continuous variables, this is most often a
0. Sometimes, imputing with 0 causes significant problems, for example, age.

Mean and Median Imputation for Continuous Variables: The next level of sophistication
is imputing with a constant value that is not predefined by the modeler. Mean imputation is by
far the most common method, but in some circumstances, if the mean and median are different
from one another, imputing with the median may be better because the median will represent
better the most typical value of the variable. However, median imputation can be more
computationally expensive, especially if the number of records in the data is large.

Imputing with Distributions: When large percentages of values are missing, the summary
statistics are affected by mean imputation. An alternative to this is, rather than imputing with a
con stant value, to impute randomly from a known distribution. For the variable AGE, if instead
of imputing with the mean (61.6), you impute using a random number generated from a normal
distribution with mean 61.6 and standard deviation 16.6 (see Table 4-4). These imputed values
will retain the same shape as the original shape of AGE.

Random Imputation from Own Distributions: A similar procedure is called “Hot Deck”
imputation, where the “deck” referred originally to the days when the Census Bureau had cards
for individuals. If someone did not respond, one could take another card from the deck at
random and use that information as a substitute. In predictive modeling terms, this is random
imputation, but instead of using a random number generator to pick a number, a random actual
value of the variable of the non-missing values is selected. Predictive modeling software rarely
provides this option, but it is easy to do.

Imputing Missing Values from a Model: The model approach to missing value imputation
begins with changing the role of the input variable with missing values to now be a target
variable. The inputs to the new model are other input variables that may predict this new target
variable well. The training data should be large enough.

Dummy Variables Indicating Missing Values: Capturing the existence of missing data can
be done with a dummy variable coded as 1 when the variable is missing and 0 when it is
populated.

Imputation for Categorical Variables: When the data is categorical, imputation is not always
necessary; the missing value can be imputed with a value that represents missing so that no cell
contains a null any longer.

DATA TRANSFORMATION AND NORMALIZATION.

Data transformation and normalization are important steps in data preparation. They help
convert raw data into a suitable format and scale, improving the accuracy and performance
of predictive models.

DATA TRANSFORMATION

Meaning

Data transformation is the process of changing the format, structure, or values of data to
make it suitable for analysis.

Types of Data Transformation

1. Scaling

Scaling adjusts the range of data so that all variables are comparable. It prevents variables with
large values from dominating the model.

Example: Converting income from ₹10,000–₹1,00,000 into a smaller range like 0–1.
2. Encoding

Encoding converts categorical (text) data into numerical form so that it can be used in models.
Most machine learning algorithms require numeric input.

Example: Converting “Male/Female” into 0 and 1.

3. Aggregation

Aggregation combines multiple data points into a summarized form. It helps in identifying
patterns and reducing data complexity.

Example: Converting daily sales into monthly sales.

4. Feature Construction

Feature construction involves creating new variables from existing data to improve model
performance. It helps capture hidden relationships.

Example: Total spending = quantity × price.

5. Log Transformation

Log transformation reduces skewness in data by compressing large values. It helps in making
data more normally distributed.

Example: Applying log to income data where a few values are extremely high.

Purpose of Data Transformation

• Makes data suitable for algorithms


• Improves model accuracy
• Reduces complexity
• Handles skewed data

DATA NORMALIZATION

Meaning

Normalization is a technique used to rescale data into a common range, usually between 0
and 1.

Why Normalization is Important?


• Ensures all variables contribute equally
• Avoids bias due to different scales
• Example: Income (₹) vs Age (years)

Techniques of Normalization

1. Min-Max Normalization

This method rescales data to a fixed range (0 to 1) while preserving relationships between
values.

𝑥 − 𝑥𝑚𝑖𝑛
𝑥′ =
𝑥𝑚𝑎𝑥 − 𝑥𝑚𝑖𝑛

Example: Marks out of 100 converted to values between 0 and 1.

2. Z-Score Normalization (Standardization)

This method transforms data so that it has a mean of 0 and a standard deviation of 1. It is useful
when data has different units.

𝑥−𝜇
𝑧=
𝜎

Example: Standardizing heights or income for comparison.

3. Decimal Scaling

This method moves the decimal point based on the maximum value in the dataset.

Example: 500 → 0.5, 1000 → 1.0 (dividing by 1000)

Key Points

• Normalization is important for algorithms like KNN (K-Nearest Neighbours), Neural


Networks
• Prevents variables with large values from dominating
• Improves model performance and accuracy
DATA PREPARATION USING EXCEL OR PYTHON/R FOR DATA CLEANING AND
PREPARATION

Data preparation is the process of converting raw data into a clean and usable format for
analysis and predictive modeling. Tools such as Excel, Python, and R are widely used for
cleaning, transforming, and preparing data because they provide powerful functions for
handling large datasets efficiently.

1. Data Preparation Using Excel

Microsoft Excel is one of the most commonly used tools for basic data cleaning and preparation
because it is simple and user-friendly.

a. Handling Missing Data in Excel

Excel helps identify and handle missing values using filters, conditional formatting, and
formulas. Missing values can be replaced using mean, median, or manually entered values.

Example: Replacing blank sales entries with the average sales value using the AVERAGE()
function.

b. Removing Duplicate Data

Excel provides the “Remove Duplicates” feature to identify and delete repeated records. This
improves data accuracy and prevents biased analysis.

Example: Removing duplicate customer IDs from a customer database.

c. Data Sorting and Filtering

Excel allows users to sort and filter data for easier analysis and organization. Data can be
arranged in ascending or descending order.

Example: Sorting customers based on purchase value.

d. Data Transformation in Excel

Excel functions such as CONCATENATE(), LEFT(), RIGHT(), and TEXT() help transform
data into the required format.

Example: Splitting full names into first name and last name columns.
e. Normalization and Calculations

Excel formulas can be used for scaling and normalization of data. Pivot tables and charts also
support summarization and visualization.

Example: Applying Min-Max normalization to sales data.

2. Data Preparation Using Python

Python is widely used in predictive analytics because of libraries such as Pandas, NumPy, and
Scikit-learn.

a. Handling Missing Data in Python

Python provides functions like fillna() and dropna() in Pandas to handle missing values
efficiently.

Example: Replacing missing customer income values with the mean income.

[Link]([Link]())

b. Removing Duplicates

Duplicate records can be removed easily using the drop_duplicates() function.

Example: Eliminating repeated customer entries in a dataset.

df.drop_duplicates()

c. Data Transformation

Python supports encoding, aggregation, and feature engineering for predictive models.

Example: Converting categorical variables such as “Male/Female” into numerical values.

pd.get_dummies(df['Gender'])

d. Data Normalization

Scikit-learn provides preprocessing tools for normalization and scaling.

Example: Scaling values between 0 and 1.

from [Link] import MinMaxScaler


e. Data Visualization

Libraries like Matplotlib and Seaborn help visualize patterns and detect outliers.

Example: Using histograms and boxplots to identify skewed data.

3. Data Preparation Using R

R is a statistical programming language widely used for data analysis and predictive modeling.

a. Handling Missing Data in R

Functions like [Link]() and [Link]() are used to identify and remove missing values.

Example: Removing incomplete records from survey data.

[Link](data)

b. Data Transformation in R

Packages like dplyr and tidyr help in reshaping and transforming datasets.

Example: Combining multiple columns into a single variable.

c. Data Normalization in R

Normalization techniques can be applied using built-in functions and packages.

Example: Standardizing customer income data for regression analysis.

d. Data Visualization in R

R provides visualization libraries such as ggplot2 for graphical analysis.

Example: Creating scatter plots to identify relationships between variables.


MODULE – 3

Statistical Concepts: Probability distributions. Hypothesis testing. Regression analysis


basics.
Building Statistical Models: Simple and multiple linear regression. Model assumptions
and diagnostics.
Introduction to Data Mining: Meaning and Process of Data Mining: CRISP-DM Model.

STATISTICAL CONCEPTS

PROBABILITY DISTRIBUTIONS

Probability distribution yields the possible outcomes for any random event. It is also defined
based on the underlying sample space as a set of possible outcomes of any random experiment.

A probability distribution is a statistical function that describes all the possible values and
likelihoods that a random variable can take within a given range. This range will be
bounded between the minimum and maximum possible values, but precisely where the possible
value is likely to be plotted on the probability distribution depends on a number of factors.
These factors include the distribution's mean (average), standard deviation, skewness,
and kurtosis. In simple terms, it tells us the pattern of probabilities for all possible outcomes.

There are two major classes of probability distributions:

1. Discrete Probability Distribution

A discrete distribution is used when the random variable can take countable values (finite or
countably infinite). Each possible value has a specific probability assigned to it. The sum of all
probabilities equals 1.

Example
Rolling a die:
• Possible values: 1, 2, 3, 4, 5, 6
• Each has probability = 1/6
Other examples: number of customers, number of defects
2. Continuous Probability Distribution

A continuous distribution is used when the variable can take any value within a range.
Probabilities are represented using a probability density function (PDF). The probability of
a single value is zero; instead, we calculate probability over intervals.

Example
• Height of people (e.g., between 150 cm and 180 cm)
• Temperature values

Important Concepts

1. Probability Mass Function (PMF)


Used for discrete variables to assign probabilities to each value.
Example: Probability of getting a 3 on a die = 1/6

2. Probability Density Function (PDF)


Used for continuous variables to describe probability over a range.
Example: Probability that height lies between 160–170 cm

3. Cumulative Distribution Function (CDF)


Shows the cumulative probability up to a certain value.
Example: Probability that a student scores less than 50 marks

Properties of Probability Distribution


• Probabilities lie between 0 and 1
• Total probability = 1
• Describes behavior of a random variable

COMMON PROBABILITY DISTRIBUTIONS

BINOMIAL DISTRIBUTION (Discrete)

The binomial distribution is a discrete distribution with a finite number of possibilities. When
observing a series of what are known as Bernoulli trials, the binomial distribution emerges. A
Bernoulli trial is a scientific experiment with only two outcomes: success or failure.
Consider a random experiment in which you toss a biased coin six times with a 0.4 chance of
getting head. If 'getting a head' is considered a ‘success’, the binomial distribution will show
the probability of r successes for each value of r.

The binomial random variable represents the number of successes (r) in n consecutive
independent Bernoulli trials.

The Bernoulli distribution is a variant of the Binomial distribution in which only one
experiment is conducted, resulting in a single observation. As a result, the Bernoulli distribution
describes events that have exactly two outcomes.

Applications of Binomial Distribution


• Number of Side Effects from Medications
• Number of Fraudulent Transactions
• Number of Spam Emails per Day
• Number of River Overflows
• Shopping Returns per Week

POISSON DISTRIBUTION (Discrete)

A Poisson distribution is a probability distribution used in statistics to show how many times
an event is likely to happen over a given period of time. To put it another way, it's a count
distribution. Poisson distributions are frequently used to comprehend independent events at a
constant rate over a given time interval. Siméon Denis Poisson, a French mathematician, was
the inspiration for the name.

Applications of Poisson Distribution


• Calls per Hour at a Call Center
• Number of Arrivals at a Restaurant
• Number of Website Visitors per Hour
• Number of Network Failures per Week
• Proportion of defects per unit length or per unit area etc.
NORMAL DISTRIBUTION (Continuous)

Normal distribution, also called Gaussian distribution, the most common distribution
function for independent, randomly generated variables. Its familiar bell-shaped curve
is ubiquitous in statistical reports, from survey analysis and quality control to resource
allocation.

The graph of the normal distribution is characterized by two parameters: the mean, or average,
which is the maximum of the graph and about which the graph is always symmetric; and
the standard deviation, which determines the amount of dispersion away from the mean.

Properties of a normal distribution


• The mean, mode and median are all equal.
• The curve is symmetric at the center (i.e. around the mean, μ).
• Exactly half of the values are to the left of center and exactly half the values are to the
right.
• The total area under the curve is 1.

The Standard Normal Model: A standard normal model is a normal distribution with a mean
of 0 and a standard deviation of 1.

Applications of Normal Distribution


• Birthweight of Babies
• Height of Males
• Shoe Sizes
• Blood Pressure
HYPOTHESIS TESTING

MEANING OF HYPOTHESES

A hypothesis (plural hypotheses) is a precise, testable statement of what the researcher(s)


predict will be the outcome of the study. It is stated at the start of the study.

A hypothesis is an assumption that is made based on some evidence. This is the initial point of
any investigation that translates the research questions into predictions. It includes components
like variables, population and the relation between the variables. A research hypothesis is a
hypothesis that is used to test the relationship between two or more variables.

TYPES OF HYPOTHESES

There are six forms of hypothesis and they are:

i. Simple Hypothesis: It shows a relationship between one dependent variable and a


single independent variable. For example – If you eat more vegetables, you will lose
weight faster. Here, eating more vegetables is an independent variable, while losing
weight is the dependent variable.

ii. Complex Hypothesis: It shows the relationship between two or more dependent
variables and two or more independent variables. Eating more vegetables and fruits
leads to weight loss, glowing skin, reduces the risk of many diseases such as heart
disease.

iii. Directional Hypothesis: It shows how a researcher is intellectual and committed to a


particular outcome. The relationship between the variables can also predict its nature.
For example- children aged four years eating proper food over a five-year period are
having higher IQ levels than children not having a proper meal. This shows the effect
and direction of effect.

iv. Non-directional Hypothesis: It is used when there is no theory involved. It is a


statement that a relationship exists between two variables, without predicting the exact
nature (direction) of the relationship.
v. Null Hypothesis: It provides the statement which is contrary to the hypothesis. It’s a
negative statement, and there is no relationship between independent and dependent
variables. The symbol is denoted by “Hₒ”.

vi. Associative and Causal Hypothesis: Associative hypothesis occurs when there is a
change in one variable resulting in a change in the other variable. Whereas, causal
hypothesis proposes a cause and effect interaction between two or more variables.

Examples of Hypothesis

Following are the examples of hypothesis based on their types:

• Null Hypothesis (𝑯𝝄 ): There is no relationship between marketing and sales.


• Alternate Hypothesis (𝑯𝟏 ): There is a relationship between marketing and sales.

• Consumption of sugary drinks every day leads to obesity is an example of a simple


hypothesis.

• All lilies have the same number of petals is an example of a null hypothesis.

• If a person gets 7 hours of sleep, then he will feel less fatigue than if he sleeps less.

CHARACTERISTICS OF HYPOTHESES

Hypothesis must possess the following characteristics:

i. Hypothesis should be clear and precise. If the hypothesis is not clear and precise, the
inferences drawn on its basis cannot be taken as reliable.

ii. Hypothesis should be capable of being tested. In a swamp of untestable hypotheses,


many a time the research programmes have bogged down. Some prior study may be
done by researcher in order to make hypothesis a testable one. A hypothesis “is testable
if other deductions can be made from it which, in turn, can be confirmed or disproved
by observation.”

iii. Hypothesis should state relationship between variables, if it happens to be a


relational hypothesis.
iv. Hypothesis should be limited in scope and must be specific. A researcher must
remember that narrower hypotheses are generally more testable and he should develop
such hypotheses.

v. Hypothesis should be stated as far as possible in most simple terms so that the same
is easily understandable by all concerned. But one must remember that simplicity of
hypothesis has nothing to do with its significance.

vi. Hypothesis should be consistent with most known facts i.e., it must be consistent
with a substantial body of established facts. In other words, it should be one which
judges accept as being the most likely.

vii. Hypothesis should be amenable to testing within a reasonable time. One should not
use even an excellent hypothesis, if the same cannot be tested in reasonable time for
one cannot spend a life-time collecting data to test it.

viii. Hypothesis must explain the facts that gave rise to the need for explanation. This
means that by using the hypothesis plus other known and accepted generalizations, one
should be able to deduce the original problem condition. Thus hypothesis must actually
explain what it claims to explain; it should have empirical reference.

FORMULATION OF HYPOTHESES

Formulating a hypothesis can take place at the very beginning of a research project, or after a
bit of research has already been done. Sometimes a researcher knows right from the start which
variables she is interested in studying, and she may already have a hunch about their
relationships. Other times, a researcher may have an interest in a particular topic, trend, or
phenomenon, but he may not know enough about it to identify variables or formulate a
hypothesis.

Whenever a hypothesis is formulated, the most important thing is to be precise about what one's
variables are, what the nature of the relationship between them might be, and how one can go
about conducting a study of them.
PROCEDURE FOR TESTING HYPOTHESES

The Steps in Hypothesis Testing are:


• State the Null Hypothesis
• State the Alternative Hypothesis
• Select the level of significance- 𝛼
• Collect Data
• Calculate a test statistic
• Construct Acceptance / Rejection regions
• Based on steps 5 and 6, draw a conclusion about 𝐻𝜊

ERRORS IN HYPOTHESES

There are two types of errors in hypothesis testing, Type I and Type II errors. Both types of
error relate to incorrect conclusions about the null hypothesis.

Type I Error: We may reject H0 when H0 is true. Type I error means rejection of hypothesis
which should have been accepted. Type I error is denoted by α (alpha) known as α error, also
called the level of significance of test

Type II Error: We may accept H0 when in fact H0 is not true. Type II error means accepting
the hypothesis which should have been rejected. Type II error is denoted by β (beta) known as
β error.

In a tabular form the said two errors can be presented as follows:

PARAMETRIC AND NON-PARAMETRIC TESTS

Hypothesis testing determines the validity of the assumption (technically described as null
hypothesis) with a view to choose between two conflicting hypotheses about the value of a
population parameter. Hypothesis testing helps to decide on the basis of a sample data, whether
a hypothesis about the population is likely to be true or false.

They are broadly classified into two types:

Parametric tests or standard tests of hypotheses

Parametric tests usually assume certain properties of the parent population from which we draw
samples. Assumptions like observations come from a normal population, sample size is large,
assumptions about the population parameters like mean, variance, etc., must hold good before
parametric tests can be used. parametric tests require measurement equivalent to at least an
interval scale.

The important parametric tests are:


• z-test
• t-test
• χ2 –test
• F-test.
• ANOVA
All these tests are based on the assumption of normality i.e., the source of data is considered to
be normally distributed.

Non-parametric tests or distribution-free test of hypotheses.

There are situations when the researcher cannot or does not want to make such assumptions.
In such situations we use statistical methods for testing hypotheses which are called non-
parametric tests because such tests do not depend on any assumption about the parameters of
the parent population. Non-parametric tests assume only nominal or ordinal data.

The important non-parametric tests are:


• Mann Whitney –U Test.
• Kruskal Wallis (K-W) Test
Difference Between Parametric and Non-Parametric Tests

Parameters of
Parametric Nonparametric
Comparison

Meaning A statistical test, in which A statistical test used in the case of


specific assumptions are made non-metric independent variables,
about the population parameter is called non-parametric test.
is known as parametric test.

Basis of test Distribution Arbitrary


statistic

Type of It is used on data that follows a It is used on data that follows any
distribution normal distribution. arbitrary distribution.

Measurement Interval or ratio Nominal or ordinal


level

Measure of Mean Median


central tendency

Information Completely known Unavailable


about population

Parametric tests have higher Nonparametric tests have lower


Statistical power
statistical power. statistical power.

Correlation test Pearson Spearman

REGRESSION ANALYSIS – BASICS

Meaning

Regression analysis is a statistical technique used to study the relationship between a dependent
variable and one or more independent variables. It is mainly used for prediction and estimation
of continuous values.

It helps in understanding how changes in independent variables influence the dependent


variable. It also quantifies the strength and direction of the relationship. In predictive analytics,
regression is widely used for forecasting and decision-making.

Example

A company wants to predict sales (Y) based on advertising expenditure (X). Regression helps
estimate how much sales will increase when advertising increases.
Key Components

• Dependent Variable (Y): The variable we want to predict or explain. Example: Sales,
profit, demand
• Independent Variable (X): The variable(s) used to predict Y. Example: Price,
advertising, income

Steps in Regression Analysis


1. Identify dependent and independent variables
2. Collect and prepare data
3. Fit the regression model
4. Evaluate model performance
5. Use model for prediction

Uses of Regression
• Sales forecasting
• Demand prediction
• Financial analysis
• Marketing effectiveness

Advantages
• Simple and easy to interpret
• Useful for prediction
• Helps understand relationships

Limitations
• Assumes linear relationship
• Sensitive to outliers
• May not capture complex patterns
BUILDING STATISTICAL MODELS

Building a regression model involves developing a mathematical equation that best fits the
data.

A. Simple Linear Regression

Meaning

Simple linear regression uses one independent variable to predict the dependent variable.

Model Equation

Y = a + bX

Explanation

• a (Intercept): Value of Y when X = 0

• b (Slope): Change in Y for a unit change in X

The model tries to fit a straight line that minimizes the error between actual and predicted
values.

Example

Predicting sales based only on advertising expenditure.

B. Multiple Linear Regression

Meaning

Multiple regression uses two or more independent variables to predict the dependent
variable.
Model Equation

𝑌 = 𝑎 + 𝑏1 𝑋1 + 𝑏2 𝑋2 + ⋯ + 𝑏𝑛 𝑋𝑛

Explanation

Each independent variable contributes to predicting Y. This model is more realistic because
business problems usually depend on multiple factors.

Example

Predicting sales based on price, advertising, and income levels.

MODEL ASSUMPTIONS

Model assumptions are the conditions that must be satisfied for a regression model to produce
valid, reliable, and unbiased results. If these assumptions are violated, the model’s predictions
and inferences may become inaccurate.

1. Linearity

This assumption states that there must be a linear relationship between the dependent variable
and independent variables. It means that a change in X should result in a proportional change
in Y. If the relationship is not linear, the model may not fit the data properly.

Example: Sales increasing steadily with advertising expenditure indicates linearity, whereas a
curved relationship violates this assumption.

2. Independence of Errors

The residuals (errors) should be independent of each other, meaning the error in one
observation should not influence another. This is especially important in time-series data where
values may be correlated over time.

Example: Daily sales errors should not depend on the previous day’s errors; otherwise,
autocorrelation exists.

3. Homoscedasticity (Constant Variance)


This assumption requires that the variance of errors remains constant across all levels of the
independent variables. If the spread of residuals increases or decreases, it leads to
heteroscedasticity, which affects the reliability of estimates.

Example: If prediction errors are small for low sales but large for high sales, the assumption
is violated.

4. Normality of Errors

Residuals should be normally distributed, especially for hypothesis testing and confidence
interval estimation. While regression can still work without perfect normality, significant
deviations may affect statistical conclusions.

Example: A bell-shaped distribution of residuals indicates that the normality assumption is


satisfied.

5. No Multicollinearity

Independent variables should not be highly correlated with each other, as this makes it difficult
to isolate their individual effects on the dependent variable. High multicollinearity can lead to
unstable coefficient estimates.

Example: If price and discount are strongly related, it becomes difficult to determine their
separate impact on sales.

6. No Significant Outliers

The dataset should not contain extreme values that disproportionately influence the regression
results. Outliers can distort the regression line and reduce model accuracy.

Example: A one-time unusually high sales value may skew the results and affect predictions.

MODEL DIAGNOSTICS

Model diagnostics refers to the set of techniques used to evaluate the performance, validity,
and reliability of a regression model. It helps in checking whether the model assumptions are
satisfied and whether the model is suitable for prediction.

1. R² (Coefficient of Determination)
R² measures the proportion of variation in the dependent variable explained by the independent
variables. Its value ranges from 0 to 1, where a higher value indicates a better fit of the model.
However, a very high R² does not always mean the model is perfect, as it may also indicate
overfitting.

Example: R² = 0.80 means 80% of variation in sales is explained by the model.

2. Adjusted R²

Adjusted R² modifies the R² value by considering the number of independent variables in the
model. It penalizes unnecessary variables that do not improve the model significantly. This
makes it more reliable for multiple regression models.

Example: If adding a new variable does not improve the model, Adjusted R² may decrease.

3. Residual Analysis

Residuals are the differences between actual and predicted values. Analyzing residuals helps
check whether assumptions like linearity, independence, and constant variance are satisfied.
Ideally, residuals should be randomly distributed without any clear pattern.

Example: A funnel-shaped pattern in residuals indicates heteroscedasticity.

4. p-value (Statistical Significance)

The p-value helps determine whether an independent variable has a significant effect on the
dependent variable. A p-value less than 0.05 generally indicates that the variable is statistically
significant.

Example: If advertising has p < 0.05, it significantly affects sales.

5. F-test (Overall Model Significance)

The F-test evaluates whether the regression model as a whole is statistically significant. It
checks if at least one independent variable affects the dependent variable.

Example: A significant F-test means the model is useful for prediction.

6. VIF (Variance Inflation Factor)

VIF is used to detect multicollinearity among independent variables. A high VIF value indicates
that variables are highly correlated, which can distort regression results.
Example: VIF > 10 suggests serious multicollinearity.

7. Outlier and Influential Points Detection

Outliers are extreme values that differ significantly from other observations. Influential points
can strongly affect the regression line. Identifying and handling them is important for
improving model accuracy.

Example: A very high sales value due to a one-time event may distort the model.

INTRODUCTION TO DATA MINING

Meaning of Data Mining

Data mining is the process of discovering useful patterns, relationships, and insights from large
volumes of data using statistical, machine learning, and analytical techniques.

In simple terms, It is the process of turning raw data into meaningful information for decision-
making.

Data mining goes beyond simple data analysis by identifying hidden patterns and trends that
are not easily visible. It plays a key role in predictive analytics by helping organizations
understand past behavior and predict future outcomes. According to Provost & Fawcett, data
mining is a core part of data-driven decision-making (DDD).

Examples
• Identifying customer buying patterns in retail
• Detecting fraudulent transactions in banking
• Predicting customer churn in telecom

PROCESS OF DATA MINING: CRISP-DM MODEL

CRISP-DM (Cross-Industry Standard Process for Data Mining) is the most widely used
framework for data mining projects. It provides a structured and iterative approach.
1. Business Understanding

This is the first step where the problem is defined from a business perspective. Objectives,
goals, and success criteria are clearly identified. It ensures that the data mining project is
aligned with organizational needs.

Example: A company wants to reduce customer churn.

2. Data Understanding

In this stage, data is collected and explored to understand its structure, quality, and relevance.
Initial analysis is performed to identify patterns and issues such as missing values.

Example: Examining customer data like age, usage, and complaints.

3. Data Preparation

This step involves cleaning, transforming, and organizing data into a usable format. It includes
handling missing data, normalization, and feature selection. It is often the most time-consuming
stage.

Example: Removing duplicate records and converting categorical data into numerical form.

4. Modeling

In this stage, various data mining techniques such as classification, regression, or clustering are
applied. Multiple models may be tested to find the best-performing one.
Example: Building a model to predict whether a customer will churn.

5. Evaluation

The model is evaluated to check whether it meets business objectives and performs accurately.
It ensures that the model is valid and reliable before implementation.

Example: Checking model accuracy and relevance for decision-making.

6. Deployment

The final model is implemented in a real-world environment. The results are used for decision-
making, and the model may be integrated into business systems.

Example: Using the model to identify customers likely to leave and targeting them with offers.
MODULE – 4

Regression Models: Advanced regression techniques (e.g., polynomial, ridge, lasso


regression). Model evaluation metrics (R², RMSE, MAE).
Classification Models: Logistic regression. Decision trees and random forests. Model
evaluation metrics (accuracy, precision, recall, F1 score).
REGRESSION MODELS

ADVANCED REGRESSION TECHNIQUES

Regression models are statistical and machine learning techniques used to study the
relationship between dependent and independent variables. They help predict continuous
numerical outcomes such as sales, profit, demand, stock prices, customer spending, and
production levels. Advanced regression techniques improve prediction accuracy and handle
complex business problems more effectively than simple linear regression.

Advanced regression techniques are used when traditional regression models are unable to
handle non-linear relationships, multicollinearity, or large numbers of variables efficiently.

POLYNOMIAL REGRESSION

Polynomial regression is an advanced form of linear regression used when the relationship
between variables is non-linear or curved rather than straight.

In many real-world business situations, the relationship between independent and dependent
variables does not follow a straight line. Polynomial regression solves this problem by adding
higher-degree terms such as 𝑋 2 , 𝑋 3 , etc., to the regression equation.

This allows the model to fit curved patterns in the data and improve prediction accuracy.
Polynomial regression is commonly used in sales forecasting, trend analysis, and production
optimization.

Equation

𝑌 = 𝑎 + 𝑏1 𝑋 + 𝑏2 𝑋 2 + 𝑏3 𝑋 3 + ⋯ + 𝑏𝑛 𝑋 𝑛

Where:

• 𝑌= Dependent variable
• 𝑋= Independent variable

• 𝑎= Intercept

• 𝑏= Regression coefficients

Example

A company may observe that increasing advertising expenditure initially increases sales
rapidly, but after a certain point, sales growth slows down. This curved relationship can be
modeled effectively using polynomial regression.

Characteristics

• Polynomial Regression is a form of regression analysis in which the relationship


between the independent variables and dependent variables are modeled in the nth
degree polynomial.

• Polynomial Regression models are usually fit with the method of least squares. The
least square method minimizes the variance of the coefficients, under the Gauss
Markov Theorem.

• Polynomial Regression is a special case of Linear Regression where we fit


the polynomial equation on the data with a curvilinear relationship between the
dependent and independent variables.

Assumptions of Polynomial Regression:

• The behavior of a dependent variable can be explained by a linear, or curvilinear,


additive relationship between the dependent variable and a set of k independent
variables (xi, i=1 to k).
• The relationship between the dependent variable and any independent variable is linear
or curvilinear (specifically polynomial).

• The independent variables are independent of each other.

• The errors are independent, normally distributed with mean zero and a constant
variance (OLS).

Why do we need Polynomial Regression?

Let’s consider a case of Simple Linear Regression.

• We make our model and find out that it performs very badly,

• We observe between the actual value and the best fit line, which we predicted and it
seems that the actual value has some kind of curve in the graph and our line is no where
near to cutting the mean of the points.

• This where polynomial Regression comes to the play, it predicts the best fit line that
follows the pattern(curve) of the data, as shown in the pic below:

• Polynomial Regression does not require the relationship between the independent and
dependent variables to be linear in the data set, This is also one of the main difference
between the Linear and Polynomial Regression.

• Polynomial Regression is generally used when the points in the data are not captured
by the Linear Regression Model and the Linear Regression fails in describing the best
result clearly.
As we increase the degree in the model, it tends to increase the performance of the model.
However, increasing the degrees of the model also increases the risk of over-fitting and under-
fitting the data.

How to find the right degree of the equation?

In order to find the right degree for the model to prevent over-fitting or under-fitting, we can
use:

1. Forward Selection: This method increases the degree until it is significant enough to
define the best possible model.

2. Backward Selection: This method decreases the degree until it is significant enough
to define the best possible model.

Math Behind Polynomial Regression

If you know what Linear Regression is then you will probably understand the maths behind the
polynomial regression too. Linear Regression is basically the first degree Polynomial.

Advantages

• Captures non-linear relationships

• Improves prediction accuracy

• More flexible than simple linear regression


Limitations

• Can lead to overfitting

• Complex interpretation

• Sensitive to extreme values

RIDGE REGRESSION & LASSO REGRESSION

Regression shrinkage, also known as regularization, is a technique used in statistical


modeling and machine learning to prevent overfitting and improve the generalization of a
model. The idea is to add a penalty term to the standard regression objective function,
encouraging the model to be simpler by discouraging overly complex or extreme parameter
values.

There are two common types of regression shrinkage techniques:

1. Lasso Regression (L1 Regularization): Lasso stands for Least Absolute Shrinkage
and Selection Operator. In lasso regression, the penalty term added to the objective
function is the absolute value of the coefficients' sum. This leads to some coefficients
becoming exactly zero, effectively performing variable selection by excluding certain
features from the model.

2. Ridge Regression (L2 Regularization): Ridge regression adds a penalty term to the
objective function that is the squared sum of the coefficients. This penalizes large
coefficients but does not force them to be exactly zero, allowing all features to be
included in the model. Ridge regression is particularly useful when dealing with
multicollinearity (high correlation among predictor variables).

Ridge Regression or (L2 Regularization) Method

Ridge regression, also known as L2 regularization, is a technique used in linear regression to


prevent overfitting by adding a penalty term to the loss function. This penalty is proportional
to the square of the magnitude of the coefficients (weights).

Ridge Regression is a version of linear regression that includes a penalty to prevent the model
from overfitting, especially when there are many predictors or not enough data.
The standard loss function (mean squared error) is modified to include a regularization term:
𝑛
Loss = MSE + 𝜆 ∑𝑖=1 𝑤𝑖2

Here, λ is the regularization parameter that controls the strength of the penalty, and wi are
the coefficients.

Example

In predicting house prices, variables such as:

• House size

• Number of rooms

• Property value

may be strongly correlated. Ridge regression helps stabilize the model and improve forecasting
accuracy.

Advantages

• Reduces overfitting

• Handles multicollinearity effectively

• Improves model stability

Limitations

• Does not eliminate variables

• Model interpretation becomes difficult

Lasso Regression or (L1 Regularization) Method

Lasso regression, also known as L1 regularization, is a linear regression technique that adds
a penalty to the loss function to prevent overfitting. This penalty is based on the absolute
values of the coefficients.

Lasso regression is a version of linear regression including a penalty equal to the absolute value
of the coefficient magnitude. By encouraging sparsity, this L1 regularization term reduces
overfitting and helps some coefficients to be absolutely zero, hence facilitating feature
selection.

The standard loss function (mean squared error) is modified to include a regularization term:

Loss = MSE + 𝜆 ∑𝑛𝑖=1 ∣ 𝑤𝑖 ∣

Here, λ is the regularization parameter that controls the strength of the penalty, and wi are
the coefficients.

Example

A telecom company predicting customer churn may use variables such as:

• Customer age

• Call duration

• Internet usage

• Subscription type

Lasso regression may automatically eliminate variables that contribute very little to churn
prediction.

Advantages

• Performs feature selection automatically

• Reduces model complexity

• Improves interpretability

Limitations

• May remove useful variables

• Less stable with highly correlated variables

Difference between Ridge Regression and Lasso Regression

The key differences between ridge and lasso regression are discussed below:
Characteristic Ridge Regression Lasso Regression

Applies L2 regularization, Applies L1 regularization, adding a


Regularization adding a penalty term penalty term proportional to
Type proportional to the square of the the absolute value of the
coefficients coefficients.

Does not perform feature


Performs automatic feature
selection. All predictors are
Feature selection. Less important predictors
retained, although their
Selection are completely excluded by setting
coefficients are reduced in size
their coefficients to zero.
to minimize overfitting

Best suited for situations


Ideal when you suspect that only
where all predictors are
a subset of predictors is important,
When to use potentially relevant, and the
and the model should focus on those
goal is to reduce overfitting
while ignoring the irrelevant ones.
rather than eliminate features

Produces a model that Produces a model that is simpler,


includes all features, but their retaining only the most significant
Output model
coefficients are smaller in features and ignoring the rest by
magnitude to prevent overfitting setting their coefficients to zero.

Reduces the magnitude of


Shrinks some coefficients
coefficients, shrinking them
to exactly zero, effectively
Impact on towards zero, but does not set
removing their influence from the
Prediction any coefficients exactly to zero.
model. This leads to a simpler
All predictors remain in the
model with fewer features
model
Characteristic Ridge Regression Lasso Regression

Generally faster as it doesn’t May be slower due to the feature


Computation
involve feature selection selection process

Use when you have many Use when you believe only some
predictors, all contributing to predictors are truly important (e.g.,
Example Use
the outcome (e.g., predicting genetic studies where only a few
Case
house prices where all features genes out of thousands are
like size, location, etc., matter) relevant).

MODEL EVALUATION METRICS (R², RMSE, MAE)

Model evaluation metrics help measure the performance and accuracy of regression models.
These metrics compare predicted values with actual values to determine how well the model
performs.

The objective of Linear Regression is to find a line that minimizes the prediction error of all
the data points.
The essential step in any machine learning model is to evaluate the accuracy of the model.
The Mean Squared Error, Mean absolute error, Root Mean Squared Error, and R-Squared or
Coefficient of determination metrics are used to evaluate the performance of the model in
regression analysis.

1. R² (Coefficient of Determination)

R² measures the proportion of variation in the dependent variable explained by the regression
model.

The value of R² ranges from 0 to 1.

R² = 0 → Model explains no variation

R² = 1 → Perfect prediction

A higher R² value indicates a better model fit because more variation in the dependent variable
is explained by the independent variables.

However, an extremely high R² may sometimes indicate overfitting.

Example

If:

R² = 0.85
then 85% of variation in sales is explained by variables such as advertising expenditure and
pricing.

2. RMSE (Root Mean Square Error)

RMSE measures the average magnitude of prediction errors in a regression model.

RMSE squares prediction errors before averaging them. Because errors are squared, larger
errors receive greater importance.

A lower RMSE indicates better prediction accuracy.

RMSE is useful when large prediction errors are highly undesirable.

Formula

∑(𝑦𝑖 − 𝑦̂𝑖 )2
𝑅𝑀𝑆𝐸 = √
𝑛

Where:

• 𝑦𝑖 = Actual value

• 𝑦̂𝑖 = Predicted value

• 𝑛= Number of observations

Example

If predicted sales are very different from actual sales, RMSE increases significantly.

Advantages

• Penalizes large errors strongly

• Useful for comparing models

Limitations

• Sensitive to outliers

• Harder to interpret than MAE


3. MAE (Mean Absolute Error)

MAE measures the average absolute difference between predicted and actual values.

Unlike RMSE, MAE does not square errors. All prediction errors are treated equally.

MAE is easier to understand because the error is expressed in the same units as the original
data.

A smaller MAE indicates better model performance.

Formula

∑ ∣ 𝑦𝑖 − 𝑦̂𝑖 ∣
𝑀𝐴𝐸 =
𝑛

Where:
• 𝑦𝑖 = Actual value
• 𝑦̂𝑖 = Predicted value
• 𝑛= Number of observations
Example

If the average prediction error in monthly sales forecasting is ₹500, then:

• MAE = 500

Advantages

• Easy to interpret

• Less sensitive to outliers than RMSE

Limitations

• Does not penalize large errors strongly

Comparison of Evaluation Metrics

Metric Purpose Best Value


R² Measures goodness of fit Closer to 1
RMSE Measures prediction error magnitude Lower value
MAE Measures average absolute error Lower value
CLASSIFICATION MODELS
Classification models are a type of predictive modeling that organizes data into predefined
classes according to feature values.
Classification models are a type of machine learning model that divides data points into
predefined groups called classes. Classifiers are a type of predictive modeling that learns class
characteristics from input data and learns to assign possible classes to new data according to
those learned characteristics. Classification algorithms are widely used in data science for
forecasting patterns and predicting outcomes. Indeed, they have an array of real-world use
cases, such as patient classification per potential health risks and spam email filtering.
Classification tasks can be binary or multiclass. In binary classification problems, a model
predicts between two classes. For example, a spam filter classifies emails as spam or not spam.
Multiclass classification problems classify data among more than two class labels. For instance,
an image classifier might classify images of pets by using a myriad of class labels, such
as dog, cat, llama, platypus and more.
Of course, each machine learning classification algorithm differs in its internal operations. All
nevertheless adhere to a general two-step data classification process:
1. Learning. In supervised learning, a human annotator assigns each data point in the
training dataset a label. These points are defined as a number of input variables (or
independent variables), which might be numerical, text strings, image features and so
forth. In mathematical terms, the model considers each data point as a tuple x. A tuple
is merely an ordered numerical sequence represented as x = (x1,x2,x3…xn). Each value
in the tuple is a given feature of the data point. The model uses each data point’s features
along with its class label to decode what features define each class. By mapping training
data per this equation, a model learns those general features (or variables) associated
with each class label.
2. Classification. The second step in classification tasks is classification itself. In this
phase, users deploy the model on a test set of unseen data. Previously unused data is
ideal for evaluating model classification in order to avoid overfitting. The model uses
its learned predicted function y=f(x) to classify the unseen data across distinct classes
according to each sample’s features. Users then evaluate model accuracy according to
the number of correctly predicted test data samples.
Predictions
Classification models output two types of predictions: discrete and continuous.
1. Discrete. Discrete predictions are the predicted class labels for each data point. For
example, we can use a predictor to classify medical patients
as diabetic or nondiabetic based on health data. The
classes diabetic and nondiabetic are the discrete categorical predictions.
2. Continuous. Classifiers assign class predictions as continuous probabilities called
confidence scores. These probabilities are values between 0 and 1, representing
percentages. Our model might classify a patient as diabetic with a .82 probability. This
means that the model believes the patient has an 82% chance of being diabetic and a
18% chance of being nondiabetic.
Researchers typically evaluate models by using discrete predictions while using continuous
predictions as thresholds. A classifier ignores any prediction under a certain threshold. For
instance, if our diabetes predictor has a threshold of .4 (40%) and classifies a patient as
diabetic with a probability of .35 (35%), then the model will ignore that label and not assign
the patient to the diabetic class.

Classification models are machine learning and statistical techniques used to predict categorical
outcomes or classes. Unlike regression models, which predict continuous numerical values,
classification models predict categories such as:
• Yes / No
• Fraud / Not Fraud
• Pass / Fail
• Disease / No Disease
• Customer Churn / No Churn
Classification models are widely used in predictive analytics for fraud detection, customer
segmentation, medical diagnosis, spam filtering, and recommendation systems.

LOGISTIC REGRESSION
Logistic regression is a supervised classification algorithm used to predict the probability of a
categorical outcome, especially binary outcomes.
Although the name contains “regression,” logistic regression is mainly used for classification
problems. It predicts probabilities between 0 and 1 using a mathematical function called the
sigmoid function.
The output is usually:
• 1 → Positive class
• 0 → Negative class
The model estimates the likelihood that a particular event will occur.
Logistic regression is suitable when:
• The dependent variable is categorical
• The problem involves binary classification

Sigmoid Function
The sigmoid curve converts predicted values into probabilities.
1
𝑃(𝑌 = 1) =
1 + 𝑒 −𝑧
Where:
• 𝑃(𝑌 = 1)= Probability of positive outcome
• 𝑒= Exponential constant
• 𝑧= Linear combination of variables

Linear Combination in Logistic Regression


The value of 𝑧 is calculated as:
𝑧 = 𝑎 + 𝑏1 𝑋1 + 𝑏2 𝑋2 + ⋯ + 𝑏𝑛 𝑋𝑛
Where:
• 𝑎= Intercept
• 𝑏= Regression coefficients
• 𝑋= Independent variables

Working of Logistic Regression


The working of logistic regression involves the following steps:
1. Collect training data
2. Identify independent and dependent variables
3. Apply the sigmoid function
4. Calculate probabilities
5. Classify observations into categories
6. Evaluate model performance
The model learns patterns from historical data and predicts the probability of future outcomes.

Types of Logistic Regression


Logistic regression is mainly classified into three types.
1. Binary Logistic Regression: Used when the dependent variable has only two categories.
Examples
• Pass / Fail
• Fraud / Not Fraud
• Yes / No

2. Multinomial Logistic Regression: Used when the dependent variable has more than two
categories without any order.
Examples
• Product preference: Brand A, Brand B, Brand C
• Transportation choice: Bus, Train, Car

3. Ordinal Logistic Regression: Used when categories have a meaningful order.


Examples
• Customer satisfaction levels:
o Poor
o Average
o Good
o Excellent
Assumptions of Logistic Regression
Logistic regression works effectively when certain assumptions are satisfied.
1. Binary or Categorical Dependent Variable: The output variable should be
categorical.
2. Independent Observations: Observations should not depend on each other.
3. No Multicollinearity: Independent variables should not be highly correlated.
4. Large Sample Size: Logistic regression performs better with larger datasets.
5. Linear Relationship with Log Odds: Independent variables should have a linear
relationship with log odds.

Applications of Logistic Regression


Logistic regression is widely used across industries.
Banking and Finance
Banks use logistic regression for:
• Credit risk analysis
• Loan default prediction
• Fraud detection
Example
Predicting whether a customer will repay a loan.

Healthcare
Hospitals use logistic regression for:
• Disease diagnosis
• Patient risk prediction
Example
Predicting whether a patient has diabetes.

Marketing
Businesses use logistic regression for:
• Customer churn prediction
• Purchase behavior analysis
Example
Predicting whether a customer will buy a product.
Human Resource Management
Organizations use logistic regression for:
• Employee attrition prediction
• Recruitment analysis

Advantages
• Simple and easy to interpret
• Produces probability outputs
• Works well for binary classification
• Computationally efficient
Limitations
• Assumes linear relationship in log odds
• Less effective for highly complex datasets
• Sensitive to outliers

DECISION TREES
Decision trees are one of the most widely used supervised machine learning algorithms for
classification and prediction problems. A decision tree is a tree-structured model that divides
data into branches based on conditions or decision rules. It helps organizations make
predictions and decisions in a simple, visual, and interpretable manner.
Decision trees are commonly used in:
• Loan approval prediction
• Customer segmentation
• Fraud detection
• Medical diagnosis
• Employee attrition analysis
• Marketing analytics
The model predicts outcomes by asking a sequence of questions and splitting data into smaller
subsets until a final decision is reached.

A decision tree is a predictive model that represents decisions and their possible outcomes in
the form of a tree-like structure.
The model starts with a root node and repeatedly splits data into branches based on conditions.
Each branch represents a possible outcome, and the final nodes represent predictions or
classifications.

Structure of a Decision Tree


A decision tree consists of several important components.
1. Root Node: The root node is the topmost node in the tree and represents the first
decision or condition used to split the dataset. Example: “Income > ₹50,000?”
2. Internal Nodes: Internal nodes represent additional decision conditions used to further
divide the data. Example: “Credit Score > 700?”
3. Branches: Branches connect nodes and represent possible outcomes of decisions.
Example: Yes, No
4. Leaf Nodes (Terminal Nodes): Leaf nodes represent the final prediction or
classification outcome. Example: Loan Approved, Loan Rejected

Working of Decision Trees


Decision trees work by repeatedly splitting the dataset into smaller subsets based on the best
decision conditions.
The process involves:
1. Selecting the best feature for splitting
2. Dividing the dataset into branches
3. Repeating the process for each branch
4. Stopping when classification becomes sufficiently accurate
The model identifies the conditions that best separate different classes.

Example of Decision Tree


A bank wants to predict whether a loan should be approved.
Independent Variables
• Income
• Credit score
• Employment status
• Existing loans
Dependent Variable
• Approved
• Rejected
The decision tree evaluates conditions step by step until a final decision is reached.
Splitting Criteria in Decision Trees
The model selects the best variable for splitting using mathematical measures.
1. Gini Index: Measures impurity or randomness in the dataset. Lower Gini value
indicates better splitting.
2. Entropy: Measures uncertainty or disorder in the dataset. Lower entropy indicates
more homogeneous groups.
3. Information Gain: Measures reduction in entropy after splitting. Higher information
gain indicates a better split.

Types of Decision Trees


Decision trees are mainly classified into two types.
1. Classification Trees: Used when the target variable is categorical.
Examples
• Fraud / Not Fraud
• Pass / Fail
• Disease / No Disease

2. Regression Trees: Used when the target variable is continuous.


Examples
• Sales prediction
• Profit forecasting
• Demand estimation

Advantages of Decision Trees


• Easy to Understand and Interpret: Decision trees are highly visual and simple to
explain, even for non-technical users.
• Handles Numerical and Categorical Data: The model can process both types of
variables effectively.
• Requires Less Data Preparation: Decision trees do not require normalization or
scaling of data.
• Useful for Decision-Making: The rule-based structure helps managers understand
business decisions clearly.
• Handles Missing Values: Some decision tree algorithms can manage missing data
effectively.

Limitations of Decision Trees


• Overfitting: Decision trees may become too complex and memorize training data
instead of learning patterns.
• Unstable Model: Small changes in data may produce different tree structures.
• Lower Accuracy: Single decision trees may be less accurate than ensemble models
such as random forests.
• Bias Toward Dominant Features: Features with many categories may dominate the
splitting process.

Pruning in Decision Trees

Pruning in tree structures, particularly in the context of decision trees, is a technique used to
prevent overfitting and improve the generalization ability of the model. Overfitting occurs
when a model learns the training data too well, capturing noise and irrelevant patterns that do
not generalize well to unseen data.
Pruning involves selectively removing parts of the tree that do not contribute significantly to
its predictive accuracy or performance on unseen data. There are mainly two types of pruning
techniques:
1. Pre-Pruning: This involves setting constraints on the growth of the tree during the
construction phase. Instead of growing the tree to its maximum depth or until all leaves
are pure (contain only one class), pre-pruning techniques stop the tree growth based on
conditions such as the maximum depth of the tree, minimum number of samples
required to split a node, or minimum improvement in impurity measure (e.g., Gini
impurity, entropy) required for a split.
2. Post-Pruning (or Cost-Complexity Pruning): Post-pruning, also known as cost-
complexity pruning, involves growing the tree to its maximum size and then iteratively
removing nodes or branches that have the least impact on the tree's performance. This
is typically done by calculating a pruning parameter (often based on impurity measures
or error rates) for each subtree and removing the subtree with the smallest increase in
error rate or impurity when pruned. The process continues until further pruning does
not improve the model's performance or until a predefined stopping criterion is met.

Applications of Decision Trees


• Banking and Finance: Used for, Loan approval prediction, Credit risk analysis, Fraud
detection
• Healthcare: Used for, Disease diagnosis, Patient risk prediction
• Marketing: Used for, Customer segmentation, Purchase prediction, Churn analysis
• Human Resource Management: Used for, Employee attrition prediction, Recruitment
decision-making

RANDOM FORESTS
Leo Breiman developed an extension of decision trees called random forests. There is publicly
available software for this method.
Random forest is an ensemble learning method that combines multiple decision trees to
improve prediction accuracy.
Instead of depending on one decision tree, random forest creates several trees using different
subsets of data and variables. Each tree gives a prediction, and the final prediction is based on
majority voting.
For classification tasks, the output of the random forest is the class selected by most trees. For
regression tasks, the mean or average prediction of the individual trees is returned (is typically
the average prediction of all individual trees in the ensemble.). Random decision forests correct
for decision trees' habit of overfitting to their training set and improves model stability.

Random forests are highly effective for:


• Classification
• Prediction
• Feature importance analysis

Example
An e-commerce company predicts whether a customer will purchase a product based on:
• Browsing history
• Purchase behavior
• Product preferences
Multiple decision trees make predictions, and the final output is determined through majority
voting.

Advantages
• High prediction accuracy
• Reduces overfitting
• Handles large datasets efficiently
• Works well with missing values

Limitations
• Computationally expensive
• Difficult to interpret
• Requires more memory and processing power
MODEL EVALUATION METRICS (ACCURACY, PRECISION, RECALL, F1
SCORE)
Model evaluation metrics are used to measure how well a classification model performs. After
building a classification model such as logistic regression, decision tree, or random forest, it is
important to evaluate whether the model is making correct predictions.

These metrics compare the predicted values with the actual values and help determine the
accuracy and reliability of the model.

The most commonly used evaluation metrics are:


• Accuracy
• Precision
• Recall
• F1 Score
These metrics are calculated using the Confusion Matrix.

Confusion Matrix

A confusion matrix is a table used to evaluate the performance of a classification model.

Predicted Class
Positive Negative
Actual Positive TP FN
Actual Negative FP TN

Where:
• TP (True Positive) → Model correctly predicts positive class
• TN (True Negative) → Model correctly predicts negative class
• FP (False Positive) → Model incorrectly predicts positive class
• FN (False Negative) → Model incorrectly predicts negative class
1. Accuracy

Accuracy measures the overall correctness of the classification model. It shows the proportion
of total predictions that were classified correctly.

Formula

𝑇𝑃 + 𝑇𝑁
𝐴𝑐𝑐𝑢𝑟𝑎𝑐𝑦 =
𝑇𝑃 + 𝑇𝑁 + 𝐹𝑃 + 𝐹𝑁

Accuracy calculates how many predictions made by the model are correct out of all predictions.
A higher accuracy value indicates better model performance. However, accuracy may become
misleading when the dataset is imbalanced. For example, if 95% of observations belong to one
class, the model may achieve high accuracy simply by predicting the majority class.

Example
Suppose a model predicts:
90 correct predictions out of 100 total predictions
Then:
Accuracy = 90%

Advantages
• Easy to understand
• Useful for balanced datasets
• Measures overall model performance

Limitations
• Misleading for imbalanced datasets
• Does not distinguish between types of errors

2. Precision

Precision measures how many predicted positive observations are actually positive.

Formula

𝑇𝑃
𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛 =
𝑇𝑃 + 𝐹𝑃
Precision focuses on the quality of positive predictions.
A high precision value means:
• Few false positives
• More reliable positive predictions
Precision becomes important when false positive errors are costly.

Example
In email spam detection:
• High precision means emails classified as spam are truly spam.
In fraud detection:
• High precision ensures that flagged transactions are actually fraudulent.

Advantages
• Reduces false alarms
• Important in fraud detection and spam filtering

Limitations
• Does not consider false negatives

3. Recall

Recall measures how many actual positive cases are correctly identified by the model.
Formula
𝑇𝑃
𝑅𝑒𝑐𝑎𝑙𝑙 =
𝑇𝑃 + 𝐹𝑁
Recall focuses on detecting all positive cases.
A high recall value means:
• Fewer false negatives
• Most positive cases are identified successfully
Recall is important when missing positive cases is risky.

Example
In disease diagnosis:
• High recall ensures that most patients with the disease are detected.
In fraud detection:
• High recall helps identify most fraudulent transactions.
Advantages
• Reduces missed positive cases
• Useful in healthcare and security applications

Limitations
• High recall may increase false positives

4. F1 Score

F1 Score is the harmonic mean of precision and recall.


Formula
𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛 × 𝑅𝑒𝑐𝑎𝑙𝑙
𝐹1 = 2 ×
𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛 + 𝑅𝑒𝑐𝑎𝑙𝑙
F1 Score balances both:
• Precision
• Recall
It is useful when:
• The dataset is imbalanced
• Both false positives and false negatives are important
A higher F1 Score indicates better classification performance.

Example
In medical diagnosis:
• F1 Score helps balance:
o Correct disease detection
o Avoiding unnecessary alarms

Advantages
• Balances precision and recall
• Effective for imbalanced datasets

Limitations
• Slightly difficult to interpret
• Does not consider true negatives directly
Comparison of Evaluation Metrics

Metric Focus Best Value Important When

Accuracy Overall correctness Higher Balanced datasets

Precision Correct positive predictions Higher False positives are costly

Recall Detecting actual positives Higher Missing positives is risky

F1 Score Balance between precision and recall Higher Imbalanced datasets


MODULE – 5

Time Series Analysis: Components of time series data. Meaning of Stationarity &
differencing, ARIMA models.
Basics of Big Data: Meaning, Characteristic’s (5V’s) and Sources of Big Data. Handling
Large Data Sets - Techniques and Challenges. Introduction to Hadoop and Spark.

TIME SERIES ANALYSIS

Time Series Analysis is a way of studying the characteristics of the response variable
concerning time as the independent variable. To estimate the target variable in predicting or
forecasting, use the time variable as the reference point. TSA represents a series of time-based
orders, it would be Years, Months, Weeks, Days, Horus, Minutes, and Seconds. It is an
observation from the sequence of discrete time of successive intervals. Since TSA involves
producing the set of information in a particular sequence, this makes it distinct from spatial and
other analyses.

Time series analysis is a statistical technique used to analyze data collected over a period of
time at regular intervals. It helps identify patterns, trends, and future movements in data. Time
series analysis is widely used in predictive analytics for forecasting sales, stock prices, weather
conditions, demand, and economic indicators.

We could predict the future using AR (Autoregressive), MA (Moving Average), ARMA


(Autoregressive Moving Average), and ARIMA (Autoregressive Integrated Moving Average)
models.

Time series data refers to observations recorded sequentially over time, such as daily, monthly,
quarterly, or yearly data.

COMPONENTS OF TIME SERIES DATA

• Trend: The long-term, sustained direction of the data. It represents continuous, gradual
increases or decreases over an extended period (e.g., an overall increase in global
population or a gradual decline in traditional cable subscriptions).
• Seasonality: Predictable, repeating patterns tied to specific, fixed timeframes. These
fluctuations occur at regular intervals (daily, weekly, monthly, or quarterly) and are
usually calendar-related (e.g., a spike in retail sales every December or increased ice
cream purchases every summer).
• Cyclical Variations: Long-term fluctuations that follow an up-and-down pattern,
typically spanning multiple years. Unlike seasonality, cycles are not fixed in duration
and are usually linked to broader macroeconomic factors or business cycles (e.g.,
economic recessions and expansions).
• Irregular Variations (Noise): Random, unpredictable short-term fluctuations that do
not fit a discernible pattern. These are unexplainable shifts caused by one-time events
(e.g., a sudden drop in tourism due to an unexpected weather event or a spike in a
specific stock price due to a sudden news announcement).

MEANING OF STATIONARITY & DIFFERENCING

Stationarity

A time series has stationarity if a shift in time doesn’t cause a change in the shape of the
distribution. Basic properties of the distribution like the mean, variance and covariance are
constant over time.
Types of Stationary

Models can show different types of stationarity:

• Strict stationarity means that the joint distribution of any moments of any degree
(e.g. expected values, variances, third order and higher moments) within the process
is never dependent on time. This definition is in practice too strict to be used for any
real-life model.

• First-order stationarity series have means that never changes with time. Any other
statistics (like variance) can change.

• Second-order stationarity (also called weak stationarity) time series have a constant
mean, variance and an autocovariance that doesn’t change with time. Other statistics in
the system are free to change over time. This constrained version of strict stationarity
is very common.

• Trend-stationary models fluctuate around a deterministic trend (the series mean).


These deterministic trends can be linear or quadratic, but the amplitude (height of one
oscillation) of the fluctuations neither increases nor decreases across the series.

• Difference-stationary models are models that need one or more differencing to


become stationary.

It can be difficult to tell if a model is stationary or not. Unlike the obvious example showing
seasonality above, you usually can’t tell by looking at a graph. If you aren’t sure about the
stationarity of a model, a hypothesis test can help.

You have several options for testing, including:

• Unit root tests (e.g. Augmented Dickey-Fuller (ADF) test or Zivot-Andrews test),

• A KPSS test (run as a complement to the unit root tests).

• A run sequence plot,

• The Priestley-Subba Rao (PSR) Test or Wavelet-Based Test, which are less common
tests based on spectrum analysis.
Importance of Stationarity

Most forecasting methods assume that a distribution has stationarity. For example,
autocovariance and autocorrelations rely on the assumption of stationarity. An absence of
stationarity can cause unexpected or bizarre behaviors, like t-ratios not following a t-
distribution or high r-squared values assigned to variables that aren’t correlated at all.

Transforming Models

Most real-life data sets just aren’t stationary. To quote Thomson (1994):

“Experience with real-world data, however, soon convinces one that both stationarity and
Gaussianity are fairy tales invented for the amusement of undergraduates.”

To put it another way, if you’ve got a real-life data set (and not a theoretical one from a class),
you’re going to need to make it stationary in order to get any useful predictions from it. A
model can sometimes be stationarized through mathematical transformation (usually
performed by software) which makes the model relatively easy to predict; It will have the same
statistical properties at a later date. The mathematical transformations are then reversed so that
the new model predicts the behavior of the original time series model. Transformations might
include:

• Difference the data: differenced data has one less point than the original data. For
example, given a series Zt you can create a new series Yi = Zi – Zi – 1.

• Fit a curve to the data, then model the residuals from that curve.

• Take the logarithm or square root (usually works for data with non-constant variance).

Some models can’t be transformed in this way — like models with seasonality. These can
sometimes be broken down into smaller pieces (a process called stratification) and individually
transformed. Another way to deal with seasonality is to subtract the mean value of the periodic
function from the data.

Differencing

Differencing is where your data has one less data point than the original data set; You’re
subtracting (or moving) a point—a “difference”. For example, given a series Zt you can create
a new series Yi = Zi – Zi – 1. As well as its general use in transformations, differencing is
widely used in time series analysis.
“A series with no deterministic component which has a stationary,
invertible ARMA representation after differencing d times is said to be integrated of order d”
In simple terms, Differencing is a technique used to make non-stationary data stationary.
It removes trends and seasonality by calculating the difference between consecutive
observations.
The first difference is calculated as:
𝑌𝑡′ = 𝑌𝑡 − 𝑌𝑡−1

Where:
• 𝑌𝑡 = current value
• 𝑌𝑡−1 = previous value

If the series is still non-stationary, second-order differencing may be applied.


Formula:


𝑌𝑡′′ = 𝑌𝑡′ − 𝑌𝑡−1

or directly,

𝑌𝑡′′ = 𝑌𝑡 − 2𝑌𝑡−1 + 𝑌𝑡−2

ARIMA MODELS (AutoRegressive Integrated Moving Average)

ARIMA is one of the most widely used statistical models for time series forecasting. It helps
predict future values based on past observations. ARIMA models are especially useful when
data shows patterns such as trends over time.

ARIMA is commonly used in:


• Sales forecasting
• Stock market prediction
• Demand forecasting
• Weather prediction
• Economic forecasting
Component Full Form Purpose
AR AutoRegressive Uses past values
I Integrated Makes data stationary using differencing
MA Moving Average Uses past forecast errors

ARIMA models are represented as:


𝐴𝑅𝐼𝑀𝐴(𝑝, 𝑑, 𝑞)
Where:
• p = Number of autoregressive terms
• d = Number of differencing operations
• q = Number of moving average terms

Many time series datasets are non-stationary, meaning:


• Mean changes over time
• Variance changes over time
• Trends exist
ARIMA converts such data into stationary form and then performs forecasting.

Components of ARIMA

1. AutoRegressive (AR) Component

The autoregressive part means that the current value depends on previous values of the same
series.
If previous observations influence current observations, the series has autoregressive behavior.
For example:
• Today's sales may depend on yesterday’s sales.
• Today's stock price may depend on past stock prices.
Mathematical Representation:
For AR(1):
𝑌𝑡 = 𝑐 + 𝜙𝑌𝑡−1 + 𝑒𝑡
Where:
• 𝑌𝑡 = Current value
• 𝑌𝑡−1 = Previous value
• 𝑝ℎ𝑖= AR coefficient
• 𝑒𝑡 = Error term

Example
If a retail store had high sales yesterday, it is likely to have relatively high sales today as well.

2. Integrated (I) Component


Integrated refers to the differencing process used to make the data stationary.
Many time series datasets contain trends and are non-stationary.
Non-stationary data can produce unreliable forecasts.
Differencing removes:
• Trends
• Seasonality
• Changing mean
First-Order Differencing
𝑌𝑡′ = 𝑌𝑡 − 𝑌𝑡−1
Example
Month Sales
January 100
February 120
Difference:
120 − 100 = 20
Order of Differencing (d)

Value of d Meaning
d=0 Already stationary
d=1 First differencing applied
d=2 Second differencing applied
Usually, d = 1 is sufficient.

3. Moving Average (MA) Component


The moving average component uses past forecast errors to improve predictions.
The current value depends not only on past observations but also on previous prediction errors.
This helps correct forecasting mistakes.
Mathematical Representation:
For MA(1):
𝑌𝑡 = 𝜇 + 𝑒𝑡 + 𝜃𝑒𝑡−1
Where:
• 𝑒𝑡 = Current error
• 𝑒𝑡−1= Previous error
• 𝑡ℎ𝑒𝑡𝑎= MA coefficient
Example
If sales predictions were too low last month, the model adjusts future predictions accordingly.

Steps in Building an ARIMA Model


1. Data Collection: Historical time series data is collected. Example, Monthly sales data
for 5 years.
2. Check Stationarity: The data is analyzed to determine whether it is stationary.
Methods: Visual inspection, ADF test (Augmented Dickey-Fuller Test)
3. Apply Differencing: If the data is non-stationary, differencing is applied.
4. Identify p and q Values: Autocorrelation techniques are used:

Method Purpose
ACF Helps identify MA(q)
PACF Helps identify AR(p)
5. Build ARIMA Model: The appropriate ARIMA(p,d,q) model is selected.
6. Evaluate the Model: Metrics used are RMSE, MAE, AIC, BIC
7. Forecast Future Values: The model predicts future observations.

Advantages of ARIMA

• Good for short-term forecasting


• Handles trends effectively
• Widely accepted statistical model
• Useful for stationary time series data
• Accurate forecasting when patterns exist

Limitations of ARIMA

• Requires stationary data


• Difficult parameter selection
• Not suitable for highly irregular data
• Cannot handle multiple seasonal patterns effectively
• Requires large historical datasets

Key Difference Summary

• AR → Focuses on previous observations


• MA → Focuses on previous errors
• ARMA → Combines AR and MA for stationary data
• ARIMA → Extends ARMA by handling non-stationary data through differencing
Difference Between AR, MA, ARMA & ARIMA

Feature AR Model MA Model ARMA Model ARIMA Model


AutoRegressive
Moving AutoRegressive
AutoRegressive Integrated
Full Form Average Moving Average
Model Moving Average
Model Model
Model
Uses past
Uses past values errors to Combines past Combines AR +
Basic Idea to predict predict values and past MA +
current values current errors differencing
values
Requires
Stationarity Requires Requires Can handle non-
stationary
Requirement stationary data stationary data stationary data
data
Main Moving
Autoregression AR + MA AR + I + MA
Components average
Model
AR(p) MA(q) ARMA(p,q) ARIMA(p,d,q)
Representation
Uses Previous
Yes No Yes Yes
Values?
Uses Previous
No Yes Yes Yes
Errors?
Uses
No No No Yes
Differencing?
Captures Improves
Corrects past Forecasts non-
dependence on forecasting using
Purpose forecasting stationary time
past both past values
errors series data
observations and errors
Stationary
Stationary series Stationary data
series with Non-stationary
Suitable For with with both AR
error data with trends
autocorrelation and MA patterns
dependence
Forecast
Stock prices
Today’s sales adjusts based Forecasting
influenced by
Example depend on on previous monthly sales
past prices and
yesterday’s sales prediction with trends
errors
errors
Complexity Simple Simple Moderate More complex
Business
Error
Common Stock prices, Economic forecasting,
correction
Applications sales forecasting demand
forecasting
prediction
BASICS OF BIG DATA

MEANING

Big Data refers to extremely large and complex datasets that cannot be efficiently processed,
stored, or analyzed using traditional data processing tools and database systems. Big Data
includes structured, semi-structured, and unstructured data generated from various digital
sources such as social media, sensors, websites, and business transactions.

Big Data analytics helps organizations extract meaningful insights, identify patterns, improve
decision-making, and support predictive analytics.

CHARACTERISTIC’S (5V’S)

Big data is defined by its five foundational characteristics. Together, the "5 Vs" explain how
organizations capture, manage, and transform massive, complex, and fast-moving information
into actionable business insights.

The 5 Vs encompass the following characteristics:

• Volume: Refers to the sheer scale and massive amount of data generated every second.
It includes terabytes and petabytes of information that cannot be processed by
traditional databases.

• Velocity: The rapid speed at which data is created, gathered, and processed. Real-time
data processing is critical for handling continuous, fast-flowing information like
financial transactions or social media streams.
• Variety: Represents the different types and formats of data. It includes traditional
structured data (e.g., SQL tables), semi-structured data (e.g., JSON logs), and
unstructured data (e.g., audio, video, and emails).

• Veracity: Measures the accuracy, reliability, and trustworthiness of the data. Dealing
with veracity involves cleaning out noise, biases, and duplicates to ensure the data is of
high quality.

• Value: The ultimate goal of big data. It refers to an organization's ability to extract
meaningful, profitable, and strategic insights from the vast amounts of raw data.

V Meaning Example

Volume Large amount of data Social media data

Velocity Speed of data generation Real-time transactions

Variety Different forms of data Text, images, videos

Veracity Data quality and reliability Accurate customer data

Value Business usefulness Better decision-making

SOURCES OF BIG DATA


1. Social Media Platforms

Social media platforms generate massive amounts of data continuously. Platforms like
Facebook, Instagram, Twitter, LinkedIn, and TikTok record posts, likes, shares, comments,
videos, and user profiles.

Example: Over 500 million tweets are posted daily. Businesses analyze this to understand
trends, customer behavior, and preferences. Social media is one of the most prominent sources
of big data today.

2. Internet of Things (IoT) Devices

Connected devices like smartwatches, fitness trackers, smart TVs, home assistants, and
connected cars generate constant streams of data. Sensors track locations, activities, and device
performance.

Example: A smart fridge monitors food inventory and usage patterns. With billions of
connected devices worldwide, IoT is among the world’s biggest sources of big data due to its
real-time updates.

3. Healthcare Systems

Hospitals, clinics, and wearable devices generate terabytes of medical data every day. This
includes patient records, diagnostic images, and device readings.

Example: MRI scans and ECG readings are recorded for millions of patients. Healthcare relies
on these sources of big data for disease prediction, personalized treatment, and research.

4. Financial Transactions

Every online purchase, card swipe, or stock trade produces high-frequency data. Banks and
financial institutions track this information for security and analysis.

Example: Visa, Mastercard, and Paytm record millions of transactions daily. Financial data is
one of the world’s biggest sources of big data because of continuous, high-speed activity.

5. E-Commerce Platforms

Online marketplaces track clicks, searches, reviews, and purchases. This helps businesses offer
personalized recommendations.
Example: Amazon monitors which products users view and their browsing time. E-commerce
platforms remain key sources of big data for understanding customer behavior.

6. Telecommunication Networks

Telecom providers gather call records, SMS logs, and internet usage data.

Example: Providers track network usage to identify dropped calls and optimize coverage.
Telecom data is an important source of big data for infrastructure planning.

7. Government and Public Records

Governments generate huge datasets, including census information, tax records, and vehicle
registrations.

Example: India’s Aadhaar system records biometric and demographic details for over a billion
people. Such records are critical sources of big data for policy-making and planning.

8. Education Systems

Student performance, online course engagement, and learning platform activity generate data
continuously.

Example: MOOCs record user progress and participation. Education data is a key source of
big data for improving teaching and learning experiences.

9. Retail Stores and Point-of-Sale Data

Retail stores collect information from barcode scans, loyalty programs, and customer footfall.

Example: Walmart processes over 1 million transactions every hour. Retail data is an important
source of big data for inventory management and sales prediction.

10. Search Engines

Search engines record billions of queries daily, reflecting user interests and trends.

Example: Google processes over 3.5 billion searches every day. Search data is one of the
world’s biggest sources of big data due to its volume, speed, and global coverage.

11. Transportation and Logistics

Data from GPS tracking, ride-hailing apps, and airline bookings provide insights for route
optimization.
Example: Uber collects ride and location data from millions of users daily. Transport data is a
vital source of big data for operational efficiency.

12. Media and Entertainment

Streaming platforms record viewing history, ratings, downloads, and preferences.

Example: Netflix uses user data to recommend shows. Media and entertainment contribute as
significant sources of big data.

13. Weather and Climate Monitoring

Satellites, temperature sensors, and environmental instruments generate continuous climate


data.

Example: NASA and ISRO track global weather patterns. These measurements are key
sources of big data for forecasting and disaster planning.

14. Manufacturing and Industry

Machines on production lines produce data on performance, faults, and efficiency.

Example: Automotive factories track assembly line operations to reduce errors. Industrial data
is a crucial source of big data for predictive maintenance.

15. Emails and Messaging Platforms

Emails and messaging apps create enormous volumes of text, attachments, and usage records.

Example: Over 300 billion emails are sent daily worldwide. Communication data is a
prominent source of big data for analysis and automation.

TECHNIQUES FOR HANDLING LARGE DATA SETS

Handling large datasets is one of the major challenges in Big Data analytics because traditional
systems are often unable to process massive volumes of structured and unstructured data
efficiently. To manage and analyze Big Data effectively, organizations use several advanced
techniques and technologies.

1. Distributed Computing
Distributed computing is a technique in which large datasets are divided and processed across
multiple computers or servers instead of relying on a single machine. Each computer in the
network performs a portion of the processing task, and the results are combined later. This
technique improves processing speed, scalability, and reliability. Distributed computing is
widely used in Big Data frameworks such as Hadoop. For example, e-commerce companies
process millions of customer transactions simultaneously using distributed systems.

2. Parallel Processing

Parallel processing refers to the execution of multiple computational tasks at the same time.
Instead of processing data sequentially, large datasets are divided into smaller parts and
processed simultaneously by multiple processors. This significantly reduces computation time
and improves efficiency. Parallel processing is commonly used in scientific research, banking
systems, and real-time analytics applications where fast processing is essential.

3. Cloud Computing

Cloud computing provides scalable storage and processing capabilities through internet-based
services. Organizations can store and process large datasets using cloud platforms such as
AWS, Microsoft Azure, and Google Cloud without investing heavily in physical infrastructure.
Cloud computing offers flexibility, scalability, cost efficiency, and remote accessibility.
Businesses use cloud computing to handle increasing data volumes and perform Big Data
analytics efficiently.

4. Data Compression

Data compression is a technique used to reduce the storage space required for large datasets.
Compression algorithms minimize file size while preserving important information. This helps
organizations save storage costs and improve data transfer speed. Data compression is
especially useful for multimedia files such as videos, images, and audio data generated by
streaming platforms and social media applications.

5. NoSQL Databases

NoSQL databases are designed to handle unstructured and semi-structured data efficiently.
Unlike traditional relational databases, NoSQL databases can store large volumes of diverse
data formats such as text, images, videos, and sensor data. Databases such as MongoDB,
Cassandra, and HBase are commonly used in Big Data environments because they provide
high scalability and flexibility.

6. Data Partitioning

Data partitioning involves dividing a large dataset into smaller, manageable segments called
partitions. Each partition can be processed independently, improving system performance and
reducing processing time. Partitioning also enhances scalability and simplifies data
management. Large organizations often partition customer or transaction data based on region,
time, or category.

7. In-Memory Processing

In-memory processing stores data in the system’s main memory (RAM) instead of reading it
repeatedly from disk storage. This significantly increases processing speed and supports real-
time analytics. Technologies such as Apache Spark use in-memory processing to perform faster
computations compared to traditional disk-based systems.

8. Data Sampling

Data sampling is a technique where a smaller representative subset of a large dataset is selected
for analysis. Instead of processing the entire dataset, analysts use samples to identify patterns
and trends more quickly. Sampling reduces computation time and is useful in predictive
analytics and statistical analysis when processing the full dataset is not feasible.

9. Batch Processing

Batch processing involves processing large volumes of data in groups or batches at scheduled
intervals instead of processing data continuously in real time. This technique is useful for
handling repetitive and large-scale processing tasks such as payroll systems, banking
transactions, and report generation.

10. Stream Processing

Stream processing is used for analyzing continuously generated real-time data. Instead of
storing data first and analyzing later, stream processing analyzes data instantly as it is
generated. This technique is widely used in stock market analysis, fraud detection, social media
monitoring, and IoT applications where immediate insights are required.
CHALLENGES IN HANDLING LARGE DATA SETS

Handling large datasets is a major challenge in Big Data analytics because the size, speed, and
complexity of data exceed the capabilities of traditional data processing systems. Organizations
face several technical, operational, and security-related difficulties while storing, processing,
and analyzing Big Data effectively.

1. Data Storage Challenges

One of the biggest challenges is storing massive amounts of data generated every second from
multiple sources such as social media, IoT devices, banking systems, and e-commerce
platforms. Traditional storage systems are often insufficient to handle such huge volumes of
data. Organizations require scalable and distributed storage infrastructure, which can be
expensive and complex to manage.

2. Data Processing Speed

Big Data is generated at very high speed, and organizations often require real-time or near real-
time analysis for decision-making. Processing such large datasets quickly is challenging
because it demands high computational power and advanced processing frameworks. Delays
in processing may reduce the usefulness of insights, especially in applications like fraud
detection and stock market analysis.

3. Data Quality Issues

Large datasets often contain incomplete, inconsistent, duplicate, or inaccurate information.


Poor data quality can produce misleading results and affect predictive analytics outcomes.
Cleaning and preparing Big Data is difficult because the datasets are extremely large and may
contain both structured and unstructured data from different sources.

4. Data Security and Privacy

Big Data frequently contains sensitive information such as financial records, customer details,
healthcare information, and business transactions. Protecting this data from cyberattacks,
unauthorized access, and data breaches is a major challenge. Organizations must implement
strong security measures, encryption techniques, and privacy policies to ensure data protection
and regulatory compliance.
5. Scalability Challenges

As organizations continue generating more data, systems must be capable of scaling efficiently
to handle increasing workloads. Expanding storage capacity and processing infrastructure
without affecting performance is a difficult task. Poor scalability may lead to slower processing
and system failures.

6. Integration of Diverse Data

Big Data comes from multiple sources and exists in different formats such as text, images,
videos, audio files, sensor data, and databases. Integrating these diverse types of data into a
unified system for analysis is a complex process. Managing structured, semi-structured, and
unstructured data together requires specialized technologies and tools.

7. High Infrastructure Cost

Managing Big Data requires advanced hardware, distributed computing systems, cloud
platforms, and skilled professionals. Setting up and maintaining such infrastructure can be
expensive, especially for small and medium-sized organizations. Continuous upgrades and
maintenance also increase operational costs.

8. Real-Time Data Management

Many businesses require instant processing and analysis of continuously generated streaming
data. Handling real-time data efficiently is challenging because it requires low-latency systems,
fast networks, and powerful processing frameworks. Applications such as online
recommendations, traffic monitoring, and financial trading depend heavily on real-time
analytics.

9. Lack of Skilled Professionals

Big Data technologies such as Hadoop, Spark, machine learning, and cloud computing require
specialized technical knowledge. Many organizations face difficulty finding skilled data
scientists, analysts, and Big Data engineers who can effectively manage and analyze large
datasets.

10. Data Governance and Compliance

Organizations must follow legal and regulatory requirements related to data usage, privacy, and
storage. Ensuring compliance with regulations while handling huge datasets is challenging.
Improper data governance can result in legal penalties, reputational damage, and loss of
customer trust.

INTRODUCTION TO HADOOP
Hadoop is an open-source framework developed for storing and processing extremely large
datasets across multiple computers in a distributed environment. It was designed to handle Big
Data efficiently when traditional database systems became insufficient for managing massive
volumes of structured and unstructured data.
Hadoop allows organizations to divide large datasets into smaller parts and process them
simultaneously using clusters of computers. This distributed approach improves processing
speed, scalability, and fault tolerance. Hadoop is widely used in industries such as banking,
healthcare, retail, telecommunications, and social media analytics.
One of the major advantages of Hadoop is its ability to process huge amounts of data at low
cost using commodity hardware instead of expensive high-end systems. Hadoop also supports
scalability, meaning additional systems can be added easily as data volume increases.

Main Components of Hadoop


Hadoop consists of four major components that work together to store and process Big Data
efficiently.
1. Hadoop Distributed File System (HDFS)
HDFS is the storage component of Hadoop. It stores large datasets across multiple computers
in a distributed manner. Data is divided into smaller blocks and replicated across different
machines to ensure reliability and fault tolerance. Even if one system fails, the data remains
available from other systems.
2. MapReduce
MapReduce is the data processing framework in Hadoop. It processes large datasets in parallel
across multiple systems.
• The Map phase divides data into smaller tasks.
• The Reduce phase combines the processed results.
This parallel processing capability helps Hadoop handle massive datasets efficiently.
3. YARN (Yet Another Resource Negotiator)
YARN is the resource management component of Hadoop. It manages system resources such
as memory and processing power and allocates them to different applications running on the
Hadoop cluster.
4. Hadoop Common
Hadoop Common contains shared libraries and utilities required for Hadoop operations. It
provides the basic services that support the functioning of other Hadoop components.

Working of Hadoop
The working of Hadoop involves storing data in HDFS and processing it using MapReduce
across multiple machines.
Large Data

Stored in HDFS

Processed using MapReduce

Results Generated

For example, a bank analyzing millions of customer transactions for fraud detection can
distribute the processing tasks across several computers using Hadoop.

Features of Hadoop
Hadoop provides several important features that make it suitable for Big Data analytics:
• Scalability – New systems can be added easily to handle increasing data volumes.
• Fault Tolerance – Data replication ensures reliability even if systems fail.
• Distributed Processing – Data is processed across multiple machines simultaneously.
• Cost Efficiency – Uses low-cost commodity hardware.
• Flexibility – Can process structured, semi-structured, and unstructured data.

Advantages of Hadoop
Hadoop enables organizations to process huge datasets efficiently and economically. It supports
parallel processing, high scalability, and distributed storage. Hadoop is also highly reliable
because data is replicated across multiple systems. It is widely used for data mining, predictive
analytics, recommendation systems, and customer behavior analysis.

Limitations of Hadoop
Although Hadoop is powerful, it has some limitations. Hadoop MapReduce processing can be
slower because it relies heavily on disk storage. It is also complex to configure and manage.
Real-time processing is difficult in Hadoop, making it less suitable for applications requiring
instant analytics.

INTRODUCTION TO SPARK
Apache Spark is an open-source Big Data processing framework designed for high-speed and
real-time data analytics. Spark was developed to overcome some limitations of Hadoop
MapReduce, especially processing speed.
Unlike Hadoop, Spark performs in-memory processing, meaning data is processed directly in
RAM instead of repeatedly reading from disk storage. This makes Spark significantly faster
than Hadoop MapReduce for many analytics tasks.
Spark is widely used in:
• Real-time analytics
• Machine learning
• Fraud detection
• Recommendation systems
• Streaming data analysis

Components of Spark
Spark consists of several important modules that support advanced analytics and data
processing.
1. Spark Core
Spark Core is the main processing engine responsible for task scheduling, memory
management, and distributed processing.
2. Spark SQL
Spark SQL is used for processing structured and semi-structured data using SQL queries.
3. MLlib
MLlib is Spark’s machine learning library used for predictive analytics, classification,
clustering, and regression tasks.
4. Spark Streaming
Spark Streaming processes real-time streaming data such as stock market transactions, sensor
data, and social media feeds.
5. GraphX
GraphX is used for graph processing and network analysis.

Features of Spark
Spark provides several advanced features that make it highly popular in Big Data analytics.
• High-Speed Processing – Faster than Hadoop MapReduce because of in-memory
computation.
• Real-Time Analytics – Supports streaming and real-time data processing.
• Machine Learning Support – Includes built-in ML libraries.
• Ease of Use – Supports programming languages such as Python, Java, Scala, and R.
• Scalability – Handles large-scale distributed data processing efficiently.

Advantages of Spark
Spark offers much faster processing compared to Hadoop MapReduce and is highly suitable
for real-time analytics applications. It supports advanced machine learning and streaming
analytics, making it ideal for predictive analytics and AI applications. Spark is also easier to
use because it supports multiple programming languages.

Limitations of Spark
Spark requires large memory resources because it performs in-memory processing. Managing
very large datasets entirely in memory can be expensive. Spark can also be more complex when
handling extremely large distributed systems.
Difference Between Hadoop and Spark

Feature Hadoop Spark


Processing Method Disk-based In-memory
Processing Speed Slower Faster
Real-Time Analytics Limited Excellent
Machine Learning Support Limited Strong
Ease of Use More complex Easier
Best Use Batch processing Real-time analytics
MODULE – 6

Predictive Analytics Tools: Basic features and comparison between Python, R, SAS, SPSS,
Tableau, Power BI, Google Big Query, Azure ML Studio.
Ethical issues in predictive analytics: Data privacy and security. Bias & Fairness in
predictive models

PREDICTIVE ANALYTICS TOOLS

Predictive analytics tools are software platforms and programming environments used for data
analysis, machine learning, forecasting, statistical modeling, visualization, and business
intelligence. These tools help organizations identify patterns from historical data and make
accurate future predictions for better decision-making.

Different tools are used based on:


• Data size
• Complexity of analysis
• Need for visualization
• Machine learning requirements
• Cloud integration
• Ease of use

Python

Python is a high-level, open-source programming language widely used in predictive analytics,


machine learning, artificial intelligence, and Big Data analytics. It has become one of the most
popular tools because of its simple syntax, flexibility, and large collection of analytical
libraries.

Python is preferred by data scientists and businesses because it can handle:


• Data cleaning
• Statistical analysis
• Predictive modeling
• Deep learning
• Automation
• Big Data processing
It supports integration with platforms such as Hadoop, Spark, TensorFlow, and cloud systems.
Major Features of Python

1. Simple and Easy Syntax: Python uses simple English-like commands, making it easier
to learn and write programs compared to many programming languages.
2. Large Collection of Libraries: Python provides powerful libraries for analytics and
machine learning, including:
• Pandas → Data manipulation
• NumPy → Numerical computing
• Matplotlib → Data visualization
• Scikit-learn → Machine learning
• TensorFlow → Deep learning
3. Machine Learning and AI Support: Python is highly suitable for advanced analytics,
artificial intelligence, and neural networks.
4. Big Data Integration: Python integrates easily with Hadoop, Spark & Cloud
platforms. This makes it useful for handling massive datasets.
5. Cross-Platform Compatibility: Python works on Windows, Linux & macOS
6. Automation Support: Python can automate repetitive analytical tasks such as Data
extraction, Report generation & Forecasting

Applications
• Fraud detection
• Customer segmentation
• Sales forecasting
• Recommendation systems
• Time series analysis

Advantages
• Open-source and free
• Flexible and scalable
• Strong community support
• Suitable for advanced predictive analytics

Limitations
• Requires programming knowledge
• Slightly slower execution speed compared to compiled languages
R

R is an open-source programming language and software environment specially developed for


statistical computing and graphical analysis. It is highly popular among statisticians,
researchers, and academic institutions.

R is widely used for:


• Statistical modeling
• Data visualization
• Hypothesis testing
• Predictive analytics
• Data mining
It provides highly advanced statistical techniques and visualization capabilities.

Major Features of R

1. Excellent Statistical Analysis: R contains numerous built-in statistical methods such


as: Regression analysis, ANOVA, Hypothesis testing, Probability distributions
2. Powerful Visualization Tools: R provides advanced graphical capabilities using
packages such as: ggplot2, lattice
These tools help create professional charts and dashboards.
3. Open-Source Platform: R is completely free and supported by a large global
community.
4. Extensive Package Support: Thousands of packages are available for:
• Machine learning
• Time series analysis
• Econometrics
• Data mining
5. Academic and Research Focus: R is heavily used in:
• Universities
• Research institutions
• Statistical studies
Applications
• Academic research
• Statistical forecasting
• Financial analytics
• Healthcare research

Advantages
• Excellent statistical capabilities
• Strong data visualization support
• Free and flexible

Limitations
• Difficult for beginners
• Slower with extremely large datasets

SAS

SAS (Statistical Analysis System) is a commercial software suite used for advanced analytics,
predictive modeling, business intelligence, and data management. It is widely used by large
organizations because of its reliability, scalability, and security.

SAS is especially popular in: Banking, Insurance, Healthcare, Government organizations.

Major Features of SA

1. Advanced Statistical Analysis: SAS provides sophisticated statistical procedures for:


• Regression
• Forecasting
• Predictive modeling
• Data mining
2. Enterprise-Level Security: SAS offers strong security and data governance features
suitable for sensitive business data.
3. Data Management Capabilities: SAS can Clean data, Transform data & Integrate
multiple data sources
4. Business Intelligence Tools: Provides reporting and dashboard features for decision-
making.
5. Scalability: Can efficiently handle large enterprise datasets.

Applications
• Risk management
• Fraud analytics
• Healthcare analytics
• Customer behavior analysis

Advantages
• Highly reliable
• Strong technical support
• Excellent for enterprise analytics

Limitations
• Expensive licensing cost
• Requires specialized training

SPSS

SPSS (Statistical Package for the Social Sciences) is a statistical software developed mainly
for data analysis, survey analysis, and research applications. It is widely used in academic,
marketing, and social science research.

SPSS is known for its user-friendly graphical interface and minimal coding requirements.

Major Features of SPSS

1. Easy-to-Use Interface: Users can perform analysis using menus and dialog boxes
instead of programming.
2. Statistical Analysis Tools: Supports Regression analysis, Correlation analysis,
Hypothesis testing & ANOVA
3. Data Management: Allows Data cleaning, Data transformation & Missing value
handling
4. Report Generation: SPSS generates tables, charts, and statistical reports
automatically.
5. Survey Data Analysis: Highly useful for questionnaire and survey analysis.

Applications
• Market research
• Academic research
• Employee surveys
• Consumer behavior studies

Advantages
• Beginner-friendly
• Minimal programming needed
• Good for statistical analysis

Limitations
• Limited advanced AI capabilities
• Expensive commercial software

Tableau

Tableau is a powerful business intelligence and data visualization tool used to create interactive
dashboards, charts, and reports.

It helps organizations convert raw data into understandable visual insights for better decision-
making.

Major Features of Tableau

1. Interactive Dashboards: Users can create dynamic dashboards with filters and drill-
down analysis.
2. Drag-and-Drop Interface: No advanced programming knowledge is required.
3. Real-Time Analytics: Supports real-time data updates and monitoring.
4. Multiple Data Source Integration: Connects with Excel, SQL databases, Cloud
platforms & Hadoop
5. Advanced Visualization: Creates Heat maps, Geographic maps, Trend charts & KPI
dashboards

Applications
• Sales analysis
• Performance tracking
• Business reporting
• Data storytelling

Advantages
• Excellent visualization quality
• Easy to use
• Fast dashboard development

Limitations
• Limited advanced predictive modeling
• Expensive enterprise version

Power BI

Power BI is a business analytics and visualization tool developed by Microsoft. It is widely


used for creating reports, dashboards, and business intelligence solutions.

Power BI integrates strongly with Microsoft products such as Excel, Azure, and SQL Server.

Major Features of Power BI

1. Interactive Dashboards: Provides highly interactive business reports and


visualizations.
2. Cloud Integration: Supports cloud-based data sharing and collaboration.
3. Microsoft Ecosystem Support: Integrates easily with Excel, Azure, Teams & SQL
Server
4. Real-Time Monitoring: Supports live dashboards and real-time analytics.
5. AI-Based Insights: Provides automated insights and trend detection.

Applications
• Financial reporting
• Sales analytics
• KPI monitoring
• Business intelligence

Advantages
• User-friendly
• Affordable
• Excellent Microsoft integration

Limitations
• Less flexible than Python for advanced analytics
• Performance limitations with very large datasets

Google BigQuery

Google BigQuery is a cloud-based Big Data analytics platform developed by Google. It is


designed for analyzing massive datasets using high-speed SQL queries.

BigQuery is serverless, meaning organizations do not need to manage hardware infrastructure.

Major Features of Google BigQuery

1. Massive Data Processing: Can process terabytes and petabytes of data quickly.
2. Serverless Architecture: No infrastructure management required.
3. High-Speed SQL Queries: Supports extremely fast analytical queries.
4. Cloud Scalability: Automatically scales based on workload.
5. Integration with Google Cloud: Works with: Google Cloud Storage, AI tools &
Machine learning services

Applications
• Big Data analytics
• Real-time business intelligence
• Customer analytics
• Predictive modeling

Advantages
• Very fast processing
• Scalable cloud infrastructure
• Easy Big Data management

Limitations
• Cloud dependency
• Cost may increase with large query usage

Azure ML Studio

Azure ML Studio is a cloud-based machine learning platform developed by Microsoft. It helps


users build, train, test, and deploy machine learning models using a visual interface.

It is widely used for predictive analytics and AI applications.

Major Features of Azure ML Studio

1. Drag-and-Drop Machine Learning: Users can build models visually without


extensive coding.
2. Automated Machine Learning: Supports automated model selection and
optimization.
3. Cloud Deployment: Models can be deployed directly to cloud applications.
4. Integration with Azure Services: Works with: Azure Storage, Power BI & SQL
databases
5. Collaboration Support: Teams can work together on predictive modeling projects.

Applications
• Predictive analytics
• AI model deployment
• Classification and forecasting
• Business intelligence

Advantages
• Easy model deployment
• Scalable cloud platform
• Supports automated machine learning

Limitations
• Requires Azure knowledge
• Internet dependency for cloud access

Tool Best Known For


Python Machine learning and AI
R Statistical analysis
SAS Enterprise-level analytics
SPSS Survey and research analysis
Tableau Interactive dashboards
Power BI Business intelligence reporting
Google BigQuery Big Data cloud analytics
Azure ML Studio Cloud-based machine learning
Comparison of Predictive Analytics Tools

Google Azure ML
Feature Python R SAS SPSS Tableau Power BI
BigQuery Studio

Cloud
Statistical Commercial Data Business
Programming Statistical Cloud Data Machine
Type of Tool Programming Analytics Visualization Intelligence
Language Software Warehouse Learning
Language Software Tool Tool
Platform

Python
SAS Tableau
Developed By Software R Foundation IBM Microsoft Google Microsoft
Institute Software
Foundation

Machine
Statistical Visualization Machine
Primary Learning & Statistical Enterprise Business Big Data
& Survey & Learning &
Purpose Data Analysis Analytics Intelligence Analytics
Analysis Dashboards AI
Analytics

Ease of Moderate to
Moderate Moderate Easy Easy Easy Moderate Moderate
Learning Difficult

Programming SQL
Yes Yes Limited Very Little No Minimal Minimal
Required Knowledge

Open Source Yes Yes No No No Partially No No

Pay-as- Subscription
Cost Free Free Expensive Expensive Expensive Affordable
you-use Based

Statistical
Very Good Excellent Excellent Excellent Limited Moderate Moderate Good
Analysis

Machine
Basic AI
Learning Excellent Very Good Good Limited Limited Good Excellent
Features
Support

Data
Good Excellent Good Moderate Excellent Excellent Moderate Moderate
Visualization

Big Data
Excellent Moderate Very Good Limited Moderate Moderate Excellent Excellent
Handling

Cloud
Good Moderate Good Limited Good Excellent Excellent Excellent
Integration

Real-Time
Good Limited Moderate Limited Good Good Excellent Excellent
Analytics
User Coding- Coding- GUI + Drag-and- Drag-and- Web Drag-and-
GUI-Based
Interface Based Based Coding Drop Drop Interface Drop

Scalability High Moderate High Moderate Moderate High Very High Very High

Large-
Advanced
Statistical Enterprise Academic Dashboards Business Scale Predictive
Best For Analytics &
Research Analytics Research & Reports Intelligence Cloud Modeling
AI
Analytics

Industries IT, Finance, Research, Banking, Academics, Business Corporate Big Data AI & Cloud
Using It Healthcare Education Insurance Surveys Analytics Reporting Companies Analytics

Ease of Massive
Flexibility & Statistical Enterprise Interactive Microsoft Cloud ML
Key Strength Statistical Data
AI Support Modeling Reliability Visualization Integration Deployment
Analysis Processing

Limited Limited
Main Requires Difficult for Limited AI Cloud
High Cost Predictive Advanced Query Cost
Limitation Coding Skills Beginners Features Dependency
Modeling Analytics

ETHICAL ISSUES IN PREDICTIVE ANALYTICS

Predictive analytics helps organizations forecast future outcomes using historical data,
statistical models, and machine learning techniques. Although predictive analytics provides
many benefits in decision-making, it also creates several ethical concerns related to privacy,
security, fairness, transparency, and responsible use of data.

Ethical issues arise when predictive models misuse personal information, produce biased
decisions, or negatively affect individuals and society. Therefore, organizations must ensure
that predictive analytics systems are accurate, transparent, secure, and fair.

The two major ethical concerns in predictive analytics are:

1. Data Privacy and Security

2. Bias and Fairness in Predictive Models


DATA PRIVACY AND SECURITY

Data privacy refers to the protection of personal and sensitive information from unauthorized
access, misuse, or disclosure. Data security refers to the methods and technologies used to
protect data from theft, cyberattacks, and data breaches.

Predictive analytics systems often use large amounts of customer, employee, medical, financial,
and behavioral data. If this data is not handled properly, it may violate privacy rights and create
ethical and legal problems.

Importance of Data Privacy in Predictive Analytics

Organizations collect data such as:


• Customer names
• Phone numbers
• Banking information
• Medical records
• Browsing history
• Purchase behavior
This information is highly sensitive. Unauthorized use or exposure of such data may lead to:
• Identity theft
• Financial fraud
• Loss of trust
• Legal penalties
Therefore, organizations must ensure that personal data is collected, stored, and used
responsibly.

Major Privacy Issues in Predictive Analytics

1. Unauthorized Data Collection: Sometimes organizations collect user data without


proper consent or awareness.
Example: Mobile applications tracking users’ locations without clearly informing
them.
2. Data Misuse: Collected data may be used for purposes other than the original intention.
Example: Customer purchase data being sold to third-party advertisers without
permission.
3. Lack of User Consent: Users may not fully understand how their data is being used in
predictive models.
Example: Social media platforms analyzing user behavior for targeted advertising.
4. Data Breaches: Hackers may steal confidential data from organizational databases.
Example: A banking database containing customer account information being hacked.
5. Excessive Data Collection: Organizations may collect unnecessary personal
information that is not relevant to the analysis.
Example: Collecting biometric information for simple online registrations.

Data Security in Predictive Analytics

Data security involves protecting information systems and databases from:


• Unauthorized access
• Cyberattacks
• Data leaks
• Malware attacks
Predictive analytics systems require strong security mechanisms because they handle massive
amounts of sensitive data.

Security Measures Used in Predictive Analytics

• Encryption: Encryption converts data into coded form so unauthorized users cannot
read it.
• Access Control: Only authorized users should be allowed to access sensitive datasets.
• Authentication Systems: Passwords, OTPs, and biometric verification help secure
systems.
• Firewalls and Antivirus Software: Protect systems from external cyber threats.
• Data Backup: Regular backups help recover data during system failures or attacks.

Ethical Guidelines for Data Privacy and Security

Organizations should:
• Obtain user consent before collecting data
• Use data only for intended purposes
• Protect sensitive information using security measures
• Follow data protection laws and regulations
• Minimize unnecessary data collection

Example of Privacy Issue

A healthcare company using patient medical records for predictive analytics without patient
permission may violate ethical and legal standards.

BIAS AND FAIRNESS IN PREDICTIVE MODELS

Bias in predictive analytics occurs when a predictive model produces unfair or discriminatory
results toward certain individuals or groups.
Fairness means ensuring that predictive models make objective and unbiased decisions for all
users.
Bias can occur because predictive models learn patterns from historical data. If historical data
contains discrimination or imbalance, the model may continue those unfair patterns.

Causes of Bias in Predictive Models

1. Biased Training Data: If historical data contains unfair patterns, the model learns and
repeats them.
Example: A hiring dataset dominated by male employees may cause the model to
prefer male candidates.
2. Incomplete Data: Missing or unrepresentative data can create inaccurate predictions
for certain groups.
Example: A medical prediction model trained mainly on adults may not perform well
for children.
3. Human Bias: Bias may enter the model through human decisions during:
• Data collection
• Feature selection
• Model design
4. Algorithmic Bias: Some algorithms may unintentionally favor one category over
another.
Types of Bias in Predictive Analytics

1. Gender Bias: When predictive models unfairly favor one gender.


Example: Recruitment systems preferring male applicants
2. Racial Bias: When models discriminate against certain racial or ethnic groups.
Example: Loan approval systems rejecting applications from certain communities
more frequently.
3. Age Bias: When predictions unfairly disadvantage specific age groups.
Example: Insurance models charging higher premiums to elderly customers unfairly.

Impact of Bias in Predictive Models


Biased predictive models may lead to:
• Unfair hiring decisions
• Discriminatory loan approvals
• Incorrect medical diagnosis
• Social inequality
• Loss of customer trust
Bias can damage both individuals and organizational reputation.

Ensuring Fairness in Predictive Analytics

Organizations can reduce bias and improve fairness using several methods:

1. Using Diverse and Balanced Data: Training data should represent all groups fairly.
2. Regular Model Testing: Models should be tested continuously for discrimination and
unfair outcomes.
3. Transparent Algorithms: Organizations should explain how predictions are made.
4. Human Oversight: Human experts should review important predictive decisions.
5. Ethical AI Practices: Organizations should adopt responsible AI policies and fairness
standards.

Example of Bias Issue

An AI recruitment system trained using past hiring data may reject female applicants because
historical hiring patterns favored male candidates.
Importance of Ethics in Predictive Analytics
Ethical predictive analytics helps organizations:
• Build customer trust
• Improve transparency
• Ensure fairness
• Protect sensitive information
• Comply with legal regulations
Ethical practices also improve the reliability and social acceptance of predictive analytics
systems.

You might also like