Data Analytics AI ML
Data Analytics AI ML
1. Explain the concept of Data Analytics and discuss its importance in data-driven
decision-making. (CO1)
Data Analytics is the systematic process of collecting, cleaning, transforming, and analyzing raw
data to extract meaningful insights, identify patterns, trends, and relationships, and support
informed decision-making. It combines statistical techniques, data processing methods, and
analytical tools—often enhanced with Artificial Intelligence (AI) and Machine Learning (ML)—to
convert data into actionable knowledge.
In simple terms, data analytics answers questions such as what happened, why it happened, what is
likely to happen next, and what actions should be taken.
Conclusion:
Data analytics plays a crucial role in modern decision-making by transforming raw data into
1/71
meaningful insights, enabling organizations to make smarter, faster, and more reliable decisions in a
data-driven world.
2. Explain the different types of Data Analytics with suitable real-world examples.
(CO1)
Data Analytics is broadly classified into four main types based on the type of insights they provide:
1. Descriptive Analytics
Purpose: Answers “What happened?”
It focuses on summarizing historical data to understand past performance.
Techniques Used:
Reports, dashboards, summary statistics, data visualization.
Real-World Example:
An e-commerce company analyzes monthly sales reports to understand total revenue, best-
selling products, and customer purchase trends.
2. Diagnostic Analytics
Purpose: Answers “Why did it happen?”
It identifies the root cause of outcomes by examining data relationships.
Techniques Used:
Data drilling, correlation analysis, root cause analysis.
Real-World Example:
A company investigates why sales dropped in a particular quarter by analyzing factors such as
pricing changes, customer feedback, or marketing campaign performance.
3. Predictive Analytics
Purpose: Answers “What is likely to happen?”
It uses historical data, statistical models, and ML algorithms to predict future outcomes.
2/71
Techniques Used:
Regression analysis, time series analysis, machine learning models.
Real-World Example:
Banks predict whether a customer is likely to default on a loan using past credit history and
financial behavior.
Weather forecasting systems predicting rainfall or temperature.
4. Prescriptive Analytics
Purpose: Answers “What should be done?”
It recommends optimal actions by evaluating multiple scenarios and outcomes.
Techniques Used:
Optimization algorithms, simulations, AI-driven decision models.
Real-World Example:
A ride-sharing app suggests surge pricing and driver allocation to maximize profit and reduce
customer wait time.
Supply chain systems recommending optimal inventory levels.
Conclusion:
Each type of data analytics builds on the previous one—from understanding past events to
predicting future outcomes and recommending actions—making them essential for intelligent, data-
driven decision-making.
3/71
1. Understanding Market Trends and Customer Behavior
Data analytics helps businesses analyze customer preferences, buying patterns, and feedback.
This enables companies to design products, services, and marketing strategies aligned with
customer needs.
2. Strategic Planning and Forecasting
By analyzing historical data, organizations can predict future demand, revenue trends, and
market growth. Predictive analytics supports long-term planning and goal setting.
3. Improving Operational Efficiency
Analytics identifies inefficiencies, process delays, and cost leakages. Businesses can optimize
workflows, reduce operational costs, and improve productivity.
4. Competitive Advantage
Data-driven insights allow businesses to respond quickly to market changes, identify
opportunities before competitors, and create differentiated offerings.
5. Risk Assessment and Management
Data analytics helps in identifying potential risks such as financial losses, customer churn, or
supply chain disruptions, allowing businesses to take preventive actions.
6. Performance Measurement
Key Performance Indicators (KPIs) and dashboards help management track progress against
strategic objectives and make timely adjustments.
Conclusion:
Data analytics empowers organizations to formulate smarter, agile, and customer-centric business
strategies by transforming raw data into actionable insights, ensuring sustainable growth and
competitive success.
4. Identify and explain various data sources used in data analytics. (CO1)
Data sources are the origins from which data is collected for analysis. These sources can be broadly
classified into internal and external data sources.
Examples:
4/71
Enterprise Resource Planning (ERP) systems
Employee records and performance data
Usage:
Used to analyze business performance, customer behavior, and operational efficiency.
Examples:
Usage:
Helps in market analysis, competitor benchmarking, and trend forecasting.
Examples:
Usage:
Easy to analyze using traditional analytical tools and queries.
Examples:
Emails
Social media posts
5/71
Images, videos, and audio files
Customer reviews and feedback
Usage:
Analyzed using AI, ML, and Natural Language Processing (NLP) techniques.
Examples:
Usage:
Used in web analytics, system monitoring, and real-time analytics.
Conclusion:
Data analytics relies on diverse data sources—internal, external, structured, semi-structured, and
unstructured—to gain comprehensive insights. Selecting appropriate data sources is essential for
accurate analysis and effective decision-making.
Example:
6/71
An organization conducts an online survey to understand customer satisfaction with its
products.
b) Interviews
Face-to-face, telephonic, or online discussions to gather in-depth information.
Example:
c) Observations
Data is collected by observing behaviors, events, or processes.
Example:
d) Experiments
Controlled experiments conducted to study cause-and-effect relationships.
Example:
A marketing team tests two different website layouts (A/B testing) to see which generates
higher conversions.
a) Internal Records
Data available within an organization.
Example:
Example:
c) Online Sources
Data obtained from websites, social media platforms, and open data portals.
7/71
Example:
Conclusion:
Different data collection methods serve different analytical needs. Selecting the appropriate method
ensures accurate, reliable, and relevant data for effective analytics.
1. Missing Data
Incomplete data due to non-responses, system errors, or data loss.
Example:
Customer records missing contact details or age information.
2. Inaccurate Data
Data that contains errors or incorrect values.
Example:
Wrong transaction amounts or incorrect customer addresses.
3. Duplicate Data
Multiple records representing the same entity.
Example:
The same customer entered multiple times in a database.
4. Inconsistent Data
Data that does not match across different datasets or systems.
Example:
Different date formats or conflicting customer IDs across systems.
5. Outdated Data
Data that is no longer relevant or current.
Example:
8/71
Old customer preferences used for current marketing campaigns.
6. Noisy Data
Data with random errors or outliers that distort analysis.
Example:
Sensor data with sudden incorrect spikes.
Conclusion:
Addressing data quality issues through data cleaning, validation, and preprocessing is essential to
ensure accurate analytics outcomes and reliable data-driven decisions.
Examples:
Qualitative Data:
Customer opinion: “The service is excellent.”
Product category: Electronics, Clothing.
Quantitative Data:
Monthly sales: ₹50,000
Number of customers: 250
9/71
Conclusion:
Qualitative data helps understand why something happens, while quantitative data explains how
much or how many, making both essential for comprehensive data analysis.
Primary Data
Advantages:
Highly accurate and relevant
Collected for a specific objective
Limitations:
High cost and time-consuming
Requires skilled resources
Secondary Data
10/71
Advantages:
Easily available and cost-effective
Useful for trend analysis and background research
Limitations:
May be outdated or biased
Not always suitable for specific research needs
Conclusion:
Both primary and secondary data collection methods have their own strengths and limitations. A
balanced combination of both often leads to more effective and reliable analytics outcomes.
1/2
9. Explain the concept of Data Governance and its significance in analytics. (CO1)
Data Governance refers to the framework of policies, standards, processes, roles, and
responsibilities that ensure data is managed properly, securely, and consistently across an
organization. It defines how data is collected, stored, accessed, used, and maintained throughout its
lifecycle to ensure data quality, integrity, and compliance.
In data analytics, data governance acts as the foundation that ensures analysts and decision-makers
work with trustworthy, well-managed data.
11/71
information.
4. Supports Regulatory Compliance
Helps organizations comply with legal and regulatory requirements related to data protection
and usage.
5. Standardizes Data Usage Across the Organization
Ensures uniform data definitions and metrics, reducing confusion and misinterpretation.
6. Facilitates Better Collaboration
Clear data ownership and policies improve coordination among teams working with data.
Conclusion:
Data governance is critical for effective analytics as it ensures high-quality, secure, and compliant
data, enabling organizations to derive reliable insights and make informed, data-driven decisions.
10. Discuss ethical issues and challenges associated with data analytics. (CO1)
Ethical issues in data analytics arise when data is collected, analyzed, or used in ways that may
violate individual rights, privacy, or fairness. With the increasing use of AI and ML, ethical
considerations have become even more important.
1. Data Privacy
Unauthorized collection or misuse of personal data can violate individual privacy.
Example:
Example:
Leakage of customer financial or health data.
3. Bias and Discrimination
Biased data can lead to unfair or discriminatory outcomes.
Example:
AI-based hiring systems favoring certain groups due to biased training data.
4. Lack of Transparency
Complex analytics models may act as “black boxes,” making decisions hard to explain.
Example:
12/71
Automated loan rejection without clear justification.
5. Informed Consent
Users may not be fully aware of how their data is collected and used.
Example:
Using customer data beyond its originally stated purpose.
6. Data Ownership
Unclear ownership can lead to misuse or conflicts.
Example:
Third-party platforms selling user data without permission.
Conclusion:
Ethical challenges in data analytics must be addressed through strong governance, transparency,
responsible AI practices, and compliance with legal and ethical standards to ensure fair, secure, and
trustworthy use of data.
11. Explain the importance of data preprocessing in the data analytics lifecycle.
(CO2)
Data preprocessing is a crucial step in the data analytics lifecycle that involves cleaning,
transforming, and preparing raw data before it is analyzed or fed into analytical and machine
learning models. Real-world data is often incomplete, inconsistent, noisy, and unstructured, making
preprocessing essential for accurate and reliable outcomes.
13/71
5. Reduces Computational Complexity
Feature selection, dimensionality reduction, and normalization reduce data size and improve
processing efficiency.
6. Enables Meaningful Insights
Proper preprocessing ensures that analytics results truly reflect real-world patterns rather than
data errors.
Conclusion:
Data preprocessing forms the backbone of the data analytics lifecycle. Without it, even the most
advanced analytics or AI models can produce misleading or incorrect insights.
12. Demonstrate how missing values can be handled using the Pandas library in
Python. (CO2)
In Pandas, missing values are usually represented as NaN . There are multiple ways to handle them
depending on the use case.
python
import pandas as pd
data = {
'Name': ['Amit', 'Riya', 'Sohan', 'Neha'],
'Age': [25, None, 30, None],
'Salary': [50000, 60000, None, 55000]
}
df = [Link](data)
print(df)
python
[Link]()
python
14/71
[Link]().sum()
python
df_drop = [Link]()
print(df_drop)
Use case: When missing values are very few and not critical.
python
df_filled = [Link](0)
print(df_filled)
python
df['Age'].fillna(df['Age'].mean(), inplace=True)
df['Salary'].fillna(df['Salary'].mean(), inplace=True)
print(df)
python
df_ffill = [Link](method='ffill')
df_bfill = [Link](method='bfill')
Conclusion:
15/71
Handling missing values is a key preprocessing task in data analytics. Pandas provides flexible and
efficient methods—such as removal, imputation, and forward/backward filling—to ensure data
completeness and analytical accuracy.
13. Explain various techniques used to detect and handle outliers in a dataset.
(CO2)
Outliers are data points that significantly differ from the majority of observations in a dataset.
They may occur due to data entry errors, measurement errors, or genuine extreme values. Detecting
and handling outliers is essential during data preprocessing to improve the accuracy and reliability
of analytics and machine learning models.
Example:
Used in financial data to detect unusually high or low transactions.
Interquartile Range (IQR) Method
IQR = Q3 − Q1
Outliers are values below (Q1 − 1.5×IQR) or above (Q3 + 1.5×IQR).
Example:
Commonly used for salary or income analysis.
2. Visualization Techniques
Box Plot
Highlights outliers as points outside the whiskers.
Scatter Plot
Useful for identifying abnormal patterns or extreme values.
3. Distance-Based Methods
Detects outliers based on their distance from other data points.
Used in clustering-based approaches.
4. Machine Learning-Based Methods
16/71
Isolation Forest
Separates anomalies by randomly partitioning data.
DBSCAN
Identifies outliers as points that do not belong to any cluster.
Conclusion:
Outlier detection and handling improve data quality and model reliability. The chosen technique
depends on the nature of data, domain knowledge, and analytical objectives.
14. Demonstrate the process of identifying and removing duplicate records using
Python. (CO2)
Duplicate records can distort analysis results and must be removed during preprocessing. Python’s
Pandas library provides efficient methods to handle duplicates.
python
import pandas as pd
data = {
'ID': [1, 2, 2, 3, 4, 4],
'Name': ['Amit', 'Riya', 'Riya', 'Sohan', 'Neha', 'Neha'],
'Marks': [85, 90, 90, 78, 88, 88]
}
17/71
df = [Link](data)
print(df)
python
[Link]()
python
[Link]().sum()
python
df[[Link]()]
python
df_cleaned = df.drop_duplicates()
print(df_cleaned)
python
df_cleaned_col = df.drop_duplicates(subset=['ID'])
Conclusion:
18/71
Removing duplicate records is a critical preprocessing step. Pandas makes it easy to identify and
clean duplicates, ensuring accurate and reliable analytics results.
1. Normalization
Normalization scales data to a fixed range, usually between 0 and 1. It is useful when features have
different scales and when distance-based algorithms are used.
X − Xmin
Xnorm =
Xmax − Xmin
Key Characteristics:
Values lie between 0 and 1
Preserves relative relationships
Sensitive to outliers
Example:
Scaling student marks or product prices before applying clustering algorithms.
2. Standardization
Standardization transforms data so that it has a mean of 0 and a standard deviation of 1.
Formula (Z-score):
X −μ
Xstd =
σ
where
μ = mean, σ = standard deviation
Key Characteristics:
19/71
Results in centered data with unit variance
Less affected by outliers compared to normalization
Suitable for algorithms assuming normal distribution
Example:
Standardizing features like age, income, and experience before applying regression or classification
models.
Conclusion:
Normalization and standardization ensure that features contribute equally to analysis and modeling,
improving accuracy and performance.
16. Demonstrate how multiple datasets can be integrated using Pandas. (CO2)
Dataset integration combines data from multiple sources into a single dataset for comprehensive
analysis. Pandas provides powerful functions such as merge() , concat() , and join() .
python
import pandas as pd
data1 = {
'ID': [1, 2, 3],
20/71
'Name': ['Amit', 'Riya', 'Sohan']
}
data2 = {
'ID': [1, 2, 3],
'Marks': [85, 90, 78]
}
df1 = [Link](data1)
df2 = [Link](data2)
python
python
python
df1.set_index('ID', inplace=True)
df2.set_index('ID', inplace=True)
joined_df = [Link](df2)
print(joined_df)
Conclusion:
21/71
Integrating multiple datasets is essential for holistic data analysis. Pandas provides flexible and
efficient methods to merge, concatenate, and join datasets based on analytical requirements.
17. Explain the concept of feature engineering and its role in improving model
performance. (CO2)
Feature engineering is the process of selecting, creating, transforming, and optimizing input
variables (features) from raw data to improve the performance of analytical and machine learning
models. It bridges the gap between raw data and effective model learning.
In many real-world projects, feature engineering has a greater impact on model performance than
the choice of algorithm itself.
Feature selection
Feature creation
Feature transformation
Feature scaling and encoding
Examples:
Creating “Total Purchase Amount” from individual transaction values
Extracting “Day”, “Month”, and “Year” from date fields
Conclusion:
Feature engineering is a critical step in analytics and machine learning pipelines, directly
22/71
influencing model accuracy, robustness, and overall performance.
python
import pandas as pd
data = {
'City': ['Pune', 'Mumbai', 'Delhi', 'Pune'],
'Gender': ['Male', 'Female', 'Female', 'Male']
}
df = [Link](data)
print(df)
1. Label Encoding
Converts categories into numeric labels.
python
le = LabelEncoder()
df['City_Label'] = le.fit_transform(df['City'])
print(df)
Use Case:
Ordinal categories with order (e.g., Low, Medium, High)
2. One-Hot Encoding
23/71
Creates binary columns for each category.
python
Use Case:
Nominal categories without order (e.g., City names)
3. Binary Encoding
Reduces dimensionality compared to one-hot encoding.
python
be = BinaryEncoder(cols=['City'])
df_binary = be.fit_transform(df)
print(df_binary)
Use Case:
High-cardinality categorical features
4. Ordinal Encoding
Assigns ordered numerical values.
python
Conclusion:
24/71
Feature encoding is essential for converting categorical data into machine-readable form. Choosing
the right encoding technique improves model performance, efficiency, and interpretability.
19. Demonstrate how to compute mean, variance, and standard deviation using
NumPy. (CO2)
NumPy is a powerful Python library used for numerical computations in data analytics and machine
learning. It provides efficient functions to calculate basic statistical measures such as mean,
variance, and standard deviation.
python
import numpy as np
python
mean_value = [Link](data)
print("Mean:", mean_value)
python
variance_value = [Link](data)
print("Variance:", variance_value)
python
25/71
std_deviation = [Link](data)
print("Standard Deviation:", std_deviation)
Output Explanation:
Mean represents the average value of the dataset.
Variance measures how far the values spread from the mean.
Standard Deviation indicates the amount of variation or dispersion in the dataset.
Conclusion:
NumPy provides simple and efficient methods to compute essential statistical measures, forming
the foundation for data analysis and machine learning tasks.
20. Explain the importance of data scaling and its impact on machine learning
algorithms. (CO2)
Data scaling is the process of transforming numerical features to a common scale without
distorting differences in the ranges of values. It is a critical preprocessing step in machine learning.
26/71
Algorithm Type Impact of Scaling
Conclusion:
Data scaling plays a vital role in ensuring optimal performance, accuracy, and efficiency of
machine learning models, especially in distance-based and gradient-based algorithms.
21. Define Exploratory Data Analysis (EDA) and explain its objectives in data
analytics. (CO3)
Exploratory Data Analysis (EDA) is an approach used to analyze datasets by summarizing their
main characteristics, often using visual methods. It helps analysts understand the structure, patterns,
relationships, and anomalies in data before applying advanced statistical or machine learning
techniques.
EDA focuses on exploring the data rather than confirming predefined hypotheses, making it a
critical step in the data analytics workflow.
27/71
Objectives of EDA:
Conclusion:
EDA provides a strong foundation for data analytics by enabling a deep understanding of the data,
reducing errors, and improving the quality of insights and models.
python
import pandas as pd
import [Link] as plt
data = {
'Marks': [55, 60, 65, 70, 75, 80, 85, 90]
}
df = [Link](data)
print(df)
28/71
Step 2: Descriptive Statistics
python
df['Marks'].describe()
python
mean = df['Marks'].mean()
median = df['Marks'].median()
mode = df['Marks'].mode()
print("Mean:", mean)
print("Median:", median)
print("Mode:", [Link])
python
[Link](df['Marks'], bins=5)
[Link]('Marks')
[Link]('Frequency')
[Link]('Histogram of Marks')
[Link]()
python
[Link](df['Marks'])
[Link]('Box Plot of Marks')
[Link]()
29/71
Conclusion:
Univariate analysis helps in understanding the distribution and spread of a single variable. It is a
fundamental EDA technique that supports better data preprocessing and modeling decisions.
A) Bivariate Analysis
1. Numerical vs Numerical
Techniques: Correlation, Scatter Plot
python
import pandas as pd
import [Link] as plt
data = {
'Hours_Studied': [1, 2, 3, 4, 5, 6],
'Marks': [40, 50, 60, 70, 80, 90]
}
df = [Link](data)
# Correlation
print([Link]())
# Scatter Plot
[Link](df['Hours_Studied'], df['Marks'])
[Link]('Hours Studied')
[Link]('Marks')
[Link]('Hours Studied vs Marks')
[Link]()
30/71
Interpretation:
Positive correlation indicates that marks increase with study hours.
2. Categorical vs Numerical
Techniques: Grouped statistics, Box Plot
python
data = {
'Gender': ['Male', 'Female', 'Female', 'Male', 'Male'],
'Score': [78, 85, 88, 72, 80]
}
df = [Link](data)
# Grouped Mean
print([Link]('Gender')['Score'].mean())
# Box Plot
[Link](column='Score', by='Gender')
[Link]('Score by Gender')
[Link]('')
[Link]()
B) Multivariate Analysis
python
data = {
'Maths': [70, 80, 90, 85, 75],
'Science': [72, 82, 88, 86, 78],
'English': [65, 70, 75, 80, 85]
}
df = [Link](data)
31/71
# Pair Plot
[Link](df)
[Link]()
# Correlation Heatmap
[Link]([Link](), annot=True, cmap='coolwarm')
[Link]('Correlation Heatmap')
[Link]()
Interpretation:
Multivariate analysis reveals how multiple variables interact with each other simultaneously.
Conclusion:
Bivariate and multivariate analyses help uncover relationships, dependencies, and patterns in data,
guiding better feature selection and modeling decisions.
24. Demonstrate how to compute descriptive statistics for a dataset using Pandas.
(CO3)
Descriptive statistics summarize the main characteristics of a dataset, such as central tendency,
dispersion, and distribution.
python
import pandas as pd
data = {
'Age': [22, 25, 30, 28, 35],
'Salary': [30000, 40000, 50000, 45000, 60000]
}
df = [Link](data)
print(df)
32/71
Step 2: Overall Descriptive Statistics
python
[Link]()
python
# Mean
[Link]()
# Median
[Link]()
# Standard Deviation
[Link]()
# Variance
[Link]()
python
data = {
'Department': ['IT', 'HR', 'IT', 'Finance', 'HR']
}
df_cat = [Link](data)
df_cat.describe()
Conclusion:
Pandas provides powerful and easy-to-use functions like describe() , mean() , median() ,
and std() to compute descriptive statistics efficiently, making it an essential tool for EDA and
data analytics.
33/71
25. Explain the concept of probability distributions and describe Normal,
Binomial, and Poisson distributions with examples. (CO3)
A probability distribution describes how the values of a random variable are distributed and the
likelihood of each possible outcome. It provides a mathematical framework to model uncertainty
and variability in real-world data and is widely used in statistics, data analytics, and machine
learning.
1. Normal Distribution
The Normal distribution is a continuous probability distribution that is symmetric and bell-
shaped, defined by its mean (μ) and standard deviation (σ).
Key Characteristics:
Example:
Distribution of student exam scores
Heights of people in a population
2. Binomial Distribution
The Binomial distribution is a discrete probability distribution that represents the number of
successes in a fixed number of independent trials, each with two possible outcomes (success or
failure).
Key Conditions:
Example:
Number of heads obtained when flipping a coin 10 times
Number of defective items in a batch
34/71
3. Poisson Distribution
The Poisson distribution models the number of events occurring in a fixed interval of time or
space, given a constant average rate.
Key Characteristics:
Example:
Number of calls received by a call center per hour
Number of accidents at a traffic junction per day
Conclusion:
Probability distributions help model real-world randomness. Normal, Binomial, and Poisson
distributions are widely used in analytics to represent continuous data, binary outcomes, and event
occurrences respectively.
26. Explain the concept of hypothesis testing and describe the null and alternative
hypotheses. (CO3)
Hypothesis testing is a statistical method used to make decisions or inferences about a population
based on sample data. It helps determine whether observed results are statistically significant or
occurred by chance.
35/71
The null hypothesis states that there is no effect, no difference, or no relationship between
variables. It represents the default assumption.
Example:
Example:
Conclusion:
Hypothesis testing provides a structured approach to decision-making using data. The null and
alternative hypotheses form the foundation for statistical inference, enabling analysts to draw
meaningful conclusions with confidence.
27. Demonstrate the application of a t-test or chi-square test using Python and
interpret the results. (CO3)
Statistical hypothesis tests help determine whether observed differences or relationships in data are
statistically significant.
36/71
A t-test is used to compare the means of two independent groups and check whether the difference
between them is significant.
Example Scenario
Compare the exam scores of Group A and Group B students.
python
import numpy as np
from [Link] import ttest_ind
python
Interpretation
p-value < 0.05 → Reject the null hypothesis (significant difference exists)
p-value ≥ 0.05 → Fail to reject the null hypothesis (no significant difference)
Conclusion (Example):
If the p-value is less than 0.05, we conclude that the average scores of Group A and Group B are
significantly different.
B) Chi-Square Test
The chi-square test checks whether there is an association between two categorical variables.
Example Scenario
Check whether Gender and Course Preference are related.
python
37/71
import pandas as pd
from [Link] import chi2_contingency
python
Interpretation
p-value < 0.05 → Significant association exists
p-value ≥ 0.05 → No significant association
Conclusion:
t-tests analyze differences between means, while chi-square tests analyze relationships between
categorical variables. Both are essential tools in data-driven decision-making.
28. Explain the concept of correlation and differentiate between its types. (CO3)
Correlation is a statistical measure that indicates the strength and direction of the relationship
between two variables. It helps understand how one variable changes in relation to another.
Types of Correlation
1. Positive Correlation
38/71
Both variables move in the same direction.
Example:
Correlation Value:
+0.1 to +1.0
2. Negative Correlation
One variable increases while the other decreases.
Example:
Correlation Value:
-0.1 to -1.0
Example:
Correlation Value:
0
a) Pearson Correlation
Measures linear relationship
Used for numerical, continuous data
b) Spearman Correlation
Measures monotonic relationship
Used for ranked or non-linear data
39/71
c) Kendall Correlation
Used for ordinal data and small samples
Conclusion:
Correlation helps identify relationships between variables but does not imply causation.
Understanding its types is essential for accurate data interpretation and analysis.
29. Explain regression analysis and describe its applications in data analytics.
(CO3)
Regression analysis is a statistical technique used to model and analyze the relationship between a
dependent variable (target) and one or more independent variables (predictors). It helps in
understanding how changes in predictor variables influence the target variable and is widely used
for prediction and forecasting.
y = β0 + β1 x
Example:
Predicting salary based on years of experience.
2. Multiple Linear Regression
Multiple independent variables
Equation:
y = β 0 + β 1 x1 + β 2 x2 + ⋯ + β n xn
Example:
Predicting house price based on size, location, and number of rooms.
3. Polynomial Regression
Models non-linear relationships using polynomial terms.
4. Logistic Regression
Used for classification problems with binary outcomes.
40/71
Example:
Predicting whether a customer will churn (Yes/No).
Conclusion:
Regression analysis is a foundational tool in data analytics that enables prediction, trend
identification, and data-driven decision-making by quantifying relationships between variables.
30. Demonstrate how data relationships can be visualized using Matplotlib and
Seaborn. (CO3)
Data visualization helps in understanding relationships, patterns, and trends between variables
effectively.
A) Using Matplotlib
python
data = {
41/71
'Hours_Studied': [1, 2, 3, 4, 5, 6],
'Marks': [45, 55, 65, 75, 85, 92]
}
df = [Link](data)
[Link](df['Hours_Studied'], df['Marks'])
[Link]('Hours Studied')
[Link]('Marks')
[Link]('Hours Studied vs Marks')
[Link]()
B) Using Seaborn
1. Regression Plot
python
python
data = {
'Maths': [70, 80, 90, 85, 75],
'Science': [72, 82, 88, 86, 78],
'English': [65, 70, 75, 80, 85]
}
df_multi = [Link](data)
[Link](df_multi)
[Link]()
42/71
3. Heatmap (Correlation Visualization)
python
Conclusion:
Matplotlib and Seaborn provide powerful visualization tools to explore and interpret data
relationships. Visual analysis enhances understanding, supports EDA, and improves data-driven
insights in analytics projects.
31. Define Machine Learning and explain its major types with suitable examples.
(CO4)
Machine Learning (ML) is a subset of Artificial Intelligence (AI) that enables computer systems
to learn patterns from data and improve their performance on a specific task without being
explicitly programmed. ML algorithms use data to build models that make predictions or decisions
based on past experience.
1. Supervised Learning
In supervised learning, the model is trained using labeled data, where both input features and output
labels are known.
Common Algorithms:
Linear Regression, Logistic Regression, Decision Trees, Support Vector Machines (SVM)
Examples:
43/71
2. Unsupervised Learning
In unsupervised learning, the model works with unlabeled data and identifies hidden patterns or
groupings.
Common Algorithms:
K-Means Clustering, Hierarchical Clustering, Principal Component Analysis (PCA)
Examples:
3. Semi-Supervised Learning
This approach uses a small amount of labeled data along with a large amount of unlabeled data.
Example:
4. Reinforcement Learning
In reinforcement learning, an agent learns by interacting with an environment and receiving
rewards or penalties.
Examples:
Conclusion:
Machine learning enables systems to automatically learn from data. Its various types allow it to
solve diverse real-world problems across industries.
32. Explain the complete machine learning workflow from data collection to model
evaluation. (CO4)
44/71
The machine learning workflow is a structured process that ensures models are accurate, reliable,
and ready for real-world deployment.
1. Data Collection
Gather relevant data from sources such as databases, sensors, APIs, or surveys.
Example:
Customer transaction data from an e-commerce platform.
2. Data Preprocessing
Clean and prepare data by handling missing values, outliers, duplicates, and scaling features.
4. Feature Engineering
Select, create, and transform features to improve model performance.
5. Model Selection
Choose suitable algorithms based on the problem type (regression, classification, clustering).
6. Model Training
Train the model using training data to learn patterns.
7. Model Evaluation
Evaluate model performance using metrics such as accuracy, precision, recall, F1-score, RMSE.
8. Model Tuning
Optimize model parameters using techniques like cross-validation and hyperparameter tuning.
45/71
9. Model Deployment
Deploy the trained model into a production environment.
Conclusion:
A well-defined machine learning workflow ensures the development of robust, scalable, and high-
performing models, making ML solutions effective in real-world applications.
python
import pandas as pd
import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression
from [Link] import mean_squared_error, r2_score
python
data = {
'Experience': [1, 2, 3, 4, 5, 6],
'Salary': [30000, 35000, 40000, 45000, 50000, 55000]
46/71
}
df = [Link](data)
python
python
model = LinearRegression()
[Link](X_train, y_train)
python
y_pred = [Link](X_test)
python
47/71
Conclusion:
Linear Regression models relationships between variables and provides reliable predictions for
continuous outcomes when relationships are linear.
Examples:
Algorithm Overview
1. Logistic Regression
Uses sigmoid function
Best for linearly separable data
48/71
Outputs probability scores
2. Decision Tree
Uses rule-based splitting
Easy to interpret
Prone to overfitting
3. Random Forest
Combines multiple decision trees
Reduces overfitting
Provides high accuracy and robustness
Conclusion:
Choosing the right classification algorithm depends on dataset size, complexity, interpretability
needs, and performance requirements.
35. Demonstrate the implementation of a Decision Tree classifier and evaluate its
performance using appropriate metrics. (CO4)
A Decision Tree classifier is a supervised learning algorithm used for classification tasks. It splits
data into branches based on feature values to make predictions.
python
import pandas as pd
from sklearn.model_selection import train_test_split
from [Link] import DecisionTreeClassifier
from [Link] import accuracy_score, confusion_matrix, classification_report
python
49/71
data = {
'Hours_Studied': [1, 2, 3, 4, 5, 6, 7, 8],
'Passed': [0, 0, 0, 1, 1, 1, 1, 1]
}
df = [Link](data)
X = df[['Hours_Studied']]
y = df['Passed']
python
python
python
y_pred = [Link](X_test)
python
50/71
print("Accuracy:", accuracy_score(y_test, y_pred))
print("Confusion Matrix:\n", confusion_matrix(y_test, y_pred))
print("Classification Report:\n", classification_report(y_test, y_pred))
Conclusion:
Decision Trees are powerful and interpretable models for classification but require proper tuning to
avoid overfitting.
36. Explain the concepts of overfitting and underfitting and discuss techniques to
address them. (CO4)
Overfitting and underfitting are common problems in machine learning that affect model
generalization.
Overfitting
Occurs when a model learns the training data too well, including noise and outliers, resulting in
poor performance on unseen data.
Symptoms:
Causes:
Complex model
Small dataset
Too many features
51/71
Underfitting
Occurs when a model is too simple to capture underlying data patterns.
Symptoms:
Causes:
Oversimplified model
Insufficient features
Inadequate training
Conclusion:
Balancing model complexity is essential to achieve good generalization. Addressing overfitting and
underfitting ensures robust and accurate machine learning models.
37. Explain the working principle of K-Means clustering and its real-world
applications. (CO4)
K-Means clustering is an unsupervised machine learning algorithm used to group data points into
K distinct clusters based on similarity. Each cluster is represented by its centroid (mean of data
points in that cluster).
52/71
3. Assign data points to nearest centroid
Each data point is assigned to the cluster whose centroid is closest (usually using Euclidean
distance).
4. Update centroids
Recalculate centroids by taking the mean of all points assigned to each cluster.
5. Repeat steps 3 and 4
Continue until centroids no longer change or maximum iterations are reached.
Key Characteristics
Distance-based algorithm
Works best with numerical data
Sensitive to initial centroid selection and outliers
Requires scaling of data
Conclusion:
K-Means is a simple yet powerful clustering algorithm widely used for grouping similar data points
and discovering hidden patterns in unlabeled datasets.
53/71
We will segment customers based on Annual Income and Spending Score.
import pandas as pd
import [Link] as plt
from [Link] import KMeans
from [Link] import StandardScaler
data = {
'Annual_Income': [15, 16, 17, 18, 19, 60, 62, 65, 70, 75],
'Spending_Score': [39, 42, 45, 48, 50, 80, 82, 85, 90, 95]
}
df = [Link](data)
scaler = StandardScaler()
scaled_data = scaler.fit_transform(df)
54/71
Step 5: Visualize Customer Segments
python
[Link](df['Annual_Income'], df['Spending_Score'],
c=df['Cluster'])
[Link]('Annual Income')
[Link]('Spending Score')
[Link]('Customer Segmentation using K-Means')
[Link]()
Interpretation
Customers are grouped into two distinct segments
One cluster represents low income–low spending
Another cluster represents high income–high spending
Conclusion:
K-Means clustering is highly effective for customer segmentation. Visualization helps businesses
clearly understand customer groups and design targeted marketing strategies.
55/71
Principal Component Analysis (PCA)
PCA is a popular linear dimensionality reduction technique that transforms original features into a
new set of uncorrelated variables called principal components, ordered by the amount of variance
they explain.
python
import pandas as pd
import numpy as np
from [Link] import StandardScaler
from [Link] import PCA
python
data = {
'Maths': [70, 80, 90, 85, 75],
'Science': [72, 82, 88, 86, 78],
'English': [65, 70, 75, 80, 85]
}
df = [Link](data)
python
56/71
scaler = StandardScaler()
scaled_data = scaler.fit_transform(df)
python
pca = PCA(n_components=2)
pca_data = pca.fit_transform(scaled_data)
python
Interpretation:
PC1 explains the maximum variance
PC2 explains the remaining variance
Dimensionality reduced from 3 → 2 while retaining most information
Conclusion:
Dimensionality reduction using PCA simplifies datasets, improves efficiency, and enhances model
performance while retaining essential patterns in the data.
40. Explain different model evaluation metrics used in machine learning. (CO4)
Model evaluation metrics are used to assess how well a machine learning model performs on
unseen data. The choice of metric depends on the problem type (classification or regression).
57/71
A) Classification Evaluation Metrics
1. Accuracy
Proportion of correctly classified instances.
TP + TN
Accuracy =
TP + TN + FP + FN
2. Precision
How many predicted positives are actually positive.
TP
P recision =
TP + FP
3. Recall (Sensitivity)
How many actual positives are correctly predicted.
TP
Recall =
TP + FN
4. F1-Score
Harmonic mean of precision and recall.
P recision × Recall
F1 = 2 ×
P recision + Recall
5. Confusion Matrix
Shows TP, TN, FP, FN in matrix form.
6. ROC–AUC
Measures the model’s ability to distinguish between classes.
58/71
Metric Selection Guide
Problem Type Common Metrics
Conclusion:
Model evaluation metrics provide quantitative insight into model performance. Choosing the right
metric ensures reliable comparison, validation, and deployment of machine learning models.
Conclusion:
Data visualization bridges the gap between data and decision-makers, ensuring insights are
59/71
accessible, actionable, and impactful.
60/71
Drag Sales → Angle
Select Pie Chart
Add Filters
Drag Region to Filters
Show filter → Apply to all worksheets
Conclusion:
61/71
Interactive dashboards in Tableau empower users to explore data dynamically, uncover insights
instantly, and make informed decisions. They are a cornerstone of modern analytics
communication.
62/71
Conclusion:
Effective data storytelling transforms raw analysis into meaningful business narratives that enable
stakeholders to understand insights quickly and act confidently.
44. Demonstrate how Tableau visualizations can be used to derive business insights.
(CO5)
Tableau visualizations help uncover patterns, trends, and opportunities by enabling interactive and
visual data exploration.
Business Scenario
A retail company wants to understand:
Which regions perform best
Which product categories generate the most profit
Sales trends over time
63/71
Business Insight:
➡ Focus marketing efforts on low-performing regions.
Business Insight:
➡ Plan inventory and promotions based on seasonal demand.
Business Insight:
➡ Prioritize high-margin products over high-volume, low-profit ones.
Insight Derived:
Enables deep-dive analysis without creating multiple reports.
Business Insight:
➡ Managers can instantly analyze region-specific performance.
64/71
Profit Margin
Insight Derived:
High-level business health snapshot.
Business Insight:
➡ Supports executive-level decision-making.
Conclusion:
Tableau visualizations convert raw business data into actionable insights by enabling interactive
exploration, quick pattern recognition, and data-driven strategic decisions.
45. Discuss recent trends in Artificial Intelligence and real-time analytics. (CO5)
Recent advancements in Artificial Intelligence (AI) combined with real-time analytics are
transforming how organizations operate, compete, and make decisions. These trends focus on
speed, automation, intelligence, and scalability.
65/71
Recent Trends in Real-Time Analytics
1. Streaming Data Processing
Continuous analysis of live data streams from sensors, applications, and user interactions.
2. Real-Time Personalization
Businesses personalize content, recommendations, and offers instantly based on live user
behavior.
3. Predictive Analytics in Real Time
Combining AI models with live data to forecast outcomes immediately.
4. Event-Driven Architectures
Systems react instantly to events such as transactions, clicks, or alerts.
5. Cloud-Based Real-Time Analytics Platforms
Scalable cloud solutions enable real-time insights without heavy infrastructure costs.
Conclusion:
The convergence of AI and real-time analytics enables organizations to shift from reactive to
proactive and autonomous decision-making, delivering speed, intelligence, and competitive
advantage.
66/71
4. Identification of Trends and Risks
Dashboards highlight patterns, anomalies, and potential risks before they escalate.
5. Alignment with Business Goals
Ensures teams stay aligned with strategic objectives by continuously tracking progress.
6. Drill-Down Analysis
Interactive dashboards allow managers to explore underlying data for deeper insights.
7. Enhanced Communication
Dashboards act as a common reference point for discussions among stakeholders.
Conclusion:
Dashboards are essential decision-support tools that transform complex data into actionable
insights, enabling managers to lead with clarity, confidence, and strategic focus.
67/71
Example:
Tracking user behavior beyond the stated purpose of data collection.
Example:
AI-based hiring or loan approval systems favoring specific groups.
Example:
Wrong medical diagnosis suggested by an AI system.
68/71
Ethical concern when AI decisions directly impact human lives
Example:
Automated surveillance or predictive policing systems.
Conclusion:
Ethical AI-driven analytics requires strong governance, transparency, bias mitigation, data privacy
protection, and human oversight to ensure fairness, trust, and social responsibility.
48. Analyze a real-world dataset and present insights using an interactive Tableau
dashboard. (CO5)
Below is a practical demonstration of analyzing a real-world business dataset and deriving
insights using Tableau.
Dataset Example
Retail Sales Dataset
Fields:
Order Date
Region
Category
Sales
Profit
Customer Segment
69/71
1. Open Tableau Desktop
2. Click Connect → Excel / CSV
3. Load the retail sales dataset
4. Verify data types (Date, Dimension, Measure)
Insight:
Identifies high- and low-performing regions
Insight:
Insight:
Reveals categories with high sales but low profit
Insight:
Highlights most valuable customer segments
70/71
Step 4: Add Interactivity
Add Filters: Region, Category, Segment
Enable Dashboard Actions (click to filter)
Add Tooltips for detailed insights
Conclusion:
An interactive Tableau dashboard transforms raw datasets into actionable business insights,
enabling stakeholders to explore data dynamically, identify opportunities, and make informed
strategic decisions.
71/71