0% found this document useful (0 votes)
2 views12 pages

Notes-Data Science Methodology

This document outlines the Data Science Methodology, detailing its lifecycle from problem definition to feedback, and includes stages such as data collection, preparation, modeling, evaluation, and deployment. It features objective and subjective questions to assess understanding of key concepts, techniques, and evaluation metrics in data science. Additionally, it emphasizes the importance of model validation and provides insights into various analytics methods.

Uploaded by

aathifa20251437
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views12 pages

Notes-Data Science Methodology

This document outlines the Data Science Methodology, detailing its lifecycle from problem definition to feedback, and includes stages such as data collection, preparation, modeling, evaluation, and deployment. It features objective and subjective questions to assess understanding of key concepts, techniques, and evaluation metrics in data science. Additionally, it emphasizes the importance of model validation and provides insights into various analytics methods.

Uploaded by

aathifa20251437
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

UNIT 1:DATA SCIENCE METHODOLOGY

Learning Objectives:

 Understand the concept of Data Science Methodology and its complete lifecycle from
problem definition to feedback.
 Explain and apply the stages of Data Science, including problem identification, data
collection, data preparation, modeling, evaluation, and deployment.
 Understand model validation techniques and interpret evaluation metrics such as
accuracy, precision, recall, and F1-score.
 Demonstrate the ability to evaluate machine learning models using Python and analyze
their performance effectively.

SECTION A (Objective Type Questions)


A. Choose the correct option.
1. Who introduced the Data Science Methodology, also known as the Foundational
Methodology for Data Science?
a. Andrew Ng b. John Rollins c. Geoffrey Hinton d. Yann LeCun
2. How many steps are included in the Data Science Methodology?
a. 5 b. 8 c. 10 d. 12
3. In which module is data collection emphasised?
a. From Problem to Approach b. From Understanding to Preparation
c. From Requirements to Collection d. From Modelling to Evaluation

4. Which framework is commonly used in the 'Problem Scoping and Defining' stage?
a. KNN Algorithm b. 5W1H Problem Canvas
c. Decision Trees d. Linear Regression
5. Which Data Analytics method is used to predict future outcomes based on historical
data?
a. Descriptive Analytics b. Diagnostic Analytics
c. Predictive Analytics d. Prescriptive Analytics
6. Which of the following is an example of Diagnostic Analytics?
a. Calculating the average marks of students.
b. Identifying why a mobile phone company's sales dropped.
c. Forecasting next year's sales based on past trends.
d. Recommending solutions to improve marketing.
7. Which of the following techniques is used in Predictive Analytics?
a. Root Cause Analysis b. Hypothesis Testing
c. Regression d. Mean, Median, and Mode
8. Which of the following is NOT a common method of data collection?
a. Forms and Questionnaires b. Focus Groups
c. Social Media Monitoring d. Fictional Storytelling
9. Which of the following is an example of a secondary data source?
a. Focus Groups b. Experiments and Observations
c. Books, Journals, and Research Papers d. Oral Histories
10. Which of the following are common forms of feedback in the AI model life cycle?
a. User reviews b. Performance reports
c. Error logs d. All of these
Answers: 1. b. 2. c. 3. c. 4. b. 5. c. 6. b. 7. c. 8. d. 9. c. 10. d.

B. Fill in the blanks.


1. The Data Science Methodology provides a ________way to collect, process, and
understand data.
2. Descriptive Analytics helps summarise ______data to identify trends and patterns.
3. Social media data is considered ________data if collected directly from owned
platforms and ________ data if sourced from third-party reports.
4. ________uses both primary methods and secondary methods for deeper insights.
5. To mitigate risks, the model is often introduced in a________ test environment or to a
limited group of users before broader implementation.
6. The ________ stage is the final step in the AI model life cycle.
7._______ techniques involve assessing a machine learning model's performance on
training and test data.
8. The ________method is used to evaluate the performance of a Machine Learning
model.
9. Instead of dividing the data into just a training set and a test set, cross-validation splits
the dataset into multiple parts called __________ .
10. _____________measures how well the model predicts positive cases correctly.
Answers: 1. structured 2. past 3. primary, secondary 4. Combination Research
5. controlled 6. feedback 7. Evaluation 8. train-test split
9. folds 10. Precision

C. State whether the following statements are true or false:


1. Correlation Analysis helps find relationships between different factors.t
2. Predictive Analytics cannot help in reducing risks.
3. Online tracking is always considered secondary data.
4. XML files are considered structured data.
5. Data scientists try different algorithms to find the most suitable model for solving a
given problem.t
6. Feature Engineering is a machine-driven process where the machine decides which
features to add.
7. Model evaluation is performed only after the model has been deployed.
8. The deployment process may involve infrastructure upgrades and continuous
monitoring.t
9. Cross-validation takes more time to run.
10. High precision means fewer False Negatives.
Answers: 1. True 2. False 3. False 4. False 5. True 6. False 7. False 8. True 9. True 10.
False
SECTION B (Subjective Type Questions)
A. Short answer type questions.
1. Name the five modules of the Data Science Methodology.
Ans. The five modules are:
a. From Problem to Approach
b. From Requirements to Collection
c. From Understanding to Preparation
d. From Modelling to Evaluation
e. From Deployment to Feedback
2. Name any two techniques used in Diagnostic Analytics to find the root cause.
Ans. Two techniques used in Diagnostic Analytics to find the root cause are:
• Root Cause Analysis – Identifies the main reason for a problem.
• Hypothesis Testing – Verifies assumptions about data.
3. How does Simulation help in Prescriptive Analytics?
Ans. Simulation helps in Prescriptive Analytics by testing different scenarios to determine
possible outcomes.
4. What is data collection?
Ans. Data collection is a systematic process of gathering information through
observations, measurements, or surveys.
5. What are mixed data sources?
Ans. Mixed data sources combine both primary and secondary data to improve accuracy,
reliability, and comprehensiveness in research.
6. What is the purpose of the Data Preparation stage in data analysis?
Ans. The Data Preparation stage includes all the activities needed to organise and process
raw data before using it in the modelling step. This process transforms data into a format
that makes it easier to work with and analyse.
7. What is model validation?
Ans. Model validation is the process of checking how well a trained model performs on
new, unseen data. It helps us measure the accuracy and reliability of the model before
using it in real-world situations.

8. What is the main difference between classification and regression problems?


Ans. The main difference between classification and regression problems are as follows:
• Classification problems aim to sort data into categories (e.g., identifying whether an
email is spam or not).
• Regression problems aim to predict continuous values (e.g., estimating house prices
based on floor area.

9. How does K-Fold Cross-Validation improve AI model evaluation?


Ans. K-Fold Cross-Validation divides data into multiple subsets (folds), training the
model on different parts each [Link] ensures the model is tested thoroughly, reducing
bias and improving generalization.
10. Identify the type of analytics in each of the following scenarios:
i. An airline company adjusts ticket prices based on travel demand to maximize profits.
ii. A bank estimates which customers are likely to apply for a loan based on their financial
history.
iii. A cricket team studies batting averages of players over the last season to identify their
best performers.
iv. A mobile phone company sees a drop in sales and analyses whether it was due to high
prices, poor marketing,or competition from new brands.
v. A teacher calculates the average marks of students in a math test to understand the
overall class performance.
Ans. i. Prescriptive ii. Predictive iii. Descriptive iv. Diagnostic v. Descriptive
B. Long answer type questions.
1. What are the key steps involved in identifying and defining the specific data
needed for a project?
Ans. The key steps involved in identifying and defining the specific data needed for a
project include:
• Identifying the types of data required – Data can be in the form of numbers (quantitative
data., words (text), or images (visual data..)
• Choosing the structure of the data – Data can be stored in different formats such as tables
(spreadsheets), text files, or databases.
• Finding sources of data – Data can be collected from company records, surveys,
websites, sensors, or public datasets.
• Preparing the data for analysis – Data often needs cleaning and organising to remove
errors, duplicates, or missing values before it is ready for use.
2. Explain primary data sources with examples.
Ans. Primary data is collected directly from its original source, making it fresh, raw, and
highly accurate. Since it has not been processed, it provides the most reliable insights for
research. Common methods of primary data collection include:
• Surveys & Interviews – Gathering opinions and feedback directly from individuals.
• Experiments & Observations – Recording firsthand data through hands-on research.
• IoT Sensors & Smart Devices – Collecting real-time data from smartwatches, weather
sensors, and automated systems.
• Marketing Campaigns & Feedback Forms – Capturing customer responses and
preferences.
• Focus Groups – Conducting group discussions to gain deeper insights.
• Oral Histories – Preserving personal accounts through recorded interviews.
3. Why is the Data Understanding stage important in data analysis?
Ans. The Data Understanding stage involves evaluating the collected dataset to ensure it is
relevant, complete, and suitable for solving the given problem. This process helps
determine whether the data truly represents the issue at hand or if additional collection or
adjustments are needed. This step is important because:
• Ensures that the collected data is relevant to the problem being addressed.
• Helps identify any missing, inaccurate, or irrelevant data before proceeding to the next
stage.
• Allows for assessing whether additional data collection or corrections are needed.
4. What are the different techniques used to explore and understand data?
Ans. There are different techniques to explore and understand the data:
a. Descriptive Statistics- It help summarise and describe the data using:
• Mean, Median, Mode – To find the average and most common values.
• Range, Variance, Standard Deviation – To check how spread out the data is.
• Pairwise Correlation – To see if two variables are related (e.g., "Do higher discounts lead
to more sales?").
b. Data Visualization- It helps us see patterns and trends in the data using:
• Histograms – Show how frequently different values appear in the data.
• Pie Charts and Bar Graphs – Help compare different categories of data.
• Scatter Plots – Show relationships between two variables.
5. What is feature engineering in data preparation, and how can it improve the
accuracy of Machine Learning models?
Provide examples.
Ans. Feature engineering is an important part of data preparation. It involves creating,
selecting, or modifying variables from raw data to improve the accuracy of Machine
Learning [Link] example, you are building an ML model to predict student exam
scores based on various factors. Your raw data includes study hours per day, attendance
percentage and number of assignments submitted.
Through feature engineering, we can create new meaningful features, such as:
• Study-to-assignment ratio = (Study hours ÷ Assignments submitted)
• Assignment Submission % = (Number of Assignment Submitted/ Total Number of
Assignments) X 100
• Consistency Score = (Attendance % + Assignment Submission %) / 2
• Engagement Score = (Weighted combination of attendance, study hours, and
assignments)
6. What are the two phases of model evaluation, and how do they ensure the model's
performance and accuracy?
Ans. The two phases of model evaluation are:
a. Diagnostic Measurement Phase – This phase assesses whether the model is performing
as expected.
• For predictive models, such as decision trees, evaluation involves checking if the
model's predictions align with the expected outcomes, helping to identify areas for
improvement.
• For descriptive models, known test cases are used to verify whether the model accurately
identifies relationships within the data. If discrepancies are found, the model is refined
accordingly.
b. Statistical Significance Test – This phase ensures that the model correctly processes and
interprets data. It helps eliminate errors by verifying that the results are based on accurate
calculations, reducing the risk of incorrect assumptions.
By following these evaluation steps, data scientists can enhance model accuracy and
reliability, ensuring that the model produces dependable results before being implemented
in real-world applications.
7. Why is model validation important in Machine Learning?
Ans. Model validation is essential to ensure that a Machine Learning model performs well
on real-world data. Without proper validation, a model may work perfectly on training
data but fail when tested on new, unseen [Link] can lead to inaccurate predictions and
unreliable results. The benefits of Model Validation include:
• Ensures Accuracy and Reliability – Helps assess how well the model generalises to new
data.
• Reduces Errors – Identifies potential issues before deployment, preventing costly
mistakes.
• Prevents Overfitting and Underfitting – Strikes a balance between model complexity and
generalisation.
• Enhances Real-World Applicability – Ensures the model is effective for practical use
cases.
8. What are the three types of evaluation techniques in machine learning models?
Ans. The three types of evaluation techniques are given below:
• Overfitting happens when a model learns too much from the training data, including
noise and unnecessary details. This makes it perform very well on training data but poorly
on new data. It's like memorizing answers instead of understanding the concept.
• Underfitting happens when the model is too simple and fails to learn important patterns
from the data. This leads to poor performance on both training and test data, like not
studying enough for an exam and getting everything wrong.
• A perfect fit occurs when a Machine Learning model achieves the right balance between
overfitting and underfitting. It means the model has learned the essential patterns from the
training data without memorizing noise or unnecessary details, allowing it to generalise
well to new, unseen data.
9. Why is train-test splitting important?
Ans. Train-test splitting is an important step in Machine Learning that ensures a model is
evaluated fairly and can generalise well to new, unseen data. It prevents misleading
performance estimates by distinguishing between the data used for learning and the data
used for evaluation. Some of the reasons of why Train-Test Splitting is to be done are as
follows:
• Prevents overfitting and underfitting – A properly split dataset prevents the model from
memorizing data(overfitting) or being too simple (underfitting).
• Provides an unbiased performance estimate – Helps measure how well the model
generalises to real-world scenarios.
• Aids in model improvement – Identifies weaknesses and allows for adjustments to
improve accuracy and reliability.
• Prepares the model for real-world use – Ensures it can make good predictions outside the
training dataset.
10. What is the Mean Squared Error (MSE) in regression?
Ans. The Mean Squared Error (MSE) is the most widely-used evaluation metric in
regression. It calculates the difference between the model's predictions and actual values,
square it, and average it across the entire dataset to get the value of MSE. MSE is given by
the equation:
MSE will never be negative as the error is always squared. MSE is useful for ensuring that
our trained model does not have any outlier predictions with significant errors because
MSE places a higher weight on these errors due to the squaring element of the function.
11. Which are the steps included in the Data Science Methodology?

C. Competency-based/Application-based questions. Skills #Critical Thinking


1. An e-commerce platform uses AI to recommend products based on customer
browsing history and previous purchases. They want to assess whether the
recommendation system is accurate and beneficial to users. Which stage of Data
Science methodology is relevant?
Ans: Evaluation
2. A government transportation agency is studying the impact of a newly launched
metro rail system on daily commuter patterns. They need to assess changes in
ridership trends, travel times, and passenger volume across different routes. Which
stage of the Data Science methodology is being applied? List the steps involved.
Ans. The agency is in the Data Collection stage. The steps involved are:
• Identify Data Sources (ticket sales, commuter surveys, GPS tracking)
• Gather Data from metro stations, mobile apps, and IoT sensors
• Clean and Prepare Data (remove inconsistencies, standardize formats)
• Analyse Data for travel patterns
• Interpret and report findings
Assertion and Reasoning questions
Direction: Questions 3-4, consist of two statements – Assertion (A) and Reasoning
(R).
Answer these questions by selecting the appropriate option given below:
a. Both A and R are true, and R is the correct explanation of A.
b. Both A and R are true, but R is not the correct explanation of A.
c. A is true, but R is false.
d. A is false, but R is true.
3. Assertion (A): Data visualization helps in understanding trends, patterns, and
relationships in data.
Reasoning (R): Data visualization reduces the need for data cleaning and preparation.
Ans. c.
4. Assertion (A): Model Evaluation ensures that an AI model provides accurate
predictions before deployment.
Reasoning (R): A model with high accuracy always performs well in real-world
applications.
Ans. c.
Statement–based questions
Direction: Questions 5-6, consist of two statements – Statement 1 and Statement 2.
Examine the statements and answer the question according to the instructions given
below:
a. Statement 1 is true, Statement 2 is false
b. Statement 2 is true, Statement 1 is false
c. Both Statements 1 and 2 are true
d. Both Statements 1 and 2 are false
5. Statement 1: There is no optimal split percentage.
Statement 2: The most common split percentage between training and testing data is 20%-
80%.
Ans. a.
6. Statement 1: The RMSE of a good model should be less than 180.
Statement 2: The lower the RMSE's value, the higher the model's performance.
Ans. C
.

You might also like