Data Science Methodology : An analytic
approach to Capstone Project
1. What is data science methodology?
It provides a structured way to collect, process and understand data.
2. Name different modules of Data Science methodology.
There are five modules of Data Science methodology
i. From Problem to Approach
ii. From Requirements to Collection
iii. From Understanding to Preparation
iv. From Modelling to Evaluation
v. From Deployment to Feedback
3. Explain five modules of Data Science methodology.
i. From Problem to Approach:
First, we find out what problem we want to solve using AI.
We try to understand where the problem comes from (like school, shop, or hospital).
We set clear goals and focus on one problem at a time.
We choose a method like machine learning or deep learning to solve it.
ii. From Requirements to Collection:
What data is needed to solve the problem.
We check what type of data we need, how much, and where to get it.
We can get data from inside (like company files) or outside (like websites).
Good data is important to make the AI work well.
We also write down how and where we got the data.
iii. From Understanding to Preparation
We look at the data to understand it.
We fix any mistakes, remove extra parts, and fill in missing things.
We pick useful parts of the data to help the model learn better.
We change the data into the right form for training.
Now the data is ready to be used.
This step ensures that the dataset is ready for modeling.
iv. From Modelling to Evaluation
We choose the right AI method to build the model.
We split the data into two parts: one for learning, one for testing.
We train the model and check how good it is using scores like accuracy.
We try different models to see which one works best.
We make small changes to improve the model.
v. From Deployment to Feedback
We use the model in real life.
We watch how it works and if it does a good job.
People give feedback and we check the results.
If needed, we fix or train the model again to make it better.
This keeps the model helpful for a long time.
4. Explain Descriptive Analytics method?
It method is used to identify trends and patterns on the basis of past data. It is the process of
understanding that has happened.
5. Explain different methods of Descriptive Analytics.
Data Aggregation: Summarizing data from multiple sources.
Data Mining: Discovering patterns in large datasets.
Statistical Measures: Mean, median, mode, standard deviation.
Data Visualization: Using charts, graphs, dashboards to present data trends.
Reporting Tools: Pre-built reports and summaries.
Example: Monthly sales reports showing revenue, region-wise performance, etc.
6. Explain Diagnostic Analytics method?
It is the process of understanding why things have happened. This method is used to analyze
the reason behind the patterns of data.
7. Explain different methods of Diagnostic Analytics.
Drill-Down Analysis: Looking at data more closely, step by step, to see small details.
Correlation Analysis: Finding out if two things are connected or move together.
Root Cause Analysis: Finding the real reason why a problem happened.
Data Discovery Tools: Easy-to-use tools that help you look at and understand data.
Hypothesis Testing: Checking if an idea or guess is true by using math and data.
Example: Analyzing customer feedback to understand why complaints increased.
8. Explain Predictive Analytics method?
It is the process of understanding what will happen in the future on the basis of past data.
9. Explain different methods of Predictive Analytics.
Regression Analysis: Predicting future on the basis of historical trends of data.
Machine Learning Models: Decision trees, neural networks, support vector machines.
Time Series Analysis: Forecasting future values using time-based data.
Classification: Categorizing data into groups
Clustering: Grouping similar data together.
Example: Predicting customer return based on usage patterns.
10. Explain Prescriptive Analytics method?
It is the process to decide what should be done after making predictions
11. Explain different methods of Prescriptive Analytics.
Optimization Algorithms: Linear programming, genetic algorithms.
Simulation: Modeling scenarios to evaluate different decisions.
Decision Analysis: Tools to evaluate possible outcomes of decisions.
Recommendation Engines: Suggesting actions using AI and ML.
Example: Suggesting the best delivery routes to minimize cost and time.
12. What do you mean by understanding data requirements?
It is the process of clearly defining what kind of data is needed to solve a problem.
13. Explain steps in understanding data requirements:
i. Identifying data types: Quantitative (numbers), textual (words), and visual (images).
ii. Choosing data structure: Formats like tables, text files, or databases.
iii. Finding data sources: From company records, surveys, websites, sensors, or public
datasets.
iv. Preparing data: Cleaning and organizing to remove errors or duplicates.
14. Explain different types of data.
Structured Data: It is organized in fixed format e.g. databases.
Semi-Structured Data: It is not fully organized, e.g., emails, XML.
Unstructured Data: No fixed format, e.g., social media posts, videos.
15. What is 5W1H?
5W1H stands for: What, Why, When, Where, Who and How
i. What?
Ask: What happened?
What is the problem or task?
Example: What is the issue we are trying to solve?
ii. Why?
Ask: Why did it happen?
Why is this important?
Example: Why is it a problem?
iii. When?
Ask: When did it happen?
When should something be done?
Example: When did the issue begin?
iv. Where?
Ask: Where did it happen?
Where does it affect?
Example: Where did the mistake take place?
v. Who?
Ask: Who is involved?
Who is affected or responsible?
Example: Who is part of the problem or solution?
vi. How?
Ask: How did it happen?
How can it be fixed or improved?
Example: How can we solve it?
16. Explain 5W+1H framework for a grocery store that wants to avoid
stock shortages by predicting demand.
Question Answer
What We want to predict the demand for grocery items.
Why To avoid running out of stock and losing customers.
When Especially during weekends, holidays, and sale days.
Where In all store locations, especially busy ones.
Who Store managers, supply chain team, and data analysts.
How By studying past sales data, seasonal trends, and using AI or data analysis tools.
17. Explain 5W +1H framework for Delayed Delivery of Online Orders.
Question Answer
What Orders from customers are being delivered late.
Why The delivery service is facing delays due to traffic and driver shortage.
When This issue started last week and occurs mainly on weekends.
Where The deliveries are being delayed in the city areas, especially in the evening.
Who The delivery team, logistics department, and customers are affected.
The delivery service is not properly scheduled, and there is a lack of enough drivers
How
for peak hours.
18. Explain 5W +1H framework for Low Employee Productivity.
Question Answer
What Employees are not meeting productivity goals.
Employees are feeling demotivated due to unclear work expectations and lack of
Why
feedback.
When The problem has been noticed over the past month.
Where It’s happening in the marketing and customer service departments.
Who Employees, team managers, and HR department are involved.
There is no regular check-in with team leaders, and there’s little recognition of
How
achievements.
19. Explain 5W+1H framework for Decreasing Sales in a Clothing Store
Question Answer
What Sales at the clothing store are decreasing.
Why Customers are not finding the latest fashion trends, and prices have increased.
When The sales began to drop over the last two months.
Where The store’s sales are particularly low in the downtown location.
Who The store manager, marketing team, and customers are involved.
The store has not updated its inventory to reflect new trends and customers are also
How
hesitant due to the price hike.
20 What challenges might arise while defining data requirements?
Incomplete Data – Some important information may not be available.
Data Privacy Issues – Collecting data of customers as per privacy laws.
Data Overload – Too much data may slow down processing and make
analysis difficult.
21. What do you mean by Data collection
It is the process of gathering relevant data needed to train, test, and validate an AI model.
22. What do you mean by Primary Data?
Primary data refers to the data collected directly from original sources.
23. What are sources of Primary Data?
Surveys and Questionnaires: Used to gather opinions or feedback.
Interviews and Focus Groups: Useful for in-depth understanding of user needs.
Sensors and IoT Devices: For real-time data like temperature, motion, etc.
Manual Observations: For analyzing behavior or activities.
Mobile Apps or Websites: Collecting user interaction data.
24. What do you mean by Secondary Data?
Secondary data refers to data that is already available, collected by others for different
purposes.
25. What are sources of Secondary Data?
Books, Research papers and journals
Websites and Online databases
Company databases and records
News articles, blogs, social media data
26. What do you mean by Mixed Data?
Mixed data sources combine both primary and secondary data to create a more
comprehensive and reliable dataset.
27. What are sources of Mixed Data?
Online Marketing Analytics : Data is collected from user surveys, experiments and existing
reports.
Online Tracking: Data is collected from first hand analytic tools and Google trends.
Social Media Monitoring: Data is collected from social media platforms.
Online Tracking (Primary): Data collected directly by you, like website visits from your
own tools.
Online Tracking (Secondary): Data taken from outside sources, like Google Trends.
Combination Research: Uses both self-collected data and existing information for better
insights.
Social Media Monitoring (Primary): Directly tracking user activity on your own social
media pages.
Social Media Monitoring (Secondary): Using reports or data summaries from external
sources.
28. What is the role of data collection in a project?
Collaboration: Data scientists, DBAs, and programmers work together to collect data
from both direct (primary) and outside (secondary) sources.
Iteration: If something is missing or unclear, the team may collect more data or make
changes.
Accuracy & Reliability: Good decisions in areas like business, health, education, and
research depend on having correct and complete data.
29. For the clothing showroom inventory management problem, analyze
the following questions:
a) What are the possible sources for collecting this data?
a. Primary Data Sources (Showroom sales records, Customer feedback or surveys)
b. Secondary Data Sources (Sales data of other clothing stores, fashion industry databases)
Ans. Both a. and b.
b). How can the collected data be categorised for better organization?
Ans. By categorising data correctly as structured data, semi-structured data and unstructured
data. This makes it easier to analyse and use for decision-making.
c). What challenges might arise during the data collection process?
Ans. Challenges during the data collection process are as follows:
Missing Data – Some key details (like customer style preferences or size demands)
may not be available.
Privacy Issues – Customer data collection must comply with data protection laws.
Inconsistent Data Formats – Data may come from different sources and need
standardization.
Data Overload – Too much data may slow down processing and require additional
storage.
30. What is the meaning of Data Understanding?
Data understanding means checking the data to make sure it is correct, complete, and useful
for solving the problem. It helps us to see if the data is good enough or if we need to collect
more data.
31. Why is Understanding Data Important?
It makes sure the data matches the problem we are trying to solve.
It helps us find mistakes, missing information, or data we don’t need.
It helps us decide if we need to collect more data.
It gives us clean and useful data for better decisions.
32. What are the Different Ways to Analyze the Data?
i. Descriptive Statistics – These help us describe the data.
Mean, Median, Mode – To find the average or most common value.
Range, Variance, Standard Deviation – To see how different the values are from
each other.
Pairwise Correlation – To check if two things are connected (like: do bigger
discounts lead to more sales?).
ii. Data Visualization – Drawing charts or graphs to see patterns:
Histograms – Show how often values appear.
Pie Charts and Bar Graphs – Compare different groups.
Scatter Plots – Show how two things are related.
iii. Handling Missing Data – Deciding what to do if some data is not there.
Finding Mistakes or Duplicates – Removing wrong or repeated data.
Identifying Outliers – Finding values that are very different from others, which may show a
mistake or something special.
33. What is the Meaning of Data Visualization?
Data visualization means showing data in pictures or graphs. It helps us quickly see patterns,
trends, and relationships in the data. This makes it easier to understand and explain the data.
34. What Are the Different Techniques of Data Visualization?
Histograms – Show how often different values appear.
Pie Charts – Show how a whole is divided into parts.
Bar Graphs – Help compare different categories.
Scatter Plots – Show the relationship between two variables.
35. What Are the Different Techniques to Clean the Data?
Handling Missing Data – Fill in missing values using methods or collect more data.
Detecting Errors and Duplicates – Find and fix wrong or repeated data.
Identifying Outliers – Check for unusual values that may be mistakes or need special
attention.
Q36. How can you check if the collected data is complete, and correct?
We can check data quality by:
Identifying missing data – Checking if important details are missing.
Detecting duplicates – Removing duplicate records .
Verifying accuracy – Ensuring that data values are correct.
Cross-checking with external data sources
Q 37. Can a correlation analysis help in grocery store?
Ans. Yes, correlation analysis can help. It shows how two things are related to each other.
For example:
If more customers visit the store, do the sales go up too?
Does bad weather cause more people to buy packaged food?
Q 38. What do you mean by data preparation?
Ans. Data preparation means organizing and cleaning raw data before using it. This step
makes data easier to understand and use in data analysis or machine learning. The main steps
are:
Cleaning Data – Fixing or removing missing values, duplicate entries, and formatting
errors.
Combining Data – Joining data from different sources like tables or files.
Transforming Data – Changing raw data into useful information for analysis.
What is Feature Engineering?
Q 39. What do you mean by feature Engineering?
Feature engineering means creating or changing data to make machine learning models
more accurate.
Q 40. How can feature engineering help in building model to predict
students’ performance?
If you’re building a model to predict students’ performance , raw data may include things like
study hours, attendance, and number of assignments submitted by the students. New features
can be created such as :
Study-to-assignment ratio = Study hours ÷ Assignments submitted
Assignment submission percentage = (Assignments submitted ÷ Total assignments)
× 100
Consistency Score = (Attendance % × Assignment submission %) ÷ 2
Engagement Score = A mix of attendance, study hours, and assignment scores
Q41. Why is Data Preparation Important?
Data preparation is important because it makes sure the data is clean and ready to use. If this
step is skipped, machine learning models might give wrong results.
It takes the most time (about 80% of the total time in a project).
It’s used across industries, since many problems need similar steps.
Automation helps – Tools powered by AI can do data preparation faster by reducing
manual work.
Q42. Case Study: Beverage Sales Analysis
For the given sample dataset of beverage sales data for three months, answer the following
questions:
Tea Coffee Juice Beverage Sale of No. of school Promotion
Month Weather
Sale Sale Sale Products Tea holidays Applied
Jan 200 180 150 250 200 3 Cold Yes
Feb ? 160 ? 300 200 5 Moderate No
Mar 180 ? 170 280 ? 2 Hot 35
a. Is there any duplicate data?
Ans: Yes, the columns “Tea Sale” and “Sale of Tea” show the same information. If they are
not different, one of them should be removed to avoid confusion.
b. Is there any incorrect data?
Ans: Yes, in the “Promotion Applied” column, the value “35” (in March) is wrong. The
column should contain “Yes” or “No”.
c. Is there any missing data?
Ans: Yes, some values are missing:
“Tea Sale” for February
“Juice Sale” for February
“Coffee Sale” for March
“Sale of Tea” for March
d. Are the names of columns clear, or are any changes suggested?
Ans: Some columns could be renamed:
“Tea Sale” and “Sale of Tea” should be merged or renamed for clarity
“Beverage Products” should specify if it includes all drinks or only some
“Promotion Applied” could be renamed to “Promotion Status” with Yes/No values
e. Did you find any irrelevant data in the dataset?
Ans: The “No. of school holidays” may not be useful unless it affects drink sales. The value
“35” in “Promotion Applied” might also be wrong or not suitable for this column.
f. What steps can be done to to clean and improve the data:
Remove or rename repeated columns
Correct the “35” value in Promotion Applied
Fill in missing values using proper methods
Clarify column names for better understanding
Ensure all data helps in analysis and decision-making
Q43 What is AI Modeling stage?
In the modeling stage, we use the cleaned data to create models that help us understand the
data and find useful information.
The type of model we choose depends on how we plan to study the data.
This process may need to be repeated, with some changes to the data, to make the model
more accurate.
Data scientists often try different methods to find the best model for solving the problem.
Q44. What is Descriptive Modeling?
What it does: Looks at past data to understand what happened.
Purpose: To summarize and explain past events or behaviors.
Example: Analyzing last year’s sales data to find which month had the highest sales.
Key Question: What happened? or Why did it happen?
Q44. What is Predictive Modeling
What it does: Uses past data to guess what might happen in the future.
Purpose: To make predictions and forecast future trends or outcomes.
Example: Predicting next month’s sales based on previous years’ data.
Key Question: What is likely to happen next?