DATA SCIENCE PROCESS: OVERVIEW
i) The data science process is a systematic approach to solving
adata problem.
ii) It defines a structured framework for articulating the problem
as a question,deciding on how to solve it and then presenting
the solution to stakeholders in an easy to understandable way.
Data Science Process
1) Setting the research goal
2) Retrieving data
3) Data preparation
4) Data exploration
5) Data modelling
6) Presentation and automation
1 Setting the rasana
1) Setting the research goal
i) Understanding the business for which the data science
project is intended to provide solution is very much
important to ensure its success.
ii) It is the first phase of data analytics project.
ii) Project charter with information such as what we are
going to rescarch, how the company benefits from that
what data and resources will be needed, a timetable, and
deliverables need to be prepared.
2) Retrieving data
i) The second step is to collect data required for the project.
ii) Data required and the source of data is already defined in
the project charter.
iii) Data from different sources are collected and
verified for
its ability, qualityand accessibility
2) Retrieving data
The second step is to collect data required for the project.
ii) Datarequired and the source of data is already defined in
the project charter.
iii) Data from different sources are collected and verified for
its ability, quality and accessibility.
iv) Data can also be delivered by third-party companies and
takes many forms ranging from Excel spreadsheets to
different types of databases.
3) Data preparation
i) Data collection is anerror-prone process as they are
collected from dillerent sources including scnsor data,
manual entry data, OCRdata.
ii) In this phase of the data science project, quality of the
data needs to be enhancedand prepare it to use in
subsequent steps.
iii) Accuracy of thedata science models highly depends on
the quality of the data used for training the models.
Three sub-phases of Data preparation
A. Data Cleansing
Remove false values rom a data SOurce and
inconsistencies across data sources.
B. Data Integra 2of 3 667 Words
Three sub-phases of Data preparation
A. Data Cleansing
Remove false values
valucs from a data source and
inconsistencies across data sources.
B. Data Integration
Enriches data sources by combining information from
multiple data sources.
C. Data transformation
Ensures that the data is in a suitable format for use in
the
models.
4) Data Exploration
i) Data exploration is concermed with the data
using data visualization and analysis
4) Data Exploration
i) Data exploration is concerned with the data analysis
using data visualization and statistical techniques to
understand the data that need to be processed.
ii) It can be describcd using dataset characterizations, such
as size, quantity, and accuracy, in order to better
understand the nature of the data.
iii) Automated data exploration software can be used to
visually explore and identify relationships between
different data variables, the structure of the dataset, the
presence of outlierS, and the distribution of data values in
order to reveal patterns and points of interest, enabling
data analysts to gain insight into the raw data.
iv) Exploratory Data Analysis (EDA) is done through
descriptive statistics, visual techniques, and simple
modelling.
5) Data modelling
i) Machine lcarning models, domain knowledge, and
insights about the data found in the previous steps are
sedto find solution for the research question.
ii) Techniques from the fields of statistics, machine learning,
operations rescarch. etc. are utilized.
ii) Building a model is an iterative process that involves
selecting the variables to build the model, executing and
training the model, and model diagnostics.
6) Presentation and Automnation
i) Final step is to present the results to business and
stakeholders.
ii) These results can take many forms ranging from
presentations to rescarch reports.
ii) Visualization helps to explore and communicate the
findings in an casy way to understand the voluminous
data.
6) Presentation and Automation
i) Final step is topresent theresults to business and
stakeholders.
ii) These results can take many forms ranging from
presentations to rescarch reports.
iii) Visualization helps to explore and communicate the
findings in an casy way to understand the voluminous
data.
iv) Sometimes we nced to automate the execution of the
process because the business will use the gaincd insights
in another project or enable an operational process to use
the outcome from the built model.
V) Data science process has to go back and rework certain
findings across the phases.
vi) For instance, outliers in the data exploration phase that
point to data import errors.
vii) As part of the data science process we gain incremental
insights, which may lead to new questions.
viii) To prevent rework, scope of the business
be defined clearlv and thawa question must