0% found this document useful (0 votes)
3 views30 pages

Introduction to Data Science Basics

Uploaded by

sharan
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views30 pages

Introduction to Data Science Basics

Uploaded by

sharan
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd

INTRODUCTION TO DATA SCIENCE

What is data science?


• Data science is an interdisciplinary field that focuses on extracting knowledge from huge
amount of data sets .
• The field involves extracting meaningful insights from raw, structured, and unstructured
data that is processed using the scientific method, different technologies, and algorithms.
• The data science uses statistics, mathematics, computer science, machine learning and
involves data cleaning, data formatting and data visualization etc
• It is a multidisciplinary field that uses tools and techniques to manipulate the data so that
you can find something new and meaningful.
• Data science uses the most powerful hardware, programming systems, and most efficient
algorithms to solve the data related problems.
• In short, we can say that data science is all about:
– Asking the correct questions and analyzing the raw data.
– Modeling the data using various complex and efficient algorithms.
– Visualizing the data to get a better perspective.
– Understanding the data to make better decisions and finding the final result.
What is data science?
Example:

Let suppose we want to travel from station A to station B by car. Now, we


need to take some decisions such as which route will be the best route to
reach faster at the location, in which route there will be no traffic jam, and
which will be cost-effective. All these decision factors will act as input data,
and we will get an appropriate answer from these decisions, so this analysis
of data is called the data analysis, which is a part of data science.
What is datafication?
• Datafication is a process of “taking all aspects of life and turning them into
digital data.”
• It is about taking previously invisible process/activity and turning it into data,
that data can be monitored, tracked, analysed and optimised leading to new
opportunities and new challenges.
• Datafication is the transformation of social action into online quantified data,
thus allowing for real-time tracking and predictive analysis.
• Example: We create data every time we talk on the phone, SMS, tweet,
email, use Facebook, watch a video, withdraw money from an ATM, use a
credit card, or even walk past a security camera.
Datafication Examples
• We are being datafied, or rather our actions are, and when we “like” someone
or something online, we are intending to be datafied.
• When we merely browse the Web, we are unintentionally, or at least
passively, being datafied through cookies that we might or might not be
aware of.
• And when we walk around in a store, or even on the street, we are being
datafied in a completely unintentional way, via sensors, cameras, or Google
glasses.
Data Science Venn Diagram

• Data Science is not one


single domain. Several fields
together make Data Science.
The Data Science Venn
diagram will teach you about a
wide variety of skills required for
using data science.

• Also, Data Science involves many


different roles, that is, a Data
Scientist needs to perform many
tasks like assemble data, prepare
data, analyze data, prepare
models, evaluate models, predict
results.

• A data scientist is somebody who


understands the subject matter
under consideration in
mathematical terms and writes
computer code to solve problems.
Data Science Venn Diagram

• Nowadays, data is everywhere. Look around


us, websites, e-Commerce, financial transactions, IoT
devices, online trading, social Network, smart watch,
etc. Every single day, people generating data. Yan
LeCunn, a director of AI Research at Facebook, said
that there will be a day that we, as a human, cannot
process all those data anymore with our brain. That’s
why we need a help from a computer.
• This huge amount of data can no longer be managed
by a standard relational database system, it needs
something big, and more powerful. This is where the
BIG DATA technology is being invented to handle this
kind of huge data which comes in different format,
structured or unstructured. This is the good news.
Now, as we already have the technology to handle
this, then what?
Data Science Venn Diagram

• Here comes the DATA SCIENCE, a multidisciplinary field of


study with goal to address the challenges in big data.
Data science is a concept to unify statistics, data analysis,
machine learning, domain knowledge and their related
methods” in order to “understand and analyze actual
phenomena” with data.
• Data science skills composed of three big area: Computer
Science, Math and Statistics and Business/Domain
Expertise. In computer science, you need coding skill, a
little bit of hacking skill will also help to find a new data
sources. You also need a machine learning skill. In Math
and Statistics, you need analytic skill, dealing with
probability, analyzing the data, diagnose problem and
selecting a proper procedure for predicting. In
Business/Domain Expertise, you will need to understand
the problem domain to finally achieve the goal. This also
to ensure that the result will be impactful and well
implemented.
Data Science Venn Diagram

• Drew Conway is the guy who came up with the idea


of the Data Science Venn Diagram. The diagram
tells you about what skills are required for being a
Data Scientist.
• He believed that Data Science is made up of mainly
three things and represented them in the form of a
Venn Diagram indicating their individual roles.
• These basic things are:
– Math and Statistics
– Computer Programming [Hacking Skills]
– Domain Knowledge
• Data Science is in the middle of this Venn Diagram
combining all these skills. The Data Science Venn
Diagram gives the visual representation of how
these areas work together in Data Science.
Data Science Venn Diagram

1. Hacking Skills [ Computer Science ]


• Hacking requires great coding skills. Coding is important
because it helps you to gather and prepare the data
because a lot of data is unstructured or present in
unusual formats. You also require programming skills to
apply statistics to your problems, handle the database,
etc. One with hacking skills can apply
very complex algorithms by computer programming.
• As the demand for Data Science is increasing in the
market and there is a lot of competition, you need to be
a good hacker. That means you must be able to alter
data in order to achieve the best results for every
problem.
• Hacking skills will help you to work very creatively with
the data and different algorithms to derive some
innovative results.
Data Science Venn Diagram

2. Math and Statistics Knowledge


• After collecting and preparing the data, the
next step is to extract insights from it.
Mathematics is essential for data analysis.
For analyzing the data, you will require
several tools from mathematics such as
probability, algebra, etc. By using various
mathematical and statistical techniques on
your data, it helps in the problem's
diagnosis.
• Hence mathematics is so important because
it enables you to select the method to solve
your problems based on the available data.
Data Science Venn Diagram

3 Domain Expertise
• For applying Data Science, you need to have
an understanding of the right questions to ask
to collect the data and discover insights from
it.
• Domain Expertise means the knowledge of the
particular field in which you are working. It may
be business, healthcare, finance, education etc.
• You must be familiar with the goals of that
field, various methods, and limitations you will
face in that field.
• Thus, having familiarity with the field will help
you to implement Data Science to your
problems more easily and effectively.
Data Science Venn Diagram

• In the Data Science Venn Diagram, there are some


areas that include the intersection of these skills
which are Machine Learning,
Traditional Research and Danger Zone.
• Machine Learning: According to the Data Science
Venn Diagram, Machine learning involves the
knowledge of Computer programming and Math
but without any domain expertise.
• This means that you can simply enter your data
into the model without knowing anything about it,
such as what data is, what it signifies, and so on,
and it will provide some results.
• If you have some prior knowledge about machine
learning then it would be easy for you to
understand the data Science Venn Diagram.
Data Science Venn Diagram

• Traditional Research: This area represents


that you have the knowledge of domain,
mathematics, statistics but do not know
coding or programming. But this is not that big
problem, because the data that you use in
Traditional Research is highly structured. Thus
you don’t need to worry about preparing the
data because the data is ready for analysis.
• Danger Zone: Danger Zone is the
combination of coding and domain knowledge
but without Math and Statistics. When Drew
Conway proposed this Data Science Venn
Diagram, he believed that this is
the rarest case and is most unlikely to happen.
Exploratory Data Analysis (EDA)
• Data analysis is the process of inspecting, cleaning, transforming, and modeling
data to extract useful insights and conclusions from it.
• Data analysis is useful for making important business decisions in real-time
situations.
• Exploratory Data Analysis is important for any business decisions. It lets data
scientists to analyze the data before reaching any conclusion.
• Exploratory data analysis (EDA) is an approach for analyzing and summarizing
data in order to understand its main characteristics, patterns, and relationships.
• The primary goal of EDA is to identify the most important features of the data and
to visualize them that allows us to make better decisions.
• Conducting EDA can help data analysts make predictions and assumptions about
data. Often, EDA involves data visualization, including creating graphs like
histograms, scatter plots and box plots.
Terminology

Before you begin exploratory data analysis, it's important to understand a few key

terms:

• Value: A data value is a piece of information, such as a number or a date.

• Variable: A data variable is a characteristic that you can measure, such as weight or

income.

• Distribution: The distribution of a dataset is how the dataset is arranged. The


distribution of a dataset can be seen by observing its shape on a graph.

• Outlier: An outlier is a data value that is significantly different from the rest of a
dataset, which is much higher or lower.

• Data model: A dataset can be organised using a data model, which shows how values
relate to one another.
Why skipping EDA is a bad idea?
 generating inaccurate models;

 generating accurate models on the wrong data;

 choosing the wrong variables for the model;

 inefficient use of the resources, including the rebuilding of the model.

Why Is EDA Important?


Exploratory Data Analysis is essential for any business. It allows data scientists to analyze
the data before coming to any assumption. It ensures that the results produced are valid and
applicable to business outcomes and goals. Importance of using EDA for analyzing data
sets is:
 Helps in identifying errors in data sets.

 Gives a better understanding of the data set.

 Helps in detecting outliers or anomalous events.

 Helps in understanding data set variables and the relationship among them.
GOALS OF EDA

• To understand the data: EDA helps to understand the underlying structure and
characteristics of the data by analyzing its features, such as distributions, ranges,
variance and correlation between variables.
• To detect anomalies and errors: EDA enables us to detect missing values, outliers,
and other data quality issues that can affect the accuracy of analysis. These
anomalies can be removed or corrected before proceeding with further analysis.
• To identify patterns and relationships: EDA seeks to identify patterns, trends, and
relationships among variables in the data set.
….Conti
• To formulate hypotheses and generate ideas: EDA can be used to test
hypotheses about the data, such as whether there are significant differences
between groups or whether certain variables are related to others.
• To prepare the data for modeling: EDA can help to identify issues with the
data, such as missing values or inconsistencies, that need to be addressed before
the data can be used for modeling.
• To communicate results: EDA enables the communication of results in a clear
and concise manner. Visualizations, such as histograms, scatter plots, and heat
maps, are often used to communicate findings to stakeholders who may not
have a technical background.
How to conduct exploratory data
analysis?
It can be easier to conduct exploratory data analysis if you break the process down
into steps. Here are six key steps that you can follow to conduct EDA:

1. Data Collection
Nowadays, data is generated in huge volumes and in various forms belonging to
every sector of human life, like healthcare, sports, manufacturing, tourism, and so
on.
Every business knows the importance of using data by properly analyzing it.
However, this depends on collecting the required data from various sources
through surveys, social media, and customer reviews.
Without collecting sufficient and relevant data, further activities cannot begin.
How to conduct exploratory data analysis?
2. Finding all Variables and Understanding Them:
When the analysis process starts, the first focus is on the available data that gives
a lot of information.
It requires first identifying the important variables which affect the outcome and
their possible impact.

3. Cleaning the Dataset


Clean the data by removing any missing or inconsistent values, outliers, or
duplicated data points. The data contains only relevant information
This will not only reduce time but also reduces the computational power from an
estimation point of view.
How to conduct exploratory data analysis?
4. Categorize your values:

After identifying any missing values, you can categorize your data to determine
which statistical and visualization approaches will work with your dataset.

You can place your values into these categories:

Categorical: Categorical variables can have a set number of values.

Continuous: Continuous variables can have an infinite number of values.

Discrete: Discrete variables can have a set number of values that must be

numeric.
How to conduct exploratory data analysis?
5. Identify Correlated Variables
 Conducting EDA can also help you in identifying the relationships between the variables
in your dataset.
Finding a correlation between variables helps to know how a particular variable is related
to another.
The correlation matrix method gives a clear picture of how different variables correlate,
which further helps in understanding vital relationships among them.

6. Statistical analysis:
 Different statistical tools are employed depending on whether the data is categorical or
numerical, the size, type of variables, and the purpose of analysis.
Using descriptive statistics such as mean, median, mode, standard deviation, and
correlation coefficients to summarize the data and identify any relationships or
dependencies.
How to conduct exploratory data analysis?
7. Data visualization:
Using graphs, charts, and other visual representations to examine the data and
identify any patterns or trends. Some common types of visualizations used in EDA
include histograms, scatter plots, box plots, and heat maps.
8. Interpretation:
Analyze the results of your EDA and draw conclusions about the data. Once the
analysis is over, the results are to be observed cautiously and carefully so that
proper interpretation can be made. Identify any limitations or assumptions in your
analysis.
The data analyst should be able to analyze and be well-versed in all analysis
techniques. The results obtained will be relevant to the data in that domain and can
be used in retail, healthcare, agriculture etc.
Types of Exploratory Data Analysis
Exploratory Data Analysis is majorly performed using the following methods:

1. Univariate
2. Bivariate
3. Multivariate

• In univariate analysis, the output is a single variable and all data collected is for it.
There is no cause-and-effect relationship at all. For example, data shows products
produced each month for a year
• .
• In bivariate analysis, the outcome is dependent on two variables, e.g., the age of an
employee, while the relation with it is compared with two variables, i.e., his salary
earned and expenses per month.

• In multivariate analysis, the outcome is more than two, e.g., type of product and
quantity sold against the product price, advertising expenses, and discounts offered. The
analysis of data is done on variables that can be numerical or categorical. The result of
the analysis can be represented in numerical values, visualization, or graphical form.
Accordingly, they could be further classified as non-graphical or graphical.
Exploratory Data Analysis Tools
1. Python
Python is used for different tasks in EDA, such as finding missing values in data collection, data
description, handling outliers, obtaining insights through charts, etc. The syntax for EDA libraries like
Matplotlib, Pandas, Seaborn, NumPy, Altair, and more in Python is fairly simple and easy to use for
beginners. You can find many open-source packages in Python, such as D-Tale, AutoViz,
PandasProfiling, etc., that can automate the entire exploratory data analysis process and save time.
2. R
R programming language is a regularly used option to make statistical observations and analyze data,
i.e., perform detailed EDA by data scientists and statisticians. Like Python, R is also an open-source
programming language suitable for statistical computing and graphics. Apart from the commonly used
libraries like ggplot, Leaflet, and Lattice, there are several powerful R libraries for automated EDA,
such as Data Explorer, SmartEDA, GGally, etc.
3. MATLAB
MATLAB is a well-known commercial tool among engineers since it has a very strong mathematical
calculation ability. Due to this, it is possible to use MATLAB for EDA but it requires some basic
knowledge of the MATLAB programming language
Life Cycle of Data Science
Life Cycle of Data Science

The life cycle of data science can be summarized in several key stages:
• Problem Definition: In this phase, the problem statement is defined and the objectives are set. It involves understanding the
business problem, identifying the data sources needed to solve the problem, and determining the success criteria.
• So, with this many question arises –
 Why we need the data?
 What kind of data is required?
 How to get the data?
 What to do with the data?
• Success of any project depends on the quality of questions asked for the dataset.
• Data Collection: In this phase, the necessary data is gathered from various sources such as databases, APIs, or web scraping.
• Data Preparation: Once the data is collected, it needs to be cleaned and preprocessed. This phase involves data cleaning,
data integration, data transformation. Data may be or may not be in required format. To perform any analytical step on the
data it needs to be in certain format. Data cleaning involves removing duplicate data, identifying missing values, outliers etc.
Life Cycle of Data Science

• Data Exploration: In this phase, the data is explored and analyzed to gain
insights into the problem. Data is explored to identify patterns, trends, and
relationships among variables. This involves statistical analysis and data
visualization
• Data Modeling: In this phase, models are built using machine learning
techniques, to solve the problem. Data models are built to predict future outcomes
or to identify patterns in the data.
• Model Evaluation: In this phase, the accuracy and effectiveness of the model are
evaluated. This involves testing the model on new data and measuring its
performance against a set of predefined metrics.
Life Cycle of Data Science

• Model Deployment: Once the models are built, they need to be deployed in a
production environment. This involves integrating the models with the business
processes and creating a pipeline for real-time data processing.
• Model Monitoring and Maintenance: After the model is deployed, it needs to
be monitored and maintained to ensure that it continues to perform well. This
involves monitoring model performance, detecting and addressing model issues,
and updating the models as needed.

Overall, the life cycle of data science is an iterative process, where each phase builds
upon the previous one, and the results of each phase inform the next.

You might also like