0% found this document useful (0 votes)
8 views47 pages

Data Analytics

The document outlines the differences between data science and data analytics, emphasizing that data science involves preparing datasets for analysis while data analytics focuses on extracting insights from data. It details four types of data analytics: descriptive, predictive, prescriptive, and diagnostic, each serving different purposes in business decision-making. Additionally, it discusses various tools and libraries used for data analysis, including Excel, SQL, Tableau, Power BI, and machine learning frameworks like TensorFlow and PyTorch.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
8 views47 pages

Data Analytics

The document outlines the differences between data science and data analytics, emphasizing that data science involves preparing datasets for analysis while data analytics focuses on extracting insights from data. It details four types of data analytics: descriptive, predictive, prescriptive, and diagnostic, each serving different purposes in business decision-making. Additionally, it discusses various tools and libraries used for data analysis, including Excel, SQL, Tableau, Power BI, and machine learning frameworks like TensorFlow and PyTorch.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Data Analytics

Data science Vs Data analytics


• Data science is the process of building, cleaning, and structuring
datasets to analyze and extract meaning.
• Data analytics, on the other hand, refers to the process and practice
of analyzing data to answer questions, extract insights, and identify
trends.
• You can think of data science as a precursor to data analysis. If your
dataset isn’t structured, cleaned, and wrangled, how will you be able
to draw accurate, insightful conclusions? Below is a deeper dive into
each field’s role in business.
What is Data Analytics?
• In a technical sense, data analytics can be described as the process of
using data to answer questions, identify trends, and extract insights.
• There are multiple types of analytics that can generate information to
drive innovation, improve efficiency, and mitigate risk.
• There are four key types of data analytics, and each answers a
different type of question:
• Descriptive analytics asks, “What happened?”
• Predictive analytics asks, “What might happen in the future?”
• Prescriptive analytics asks, “What should be done next?”
• Diagnostic analytics asks, “Why did this happen?”
Value Proposition of Each type of Data Analytics
Descriptive Analytics
• Descriptive analytics primarily uses observed data to identify key
characteristics of a data set. It relies solely on historical data to
provide reports on past events.
• This type of analysis is also used to generate ad hoc (as needed)
reports that summarize large amounts of data to answer simple
questions like “how much?” or “how many?” It can also be used to
ask deeper questions about a specific problem.
• Descriptive analytics is not used to draw inferences or predictions
from its findings; it is just a starting point used to inform decisions or
to prepare data for further analysis.
Descriptive Analytics Cont.
• The descriptive analytics process is as follows:
• Ask a historical question that needs an answer, such as “How much of product
X did we sell last year?”
• Identify required data to answer the question
• Collect and prepare data
• Analyze data
• Present results
• It’s sometimes called the simplest form of data analysis because it
describes trends and relationships but doesn’t dig deeper.
Examples of Descriptive Analytics use-cases
• Traffic and Engagement Reports: This involves tracking engagement
in the form of social media likes, comments tec.
• For example, you may be responsible for reporting on which media
channels drive the most traffic to the product page of your company’s
website.
• Using descriptive analytics, you can analyze the page’s traffic data to
determine the number of users from each source.
• You may decide to take it one step further and compare traffic source
data to historical data from the same sources. This can enable you to
update your team on movement; for instance, highlighting that traffic
from paid advertisements increased 20 percent year over year.
Examples of Descriptive Analytics use-cases
• Financial Statement Analysis: There are several types of financial
statements, including the balance sheet, income statement, cash flow
statement, and statement of shareholders’ equity. Each caters to a specific
audience and conveys different information about a company’s finances.
• Financial Statement analysis can be done in three ways:
• Vertical analysis: involves reading a statement from top to bottom and comparing
each item to those above and below it. This helps determine relationships between
variables. For instance, if each line item is a percentage of the total, comparing them
can provide insight into which are taking up larger and smaller percentages of the
whole.
• Horizontal analysis: involves reading a statement from left to right and comparing
each item to itself from a previous period. This type of analysis determines change
over time.
• Ratio analysis: involves comparing one section of a report to another based on their
relationships to the whole. This directly compares items across periods, as well as
your company’s ratios to the industry’s to gauge whether yours is over- or
underperforming.
Other Examples of Descriptive Analytics use-
cases
• Summarizing historical events such as sales, inventory, or operations
data
• Reporting general trends like revenue growth, customer demands, or
employee injuries
• Collating survey results
• Reporting country’s trends like population growth, economic growth
etc.
• Reporting on progress toward key performance indicators (KPIs)
Predictive Analytics
• Predictive analytics utilizes real-time and/or past data to make
predictions based on probabilities. It can also be used to infer missing
data or establish a predicted future trend.
• Predictive analytics uses simulation models and forecasting to suggest
what could happen going forward, which can guide realistic goal
setting, effective planning, management of performance
expectations, and avoiding risks.
• This information can empower executives and managers to take a
proactive and fact-based approach to strategy and decision making.
Predictive Analytics Cont.
• The predictive analytics process is as follows:
• Ask a forward-thinking question, such as “Can we predict how much product
X we will sell next year?”
• Collect and prepare data
• Develop predictive analytics models
• Apply models to the prepared data
• Review models and present results
Examples of predictive analytics include:
• Forecasting customer behavior, purchasing patterns, and identifying
sales trends
• Predicting customer preferences and recommending products to
customers based on past purchases and search history
• Predicting the likelihood that a given customer will purchase another
product or leave the store
• Identifying possible security breaches that require further
investigation
• Predicting staffing and resourcing needs
Prescriptive Analytics
• Prescriptive analytics builds on descriptive and predictive analysis by
recommending courses of action that will reap the greatest benefit
for the organization.
• In short, prescriptive analytics tells you what should be done in a
given situation. It helps executives, managers, and employees make
the best decisions based on available data.
• A good example of prescriptive analytics is the field of GPS-based
map and direction applications. These applications provide route
options to a destination based on traffic volume, road conditions, and
maximum speed. It can then prescribe the best route based on user-
defined objectives such as shortest distance or quickest time.
Diagnostic Analytics
• Diagnostic analytics enhances the descriptive analytics process by
digging in deeper and attempting to discover the cause(s).
• The diagnostic analytics process is as follows:
• Identify anomalies (inconsistencies) in data sets
• Collect data related to the anomalies
• Use statistical techniques to uncover relationships and trends that could
explain the anomalies
• Present possible causes
Example of Diagnostic Analytics
• Using subscription cancellations, correlated with customer comments
and ratings, to determine the most common reasons why users
cancel subscriptions.
• Another example would be determining whether there is a
correlation between the demographics of consumers and their
purchasing patterns at specific times of year.
Tools for Data Analysis
Tools for Data Analysis
• Excel
• Tablue
• Power BI
• Frameworks
• Libraries
Excel
• Excel is a powerful tool suitable for small datasets and quick data
analysis. With Excel, you can manipulate data, summarize it with pivot
tables, visualize it, and perform quick statistics to summarize it.
• Why it’s important to know:
• Excel is powerful and very popular for performing small-scale data
analysis, calculations, data summaries, and data visualizations.
Critical Excel Skills in Data Analysis
• Perform data cleaning by removing blank spaces as well as incorrect
and outdated information
• Format and adjust data using conditional formatting
• Perform data calculations using formulas
• Organize data using sorting and filtering
• Create visualizations using graphing and charting
• Calculate, summarize, and analyze data using pivot tables
• Aggregate data for analysis
Structured Query Language (SQL)
• SQL, which stands for Structured Query Language, is a powerful
database management tool that allows data analysts to retrieve and
interact with selections of data that are stored in relational databases.
• Relational databases have a defined structure and contain multiple
interrelated data tables that need to be queried with a language like
SQL to be useful.
• SQL is fast and can handle data sets much larger than Excel can. As a
data analyst you will use SQL to access, read, manipulate, and analyze
the data stored in a relational database to generate useful insights to
drive a data-informed decision-making process.
SQL Cont.
• Why it’s important to know:
• Popular big data systems make use of SQL for maintaining relational
databases and processing structured data.
• It is used for carrying out data analytics with data stored in relational
database management systems such as Oracle, Microsoft SQL, and
MySQL.
• Critical SQL Skills for Data Anlysis
• Creating tables
• Retrieving data using SQL index
• Retrieving data using SQL queries
• Aggregating data with SQL joins
Tableau
• Tableau is one of the most used data analytics and visualization tools on
the market. Visualizations are an important way to present data in a format
that can easily be understood by non-technical decision makers and
stakeholders.
• Why it’s important to know:
• Tableau is a data analytics market leader due to the depth and quality of its data
visualizations.
• Tableau can extract and combine data from multiple sources including Excel
spreadsheets and SQL databases. It can also access large data storage locations,
known as data warehouses, as well as cloud-based data repositories.
• Critical Tableau Skills
• Comparing data from multiple views using Tableau dashboards
• Creating visualizations using Tableau visualization tools
Power BI
• Power BI is a business analytics service provided by Microsoft. It
offers tools for aggregating, analyzing, visualizing, and sharing data.
• Why it is important
• Power BI is a powerful business analytics tool provided by Microsoft, essential
for aggregating, analyzing, visualizing, and sharing data. It allows users to
create interactive and shareable dashboards, integrate various data sources,
and perform real-time data analysis.
• Its importance lies in its ability to provide detailed insights, support data-
driven decision-making, and facilitate collaboration within organizations.
• Power BI’s robust features, including AI capabilities, mobile access, and
embedded analytics, make it a critical tool for modern businesses aiming to
harness their data effectively.
Frameworks and Libraries
Sckitlearn
• Scikit-learn is an open-source machine learning library for the Python
programming language. It is built on top of other popular Python
libraries such as NumPy, SciPy, and matplotlib.
• Scikit-learn is a community project, developed by a large group of
people, all across the world.
• Simple and efficient tools for predictive data analysis
• Accessible to everybody, and reusable in various contexts
• Built on NumPy, SciPy, and matplotlib
• Open source, commercially usable - BSD license
Sckitlearn Cont.
• Simple and Efficient Tools: Scikit-learn provides simple and efficient
tools for data mining and data analysis, making it accessible to both
beginners and experienced practitioners.
• Integration with Python Ecosystem: Scikit-learn integrates well with
other popular Python libraries such as pandas (for data
manipulation), NumPy (for numerical computations), and matplotlib
(for plotting).
• Performance: While scikit-learn is not optimized for large-scale deep
learning (like TensorFlow or PyTorch), it is highly efficient for a wide
range of classical machine learning tasks and works well for
moderately sized datasets.
Sckitlearn Cont.
• Wide Range of Algorithms: It includes a wide range of supervised and
unsupervised learning algorithms, such as:
• Classification (e.g., support vector machines, k-nearest neighbors, logistic
regression)
• Regression (e.g., linear regression, ridge regression, lasso regression)
• Clustering (e.g., k-means, hierarchical clustering, DBSCAN)
• Dimensionality Reduction (e.g., principal component analysis, singular value
decomposition)
• Model Selection (e.g., grid search, cross-validation)
• Preprocessing (e.g., standardization, normalization, encoding categorical
features)
• [Link]
TensorFlow
• TensorFlow is an open-source machine learning framework developed by
the Google Brain team.
• It is widely used for building and deploying machine learning models,
particularly deep learning models.
• Key features and components of TensorFlow
• High-Level APIs:
• Keras: A high-level API for building and training neural networks, integrated within
TensorFlow for ease of use.
• Estimators: Simplifies the model-building process for tasks like training, evaluation,
and prediction.
• TensorFlow allows for both high-level abstractions and low-level
operations, giving users the flexibility to build complex models from scratch
if needed.
TensorFlow Cont.
• Flexibility: TensorFlow allows for building and deploying machine learning
models across different platforms, such as desktops, servers, mobile
devices, and edge devices. It supports both high-level APIs, like Keras, for
quick model building, and low-level APIs for more complex model
customization.
• Scalability: TensorFlow can handle large-scale machine learning tasks,
making it suitable for both small experiments and large production
deployments. It supports distributed computing, enabling training on
multiple GPUs or across multiple machines.
• Community and Support: TensorFlow has a large and active community,
which means there is a wealth of tutorials, documentation, and third-party
resources available. The community also contributes to its ongoing
development and improvement.
TensorFlow Cont.
• Applications: TensorFlow is used in a wide range of applications, from
image and speech recognition to natural language processing and
reinforcement learning. It supports various types of neural networks,
including convolutional neural networks (CNNs), recurrent neural
networks (RNNs), and transformers.
• Ecosystem: TensorFlow has a rich ecosystem of tools and libraries,
including TensorFlow Extended (TFX) for end-to-end machine learning
pipelines, TensorFlow Lite for deploying models on mobile and IoT
devices, and [Link] for running models in the browser.
Pytorch
• PyTorch is an open-source machine learning library developed by
Facebook's AI Research lab (FAIR).
• It is widely used for deep learning applications and provides flexibility and
speed in building and experimenting with neural network models.
• Ecosystem: PyTorch has a rich ecosystem of tools and libraries, including:
• TorchVision: For computer vision tasks.
• TorchText: For natural language processing.
• TorchAudio: For audio processing.
• PyTorch Lightning: For high-level, structured PyTorch code.
• fastai: For simplifying complex deep learning tasks.
• [Link]
Pandas:
• pandas is a fast, powerful, flexible and easy to use open source data analysis and
manipulation tool, built on top of the Python programming language.
• A fast and efficient DataFrame object for data manipulation with integrated
indexing;
• Tools for reading and writing data between in-memory data structures and
different formats: CSV and text files, Microsoft Excel, SQL databases, and the fast
HDF5 format;
• Intelligent data alignment and integrated handling of missing data: gain
automatic label-based alignment in computations and easily manipulate messy
data into an orderly form;
• Flexible reshaping and pivoting of data sets;
• Intelligent label-based slicing, fancy indexing, and subsetting of large data sets;
• Columns can be inserted and deleted from data structures for size mutability;
Pandas Cont.
• Aggregating or transforming data with a powerful group by engine allowing
split-apply-combine operations on data sets;
• High performance merging and joining of data sets;
• Hierarchical axis indexing provides an intuitive way of working with high-
dimensional data in a lower-dimensional data structure;
• Time series-functionality: date range generation and frequency conversion,
moving window statistics, date shifting and lagging. Even create domain-
specific time offsets and join time series without losing data;
• Python with pandas is in use in a wide variety of academic and
commercial domains, including Finance, Neuroscience, Economics,
Statistics, Advertising, Web Analytics, and more.
• [Link]
Numpy
• NumPy is the fundamental package for scientific computing in
Python.
• It is a Python library that provides a multidimensional array object,
various derived objects (such as masked arrays and matrices), and an
assortment of routines for fast operations on arrays, including
mathematical, logical, shape manipulation, sorting, selecting, I/O,
discrete Fourier transforms, basic linear algebra, basic statistical
operations, random simulation and much more.
• Learn more about Numpy here: [Link]
Matplotlib
• Matplotlib is a comprehensive library for creating static, animated,
and interactive visualizations in Python. Matplotlib makes easy things
easy and hard things possible.
• It can be used to.
• Create publication quality plots.
• Make interactive figures that can zoom, pan, update.
• Customize visual style and layout.
• Export to many file formats.
• Embed in JupyterLab and Graphical User Interfaces.
• Use a rich array of third-party packages built on Matplotlib.
• [Link]
Seaborn
• Seaborn is a library for making statistical graphics in Python. It builds on
top of matplotlib and integrates closely with pandas data structures.
• Seaborn helps you explore and understand your data. Its plotting functions
operate on dataframes and arrays containing whole datasets and internally
perform the necessary semantic mapping and statistical aggregation to
produce informative plots.
• Its dataset-oriented, declarative API lets you focus on what the different
elements of your plots mean, rather than on the details of how to draw
them.
• There is no universally best way to visualize data. Different questions are
best answered by different plots. Seaborn makes it easy to switch between
different visual representations by using a consistent dataset-oriented API.
• [Link]
Beautiful Soup
• Beautiful Soup is a Python library for pulling data out of HTML and
XML files. It works with your favorite parser to provide idiomatic ways
of navigating, searching, and modifying the parse tree. It commonly
saves programmers hours or days of work.
• Highly used in wescarapping/ webcrawling.
• Beautiful soup documentation: [Link]
[Link]/en/latest/
• Beautiful soup practice exercise with codes and narration:
[Link]
Jupyter Notebook
• Jupyter Notebook is an open-source web application that allows you
to create and share documents that contain live code, equations,
visualizations, and narrative text. It's widely used in data science,
machine learning, academic research, and more.
• You can get Jupyter notebook in the following ways
• As part of Anaconda distribution environment (see next slides)
• Jupyter notebook website
• Cloud-based Google Colaboratory (see next slides)
• Kaggle Notebooks
Anaconda Distribution Environment
• The Anaconda distribution is a popular choice for data science and
machine learning projects. It includes a few key components that
make it useful for this field:
• Conda: This is a package and environment manager that allows you to
install specific versions of software packages. This is crucial in data
science, where different projects may require different versions of
libraries to function correctly.
• Conda creates isolated environments where you can have specific
versions of libraries installed without affecting other projects on your
system.
Anaconda Distribution Environment Cont.
• Anaconda Navigator: This is a graphical user interface (GUI) built on top of
conda. It allows you to manage environments, packages, and even launch
other development applications, all from a user-friendly interface. This is a
good option for those who prefer not to use the command line for
managing their environments.
• Pre-installed packages: The Anaconda distribution comes with over 300
pre-installed packages that are commonly used in data science and
machine learning. This saves you time from having to install each package
individually and ensures compatibility between them.
• Anaconda Repository: In addition to the pre-installed packages, Anaconda
provides access to a vast repository (over 8000 packages) of open-source
data science and machine learning packages. You can use conda to search
for and install any additional packages you may need for your project.
Google Colab
• Google Colab, or Google Colaboratory, is a cloud-based platform
provided by Google that allows users to write and execute Python
code in a Jupyter Notebook environment.
• It's particularly popular among data scientists, machine learning
practitioners, and researchers for its ease of use and powerful
features.
• Key aspects and features of Google Colab:
• Cloud-Based: Google Colab is entirely cloud-based, which means you
don't need to install any software on your local machine. You can
access and run notebooks from anywhere with an internet
connection
Google Colab Cont.
• Free GPU and TPU Access: One of the most significant advantages of
Google Colab is the availability of free access to powerful GPUs and
TPUs. This is particularly useful for training large machine learning
models that require significant computational resources.
• Integration with Google Drive: Colab is seamlessly integrated with
Google Drive, allowing you to save, access, and share your notebooks
directly from your Google Drive account. This makes collaboration
and version control easy.
• Interactive Code Execution: Like Jupyter Notebooks, Colab supports
interactive code execution. You can run code cells independently, see
the output immediately, and make changes on the fly.
Google Colab Cont.
• Pre-installed Libraries: Colab comes with many popular Python
libraries pre-installed, including TensorFlow, Keras, PyTorch, NumPy,
pandas, and more. This saves time and effort in setting up your
environment.
• Free and Paid Tiers: While Google Colab provides free access to
computational resources, there are also paid tiers (Colab Pro and
Colab Pro+) that offer enhanced features, such as longer runtimes,
more memory, and priority access to faster GPUs and TPUs.
Practical Exercises
Preliminaries
• Practical exercises can also be accessed via the youtube links
provided for purposes of practice.
• The codes have also been uploaded to Gitihub and can be
downloaded via provided links.
Practical Data Analysis
• Netflix Stock analysis -Download dataset from Kaggle
[Link]
set-20022022/code
• practice using the videos in the links below
• [Link]
• [Link]

• Wine quality datasets- [Link]


quality-dataset
• [Link]

• Other open access datasets-


GitiHub Link to Codes and Datasets
• [Link]
• Note that the codes were created using jupyter notebooks then
uploaded to Gitihub.
• If your machine does have the capacity to handle anaconda try using
Kaggle Notebooks

You might also like