Contents
Module 1..................................................................................................................... 1
Modern Data Ecosystem and the Role of Data Analytics........................................1
Data Analytics vs. Data Analysis..........................................................................1
Summary and Highlights......................................................................................1
The Data Analyst Role.............................................................................................2
Generative AI: An essential Skill for today's Data Analysts.................................2
Summary and Highlights......................................................................................... 7
Module 2..................................................................................................................... 8
Module 1
Modern Data Ecosystem and the Role of Data Analytics
Data Analytics vs. Data Analysis
The terms Data Analysis and Data Analytics are often used interchangeably,
including in this course.
However it is important to note that there is a subtle difference between the terms
and meaning of the words Analysis and Analytics. In fact some people go far as
saying that these terms mean different things and should not be used
interchangeably. Yes, there is a technical difference...
The dictionary meanings are:
Analysis - detailed examination of the elements or structure of something
Analytics - the systematic computational analysis of data or statistics
Analysis can be done without numbers or data, such as business analysis
psycho analysis, etc. Whereas Analytics, even when used without the prefix
"Data", almost invariably implies use of data for perfoming numerical
manipulation and inference.
Some experts even say that Data Analysis is based on inferences based on
historical data whereas Data Analytics is for predicting future performance. The
design team of this course does not subscribe to this view, and you will see why later
in the course as you become familiar with the terms like predictive analytics,
prescriptive analytics, etc.
So in this course we take a more liberal view, and use the terms Data Analysis and
Data Analytics to mean the same thing. For example, an earlier video is titled
Defining Data Analysis, whereas the preceeding video with the viewpoints of several
data professionals is titled What is Data Analytics. The difference in these titles is not
intentional.
Summary and Highlights
In this lesson, you have learned the following information:
A modern data ecosystem includes a network of interconnected and continually
evolving entities that include:
Data that is available in a host of different formats, structure, and sources.
Enterprise Data Environment in which raw data is staged so it can be organized,
cleaned, and optimized for use by end-users.
End-users such as business stakeholders, analysts, and programmers who consume
data for various purposes.
Emerging technologies such as Cloud Computing, Machine Learning, and Big
Data, are continually reshaping the data ecosystem and the possibilities it offers.
Data Engineers, Data Analysts, Data Scientists, Business Analysts, and Business
Intelligence Analysts, all play a vital role in the ecosystem for deriving insights and
business results from data.
Based on the goals and outcomes that need to be achieved, there are four primary
types of Data Analysis:
1. Descriptive Analytics, that helps decode “What happened.”
2. Diagnostic Analytics, that helps us understand “Why it happened.”
3. Predictive Analytics, that analyzes historical data and trends to suggest “What
will happen next.”
4. Prescriptive Analytics, that prescribes “What should be done next.”
The Data Analysis process involves:
1. Developing an understanding of the problem and the desired outcome.
2. Setting a clear metric for evaluating outcomes.
3. Gathering, cleaning, analyzing, and mining data to interpret results.
4. Communicating the findings in ways that impact decision-making.
The Data Analyst Role
Generative AI: An essential Skill for today's Data Analysts
Introduction
As a beginner in data analytics, you’re stepping into a field that’s rapidly evolving.
Generative AI is becoming an essential tool for data analysts, allowing them to
create new content and gain deeper insights. Let’s explore what generative AI is and
how it can enhance your skills.
What is generative AI?
Generative AI refers to a class of artificial intelligence models that create new
content such as text, images, music, and more by learning patterns from existing
data.
Generative AI can respond naturally to human conversation and serve as a tool for
customer service and personalization of customer workflows. For example, you can
use AI-powered chatbots, voice bots, and virtual assistants that respond more
accurately to customers for first-contact resolution.
How does generative AI work?
Generative AI starts with a prompt that could be in the form of a text, an image, a
video, a design, musical notes, or any input that the AI system can process. Various
AI algorithms then return new content in response to the prompt. Content can
include essays, solutions to problems, or realistic fakes created from pictures or
audio of a person.
Early versions of generative AI required submitting data via an API or an otherwise
complicated process. Developers had to familiarize themselves with special tools
and write applications using languages such as Python.
Now, pioneers in generative AI are developing better user experiences that let you
describe a request in plain language. After an initial response, you can also
customize the results with feedback about the style, tone, and other elements you
want the generated content to reflect.
Key techniques in generative AI:
1. Generative adversarial networks (GANs): GANs consist of two neural
networks: the generator and the discriminator. The generator creates new
data, whereas the discriminator evaluates it. Over time, the generator
improves to produce realistic data.
2. Variational auto encoders (VAEs): VAEs encode input data into a compressed
format and then decode it back, generating new data points similar to the
input data.
3. Transformers: Used primarily in natural language processing (NLP),
transformers generate human-like text by predicting the next word in a
sequence. Generative Pre-trained Transformer 3 (GPT-3) is a notable
example.
Generative AI models
Generative AI models combine various AI algorithms to represent and process
content. For example, to generate text, various NLP techniques transform raw
characters (e.g., letters, punctuation, and words) into sentences, parts of speech,
entities, and actions, which are represented as vectors using multiple encoding
techniques. Similarly, images are transformed into various visual elements, also
expressed as vectors. One caution is that these techniques can also encode the
biases, racism, deception, and puffery contained in the training data.
Once developers settle on a way to represent the world, they apply a particular
neural network to generate new content in response to a query or prompt.
Techniques such as GANs and VAEs—neural networks with a decoder and encoder
—are suitable for generating realistic human faces, synthetic data for AI training, or
even facsimiles of particular humans.
Recent progress in transformers, such as Google’s Bidirectional Encoder
Representations from Transformers (BERT), OpenAI’s GPT, and Google AlphaFold,
have also resulted in neural networks that can not only encode language, images,
and proteins but also generate new content.
What are the use cases for generative AI?
Generative AI can be applied in various use cases to generate virtually any kind of
content. The technology is becoming more accessible to users of all kinds thanks to
cutting-edge breakthroughs like GPT that can be tuned for different applications.
Some of the use cases for generative AI include the following:
1. Implementing chatbots for customer service and technical support.
2. Deploying deepfakes for mimicking people or even specific individuals.
3. Improving dubbing for movies and educational content in different languages.
4. Writing email responses, dating profiles, resumes, and term papers.
5. Creating photorealistic art in a particular style.
6. Improving product demonstration videos.
7. Suggesting new drug compounds to test.
8. Designing physical products and buildings.
9. Optimizing new chip designs.
10. Writing music in a specific style or tone.
What are the benefits of generative AI?
Generative AI can be applied extensively across many areas of the business. It can
make it easier to interpret and understand existing content and automatically create
new content. Developers are exploring ways that generative AI can improve existing
workflows, with an eye to adapting workflows entirely to take advantage of the
technology. Some of the potential benefits of implementing generative AI include the
following:
Automating the manual process of writing content.
Reducing the effort of responding to emails.
Improving the response to specific technical queries.
Creating realistic representations of people.
Summarizing complex information into a coherent narrative.
Simplifying the process of creating content in a particular style
What are the limitations of generative AI?
Early implementations of generative AI vividly illustrate its many limitations. Some of
the challenges generative AI presents result from the specific approaches used to
implement particular use cases. For example, a summary of a complex topic is
easier to read than an explanation that includes various sources supporting key
points. The readability of the summary, however, comes at the expense of a user
being able to vet where the information comes from.
Here are some of the limitations to consider when implementing or using a
generative AI app:
It does not always identify the source of content.
It can be challenging to assess the bias of original sources.
Realistic-sounding content makes it harder to identify inaccurate information.
It can be difficult to understand how to tune in to new circumstances.
Results can gloss over bias, prejudice, and hatred.
What are the concerns surrounding generative AI?
The rise of generative AI is also fueling various concerns. These relate to the quality
of results, the potential for misuse and abuse, and the potential to disrupt existing
business models. Here are some of the specific types of problematic issues posed
by the current state of generative AI:
It can provide inaccurate and misleading information.
It is more difficult to trust without knowing the source and provenance of
information.
It can promote new kinds of plagiarism that ignore the rights of content
creators and artists of original content.
It might disrupt existing business models built around search engine
optimization and advertising.
It makes it easier to generate fake news.
It makes it easier to claim that real photographic evidence of wrongdoing was
just an AI-generated fake.
It could impersonate people for more effective social engineering
cyberattacks.
Given the newness of GenAI tools and their rapid adoption, enterprises should
prepare for the inevitable “trough of disillusionment” that’s part and parcel of
emerging technology by adopting sound AI engineering practices and making
responsible AI a cornerstone of their GenAI efforts, ensuring transparency, ethical
considerations, and long-term sustainability in their AI implementations.
What are some examples of generative AI tools?
Generative AI tools exist for various modalities, such as text, imagery, music, code,
and voices. Some popular AI content generators to explore include the following:
Text generation tools include GPT, Jasper, AI-Writer, and Lex.
Image generation tools include Dall-E 2, Midjourney, and Stable Diffusion.
Music generation tools include Amper, Dadabots, and MuseNet.
Code generation tools include codeStarter, Codex, GitHub Copilot, and
Tabnine.
Voice synthesis tools include Descript, Listnr, and [Link].
AI chip design tool companies include Synopsys, Cadence, Google, and
NVIDIA.
Applications of generative AI in data analytics
Generative AI has many applications that can enhance your data analytics work:
Data augmentation: Create synthetic data to augment existing data sets,
which is especially useful when data is scarce or imbalanced. This can
improve predictive model performance.
Anomaly Detection: Identify anomalies or outliers by understanding the
distribution of normal data. This is valuable in fraud detection, network
security, and quality control.
Text and image generation: Generate realistic text and images for marketing,
content creation, and customer engagement, such as automatic product
descriptions and marketing visuals.
Simulation and forecasting: Simulate scenarios and forecast future events by
generating potential outcomes from historical data. This is crucial in financial
planning, supply chain management, and strategic decision-making.
Conclusion
Generative AI is a transformative technology that can significantly enhance your
capabilities as a data analyst. By mastering generative AI techniques, you can
unlock new possibilities in data augmentation, anomaly detection, content creation,
and forecasting. As you embark on this journey, remember to balance innovation
with ethical responsibility, ensuring that AI is used positively.
Summary and Highlights
In this lesson, you have learned the following information:
The role of a Data Analyst spans across:
Acquiring data that best serves the use case.
Preparing and analyzing data to understand what it represents.
Interpreting and effectively communicating the message to stakeholders who
need to act on the findings.
Ensuring that the process is documented for future reference and
repeatability.
In order to play this role successfully, Data Analysts need a mix
of technical, functional, and soft skills.
Technical Skills include varying levels of proficiency in using spreadsheets, statistical
tools, visualization tools, programming and querying languages, and the ability to
work with different types of data repositories and big data platforms.
An understanding of Statistics, Analytical techniques, problem-solving, the ability to
probe a situation from multiple perspectives, data visualization, and project
management skills – all of which come under Functional Skills a Data Analyst needs
in order to play an effective role.
Soft Skills include the ability to work collaboratively, communicate effectively, tell a
compelling story with data, and garner support and buy-in from stakeholders.
Curiosity to explore different pathways and intuition that helps to give a sense of the
future based on past experiences are also essential skills for being a good Data
Analyst.
Module 2
The Data Ecosystem and Languages for Data Professionals
In this lesson, you have learned the following information:
A data analyst ecosystem includes the infrastructure, software, tools, frameworks,
and processes used to gather, clean, analyze, mine, and visualize data.
Based on how well-defined the structure of the data is, data can be categorized as:
1. Structured Data, that is data which is well organized in formats that can be
stored in databases.
2. Semi-Structured Data, that is data which is partially organized and partially
free form.
3. Unstructured Data, that is data which can not be organized conventionally into
rows and columns.
Data comes in a wide-ranging variety of file formats, such as delimited text
files, spreadsheets, XML, PDF, and JSON, each with its own list of benefits and
limitations of use.
Data is extracted from multiple data sources, ranging from relational and non-
relational databases to APIs, web services, data streams, social platforms, and
sensor devices.
Once the data is identified and gathered from different sources, it needs to be staged
in a data repository so that it can be prepared for analysis. The type, format, and
sources of data influence the type of data repository that can be used.
Data professionals need a host of languages that can help them extract, prepare,
and analyze data. These can be classified as:
Querying languages, such as SQL, used for accessing and manipulating data
from databases.
Programming languages such as Python, R, and Java, for developing
applications and controlling application behavior.
Shell and Scripting languages, such as Unix/Linux Shell, and PowerShell, for
automating repetitive operational tasks.
Understanding Data Repositories and Big Data Platforms
In this lesson, you have learned the following information:
A Data Repository is a general term that refers to data that has been collected,
organized, and isolated so that it can be used for reporting, analytics, and also for
archival purposes.
The different types of Data Repositories include:
Databases, which can be relational or non-relational, each following a set of
organizational principles, the types of data they can store, and the tools that
can be used to query, organize, and retrieve data.
Data Warehouses, that consolidate incoming data into one
comprehensive storehouse.
Data Marts, that are essentially sub-sections of a data warehouse, built to
isolate data for a particular business function or use case.
Data Lakes, that serve as storage repositories for large amounts of structured,
semi-structured, and unstructured data in their native format.
Big Data Stores, that provide distributed computational and storage
infrastructure to store, scale, and process very large data sets.
ETL, or Extract Transform and Load, Process is an automated process that converts
raw data into analysis-ready data by:
Extracting data from source locations.
Transforming raw data by cleaning, enriching, standardizing, and validating it.
Loading the processed data into a destination system or data repository.
Data Pipeline, sometimes used interchangeably with ETL, encompasses the entire
journey of moving data from the source to a destination data lake or
application, using the ETL process.
Big Data refers to the vast amounts of data that is being produced each moment of
every day, by people, tools, and machines. The sheer velocity, volume, and variety
of data challenge the tools and systems used for conventional data. These
challenges led to the emergence of processing tools and platforms designed
specifically for Big Data, such as Apache Hadoop, Apache Hive, and Apache Spark.
Module 3
Gathering Data
In this lesson, you have learned:
The process of identifying data begins by determining the information that
needs to be collected, which in turn is determined by the goal you seek to
achieve.
Having identified the data, your next step is to identify the sources from which
you will extract the required data and define a plan for data collection.
Decisions regarding the timeframe over which you need your data set, and
how much data would suffice for arriving at a credible analysis also weigh in
at this stage.
Data Sources can be internal or external to the organization, and they can
be primary, secondary, or third-party, depending on whether you are obtaining
the data directly from the original source, retrieving it from externally available
data sources, or purchasing it from data aggregators.
Some of the data sources from which you could be gathering data include
databases, the web, social media, interactive platforms, sensor devices, data
exchanges, surveys and observation studies.
Data that has been identified and gathered from the various data sources is
combined using a variety of tools and methods to provide a single
interface using which data can be queried and manipulated.
The data you identify, the source of that data, and the practices you employ
for gathering the data have implications for quality, security, and privacy,
which need to be considered at this stage.
Wrangling Data
In this lesson, you have learned the following information:
Once the data you identified is gathered and imported, your next step is to make
it analysis-ready. This is where the process of Data Wrangling, or Data Munging,
comes in. Data Wrangling is an iterative process that involves data
exploration, transformation, and validation.
Transformation of raw data includes the tasks you undertake to:
Structurally manipulate and combine the data using Joins and Unions.
Normalize data, that is, clean the database of unused and redundant data.
Denormalize data, that is, combine data from multiple tables into a single
table so that it can be queried faster.
Clean data, which involves profiling data to uncover quality issues, visualizing
data to spot outliers, and fixing issues such as missing values, duplicate
data, irrelevant data, inconsistent formats, syntax errors, and outliers.
Enrich data, which involves considering additional data points that could add
value to the existing data set and lead to a more meaningful analysis.
A variety of software and tools are available for the Data Wrangling process. Some
of the popularly used ones include Excel Power Query, Spreadsheets, OpenRefine,
Google DataPrep, Watson Studio Refinery, Trifacta Wrangler, Python, and R, each
with their own set of characteristics, strengths, limitations, and applications.
Module 4
Analyzing and Mining Data
In this lesson, you have learned the following information:
Statistics is a branch of mathematics dealing with the collection, analysis,
interpretation, and presentation of numerical or quantitative data.
Statistical Analysis involves the use of statistical methods in order to develop an
understanding of what the data represents.
Statistical Analysis can be:
Descriptive; that which provides a summary of what the data represents.
Common measures include Central Tendency, Dispersion, and Skewness.
Inferential; that which involves making inferences, or generalizations, about
data. Common measures include Hypothesis Testing, Confidence Intervals,
and Regression Analysis.
Data Mining, simply put, is the process of extracting knowledge from data. It
involves the use of pattern recognition technologies, statistical analysis, and
mathematical techniques, in order to identify
correlations, patterns, variations, and trends in data.
There are several techniques that can help mine data, such
as, classifying attributes of data, clustering data into groups, establishing
relationships between events, variables, and input and output.
A variety of software and tools are available for analyzing and mining data. Some of
the popularly used ones include Spreadsheets, R-Language, Python, IBM SPSS
Statistics, IBM Watson Studio, and SAS, each with their own set of characteristics,
strengths, limitations, and applications.
Communicating data analysis findings
In this lesson, you have learned the following information:
Data has value through the stories that it tells. In order to communicate your findings
impactfully, you need to:
Ensure that your audience is able to trust you, understand you, and relate to
your findings and insights.
Establish the credibility of your findings.
Present the data within a structured narrative.
Support your communication with strong visualizations so that the message is
clear and concise, and drives your audience to take action.
Data visualization is the discipline of communicating information through the use
of visual elements such as graphs, charts, and maps. The goal of visualizing data is
to make information easy to comprehend, interpret, and retain.
For data visualization to be of value, you need to:
Think about the key takeaway for your audience.
Anticipate their information needs and questions, and then plan the
visualization that delivers your message clearly and impactfully.
There are several types of graphs and charts available for you to be able to plot any
kind of data, such as bar charts, column charts, pie charts, and line charts.
You can also use data visualization to build dashboards. Dashboards organize and
display reports and visualizations coming from multiple data sources into a single
graphical interface. They are easy to comprehend and allow you to generate reports
on the go.
When deciding which tools to use for data visualization, you need to consider the
ease-of-use and purpose of the visualization. Some of the popularly used tools
include Spreadsheets, Jupyter Notebook, Python libraries, R-Studio and R-
Shiny, IBM Cognos Analytics, Tableau, and Power BI.
Module 5
Opportunities and learning path
In this lesson, you have learned the following information:
Data Analyst roles are sought after in every industry, be it Banking and Finance,
Insurance, Healthcare, Retail, or Information Technology.
Currently, the demand for skilled data analysts far outweighs the supply, which
means companies are willing to pay a premium to hire skilled data analysts.
Data Analyst job roles can be broadly classified as follows:
Data Analyst Specialist roles - On this path, you start as a Junior Data Analyst
and move up to the level of a Principal Analyst by continually advancing your
technical, statistical, and analytical skills from a foundational level to an expert
level.
Domain Specialist roles - These roles are for you if you have acquired
specialization in a specific domain and want to work your way up to be seen
as an authority in your domain.
Analytics-enabled job roles - These roles include jobs where having analytic
skills can up-level your performance and differentiate you from your peers.
Other Data Professions - There are several other roles in a modern data
ecosystem, such as Data Engineer, Big Data Engineer, Data Scientist,
Business Analyst, or Business Intelligence Analyst. If you upskill yourself
based on the required skills, you can transition into these roles.
There are several paths you can consider in order to gain entry into the Data
Analyst field. These include:
An academic degree in Data Analytics or disciplines such as Statistics and
Computer Science.
Online multi-course specializations offered by learning platforms such as
Coursera, edX, and Udacity.
Mid-career transition into Data Analysis by upskilling yourself. If you have a
technical background, for example, you can focus on developing the technical
skills specific to Data Analysis. If you do not have a technical background, you
can plan to skill your self in some basic technologies and then work your way
up from an entry-level position.
Final Assignment: Data Analysis in Action
Using Data Analysis for Detecting Credit Card Fraud
Companies today are employing analytical techniques for the early detection of
credit card frauds, a key factor in mitigating fraud damage. The most common type
of credit card fraud does not involve the physical stealing of the card, but that of
credit card credentials, which are then used for online purchases.
Imagine that you have been hired as a Data Analyst to work in the Credit Card
Division of a bank. And your first assignment is to join your team in using data
analysis for the early detection and mitigation of credit card fraud.
In order to prescribe a way forward, that is, suggest what should be done in order for
fraud to get detected early on, you need to understand what a fraudulent transaction
looks like. And for that you need to start by looking at historical data.
Here is a sample data set that captures the credit card
transaction details for a few users.
Descriptive techniques of analysis, that is, techniques that help you gain an
understanding of what happened, include the identification of patterns and anomalies
in data. Anomalies signify a variation in a pattern that seems uncharacteristic, or, out
of the ordinary. Anomalies may occur for perfectly valid and genuine reasons, but
they do warrant an evaluation because they can be a sign of fraudulent activity.
Past studies have suggested that some of the common events
that you may need to watch out for include:
A change in frequency of orders placed, for example, a customer who
typically places a couple of orders a month, suddenly makes numerous
transactions within a short span of time, sometimes within minutes of the
previous order.
Orders that are significantly higher than a user’s average transaction.
Bulk orders of the same item with slight variations such as color or size—
especially if this is atypical of the user’s transaction history.
A sudden change in delivery preference, for example, a change from home or
office delivery address to in-store, warehouse, or PO Box delivery.
A mismatched IP Address, or an IP Address that is not from the general
location or area of the billing address.
Before you can analyze the data for patterns and anomalies, you
need to:
Identify and gather all data points that can be of relevance
to your use case. For example, the card holder’s details, transaction
details, delivery details, location, and network are some of the data points that
could be explored.
Clean the data. You need to identify and fix issues in the data that can
lead to false or incomplete findings, such as missing data values and incorrect
data. You may also need to standardize data formats in some cases, for
example, the date fields.
Finally, when you arrive at the findings, you will create appropriate visualizations that
communicate your findings to your audience. The graph below samples one such
visualization that you would use to capture a trend hidden in the sample data set
shared earlier on in the case study.
In the next section you will be asked to answer the following 5
(five) questions based on this case study:
1. List at least 5 (five) data points that are required for the analysis and detection
of a credit card fraud. (3 marks)
2. Identify 3 (three) errors/issues that could impact the accuracy of your findings,
based on a data table provided. (3 marks)
3. Identify 2 (two) anomalies, or unexpected behaviors, that would lead you to
believe the transaction may be suspect, based on a data table provided. (2
marks)
4. Briefly explain your key take-away from the provided data visualization chart.
(1 mark)
5. Identify the type of analysis that you are performing when you are analyzing
historical credit card data to understand what a fraudulent transaction looks
like. [Hint: The four types of Analytics include: Descriptive, Diagnostic,
Predictive, Prescriptive] (1 mark)