Google Data Analytics Course Overview
Google Data Analytics Course Overview
1. Discovery
2. Pre-processing data
3. Model planning
4. Model building
5. Communicate results
6. Operationalize
EMC Corporation is now Dell EMC. This model, created by David Dietrich, reflects the
cyclical nature of typical business projects. The phases aren’t static milestones; each
step connects and leads to the next, and eventually repeats.
An iterative data analysis process was created by a company called SAS, a leading
data analytics solutions provider. It can be used to produce repeatable, reliable, and
predictive results:
1. Ask
2. Prepare
3. Explore
4. Model
5. Implement
6. Act
7. Evaluate
The SAS model emphasizes the cyclical nature of their model by visualizing it as an
infinity symbol. Its process has seven steps, many of which mirror the other models, like
ask, prepare, model, and act. But this process is also a little different; it includes a step
after the act phase designed to help analysts evaluate their solutions and potentially
return to the ask phase again.
3. Pre-processing data
5. Visualizing data
Big data analytics process
Authors Thomas Erl, Wajid Khattak, and Paul Buhler proposed a big data analytics
process in their book, Big Data Fundamentals: Concepts, Drivers &
Techniques. Their process suggests phases divided into nine steps:
2. Data identification
4. Data extraction
7. Data analysis
8. Data visualization
This process appears to have three or four more steps than the previous models. But in
reality, they have just broken down what has been referred to as prepare and process
into smaller steps. It emphasizes the individual tasks required for gathering, preparing,
and cleaning data before the analysis phase.
An ecosystem is a group of elements that interact with one another. Ecosystems can
be large, like the jungle in a tropical rainforest or the Australian outback. Or, tiny, like
tadpoles in a puddle, or bacteria on your skin. And just like the kangaroos and koala
bears in the Australian outback, data lives inside its own ecosystem too. Data
ecosystems are made up of various elements that interact with one another in order to
produce, manage, store, organize, analyze, and share data. These elements include
hardware and software tools, and the people who use them.
The cloud is a place to keep data online, rather than on a computer hard drive. So
instead of storing data somewhere inside your organization's network, that data is
accessed over the internet. So the cloud is just a term we use to describe the virtual
location. The cloud plays a big part in the data ecosystem, and as a data analyst, it's
your job to harness the power of that data ecosystem, find the right information, and
provide the team with analysis that helps them make smart decisions.
Agricultural companies regularly use data ecosystems that include information
including geological patterns in weather movements. Data analysts can use this data to
help farmers predict crop yields. Some data analysts are even using data ecosystems to
save real environmental ecosystems
At the Scripps Institution of Oceanography, coral reefs all over the world are monitored
digitally, so they can see how organisms change over time, track their growth, and
measure any increases or declines in individual colonies.
Data science is defined as creating new ways of modeling and understanding the
unknown by using raw data. Here's a good way to think about it. Data scientists create
new questions using data, while analysts find answers to existing questions by creating
insights from data sources.
Data analysis and data analytics sound the same, but they're actually very different
things. Let's start with analysis. You've already learned that data analysis is the
collection, transformation, and organization of data in order to draw conclusions, make
predictions, and drive informed decision-making. Data analytics in the simplest terms is
the science of data. It's a very broad concept that encompasses everything from the job
of managing and using data to the tools and methods that data workers use each and
every day. So when you think about data, data analysis and the data ecosystem,
it's important to note that no matter how valuable data-driven decision-making is, data
alone will never be as powerful as data combined with human experience, observation,
and sometimes even intuition. To get the most out of data-driven decision-making, it's
important to include insights from people who are familiar with the business problem.
These people are called subject matter experts, and they have the ability to look at the
results of data analysis and identify any inconsistencies, make sense of gray areas, and
eventually validate choices being made.
At the heart of data-driven decision making is data. Therefore, it's essential that data
analysts focus on the data to ensure they make informed decisions. If you ignore data
by preferring to make decisions based on your own experience, your decisions may be
biased. But even worse, decisions based on gut instinct without any data to back them
up can cause mistakes.
Blending data with business knowledge, plus maybe a touch of gut instinct, will be a
common part of your process as a junior data analyst. The key is figuring out the exact
mix for each particular project. A lot of times, it will depend on the goals of your
analysis. That is why analysts often ask, “How do I define success for this project?”
Key takeaways
Analytical skills are qualities and characteristics associated with solving problems
using facts.
Your inherent analytical skills are essential for conducting data analysis and will be
even more critical when you combine them with the tools and techniques from this
program. Understanding how to use these skills in business scenarios is the first step
toward developing them further and using them effectively in your career.
Analytical thinking involves identifying and defining a problem and then solving it by
using data in an organized, step-by-step manner
Visualization is important because visuals can help data analysts understand and
explain information more effectively. Think about it like this. If you are trying to explain
the Grand Canyon to someone, using words would be much more challenging than
showing them a picture.
Strategizing helps data analysts see what they want to achieve with the data and how
they can get there.
correlation is like a relationship. You can find all kinds of correlations in data. Maybe
it's the relationship between the length of your hair and the amount of shampoo you
need. Or maybe you notice a correlation between a rainier season leading to a high
number of umbrellas.
big-picture thinking. This means being able to see the big picture as well as the
details. A jigsaw puzzle is a great way to think about this. Big-picture thinking is like
looking at a complete puzzle.
A root cause is the reason why a problem occurs. If we can identify and get rid of a
root cause, we can prevent that problem from happening again. A simple way to wrap
your head around root causes is with the process called the Five Whys. In the Five
Whys you ask "why" five times to reveal the root cause.
Gap analysis lets you examine and evaluate how a process works currently in order to
get where you want to be in the future. Businesses conduct gap analysis to do all kinds
of things, such as improve a product or become more efficient. The general approach to
gap analysis is understanding where you are now compared to where you want to be.
Then you can identify the gaps that exist between the current and future state and
determine how to bridge them.
A root cause is the reason why a problem occurs. So, by identifying and eliminating the
root cause, data professionals can help stop that problem from occurring again.
Analytical skills: Qualities and characteristics associated with using facts to solve
problems
Data analyst: Someone who collects, transforms, and organizes data in order to
draw conclusions, make predictions, and drive informed decision-making
Data ecosystem: The various elements that interact with one another in order to
produce, manage, store, organize, analyze, and share data
Data science: A field of study that uses raw data to create new ways of modeling
and understanding the unknown
Data strategy: The management of the people, processes, and tools used in data
analysis
Gap analysis: A method for examining and evaluating the current state of a process
in order to identify opportunities for improvement in the future
Technical mindset: The ability to break things down into smaller steps or pieces
and work with them in an orderly and logical way
Module 2
The life cycle of data is plan, capture, manage, analyze, archive and destroy.
Planning, a business decides what kind of data it needs, how it will be managed
throughout its life cycle, who will be responsible for it, and the optimal outcomes. For
example, let's say an electricity provider wanted to gain insights into how to save people
energy.
Capture data This is where data is collected from a variety of different sources and
brought into the organization. With so much data being created everyday, the ways to
collect it are truly endless. One common method is getting data from outside resources.
For example, if you were doing data analysis on weather patterns, you'd probably get
data from a publicly available dataset like the National Climatic Data Center.
Manage Data. Here we're talking about how we care for our data, how and where it's
stored, the tools used to keep it safe and secure, and the actions taken to make sure
that it's maintained properly. This phase is very important to data cleansing, which we'll
cover later on.
Analyze your data. This is where data analysts really shine. In this phase, the data is
used to solve problems, make great decisions, and support business goals. For
example, one of our electricity company's goals might be to find ways to help customers
save energy. Moving along the data life cycle now evolves to the archive phase.
Archiving means storing data in a place where it's still available, but may not be used
again. During analysis, analysts handle huge amounts of data. Can you imagine if we
had to sort through all of the available data that's out there, even if it was no longer
useful and relevant to our work? It makes way more sense to archive it than to keep it
around.
Destroy it, the company would use a secure data erasure software. If there were any
paper files, they would be shredded too. This is important for protecting a company's
private information, as well as private data about its customers. And there you have it,
the data life cycle.
You have learned that there are six stages to the data life cycle. Here's a recap:
1. Plan: Decide what kind of data is needed, how it will be managed, and who will
be responsible for it.
2. Capture: Collect or bring in data from a variety of different sources.
3. Manage: Care for and maintain the data. This includes determining how and
where it is stored and the tools used to do so.
4. Analyze: Use the data to solve problems, make decisions, and support business
goals.
5. Archive: Keep relevant data stored for long-term and future reference.
6. Destroy: Remove data from storage and delete any shared copies of the data.
In the prepare phase, the emphasis is on identifying and locating data you can use to
answer your questions. In an upcoming course, you'll learn more about the different
types of data and how to identify which kinds of data are most useful for solving a
particular problem.
In this phase, the aim is to refine the data. Data analysts find and eliminate any errors
and inaccuracies that can get in the way of results. This usually means:
Cleaning data
With a solid foundation of well-defined questions and clean data, you’ll delve into the
analyze phase. This is when you turn the data you’ve gathered, prepared, and
processed into actionable information. Data analysts use many powerful tools in their
work
This phase is exactly what it sounds like: It’s time to share what you’ve learned with
your stakeholders! In this part of the program, you'll learn how data analysts interpret
results and share them with others to help stakeholders make effective, data-driven
decisions. In the share phase, visualization is a data analyst's best friend.
The data analysis journey culminates in the act phase, when data insights are put to
work. For you, this action involves preparing for your job search and having the chance
to complete a case study project. It's a great opportunity for you to bring together
everything you've worked on throughout this course. Plus, adding a case study to your
portfolio helps you stand out from other candidates
Data analysts rely on spreadsheets to collect and organize data. Two popular
spreadsheet applications you will probably use a lot in your future role as a data analyst
are Microsoft Excel and Google Sheets.
Identify patterns and piece the data together in a way that works for each specific data
project
Query languages
Make it easier for you to learn and understand the requests made to databases
Allow analysts to select, create, add, or download data from a database for analysis
Visualization tools
Data analysts use a number of visualization tools, like graphs, maps, tables, charts, and
more. Two popular visualization tools are Tableau and Looker.
These tools
Help stakeholders come up with conclusions that lead to informed decisions and
effective business strategies
worksheets
- Looker communicates directly with a database, allowing you to connect your data
right to the visual
Terms and definitions for Course 1, Module 2
Stakeholders: People who invest time and resources into a project and are
interested in its outcome
What is a query?
A query is a request for data or information from a database. When you query
databases, you use SQL to communicate your question or request.
Syntax. Syntax is the predetermined structure of a language that includes all required
words, symbols, and punctuation, as well as their proper placement.
Use FROM to choose the tables where the columns you want are located.
SELECT
FROM
WHERE
Next it is time to refine the data and present the results of your analysis. Right now, you
have a lot of data spread across several different tables, which isn’t an ideal way to
share your results with management and the marketing team
Now that you have decided what kind of information and insights you want to display, it
is time to start creating the actual visualizations. Keep in mind that creating the right
visualization for a presentation or to share with stakeholders is a process.
There are many different tools you can use for data visualization.
You can use the visualizations tools in your spreadsheet to create simple visualizations
such as line and bar charts.
You can use more advanced tools such as Tableau that allow you to integrate data into
dashboard-style visualizations.
If you’re working with the programming language R you can use the visualization tools
in RStudio.
Spreadsheets (Microsoft Excel or Google Sheets)
In our example, the built-in charts and graphs in spreadsheets made the process of
creating visuals quick and easy. Spreadsheets are great for creating simple
visualizations like bar graphs and pie charts, and even provide some advanced
visualizations like maps, and waterfall and funnel diagrams (shown in the following
figures).
Tableau is a popular data visualization tool that lets you pull data from nearly any
system and turn it into compelling visuals or actionable insights.
A lot of data analysts work with a programming language called R. Most people who
work with R end up also using RStudio, an integrated developer environment (IDE), for
their data visualization needs. As with Tableau, you can create dashboard-style data
visualizations using RStudio.
Module 4
Fairness means ensuring your analysis doesn't create or reinforce bias. This can be
challenging, but if the analysis is not objective, the conclusions can be misleading and
even harmful.
Glossary terms from course 1, module 4
Terms and definitions for Course 1, Module 4
Business task: The question or problem data analysis resolves for a business
Fairness: A quality of data analysis that does not create or reinforce bias
Oversampling: The process of increasing the sample size of nondominant groups in a population. This can help
you better represent them and address imbalanced datasets
Self-reporting: A data collection technique where participants provide information about themselves
Module 1 Course 2
Here's an example that breaks down the thought process of turning a problem question
into one or more SMART questions using the SMART method: What features do
people look for when buying a new car?
Relevant: Does the question identify which features make or break a potential car
purchase?
Time-bound: Does the question validate data on the most popular features from the
last three years?
Questions should be open-ended. This is the best way to get responses that will
help you accurately qualify or disqualify potential solutions to your specific problem. So,
based on the thought process, possible SMART questions might be:
Cloud: A place to keep data online, rather than a computer hard drive
Data analysis process: The six phases of ask, prepare, process, analyze, share, and act
whose purpose is to gain insights that drive informed decision-making
Data life cycle: The sequence of stages that data experiences, which include plan, capture,
manage, analyze, archive, and destroy
Problem types: The various problems that data analysts encounter, including categorizing
things, discovering connections, finding patterns, identifying themes, making predictions, and
spotting something unusual
Specific question: A question that is simple, significant, and focused on a single topic or a
few closely related ideas
Data sources
If you don’t collect the data using your own resources, you might get data from second-
party or third-party data providers. Second-party data is collected directly by
another group and then sold. Third-party data is sold by a provider that didn’t
collect the data themselves. Third-party data might come from a number of different
sources.
Datasets can show a lot of interesting information. But be sure to choose data that can
actually help solve your problem question. For example, if you are analyzing trends over
time, make sure you use time series data — in other words, data that includes dates.
If you are collecting your own data, make reasonable decisions about sample size. A
random sample from existing data might be fine for some projects. Other projects might
need more strategic data collection to focus on certain criteria. Each project has its own
needs.
1. Conceptual data modeling gives a high-level view of the data structure, such as
how data interacts across an organization. For example, a conceptual data model may
be used to define the business requirements for a new database. A conceptual data
model doesn't contain technical details.
3. Physical data modeling depicts how a database operates. A physical data model
defines all entities and attributes used; for example, it includes table names, column
names, and data types for the database.
Audio file: Digitized audio storage usually in an MP3, AAC, or other compressed
format
Boolean data: A data type with only two possible values, usually true or false
Continuous data: Data that is measured and can have almost any numeric value
Cookie: A small file stored on a computer that contains information about its users
Data model: A tool for organizing data elements and how they relate to one another
Discrete data: Data that is counted and has a limited number of values
First-party data: Data collected by an individual or group using their own resources
Long data: A dataset in which each row is one time point per subject, so each
subject has data in multiple rows
Nominal data: A type of qualitative data that is categorized without a set order
Ownership: The aspect of data ethics that presumes individuals own the raw data
they provide and have primary control over its usage, processing, and sharing
Pixel: In digital imaging, a small area of illumination on a display screen that, when
combined with other adjacent areas, forms a digital image
Record: A collection of related data in a data table, usually synonymous with row
Second-party data: Data collected by a group directly from its audience and then
sold
Social media: Websites and applications through which users create and share
content or participate in social networking
String data type: A sequence of characters and punctuation that contains textual
information (Refer to Text data type)
Structured data: Data organized in a certain format such as rows and columns
Text data type: A sequence of characters and punctuation that contains textual
information (also called string data type)
Unstructured data: Data that is not organized in any easily identifiable manner
Video file: A collection of images, audio files, and other data usually encoded in a
compressed format such as MP4, MV4, MOV, AVI, or FLV
Wide data: A dataset in which every data subject has a single row with multiple
columns to hold the values of various attributes of the subject
In a non-relational table, you will find all of the possible variables you might be
interested in analyzing all grouped together. This can make it really hard to sort
through. This is one reason why relational databases are so common in data analysis:
they simplify a lot of analysis processes and make data easier to find and use across an
entire database.
Module 3 Course 3
The key to relational databases
Tables in a relational database are connected by the fields they have in common. You
might remember learning about primary and foreign keys before. As a quick refresher, a
primary key is an identifier that references a column in which each value is unique.
In other words, it's a column of a table that is used to uniquely identify each record
within that table. The value assigned to the primary key in a particular row must be
unique within the entire table. For example, if customer_id is the primary key for the
customer table, no two customers will ever have the same customer_id.
By contrast, a foreign key is a field within a table that is a primary key in another
table. A table can have only one primary key, but it can have multiple foreign keys.
These keys are what create the relationships between tables in a relational database,
which helps organize and connect data across multiple tables in the database.
Some tables don't require a primary key. For example, a revenue table can have
multiple foreign keys and not have a primary key. A primary key may also be
constructed using multiple columns of a table. This type of primary key is called a
composite key. For example, if customer_id and location_id are two columns of a
composite key for a customer table, the values assigned to those fields in any given row
must be unique within the entire table.
Understanding Metadata
Metadata is information that describes data, helping to provide context and meaning.
It is not the data itself but rather data about the data, like labels on a box of toys.
Types of Metadata
Descriptive Metadata: Identifies and describes a piece of data, such as a book's title and
author.
Structural Metadata: Indicates how data is organized, like the chapters in a book.
Administrative Metadata: Provides technical details about a digital asset, such as file type
and creation date.
Importance of Metadata
Metadata helps data analysts interpret data within a database, enabling effective problem-
solving and data-driven decision-making.
It is crucial for organizing, protecting, and understanding data in various contexts, such as
photos and emails.
If you have any further questions or need clarification on any point, feel free to ask!
Data analysts use reliable and high-quality data to identify the root causes of any
problems that might occur during analysis and to improve their results. If the data being
used to solve a problem or to make a data-driven decision is unreliable, there’s a good
chance the results will be unreliable as well.
Metadata helps data analysts confirm their data is reliable by making sure it is:
Accurate
Precise
Relevant
Timely
Consistency
Data analysts thrive on consistency and aim for uniformity in their data and
databases, and metadata helps make this possible. For example, to use survey data
from two different sources, data analysts use metadata to make sure the same
collection methods were applied in the survey so that both datasets can be compared
reliably
Understanding Metadata
Data Governance
When internal data is insufficient, analysts can turn to external data for a broader
perspective, often collaborating with other organizations.
Open data initiatives, such as those by the U.S. government, provide public
access to various datasets, promoting transparency and innovation.
The content discusses how to handle situations where there is insufficient data
for analysis, which is a common challenge for data analysts.
Insufficient data can arise from limited time frames, such as only having current
year's data, which may not reveal seasonal trends.
Data from a single source can limit insights; for example, using only one booking
site may miss broader trends.
Analysts can wait for more data to arrive or adjust their analysis objectives based
on available data.
Engaging with stakeholders to redefine objectives can help in cases of
incomplete or outdated data.
Outdated data may not reflect current trends, necessitating the search for new
datasets.
Geographically-limited data can skew results, especially for global companies
that require comprehensive datasets.
Overall, understanding and addressing data limitations is crucial for effective analysis
and decision-making.
Understanding Sample Size
A population includes all possible data values in a dataset, but analyzing the
entire population can be impractical due to time and cost constraints.
Sample size allows analysts to use a representative portion of the population to
make predictions or conclusions, making the analysis more efficient.
Challenges of Sampling
Using a small sample can lead to uncertainty and potential sampling bias, where
certain groups may be overrepresented or underrepresented.
An example is given where a survey might exclude cat owners without
smartphones, leading to biased results.
Random Sampling
Random sampling ensures that every member of the population has an equal
chance of being selected, helping to mitigate sampling bias.
This method allows for a more accurate representation of the population, which is
crucial for effective data analysis.