Module 1
To track people's online activities and interests, which method of data collection is most
effective? -> Cookies
Select the right data
Following are some data-collection considerations to keep in mind for your analysis:
How the data will be collected
Decide if you will collect the data using your own resources or receive (and possibly purchase it)
from another party. Data that you collect yourself is called first-party data.
Data sources
If you don’t collect the data using your own resources, you might get data from second-party or
third-party data providers. Second-party data is collected directly by another group and then sold.
Third-party data is sold by a provider that didn’t collect the data themselves. Third-party data
might come from a number of different sources.
Solving your business problem
Datasets can show a lot of interesting information. But be sure to choose data that can actually
help solve your problem question. For example, if you are analyzing trends over time, make sure
you use time series data — in other words, data that includes dates.
How much data to collect
If you are collecting your own data, make reasonable decisions about sample size. A random
sample from existing data might be fine for some projects. Other projects might need more
strategic data collection to focus on certain criteria. Each project has its own needs.
Time frame
If you are collecting your own data, decide how long you will need to collect it, especially if you
are tracking trends over a long period of time. If you need an immediate answer, you might not
have time to collect new data. In this case, you would need to use historical data that already
exists.
Use the flowchart below if data collection relies heavily on how much time you have:
- What are cookies? -> Cookies are small files stored on computers that contain
information about users.
- For data analytics projects, first-party data is typically preferred because users know it
originated within the organization. First-party data is collected by an individual or group
using their own resources.
- A grocery store chain purchases customer data from a credit card company. The grocer
uses this data to identify its most loyal customers and offer them special promotions and
discounts. What type of data is being used in this scenario? -> Second-party data is being
used. Second-party data is collected by a group directly from its audience and then sold.
- In data analytics, what term refers to all possible data values in a dataset? -> Population
+ Population refers to the entire set of possible data points or values that could be
observed for a given variable or study.
+ Sample is a subset of the population used for analysis.
+ Source refers to where the data comes from (e.g., database, API).
+ Representation means how data is modeled or displayed, not the full set of values.
Discrete data isn't limited to dollar amounts. Examples of other discrete data are stars and
points. When partial measurements (half-stars or quarter-points) aren't allowed, the data
is discrete. If you don't accept anything other than full stars or points, the data is
considered discrete.
Continuous data: data that is measured and can have almost any numeric value
Its value can be shown as a decimal with several places
Nominal data: A type of qualitative data that is categorized without a set order
This data doesn’t have a sequence
Ex: collecting data about movies: You ask people if they've watched a given
movie. Their responses would be in the form of nominal data. They could
respond "Yes," "No," or "Not sure." These choices don't have a particular
order.
Ordinal data: A type of qualitative data with a set order or scale
If you asked a group of people to rank a movie from 1 to 5, some might rank it
as a 2, others a 4, and so on. These rankings are in order of how much each
person liked the movie.
Internal data: Data that lives within a company’s own system
For example, if a movie studio had compiled all of the data in the spreadsheet
using only their own collection methods, then it would be their internal
data. The great thing about internal data is that it's usually more reliable and
easier to collect, but in this spreadsheet, it's more likely that the movie
studio had to use data owned or shared by other studios and sources because it
includes movies they didn't make.
External data: Data that lives and is generated outside of an organization
External data becomes particularly valuable when your analysis depends on as
many sources as possible.
Structured data: Data organized in a certain format such as rows and columns
Spreadsheets and relational databases are two examples of software that can
store data in a structured way.
Unstructured data: Data that is not organized in any easily identifiable manner
Audio and video files are examples of unstructured data because there's no
clear way to identify or organize their content. Unstructured data might have
internal structure, but the data doesn't fit neatly in rows and columns like
structured data.
Data formats in practice
Structured Data: Data organized in a defined format, like rows and columns in a
database.
Qualitative Data: Subjective data that describes qualities or characteristics, such as
customer preferences.
Continuous Data: Data that can take any numeric value, such as height or temperature.
Unstructured Data: Data that cannot be easily organized into a structured format, such
as social media posts or videos.
Secondary Data: Data gathered by others or from previous research, such as census data
or purchased datasets.
Primary Data: Data collected directly by a researcher from first-hand sources, such as
interviews or surveys.
Internal Data: Data stored within a company's own systems, like employee wages or
sales data.
When you think about the word "format," a lot of things might come to mind. Think of an
advertisement for your favorite store. You might find it in the form of a print ad, a billboard, or
even a commercial. The information is presented in the format that works best for you to take it
in. The format of a dataset is a lot like that, and choosing the right format will help you manage
and use your data in the best way possible.
Data format examples
As with most things, it is easier for definitions to click when you can pair them with examples
you might encounter on a daily basis. Review each data format’s definition first and then use the
examples to lock in your understanding.
Primary versus secondary data
The following table highlights the differences between primary and secondary data and presents
examples of each.
Internal versus external data
The following table highlights the differences between internal and external data and presents
examples of each.
Continuous versus discrete data
The following table highlights the differences between continuous and discrete data and presents
examples of each.
Qualitative versus quantitative data
The following table highlights the differences between qualitative and quantitative data and
presents examples of each.
Nominal versus ordinal data
The following table highlights the differences between nominal and ordinal data and presents
examples of each.
Structured versus unstructured data
The following table highlights the differences between structured and unstructured data and
presents examples of each.
Unstructured data examples: Audio files, Video files, Emails, Photos, Social media
Data Model: A model that is used for organizing data elements and how they relate to one
another -> Help to keep data consistent and provide a map of how data is organized
The effects of different structures
Data is everywhere and it can be stored in lots of ways. Two general categories of data are:
Structured data: Organized in a certain format, such as rows and columns.
Unstructured data: Not organized in any easy-to-identify way.
For example, when you rate your favorite restaurant online, you're creating structured data. But
when you use Google Earth to check out a satellite image of a restaurant location, you're using
unstructured data.
Here's a refresher on the characteristics of structured and unstructured data:
Structured data: - Defined data types - Most often quantitative data - Easy to organize - Easy to
search - Easy to analyze - Stored in relational databases - Contained in rows and columns -
Examples: Excel, Google Sheets, SQL, customer data, phone records, transaction history
Unstructured data: - Varied data types - Most often qualitative data - Difficult to search -
Provides more freedom for analysis - Stored in data lakes and NoSQL databases - Can't be put in
rows and columns - Examples: Text messages, social media comments, phone call transcriptions,
various log files, images, audio, video
Structured data
As we described earlier, structured data is organized in a certain format. This makes it easier to
store and query for business needs. If the data is exported, the structure goes along with the data.
Unstructured data can’t be organized in any easily identifiable manner. And there is much more
unstructured than structured data in the world. Video and audio files, text files, social media
content, satellite imagery, presentations, PDF files, open-ended survey responses, and websites
all qualify as types of unstructured data.
The fairness issue
The lack of structure makes unstructured data difficult to search, manage, and analyze. But
recent advancements in artificial intelligence and machine learning algorithms are beginning to
change that. Now, the new challenge facing data scientists is making sure these tools are
inclusive and unbiased. Otherwise, certain elements of a dataset will be more heavily weighted
and/or represented than others. And as you're learning, an unfair dataset does not accurately
represent the population, causing skewed outcomes, low accuracy levels, and unreliable analysis.
Data collected by an individual or group using their own resources – First-party data
Data sold by a provider that didn’t collect the data themselves – Third-party data
Data collected by another group and then sold – Second-party data
Data that is measured and can have almost any numeric value – Continuous data
Data organized in a certain format such as rows and columns – Structured data
Data that is counted and has a limited number of values – Discrete data
Qualitative data that is categorized without a set order – Nominal data
Data that lives within a company’s own systems – Internal data
Qualitative data with a set order or scale – Ordinal data
Data that lives and is generated outside of an organization – External data
Data modeling levels and techniques
Unified Modeling Language (UML): UML diagrams are detailed representations
that describe a system's structure, including entities, attributes, operations, and
their relationships.
Data Modeling: Data modeling is the process of creating diagrams that visually
represent how data is organized and structured, serving as a blueprint for
understanding data relationships.
Entity Relationship Diagram (ERD): ERDs are visual representations that illustrate
the relationships between entities in a data model, aiding in understanding data
structure.
Physical Data Modeling: Physical data modeling depicts the operational aspects of a
database, defining entities, attributes, table names, column names, and data types.
Conceptual Data Modeling: Conceptual data modeling provides a high-level view of
data structure and interactions across an organization, focusing on business
requirements without technical details.
Logical Data Modeling: Logical data modeling details the technical aspects of a
database, including relationships and attributes, but does not specify actual
database table names.
This reading introduces you to data modeling and different types of data models. Data
models help keep data consistent and enable people to map out how data is organized. A
basic understanding makes it easier for analysts and other stakeholders to make sense of
their data and use it in the right ways.
Important note: As a junior data analyst, you won't be asked to design a data model. But
you might come across existing data models your organization already has in place.
What is data modeling?
Data modeling is the process of creating diagrams that visually represent how data is
organized and structured. These visual representations are called data models. You can
think of data modeling as a blueprint of a house. At any point, there might be electricians,
carpenters, and plumbers using that blueprint. Each one of these builders has a different
relationship to the blueprint, but they all need it to understand the overall structure of the
house. Data models are similar; different users might have different data needs, but the
data model gives them an understanding of the structure as a whole.
Levels of data modeling
Each level of data modeling has a different level of detail.
1. Conceptual data modeling gives a high-level view of the data structure, such as how
data interacts across an organization. For example, a conceptual data model may be
used to define the business requirements for a new database. A conceptual data
model doesn't contain technical details.
2. Logical data modeling focuses on the technical details of a database such as
relationships, attributes, and entities. For example, a logical data model defines how
individual records are uniquely identified in a database. But it doesn't spell out
actual names of database tables. That's the job of a physical data model.
3. Physical data modeling depicts how a database operates. A physical data model
defines all entities and attributes used; for example, it includes table names, column
names, and data types for the database.
More information can be found in this comparison of data models.
Data-modeling techniques
There are a lot of approaches when it comes to developing data models, but two common
methods are the Entity Relationship Diagram (ERD) and the Unified Modeling Language
(UML) diagram. ERDs are a visual way to understand the relationship between entities in
the data model. UML diagrams are very detailed diagrams that describe the structure of a
system by showing the system's entities, attributes, operations, and their relationships. As a
junior data analyst, you will need to understand that there are different data modeling
techniques, but in practice, you will probably be using your organization’s existing
technique.
You can read more about ERD, UML, and data dictionaries in this data modeling
techniques article.
Data analysis and data modeling
Data modeling can help you explore the high-level details of your data and how it is related
across the organization’s information systems. Data modeling sometimes requires data
analysis to understand how the data is put together; that way, you know how to map the
data. And finally, data models make it easier for everyone in your organization to
understand and collaborate with you on your data. This is important for you and everyone
on your team!
What type of data is the height of a skyscraper? -> Continuous
In data analytics, what is the term for data that is generated from, and lives, outside of an
organization? -> External
Data type: A specific kind of data attribute that tells what kind of value the data is -> Tells you
what kind of data you’re working with
Data types in spreadsheet
- Number
- Text or string data type: A sequence of characters and punctuation that contains textual
information: Một chuỗi ký tự và dấu câu chứa thông tin văn bản (that
information would be the treats and people's names. These can also include numbers,
like phone numbers or numbers in street addresses. But these numbers wouldn't be used
for calculations. In this case they're treated like text, not numbers)
- Boolean data type: A data type with only two possible value such as TRUE or FALSE
Use Boolean logic
OR Operator: The OR operator allows for the overall statement to be true if at least one
of the conditions is true.
Truth Table: A truth table is a mathematical table used to determine the truth values of
logical expressions based on their inputs.
NOT Operator: The NOT operator negates a condition, meaning it filters out results that
meet the specified condition.
AND Operator: The AND operator requires that both conditions in a statement must be
true for the overall statement to be true.
Boolean Logic: Boolean logic is a form of algebra that uses truth values (true and false)
to create logical statements through operators.
In this reading, you will explore the basics of Boolean logic and learn how to use single and
multiple conditions in a Boolean statement. These conditions are created with Boolean
operators, including AND, OR, and NOT. These operators are similar to mathematical
operators and can be used to create logical statements that filter your results. Data analysts
use Boolean statements to do a wide range of data analysis tasks, such as writing queries for
searches and checking for conditions when writing programming code.
Boolean logic example
Imagine you are shopping for shoes, and are considering certain preferences:
You will buy the shoes only if they are any combination of pink and grey
You will buy the shoes if they are entirely pink, entirely grey, or if they are pink and grey
You will buy the shoes if they are grey, but not if they have any pink
These Venn diagrams illustrate your shoe preferences. AND is the center of the Venn
diagram, where two conditions overlap. OR includes either condition. NOT includes only
the part of the Venn diagram that doesn't contain the exception.
The intersection of these circles is highlighted to indicate the AND condition requires shoes
to be both grey and pink. The Venn diagram that represents OR includes a circle labeled grey
shoes overlapping with a circle labeled pink shoes. The entirety of both circles is highlighted
to indicate the OR condition means any shoe with grey, pink, or some combination satisfies
the requirement. The Venn diagram that represents NOT includes a circle labeled grey shoes
overlapping with a circle labeled pink shoes. The portion of the grey shoes circle that does
not intersect with the pink shoes circle is highlighted to indicate the NOT condition requires
shoes to not include pink.
Use Boolean logic in statements
In queries, Boolean logic is represented in a statement written with Boolean operators. An
operator is a symbol that names the operation or calculation to be performed. Read on to
discover how you can convert your shoe preferences into Boolean statements.
The AND operator
Your condition is “If the color of the shoe has any combination of grey and pink, you will
buy them.” The Boolean statement would break down the logic of that statement to filter
your results by both colors. It would say IF (Color="Grey") AND (Color="Pink") then
buy them
The AND operator lets you stack both of your conditions.
Below is a simple truth table that outlines the Boolean logic at work in this statement. In the
Color is Grey column, there are two pairs of shoes that meet the color condition. And in the
Color is Pink column, there are two pairs that meet that condition. But in the If Grey AND
Pink column, only one pair of shoes meets both conditions. So, according to the Boolean
logic of the statement, there is only one pair marked true. In other words, there is one pair of
shoes that you would buy.
Color is Color is If Grey AND Pink, Boolean Logic
Grey Pink then Buy
Grey/True Pink/True True/Buy True AND True =
Color is Color is If Grey AND Pink, Boolean Logic
Grey Pink then Buy
True
Grey/True Black/False False/Don't buy True AND False
= False
Red/False Pink/True False/Don't buy False AND True
= False
Red/False Green/False False/Don't buy False AND False
= False
The OR operator
The OR operator lets you move forward if either one of your two conditions is met. Your
condition is “If the shoes are grey or pink, you will buy them.” The Boolean statement would
be IF (Color="Grey") OR (Color="Pink") then buy them.
Notice that any shoe that meets either the Color is Grey or the Color is Pink condition is
marked as true by the Boolean logic. According to the truth table below, there are three pairs
of shoes that you can buy.
Color is Color is If Grey OR Pink, Boolean Logic
Grey Pink then Buy
Red/False Black/False False/Don't buy False OR False =
False
Black/False Pink/True True/Buy False OR True =
True
Grey/True Green/False True/Buy True OR False =
True
Grey/True Pink/True True/Buy True OR True =
True
The NOT operator
Finally, the NOT operator lets you filter by subtracting specific conditions from the results.
Your condition is "You will buy any grey shoe except for those with any traces of pink in
them." Your Boolean statement would be IF (Color="Grey") AND (Color=NOT "Pink")
then buy them
Now, all of the grey shoes that aren't pink are marked true by the Boolean logic for the NOT
Pink condition. The pink shoes are marked false by the Boolean logic for the NOT Pink
condition. Only one pair of shoes is excluded in the truth table below.
Color is Color is Boolean If Grey Boolean
Grey Pink Logic for AND (NOT Logic
NOT Pink), then
Pink Buy
Grey/True Red/False Not False True/Buy True
= True AND
True =
True
Grey/True Black/False Not False True/Buy True
= True AND
True =
True
Grey/True Green/False Not False True/Buy True
= True AND
True =
True
Grey/True Pink/True Not True False/Don't True
= False buy AND
False =
False
The power of multiple conditions
For data analysts, the real power of Boolean logic comes from being able to combine
multiple conditions in a single statement. For example, if you wanted to filter for shoes that
were grey or pink, and waterproof, you could construct a Boolean statement such as: “IF
((Color = "Grey") OR (Color = "Pink")) AND (Waterproof="True")
Notice that you can use parentheses to group your conditions together.
Key takeaways
Operators are symbols that name the operation or calculation to be performed. The operators
AND, OR, and NOT can be used to write Boolean statements in programming languages.
Whether you are doing a search for new shoes or applying this logic to queries, Boolean
logic lets you create multiple conditions to filter your results. Now that you know a little
more about Boolean logic, you can start using it!
Resources for more information
Learn about who pioneered Boolean logic in this historical article: Origins of Boolean
Algebra in the Logic of Classes.
Find more information about using AND, OR, and NOT from these tips for searching with
Boolean operators.
When discussing structured databases, data analysts refer to the data contained in a row as
a record. How do they refer to the data contained in a column? -> field
What you’ll need
If you would like to access the spreadsheets the instructor uses in this video, select the link
to a dataset to create a copy. If you don’t have a Google account, download the data
directly from the attachments below.
Link to population datasets:
Population, Latin, and Caribbean Countries, 2010–2019, wide format
Population, Latin, and Caribbean Countries, 2010–2019, long format
Example 1: Examine wide data
Wide data is a dataset in which every data subject has a single row with multiple columns
to hold the values of various attributes of the subject. It is helpful for comparing specific
attributes across different subjects.
1. Open the Population, Latin, and Caribbean Countries, 2010–2019, wide format
spreadsheet.
2. Each row contains all population data for one country.
3. The population data for each year is contained in a column.
4. Find the annual population of Argentina in row 3.
5. In this wide format, you can quickly compare the annual population of Argentina to
the annual populations of Antigua and Barbuda, Aruba, the Bahamas, or any other
country.
Find the country with the highest population in 2010
1. Select column E, which contains each country’s 2010 population data.
2. Right-click column header E and choose Sort Z to A.
3. Notice that Brazil is now at the top of the list because it had the highest population
in the year 2010.
Find the country with the lowest population in 2013
1. Select column H.
2. Right-click column header H and choose Sort A to Z.
3. Notice that the British Virgin Islands are now at the top because they had the lowest
population of all countries in 2013.
Example 2: Examine long data
Long data is data in which each row represents one observation per subject, so each subject
will be represented by multiple rows. This data format is useful for comparing changes
over time or making other comparisons across subjects.
1. Open the Population, Latin, and Caribbean Countries, 2010–2019, long format
spreadsheet.
2. Notice the data is no longer organized into columns by year. All of the years are now
in one column.
3. Find Argentina’s population data in rows 12-21. Each row contains one year of
Argentina’s population data.
Wide data: Data in which every data subject has a single row with multiple columns to hold the
values of various attribute of the subject
Long data: data in which each row is one time point per subject, so each subject will have data in
multiple rows
Transforming data
Data Merging: Data merging involves combining datasets from different sources into a
single dataset, often requiring data transformation for compatibility.
Wide Data: Wide data is structured so that each row contains multiple data points for
particular items, making it easier to read and compare.
Long Data: Long data is structured such that each row contains a single data point for a
particular item, making it suitable for detailed analysis.
Data Organization: Data organization refers to structuring data in a way that enhances its
usability and accessibility for analysis.
Data Transformation: Data transformation is the process of changing the data’s format,
structure, or values to facilitate analysis.
What is data transformation?
A woman presenting data, a hand holding a medal, two people chatting, a ship's wheel being
steered, two people high-fiving each other
In this reading, you will explore how data is transformed and the differences between wide
and long data. Data transformation is the process of changing the data’s format, structure,
or values. As a data analyst, there is a good chance you will need to transform data at some
point to make it easier for you to analyze it.
Data transformation usually involves:
Adding, copying, or replicating data
Deleting fields or records
Standardizing the names of variables
Renaming, moving, or combining columns in a database
Joining one set of data with another
Saving a file in a different format. For example, saving a spreadsheet as a comma
separated values (.csv) file.
Why transform data?
Goals for data transformation might be:
Data organization: better organized data is easier to use
Data compatibility: different applications or systems can then use the same data
Data migration: data with matching formats can be moved from one system to another
Data merging: data with the same organization can be merged together
Data enhancement: data can be displayed with more detailed fields
Data comparison: apples-to-apples comparisons of the data can then be made
Data transformation example: data merging
Mario is a plumber who owns a plumbing company. After years in the business, he buys
another plumbing company. Mario wants to merge the customer information from his newly
acquired company with his own, but the other company uses a different database. So, Mario
needs to make the data compatible. To do this, he has to transform the format of the acquired
company’s data. Then, he must remove duplicate rows for customers they had in common.
When the data is compatible and together, Mario’s plumbing company will have a complete
and merged customer database.
Data transformation example: data organization (long to wide)
To make it easier to create charts, you may also need to transform long data to wide data.
Consider the following example of transforming stock prices (collected as long data) to wide
data.
Long data is data where each row contains a single data point for a particular item. In the
long data example below, individual stock prices (data points) have been collected for Apple
(AAPL), Amazon (AMZN), and Google (GOOGL) (particular items) on the given dates.
Long data example: Stock prices
Wide data is data where each row contains multiple data points for the particular items
identified in the columns.
Wide data example: Stock prices
With data transformed to wide data, you can create a chart comparing how each company's
stock changed over the same period of time.
You might notice that all the data included in the long format is also in the wide format. But
wide data is easier to read and understand. That is why data analysts typically transform long
data to wide data more often than they transform wide data to long data. The following table
summarizes when each format is preferred:
Wide data is preferred when Long data is preferred when
Creating tables and charts with a Storing a lot of variables about each subject. For
few variables about each subject example, 60 years worth of interest rates for each
bank
Comparing straightforward line Performing advanced statistical analysis or
graphs graphing
Terms and definitions for Course 3, Module 1
Agenda: A list of scheduled appointments
Audio file: Digitized audio storage usually in an MP3, AAC, or other compressed format
Boolean data: A data type with only two possible values, usually true or false
Continuous data: Data that is measured and can have almost any numeric value
Cookie: A small file stored on a computer that contains information about its users
Data element: A piece of information in a dataset
Data model: A tool for organizing data elements and how they relate to one another
Digital photo: An electronic or computer-based image usually in BMP or JPG format
Discrete data: Data that is counted and has a limited number of values
External data: Data that lives, and is generated, outside of an organization
Field: A single piece of information from a row or column of a spreadsheet; in a data table,
typically a column in the table
First-party data: Data collected by an individual or group using their own resources
Long data: A dataset in which each row is one time point per subject, so each subject has
data in multiple rows
Nominal data: A type of qualitative data that is categorized without a set order
Ordinal data: Qualitative data with a set order or scale
Ownership: The aspect of data ethics that presumes individuals own the raw data they
provide and have primary control over its usage, processing, and sharing
Pixel: In digital imaging, a small area of illumination on a display screen that, when
combined with other adjacent areas, forms a digital image
Population: In data analytics, all possible data values in a dataset
Record: A collection of related data in a data table, usually synonymous with row
Sample: In data analytics, a segment of a population that is representative of the entire
population
Second-party data: Data collected by a group directly from its audience and then sold
Social media: Websites and applications through which users create and share content or
participate in social networking
String data type: A sequence of characters and punctuation that contains textual information
(Refer to Text data type)
Structured data: Data organized in a certain format such as rows and columns
Text data type: A sequence of characters and punctuation that contains textual information
(also called string data type)
United States Census Bureau: An agency in the U.S. Department of Commerce that serves
as the nation’s leading provider of quality data about its people and economy
Unstructured data: Data that is not organized in any easily identifiable manner
Video file: A collection of images, audio files, and other data usually encoded in a
compressed format such as MP4, MV4, MOV, AVI, or FLV
Wide data: A dataset in which every data subject has a single row with multiple columns to
hold the values of various attributes of the subject