Machine Learning Basics: QA & Profiling
Machine Learning Basics: QA & Profiling
PART 1:
This is Part 1 of a 4-Part series designed to help you build a deep, foundational understanding of
machine learning, including data QA & profiling, classification, forecasting and unsupervised learning
1 ML Intro & Landscape Machine Learning introduction, definition, process & landscape
Tools to explore data quality (variable types, empty values, range &
2 Preliminary Data QA count calculations, table structure, left/right censored data, etc.)
• Anyone who wants to understand • Anyone who would rather copy and
WHEN, WHY, and HOW to deploy paste code than become fluent in the
machine learning tools & techniques underlying algorithms
ML INTRO & LANDSCAPE
noun
*[Link]
COMMON ML QUESTIONS
Which customers are most What will sales look like for What patterns do we see in
likely to churn next month? the next 12 months? terms of product cross-selling?
How can we use online When we adjusted tactics Which product is customer X
customer reviews to monitor last month, did we drive any most likely to purchase next?
changes in sentiment? incremental revenue?
WHEN IS ML THE RIGHT FIT?
QUALITY ASSURANCE
UNIVARIATE PROFILING
MULTIVARIATE PROFILING
MACHINE LEARNING
Quality Assurance (QA) is about preparing & cleaning data prior to analysis. We’ll cover common QA topics
including variable types, empty/missing values, range & count calculations, censored data, etc.
MACHINE LEARNING PROCESS
QUALITY ASSURANCE
UNIVARIATE PROFILING
MULTIVARIATE PROFILING
MACHINE LEARNING
Univariate profiling is about exploring individual variables to build an understanding of your data. We’ll cover
common topics like normal distributions, frequency tables, histograms, etc.
MACHINE LEARNING PROCESS
QUALITY ASSURANCE
UNIVARIATE PROFILING
MULTIVARIATE PROFILING
MACHINE LEARNING
Multivariate profiling is about understanding relationships between multiple variables. We’ll cover common
tools for exploring categorical & numerical data, including kernel densities, violin & box plots, scatterplots, etc.
MACHINE LEARNING PROCESS
QUALITY ASSURANCE
UNIVARIATE PROFILING
MULTIVARIATE PROFILING
MACHINE LEARNING
Machine learning is a natural extension of multivariate profiling, and uses statistical models and methods to
answer questions which are too complex to solve using simple visual analysis or trial-and-error
MACHINE LEARNING LANDSCAPE
MACHINE LEARNING
Clustering/Segmentation
Classification Regression Reinforcement Learning
K-Means (Q-learning, deep RL, multi-armed-bandit, etc.)
QUALITY ASSURANCE
UNIVARIATE PROFILING
MULTIVARIATE PROFILING
MACHINE LEARNING
Quality Assurance (QA) is about preparing & cleaning data prior to analysis. We’ll cover common QA topics
including variable types, empty/missing values, range & count calculations, censored data, etc.
PRELIMINARY DATA QA
Data QA (otherwise known as Quality Assurance or Quality Control) is the first step in
the analytics and machine learning process; QA allows you to identify and correct
underlying data issues (blanks, errors, incorrect formats, etc.) prior to analysis
Are there any missing or Was the data our client Is there any risk that the Are there any outliers
empty values in the data captured from the online data capture process was that might skew the
shared by the HR team? survey encoded properly? biased in some way? results of our analysis?
VARIABLE TYPES
Range Calculations
Common variable types include:
Left/Right Censored
Variable Types Investigating empty values, and how they are recorded in your data,
is a prerequisite for every single analysis
Empty Values
Empty values can be recorded in many ways (NA, N/A, #N/A, NaN,
Null, “-”, “Invalid”, blank, etc.), but the most common mistake is
Range Calculations turning empty numerical values into zeros (0)
Count Calculations
Left/Right Censored
Table Structure
For a missing Retail Price, you would
likely be able to impute the value
since you know the product name/ID
RANGE CALCULATIONS
Variable Types One of the simplest QA tools for numerical variables is to calculate
the range of values in a column (minimum and maximum values)
min(height) = -10
max(height = 10 Is your variable normalized around
Table Structure a central value (i.e. 0)?
COUNT CALCULATIONS
Range Calculations
Distinct counts can be particularly useful for QA, and help to:
• Understand the granularity or “grain” of your data
Count Calculations
• Identify how many unique values a field contains
• Ensure consistency by identifying misspellings or categorization errors
which might otherwise be difficult to catch (i.e. leading or trailing spaces)
Left/Right Censored
PRO TIP: For numerical variables with many unique values (i.e. long decimals), use a
Table Structure histogram to plot frequency based on custom ranges or “bins” (more on that soon!)
LEFT/RIGHT CENSORED
Variable Types
When data is left or right censored, it means that due to some
circumstance the min or max value observed is not the natural
minimum or maximum of that metric
Empty Values • This can be difficult to spot unless you are aware of how the data is being
recorded (which means it’s a particularly dangerous issue to watch out for!)
Count Calculations
Left/Right Censored
Table Structure
Mall Shopper Survey Results Ecommerce Repeat Purchase Rate
Only tracks shoppers over the age of 18 due to legal Sharp drop as you approach the current date has nothing to do
reasons, so anyone under 18 is excluded (even though with customer behavior, but the fact that recent customers
there are plenty of mall shoppers under 18) haven’t have the opportunity or need to repurchase yet
TABLE STRUCTURE
Variable Types Table structures generally come in two flavors: long or wide
Range Calculations
PIVOT
Count Calculations
UNPIVOT
Left/Right Censored
Long tables typically contain a single, distinct column for each field (Date,
Variable Types
Product, Category, Quantity, Profit, etc.)
• Easy to see all available fields and variable types
Empty Values • Great for exploratory data analysis and aggregation (i.e. PivotTables)
Range Calculations Wide tables typically split the same metric into multiple columns or
categories (i.e. 2018 Sales, 2019 Sales, 2020 Sales, etc.)
• Typically not ideal for human readability, since wide tables may contain thousands
Count Calculations of columns (vs. only a handful if pivoted to a long format)
• Often (but not always) the best format for machine learning model input
Left/Right Censored • Great format for visualizing categorical data (i.e. sales by product category)
Table Structure There’s no right or wrong table structure; each type has strengths & weaknesses!
CASE STUDY: PRELIMINARY QA
THE You’ve just been hired as a Data Analyst for Maven Market, a local grocery
SITUATION store looking for help with basic data management and analysis.
The store manager would like you to conduct some analyses on product
THE inventory and sales, but the data is a mess.
ASSIGNMENT You’ll need to explore the data, conduct a preliminary QA, and help clean it up
to prepare the data for further analysis.
Review all fields to ensure that variable types are configured for proper
analysis (i.e. no dates formatted as strings, text formatted as values, etc.)
Remember that NA and 0 do not mean the same thing! Think carefully about
how to handle missing data and the impact it may have on your analysis
Run basic diagnostics like Range, Count, and Left/Right Censored checks
against all columns in your data set...every time
Understand your table structure before conducting any analysis to reduce the
risk of double counting, inaccurate calculations, omitted data, etc.
UNIVARIATE PROFILING
QUALITY ASSURANCE
UNIVARIATE PROFILING
MULTIVARIATE PROFILING
MACHINE LEARNING
Univariate profiling is about exploring individual variables to build an understanding of your data. We’ll cover
common topics like normal distributions, frequency tables, histograms, etc.
UNIVARIATE PROFILING
Univariate profiling is the next step after preliminary QA; think of univariate profiling as
conducting a descriptive analysis of each variable by itself
VARIABLE TYPES
Categorical Variables
Terms like discrete, categorical, multinomial, and classes may all be used interchangeably.
Binary is a special type of categorical variable which takes only 1 of 2 cases: true for false (or 1
Data Profiling or 0) and is also known as a logical variable or a binary flag
DISCRETIZATION
Categorical
Distributions
Numerical Variables
Discretization Rules:
If Price <100 then Price Level = Low
Histograms & If Price >=100 & Price <500 then Price Level = Med
Kernel Densities If Price >=500 then Price Level = High
Normal Distribution
Data Profiling
NOMINAL VS. ORDINAL VARIABLES
VARIABLE TYPES
Categorical Variables
Histograms &
Kernel Densities
There are two types of categorical variables: nominal and ordinal
Normal Distribution • Nominal variables contain categories with no inherent logical rank, which can be
re-ordered with no consequence (i.e. Product Type = Camping, Biking or Hiking)
• Ordinal variables contain categories with a logical order (i.e. Size = Small, Medium,
Data Profiling Large), but the interval between those categories has no logical interpretation
CATEGORICAL DISTRIBUTIONS
Histograms & • Heat maps: Formatted to visualize patterns (typically used for multiple variables)
Kernel Densities
Understanding categorical distributions will help us gather knowledge for
Normal Distribution building accurate machine learning models (more on this later!)
Data Profiling
CATEGORICAL DISTRIBUTIONS
Categorical Variables
Section Distribution:
Camping Biking
Categorical Frequency table
14 6
Distributions
Camping Biking
Proportions table
Numerical Variables 70% 30%
Histograms &
Kernel Densities Size & Section Distribution:
Camping Biking
S 6 4
Normal Distribution Heat Map
L 8 2
Data Profiling
NUMERICAL VARIABLES
VARIABLE TYPES
Categorical Variables
You may hear numeric variables described further as interval and ratio, but the distinction is trivial
and rarely makes a difference in common use cases
Data Profiling
HISTOGRAMS
Categorical Variables
Histograms are used to plot a single, discretized numerical variable
Numerical Variables
Age Values:
8 29 45 8
Histograms & 25 33 37
7
Frequency
Kernel Densities 6
19 43 21 5
28 32 40 4
Normal Distribution 24 17 28 3
2
5 22 39
1
15 47 12
Data Profiling 0-10 11-20 21-30 31-40 41-50
Age Range
KERNEL DENSITIES
Categorical Variables
Kernel densities are “smooth” versions of histograms, which can help
to prevent users from over-interpreting breaks between bins
Numerical Variables
Age Values:
8 29 45 8
Histograms & 25 33 37
7
Frequency
Kernel Densities 6
19 43 21 5
28 32 40 4
Normal Distribution 24 17 28 3
2
5 22 39
1
15 47 12
Data Profiling 0-10 11-20 21-30 31-40 41-50
Age Range
HISTOGRAMS & KERNEL DENSITIES
PRO TIP: If your data is relatively symmetrical (not skewed), you can use Sturge’s Rule as a
Data Profiling quick “rule of thumb” to determine an appropriate number of bins: K = 1 + 3.322 log (N)
(where K = number of bins, N = number of observations)
CASE STUDY: HISTOGRAMS
THE You’ve just been promoted as the new Pit Boss at The Lucky Roll Casino.
SITUATION Your mission? Use data to help expose cheats on the casino floor.
Profits at the craps tables have been unusually low, and you’ve been asked to
THE investigate the possible use of loaded die (weighted towards specific numbers).
ASSIGNMENT Your plan is to track the outcome of each roll, then compare your results
against the expected probability distribution to see how closely they match.
Histograms &
Kernel Densities
Normal Distribution
Data Profiling
You can find normal distributions in many real-world examples: heights, weights, test scores, etc.
NORMAL DISTRIBUTION
Categorical
Distributions
1 1 𝑥−𝜇 2
Numerical Variables −
𝑒 2 𝜎
Histograms &
Kernel Densities 𝜎 2𝜋 Turn the parabola
upside down
Normal Distribution
Make the tails
flare out
Data Profiling
CASE STUDY: NORMAL DISTRIBUTION
THE It’s August 2016, and you’ve been invited to Rio de Janeiro as a Data Analyst
SITUATION for the Global Olympic Committee.
Your job is to collect demographic data for all female athletes competing in
THE
the Summer Games and determine how the distribution of Olympic athlete
ASSIGNMENT heights compares against the general public.
1. Gather heights for all female athletes competing in the 2016 Games
THE 2. Plot height frequencies using a Histogram, and test various bin widths
OBJECTIVES 3. Determine if athlete heights follow a normal distribution, or “bell curve”
4. Compare the distributions for athletes vs. the general public
DATA PROFILING
Numerical Variables
Mode of City = “Houston”
Mode of Sessions = 24
Histograms &
Kernel Densities Mode of Gender = F, M
(this is a bimodal field!)
Normal Distribution
Common uses:
Data Profiling • Understanding the most common values within a dataset
• Diagnosing if one variable is influenced by another
MODE
Categorical Variables While modes typically aren’t very useful on their own, they can provide
helpful hints for deeper data exploration
Categorical • For example, the right histogram below shows a multi-modal distribution, which
Distributions indicates that there may be another variable impacting the age distribution
8 8
Histograms & 7 7
Kernel Densities 6 6
Frequency
Frequency
5 5
4 4
Normal Distribution 3 3
2 2
1 1
0-10 11-20 21-30 31-40 41-50 0-10 11-20 21-30 31-40 41-50
Data Profiling
Age Range Age Range
MEAN
Categorical Variables
The mean is the calculated “central” value in a discrete set on numbers
• Mean is what most people think of when they hear the word “average”, and is
calculated by dividing the sum of all values by the count of all observations
Categorical
Distributions • Means can only be applied to numerical variables (not categorical)
Numerical Variables
𝑠𝑢𝑚 𝑜𝑓 𝑎𝑙𝑙 𝑣𝑎𝑙𝑢𝑒𝑠
𝑚𝑒𝑎𝑛 =
𝑐𝑜𝑢𝑛𝑡 𝑜𝑓 𝑜𝑏𝑠𝑒𝑟𝑣𝑎𝑡𝑖𝑜𝑛𝑠
Histograms &
Kernel Densities 5,220
=
5
= 𝟏, 𝟎𝟒𝟒
Normal Distribution
Common uses:
• Making a “best-guess” estimate of a value
Data Profiling
• Calculating a central value when outliers are not present
MEDIAN
The median is the middle value in a list of values sorted from highest to
Categorical Variables
lowest (or vice versa)
• When there are two middle-ranked values, the median is the average of the two
Categorical
• Medians can only be applied to numerical variables (not categorical)
Distributions
Numerical Variables
Normal Distribution
Common uses:
Data Profiling • Identifying the “center” of a distribution
• Calculating a central value when outliers may be present
PERCENTILE
Numerical Variables
Histograms &
Kernel Densities
Numerical Variables
Variance = 5
Normal Distribution
Common uses:
Data Profiling
• Comparing the numerical distributions of two different groups (i.e. prices of products
ordered online vs. in store)
VARIANCE
Categorical Variables 𝑛 2
σ𝑖=1(𝑥𝑖 − 𝜇)
Categorical
Distributions 𝑛−1
Numerical Variables Calculation Steps:
Histograms &
Kernel Densities
Common uses:
Data Profiling • Comparing segments for a given metric (i.e. time on site for mobile users vs. desktop)
• Understanding how likely certain values are bound to occur
SKEWNESS
Categorical Variables Skewness tells us how a distribution varies from a normal distribution
• This is commonly used to mathematically describe skew to the left or right
Categorical
Distributions
Left skew Normal Distribution Right skew
Numerical Variables
Histograms &
Kernel Densities
Normal Distribution
Common uses:
Data Profiling
• Identifying non-normal distributions, and describing them mathematically
BEST PRACTICES: UNIVARIATE PROFILING
Make sure you are using the appropriate tools for profiling categorical
variables vs. numerical variables
QA still comes first! Profiling metrics are important, but can lead to
misleading results without proper QA (i.e. handling outliers or missing values)
QUALITY ASSURANCE
UNIVARIATE PROFILING
MULTIVARIATE PROFILING
MACHINE LEARNING
Multivariate profiling is about understanding relationships between multiple variables. We’ll cover common
tools for exploring categorical & numerical data, including kernel densities, violin & box plots, scatterplots, etc.
MULTIVARIATE PROFILING
Multivariate profiling is the next step after univariate profiling, since single-metric
distributions are rarely enough to draw meaningful insights or conclusions
Categorical-Numerical
Distributions This is one of the simplest forms of multivariate profiling, and leverages
the same tools we used to analyze univariate distributions:
Multivariate Kernel
Densities • Frequency tables: Show the count (or frequency) of each distinct combination
• Proportions tables: Show the count of each combination as a % of the total
Violin & Box Plots • Heat maps: Frequency or proportions table formatted to visualize patterns
THE You’ve just been hired by the New York Department of Transportation
SITUATION (DOT) to help analyze traffic accidents in New York City from 2019-2020
THE 1. Create a table to plot accident frequency by time of day and day of week
OBJECTIVES 2. Apply conditional formatting to the table to create a heatmap showing the
days and times with the fewest (green) and most (red) accidents in the sample
CATEGORICAL-NUMERICAL DISTRIBUTIONS
Multivariate Kernel Teal class has a mean of ~15 and relatively low
Densities variance (highly concentrated around the mean)
Violin & Box Plots Yellow class has a mean of ~20 and moderate
variance relative to other categories
Numerical-Numerical
Distributions Purple class has a mean of ~25, overlaps with
yellow, and has relatively high variance
Numerical-Numerical
Distributions
Categorical profiling works for simple cases, but breaks down quickly
Categorical-Categorical
Distributions • Humans are pretty good at visualizing 1, 2, or maybe even 3 variables, but how
would you visualize a joint distribution for 10 variables? 100?
Categorical-Numerical
Distributions
Multivariate Kernel
Densities
?
Violin & Box Plots Categorical profiling can’t answer prescriptive or predictive questions
• Suppose you randomized several elements on your sales page (font, image, layout,
button, copy, etc.) to understand which ones drive conversions
Numerical-Numerical
Distributions • You could count conversions for individual elements, or some combinations of
elements, but categorical distribution alone can’t measure causation
Scatter Plots &
Correlation
This is when you need machine learning!
NUMERICAL-NUMERICAL DISTRIBUTIONS
Categorical-Numerical They are typically visualized using scatter plots, which plot points along
Distributions
the X and Y axis to show the relationship between two variables
Multivariate Kernel • Scatter plots allow for simple, visual intuition: when one variable increases or
Densities decreases, how does the other variable change?
• There are many possibilities: no relationship, positive, negative, linear, non-linear,
Violin & Box Plots cubic, exponential, etc.
Numerical-Numerical
Distributions
Multivariate Kernel
Densities
σ𝑛𝑖=1(𝑥𝑖 − 𝜇) 2 σ𝑛𝑖=1(𝑥𝑖 − 𝑥)(𝑦𝑖 − 𝑦)
Violin & Box Plots 𝑛−1 (𝑛 − 1)𝑠𝑥 𝑠𝑦
Numerical-Numerical
Distributions • Here we multiply variable X’s difference from its mean with variable Y’s difference
from its mean, instead of squaring a single variable (like we do with variance)
Scatter Plots &
Correlation • Sx and Sy are the standard deviations of X and Y, which puts them on the same scale
CORRELATION VS. CAUSATION
Categorical-Categorical
CORRELATION
Distributions
Categorical-Numerical
Distributions
Multivariate Kernel
Densities DOES NOT IMPLY
Violin & Box Plots
Numerical-Numerical
Distributions CAUSATION
Scatter Plots &
Correlation
CORRELATION VS. CAUSATION
Categorical-Categorical
Drowning Deaths
Distributions
Categorical-Numerical
Distributions
Multivariate Kernel
Densities
Ice Cream Cones Sold
Violin & Box Plots Consider the scatter plot above, showing daily ice cream sales and
drowning deaths in a popular New England vacation town
Numerical-Numerical • These two variables are clearly correlated, but do ice cream cones CAUSE people to
Distributions drown? Do drowning deaths CAUSE a surge in ice cream sales?
Categorical-Categorical Scatter plots show two dimensions by default (X and Y), but using
Distributions symbols or color allows you to visualize additional variables and
expose otherwise hidden patterns or trends
Categorical-Numerical
Distributions
Multivariate Kernel
Densities
Numerical-Numerical
Distributions
THE You’ve just landed your dream job as a Marketing Analyst at Loud & Clear, the
SITUATION hottest ad agency in San Diego.
Your client would like to understand the impact of their digital media spend, and
THE how it relates to website traffic, offline spend, site load time, and sales.
ASSIGNMENT Your role is to collect and visualize these metrics at the weekly-level in order to
begin exploring the relationships between them.
Use categorical variables to filter or “cut” your data and quickly compare
profiling metrics or distributions across classes
Remember that correlation does not imply causation, and that variables can
be related without one causing a change in the other
LOOKING AHEAD
CONGRATULATIONS!
Now that you’ve completed Part 1: QA & Data Profiling, you should have a strong grasp of QA
techniques, univariate & multivariate distributions, and common data profiling metrics.
In Part 2 we’ll dive into supervised machine learning and explore powerful classification
techniques like K-Nearest Neighbors, Naïve Bayes, Decision Trees, Logistic Regression,
Sentiment Analysis and more.