Article
Article
We thank Gábor Békés, Erin Conlon, Ezgi Cengiz, Ina Ganguli, Laura Giuliano,
Carl Nadler, Hasan Tekgüç, Michael Reich, Jesse Rothstein, and participants at the
S203
S204 Cengiz et al.
I. Introduction
A long-standing question in economics centers around understanding how
minimum wages affect low-wage labor markets. A key challenge to convinc-
ingly answer this question comes from the difficulty in successfully identify-
ing most workers who are actually affected by the policy. While we can easily
locate workers who are currently earning the minimum wage, it is difficult to
identify all potential workers who also may have been working had the min-
imum wage been different. This difficulty has led many researchers to focus
on specific industries or demographic groups, such as teens (Card 1992; Neu-
mark and Wascher 1992; Giuliano 2013; Neumark, Salas, and Wascher 2014;
Allegretto et al. 2017; Totty 2017), younger workers with lower educational
credentials (Sabia, Burkhauser, and Hansen 2012; Manning 2016; Clemens
and Strain 2017; Clemens and Wither 2019), and individuals without a high
school degree (Addison and Blackburn 1999; Addison, Blackburn, and Cotti
2011). However, these groups constitute relatively small shares of all mini-
mum wage workers. As a result, there is a tension between what is often an-
alyzed (e.g., minimum wage effects on teens) and what is argued (effects of
the policy on affected workers largely composed of adults; Belman, Wolfson,
and Nawakitphaitoon 2015; Manning 2016).1
In this paper, we use machine learning (ML) tools to predict which individ-
uals were likely affected by minimum wage increases and then estimate the
Labor and Employment Relations Association 70th Annual Meeting, Institute for
Research on Labor and Employment Research Presentation seminar, 44th Eastern
Economic Association Conference, 2018 New School–University of Massachusetts
Economics Graduate Student Workshop, and the authors’ conference in honor of
Alan Krueger for very helpful comments. We are also grateful to Jon Piqueras for
outstanding research assistance. The previous version of the paper was circulated un-
der the title “Seeing Beyond the Trees: Using Machine Learning to Estimate the Im-
pact of Minimum Wages on Affected Individuals.” Attila Lindner acknowledges fi-
nancial support from the Economic and Social Research Council (new investigator
grant, ES/T008474/1) and from the European Research Council (ERC) under the
European Union’s Horizon 2020 Research and Innovation Programme (grant agree-
ment 949995). David Zentler-Munro acknowledges financial support from the Eco-
nomic and Social Research Council (ESRC) under grant ES/R005745/1. Contact the
corresponding author, Arindrajit Dube, at adube@[Link]. Information concern-
ing access to the data used in this paper is available as supplemental material online.
1
The discrepancy is particularly relevant when the measured outcome is the teen
employment rate, which has been the subject of extensive research in the United
States. Belman and Wolfson (2014) consider 30 studies that examined the employment
effects of the minimum wage on some demographic groups between 2001 and early
2013, and they find that 17 of them had teen employment as the dependent variable.
Neumark (2017) shows that 12 of 13 studies that examined minimum wage effects on
“lower-skilled” employment between 2010 and 2016 focused on teens (see his table 1).
However, teens are less likely to be in the affected group than nonteen adult minimum
wage workers (Lundstrom 2016). Compared with affected nonteens, only a relatively
small share of teens live in poverty. According to the 2016 American Community Sur-
vey, 18.4% of teens were in households with incomes under the poverty level.
Seeing beyond the Trees S205
2
We are aware of only one previous publication that applied this method (Cengiz
et al. 2019), and the list of coauthors includes three of the authors of this paper. That
paper utilized the Card and Krueger prediction-based approach primarily to show
the differences between that method and the bunching method developed in Cengiz
et al. (2019).
S206 Cengiz et al.
3
In the prediction exercise, we restrict the sample to states instituting prominent
minimum wage hikes and periods preceding those policy changes. The full set of
predictors and how they are coded are reported in app. B (apps. A–E are available
online). In the prediction model, we do not use variables related to past employ-
ment status or occupation/industry in the CPS Outgoing Rotation Group (CPS-
ORG). They are sometimes missing even if the individual is currently in the labor
force and looking for a job. We prefer to keep the observations with missing infor-
mation on these variables in the sample, as they potentially carry information about
labor supply effects of the policy. In our preferred prediction model, we do not use
state of residence or year information either. This choice is primarily to be able to
build samples that are comparable and consistent across time and space.
Seeing beyond the Trees S207
workers who are classified correctly) are sizeable when we limit attention to
nonteen workers—a group that is of particular interest to policy makers.
Armed with the prediction model, we implement an event study analysis
that exploits 172 prominent state-level minimum wage increases between
1979 and 2019. We assess the impact of the policy on various groups formed
on the basis of the predicted exposure probability. The high-probability
group comprises the 10% of the population with the highest likelihood of
being affected by the policy. We also study the impact of the policy on the
high-recall group, which captures 75% of all minimum wage workers.
For both groups, we find a considerable increase in wages after the policy
change; as expected, the wage increase is somewhat lower for the high-recall
group. At the same time, we detect a small, positive, statistically insignificant
effect on employment for both groups. The implied employment elasticity
with respect to own wage—the labor demand elasticity in the standard com-
petitive model of the labor market —is 0.29 (SE, 0.32) for the high-probability
group and 0.14 (SE, 0.25) for the high-recall group. The confidence bounds on
both of these estimates can rule out anything more than modest negative
disemployment effects at the conventional significance levels.
We find no evidence of substantial changes in the unemployment or par-
ticipation rates in response to the policy. We are also not able to detect any
economically meaningful (or statistically significant) changes in labor mar-
ket transitions between employment, unemployment, and nonparticipation.
This lack of response on the LFP margin provides new evidence that mini-
mum wages have a limited impact on search effort when we focus on indi-
viduals who are most likely to be affected by the policy.
Our results are robust to controlling for time-varying heterogeneity in a
wide variety of ways. Moreover, the increase in wages lines up well with the
timing of the minimum wage increases, and the effects emerge only in the group
of individuals likely to be exposed to the minimum wage increase. We find
no significant differences in labor market outcomes for the low-probability
group—suggesting that no unusual changes took place in the states’ labor
markets around the minimum wage increases we study here. All of these find-
ings reinforce the credibility of our research design.
Furthermore, we also study whether the responses to the policy vary
across demographic groups. Most importantly, we study whether differen-
tial responses can be detected on employment, unemployment, and partici-
pation margins for workers who are thought to have larger extensive margin
labor supply elasticities—such as teens, older workers, and single mothers.
In addition, we also assess the impact of the policy by the likelihood of mov-
ing into or out of the labor force. We use demographic information and ap-
ply ML tools again to classify workers as being more likely to move into and
out of the labor force. Even when we focus on the group of workers with the
highest predicted transition probabilities, we find no evidence of substantial
change in the unemployment or participation rates.
S208 Cengiz et al.
4
See Belman and Wolfson (2014) for a thorough literature review on the subject.
5
Assigning workers according to their baseline wages requires panel data. Panel data
sets, such as the National Longitudinal Survey of Youth (Currie and Fallick 1996) and
Survey of Income and Program Participation (Clemens and Wither 2019), are often
smaller than the CPS applied here and cover fewer years. The prediction probability ap-
proach can be applied to cross-sectional data, so we can include many more prominent
minimum wage changes in our analysis.
Seeing beyond the Trees S209
III. Data
The primary data we use throughout the analysis come from the CPS.
We use the 1979–2019 CPS-ORG sample for the hourly wage and weekly
earnings variables. This is a subset of the Basic Monthly CPS (CPS-Basic), a
monthly survey of approximately 60,000 households in the United States. The
Seeing beyond the Trees S211
CPS-ORG includes only the fourth and eighth sample months, when usual
hourly wages, weekly earnings, and weekly hours worked are asked. These
variables are of primary importance for the prediction as well as for the esti-
mation, and thus we rely on their accuracy. For this reason, we exclude obser-
vations with imputed hourly wages, imputed weekly earnings, or imputed
hours worked. For hourly workers we use the reported hourly wage, and
for other workers we define the hourly wage to be their usual weekly earnings
divided by usual weekly hours. We also use a range of demographic variables
in the data set when predicting an individual’s likelihood of having a wage
close to the minimum. These variables indicate an individual’s age, race, His-
panic status, gender, education, veteran status, marital status, and rural status
of the residency (for the exact definitions, see app. B).
We use the 1979–2019 CPS-Basic files for the employment, unemployment,
and LFP variables as well as for a number of secondary variables describing the
nature of employment (part-time, overtime, and self-employment). Unlike the
CPS-ORG, CPS-Basic contains observations for every month that a respon-
dent is surveyed. Therefore, using the CPS-Basic to estimate employment, un-
employment, and LFP effects of the minimum wage results in greater precision
than using the CPS-ORG to estimate these effects. It also allows us to estimate
the impacts of the minimum wage on transitions between employment, unem-
ployment, and inactivity. We obtain the minimum wage data from Vaghul and
Zipperer (2016), which has been extended through 2019 by the authors.
6
Our results are not sensitive to the definition of minimum wage workers based
on alternative cutoff values. Setting the thresholds to 3% above the minimum wage
or 200% above the minimum wage produces virtually the same ordering of obser-
vations according to predicted probabilities, suggesting that the specific definition
that we use has essentially no bearing on the conclusions.
S212 Cengiz et al.
and (2) there is a prominent minimum wage change in the next 12 quarters.
The former criterion ensures that we are not training the model using work-
ers in states/quarters where the wage distribution may not have stabilized
following a minimum wage event. The latter criterion ensures that we are
training the model on workers who will experience a minimum wage event
in the near future and are therefore pertinent to our analysis. There are
469,174 worker-level observations in the CPS that satisfy this screen be-
tween 1979 and 2019.
Second, we divide the 469,174 observations into two mutually exclusive
samples: a training sample and a test sample. To create the training sample
we randomly draw 150,000 observations. We apply various learning tools,
such as random forests, tree boosting, basic logistic, elastic net, and the linear
probability model along the lines of Card and Krueger (1995), and fit each
model on the training sample. In the next section we describe the key idea be-
hind each prediction algorithm. For further details on these prediction models,
we refer the reader to appendix C and Friedman, Hastie, and Tibshirani (2009).
The test sample is composed of the complement of the training sample.
We use the test sample to compare the performance of the prediction models
by plotting the precision-recall curves (explained below) along with other
descriptive statistics.
Once we have optimized over the prediction models, we use the preferred
(best-performing) model to calculate the predicted probability in the full
data set that includes all time periods and states between 1979 and 2019.
We use that full data set to investigate the causal effects of the minimum wage
on various predicted probability groups.
A. Prediction Algorithms
Decision Trees
A single decision tree lies at the heart of many learning techniques, including
random forests and gradient boosting. A decision tree recursively divides the
feature (predictor) space into two in a way that reduces the prespecified loss
function the most.7 More concretely, in the beginning the algorithm tries every
possible split to divide the entire sample space into two and picks the one that
diminishes the loss function the most. Subsequently, each subsample is treated
as the new sample, and the first step is repeated. Once the splitting is over, it
predicts the class of every observation according to the majority vote in the
subspace (terminal node) to which the observation belongs.
This procedure requires a decision on when to stop the splitting. In prin-
ciple, the splitting could continue until there is only one data point at each
7
The loss function is the deviance, defined as 22om ok nmk logð p^mk Þ, where nmk
indicates the number of observations at terminal node m that belongs to class k and
^pmk is the share of observations at terminal node m that belongs to class k.
Seeing beyond the Trees S213
terminal node. Such a tree fits the training sample perfectly but would suffer
from overfitting. To overcome the problem, it is common to use cross val-
idation to determine the complexity of the tree. For a more accurate predic-
tion, we collapse some internal nodes (“prune the tree”) and decrease the
prediction variance at the expense of bias.
Decision trees are not among the most successful learners, yet they are
relatively easy to interpret. In figure 1 we plot a pruned decision tree pro-
duced to predict whether a worker has an hourly wage of less than 125% of
the statutory minimum wage using demographic and educational character-
istics. The tree predicts that the only group in the training sample with
hourly wages less than the threshold is the one with those who are 19 years
old or younger. The majority vote in all of the other terminal nodes is
“false,” indicating that nonteen observations are expected to work for hourly
wages higher than the threshold.
It is noteworthy that the recommendation based on a simple decision tree
is to proxy minimum wage workers with teens—which happens to be the
most common approach taken in the literature. However, as we show below,
it is possible to obtain much better predictions by combining multiple deci-
sion trees. The two most common ways to do so are the random forest by
Breiman (2001) and the gradient-boosting trees by Friedman (2001).
Random Forest
The random forest is a tree-based ensemble learning technique. It provides
a way to overcome the bias-variance trade-off of a single tree. It constructs a
multitude of fully grown decision trees formed using different training boot-
strap samples. Each tree produces unbiased predictions that have large var-
iances. We calculate the average of the predictions, thereby reducing the var-
iance. To further reduce the variance, we decrease the correlation among
trees by employing a randomly selected portion of the predictors at each
split. Although this results in the loss of the interpretability of individual
trees, it has no impact on the bias, since individual trees are still fully grown.8
Our fivefold cross validation finds that the optimum random forest is
achieved with 2,000 trees and only two predictors tried at each split.
Boosting
Boosting approaches the problem of how to combine multiple trees from
a different angle. Instead of producing many fully grown trees and averaging
them, the trees in this model are grown sequentially where subsequent trees
attempt to fix the errors of the preceding ones. As a result, while the first tree
8
Note that if trees are perfectly correlated, the reduction of the variance would
be nil. If they are independent, the variance of the final model would be j2 =B, where
j2 is the prediction variance of a single tree and B indicates the number of trees.
S214 Cengiz et al.
FIG. 1.—Minimum wage workers according to pruned trees. The figure plots a
pruned decision tree produced to predict whether an individual is a minimum wage
worker (has an hourly wage of less than 125% of the statutory minimum wage) using
demographic and educational characteristics. In the beginning the tree tries every pos-
sible split to divide the entire sample space into two and picks the one that diminishes
the loss function the most. Then each subspace is treated as the new feature space, and
the first step is repeated. “True” indicates that the tree predicts that workers in the ter-
minal node are minimum wage workers and “false” indicates otherwise. While the tree
explores characteristics such as gender, marital status, veteran status, and rural residency
status, it picks only age and education to make splits. LTHS 5 less than high school;
HSG 5 high school graduate; SC 5 some college; CG 5 college graduate.
the outcome variable (e.g., using the residual as the outcome variable) or slightly
changing the loss function (weighting the misclassified observations more
heavily). After building the subsequent tree, we combine the predictions of
all trees through a weighted majority vote. On the basis of our fivefold cross
validation, the optimum boosted tree model is obtained with the following
parameters: number of trees 5 4,000, shrinkage factor 5 0:005, depth of
tree 5 6, and minimum observations in a node 5 10.
Elastic Net
We use the elastic net regularization developed by Zou and Hastie (2005).
The underlying model is very similar to the logistic regression except that the
elastic net model penalizes model complexity. The penalty term is a linear
combination of the lasso and ridge methods; lasso tends to drop poor predic-
tors, while ridge tends to shrink their coefficients toward zero, so elastic net
combines both. As opposed to tree-based models, the elastic net regression re-
quires prespecification of the exact functional form for the predictors in the
prediction equation. Therefore, we purposefully build a fairly complex model,
where we include all of the predictors, their four-way interactions, and all of
the interactions with the quadratic, cubic, and quartic terms of the age variable.
We rely on the regularization to simplify the model and prevent overfitting.
9
We also tried to implement neural networks and support vector machines. While
the model constructed using the neural networks performs slightly worse than the
boosted tree model, the models using the support vector machines fail to provide a
well-performing prediction model.
S216 Cengiz et al.
“Recall” refers to the share of true minimum wage workers who we correctly
classify as being in the predicted group. For instance, if a predicted group has
only one observation and the observation is a true minimum wage worker, then
the precision is 1; however, here the recall is very small, as the sample will cover
only a minuscule fraction of the minimum wage workers in the population. On
the other hand, if the predicted group contains every observation in the popu-
lation, then the recall rate is 1, as the sample, by construction, includes all min-
imum wage workers in the population. However, here the precision is going to
be small, since the predicted group also includes all the non-minimum-wage
workers in the population. The ideal is to construct a predicted group that in-
cludes all the minimum wage workers and none of the non-minimum-wage
workers so that both the precision and the recall are 1. Generally, the higher
the precision for a given recall rate, the better the performance of the model.10
Figure 2A shows the precision-recall curves corresponding to the various
prediction algorithms. We also estimate and report the performance of a ba-
sic logistic model with age and the categorical education variables as predic-
tors for comparison. To plot the curve, we calculate the predicted probabil-
ities for each individual in the test sample. We then define the predicted
group for alternative probability thresholds, where all workers in the group
have a predicted probability greater than the threshold. We calculate the pre-
cision and the recall for each of these groups and obtain the curve. In other
words, each point on the curves corresponds to a separate predicted group.
When we raise the threshold, we expect the precision to increase but at the
cost of a reduced recall rate. How strong this trade-off is between the preci-
sion and recall rates for various prediction models is shown in the figure.
The figure shows that the boosted tree model (solid line) outperforms other
prediction models, since it provides the highest precision at almost all recall
levels. For comparison, in figure 2B we report the other prediction models rel-
ative to the boosted tree model. The boosted tree model (and also the other
prediction algorithms) improves precision considerably relative to the basic
logistic model. Nevertheless, the differences between the other prediction
models and the boosted tree model are relatively small, especially at higher
recall rate levels. The random forest model achieves almost the same result
as the boosted tree model. It is also notable that the Card and Krueger sub-
jective judgment approach does almost as well as the elastic regularization of
the logistic model, and the performance of their model is not far behind the
best-performing prediction model.11
10
Another approach commonly used to compare models is to plot the receiver operat-
ing characteristic (ROC) curve. The ROC curve plots the recall against the false positive
rate, the latter defined as the number of non-minimum-wage observations as a proportion
of the number of non-minimum-wage workers in the population. In our case, we reach the
same conclusion whether we use the ROC curve or the precision-recall curve.
11
An alternative way to assess model performance is to compare the fraction of true
minimum wage workers in each predicted probability decile. Table A.1 (tables A.1–A.5,
Seeing beyond the Trees S217
C.1, E.1 are available online) shows that the boosted tree model has a slightly higher
fraction of true minimum wage workers in the most likely predicted deciles and a
lower fraction in the least likely predicted deciles. This provides further support of
the slightly better performance of the boosted tree model.
12
Of course, it is possible that someone is directly interested in the impact of the
policy on the labor market outcomes of certain demographic groups or industries.
Nevertheless, in most cases researchers pick specific subgroups (e.g., teens) or sec-
tors (e.g., restaurants) not because they are the main subjects of interest but because
these are subgroups where the fraction of minimum wage workers is high. Further-
more, the prediction approach can also be applied if someone is specifically inter-
ested in the impact of the minimum wage on some subgroups (see table 5).
FIG. 2.—Precision-recall curves. A plots the precision-recall curves for various
prediction models described in section IV.A and for a basic logistic model that we
estimate using (linear) age and categorical education variables. We obtain the precision-
recall curves in the following way: we use our prediction model to calculate the prob-
ability that someone is a minimum wage worker, and we assign all individuals to the
predicted group if that probability is above a certain threshold. The figure shows
the estimated precision and recall rates obtained when we vary the threshold value.
The horizontal dotted line shows the average share of minimum wage workers in the
sample. The areas below the precision-recall curves are 0.449 for boosted tree, 0.445 for
random forest, 0.443 for elastic net, 0.435 for Card and Krueger’s linear probability
model, 0.342 for basic logistic, and 0.269 for single tree. B shows the difference in pre-
cision rate between the best-performing model—the boosted tree model—and the
other models at each recall rate. A color version of this figure is available online.
Seeing beyond the Trees S219
Table 1
Demographic Characteristics for Each Predicted Probability Decile
Black or
Teen 20 ≤ Age < 30 LTHS HSG Female White Hispanic
(1) (2) (3) (4) (5) (6) (7)
Most likely decile .719 .038 .752 .145 .592 .837 .244
Probability decile 9 .047 .405 .534 .238 .674 .847 .359
Probability decile 8 .004 .341 .344 .437 .594 .834 .243
Probability decile 7 .004 .298 .187 .575 .571 .833 .351
Probability decile 6 .000 .191 .085 .660 .673 .873 .150
Probability decile 5 .000 .187 .100 .475 .492 .784 .253
Probability decile 4 .000 .178 .067 .236 .512 .794 .237
Probability decile 3 .000 .162 .004 .297 .404 .865 .175
Probability decile 2 .000 .088 .000 .143 .385 .848 .122
Least likely decile .000 .015 .000 .039 .314 .741 .134
NOTE.—The table shows some demographic characteristics for each predicted probability decile. The pre-
dicted probability refers to the probability that an individual has an hourly wage lower than 125% of the min-
imum wage preceding the minimum wage hike and is calculated according to the best-performing prediction
model—the boosted tree model. Each row shows the average characteristics of individuals in the particular
predicted probability decile. The top (bottom) row shows the characteristics at the top (bottom) decile, which
consists of individuals that are most (least) likely exposed to the minimum wage according to our prediction
model. Each cell shows the share of the selected demographic group: col. 1, the share of teens (i.e., those youn-
ger than 20); col. 2, the share of individuals who are between 20 and 30 years of age; col. 3, the share of indi-
viduals with less than high school (LTHS) education; col. 4, the share of high school graduates (HSGs) with no
college education; col. 5, the share of females; col. 6, the share of white individuals; and col. 7, the share of
Black or Hispanic individuals.
without some college education in the lowest deciles. Another finding worth
noting is that the share of women workers is high in the top deciles (e.g.,
67.4% in decile 9) and is lower in the bottom deciles (31.4% in the least likely
decile), indicating that an individual’s gender also plays an important role.
Last, Black/Hispanic individuals’ share in the least likely decile is 13.4%,
whereas they make up at least 24% of the top two deciles.
An alternative way to examine who are the minimum wage workers is to
consider the relative importance of each predictor. In figure 4 we plot the “rel-
ative influences” of the variables calculated following Friedman (2001).13 The
The figure shows the reduction in the loss function caused by each variable
13
used in the nonterminal nodes. We normalize the relative influences so that they
sum up to 100. The average importance of each variable is
1 M 2
I 2l 5 o I l ðTm Þ,
M m51
where I l ðTm Þ is the reduction in the loss function due to the use of variable l in the
nonterminal nodes of tree m. However, we wish to caution against interpreting
them directly. First, the importance is in terms of prediction, not explanation. Sec-
ond, there are cases where one variable needs to be interacted with another one for
high predictive power. In those cases, only one of them is deemed to have a strong
influence, whereas both are essential.
Seeing beyond the Trees S221
FIG. 4.—Relative influences of the predictors in the boosted tree prediction model.
We plot relative influences of the variables in the best-performing prediction model—
the boosted tree model—calculated as in Friedman (2001; for details, see n. 13). The
bars, which indicate the decline in the loss function associated with the corresponding
variable, are normalized so that they sum up to 100.
figure largely confirms our previous observations. It finds age as the most im-
portant predictor in the sample with a very large margin. The variable for ed-
ucational credentials comes after age.14 Gender variables are also relatively im-
portant in the prediction. The indicator variables for Hispanic, rural, race, and
veteran status appear to have less influence on the prediction.
14
In fact, dropping teen observations from the sample decreases the relative im-
portance of the age variable substantially. It renders age to be the close second-most
important variable in the prediction. The educational credentials variable of the ob-
servation becomes the most important predictor.
S222 Cengiz et al.
are indeed minimum wage workers. The associated recall rate is 36%, which
means that this group covers around 36% of all minimum wage workers.
Since the high-probability group covers just over a third of all minimum
wage workers, we also study the impact of the minimum wage on a more
broadly defined group. In the high-recall group, we set a threshold probabil-
ity such that 75% of all minimum wage workers are captured. In practice,
this leads us to set the threshold at 12%; at that level we achieve a 35% pre-
cision rate. This high-recall group covers just more than 40% of all workers
in the data. To study the impact of the policy on workers unlikely to be af-
fected by the policy, we also define a group for whom the predicted proba-
bility is less than 12%. Throughout the paper, we refer to that group as the
low-probability group.
o b treat t
g
Y st 5 t st 1 Qst 1 ms 1 rt 1 ust , (1)
t523
g
where Y st is the labor market outcome (e.g., employment rate, unemploy-
ment rate, participation rate) in state s and at quarter t for group g. Unlike
when estimating the prediction model, here we use all states and quarters
available in the CPS data. As we discussed above, we study the impact of
the minimum wage on various groups defined by the prediction model, such
15
Appendix G in Cengiz et al. (2019) shows how this event study approach is
related to alternative methods applied in the literature like the two-way fixed effects
estimator with log minimum wage.
16
We show the graphical distribution of prominent minimum wage changes and
the number of such changes in each year in fig. A.1 (figs. A.1–A.6, C.1, C.2 are avail-
able online). In fig. A.2 we plot the change in state-level log (real) minimum wage fol-
lowing a prominent minimum wage increase.
Seeing beyond the Trees S223
as the high-probability group and the high-recall group. Here, treattst is a bi-
nary variable that takes on the value of 1 if the minimum wage was increased
t years from date t in state s. This definition implies that t 5 0 represents the
first year following the minimum wage increase (i.e., the quarter of treatment
and the subsequent three quarters), and t 5 21 is the year (four quarters)
before treatment. Our benchmark specification controls for state and period
fixed effects, ms and rt, and we also include controls for small or federal in-
creases, Qst.17 We cluster our standard errors by state, which is the level at
which policy is assigned.
Our baseline approach uses staggered variation of minimum wage increase.
As shown in many recent papers, this can lead to negative weighting bias in
the presence of heterogeneous treatment effects (see, e.g., Sun and Abraham
2020). To alleviate these concerns, in table E.1 we present estimates with a
stacked regression approach following Cengiz et al. (2019). In this approach,
we align events by event time (and not calendar time) and use only within-
event variation (between the treated unit and clean control units), which is
equivalent to a setting where all of the events happened all at once and were
not staggered. Gardner (2021) derives the implicit weights for such a stacked re-
gression approach and shows that they do not suffer from negative weighting.
Main Results
Table 2 shows the estimated effects on the high-probability, the high-recall,
and the low-probability groups. We report 5-year averaged posttreatment
estimates for the key labor market outcomes (wages, employment, unem-
ployment, and LFP) relative to the t 5 21, formally ð1=5Þo4t50 ðbt 2 b21 Þ.
Columns 1 and 3 establish that the minimum wage has a significant positive
impact on wages for groups of workers predicted to be exposed to the min-
imum wage. In the high-probability group, wages increased by around 2.3%
(SE, 0.3%), while in the high-recall group—which captures 75% of the min-
imum wage workers—wage increase was a little smaller but still significant
(1.6%; SE, 0.3). In contrast, column 5 shows no indication of any significant
wage effects for the low-probability group. This confirms that the wage
growth occurred only for workers exposed to the minimum wages and not
for individuals unlikely to be directly exposed to the minimum wage shock.
Columns 2, 4, and 6 of table 2 show analogous estimates but classifying
workers according to the Card and Krueger linear probability model.18 It is
17
The variables we use to control for federal and small events are the same as the
ones employed in Cengiz et al. (2019). We collapse the windows for small and federal
events into three periods: EARLY, PRE, and POST. EARLY is for 3 and 2 years be-
fore, and PRE is for 1 year before the event. POST is for 0–4 years after the event.
18
Table A.2 shows that the boosted tree prediction model picks a sample that is a
bit older, more educated, more female, and less white than the sample selected by
the Card and Krueger model.
Table 2
Impact of the Minimum Wage on Labor Market Outcomes
S224
(1) (2) (3) (4) (5) (6)
D wage (%) .023*** .025*** .016*** .014*** 2.001 2.001
(.003) (.004) (.003) (.003) (.003) (.003)
D employment (pp) .002 .001 .001 .001 .001 .000
(.002) (.002) (.001) (.001) (.001) (.001)
D unemployment (pp) 2.001 2.001 2.000 2.001* 2.000 .000
(.001) (.001) (.000) (.000) (.000) (.000)
D participation (pp) .002 .001 .000 .001 .001 .001
(.002) (.002) (.001) (.001) (.001) (.001)
Employment elasticity with respect
to minimum wage .071 .036 .020 .028 .012 .006
(.076) (.066) (.038) (.032) (.011) (.013)
Employment elasticity with respect
to wage .286 .138 .114 .191 NA NA
(.316) (.245) (.216) (.211)
Table 2 (Continued)
(1) (2) (3) (4) (5) (6)
Number of events 172 172 172 172 172 172
Number of observations 7,854 7,854 7,854 7,854 7,854 7,854
Number of individuals in sample 6,639,492 5,812,367 20,917,455 24,737,455 29,370,470 25,511,485
Mean employment .338 .399 .415 .439 .741 .771
Mean unemployment .058 .064 .046 .043 .030 .030
Mean participation .395 .462 .460 .481 .771 .801
Group High probability High probability High recall High recall Low probability Low probability
Prediction model Boosted tree CK linear Boosted tree CK linear Boosted tree CK linear
NOTE.—The table reports the effects of the minimum wage on labor market outcomes according to the event study analysis (see eq. [1]) using 172 state-level minimum wage changes
between 1979 and 2019. The table reports 5-year averaged posttreatment estimates for each key labor market outcome: percent change in wages and the change in employment to pop-
ulation, unemployment to population, and labor force participation rate. We also report the employment elasticity with respect to the minimum wage and the employment elasticity with
respect to the wage, which is the ratio of the percent change in employment and wage. To calculate the percent change in employment, we divide the change in employment to population
by the mean employment to population rate preceding the minimum wage hikes (reported at the bottom of the table). “Number of observations” refers to the number of quarter-state cells
used for estimation, while “Number of individuals” refers to the underlying CPS sample used to calculate labor market outcomes in these cells. Columns 1 and 2 show estimates for the
high-probability group, which captures the 10% of the population with highest predicted probability. Columns 3 and 4 show estimates for the high-recall group, which consists of in-
dividuals whose predicted probability is above 12%—a threshold that leads to a 75% recall rate of minimum wage workers. Columns 5 and 6 show the estimates for workers whose
predicted probability is below 12%. Columns 1, 3, and 5 use the best-performing prediction model—the boosted tree model. Columns 2, 4, and 6 use the Card and Krueger (CK) linear
prediction model (here, the high-recall group is defined by having a predicted probability, using the linear prediction model, above 12%, which again generates a 75% recall rate). All
S225
regressions are weighted by state-quarter population. Robust standard errors in parentheses are clustered by state. pp 5 percentage point.
* p < .10.
*** p < .01.
S226 Cengiz et al.
worth noting that the wage effects are almost the same for the best-performing
prediction model and for the Card and Krueger approach and suggests that
the Card and Krueger model performs quite well in this setting.19
Table 2 also reports the effect of the minimum wage on employment. Con-
sidering the high-probability and high-recall groups, we find a small and statis-
tically insignificant positive effect on employment. The employment elasticity
with respect to minimum wage is around 0.07 (SE, 0.08) and 0.02 (SE, 0.04),
respectively. The 90% confidence intervals around these estimates can rule
out an elasticity of 20.1, the lower bound (in magnitude) of the range sug-
gested by Neumark and Wascher (2008).
Table 2 also shows that the employment effects are somewhat smaller for
the high-recall group than for the high-probability group, which is in line with
the wage effects. This leads to a similar elasticity of employment with respect
to own wage, which would be the labor demand elasticity in the neoclassical
model, in the two groups. When we calculate the employment elasticity with
respect to own wage, we obtain an elasticity of 0.29 (SE, 0.32) for the high-
probability group and 0.11 (SE, 0.22) for the high-recall group. The estimates
are quite precise, especially those from the high-recall group, which can rule
out all but a modest negative impact of the policy on employment.
A key advantage of the probability-based approach is that we can study
outcomes other than employment and wages. In table 2 we also report the ef-
fect of the minimum wage on unemployment rate and on participation rate.
For the high-probability group (col. 1), we find a slight decrease in unemploy-
ment and a slight increase in the participation rate. Importantly, however,
none of these changes in unemployment and participation rates are statistically
significant. The estimates for the high-recall group (col. 3) show no change in
either the unemployment rate or the participation rate.
The estimated slight decrease in unemployment or the slight increase in par-
ticipation (or just unchanged values of both) suggests that search effort is un-
likely to fall in response to the policy. This set of empirical findings is difficult
to reconcile with a Flinn-type search-and-matching model (Flinn 2006), where
adjustments on the participation margin play a vital role, or with models
predicting a considerable increase in the unemployment rate in response to
the policy (Drazen 1986; Lang 1987; Swinnerton 1996). Furthermore, as we
discussed in section II, the lack of a visible drop in participation implies that
the policy did not have a significant negative effect on the welfare of workers
at the margin of LFP, even as it raised wages for inframarginal workers.
19
However, the key advantage of applying ML tools is that someone can select
the predictors in a data-driven way without knowing much about the context. Even
if the functional form chosen by Card and Krueger (1995) performs very well, it is
unclear how someone with less knowledge about US labor markets could come up
with that functional form. Moreover, there is no guarantee that it would perform
well in all contexts.
Seeing beyond the Trees S227
20
Since we do not fully observe the impact of the policy 5 years after the minimum
wage increase for the most recent events, those events impact the estimates only in
earlier posttreatment years. In figs. A.4 and A.5 we assess the timing of the policy
when we focus only on minimum wage changes where we see responses for the entire
event window. For these events with a balanced sample, which are not subject to the
composition effect, we find no decline in wage effects over time.
S228 Cengiz et al.
FIG. 5.—Impact of the minimum wage for alternative predicted probability thresh-
old values. The figure shows the main results from our event study analysis (see eq. [1])
using alternative predicted probability threshold values. We exploit 172 state-level min-
imum wage changes between 1979 and 2019. The figure shows the effect of a minimum
wage increase on wages (A), on employment to population (B), on unemployment to
population (C), and on labor force participation rate (D). In each panel the solid line
shows the 5-year averaged posttreatment estimates for individuals whose predicted
probability is above the minimum predicted probability threshold. On the x-axis we
also report the corresponding recall rate (the fraction of minimum wage workers re-
trieved by the prediction model if the particular minimum predicted probability thresh-
old is applied) and the precision rate (the fraction of minimum wage workers in the pre-
dicted group if the particular minimum predicted probability threshold is applied). We
also plot the thresholds corresponding to the high-probability group capturing the
10% of the population with the highest predicted probability and to the high-recall
group capturing 75% of all minimum wage workers. To calculate the predicted prob-
abilities, we use the best-performing prediction model—the boosted tree model. The
shaded areas show the 95% confidence interval based on standard errors that are
clustered at the state level. A color version of this figure is available online.
Panel B shows the impact of the policy on employment around the min-
imum wage hike. For both the high-recall group and the high-probability
group, we see a similar pattern: there is no clear evidence of preexisting trends
in employment, although there is a slight dip in employment 2 years before
the minimum wage increase when we look at the high-recall group (fig. 6).
Nevertheless, there is no unusual employment change if we look at the longer
Seeing beyond the Trees S229
FIG. 6.—Impact of the minimum wage over time, high-recall group. The figure
shows the main results from our event study analysis (see eq. [1]) using 172 state-
level minimum wage changes between 1979 and 2019. The figure shows the effect of
a minimum wage increase on wages (A), on employment to population (B), on un-
employment to population (C), and on labor force participation rate (D) for the high-
recall group. The high-recall group consists of all workers whose predicted probability
is above 12%—a threshold that corresponds to a 75% recall rate. To calculate the pre-
dicted probabilities, we use the best-performing prediction model—the boosted tree
model. We also show the 95% confidence interval based on standard errors that are clus-
tered at the state level. A color version of this figure is available online.
horizon between 1 and 3 years preceding the minimum wage increase. Fur-
thermore, the small drop in employment between 1 and 2 years preceding the
minimum wage increase would imply that the economy slightly deteriorates
before an average minimum wage hike, so we would expect a decrease in
employment rate after the policy change. In contrast, we see no clear break
in employment after the policy change—if anything, there is a small, statisti-
cally insignificant increase.
Panel C shows the impact on the unemployment rate. There is neither any
preexisting trend nor any break after the minimum wage increase. Panel D
shows the impact of the policy on participation rate, which also shows a flat
response following the policy change. Overall, the preexisting trends are re-
assuring, and we find no indication of any break in the participation rate at
the time of the minimum wage increases. Moreover, we do not find evidence
S230 Cengiz et al.
FIG. 7.—Impact of the minimum wage over time, high-probability group. The fig-
ure shows the main results from our event study analysis (see eq. [1]) using 172 state-
level minimum wage changes between 1979 and 2019. The figure shows the effect of a
minimum wage increase on wages (A), on employment to population (B), on unem-
ployment to population (C), and on labor force participation rate (D) for the high-
probability group. The high-probability group consists of the 10% of the population
with the highest likelihood of being affected by the policy. To calculate the predicted
probabilities, we use the boosted tree model. We also show the 95% confidence in-
terval based on standard errors that are clustered at the state level. A color version of
this figure is available online.
FIG. 8.—Impact of the minimum wage by predicted probability quintiles. The fig-
ure shows the effect of the minimum wage separately for each predicted probability
quintile. The highest quintile comprises individuals with predicted probabilities (of
being minimum wage workers) in the top 20%. We estimate equation (1) for each
quintile separately and report the 5-year averaged posttreatment estimates. We use
172 state-level minimum wage changes between 1979 and 2019. The figure shows
the effect of a minimum wage increase on wages (A), on employment to population
ratio (B), on unemployment rate (C), and on labor force participation rate (D). We
also show the 95% confidence intervals based on standard errors that are clustered
at the state level. A color version of this figure is available online.
Robustness
In table 3 we assess the robustness of the main results shown in table 2 to the
inclusion of various versions of time-varying heterogeneity for the high-recall
group, while we report the same robustness checks for the high-probability
group in table A.3. In column 1 we report the estimates for the baseline spec-
ification shown in table 2. Column 2 allows the period effects to vary by the
nine census divisions. The results are similar to the baseline specification: we
find a positive and statistically significant wage effect and a positive employ-
ment effect—which comes from an increase in the participation rate.
Table 3
Impact of the Minimum Wage on Labor Market Outcomes, Robustness to Alternative Specifications (High-Recall Group)
(1) (2) (3) (4) (5) (6) (7)
D wage (%) .016*** .016*** .018*** .012*** .016*** .016*** .016***
(.003) (.003) (.003) (.003) (.005) (.002) (.002)
D employment (pp) .001 .002 .001 2.001 .002 .000 .000
(.001) (.002) (.002) (.001) (.002) (.001) (.001)
D unemployment (pp) 2.000 .000 2.001 .000 2.001* 2.000 2.000
S232
(.000) (.000) (.000) (.001) (.000) (.000) (.000)
D participation (pp) .000 .002 .000 2.000 .001 .000 .000
(.001) (.002) (.002) (.001) (.002) (.001) (.001)
Employment elasticity with respect
to minimum wage .020 .047 .018 2.014 .033 .010 .007
(.038) (.038) (.040) (.037) (.052) (.021) (.035)
Employment elasticity with respect
to wage .114 .284 .096 2.109 .233 .071 .048
(.216) (.242) (.208) (.285) (.330) (.141) (.205)
Number of events 172 172 406 172 99 172 172
Number of observations 7,854 7,854 7,854 7,854 7,854 7,854 7,854
Number of individuals in sample 20,917,455 20,917,455 20,917,455 20,917,455 20,917,455 20,917,455 20,917,455
Mean employment .415 .415 .425 .419 .426 .415 .415
Mean unemployment .046 .046 .049 .045 .050 .046 .046
Mean participation .460 .460 .473 .464 .476 .460 .460
Table 3 (Continued)
(1) (2) (3) (4) (5) (6) (7)
Controls:
State fixed effects Y Y Y Y Y Y Y
Quarter fixed effects Y Y Y Y Y Y Y
Division-quarter fixed effects Y
State federal events Y
Unweighted Y
Number of events after 2014:Q1 Y
State employment control: all Y
State unemployment control: all Y
State employment control: low-probability group Y
State unemployment control: low-probability group Y
NOTE.—The table reports the effects of the minimum wage on labor market outcomes according to the event study analysis (see eq. [1]) using 172 minimum wage changes between
1979 and 2019. We assess the impact of the minimum wage on the high-recall group. The high-recall group consists of individuals whose predicted probability is above 12%—a thresh-
old that leads to a 75% recall rate of minimum wage workers. The table reports 5-year averaged posttreatment estimates for each key labor market outcome: percent change in wages and
the change in employment to population, unemployment to population, and labor force participation rate. We also report the employment elasticity with respect to the minimum wage
and the employment elasticity with respect to the wage, which is the ratio of the percent change in employment and wage. To calculate the percent change in employment, we divide the
change in employment to population by the mean employment to population rate preceding the minimum wage hikes (reported at the bottom of the table). “Number of observations”
S233
refers to the number of quarter-state cells used for estimation, while “Number of individuals” refers to the underlying CPS sample used to calculate labor market outcomes in these cells.
In all of the regressions, we use the best-performing prediction model—the boosted tree model. Column 1 shows the preferred benchmark estimate reported in col. 3 of table 2. Column 2
augments the baseline model with division-by-quarter fixed effects. Column 3 reports estimates using 406 state or federal minimum wage increases. All regressions are weighted by state-
quarter population except col. 4, where we report unweighted estimates. Column 5 considers only minimum wage events that happened on or before 2014:Q1 to ensure a full 5-year post-
treatment period. Column 6 controls for state-level unemployment and employment rates (as a fraction of population), while col. 7 controls for the employment and unemployment rates of
individuals with low predicted probability of being a minimum wage worker (less than 12%). Robust standard errors in parentheses are clustered by state. pp 5 percentage point.
* p < .10.
*** p < .01.
S234 Cengiz et al.
Table 4
Impact of the Minimum Wage on Labor Market Transitions
(1) (2) (3)
D E-U flow as a share of employment (pp) .000 .000 .000
(.001) (.000) (.000)
D E-I flow as a share of employment (pp) 2.000 .001 .000
(.001) (.001) (.000)
D U-E flow as a share of unemployment (pp) .006 .004 .004
(.005) (.004) (.003)
D U-I flow as a share of unemployment (pp) .002 .003 .000
(.005) (.004) (.003)
D I-E flow as a share of inactivity (pp) .001 .000 .001
(.001) (.000) (.001)
D I-U flow as a share of inactivity (pp) .000 .000 2.000
(.001) (.000) (.000)
Number of events 172 172 172
Number of observations 7,854 7,854 7,854
Number of individuals in sample 4,701,665 14,856,017 21,156,039
Mean E-U flow as a share of employment (%) .028 .022 .009
Mean E-I flow as a share of employment (%) .107 .061 .020
Mean U-E flow as a share of unemployment (%) .230 .246 .258
Mean U-I flow as a share of unemployment (%) .369 .295 .190
Mean I-E as a share of inactivity (%) .057 .042 .053
Mean I-U as a share of inactivity (%) .036 .024 .023
Group High High recall Low
probability probability
Prediction model Boosted tree Boosted tree Boosted tree
NOTE.—The table reports the effects of the minimum wage on labor market transition rates according to
the event study analysis (see eq. [1]) using 172 state-level minimum wage changes between 1979 and 2019.
The table reports 5-year averaged posttreatment estimates for percentage point changes in the employment-to-
unemployment (E-U) transition rate as a share of employment (row 1), the employment-to-inactivity (E-I)
transition rate as a share of employment (row 2), the unemployment-to-employment (U-E) transition rate
as a share of unemployment (row 3), the unemployment-to-inactivity (U-I) transition rate as a share of un-
employment (row 4), the inactivity-to-employment (I-E) transition rate as a share of inactivity (row 5),
and the inactivity-to-unemployment (I-U) transition rate as a share of inactivity (row 6). We also report
the mean levels of each of these variables for the period preceding the minimum wage hikes. “Number of obser-
vations” refers to the number of quarter-state cells used for estimation, while “Number of individuals” refers to
the underlying CPS sample used to calculate labor market outcomes in these cells. Column 1 shows estimates for
the high-probability group, which captures the 10% of the population with highest predicted probability. Column 2
shows estimates for the high-recall group, which consists of individuals whose predicted probability is above
12%—a threshold that leads to a 75% recall rate of minimum wage workers. Column 3 shows the estimates
for workers whose predicted probability is below 12%. Columns 1–3 use the best-performing prediction
model—the boosted tree model. All regressions are weighted by state-quarter population. Robust standard
errors in parentheses are clustered by state. pp 5 percentage point.
S236
(1) (2) (3) (4) (5) (6) (7) (8)
D wage (%) .016*** .012*** .015*** .001 .027*** .016** .015*** .014***
(.003) (.003) (.004) (.005) (.003) (.007) (.004) (.003)
D employment (pp) .001 2.003 .001 2.001 .004 2.002 .001 .000
(.001) (.003) (.001) (.002) (.003) (.004) (.002) (.001)
D unemployment (pp) 2.000 2.001 2.001* 2.000 2.001 .000 2.000 2.000
(.000) (.001) (.000) (.000) (.002) (.001) (.000) (.000)
D participation (pp) .000 2.004 .000 2.002 .003 2.001 .000 .000
(.001) (.002) (.001) (.002) (.002) (.004) (.002) (.001)
Employment elasticity with respect
to minimum wage .020 2.069 .020 2.044 .112 2.072 .019 .013
(.038) (.056) (.039) (.050) (.088) (.147) (.058) (.038)
Employment elasticity with respect
to wage .114 2.526 .128 NA .396 2.430 .120 .090
(.216) (.431) (.248) (.318) (.927) (.356) (.258)
Table 5 (Continued)
(1) (2) (3) (4) (5) (6) (7) (8)
Number of events 172 172 172 172 172 172 172 172
Number of observations 7,854 7,841 7,854 7,854 7,854 7,854 7,854 7,854
Number of individuals in sample 20,917,455 5,680,302 13,357,475 4,883,626 3,763,811 2,163,408 9,118,096 16,755,128
Mean employment .415 .499 .385 .347 .333 .264 .361 .405
Mean unemployment .046 .061 .038 .024 .068 .015 .050 .049
Mean participation .460 .560 .423 .371 .401 .279 .412 .453
Group High recall High recall High recall High recall High recall High recall High recall High recall
Demographic group All Black or Hispanic Female Married female Teen Aged 60–70 LTHS HSL
NOTE.—The table reports the effects of the minimum wage on labor market outcomes by demographic group according to the event study analysis (see eq. [1]). We exploit
172 state-level minimum wage changes between 1979 and 2019. We assess the impact of the minimum wage on the high-recall group including individuals whose predicted probability
is above 12%—a threshold that leads to a 75% recall rate of minimum wage workers. The table reports 5-year averaged posttreatment estimates for each key labor market outcome:
percent change in wages and the change in employment to population, unemployment to population, and labor force participation rate. We also report the employment elasticity with
respect to the minimum wage and the employment elasticity with respect to the wage, which is the ratio of the percent change in employment and wage. To calculate the percent
change in employment, we divide the change in employment to population by the mean employment to population rate preceding the minimum wage hikes (reported at the bottom
of the table). “Number of observations” refers to the number of quarter-state cells used for estimation, while “Number of individuals” refers to the underlying CPS sample used to
calculate labor market outcomes in these cells. In all of the regressions, we use the best-performing prediction model—the boosted tree model. The demographic subgroups are Black
or Hispanic, woman, married woman, teen, aged 60 and older and less than 70, less than high school (LTHS) education, and high school or less (HSL) education. All estimates are
weighted by the corresponding subgroup’s population in the state. Robust standard errors in parentheses are clustered by state. pp 5 percentage point.
S237
* p < .10.
** p < .05.
*** p < .01.
S238 Cengiz et al.
else, in each subgroup, we focus on workers who are in the high-recall group.21
Restricting the sample to workers who are likely to be minimum wage workers
is also necessary for getting first-stage wage effects, which would not be pos-
sible if we had all workers in the sample (see table A.1 in the online appendix of
Cengiz et al. 2019).
For most subgroups in table 5, we find a clear and significant impact on
wages, which confirms a key advantage of restricting the sample to groups
that are likely to be exposed to the minimum wage. In column 1 we simply
reproduce the benchmark estimates for the overall high-recall group for
comparability. In column 2 we report the estimates for workers who are
Black or Hispanic.22 While the estimates are somewhat noisy, the point es-
timates indicate a nontrivial drop in employment and participation. This
suggests that in the Black or Hispanic group, some workers (who were at
the participation margin) may have been made worse off by the policy.
However, the standard errors are too large to draw a clear conclusion on this
question.
In columns 3 and 4 we examine the impact of the policy on all women and
on married women, respectively. Since the extensive margin labor supply elas-
ticities are often found to be larger for women than for men, it might be the
case that minimum wages lead to a greater increase in women’s participa-
tion. Furthermore, married women are also typically thought to have larger
responsiveness on the extensive margin. However, our estimated effects for
employment and participation outcomes for women and married women
are very similar to the benchmark specification; however, we note that there
is no statistically (or economically) significant wage effect for married women
in our sample, which makes it difficult to draw any strong conclusions. (We
find more informative evidence when we consider married mothers with
younger kids below.)
Columns 5 and 6 report the impact on teens and on older individuals (aged 60–
70), respectively. Both younger individuals and older individuals have lower
participation rates than prime-age individuals, and they are also thought to
have more elastic labor supply (Blundell, Bozio, and Laroque 2011). For
teens, we find larger wage effects and slightly larger employment increase
21
We use the predicted probabilities estimated on all workers. One could estimate
a separate prediction model for each subgroup and then use those predicted proba-
bilities. However, it is unlikely that there are substantial gains from estimating a sep-
arate prediction model for each subgroup. If there were substantial gains from using
some predictors differently within a particular subgroup, then the boosting algo-
rithm would tend to detect and incorporate this fact into the prediction model—even
if it were estimated using the full sample. We also explored separately estimating pre-
diction models for some specific demographic groups but found negligible changes in
the precision of our estimates.
22
We attempted to estimate the effect of the policy on Black and Hispanic indi-
viduals separately, but the estimates were too imprecise to be informative.
Seeing beyond the Trees S239
than for the overall sample. The increase in employment comes from the
changes in participation, which is in line with the idea that more teens are
on the participation margin and also with the evidence presented by Laws
(2018). For older individuals, the wage effects are imprecise (although the
point estimates are close to the overall sample). This makes interpreting the
results for older individuals difficult. Still, we document a slight positive (sta-
tistically insignificant) employment and participation effect, which is in line
with Borgschulte and Cho (2019), who document a similar-sized response
in employment for those aged between 62 and 70.
Columns 7 and 8 show the impact on those with lower educational cre-
dentials. The labor market impact of the policy on these education groups is
very similar to the impact of the policy on the overall sample. These findings
suggest that workers with lower education credentials seem to benefit from
the minimum wage policies.
Table 6 presents additional heterogeneity analysis. Columns 2–4 show the
estimates by the predicted probability of moving into and out of the labor
force. The goal of this exercise is to help us assess any heterogeneity in the
effects by how likely workers are to be at the margins of LFP. We predict
the probability that an individual changes their LFP status—either from non-
participation to participation or from participation to nonparticipation—by
applying the boosting tree ML method. For estimating this prediction model,
we use all of the demographic variables that we had used before to predict
minimum wage exposure but also add the number of children under the
age of 5, since it is likely to be an important predictor of LFP. The relative
importance of the predictors is shown in figure A.6. Similar to the prediction
model on minimum wage exposure, age, education, and gender are the three
most important predictors of changes in participation. In addition, race and
number of children under the age of 5 also substantially influence the predic-
tion model.
Since the number of children is coded consistently only since 1986 in the
CPS, we report estimates using the 1986–2018 period in table 6. Column 1
reports estimates for all workers using only that period, and they are very
similar to the estimates using the 1979–2019 sample. Column 2 shows the es-
timates for the group of workers that has a high predicted probability of
changing LFP status. The estimated impacts for this group are very similar
to those from the overall sample. Columns 3 and 4 show the estimates with
lower probability of switching. Again, we find responses similar to those
from the overall sample. They key takeaway is that even when we look at
individuals who are more likely to switch LFP, we find no meaningful dif-
ference in the causal effects of minimum wage policies. These findings cast
doubt that there was much of an impact from minimum wage changes in
our sample on job search and participation behavior, even among groups
that are likely to be at the margin of participation. In columns 5 and 6 we also
report estimates on single and married mothers, focusing on those with kids
Table 6
Impact of the Minimum Wage on Labor Market Outcomes by Demographic Group II
S240
(1) (2) (3) (4) (5) (6)
D wage (%) .015*** .017*** .012*** .011 .025** .006
(.003) (.003) (.003) (.008) (.012) (.009)
D employment (pp) .001 .000 .002 2.001 2.000 2.005
(.001) (.002) (.002) (.001) (.006) (.004)
D unemployment (pp) 2.001 2.001 2.001 .001* 2.001 2.001
(.000) (.001) (.001) (.000) (.003) (.001)
D participation (pp) .000 2.001 .001 2.001 2.001 2.006
(.001) (.002) (.002) (.001) (.007) (.004)
Employment elasticity with respect
to minimum wage .017 .006 .035 2.084 2.005 2.115
(.035) (.035) (.038) (.071) (.113) (.097)
Employment elasticity with respect
to wage .107 .033 .276 NA 2.021 NA
(.217) (.193) (.339) (.442)
Table 6 (Continued)
(1) (2) (3) (4) (5) (6)
Number of events 156 156 156 156 156 156
Number of observations 6,222 6,222 6,222 6,222 6,222 6,222
Number of individuals in sample 15,760,550 7,883,899 2,915,439 4,961,212 525,347 618,103
Mean employment .417 .512 .561 .148 .520 .417
Mean unemployment .047 .068 .042 .008 .091 .040
Mean participation .463 .581 .604 .155 .612 .457
Probability group High recall High recall High recall High recall High recall High recall
Demographic group All High LFP switch Medium LFP switch Low LFP switch Single mother Married mother
kids under 5 kids under 5
NOTE.—The table reports the effects of the minimum wage on labor market outcomes by demographic group according to the event study analysis (see eq. [1]). We exploit 156 state-
level minimum wage changes between 1986 and 2018. We assess the impact of the minimum wage on the high-recall group including individuals whose predicted probability is above 12%—
a threshold that leads to a 75% recall rate of minimum wage workers. The table reports 5-year averaged posttreatment estimates for each key labor market outcome: the change in em-
ployment to population, unemployment to population, and labor force participation rate. “Number of observations” refers to the number of quarter-state cells used for estimation, while
“Number of individuals” refers to the underlying CPS sample used to calculate labor market outcomes in these cells. In all of the regressions, we use the best-performing prediction model—
the boosted tree model. Column 1 shows the estimates on overall employment using the 1986–2018 periods. Columns 2–4 show the estimates for individuals with the different predicted
probability of moving into or out of the labor force. Column 5 shows the estimates for single mothers with children under the age of 5, while col. 6 shows the estimates for married mothers
with children under the age of 5. All estimates are weighted by the corresponding subgroup’s population in the state. Robust standard errors in parentheses are clustered by state. pp 5
percentage point.
S241
* p < .10.
** p < .05.
*** p < .01.
S242 Cengiz et al.
VI. Conclusion
In this paper we estimate the impact of the minimum wage on various la-
bor market outcomes using 172 prominent minimum wage changes between
1979 and 2019. To capture the impact of the policy on a broad group of af-
fected workers, we utilize modern ML techniques to estimate the likelihood
that someone is a minimum wage worker. While the best-performing predic-
tion model does better than the linear prediction model of Card and Krueger
(1995), the gap is not large. One implication of these findings is that mini-
mum wage researchers who are not interested in investing in a ML approach
may do fairly well by simply applying the Card and Krueger linear proba-
bility prediction model. Of course, the advantage of the ML approach is that
23
We find an effect of minimum wages on employment for single mothers with
young children that is close to zero and statistically insignificant. This differs from
Godoy, Reich, and Allegretto (2020), who find a large, statistically significant in-
crease in employment for single mothers with kids under 5 in response to a mini-
mum wage increase. Table A.5 provides a reconciliation of the discrepancy between
these findings in greater detail. There are a number of differences between the two
papers, but the primary reason why our estimates differ from Godoy, Reich, and
Allegretto (2020) is that we use the CPS-Basic files while they use the (smaller)
CPS-ORG data. While use of the CPS-Basic versus CPS-ORG data by itself does
not make a difference in the overall sample (see cols. 1–5 in table A.5), the results are
more sensitive to the data sets for the single mother sample, which is much smaller
(see cols. 6–10 in table A.5).
Seeing beyond the Trees S243
Table 7
Impact of the Minimum Wage on Alternative Labor Market Outcomes
(1) (2) (3)
D self-employment as share of employment (pp) .001 2.000 2.001
(.001) (.001) (.001)
D part-time as share of employment (pp) 2.005** 2.001 .000
(.002) (.001) (.000)
D overtime as share of employment (pp) .001 .000 2.001
(.001) (.001) (.001)
Number of events 172 172 172
Number of observations 7,854 7,854 7,854
Number of individuals in sample 6,639,492 20,917,455 29,370,470
Mean self-employment .029 .066 .129
Mean part-time .414 .216 .080
Mean overtime .029 .067 .196
Group High High recall Low
probability probability
Prediction model Boosted tree Boosted tree Boosted tree
NOTE.—The table reports the effects of the minimum wage on alternative labor market outcomes according
to the event study analysis (see eq. [1]) using 172 state-level minimum wage changes between 1979 and 2019.
The table reports 5-year averaged posttreatment estimates for the following labor market outcomes: self-
employment, part-time (working fewer than 30 hours per week), and overtime (working more than 40 hours
per week). Each of these variables is expressed as a share of total employment. Column 1 shows estimates for
the high-probability group, which captures the 10% of the population with the highest predicted probability.
Column 2 shows estimates for the high-recall group, which consists of individuals whose predicted probabil-
ity is above 12%—a threshold that leads to a 75% recall rate of minimum wage workers. Column 3 shows the
estimates for workers whose predicted probability is below 12%. All columns use the best-performing pre-
diction model—the boosted tree model. All of the regressions are weighted by state-quarter population. Ro-
bust standard errors in parentheses are clustered by state. pp 5 percentage point.
** p < .05.
References
Adams, Camilla, Jonathan Meer, and CarlyWill Sloan. 2018. The minimum
wage and search effort. NBER Working Paper no. 25128, National Bu-
reau of Economic Research, Cambridge, MA.
Addison, John T., and McKinley L. Blackburn. 1999. Minimum wages and
poverty. Industrial and Labor Relations Review 52, no. 3:393–409.
Addison, John T., McKinley L. Blackburn, and Chad Cotti. 2011. Mini-
mum wage increases under straightened circumstances. IZA Discussion
Paper no. 6036, Institute of Labor Economics, Bonn.
Ahn, Tom, Peter Arcidiacono, and Walter Wessels. 2011. The distributional
impacts of minimum wage increases when both labor supply and labor
demand are endogenous. Journal of Business and Economic Statistics 29,
no. 1:12–23.
Allegretto, Sylvia, Arindrajit Dube, Michael Reich, and Ben Zipperer. 2017.
Credible research designs for minimum wage studies: A response to
Neumark, Salas, and Wascher. ILR Review 70, no. 3:559–92.
Autor, David H., John J. Donohue III, and Stewart J. Schwab. 2006. The
costs of wrongful-discharge laws. Review of Economics and Statistics
88, no. 2:211–31.
Belman, Dale, and Paul J. Wolfson. 2014. What does the minimum wage do?
Kalamazoo, MI: W.E. Upjohn Institute.
Belman, Dale, Paul Wolfson, and Kritkorn Nawakitphaitoon. 2015. Who is
affected by the minimum wage? Industrial Relations: A Journal of Econ-
omy and Society 54, no. 4:582–621.
Blundell, Richard, Antoine Bozio, and Guy Laroque. 2011. Labor supply
and the extensive margin. American Economic Review 101, no. 3:482–86.
[Link]
Seeing beyond the Trees S245
Borgschulte, Mark, and Heepyung Cho. 2019. Minimum wages and retire-
ment. ILR Review 73, no. 1:153–77.
Breiman, Leo. 2001. Random forests. Machine Learning 45, no. 1:5–32.
Card, David. 1992. Do minimum wages reduce employment? A case study
of California, 1987–89. ILR Review 46, no. 1:38–54.
Card, David, and Alan B. Krueger. 1995. Myth and measurement: The new
economics of the minimum wage. Princeton, NJ: Princeton University
Press.
Cengiz, Doruk, Arindrajit Dube, Attila Lindner, and Ben Zipperer. 2019.
The effect of minimum wages on low-wage jobs. Quarterly Journal of
Economics 134, no. 3:1405–54.
Clemens, Jeffrey, Lisa B. Kahn, and Jonathan Meer. 2018. The minimum wage,
fringe benefits, and worker welfare. NBER Working Paper no. 24635, Na-
tional Bureau of Economic Research, Cambridge, MA.
Clemens, Jeffrey, and Michael R. Strain. 2017. Estimating the employment
effects of recent minimum wage changes: Early evidence, an interpretative
framework, and a pre-commitment to future analysis. NBER Working
Paper no. 23084, National Bureau of Economic Research, Cambridge, MA.
Clemens, Jeffrey, and Michael Wither. 2019. The minimum wage and the
Great Recession: Evidence of effects on the employment and income tra-
jectories of low-skilled workers. Journal of Public Economics 170:53–67.
Currie, Janet, and Bruce C. Fallick. 1996. The minimum wage and the em-
ployment of youth evidence from the NLSY. Journal of Human Re-
sources 31, no. 2:404–28.
Drazen, Allan. 1986. Optimal minimum wage legislation. Economic Journal
96, no. 383:774–84. [Link]
Dustmann, Christian, Attila Lindner, Uta Schönberg, Matthias Umkehrer,
and Philipp vom Berge. 2022. Reallocation effects of the minimum wage.
Quarterly Journal of Economics 137, no. 1:267–328.
Flinn, Christopher J. 2002. Interpreting minimum wage effects on wage dis-
tributions: A cautionary tale. Annales d’Economie et de Statistique 67/
68:309–55. [Link]
———. 2006. Minimum wage effects on labor market outcomes under search,
matching, and endogenous contact rates. Econometrica 74, no. 4:1013–62.
[Link]
———. 2011. The minimum wage and labor market outcomes. Cambridge,
MA: MIT Press.
Friedman, Jerome H. 2001. Greedy function approximation: A gradient
boosting machine. Annals of Statistics 29, no. 5:1189–232.
Friedman, Jerome, Trevor Hastie, and Robert Tibshirani. 2009. The elements
of statistical learning: Data mining, inference, and prediction. New York:
Springer.
Gardner, John. 2021. Two-stage differences in differences. Working paper,
University of Mississippi.
S246 Cengiz et al.