3 - Algorithm Reliance Fast and Slow. Management Science.
3 - Algorithm Reliance Fast and Slow. Management Science.
In algorithm-augmented service contexts where workers have decision authority, they face two decisions
about the algorithm: whether to follow its advice, and how quickly to do so. The pressure to work quickly
increases with the speed of arriving customers. In this paper, we ask: how do workers use algorithms to
manage system loads? With a laboratory experiment, we find that superior algorithm quality and high system
loads increase participants’ willingness to use their algorithm’s advice. Consequently, participants with the
superior algorithm make higher-quality recommendations than those with no algorithm (participants with the
inferior algorithm make slightly lower-quality recommendations than those without). However, participants
do not necessarily speed up by using algorithms’ advice; their throughput times only decrease compared
to the no-algorithm baseline when the system load is high and algorithm quality is superior, although
participants would benefit from working faster in all treatments. This happens in part because participants
in the high-load, superior-algorithm treatment serve customers more quickly than participants in the other
treatments, conditional on using the algorithm. Participants in the high-load, superior-algorithm treatment
work especially quickly in later periods as they increasingly default to their algorithm’s advice. Our findings
show that algorithms can have benefits for both decision quality and speed. Quality benefits come from
workers’ decision to use their algorithms’ advice, while speed benefits depend on workers’ algorithm use and
the time they spend deliberating about their algorithm use. Ultimately, algorithm quality and system load are
mutually reinforcing factors that influence both service quality and especially speed.
Key words : behavioral operations; human-algorithm interaction; service operations; queueing systems
1. Introduction
More and more organizations are embracing decision-support algorithms for their workers, recog-
nizing the potential value for operational efficiency and quality (Pettey 2016). Algorithms that
support human decision-making (instead of replacing it) are particularly promising for customer
queueing systems in service settings; customers often value human touch, which fully-automated
systems lack (Buell 2018). However, when human workers are granted ultimate decision authority,
their behaviors—if, when, and how they use algorithmic decision support—determine algorithms’
effects on service efficiency and quality. Little is currently known about how workers in queueing
systems will use such algorithms, or how these behaviors will affect service outcomes.
Research about human-algorithm collaboration shows that people are often averse to using algo-
rithms outside of queueing systems (e.g., Dietvorst et al. 2015). This might suggest that workers
within queuing systems will also choose not to follow algorithms’ advice, even when it would ben-
efit them to do so. Yet, other research shows that contextual factors, such as pressure from an
1
Snyder, Keppler, Leider: Algorithm Reliance, Fast and Slow
2
explicit time limit, can reduce algorithm aversion (Jung and Seiter 2021). Attributes of queueing
systems, like time pressure from system load—the workload created by incoming customer arrivals
(Delasay et al. 2019)—might similarly make workers there more inclined to use algorithms. After
all, workers in traditional service settings (that is, without algorithmic decision support) respond
to pressure from higher system loads by working more quickly (e.g., KC and Terwiesch 2009).
Even if workers do use algorithms to serve customers in a queueing system, this behavior alone
does not mean that service companies will see improved efficiency or quality outcomes from intro-
ducing new decision-support algorithms. The outcomes also depend on how workers use the algo-
rithms. Namely, workers might follow algorithms’ advice, but fail to do so quickly. Most previous
research about human-algorithm interaction has focused on its effects for quality only, so we do not
know that algorithm use necessarily affects worker decision-making speed. Workers in customer
queueing systems might use algorithms to speed up, or they might not. In practice, designing algo-
rithms to minimize deviations from them makes warehouse workers faster at packing boxes (Sun
et al. 2022), but in theory, algorithms may not save workers any time if they “induce the human
to exert more cognitive efforts” (Boyacı et al. 2023, page 1). Algorithms could induce workers to
exert cognitive efforts if workers only follow their advice sometimes, and take time deliberating
about when to do so. If workers do use algorithms in a way that improves their speed, this could
be problematic for service quality if the behavior is mostly driven by automation bias—a bias that
leads people to blindly rely on even bad algorithmic advice, especially under pressure from time
limits (Goddard et al. 2012).
These possibilities lead us to ask: how do workers in queueing systems engage with algorithmic
decision-support to manage system loads? What are the implications of this behavior for system
performance? In this paper, we investigate these questions with a novel laboratory experiment, and
use our answers to identify lessons for managers about improving worker-algorithm interactions.
Participants in our experiment take the role of workers in an M/G/1 queueing system (i.e.,
customer arrivals follow a Poisson process and each queue has a single server). They are tasked
with providing a joke recommendation to each unique customer in the queue, with support from
an algorithm (a task inspired by Yeomans et al. 2019, who study these recommendations outside
of queueing systems). Joke recommendation captures key aspects of personalized services: there is
scope for participant expertise, as subjects should have experience with or intuition about humor,
and customers have unique, heterogeneous preferences. Participants are directly incentivized for
recommendation quality, and to a lesser degree, recommendation speed. They repeat the joke
recommendation task over one practice and five paid periods, giving them the opportunity to
serve over 140 customers in total. We vary two features of the experimental setting to answer our
Snyder, Keppler, Leider: Algorithm Reliance, Fast and Slow
3
research questions: the system load (high—λload = 7.5 customers/minute, or low—λload = 5 cus-
tomers/minute) and the algorithm itself (no algorithm, superior-to-worker average recommendation
quality, or inferior-to-worker average recommendation quality). In an extension, we introduce two
interventions designed to improve performance outcomes in the superior-algorithm contexts.
We find that high system load and superior algorithm quality both induce greater algorithm
reliance. Because participants use their algorithm for at least some service decisions, their rec-
ommendations are best in the superior-algorithm conditions, and worst in the inferior-algorithm
conditions. More significantly, we find that greater reliance does not necessarily mean greater
worker speed. Participants’ throughput times are only significantly faster than the no-algorithm
baseline in the high-load, superior-algorithm condition, although participants facing low system
loads could earn 13% more by serving their customers more quickly. Further, only about half (54%)
of the throughput time improvement we see in this treatment can be explained by participants’
higher rates of algorithm use. The remainder of the improvement occurs because participants in
the high-load, superior-algorithm treatment follow their algorithm’s advice faster. Conditional on
using the algorithm’s advice, participants’ service speeds are significantly shorter in the high-load,
superior-algorithm treatment than in treatments where the system load is low or the algorithm’s
quality is inferior (conditional on not using the algorithm’s advice, participants’ service speeds
are shorter under high loads, regardless of algorithm quality). Moreover, subjects in the high-
load, superior-algorithm condition increasingly default to the algorithm in later periods. With
our follow-up treatments, we show that our interventions to improve system performance increase
participants’ rates of algorithm use and also their fast algorithm use—especially in the low-load
setting.
Our findings reveal three insights about how workers and algorithms interact in customer queue-
ing systems, specifically related to service efficiency and speed. First, in service settings, algorithms
have the potential to improve both decision quality and speed. We demonstrate multiple contexts
in which workers use algorithms to achieve better and faster service. Second, the benefits to service
speed depend on both workers’ willingness to follow a (superior) algorithms’ advice and the time
they spend thinking about that decision; in contrast, the benefits to service quality come primarily
from the reliance decision. In other words, algorithm reliance can be fast or slow, depending on
whether decision-makers default to the algorithm or deliberate; either type of reliance has the same
effect on service quality, but only faster algorithm reliance improves speed. Third, algorithm quality
and system load are mutually-reinforcing factors that influence both service quality and especially
speed. Algorithm reliance increases when algorithm quality improves or system load increases, but
either dial alone may not make workers faster—it takes a combination of both to realize speed
improvements. These three insights complement a large and growing body of work about the effects
Snyder, Keppler, Leider: Algorithm Reliance, Fast and Slow
4
of discretionary algorithm use on decision quality (e.g., Balakrishnan et al. 2022, Dietvorst et al.
2015, Lin et al. 2021), by relating the same discretion to another operationally-important outcome,
time.
2. Literature Review
Here we provide more detail about what is currently known about human-algorithm interactions,
and why queueing systems present an interesting new context in which to study these interactions.
example, pilots in flight simulations, under pressure to act quickly, over-rely on inaccurate decision
support (Sarter and Schroeder 2001). Ultimately, people use algorithms even when they are bad,
and avoid them even when they are good, in part because it is often difficult to evaluate algorithms’
quality—doing so requires “cognitive efforts” (Boyacı et al. 2023). This difficulty has been the topic
of emerging research in operations management (e.g., Bastani et al. 2020).
3. Experiment Design
3.1. Standard Behavioral Queueing Notation
Let us introduce some queueing system notation, adapted from Allon and Kremer (2018). Figure
1 diagrams the queueing system we study in this paper. Four key steps make up this system:
A. Customers arrive to the system at an average rate λload . In this paper, we say λload determines
overall system load; higher arrival rates mean higher system loads.
B. Customers join the queue and wait to be served. Their time waiting in the queue is TW .
C. Customers arrive to be served. Their time in service is TS . Service quality, which is customers’
value of the service, is v.
Snyder, Keppler, Leider: Algorithm Reliance, Fast and Slow
7
D. Customers exit the system. The total time customers spend in the system, their throughput
time, is T = TS + TW . In practice, workers may see feedback about customer satisfaction from
surveys or later interactions after this step. However, they typically cannot learn counterfactual
information about how customers would have responded to other service decisions, and so
they cannot know what the “optimal” service decision would be or how satisfied customers
would be with this decision (v ∗ ).
B. C.
A. 𝜆𝜆𝑙𝑙𝑙𝑙𝑙𝑙𝑙𝑙 Service D. Exit
Wait
Customers prefer shorter throughput times (smaller T ), but they also enjoy high service quality
(larger v). Therefore, customer satisfaction (“net utility”) depends on both outcomes, although
not necessarily equally. Allon and Kremer (2018) conceptualize aggregate-level system welfare “as
the product of customers’ net utility and system throughput” (page 362).
customer i+1
Current Recommendation Options D.
B. • Joke E #
• Joke G #
• Joke F #
• Joke H #
customer i-1
subjectivity and human intuition. In many service settings, workers employ empathy and past
experiences to inform their customer interactions and we expect participants will similarly have
some experience with predicting other people’s senses of humor from everyday exchanges. Even
the most inexperienced joke tellers among them can develop some human intuition for the task
by reading and comprehending each joke—something our algorithm does not do. We expect our
participants to be wary of the algorithm’s ability to provide helpful advice without this “human”
experience, thus creating a perceived tension between the benefits of algorithmic advice for service
speed versus service quality.
1
We use a step function conversion to help participants understand the incentives and interpret feedback.
2
Queue lengths are stable and short across periods. E.g., the average queue length is 1.0 customers in the low-load,
no-algorithm condition and 2.2 customers in the high-load, no-algorithm condition. See Table B.3 in the Appendix.
Snyder, Keppler, Leider: Algorithm Reliance, Fast and Slow
10
outcome. They do not see counterfactual information about how well they could have performed by
making other recommendations (e.g., v ∗ )—we, the researchers, do know this information. Partici-
pants see detailed feedback only about the most recently served customer; upon serving customer
i + 1, the display will refresh to show feedback about customer i + 1. Participants do see a summary
of the number of customers they have served and their total earnings for the period. Figure A.1 in
the Appendix shows how this feedback is presented to participants via our interface.
3.3. Treatments
We implement a 2 × 3, between-subjects design that varies participants’ system load and their
algorithm: (high load, low load) × (no algorithm, superior algorithm, inferior algorithm).
3.3.2. Algorithm
We manipulate both the availability and the quality of algorithm’s advice with our algorithm
treatment dimension. The no-algorithm treatments serve as a control that shows how participants
serve customers under high and low system loads without any algorithmic decision support. The
superior-to-worker- and inferior-to-worker-algorithm treatments vary the average advice quality
that participants see. Participants in our original treatments are not told anything about their
algorithm’s rating performance (we describe the effect of similar up-front information in an exten-
sion, see Section 5). Participants can reveal their algorithm’s advice for any customer with the
click of a button3 . We require participants to click a button to view the algorithm’s advice to get
a more precise measure of algorithm use. Clicking the button automatically selects the algorithm’s
advice for recommendation but does not automatically submit it. Participants can deviate from
the algorithm’s recommendation if they so choose. Figure A.1 in the Appendix shows what our
interface looks like just before and just after this button is clicked.
3
Results from an alternative design without the button suggest that the presence of the button in our experiment
does not meaningfully change the distribution of our algorithm use results—see Table D.4 in the Appendix.
Snyder, Keppler, Leider: Algorithm Reliance, Fast and Slow
11
Algo. performance
4
(i.e., 100% algorithm use)
2
1
0
No algo. Superior algo. only Inferior algo. only
Behind the scenes, both the superior and the inferior algorithms use OLS regressions on the
four historical joke ratings to predict each customer’s ratings of each joke option. The superior
algorithm selects the joke option with the highest predicted rating, and this algorithm works well
(see Figure A.3 in the Appendix for a comparison to alternative algorithms). The inferior algorithm
selects the joke with the lowest predicted rating with highest probability, and the joke with the
highest predicted rating with lowest (but non-zero) probability. Naturally, the superior algorithm
(average recommendation rating: 3.1) is significantly better than the inferior algorithm (average
rating: 0.8). What is more, the superior algorithm is also on the whole significantly better than our
participants at recommending jokes and the inferior algorithm is on the whole worse, though to a
lesser degree. Figure 3 compares the performance of the superior and inferior algorithms’ advice
(in other words, the average rating performance of a participant who follows the algorithm’s advice
for every recommendation) against a human benchmark—the average rating performance of our
participants in the no-algorithm conditions.
4
We collected data for the no- and superior-algorithm treatments between Feb.–Jun. of 2021, and the inferior-
algorithm treatments between Oct.–Nov. of 2022 in response to reviewer comments. Due to subject pool constraints,
we prioritized recruiting for the algorithm treatments, where algorithm use is measurable. Table D.1 in the Appendix
shows that participant controls are largely consistent over both periods of data collection.
Snyder, Keppler, Leider: Algorithm Reliance, Fast and Slow
12
the instructions (Section 1 of the Supplementary Appendix). Participants then answered several
comprehension questions before beginning the joke recommendation task described above. At the
end of the session, they completed a short follow-up questionnaire (Section 3 of the Supplementary
Appendix). Sessions lasted no longer than 90 minutes, and participants earned $20 on average,
including a $5 show-up fee.
5
Our results are generally robust to other thresholds for defaulting (e.g., see Table D.3 in the Appendix).
Snyder, Keppler, Leider: Algorithm Reliance, Fast and Slow
13
4. Results
4.1. Performance Outcomes by Treatment
Figure 4 summarizes both of our incentivized outcome measures—recommendation rating (the y-
axis of this plot) and throughput time (the x-axis)—by treatment. Note that because participants
earn more from faster throughput times, the x-axis is descending. Data points closer to the upper
right corner of the plot reflect higher overall performance along both dimensions. For example,
participants in the low-system-load, no-algorithm treatment (LL × NA in the figure) outperform
participants in the high-load, no-algorithm treatment (HL × NA); the former see on average a
rating of 1.8 and a throughput time of 13.6 seconds, while the latter see on average a rating of 1.7
and a throughput time of 17.3 seconds (only the difference in throughput time is significant, p <
0.05). Similarly, participants in the low-load, inferior-algorithm treatment (LL × IA) outperform
participants in the high-load, inferior-algorithm treatment (HL × IA); the former see on average a
rating of 1.5 and a throughput time of 13.8 seconds, while the latter see on average a rating of 1.4
and a throughput time of 17.5 seconds (again, only the difference in throughput time is significant,
p < 0.05). By contrast, participants in the low-load, superior-algorithm treatment (LL × SA) do
not outperform participants in the high-load, superior-algorithm treatment (HL × SA); the former
see on average a rating of 2.4 and a throughput time of 13.3 seconds, while the latter see on
average a rating of 2.4 and a throughput time of 12.5 seconds. As expected, participants’ algorithm
treatment assignment affects their rating performance. Participants with the inferior algorithm
perform slightly worse than participants with no algorithm on this dimension (p < 0.001), while
participants with the superior algorithm perform much better (p < 0.001).
2.5
LL x SA HL x SA
2
Rating
LL x NA
HL x NA
1.5
LL x IA
HL x IA
1
18 16 14 12
Throughput time (seconds)
Most interestingly, Figure 4 shows that while better algorithms improve throughput time a little,
better algorithms plus high system loads improve throughput times a lot. On average, customers of
Snyder, Keppler, Leider: Algorithm Reliance, Fast and Slow
14
participants in the low-load conditions spend a similar amount of time, between 13 and 14 seconds,
in the system regardless of the participants’ algorithm treatment assignment. On the other hand,
throughput times in the high-load condition are over 4.5 seconds faster on average for participants
with advice from the superior algorithm, compared to the other two algorithm conditions. The high-
load, superior-algorithm average throughput time outcome is especially striking in contrast to the
low-load, superior-algorithm average throughput time. Throughput times are almost four seconds
faster in the low-load than the high-load conditions, except in the superior-algorithm treatments
where the comparison reverses. As a result, participants earn significantly less per customer under
high loads than low loads, except with the superior algorithm (see Figure B.1 in the Appendix for
a graph of participants’ earnings). In the following sections we unpack why this is happening.
P(Algo. usei,j = 1) = F (β0 + β1 HLj + β2 SAj + β3 HLj × SAj + β4 ITi,j + β5 Ci + β6 Pj + ϵi,j ). (1)
In this equation, P(Algo. usei,j = 1) denotes participant j’s probability of using the algorithm
to serve customer i. HLj is an indicator variable denoting participant j’s system-load treatment
assignment (1 if j is assigned the high system load and 0 if the low load). SAj is an indicator
variable denoting participant j’s algorithm treatment assignment (1 if j is assigned the superior
algorithm and 0 if the inferior). ITi,j equals the time between customer i − 1 and customer i’s
arrivals to the system, minus the average inter-arrival time of customers in participant j’s system
load treatment assignment. Ci represents customer controls: the average and range of customer i’s
four sample joke ratings. Pj represents participant controls: attributes of participant i, including
gender and college major, collected from the post-experiment questionnaire (described in Section
3.5). Across the regressions, we use robust standard errors clustered at the participant level unless
otherwise specified. We exclude the no-algorithm control treatment results from this regression
because it is not possible for participants to use the algorithm’s advice there.
Table 3 shows the results from this regression model for our three measures of algorithm use:
consulting, following, and defaulting to the algorithm (see Table B.2 in the Appendix for the full
customer and participant control details). The regressions confirm that system load significantly
affects algorithm use, for any definition of use. Participants are significantly, and substantially, more
likely to use both algorithms under higher system loads. The odds of a participant using the inferior
algorithm are more than 1.7 times higher if this participant faces a high system load than a low
Snyder, Keppler, Leider: Algorithm Reliance, Fast and Slow
15
Odds ratios
High system load treatment (0/1) 1.862 1.729 2.238
Superior algorithm treatment (0/1) 1.582 1.650 1.590
Load × algorithm treatment interaction .8623 .948 1.114
system load (β1 : p < 0.05) and the results are similar for participants with the superior algorithm
(β1 + β3 : p < 0.05). Relatedly, participants are reactive to variations in arrivals within treatments,
driven by random inter-arrival times. They are significantly more likely to use their algorithm to
serve a customer if this customer arrived quickly after the previous customer (p < 0.01).
Participants also respond to the quality of their algorithm’s advice. However, their exact response
to algorithm quality varies by system load; while algorithm quality affects participants’ choice
to consult and follow the algorithm under low loads, it affects participants’ choice to default to
the algorithm under high loads. Participants facing low system loads are significantly more likely
to consult and follow the superior than the inferior algorithm (β2 : p < 0.05 for these measures),
while participants facing high system loads are significantly more likely to default to the superior
algorithm’s advice than the inferior algorithm’s (β2 + β3 : p < 0.05 for this measure, β2 + β3 : p < 0.1
for the “follow” measure). This suggests that the interaction of algorithm quality and system load
affects not only whether, but also how participants use the algorithm.
Snyder, Keppler, Leider: Algorithm Reliance, Fast and Slow
16
load on superior algorithm use, this behavior does not result in comparable time savings. In fact,
participants’ average throughput times are actually slightly (insignificantly) slower in the inferior-
than in the no-algorithm treatments. There is another, additional explanation for our findings; par-
ticipants in the high-load, superior-algorithm treatment are using the algorithm differently, namely
faster, than participants in the other treatments7 .
We test this conclusion with a counterfactual simulation of throughput times in the high-load,
inferior-algorithm treatment to estimate the outcome we would expect if participants in that treat-
ment used their algorithm in the same way (but not at the same rate) as participants in the
high-load, superior-algorithm treatment. To construct counterfactuals, we assume that partici-
pants’ service times would follow the same distribution in both treatments, conditional on their
choice to follow the algorithm’s advice. Based on this assumption, we replace the service time for
every instance in which a participant in the high-load, inferior-algorithm treatment follows the algo-
rithm’s advice with a random draw from an exponential distribution with mean 3.6 seconds. This
random draw approximates participants’ actual service times in the high-load, superior-algorithm
treatment, given they followed the algorithm’s advice. We use the updated (hypothetical) service
times and customers’ arrival times to calculate every resulting waiting and throughput time.
According to our simulation, if participants in the high load, inferior algorithm treatment followed
the algorithm in the same way as participants in the high load, superior algorithm treatment do (but
still at a rate of 39%), their average throughput time would be about 15.2 seconds, instead of 17.5
seconds. In other words, the actual difference in throughput times between the high-load, inferior-
17.5−12.5
and high-load, superior-algorithm treatments is 85% larger (= 15.2−12.5
) than we would expect if
participants used both algorithms in the same way. Participants’ rate of following the algorithm
15.2−12.5
explains only half (54% = 17.5−12.5
) of the increase in throughput times from the superior relative
to the inferior algorithm. The remaining half (46%) of this difference is explained by the speed with
which participants in the high-load, superior-algorithm treatment follow their algorithm’s advice.
7
See Table B.4 in the Appendix for summary statistics about participants’ service times.
Snyder, Keppler, Leider: Algorithm Reliance, Fast and Slow
18
and service time variation) but generally much longer TW than their low-load counterparts, and
T is generally much longer as a result. This result is unsurprising based on what we know from
previous behavioral queueing research; workers respond to system load by working more quickly,
but queues are longer because of the faster customer arrival rate.
Waiting, 𝑇𝑇𝑤𝑤
Service, 𝑇𝑇𝑠𝑠
However, there is one exception to our finding: when participants have advice from the superior
algorithm, TW is only 0.89 seconds longer in the high-load than the low-load condition, and this
difference is not significant (Mann-Whitney U test for subject-level averages: p = 0.711). For com-
parison, the difference between TW in the high- versus low-load treatments is more than six seconds
in the no-algorithm and inferior-algorithm treatments (p = 0.054 for the no-algorithm treatment,
p = 0.018 for the inferior algorithm treatment). This result is driven by changes to service times
in the high-load treatment. For participants facing the low system load, there is no significant
or substantial difference in TS , TW , or T when participants have the superior algorithm versus
no algorithm (e.g., for TW , p = 0.195), or versus the inferior algorithm (p = 0.123). In contrast,
participants facing the high system load are approximately one second faster when they have the
superior algorithm than no algorithm or the inferior algorithm (p < 0.006 for both). This one-
second improvement to service times translates to a more than 4.5-second improvement to wait
times compared to both the no-algorithm and inferior-algorithm treatments (again, p < 0.006).
Participants’ service times have implications for other queueing system-level outcomes like idle
time and queue lengths. We report summary statistics about these measures in Table B.3 in the
Appendix. As expected, utilization is consistently lower in the low-load conditions than in the high-
load conditions (between 53-58% versus between 61-73%, respectively)—relatedly, participants
have more idle time in the low-load conditions than in the high-load conditions. However, partici-
pants in the high-load, superior-algorithm condition experience experience relatively more idle time
Snyder, Keppler, Leider: Algorithm Reliance, Fast and Slow
19
38.6%−28.8%
than participants in the other high-load conditions (about 34%= 28.8%
more), because their
service times are faster. The queueing system is relatively stable in both the low- and high-load
conditions; the longest average queues are in the high-load, inferior-algorithm condition and even
there the average queue length does not get above 2.3 customers.
We might assume that participants under low load are not working any more quickly with the
(superior) algorithm because they already make decisions sufficiently quickly to earn the maximum
time payoffs. However, even participants under low system loads have room for earnings improve-
ment. Participants could earn over 13% more than they do by serving every customer i quickly
enough to earn the maximum time payoff amount ($0.02)8 . They would do better by using their
algorithm more, and faster, like participants in the high-load, superior-algorithm condition.
The variable names here follows the notation defined for the logistic regression equation, Equation
1 above, except that our outcome of interest is participant j’s service time for customer i, TSi,j .
Table 4 shows the results of this regression, conditioned on participant j’s choice to consult (or
not) the algorithm’s advice for customer i. We separate the regressions in this way to disentangle
participants’ service times when using the algorithm from the rate at which they use the algo-
rithm; across all algorithm treatments, participants’ service times are faster when they consult the
algorithm than when they do not (see Table B.4 in the Appendix). Note that for both regressions,
we include only participants with the superior or inferior algorithm for the sake of comparison, as
participants in the no-algorithm treatment cannot ever consult the algorithm’s advice9 .
Table 4 shows participants take significantly less time (more than one second less) to serve
customers in the high-system-load treatment than the low-load treatment when they do not consult
the algorithm’s advice (p < 0.01). This result is expected, as participants must work more quickly
to keep pace with customer arrivals under high system load. Also unsurprisingly, the quality of
the algorithm has no effect on the time it takes participants to serve customers, assuming they
are not consulting its advice. However, conditional on consulting the inferior algorithm’s advice,
participants are not significantly faster at serving customers in the high-load treatment than the
low-load treatment (p = 0.290). Nor are they faster at serving customers when they have superior
8
A small number of participants in each treatment earn this amount for every customer, indicating that it is possible.
9
We replicate the “Consult==0” regression for all treatments in Table B.5 in the Appendix.
Snyder, Keppler, Leider: Algorithm Reliance, Fast and Slow
20
algorithmic advice than inferior advice under low loads (p = 0.368). However, participants in the
high-load, superior-algorithm treatment are significantly faster than participants in the low-load,
superior-algorithm treatment (β1 + β3 : p < 0.01) and than participants in the high-load, inferior-
algorithm treatment (β2 + β3 : p < 0.01)—conditional on consulting the algorithm. That is, it is
the interaction of system load and algorithm quality that induces faster algorithm use. The same
findings persist when we condition the regression on following (or not) the algorithm’s advice.
Replicating the “Consult==0” regression, and also a “Follow==0,” regression for all participants
(including participants in the no-algorithm condition) gives similar results (see Table B.5 in the
Appendix), with one new finding: conditional on not following the algorithm’s advice, participants’
service times are about 0.8 seconds faster in the no-algorithm treatment (p < 0.075) than either of
the algorithm treatments. It may be that the presence of the algorithm creates additional cognitive
load; subjects with an algorithm have two decisions—whether to consult their algorithm and which
recommendation to give to their customer—whereas subjects without the algorithm have only one.
A caveat to any interpretation of these regressions is that participants’ decision to consult and
follow the algorithm is not random, and this choice may be associated with the time they are
willing to spend on their service decision.
Snyder, Keppler, Leider: Algorithm Reliance, Fast and Slow
21
Follow
Default
We test this more formally with logit regressions. These regressions follow the same form as our
regressions of algorithm use on treatment assignment (Equation 1) but with the addition of a new
indicator variable, periodi,j , which equals 1 if participant j serves customer i in the second half
(Periods 3–5) of the experiment. We interact this term with HLj , SAj , and HLj × SAj to learn
how algorithm use over time varies by treatment. Table B.6 in the Appendix shows the results for
this regression, and these results reinforce Figure 6; the effect of time (as represented by periodi,j )
on participants’ defaulting behavior is significantly greater in the high-load, superior-algorithm
treatment than in the low-load, superior-algorithm treatment (β5 + β7 = 0.396 : p < 0.05) or the
high-load, inferior-algorithm treatment (β6 + β7 = 0.352 : p < 0.05). This means that participants in
the high-load, superior-algorithm treatment serve customers even more quickly than participants
in the low-load, inferior-algorithm treatment (β5 + β7 = −0.378 : p = 0.086) and the high-load,
Snyder, Keppler, Leider: Algorithm Reliance, Fast and Slow
22
inferior-algorithm treatment (β6 + β7 = −0.346 : p = 0.096) do over time (see Table B.7 in the
Appendix).
5.3. Participants
In this extension, everything else about our experiment design is kept the same (see Section 3),
and all participants receive advice from the superior algorithm. We recruited 167 new participants
for this extension between September 2021 and January 2022, following the same protocol described
in Section 3.410 . Table 5 shows the division of these participants into the four treatments11 .
2.75
HL x UI
LL x UI
HL x SA LL x ME
LL x SA
2.25 HL x ME
Rating
LL x NA
HL x NA
1.75
LL x IA
HL x IA
1.25
18 15 12 9
Throughput time
Figure 7 Performance Across All Treatments: Rating Quality and Throughput Time
Note. HL and LL are load abbreviations. NA, SA, and IA are algorithm abbreviations for our
original treatments. UI and ME are abbreviations for the up-front information and mandated
experience interventions (participants exposed to these interventions get advice from the SA).
of 2.7 and a throughput time of 9.5 seconds under low load (LL × UI). These participants’ rec-
ommendations are rated significantly higher than participants’ recommendations in the original
superior-algorithm treatments (p = 0.07 in the high-load comparison and p = 0.002 in the low-load
comparison). The up-front information intervention also significantly improves throughput times
under low load (p = 0.004). Participants with mandated experience using the algorithm see on
average a rating of 2.3 and a throughput time of 12.4 seconds under high load (HL × ME), and
a rating of 2.6 and a throughput time of 10.7 seconds under low load (LL × ME). Under high
loads, mandated experience has no significant effect on rating or throughput time compared to
the original superior-algorithm treatment. Under low loads, mandated experience results in higher
ratings (p < 0.05) and faster throughput times (p < 0.05). See Tables C.1 and C.2 in the Appendix
for details. In summary, both interventions yield some performance improvement compared to the
original, no-algorithm results, except mandated experience under high load.
The two interventions improve performance because they lead to greater rates of defaulting to
the algorithm, especially by participants in the low-load setting. To formally analyze the effect of
the two interventions on participants’ choices to consult, follow, and default to the superior algo-
rithm, we estimate logit regressions (described in Table C.3 in the Appendix). Up-front information
changes the way participants rely on the algorithm by encouraging greater use from the start. In
the low-load condition, up-front information significantly boosts participants’ rates of consulting
(β3 = 1.144, p < 0.01), following (β3 = 0.975, p < 0.01), and defaulting to (β3 = 0.967, p < 0.01) the
superior algorithm, almost to the level of their high-load counterparts. In the high-load condi-
tion, up-front information leads participants to default significantly more to the superior algorithm
Snyder, Keppler, Leider: Algorithm Reliance, Fast and Slow
26
(β3 + β5 = 0.585, p < 0.05). In fact, this intervention leapfrogs the dynamics of the early periods
such that participants in the high-load treatment are immediately defaulting to the superior algo-
rithm’s advice as if they were already in a later period (participants in the original high-load,
superior-algorithm treatment defaulted to the algorithm’s advice for 30% of customers in Period 5,
up from 21% of customers in Period 1, whereas participants in the high-load, up-front-information
treatment defaulted to the algorithm’s advice for 38% of customers already by Period 1). On the
other hand, only participants in the low-load condition respond to our mandated-experience inter-
vention by significantly changing (increasing) their use of the superior algorithm. Participants in
that treatment do default to (β2 = 0.623, p < 0.05) the superior algorithm more than participants
in the original, no-intervention treatment, although the magnitude of the effect is smaller than the
magnitude of the effect of up-front information. However, participants in the high-load, mandated-
experience treatment do not respond to this intervention by changing their algorithm use behavior
in a meaningful way (for example, for defaulting to the algorithm: β2 + β4 = −0.120, p = 0.663).
6. Discussion
We have shown algorithm quality and system load are mutually reinforcing for throughput time
performance, so superior algorithms may not always improve worker speed and efficiency—even
with explicit incentives. Here, we describe theoretical and practical implications of this conclusion.
Snyder, Keppler, Leider: Algorithm Reliance, Fast and Slow
27
4.5.1, participants treat their heterogeneous customers differently, spending more time on certain
customers and using their algorithm more often to serve others, but this strategy is ineffective. This
finding supports previous research which shows that even when people can identify algorithms that
are on average good or on average bad, they often cannot identify specifically when an algorithm’s
advice will be helpful; they follow it too much when the advice is poor and follow it too little
when it is good (Balakrishnan et al. 2022). However, we also see that as participants facing high
system loads move towards defaulting to the superior algorithm in later periods, they react less
to customer’s sample ratings. In some sense, this may reflect a movement away from efforts to
personalize the service, as workers themselves spend little or no time reviewing customer details
before outsourcing the decision to the algorithm.
Our findings offer managerial insights about how to implement algorithms for service scale.
Managers ought not assume that workers will use algorithms for efficiency gains under any system
load. If initial algorithm pilots reveal disappointing throughput time improvements, this does not
necessarily mean that algorithms cannot help the company increase its scale. It may simply mean
that system loads are too low to see results, or that the algorithm needs improvement. Likewise,
managers are unlikely to benefit from introducing a mediocre algorithm for the sake of service
efficiency and scale. This could bring about the worst of both worlds—poor service quality without
faster service. We offer one caveat to the claim that higher system loads will improve efficiency
gains: in our experiment, participants were no better than random at the task—therefore they did
not evidence the speed-quality tradeoff present in many service decisions. In settings where workers’
decision quality improves with decision time, increasing system loads may reduce workers’ efforts
to provide valuable personalization. Alternatively, workers may prioritize quality, so the efficiency
gains from higher system loads could be less.
7. Conclusion
In this paper we investigate how workers in queueing systems engage with algorithmic decision-
support to manage system loads. We have shown with an experiment that high system loads and
superior algorithm quality both separately cause subjects to use the algorithm more for service
decisions. However, these two effects cannot fully explain why better algorithms plus high system
loads improve participants’ throughput times a lot, while better algorithms alone improve through-
put times only a little. Another part of the explanation is that the interaction of high system loads
and superior algorithm quality influences how participants use the algorithm; high loads plus supe-
rior algorithms leads to fast algorithm use (defaulting to the algorithm). In an extension, we show
that interventions for better system performance also induce defaulting behavior, especially when
system loads are low. Our results suggest that participants learn about their algorithm’s quality
as they use it, and this, with our other findings, has implications for practice including about how
to pilot new algorithms and how to use algorithms effectively to scale up services.
We view our study as a first step to understanding the managerial and theoretical implications
of worker-algorithm interactions within service queueing systems. There is much more to learn.
One interesting extension would be to manipulate the rating and time payoffs directly, to shed
more light on the question of which dimension (decision quality or efficiency) is the most important
driver of algorithm use. Another would be to replace the joke-recommendation task with another
that creates a speed-quality trade-off. Follow-up experiments varying system load over time are yet
another way to extend our work; we see from our experiment that high loads induce workers to
learn to default to the superior algorithm, but it remains to be seen if workers continue to default
Snyder, Keppler, Leider: Algorithm Reliance, Fast and Slow
30
to this algorithm when system loads decrease, and we cannot know this from our results alone. If
short-term exposure to high system loads does promote fast algorithm use, then companies might
benefit from implementing algorithms while simultaneously increasing system loads. However, this
may create a problem: if the algorithm’s advice is poor, then increasing system load also increases
the number of customers exposed to worse service. How companies should weigh this risk against
the possible benefits from learning presents another direction for future experiments and theory.
References
Allon, G, M Kremer. 2018. Behavioral foundations of queueing systems. The handbook of behavioral opera-
tions 9325 325–366.
Avshalomov, Z, K VanVliet, S Horn. 2022. Data processing system and method for dynamic assessment,
classification, and delivery of adaptive personalized recommendations. U.S. Patent 11,256,873 B2.
Balakrishnan, M, K Ferreira, J Tong. 2022. Improving human-algorithm collaboration: Causes and mitigation
of over- and under-adherence. Working Paper 1–30.
Bastani, H, O Bastani, WP Sinchaisri. 2020. Learning best practices: Can machine learning improve human
decision-making? Working Paper 1–30.
Batt, RJ, C Terwiesch. 2017. Early task initiation and other load-adaptive mechanisms in the emergency
department. Management Science 63(11) 3531–3551.
Beer, R, A Qi, I Rı́os. 2022. Behavioral externalities of process automation. Available at SSRN 4295527 .
Boyacı, T, C Canyakmaz, F de Véricourt. 2023. Human and machine: The impact of machine input on
decision making under cognitive limitations. Management Science .
Buell, R. 2018. The parts of customer service that should never be automated. Harvard Business Review .
Cao, X, D Zhang. 2020. The impact of forced intervention on ai adoption. Available at SSRN 3640862 .
Caro, F, A Saez de Tejada Cuenca. 2022. Believing in analytics: Managers’ adherence to price recommen-
dations from a dss. Manufacturing & Service Operations Management, forthcoming .
Castelo, N, MW Bos, DR Lehmann. 2019. Task-dependent algorithm aversion. Journal of Marketing Research
56(5) 809–825.
Dai, T, M Abramoff. 2023. Incorporating artificial intelligence into healthcare workflows: Models and insights.
INFORMS TutORials in Operations Research .
Delasay, M, A Ingolfsson, B Kolfal, K Schultz. 2019. Load effect on service times. European Journal of
Operational Research 279(3) 673–686.
Dietvorst, B, J Simmons, C Massey. 2015. Algorithm aversion: People erroneously avoid algorithms after
seeing them err. Journal of Experimental Psychology: General 144(1) 114–126.
Dietvorst, BJ, JP Simmons, C Massey. 2018. Overcoming algorithm aversion: People will use imperfect
algorithms if they can (even slightly) modify them. Management Science 64(3) 1155–1170.
Do, HT, M Shunko, MT Lucas, DC Novak. 2018. Impact of behavioral factors on performance of multi-server
queueing systems. Production and Operations Management 27(8) 1553–1573.
Snyder, Keppler, Leider: Algorithm Reliance, Fast and Slow
31
Duch, ML, MRP Grossmann, T Lauer. 2020. z-Tree unleashed : A novel client-integrating architecture for
conducting z-tree experiments over the internet. Journal of Behavioral and Experimental Finance 28
100400.
Ferreira, K, CT Ryan, S Mehta. 2023. Reup education: Can ai help learners return to college? Harvard
Business School case 624-007 1–25.
Filiz, I, JR Judek, M Lorenz, M Spiwoks. 2021. Reducing algorithm aversion through experience. Journal
of Behavioral and Experimental Finance 31 100524.
Fischbacher, U. 2007. z-tree: Zurich toolbox for ready-made economic experiments. Exp Econ 10 171–178.
Gans, N, N Liu, A Mandelbaum, H Shen, H Ye. 2010. Service times in call centers: Agent heterogeneity
and learning with some operational consequences. Borrowing strength: theory powering applications–A
Festschrift for Lawrence D. Brown, vol. 6. Institute of Mathematical Statistics, 99–124.
Gillis, T, B McLaughlin, J Spiess. 2023. On the fairness of machine-assisted human decisions: Theory and
experimental evidence. arXiv preprint arXiv:2110.15310 .
Goddard, K, A Roudsari, JC Wyatt. 2012. Automation bias: a systematic review of frequency, effect medi-
ators, and mitigators. Journal of the American Medical Informatics Association 19(1) 121–127.
Greiner, B. 2004. The Online Recruitment System ORSEE 2.0 - A Guide for the Organization of Experiments
in Economics. Working Paper Series in Economics 10, University of Cologne, Department of Economics.
Hathaway, BA, E Kagan, M Dada. 2022. The gatekeeper’s dilemma: “when should i transfer this customer?”.
Operations Research 0(0) null.
Hoffman, M, L Kahn, D Li. 2018. Discretion in hiring. The Quarterly Journal of Economics 133(2) 765–800.
Hopp, WJ, SMR Iravani, GY Yuen. 2007. Operations systems with discretionary task completion. Manage-
ment Science 53(1) 61–77.
Huang, MH, RT Rust. 2018. Artificial intelligence in service. Journal of Service Research 21(2) 155–172.
Ibanez, MR, JR Clark, RS Huckman, BR Staats. 2018. Discretionary task ordering: Queue management in
radiological services. Management Science 64(9) 4389–4407.
Ibrahim, R, SH Kim, J Tong. 2021. Eliciting human judgment for prediction algorithms. Management
Science 67(4) 2314–2325.
Jung, M, M Seiter. 2021. Towards a better understanding on mitigating algorithm aversion in forecasting:
an experimental study. Journal of Management Control 32(4) 495–516.
Jussupow, E, I Benbasat, A Heinzl. 2020. Why are we averse towards algorithms? A comprehensive literature
review on algorithm aversion. Proceedings of the 28th European Conference on Information Systems
(ECIS). European Conference on Information Systems, 1–16.
Kagan, E, M Dada, B Hathaway. 2022. Ai chatbots in customer service: Adoption hurdles and simple
remedies. Working Paper 1–44.
Kawaguchi, K. 2021. When will workers follow an algorithm? A field experiment with a retail business.
Management Science 67(3) 1670–1695.
KC, DS, C Terwiesch. 2009. Impact of workload on service time and patient safety: An econometric analysis
of hospital operations. Management Science 55(9) 1486–1498.
Snyder, Keppler, Leider: Algorithm Reliance, Fast and Slow
32
KC, DS, C Terwiesch. 2012. An econometric analysis of patient flows in the cardiac intensive care unit.
Manufacturing & Service Operations Management 14(1) 50–65.
Kesavan, S, T Kushwaha. 2020. Field experiment on the profit implications of merchants’ discretionary
power to override data-driven decision-making tools. Management Science 66(11) 5182–5190.
Kremer, M, F de Véricourt. 2023. Mismanaging diagnostic accuracy under congestion. Operations Research
71(3) 895–916.
Kwon, C, A Raman, J Tamayo. 2022. Human-computer interactions in demand forecasting and labor
scheduling decisions. Available at SSRN 4296344 .
Lehmann, CA, CB Haubitz, A Fügener, UW Thonemann. 2022. The risk of algorithm transparency: How
algorithm complexity drives the effects on the use of advice. Production and Operations Management
31(9) 3419–3434.
Li, J, S Leider, D Beil, I Duenyas. 2021. Running online experiments using web-conferencing software.
Journal of the Economic Science Association 7(2) 167–183.
Lin, W, S-H Kim, J Tong. 2021. Does algorithm aversion exist in the field? an empirical analysis of algorithm
use determinants in diabetes self-management. An Empirical Analysis of Algorithm Use Determi-
nants in Diabetes Self-Management (July 23, 2021). USC Marshall School of Business Research Paper
Sponsored by iORB, No. Forthcoming .
Liu, N, SR Finkelstein, ME Kruk, D Rosenthal. 2018. When waiting to see a doctor is less irritating: Under-
standing patient preferences and choice behavior in appointment scheduling. Management Science
64(5) 1975–1996.
Pettey, C. 2016. Five keys to understanding algorithmic business. Gartner .
Pisano, GP, RMJ Bohmer, AC Edmondson. 2001. Organizational differences in rates of learning: Evidence
from the adoption of minimally invasive cardiac surgery. Management Science 47(6) 752–768.
Sarter, NB, B Schroeder. 2001. Supporting decision making and action selection under time pressure and
uncertainty: The case of in-flight icing. Human factors 43(4) 573–583.
Schultz, KL, DC Juran, JW Boudreau, JO McClain, LJ Thomas. 1998. Modeling and worker motivation in
JIT production systems. Management Science 44(12-part-1) 1595–1607.
Shunko, M, J Niederhoff, Y Rosokha. 2018. Humans are not machines: The behavioral impact of queueing
design on service time. Management Science 64(1) 453–473.
Smith, A. 2018. Public attitudes toward computer algorithms. Pew Research Center, Washington, D.C. .
Sun, J, DJ Zhang, H Hu, JA Van Mieghem. 2022. Predicting human discretion to adjust algorithmic
prescription: A large-scale field experiment in warehouse operations. Management Science 0(0) null.
Tan, TF, S Netessine. 2014. When does the devil make work? An empirical study of the impact of workload
on worker productivity. Management Science 60(6) 1574–1593.
Van Donselaar, KH, V Gaur, T Van Woensel, RACM Broekmeulen, JC Fransoo. 2010. Ordering behavior
in retail stores and implications for automated replenishment. Management Science 56(5) 766–784.
Wickelgren, W. 1977. Speed-accuracy tradeoff and information processing dynamics. Acta Psychologica
41(1) 67–85.
Snyder, Keppler, Leider: Algorithm Reliance, Fast and Slow
33
Appendices
A. Experiment Design Details
In this section of the Appendix, we provide more details and context about our experiment design.
Figure A.1 shows the experimental interface participants see in the original algorithm treatments,
before and after consulting the algorithm’s advice. We describe details of the interface in Section
3.2.1 in the main body of the paper. Figure A.2 shows the experimental interface that participants
in the mandated experience treatment see for the first three customers in each period, when they
must follow the algorithm’s advice. We describe details of this interface in Section 5.2.
Figure A.2 Experimental Interface: Mandated Experience Intervention (first three customers)
Tables A.1 and A.2 show the outcome-(rating and throughput time, respectively)-to-payoff con-
version we use to incentivize decisions in our experiment. We describe the design of our payoff
scheme in detail in Section 3.2.3 in the main body of the paper.
The customers that participants see over the course of the experiment are heterogeneous and rate
the 12 jokes differently from one another. Table A.3 shows summary statistics about the features of
the customers that participants can see from the four sample ratings, and the standard deviation
in these measures for all customers. Because participants see more customers in the high-load
treatments, the summary statistics are slightly different in the high-load treatments than the low-
load treatments. Table A.4 shows a summary of the different ratings customers gave to different
jokes. We discuss customer heterogeneity in more detail in Section 3.2.2 in the paper.
As we describe in Section 3.3.2 in the paper, we used an OLS design to create our algorithms.
Although the design is relatively simple, our superior algorithm performs well against alternative
designs, as Figure A.3 shows.
Figure B.1 shows the average per-customer payoff in each of the six treatments. We provide
details about participants’ rating and throughput time performances underlying these payoffs in
Section 4.1 of the paper. As expected, participants earn more with the superior algorithm and
less with the inferior algorithm. Holding algorithm treatment fixed, participants earn more per
customer under low loads, except in the superior-algorithm conditions.
Table B.2 is the complete version of Table 3 in the paper, including all participant and customer
control results. As a reminder, this table shows results from logit regressions of the form:
P(Algorithm usei,j = 1) denotes participant j’s probability of using the algorithm to serve customer
i. HLj is an indicator variable denoting participant j’s system-load treatment assignment (1 if
j is assigned the high system load and 0 if the low load). SAj is an indicator variable denoting
participant j’s algorithm treatment assignment (1 if j is assigned the superior algorithm and 0 if the
inferior). ITi,j equals the time between customer i − 1 and customer i’s arrivals to the system, minus
the average inter-arrival time of customers in participant j’s system load treatment assignment.
Ci represents customer controls: the average and range of customer i’s four sample joke ratings.
Pj represents participant controls: attributes of participant i including gender and college major
collected from the post-experiment questionnaire. Across the regressions, we use robust standard
errors clustered at the participant level. We exclude the no-algorithm control treatment results
from this regression because it is not possible for participants to use the algorithm’s advice there.
The regressions confirm that system load significantly affects algorithm use, for any definition
of use. Participants are significantly, and substantially, more likely to use both algorithms under
higher system loads. Participants also respond to the quality of their algorithm’s advice. However,
their exact response to algorithm quality varies by system load. Customer controls and participants’
major and gender are also predictive of algorithm use.
Snyder, Keppler, Leider: Algorithm Reliance, Fast and Slow
39
Odds ratios
High system load treatment (0/1) 1.862 1.729 2.238
Superior algorithm treatment (0/1) 1.582 1.650 1.590
Load × algorithm treatment interaction .8623 .948 1.114
Tables B.3 and B.4 contain summary statistics related to the queueing system, for each of the
original treatments. Because B.4 shows some results about participants’ service times conditional
on consulting the algorithm’s advice, these statistics are not applicable for the no-algorithm treat-
ments. We discuss the queueing system in detail in Section 4.3 and Section 4.4 in the main body
of the paper.
Table B.3 shows that the system is relatively stable and queue lengths to not explode even in the
high-load conditions. Table B.4 reinforces that participants’ service times are particularly fast in
the high-load, superior-algorithm condition when participants do consult their algorithm’s advice.
Table B.3 Queueing System Summary Statistics, Including Idle Time and Queue Length
Algorithm No Algorithm Superior Inferior
System Load High Low High Low High Low
Table B.4 Service Time Summary Statistics, Including Times Pre- and Post- Consulting the Algorithm
Algorithm No Algorithm Superior Inferior
System Load High Low High Low High Low
Table B.5 replicates a regression (“Consult==0”) presented in Table 4 in the main body of the
text for all treatments, including the no-algorithm treatments. This regression is an OLS regression
of the form:
The variable names here follows the notation defined for Table B.2 above, except that our outcome
of interest is participant j’s service time for customer i, TSi,j . The regression is conditioned on
participant j’s choice not to consult the algorithm’s advice for customer i. A second regression
in the table takes the same form, but conditioned on participant j’s choice not to follow the
algorithm’s advice.
Replicating the “Consult==0” regression, and also a “Follow==0,” regression for all participants
(including participants in the no-algorithm condition) gives similar results, with one new finding:
conditional on not following the algorithm’s advice, participants’ service times are about 0.8 seconds
faster in the no-algorithm treatment (p < 0.075) than either of the algorithm treatments.
Snyder, Keppler, Leider: Algorithm Reliance, Fast and Slow
42
Table B.5 Service Time Linear Regressions, Including the No-Algorithm Treatment
Consult==0 Follow==0
VARIABLES Service time Service time
Table B.6 below shows algorithm use logit regressions following the same form as the regression
shown in Table B.2, with the addition of a new indicator variable, periodi,j , which equals 1 if
participant j serves customer i in the second half (Periods 3–5) of the experiment. We interact this
term with HLj , SAj , and HLj × SAj to learn how algorithm use over time varies by treatment.
This gives us insight into the effect of time on participants’ algorithm use behavior, which we
discuss in more detail in Section 4.5 of the paper. Similarly, Table B.7 shows results from an OLS
regression of the same form with service times as the dependent variable.
Table B.6 shows that the effect of time (as represented by periodi,j ) on participants’ defaulting
behavior is significantly greater in the high-load, superior-algorithm treatment than in the low-
load, superior-algorithm treatment (β5 + β7 = 0.396: p < 0.05) or the high-load, inferior-algorithm
treatment (β6 + β7 = 0.352: p < 0.05). As Table B.7 shows, this means that participants in the
high-load, superior-algorithm treatment serve customers even more quickly than participants in
Snyder, Keppler, Leider: Algorithm Reliance, Fast and Slow
44
the low-load, inferior-algorithm treatment (β5 + β7 = −0.378: p = 0.086) and the high-load, inferior-
algorithm treatment (β6 + β7 = −0.346: p = 0.096) do over time.
β5 + β 7 -0.378*
(0.219)
β6 + β 7 -0.346*
(0.207)
Observations 45,894
Number of participants 311
R-squared 0.051
Robust standard errors clustered at the subject level
*** p<0.01, ** p<0.05, * p<0.1
In Table B.8, we show a regression of relative superior algorithm performance on customer type.
In particular, we measure the difference between customers’ ratings for the superior algorithm’s
recommendation and their ratings for recommendations by participants in the no-algorithm treat-
ments. We define customer types according to their sample rating average and range; high averages
and ranges are defined as being above the median while low averages and ranges are those below
the median value. This relates to our discussion of participants’ algorithm use strategy for hetero-
geneous customers in Section 4.5.1 of the main text—the algorithm is not relatively better than
participants for any customer type.
Finally, Table B.9 shows a version of the algorithm use regressions (e.g., see Table B.2 above)
where period is interacted with the customer rating average control. This again gives information
Snyder, Keppler, Leider: Algorithm Reliance, Fast and Slow
45
about participants’ algorithm use behavior over time for heterogeneous customers: the effect of the
customer rating × period interaction is positive and significant—0.213**—only in the regression
of participants’ algorithm use for the high-load, superior-algorithm treatment.
Observations 12,576
R-squared 0.002
Standard errors clustered at the customer level
*** p<0.01, ** p<0.05, * p<0.1
Table B.9 Algorithm Use Logit Regressions with Customer Rating-Period Interaction
HL × SA LL × SA HL × IA LL × IA
VARIABLES Follow (0/1) Follow (0/1) Follow (0/1) Follow (0/1)
Table C.1 Recommendation Rating Regressions: Intervention versus Original Superior-Algorithm Results
HL × UI LL × UI HL × ME LL × ME
VARIABLES Rating Rating Rating Rating
Table C.2 Throughput Time Regressions: Intervention versus Original Superior-Algorithm Results
HL × UI LL × UI HL × ME LL × ME
VARIABLES Thruput time Thruput time Thruput time Thruput time
To analyze the effects of the two interventions and system load on algorithm use more formally,
we estimate logit regressions of the form:
P(Algo. usei,j = 1) = F (β0 + β1 HLj + β2 MEj + β3 UIj + β4 HLj × MEj + β5 HLj × UIj + ...).
This regression follows the same form outlined in Equation 1 in our paper, our original “algo-
rithm use” logistic regression, with two new variables in place of the algorithm quality treatment
indicator—we measure the effect of interventions only on participants with the superior algorithm:
MEj is an indicator variable denoting whether participant j is assigned to the mandated experi-
ence intervention, and UIj is an indicator variable denoting whether participant j is assigned to
the up-front information intervention. We exclude participants in the original no-algorithm and
inferior-algorithm treatments from this regression. The results of this regression are as follows:
Table C.3 reveals that up-front information changes the way participants rely on the algorithm
by encouraging greater use from the start. On the other hand, only participants in the low-load
condition respond to our mandated-experience intervention by significantly changing (increasing)
their use of the superior algorithm.
We repeat this analyses for a randomly-selected subset of 36 participants from the low-load,
mandated-experience treatment, to ensure our results are not driven by the randomly larger number
of subjects in that treatment (due to some randomness in the recruiting process). Table C.4 shows
the results (and it confirms that the results are not driven by the subject count).
Table C.4 Algorithm Use Logit Regressions: 36 (Random) Subjects from the LL × ME Treatment
VARIABLES Consult (0/1) Follow (0/1) Default (0/1)
We do not see significant differences in most customer controls except the “STEM” indicator,
which is only marginally significant (p = 0.073) and not large in magnitude. Note that participants
in the inferior-algorithm treatments are slightly more likely to be STEM majors and participants
majoring in STEM are more likely to use their algorithm’ advice; if anything, this means the effect
of algorithm quality on algorithm use is likely to be even larger than what our results show.
We collected data about our interventions in one wave after our original superior-algorithm
treatments. However, in our analyses (see Section 5.4 in the paper) we compare participant behavior
in the high-load, superior-algorithm treatment from our original study to participant behavior in
our intervention treatments (collected several months later). To confirm that our subject pool did
not change in measurable ways over this time, we conduct balance checks for each of the participant
control variables for all participants with the superior algorithm. Specifically, we regress each
participant control on the superior-algorithm (SA), no intervention treatment indicator (which we
choose because it also serves as a “first wave” indicator). Table D.2 shows the results from these
regressions; we do not see significant differences for any customer control.
In addition to balance checks, we conducted other analyses about our experiment design. In
developing the experiment we considered alternative design choices; the following results show that
our results are generally robust to these choices. First, D.3 shows that our results are generally
robust to our definition of “defaulting”—we find similar results to the ones we describe in Table B.2
above for service-time cutoffs of two seconds and four seconds (versus three seconds, the threshold
we use to define “defaulting” in our experiment). Second, Table D.4 shows that our results are
consistent when there is no button hiding the algorithm’s advice. We tested a version of our
experiment where the algorithm’s advice was automatically made visible to participants without
any button click for the superior algorithm treatments. In these treatments, we see that the high
load treatment induces defaulting, just as it does when the algorithm’s advice is hidden.
Table D.3 Algorithm Defaulting Behavior: Logit Regressions—Alternate Ts i Thresholds for Defaulting
VARIABLES < 2 seconds < 4 seconds
Observations 19,213
Number of participants 109
Robust standard errors in parentheses
*** p<0.01, ** p<0.05, * p<0.1
Snyder, Keppler, Leider: Algorithm Reliance, Fast and Slow
52
As each customer arrives to receive a joke recommendation, participants in our experiment (in
the superior- and inferior-algorithm treatments) are immediately faced with a choice: they can
default to the algorithm straight away, or they can plan to consult the algorithm and deliberate—
weighing the algorithm’s advice against their own planned recommendation, or they can ignore
the algorithm outright. Participants are incentivized to make the choice that will give them the
highest expected payoff (from their recommendation quality and speed) for each customer. Of
course, participants’ real choice is more complicated; they can spend time deciding whether they
will ignore the algorithm or not (and indeed, Table B.4 suggests participants do take time before
consulting their algorithm)—and they can decide how much time to spend at this stage. However,
by representing the set of choices as finite, we can use a discrete choice model learn something
about how participants view the choice to consult and default to their algorithm under different
system loads and with algorithms of different quality (e.g., see Liu et al. 2018).
We define participants’ initial three options as: 1) default to the algorithm, 2) consult the algo-
rithm without defaulting to it, and 3) ignore the algorithm. We define the options as such because
participants’ ultimate choices are observable and because we find that fast decisions (in particular,
the choice to default to the algorithm) play an important role in participants’ throughput time
performance in each treatment. We do not consider separately participants’ choices to follow or
deviate from the algorithm after consulting it because it is difficult to know whether a participant’s
choice to follow the algorithm is driven by their intention to follow the algorithm’ advice or if it is
just a coincidence that the algorithm’s recommendation matched their own joke recommendation.
We estimate the expected utility of each option y as seen by participant j as:
In this equation, payjy is participant j’s expected total payoff for choosing option y—calculated as
the average payoff from choosing option y based on experience with previous customers. algoy is an
indicator that option y involves consulting the algorithm; algoy is 1 for the options to default and
consult the algorithm, and 0 for the option to ignore the algorithm. defaulty is an indicator that y
is the defaulting option (1 if so, 0 otherwise). We do not include treatment indicators directly in the
model specification because we do not want treatment-specific coefficients. Instead, we incorporate
the treatments via payjy . The expected payoff from each choice depends on whether the algorithm’
advice is on average superior or inferior, and whether the system load is high or low.
Snyder, Keppler, Leider: Algorithm Reliance, Fast and Slow
53
We calculate payjy using past average rating and time payoff values for each choice, including
values seen during the practice period (although we exclude this unincentivized period from our
analysis). Because participants see separate feedback about their rating and time performance, we
assume that they can keep track of these values separately. Therefore, we assume that participants
can infer something about the value of defaulting to the algorithm even if they have only previously
followed it after deliberation, because they have information about the rating payoff that resulted
from matching the algorithm’s advice. To simplify the data generation process, we also assume
that participants do not consider the downstream effect of making choice y to serve customer i on
the time payoff for customer i + 1 (or customer i + 2, etc.). Further, we assume that participants’
prior belief about the rating and time payoffs associated with each option are $0.00. We use these
assumptions in creating our dataset for the model.
The choice to consult the algorithm with deliberation includes the possibility of following or
defaulting to the algorithm, and so we calculate payjy for this option with a weighted average:
we calculate for each participant the probability that they will follow the algorithm’s advice after
consulting it (based on their past behavior), and multiply this probability by the expected payoff
of following the algorithm after consulting it. We add this value to the product of the probability
that the participant deviates from the algorithm’s advice and the expected payoff from deviation.
For our analysis, we fit a multinomial log-linear model using the package nnet and the function
multinom in R. We present the results below in Table E.1. Unsurprisingly, the coefficient β1 is
β1 : payjy 30.209***
(0.294)
β2 : algoy indicator (0/1) -1.345***
(0.015)
β3 : defaulty indicator (0/1) -0.385***
(0.017)
Constant -3.581***
(0.028)
positive and significant. This means that all else being equal, the probability a participant chooses
Snyder, Keppler, Leider: Algorithm Reliance, Fast and Slow
54
y goes up as the historical average earnings associated with y increases. The magnitude of this
coefficient is large, but recall that participants’ earnings are small, on average less than $0.11. More
notably, both β2 and β3 are negative and significant; the magnitude of these coefficients means
that for participants, consulting the algorithm’s advice has a disutility equal to $0.045, (which
0.045
is 43.6% = .102
of participants’ average per-customer payoff in the high-load, superior-algorithm
treatment). Defaulting to the algorithm’s advice has a disutility equal to $0.057 (which is 56.1%
0.057
= .102
of participants’ average per-customer payoff in the high-load, superior-algorithm treatment).
This suggests that participants are averse to consulting the algorithm’s advice, and even more
averse to defaulting to the algorithm’s advice.
As a byproduct of creating the dataset for the discrete choice model, we have records of par-
ticipants’ historical average earnings for each choice over time. We present a lowess smoothing of
these records, as they evolve from the start of Period 1 to the end of Period 5 in each treatment,
in Figure E.1 below. Note that the “Consult” line in each graph indicates the choice to consult the
algorithm without defaulting to it.
In all the graphs, it appears that participants earn more from consulting the algorithm without
defaulting to it than they do from defaulting to it, which would be surprising given our results that
participants benefit from defaulting to the superior algorithm. In reality, this result is driven by
the fact that many participants rarely consult the algorithm’s advice, and therefore do not update
their prior belief about the expected payoff from following the algorithm’s advice. Participants
who infrequently consult the algorithm’s advice will have a higher expected payoff from consulting
the algorithm than defaulting to it because consulting the algorithm leaves open the possibility of
deviation.
Figure E.1a, which shows the results for the high-load, superior-algorithm condition, indicates
that participants learn from the practice period that they will not earn less from consulting or
defaulting to the algorithm as compared to ignoring it. As time goes on, participants learn from
feedback that consulting and defaulting to the algorithm’s advice are actually better than ignoring
it. This learning combats participants’ bias against consulting and defaulting to the algorithm.
Figure E.1b, which shows the results for the low-load, superior-algorithm condition, reveals that
participants have a higher expected earning from ignoring the algorithm compared to the high-
load, superior-algorithm condition, which makes it harder to see that following the algorithm’s
advice is beneficial. Figures E.1c and E.1d show that participants see as early as the practice period
that they will earn more from ignoring the algorithm’s advice, and the gap between each choice is
relatively stable over time—which is unsurprising given our finding that participants do not change
their rates of algorithm use after the practice period.
Snyder, Keppler, Leider: Algorithm Reliance, Fast and Slow
55
.1
.05 .06 .07 .08 .09
.04
0 50 100 150 200 0 50 100 150
Customer number Customer number
(a) High Load, Superior Algorithm (b) Low Load, Superior Algorithm
.1
.05 .06 .07 .08 .09
.04
(c) High Load, Inferior Algorithm (d) Low Load, Inferior Algorithm
Figure E.1 Lowess Smoothing: Historical Avg. Payoff for Each Choice Across Periods 1-5, by Treatment
Snyder, Keppler, Leider: Algorithm Reliance, Fast and Slow
56
ReUp Education (ReUp) is a success coaching service that helps students who have left college
(called college stop-outs) to re-enroll and graduate with the help of one-on-one support from its
coaches. Our relationship with ReUp began in 2020, during the idea-generation stage of this paper.
At that time, ReUp had implemented an algorithm behind-the-scenes to help coaches manage their
workloads, and was preparing to implement a second algorithm, the Personas algorithm (Avshalo-
mov et al. 2022). The Personas algorithm was designed to help coaches make decisions about how
to deliver personalized service to each customer, by classifying students (customers) into Personas
and using these Personas to develop personalized support-plans which would include recommenda-
tions for what to discuss with each student in each conversation. In theory, the Personas algorithm
would be valuable to ReUp Education because it was designed to promote efficient, high-quality,
personalized service. ReUp’s coaches could be assigned to hundreds or thousands of students, so
scale and efficiency were particularly salient issues there.
In the summer of 2020 we conducted formal, semi-structured interviews with two managers
and three coaches at the company to better understand their algorithm tools and day-to-day
operations (ReUp Education was a mid-stage startup with approximately 30 coaches at the time
of our interviews). At the time, managers expressed concerns that coaches were not relying on the
company’s algorithms enough: “I would say, we’re still at a point where coaches are using [the
algorithms] for varying amounts of time in their week,” said one. The coaches reinforced this. They
often explained that they personally do not follow algorithms’ advice for all of their decisions—one
coach disclosed to us that “I’m the person that deviates [from the algorithms] a little bit.” However,
coaches were largely positive about the algorithms’ potential for coaches in general: “I think the
Personas thing is super exciting. It’s kind of [new] for us to see how it will fully impact our work...
But given what I think it’s going to do and given the preliminary engagement that I’ve had with
it and how successful it’s been just even in the version that it’s in now, it’s really exciting to know
that we could better engage with students,” said one.
After our interviews, ReUp conducted a more formal, internal test of the Personas algorithm.
Unfortunately, despite initial excitement and optimism about the technology, ReUp found that the
algorithm did not lead to any improvements in service outcomes (Ferreira et al. 2023).