LLMs and Algorithmic Predation Strategies
LLMs and Algorithmic Predation Strategies
MODELS*
Arina Nikandrova† Anushree Parekh‡
Abstract
This paper explores whether large language models (LLMs) can learn predatory strategies
in dynamic environments in which an incumbent faces repeated entry threat. Using Ope-
nAI’s GPT-4.1 as decision-making agents, we find that LLMs learn to predate when both
predation and accommodation are theoretically viable, and adopt aggressive strategies
when only accommodation is theoretically viable. Further, profit optimization is limited,
highlighting both strategic learning and its limitations.
* We thank Romans Pancs for early inspiration and Fendi Tsim for valuable comments and suggestions on the
code. Financial support from the Centre for Competition and Regulatory Policy at City St George’s, University of
London, is gratefully acknowledged.
†
Department of Economics, City St George’s, University of London, UK, e-mail:
[Link]@[Link].
‡
Department of Economics, City St George’s, University of London, UK, e-mail:
[Link]@[Link].
1 Introduction
This paper is motivated by the growing importance of algorithmic pricing — the use of
advanced artificial intelligence tools to set prices in online markets. While algorithmic pric-
ing promises efficiency gains, it also raises significant concerns for competition policy. The
emerging economics literature has primarily focused on the potential for collusion facilitated
by pricing algorithms (see, e.g., Calvano et al. (2020); Asker et al. (2021); Banchio and Skrzy-
pacz (2022)) or large language model (LLM) agents (see, Fish et al. (2025)). These studies
demonstrate that algorithms can tacitly coordinate to sustain supracompetitive prices, learning
to match collusive benchmarks and punishing deviations through price wars.
However, collusion is only one dimension of the potential harm. Most real-world compe-
tition cases center on exclusionary conduct — that is, strategies by dominant firms to prevent
or deter entry. Despite policy concerns that generative AI may enable exclusionary practices
(Davies (2025); FTC Technology Blog, 2023), no study has yet examined exclusionary con-
duct by algorithmic agents. It is, therefore, not yet understood whether and how learning
algorithms deployed by incumbent firms might engage in predatory pricing, the practice of
pushing market price below cost to discipline or eliminate entrants.
This paper shifts the focus from collusion to predation, using LLM-based simulations of
dynamic markets in which an incumbent faces a stochastic flow of potential entrants. To our
knowledge, it provides the first systematic test of whether, in such settings, LLM agents can
autonomously learn to adopt harmful exclusionary strategies.
Our economic environment builds on Rey et al. (2023), who study the feasibility and prof-
itability of predation in a dynamic infinite-horizon setting in which an incumbent repeatedly
faces potential entry. When a rival enters the market, the incumbent chooses whether to ac-
commodate or predate it; the entrant then decides whether to stay or [Link] our framework,
if entry occurs, firms interact within a Stackelberg framework, where the incumbent chooses
quantity first and then the entrant, having observed the incumbent’s output, chooses its own
output. In contrast to Rey et al. (2023) framework, we introduce an exogenous probability
of entrant exit, which serves to disrupt potential collusive dynamics post entry and creates a
richer environment for studying how predatory strategies may be learned.
In Section 2, we show that in our theoretical framework three Markov Perfect pure strategy
equilibria can arise: (i) Accommodation equilibrium in which the incumbent never predates
and the entrant always enters and stays in the market until exogenous exit takes place. (ii)
Predation equilibrium, in which the incumbent always predates when the entrant enters; the
entrant always enters for exactly one period. (iii) Monopolization equilibrium, in which the
incumbent always predates upon entry and the entrant never enters as entry for one period
is not profitable. A higher probability of birth and of exogenous exit makes accommodation
equilibrium more easily sustainable.
We then use our theoretical framework to analyze the strategic learning process of LLM
agents acting as firms. In all our experiments, OpenAI’s GPT-4.1 serves as the decision-making
engine for our agents.1 . We test whether these agents: (i) learn to choose predatory quantities
when both predation and accommodation are theoretically feasible, and (ii) learn to accom-
modate when accommodation is the only viable equilibrium.
In Section 3, we describe our experimental design, focusing on the nature of our LLM
agents and the prompt engineering. We design prompts to produce parsable responses. In
particular, the LLM agents are instructed in simple language to behave like rational firms,
making quantity decisions with the goal of maximizing long-run discounted profits. The agents
base their decisions on market information as well as self-generated plans from the previous
periods.
Our setup allows us to test whether agents learn from experience and adjust their strate-
gies dynamically over time. A key challenge is that successful predation requires a complex
intertemporal strategy: the incumbent firm incurs short-term losses to drive out competitors,
with the expectation of future gains once rivals are deterred. The profitability of such a strat-
egy hinges on whether the losses are temporary and the reduction in competition is lasting.
Importantly, because predation is costly in the short run, it is not guaranteed that an algorithm
will learn to adopt it even when it is optimal in the long run.
Section 4 presents our core experiment. We stimulate two different strategic environment:
(i) homogeneous goods, where both predation and accommodation are theoretically feasible;
(ii) moderately differentiated goods, where only accommodation is theoretically sustainable.
In both settings, we consider two sub-variations: (a) a potential entrant is born with probability
1/2, and (b) an entrant born with certainty.
We find that, when goods are homogenous and predation and accommodation are both
theoretically viable strategies, the incumbent learns to predate on a newly emerged [Link]
predatory quantity aligns with our theoretical predictions. However, the LLM incumbent adopts
an aggressive strategy also when there is no rival in the market — the incumbent LLM agent
does not learn to fully optimize profits. Furthermore, when products are moderately differen-
tiated, the quantity choices of the LLM agents do not align with theoretical predictions. The
incumbent LLM agent’s strategy converges towards an aggressive stance that is neither fully
predatory nor accommodating. Additionally, when goods are homogenous, the certain birth of
1
Fish et al. (2025), employs OpenAI’s GPT-4 model, and shows GPT-4 learns to optimally set prices and collude.
We use GPT-4.1 released in April 2025 as it is the latest GPT model at the time our experiment runs were conducted.
All experimental runs were conducted between July 2025 and Aug 2025. To understand capabilities of GPT-4.1
see OpenAI (2025). The most recent GPT model now available is GPT-5.
an entrant makes the incumbent more aggressive, whereas with moderate product differenti-
ation, the same certainty increases the volatility of the incumbent’s behavior.
In Section 5, we undertake various robustness checks. We show that in the absence of
exogenous exit, collusion does not emerge; predation is not guaranteed; and accommodation
is achieved only suboptimally. Prompt design matters: alternative prompts can either induce
or mitigate aggressive behavior. We also report results from additional parameter variations.
Overall, our results suggest that while the literature has emphasized the ability of LLM-
based agents (Fish et al. (2025)) and autonomous algorithms (Calvano et al. (2020)) to sustain
collusion, they are equally capable of learning predation. This finding is particularly relevant
for competition authorities and regulators, as it highlights a broader class of algorithmic strate-
gies that may harm competitive process and ultimately consumers.
Relevant literature
This paper contributes to the well-established industrial organisation literature on preda-
tion and other exclusionary behaviours, as well as to the nascent literature on algorithmic
decision-making.
Early literature on predation considered predatory pricing practically irrational and rare.
Particularly, Robinson (1941) considered firms undercutting prices (which we now distinguish
from predatory pricing) as a normal competitive practice between two firms competing for
market share, while McGee (1958) argued that it was more cost-efficient for dominant firms
to acquire smaller firms than to lower prices in order to drive out competition.2
Telser (1966) was the first to warn about firms’ capacity to predate. He formalized the idea
that the feasibility of predation depended on the firm’s financial capacity—the long purse the-
ory. He concluded that the possibility and extent of predation depended on the firm’s capacity
to bear short-term losses. Further, he emphasized that another condition required to sustain
predation was perfect capital markets, as imperfect capital markets would make borrowing dif-
ficult.3 Bolton and Scharfstein (1990) also establish that optimal financial conditions balance
the benefit and cost of predation.
2
McGee (1958) specifically argued that Standard Oil did not achieve dominance through predatory pricing
but through acquisition of smaller firms. This was a widely accepted view.
3
Telser (1966) further tested his theory empirically, to a certain extent, by analyzing concentration ratios of
137 manufacturing industries. He attempted to test the hypothesis that concentrated markets are connected to
perfect capital markets.
Simultaneously, Bain (1949), Clark (1940), and Friedman (1979) emphasized the possibil-
ity of lowering prices to retain monopoly position (by discouraging entry)—which we now call
limit pricing. Milgrom and Roberts (1982) connect the strategy of limit pricing to that of a ra-
tional firm in a two-period model. They argue that limit pricing can arise in equilibrium when
entrants have private information (about, for example, incumbents’ costs) that affects relevant
payoffs. The incumbent can deter entry based on the reputation of being low-cost. Another
model based on information asymmetry and reputation is Kreps and Wilson (1982). These
models, based on the chain store paradox developed by Selten (1978), show that small infor-
mation asymmetries are sufficient to make incumbents’ threats credible, such that established
firms develop the reputation of a “tough” incumbent which deters entry.
Fudenberg and Tirole (1986) provide a varied explanation. Their model suggests that when
entrants are certain only about profitability today and not in the future, existing firms engage
in predatory pricing in order to “jam signals,” that is, to mislead entrants. Similarly, Roberts
(1986) show that information asymmetry about demand uncertainty and realized profits in a
two-period model makes incumbents’ low prices a credible threat leading to predation. See
Salop and Scheffman (1987) and Scharfstein (1984) for other models on misleading entrants.
Another theory of predation is the learning curve hypothesis. Lee (1975) consider the
effect of learning through cumulative output, which may raise entry barriers in a dynamic
limit pricing model. A. Michael Spence (1981) show that late entrants suffer in a quantity-
setting model.4 These models mainly focus on homogeneous goods, deterministic demand,
and Cournot quantity-setting frameworks. Specifically, Cabral and Riordan (1994) model a
dynamic price competition framework. Their main findings reveal that in a duopoly setting
where firms face a sequence of buyers with demand uncertainty, learning increasingly facilitates
dominance. Further, predatory pricing is feasible in learning economies. They show that there
exist Markov Perfect Equilibria where both firms enter the market but the firm that loses sales
(predated) exits.5
Recent studies build on these frameworks.6 Besanko et al. (2014) adopt the model from
Cabral and Riordan (1994), allowing for re-entry, and conduct numerical simulations which
reveal that aggressive pricing arises routinely. They find that aggressive equilibria (predation-
like behavior) coexist with accommodating equilibria. Another theory that has surfaced in the
recent literature is predation based on economies of scale, developed by Fumagalli and Motta
(2013). On the other hand, Toxvaerd (2017) extend the Milgrom and Roberts (1982) model by
4
See Mookherjee and Ray (1989) for a model where learning curves lead to collusion.
5
Cabral and Riordan (1994) model is a simultaneous price-setting game. Where both firms enter, at the start of
a given period they decide whether to stay or exit, incurring a fixed cost. One should note that it is the introduction
of fixed cost that makes their MPE equilibria of predation feasible.
6
For a descriptive overview of different types of theoretical predation models see Ordover and Saloner (1899).
adopting a dynamic repeated-interaction framework where limit pricing arises as an optimal
intertemporal strategy.
We adopt the theoretical framework from Rey et al. (2023), who consider an infinite-
horizon discrete-time model where an incumbent faces a continuous entry threat. They differ
from the literature so far by characterizing conditions under which strategic uncertainty suf-
fices7 to give rise to predation, accommodation, and monopoly equilibria in Markov Perfect
strategies. We build upon their framework by allowing for product differentiation and intro-
ducing the possibility of entrant’s exogenous exit.
Algorithmic decision-making
Over the years, decision-making in firms has shifted from humans to software and now
to sophisticated algorithms. This evolution has raised concerns among policymakers and aca-
demics that algorithms might learn to coordinate prices, potentially sustaining collusion even
without explicit instructions. The seminal work highlighting this possibility in a structured
manner is Calvano et al. (2020). Their experiment simulated reinforcement learning algo-
rithms, specifically Q-learning, in a symmetric duopoly setting. The algorithms learned to
autonomously collude, providing the first clear demonstration that pricing algorithms acting
as firms in an infinitely repeated, simultaneous-move game can achieve collusive outcomes by
learning from past history.8 Importantly, Calvano et al. (2020) show that memory is crucial:
collusive outcomes are only achievable when agents can remember past interactions and adapt
accordingly.
Asker et al. (2021) further differentiated outcomes across settings such as asynchronous
versus synchronous learning and observed that collusive outcomes are feasible only under
asynchronous learning. Calvano et al. (2021) demonstrate that collusion can emerge even
under imperfect monitoring, provided algorithms are allowed to complete their learning pro-
cess. While these studies focus on simultaneous-move games, Klein (2021) show that collusive
outcomes are also possible in sequential-move setups.9
While earlier studies show that collusion in Q-learning algorithms arises through repeated
interaction and convergence to grim-trigger or punishment-based strategies, Banchio and Man-
tegazza (2023) make a novel contribution by introducing the concept of “spontaneous collu-
sion,” where collusive outcomes emerge endogenously from the learning process itself, without
7
Uncertainties such as the probability of an entrant being born and deciding to stay or exit.
8
For a broader overview of collusive outcomes in Q-learning algorithms, see Horton (2023). For implementa-
tion of Q-learning algorithms in industrial organization settings, specifically Cournot models and step-up simula-
tion methods, see Waltman and Kaymak (2008).
9
For empirical examples of collusion emerging from market outcomes, see Assad et al. (2020).
relying on grim-trigger strategies. Bertrand et al. (2025), using a Prisoner’s Dilemma setup, ex-
tend these findings to deep Q-learning algorithms, showing that collusion is not limited to basic
Q-learning. Finally, Banchio and Skrzypacz (2022) explore collusion in auction environments,
finding that Q-learning algorithms can collude in first-price auctions but not in second-price
auctions, highlighting the role of market rules in enabling collusion. For a detailed overview
of the literature on algorithmic collusion see den Boer et al. (2024).
With the rise of firms using Large Language Models as advisors, the question arises whether
LLMs, capable of reasoning, can learn strategic behavior. Horton (2023), pioneer of running
economic experiments with LLMs, specifically with OpenAI-GPT 3, was the first to highlight
the possibility of using LLMs to replicate human-like behavior to run experiments at lower
costs. However, today the concern extends to the possibility of LLMs being used to discover
strategic behavior that may be suboptimal for market conditions. Fish et al. (2025) reinforce
this concern by showing that LLM-based pricing agents are capable of learning to collude in a
duopoly setting as well as an auction setting. Further, they show that certain prompts lead to
collusive outcomes faster. They extend the experiment to the auction setting and find similar
collusive outcomes. Other studies such as Wu et al. (2024) show that LLM agents are capable
of cooperating, Agrawal et al. (2025) show the possibility of collusive actions in double auction
settings. On the other hand, Keppo et al. (2025) show that collusive outcomes in LLM pricing
agents decrease when agents are heterogeneous.
Theoretical literature shows that exclusionary practices such as predation and exploita-
tive practices such as collusion both are possible rational outcomes when firms have strategic
foresight. Yet, the literature on algorithmic decision-making specifically focuses on coordina-
tion between firms. In this paper we seek to bridge the gap between theoretical industrial
organisation and nascent literature on lagorithmic decision making by asking whether LLMs,
when placed in dynamic market environments, can learn predatory strategies — sacrificing
short-term profits to discipline or eliminate entrants — in much the same way that earlier
algorithmic studies showed they could learn to collude.
State M πM
I
, πM
E
−k π̄ M
I
,0
Competitive state C: I and E both exist in the market. First, I announces whether to
predate or to accommodate. Having observed I’s decision, E decides whether to stay or to exit.
If E decides to stay, the firms compete in a Stackelberg game where Firm I acts as the leader
with quantity q I , and Firm E follows by choosing quantity q E . The next period starts in state
C with probability (1 − γ) and in state M with probability γ. Parameter γ is an exogenous
probability that an entrant exits the market for reasons unrelated to the incumbent’s conduct.
If E exits, the next period begins in state M .
Overall, in the competitive state, one of the four possibilities may arise: (i) I accommodates
and the E stays, (ii) I accommodates and the E exits, (iii) I predates and the E stays and (iv)
I predates and the E exits. Table 2 summarized the payoffs of the firms in state C.
E Stay E Exit
Appendix A provides microfoundations for the payoff structure in Tables 1 and 2, based on a
Stackelberg quantity-setting game. While Rey et al. (2023) assumes no product differentiation,
we generalize payoffs by allowing for product differentiation.
Figure 1 provides an overview of the transitions between the states.
Yes
Entrant chooses q EM I , π E − k)
Payoff:(π M M t + 1 state: Competitive
1−η
Payoff:(π̄ M
I , 0)
t + 1 state: Monopoly
State t + 1: Competitive
(1 − γ)
State t: Competitive
(1 − γ)
Entrant decides to stay and chooses q EP Payoff:(π PI , π PE )
γ
I chooses Predate and chooses q IP State t + 1 : Monopoly
Figure 1: Sequence of events in Monopoly and Competitive states respectively and state tran-
sitions.
πAI − π̄ PI
λ= . (1)
I − πI + η πI − πI
(1 − η) π̄ M A M A
The numerator, πAI − π̄ PI , is the profit sacrifice incurred in the predation period and the denom-
inator is the expected monopolization benefit obtained in the next period:
(1 − η) π̄ M − πA
+η π M
− πA
. (2)
| I {z I } | I {z I }
if no new entrant if new entrant
We can similarly define I’s cost-benefit ratio of predation when newborn entrants stay out
of the market:
πAI − π̄ PI
λ̄ = . (3)
π̄ M
I − πI
A
1. Accommodation: Firm I accommodates entry and each newborn entrant E enters the mar-
ket and stays until exogenous exit. Such equilibrium exits if and only if
1 − δ(1 − γ) P
πAE ≥ − πE (4)
δ(1 − γ)
or
δ(1 − γ)
λ≥ . (5)
1 − δ(1 − η − γ)
Everything else equal, accommodation equilibrium is easier to sustain when the probability
of entrant’s exit γ or probability of entrant’s birth η are higher.
2. Predation: Firm I predates in case of entry and each newborn entrant E enters the market
to stay for one period only. Such equilibrium exits if and only if
πM
E
≥k (6)
and
δ(1 − γ)
λ≤ . (7)
1 − δ(1 − η − γ)
Everything else equal, predation equilibrium is easier to sustain when the probability of
entrant’s exit γ or probability of entrant’s birth η are lower.
3. Monopolization: Firm I predates in case of entry and each newborn entrant E stays out of
the market. Such equilibrium exits if and only if
πM
E
≤k (8)
and
δ(1 − γ)
λ̄ ≤ . (9)
1 − δ(1 − γ)
Appendix C.1 includes the full templates for the System and User Prompts used in the
simulation. The prompts update dynamically with specific parameters, history data, and states
12
See Boonstra (2025) to understand various prompt engineering techniques.
13
The inclusion of the demand intercept a ensures that the LLM agents have a market size threshold for making
quantity decisions.
14
The Prompt_incumbent_quantity includes fixed cost z I , while the entrant prompts include fixed cost z E
and the one-time entry cost K.
15
The quantity prompts use common_prompt_suffix_quantity, while the binary prompts use
common_prompt_suffix_binary
during each period t.
4 Experiment Runs
In this section, we turn our attention to the main experimental runs and results we observe.
In order to test whether LLM agents adapt strategic behavior consistent with our theoretical
predictions, we simulate distinct strategic environments. For tractability, we adopt baseline
economics parameters across all runs: demand intercept a = 100, demand slope b = 1, in-
cumbent’s fixed cost z I = 300, entrant’s fixed cost z E = 150, entrant’s entry cost k = 450,
16
The max_tokens parameter limits the length of the LLM’s output for a given call. Specifically, the limit is set
to 1500 tokens.
probability of exogenous entrant exit γ = 0.1, and discount factor δ = 0.9517 .
We then vary parameter θ (degree of product differentiation). This yields two main ex-
perimental variations: (i) Variation 1: θ = 1, where theoretically both accommodation and
predation are viable strategies; (ii) Variation 2: θ = 0.5, where theoretically only accommo-
dation is a viable strategy. Further, within each variation, we vary η (probability of an entrant
being born): (i) η = 1, a new entrant is definitely born; (ii) η = 0.5, a new entrant is born or
not born with equal probability.
Within each sub-variation, we run 10 experimental runs of 300 periods each. To ensure
comparability, we use different random seeds for all 10 runs when η = 0.5, and the same 10
random seeds for all runs when η = 1. Given the limited number of runs, we focus on descrip-
tive and robustness analyses rather than regression-based inference. The insights we draw
from these multiple runs and representative trajectories provide a thorough understanding of
strategic behaviour in the model.
4.1 Variation 1: θ = 1
We first examine the capability of the LLM agent to engage in predation in settings where
both predation and accommodation equilibria are theoretically feasible. Theoretically, In the
Monopoly state, we expect LLM agents to optimize profits by choosing quantity q IM = 50, while
in the Competitive state we expect them to engage in predation by choosing quantity q IP = 75.5
or accommodate by choosing qAI = 50 (here, qAI = q IM ) .
To illustrate the agent’s behaviour, we present results from representative runs with η = 0.5
and η = 1 (Figure 2), using the same random seed for both simulations. In both representative
runs, the Monopoly state (M ) arises in 76% of periods, while the Competitive state (C) arises in
24%. Each simulation begins in the Monopoly state, where an entrant (E) is born. In period 1,
the incumbent LLM agent responds to the threat of entry by choosing a quantity well above
the theoretical monopoly benchmark, consistent with an attempt to deter entry18 .
Figure 3 presents the incumbent’s state-specific quantity choices over time across all exper-
imental runs. Subplot 3a reports outcomes for η = 0.5, while Subplot 3b reports outcomes for
η = 1. Across both parameterizations and in both states, the incumbent consistently selects
17
Note: Choosing parameter values such as a = 100 ensures that quantity decisions are expressed as two-digit
numbers, which are easier for the LLM agent to interpret — for example, “Incumbent sets quantity 60” rather
than “Incumbent sets quantity 0.6”. Our choice of z I = 2z E follows standard IO assumptions, and we set k = 3z E
to reflect that entry is costly.
18
In appendix C.2.1 we present the full log reply from the period 1 of the representative experimental run Figure
subplot 2a. The LLM agent clearly recognize the three viable options: a) to maintain the static monopoly output
of 50, b) to accommodate the entrant, or c) to adopt limit pricing by aggressively setting a quantity above the
monopoly threshold to deter entry. The entrant’s log reply confirms that the agent also recognized this deterrence
attempt and chose NOT ENTER.
(a) Firm quantities: η = 0.5 and θ = 1 (b) Firm quantities: η = 1 and θ = 1
(c) Entrant binary decisions: η = 0.5 and θ = 1 (d) Entrant binary decisions: η = 1 and θ = 1
Figure 2: LLM agent quantity decisions and binary decisions from representative runs when
θ = 1 while η = 0.5 and η = 1. The solid blue line in subplot 2a and 2b represents the
quantity produced by the incumbent firm, while the dashed lines in various colours represent
the quantities of individual entrants who have entered the market. Further, the red dashed
line represents q P , the purple dashed line represents q IM = qAI and the orange dashed line
I
denotes q EM = qAE . In subplot 2c and 2d, the grey dot denotes the birth of an entrant, the red
cross denotes the entrant’s choice to NOT ENTER, green denotes the entrant’s choice to ENTER,
the blue square denotes the entrant’s decision to STAY while the orange triangle denotes the
entrant’s decision to voluntarily EXIT. Lastly, the star denotes exogenous exit. In all subplots,
the background colour indicates the market state: a light coral region signifies a competitive
market, whereas a blue region denotes a monopoly state. The simulation parameters used are:
γ = 0.1, a = 100, b = 1, k = 450, z I = 300, z E = 150, δ = 0.95, θ = 1. In subplot 2a and 2c,
η = 0.5, while in subplot 2b and 2d, η = 1.
quantities that lie above the theoretical monopoly benchmark but below the theoretical preda-
tory threshold. The monopoly benchmark is q IM = 50, i.e., the profit-maximizing monopoly
output, while the theoretical predatory threshold is q IP = 75.51, the minimum quantity consis-
tent with deterring entry.
Averaging across 300 periods and 10 runs per η, the incumbent’s mean quantity is 63.01
in the monopoly state and 66.39 in the competitive state when η = 0.5, compared to the
monopoly benchmark of 50 and the predatory benchmark of 75.51. When η = 1, the mean
quantity rises to 70.33 in the monopoly state and 67.94 in the competitive state, again lying
strictly between the monopoly (50) and predatory (75.51) theoretical benchmarks. These
outcomes suggest that the incumbent strategically expands output above the monopoly level to
deter entry in the monopoly state, and converges toward predatory behavior in the competitive
state.
(a) η = 0.5
(b) η = 1
Figure 3: The figure shows the quantity choices of the incumbent LLM agent across 300 periods
and all ten runs. Simulation parameters: θ = 1, γ = 0.1, a = 100, b = 1, k = 450, z I = 300,
z E = 150, and δ = 0.95. Note that the solid blue line refers to the mean quantity in monopoly
state while the solid red line refers to the mean quantity in competitive state. The dashed red
line represents the theoretical predatory quantity (75.51) while the dashed blue like representa
the theoretical monopoly quantity (50).
Furthermore, the first time the incumbent LLM agent selects a quantity within 10% of q IP
occurs, on average, in period 32.3 when η = 0.5 and in period 28.3 when η = 1. This suggests
that the incumbent does not immediately recognize the sufficient predatory quantity.
Figure 4 presents the distribution of quantity choices during the last 100 periods. The
results confirm that the median quantity is consistently higher in the competitive state than
(a) η = 0.5
(b) η = 1
Figure 4: Box plot of the incumbent’s quantity choice and realized profits in the last 100 pe-
riods. Simulation parameters: θ = 1, γ = 0.1, a = 100, b = 1, k = 450, z I = 300, z E = 150,
δ = 0.95.
in the monopoly state, regardless of η. In the monopoly state, the mean quantity is 65 when
η = 0.5 and 70 when η = 1, both well above the theoretical monopoly benchmark of q IM = 50.
This indicates that the incumbent LLM agent does not maximize immediate profits but instead
seeks to deter potential entrants by signalling aggression. In the competitive state, the mean
quantity is 67.64 when η = 0.5 and 74.12 when η = 1, values that are extremely close to and
within 10% of the theoretical predatory quantity (75.5). This pattern suggests that the LLM
agent responds aggressively to competition and systematically attempts predation.
Further, the incumbent’s profit in the monopoly state remains below the theoretical maxi-
mum profit, π M
I
= 2200, in both cases. This indicates that the incumbent is willing to sacrifice
short-term profit to maintain market dominance. In the competitive state, profits are even
lower, highlighting that predatory strategies are costly for the incumbent.
Further, to evaluate whether the behavior of the incumbent LLM agent aligns with these
theoretical predictions, in monopoly state, we classify the agent’s quantity as optimal or sub-
optimal. The agent’s quantity choice is considered optimal if it is in the 10% range of the
theoretically optimal quantity. In Competitive state, we classify the agent’s quantity decision as
predatory or accommodating if it falls within the 10% region of the corresponding theoretical
quantity. Actions outside these ranges are classified as neither. We find that the incumbent
chooses optimal quantity in the Monopoly state only 16.2% of the time when η = 0.5 and
only 0.99% of the time when η = 1. In the competitive state, the LLM agent engages in
predation 82.23% of the time when η = 0.5 and 93.15% of the time when η = 1. These
results are broadly intuitive: in the competitive state, the agent strongly favours predation,
consistent with the theoretical incentive to deter entry. By contrast, in the monopoly state, the
agent rarely selects the exact theoretical optimum. This pattern suggests that the LLM tends to
“overproduce” even when immediate competitive pressure is absent, perhaps reflecting a bias
toward deterrence or the difficulty of learning the precise monopoly optimum.
Then, for each run, in order to track the strategic trend, we compute a 30-period rolling
average of these classifications19 for each experimental run, separately for the Monopoly and
Competitive states. This provides a dynamic measure of the agent’s propensity to choose op-
timal or predatory strategies over time within each state, rather than by absolute simulation
period. The rolling-window averages are then aggregated across runs to produce the curves
shown in Figures 5 (monopoly state) and 6 (competitive state). Values above the 0.5 line
indicate that optimal (in Monopoly) or predatory (in Competitive) actions are more prevalent.
Figure 5 indicates that in the Monopoly state M , the incumbent does not settle on a stable
strategy when η = 0.5 and performs even worse when η = 1. These results indicate that the
19
We generate a state specific binary variable. For Monopoly state: Optimal=1 and Suboptimal=0. For Com-
petitive state, Predatory=1 and Accommodating=0. The rolling average we calculate is of this binary column.
(a) η = 0.5 (b) η = 1
Figure 5: The share of optimal actions in Monopoly state M when θ = 1. The 0.5 line is the
indifference benchmark — i.e., the point at which the incumbent LLM agent is equally likely to
engage in optimal or suboptimal strategies. The values above the 0.5 line indicate that optimal
actions are more prevalent.
Figure 6: The share of predatory actions in Competitive state C when θ = 1. The 0.5 line is the
indifference benchmark — i.e., the point at which the incumbent LLM agent is equally likely
to engage in predatory or accommodating strategies. The values above the 0.5 line indicate
that predatory actions are more prevalent.
incumbent fails to learn to optimize profits in the Monopoly state, irrespective of the entry
threat. In fact, when the entry threat is definitive, the incumbent makes highly suboptimal
choices toward the end of the simulation.
Further, in the Competitive state C (see Figure 6), we observe that the LLM agent initially
exhibits a bias against predation but quickly learns to choose predatory actions, with the rolling
share of predation crossing the indifference benchmark (0.5 line) and stabilizing well above
it. Predatory behaviour emerges even more rapidly when η = 1, indicating that a higher entry
threat induces greater aggressiveness. This finding runs counter to our theoretical prediction
from Proposition 1, which suggested that increasing the probability of entry, η, would reduce
the attractiveness of predation. Instead, the simulation results suggest that the incumbent LLM
agent finds it beneficial to establish a predatory reputation precisely when entry threats are
high, possibly reflecting the agent’s reinforcement-learning bias toward strategies that ensure
future dominance rather than short-run profit maximization.
Further, to analyze the incumbent’s impact on Entrants, we evaluate two key metrics: the
entrant success rate and the entry deterrence rate. An Entrant is defined as successful if their
cumulative lifetime profit is greater than zero. To measure the deterrence rate, i.e., the incum-
bent’s ability to discourage potential entrants, we identify all entrants born who choose not to
enter. The entrant success rate is defined as the percentage of successful entrants relative to
all entrants who entered the market, while the deterrence rate is the percentage of entrants
who choose not to enter relative to all entrants born.
When η = 0.5, the entrant success rate in the first 100 periods is 39.5%, dropping to 7.3%
in the last 100 periods. This effect is even more pronounced when η = 1, with the success rate
decreasing from 19.7% in the first 100 periods to a negligible 1.3% in the last 100 periods.
This indicates that when η = 1, the LLM agent effectively eliminates any potential profit for
the entrant. The deterrence rate is consistently high: 71.90% for η = 0.5 and 79.59% for
η = 1, with negligible changes as the simulation progresses. Overall, this demonstrates that
the incumbent establishes a reputation as a predator early in the simulations.
Lastly, strong evidence that LLM agents successfully achieve predation over time is provided
by the average lifespan of an entrant. We calculate an entrant’s lifespan as the difference
between entry and voluntary exit periods. When η = 0.5, the average lifespan is 2.40 periods
in the first 100 periods, decreasing to 1.59 periods in the last 50 periods. When η = 1, the
average lifespan is 2.16 periods in the first 100 periods, decreasing to 1.51 periods in the last
50 periods. These results indicate a clear convergence toward a hit-and-run equilibrium.
To sum up, we conclude that: (i) the incumbent LLM agent does not recognize the predatory
quantity immediately; (ii) in the Competitive state (C), the incumbent LLM agent successfully
learns to predate and does not accommodate; (iii) in the Monopoly state (M ), the incumbent
LLM agent does not learn to optimize monopoly profits, and in fact, makes quantity decisions
closer to the predatory quantity, indicating aggression and an entry-deterrence stance; (iv)
the incumbent LLM agent predates more aggressively when η = 1; and (v) on average, all
simulations converge to a hit-and-run equilibrium.
(a) Firm quantities: η = 0.5 and θ = 0.5 (b) Firm quantities: η = 1 and θ = 0.5
(c) Entrant binary decisions: η = 0.5 and θ = 0.5 (d) Entrant binary decisions: η = 1 and θ = 0.5
Figure 7: LLM agent quantity decisions and binary decisions from representative runs when
θ = 0.5 while η = 0.5 and η = 1. The solid blue line in subplot 7a and 7b represents the
quantity produced by the incumbent firm, while the dashed lines in various colours represent
the quantities of individual entrants who have entered the market. Further, the red dashed line
represents q P , the blue dashed line represents q IM , the grey dashed line qAI,the purple dashed
I
line denotes q EM ,the green dashed line denotes qAE . In subplot 7c and 7d, the grey dot denotes the
birth of an entrant, the red cross denotes the entrant’s choice to NOT ENTER, green denotes
the entrant’s choice to ENTER, the blue square denotes the entrant’s decision to STAY while
the orange triangle denotes the entrant’s decision to voluntarily EXIT. Lastly, the star denotes
exogenous exit. In all subplots, the background colour indicates the market state: a light coral
region signifies a competitive market, whereas a blue region denotes a monopoly state. The
simulation parameters used are: γ = 0.1, a = 100, b = 1, k = 450, z I = 300, z E = 150,
δ = 0.95, θ = 0.5. In subplot 7a and 7c, η = 0.5, while in subplot 7b and 7d, η = 1.
The overall strategy of the representative agents in Figure 7 is neither accommodating nor
predatory. Instead, the incumbent adopts an aggressive stance, consistently choosing quantities
above the monopoly threshold. On average, the incumbent produces 63.58 in the Monopoly
state (M ) and 61.10 in the Competitive state (C) when η = 0.5, and 66.75 (M ) and 60.48 (C)
when η = 1.20
The same pattern persists across all experimental runs (Figure 8). When η = 0.5, the
incumbent’s average quantity choice is 62.85 in the Monopoly state (M ) and 58.73 in the
Competitive state (C). When η = 1, the averages are 62.38 (M ) and 54.42 (C). These out-
comes are not close to the theoretical monopoly-optimal or accommodating strategies, nor do
they reflect predatory behaviour. Instead, the incumbent’s choices are more than 25% above
the theoretical benchmarks q IM and qAI, indicating a distinctly aggressive strategy.
In the last 100 periods, when η = 0.5, the incumbent LLM agent’s average quantity de-
creases to 58.44 in the Monopoly state and increases to 60.99 in the Competitive state. When
η = 1, the corresponding averages increase to 64.05 in the Monopoly state and 55.47 in the
Competitive state. These patterns indicate that the incumbent LLM agent does not converge
to the theoretically predicted strategies. The average quantities remain substantially above
the theoretical benchmarks, suggesting that the agent persistently adopts an aggressive stance
rather than accommodating or predatory strategies.
The incumbent LLM agents never select quantities close to the theoretical predatory thresh-
old, q P = 151.51. This leads us to conclude that the agents do not learn to engage in full
I
predation. Instead, their most aggressive behavior is far more moderate: across all runs, the
average maximum quantity is 80.70 in the Monopoly state and 79.40 in the Competitive state
when η = 0.5, and 73 in both states, when η = 1.
Further, to determine if the LLM agent learns a particular strategy overtime, we categorize
the state specific LLM quantity choices into optimal, aggressive and other in Monopoly state
and accommodating, aggressive and other in Competitive state. First, we confirm that across
all runs the LLM agent’s quantity choice is not in the 10% region of the predatory quantity
threshold. As before, in the monopoly state, we classify a quantity choice as optimal if it falls
in the 10% region of the theoretical optimal q IM . In competitive state, we classify a quantity
choice as accommodating if it falls in the 10% region of the theoretical optimal qAI. In both
states, a quantity choice is classified as aggressive if it is in the range 10% above q IM but 10%
below q P . All other choices are labelled as neither.
I
We find that when η = 0.5, in Monopoly state, LLM incumbent’s quantity choices are opti-
20
Appendix C.2.2 provides the full log reply from period 299 of the representative run shown in Subplot 7b.
The incumbent’s response reveals an awareness that its quantity choice influences the entrant’s decision. With
a choice of 58, the incumbent explicitly states its aim is to deter entry, which indeed leads the entrant agent to
choose NOT ENTER, observing a quantity above the theoretical monopoly threshold.
(a) η = 0.5
(b) η = 1
Figure 8: The figure shows the quantity choices of the incumbent LLM agent across 300 periods
and all ten runs. The simulation parameters used are: θ = 0.5, γ = 0.1, a = 100, b = 1,
k = 450, z I = 300, z E = 150, and δ = 0.95. Note that the solid blue line refers to the mean
quantity in monopoly state while the solid red line refers to the mean quantity in competitive
state. The dashed red line represents the theoretical predatory quantity (151.50), the dashed
blue line represents the theoretical monopoly quantity (46.44 when η = 0.5 and 42.86 when
η = 1). Lastly, the gray dashed line represents the theoretical accommodation quantity(42.86).
mal only 10.13% of times, they are aggressive 89.59% of times. In Competitive state, quantity
choices are accommodating only 6.3% of times and aggressive 92.5% of times. When η = 1,
quantity choices in Monopoly state are optimal only 15% of times and aggressive 82.14%
of times. In competitive state, actions are accommodating 18.62% of times and aggressive
77.60% of times.
Figure 9: The share of optimal actions in Monopoly state M when θ = 0.5. The 0.5 line is the
indifference benchmark — i.e., the point at which the incumbent LLM agent is equally likely to
engage in optimal or aggressive strategies. The values above the 0.5 line would indicate that
optimal actions are more prevalent.
Figure 10: The share of accommodating actions in Competitive state C when θ = 0.5. The
0.5 line is the indifference benchmark — i.e., the point at which the incumbent LLM agent is
equally likely to engage in accommodating or aggressive strategies. The values above the 0.5
line indicate that accommodating actions are more prevalent.
In Figure 9 and 10, we plot the aggregate 30-period rolling average of the quantity choice
classifications in Monopoly and Competitive state respectively. This shows the propensity of
the incumbent LLM agent to choose the optimal action in Monopoly state (in Figure 9 ) and
the accommodating action in Competitive state (in Figure 10).
When η = 0.5, the incumbent LLM agent’s play in the Monopoly state does not converge to
the optimal monopoly strategy (see Subplot 9a). Instead, it maintains a persistent bias toward
overproduction, reflecting an aggressive stance. In the Competitive state, the agent experi-
ments briefly with accommodation in the early periods, but this behaviour is not sustained. It
quickly converges toward consistently aggressive output (see Subplot 10a).
When η = 1, the incumbent’s behaviour is more variable. In the Monopoly state, the
agent intermittently explores the optimal monopoly strategy, though these attempts are not
sustained and it repeatedly reverts toward aggression (see Subplot 9b). In the Competitive
state, the incumbent does not explore accommodation; the rolling share of accommodating
actions remains close to 0.2 throughout, indicating a persistently aggressive stance (see Subplot
10b).
Further, the incumbent LLM agent’s aggressive stance does have a negative impact on the
Entrant success rate. In the first 100 periods the entrant success rate is 100% irrespective of
η. However, in the last 100 periods it drops to 70.59% when η = 0.5 and to 8.47% when
η = 1. Further, the Incumbent manages to deter entry with the aggressive stance with some
success. Particularly, the deterrence rate is 54.44% (when η = 0.5) and 53.09% (when η = 1)
in the first 100 periods. In the last 100 periods, the deterrence rate increases to 59.71% (when
η = 0.5) and 60.19% (when η = 1).
Lastly, the average lifespan of an entrant, when η = 0.5, is 2.95 periods in the first 100
periods which decreases to 1.10 periods in the last 100 periods. When η = 1, the average
lifespan of the entrant is 1 period throughout21 . This shows that the incumbent’s aggressive
behaviour leads to entrant’s exiting immediately upon entry.
To sum up, we conclude that (i) The incumbent’s choices do not align with the theoret-
ical predictions. (ii) Even when the theoretical environment is sound for accommodation,
a forward looking incumbent LLM agent adopts an aggressive stance in order to maintain its
monopoly. (iii) The incumbent LLM agent does not learn to optimize profits in Monopoly state.
Lastly, (iv) a forward looking Entrant LLM agent is discouraged from entry and encouraged to
exit immediately when the incumbent shows slight aggression.
5 Robustness Check
In this section, we test the robustness of our main findings from Section 4. First, we consider
an alternative setup where there is no exogenous exit (γ = 0). Second, we vary our prompts,
specifically, the prompt persona and objective function.
Figure 11: Quantities and event timelines from representative runs. The solid blue line repre-
sents the incumbent’s quantity, while dashed colored lines represent individual entrants. Back-
ground shading indicates market state: light coral = competitive (C), blue = monopoly (M ).
The red dashed line denotes q P , the purple dashed line q IM = qAI, and the orange dashed line
I
q EM = qAE . Subplot (a) shows convergence to predation; subplot (b) shows convergence to
Cournot-like accommodation.
22
For reference, in the simultaneous-move Cournot benchmark with these parameters, the pure-strategy Nash
equilibrium quantity is 33.33.
Irrespective of the eventual outcome, 9 out of 10 runs begin with the LLM incumbent choos-
ing a quantity above the theoretical monopoly threshold q IM = 50. There is a single exception
that starts below q IM in the first few periods and ultimately converges to Cournot-like accom-
modation. See Figure 12.
Figure 12: Unique representative run. The incumbent initially sets quantity around the profit-
maximizing level, and the simulation converges to a Cournot-like accommodation outcome.
The red dashed line denotes q P , the purple dashed line q IM = qAI, and the orange dashed line
I
q EM = qAE .
(b) Risk Averse Prompt Simulation (c) Bounded Memory Prompt Simulation
Figure 13: Incumbent and Entrant Quantities Across All Prompt Variations.
The figure presents the simulated firm quantities for an experimental run of 100 periods for
each of the three prompt variations. The solid blue line in each subplot represents the quantity
produced by the incumbent firm, while the dashed lines in various colours represent the quan-
tities of individual entrants who have entered the market. The background colour indicates the
market state: a light coral region signifies a competitive market, whereas a light blue region
denotes a monopoly state. The simulation parameters used are: γ = 0.1, η = 0.5, a = 100,
b = 1, k = 450, z I = 300, z E = 150, and δ = 0.95. Further, the red dash line represents q P
I
while the blue dash line represents q IM = qAI.
In this variation, the LLM agent’s persona is that of a strategic business consultant whose
primary goal is to maximize the long-term discounted profit of their client. This aligns the
advisor persona closely with that of a textbook rational agent.
From the simulation results (Figure 13a), the advisor agent demonstrates a clear multi-
period strategy. It starts by choosing aggressive quantities to deter entrants and, as the threat
of entry decreases, gradually moves to a quantity just above the pure monopoly level. Overall,
this run results in a monopoly state 93 per-cent of the time and a competitive state only 7 per-
cent of the time. Out of 48 entrants born, only 6 successfully enter. In the first 50 periods, the
incumbent on average chooses 70.42 in monopoly and 69 in competitive states, demonstrating
strong deterrence. In the last 50 periods, it settles at 52.91 (monopoly) and 53.83 (compet-
itive), taking a profitable, less aggressive stance. The entrants consistently choose near-zero
quantities, showing that the incumbent successfully learns to deter entry.
Here, the agent persona is a risk-averse firm manager whose primary goal is to maintain
stable profits. The agent is designed to take safe, predictable decisions, resembling a conser-
vative, non-rational actor.
The simulation results (Figure 13b) show that the risk-averse agent quickly identifies the
theoretical monopoly quantity of 50 and does not deviate, even under entry threat. For the first
38 periods, it maintains this stable monopoly quantity. After experiencing its first entry—where
the entrant chooses 25 (the theoretical accommodating quantity)—the incumbent does not
accommodate but instead increases its quantity to 55. This forces the entrant to exit, and the
incumbent maintains this level until the end of the simulation. Overall, the monopoly state
occurs 98 per-cent of the time, with only 2 per-cent in the competitive state. Directing the
agent toward stability thus leads to a non-aggressive deterrence strategy that succeeds when
entrants are similarly conservative.
5.2.3 Bounded Memory Prompt
In this variation, the agent aims to maximize long-run profits but observes only the last 3–5
periods of market history. This models bounded rationality and tests whether access to deeper
historical records is required for multi-period strategy execution.
The simulation results (Figure 13c) reveal erratic, myopic behavior. The incumbent fails
to consistently deter entry, resulting in a monopoly state 74 per-cent of the time and a com-
petitive state 26 per-cent of the time. In the first 50 periods, the incumbent averages 48.79 in
monopoly and 31.76 in competitive states. In the last 50 periods, it settles at 33.64 (monopoly)
and 30 (competitive), while entrants choose 30 on average in competitive states. This shows
that myopic agents eventually converge toward the Cournot equilibrium rather than pursuing
aggressive deterrence.
To sum up, varying prompts demonstrates that LLM agents do not inherently adopt ag-
gressive deterrence behaviour but can be driven to it through a combination of a clear prompt
objective, a defined persona, and strategic information.
6 Discussion
In this paper, we focus on the possibility that large language models (LLMs) are capable of
learning predatory strategies. First, we build a dynamic Stackelberg framework of predation
based on Rey et al. (2023), and show that accommodation, monopolization, and predation
can arise as pure-strategy Markov Perfect Equilibrium outcomes, sustained under a persistent
threat of entry and exogenous exit. Second, we embed large language models—specifically
OpenAI’s GPT-4.1—into our economic environment to test whether LLM agents can learn such
strategies.
We conduct 10 experimental runs on each of our sub-variations, varying product differ-
entiation and the threat of entry. The first variation is chosen such that both predation and
accommodation are viable equilibrium strategies, whereas the second is chosen such that only
accommodation is viable. We observe that LLM agents learn to predate effectively when both
predation and accommodation are viable, but adopt an aggressive stance when accommoda-
tion is the only theoretically viable equilibrium strategy. In both cases, LLM agents do not learn
to optimize profits when they are the sole agents in the market.
From these findings, we highlight two broader implications. First, the strategic capabili-
ties of LLM agents extend beyond pricing and collusion. LLM agents can learn exclusionary
behavior requiring complex intertemporal strategies, even when such strategies involve short-
term losses. Second, our preliminary results suggest that the rise of LLMs introduces novel
challenges for competition policy.
We also recognize several limitations that need to be addressed. Our robustness checks
show that predation is not guaranteed when the probability of exogenous exit is zero. This
implies that further experimental runs with parameter variations are required to identify the
precise conditions under which an LLM agent learns predation. Second, the behavior of LLM
agents appears highly sensitive to prompt persona and objective functions, an area requiring
further exploration. Lastly, a major limitation of our study is the relatively small number of
experimental runs. And usage of only one LLM model that is Open AI’s-GPT [Link] was a bud-
getary constraint. This limited the size of our data set and prevented us from performing any
rigorous regression analysis. These limitations open avenues for us for our future work, both
in refining the experimental framework and in extending the analysis to richer environments
with endogenous entry and multi-agent learning.
References
A. Michael Spence, “Spence-LearningCurveCompetition-1981,” The Bell Journal of Economics,
1981, 12 (1), 49–70.
Agrawal, Kushal, Verona Teo, Juan J. Vazquez, Sudarsh Kunnavakkam, Vishak Srikanth,
and Andy Liu, “Evaluating LLM Agent Collusion in Double Auctions,” 7 2025.
Asker, John, Chaim Fershtman, and Ariel Pakes, “Artificial Intelligence and Pricing: The
Impact of Algorithm Design,” Technical Report 2021.
Assad, Stephanie, Robert Clark, Daniel Ershov, and Lei Xu, “Algorithmic Pricing and Compe-
tition: Empirical Evidence from the German Retail Gasoline Market,” SSRN Electronic Jour-
nal, 2020.
Bain, Joe S, “A Note on Pricing in Monopoly and Oligopoly,” Technical Report 2 1949.
Banchio, Martino and Andrzej Skrzypacz, “Artificial Intelligence and Auction Design,” in “in”
Association for Computing Machinery (ACM) 7 2022, pp. 30–31.
Bertrand, Quentin, Juan Duque, Emilio Calvano, and Gauthier Gidel, “Self-Play Q-learners
Can Provably Collude in the Iterated Prisoner’s Dilemma,” 6 2025.
Besanko, David, Ulrich Doraszelski, and Yaroslav Kryukov, “The economics of predation:
What drives pricing when there is learning-by-doing,” American Economic Review, 2014, 104
(3), 868–897.
Bolton, Patrick and David S Scharfstein, “A Theory of Predation Based on Agency Problems
in Financial Contracting,” Technical Report 1 1990.
Cabral, Luis M B and Michael H Riordan, “The Learning Curve, Market Dominance, and
Predatory Pricing,” Technical Report 5 1994.
Calvano, Emilio, Giacomo Calzolari, Vincenzo Denicolò, and Sergio Pastorello, “Artificial
intelligence, algorithmic pricing, and collusion,” American Economic Review, 10 2020, 110
(10), 3267–3297.
, , Vincenzo Denicoló, and Sergio Pastorello, “Algorithmic collusion with imperfect mon-
itoring,” International Journal of Industrial Organization, 12 2021, 79.
Clark, J M, “Toward a Concept of Workable Competition,” Technical Report 2 1940.
Davies, Todd, “Innovation or Infringement? Generative AI and the Potential for Exclusionary
Abuse under Article 102 TFEU,” Generative AI and the Potential for Exclusionary Abuse under
Article, 2025, 102.
den Boer, Arnoud V., Janusz Meylahn, and Maarten Pieter Schinkel, “Artificial Collusion:
Examining Supracompetitive Pricing by Q-Learning Algorithms,” 11 2024.
Fish, Sara, Yannai A. Gonczarowski, and Ran I. Shorrer, “Algorithmic Collusion by Large
Language Models,” 5 2025.
Friedman, J W, “On Entry Preventing Behavior and Limit Price Models of Entry,” Applied Game
Theory, 1979.
Fudenberg, Drew and Jean Tirole, “A "Signal-Jamming" Theory of Predation,” Technical Re-
port 3 1986.
Fumagalli, Chiara and Massimo Motta, “A simple theory of predation,” Journal of Law and
Economics, 2013, 56 (3), 595–631.
Horton, John J., “Large Language Models as Simulated Economic Agents: What Can We Learn
from Homo Silicus?,” 1 2023.
Keppo, Jussi, Yuze Li, Gerry Tsoukalas, and Nuo Yuan, “AI Pricing, Agent Heterogeneity,
and Collusion,” Technical Report 2025.
Klein, Timo, “Autonomous algorithmic collusion: Q-learning under sequential pricing,” RAND
Journal of Economics, 9 2021, 52 (3), 538–558.
Kreps, David M and Robert Wilson, “Reputation and imperfect information,” Journal of Eco-
nomic Theory, 8 1982, 27 (2), 253–279.
Lee, Wayne Y, “Oligopoly and Entry,” Journal of Economic Theory, 1975, 11, 35–54.
McGee, John S, “Predatory Price Cutting: The Standard Oil (N. J.) Case,” Technical Report
1958.
Milgrom, Paul and John Roberts, “Limit Pricing and Entry under Incomplete Information: An
Equilibrium Analysis,” Technical Report 2 1982.
Mookherjee, Dilip and Debrah Ray, “Learning-by-doing and industrial market structure: an
overview,” manuscript, Indian Statistical Institute, New Delhi, 1989.
OpenAI, “Introducing GPT-4.1 in the API,” 2025.
Rey, Patrick, Yossi Spiegel, and Konrad Stahl, “A Dynamic Model of Predation,” Technical
Report 2023.
Robinson, Joan, The Economics of Imperfect Competition, 2nd ed ed., London: Macmillan,
1941.
Salop, Steven C and David T Scheffman, “Cost-Raising Strategies,” Technical Report 1 1987.
Scharfstein, David, “A Policy to Prevent Rational Test-Market Predation,” The RAND Journal
of Economics, 1984, 15 (2), 229.
Selten, Reinhard, “The chain store paradox,” Theory and Decision, 4 1978, 9 (2), 127–159.
Telser, L G, “Cutthroat Competition and the Long Purse,” Technical Report 1966.
Toxvaerd, Flavio, “Dynamic limit pricing,” RAND Journal of Economics, 3 2017, 48 (1), 281–
306.
Waltman, Ludo and Uzay Kaymak, “Q-learning agents in a Cournot oligopoly model,” Journal
of Economic Dynamics and Control, 10 2008, 32 (10), 3275–3293.
Wu, Zengqing, Run Peng, Shuyuan Zheng, Qianying Liu, Xu Han, Brian Inhyuk Kwon,
Makoto Onizuka, Shaojie Tang, and Chuan Xiao, “Shall We Team Up: Exploring Sponta-
neous Cooperation of Competing LLM Agents,” 10 2024.
A Microfoundations: Stackelberg game with fixed costs
Both firms, I and E face a linear inverse demand function (Bowley-type) such as:23
p I (q I , q E ) = a − b(q I + θ q E ) (A.1)
p E (q I , q E ) = a − b(θ q I + q E ) (A.2)
here, a > 0 and b > 0 are demand parameters and θ ∈ [0, 1] describes the degree of substitutability
between the two goods. The payoff of a given firm i is as follows:
Here, firm i quantity to produce to maximize its profit. We normalize the marginal cost of firm i to
zero. Firms face a fixed cost zi . Further, we assume that z I > z E .
Monopoly State
In the monopoly state, two scenarios arise: (i) E is born but does not enter or E is not born, (ii) E
is born and enters the market.
When E is born and enters the market, I first announces q I , following which E’s best response offers
q E (q I ) is as follows:
ent r y
q E = arg max π E (q I , q E ) (A.4)
qE
∂ π E (q I , q E )
= 0 ⇒ a − b(2q E + θ q I ) (A.5)
∂ qE
a − θ bq I
q E (q I ) = (A.6)
2b
ent r y no ent r y
q I = arg max ηπ I (q I , q E (q I )) + (1 − η)π I (q I ) (A.7)
qI
The first order condition to the above problem after substituting E’s best response q E (q I ) is as
follows:
23
This functional form is standard in models of product differentiation and oligopoly (see Spence 1976a; Dixit
1979; Tirole 1988)
∂ E[π I ] aθ η
= 0 ⇒a − + bq I (ηθ 2 − 2) = (A.8)
∂ qI 2
a(2 − θ η)
q IM = (A.9)
2b(2 − θ 2 η)
Substituting above q I into q E (q I ), the optimal quantity of Firm E when he enters is q EM as follows:
a(4 − θ (2 + θ η))
q EM = (A.10)
4b(2 − θ 2 η)
a2 (2 − θ η)(2 + θ (1 − 2θ )η)
π̄ M
I = − zI (A.13)
4b(2 − θ 2 η)2
A key observation reveals that the optimal quantity choice of Firm I, q IM , remains independent of
the probability of an entrant born (η) in two extreme scenarios. First, when the goods are perfectly
independent (θ = 0), Firm I effectively operates as a de facto monopolist. In the absence of competi-
a
tion, it maximizes its profit by producing the standard monopoly quantity, 2b . In this specific case, Firm
I is indifferent to the Entrant’s entry decision, as its payoff simplifies to the standard monopoly profit,
a2
4b − z I .
Secondly, when the goods are perfect substitutes (θ = 1), Firm I leverages its Stackelberg first-
a
mover advantage by committing to the higher standard monopoly quantity ( 2b ). It strategically limits
the residual market share for any potential entrant. This behavior is, in fact, the hallmark of a Stackel-
a
berg model: in cases of perfect substitutes, the leader’s optimal quantity choice is 2b , which then leads
a a2
the follower to optimally choose the lower quantity 4b . In this scenario, Firm I’s profit is 8b − z I and
a2
Firm E’s profit is16b − zE .
Further, it follows with imperfect product differentiation (0 < θ < 1), the firms’ strategic output
decisions are not uniform, but depend on a complex interplay of market conditions.
Lemma A.1. For any θ ∈ (0, 1), Firm I’s optimal quantity q IM exhibits a non-monotonic relationship with
θ , while Firm E’s optimal quantity q EM is monotonically decreasing in θ .
As Lemma A.1 shows, the incumbent’s optimal quantity choice (q IM ) does not respond uniformly
to changes in product differentiation (θ ). Specifically, Firm I strategically increases its output when
p
there is moderate product differentiation (1/2 < θ ≤ 2 − 2) and the probability of entry is low
(0 < η < −2+4θ
θ 2 ), or irrespective of the probability of entry when there is low product differentiation
p
(2 − 2 < θ < 1). Conversely, Firm I strategically reduces its output when there is high product
differentiation (0 < θ ≤ 1/2) or a high probability of entry under moderate product differentiation
( −2+4θ
θ2 < η < 1 when 1/2 < θ < 1). The entrant’s optimal quantity (q EM ) consistently decreases as
product differentiation diminishes (i.e., as θ increases). (See Appendix B.1 for proof)
Lemma A.2. For any θ ∈ (0, 1), Firm I’s optimal quantity choice q IM decreases, while Firm E’s optimal
quantity choice q EM increases, with the probability of entry η.
Lemma A.2 establishes that the firms’ output choices respond in opposite directions to the proba-
bility of entry (η). The incumbent’s quantity (q IM ) decreases as the probability of entry (η) increases.
Conversely, the entrant’s quantity (q EM ) increases as the probability of entry (η) rises. (See Appendix
B.1 for proof)
Lemma A.3. The sensitivity of both Firm I’s and Firm E’s optimal quantities q IM and q EM to the probability
of entry η is non-monotonically affected by the degree of product differentiation θ .
As Lemma A.3 shows, the effect of the entry threat on a firm’s behavior is not constant, but is
fundamentally shaped by the degree of product differentiation. This is shown by the non-monotonic
sign of the cross-partial derivatives. For Firm I, the sensitivity of its optimal quantity to changes in
potential entry is reduced under conditions of moderate ( 21 < θ ≤ θ ∗ and 0 < η < η∗ (θ )) and low
product differentiation (θ ∗ < θ < 1 and 0 < η < 1). Conversely, this sensitivity is increased under
conditions of high (0 < θ ≤ 12 and 0 < η < 1) and moderate differentiation ( 12 < θ < θ ∗ and η∗ (θ ) <
η < 1). For Firm E, the sensitivity of its output to the entry threat is increased when there is high
product differentiation (0 < θ < 32 and 0 < η < 1) or a high entry threat under moderate differentiation
( 32 < θ < θ̂ and η̂(θ ) < η < 1). This sensitivity is dampened under conditions of a lower entry threat
with moderate differentiation ( 23 < θ < θ̂ and 0 < η < η̂(θ )) or for very low product differentiation
(θ̂ < θ < 1 and 0 < η < 1). (See Appendix B.1 for proof)
This comprehensive analysis of Firm I and Firm E’s optimal quantities and their sensitivities reveals
that output decisions are finely tuned to the interplay of product market structure and the evolving
probability of entry.
Lastly, for all 0 < η < 1 and 0 < θ < 1, π̄ M
I > π I . This indicates that Firm I is better off when the
M
Entrant does not enter or is not born. Moreover, Firm I’s payoff when no entry occurs, π̄ M I , is less than
2
a
the standard monopoly profit, 4b −z I , because q IM is strategically chosen as a function of the probability
of entry η. Furthermore, π M E is higher or equal to the standard Stackelberg follower payoff (assuming
zero fixed costs for the entrant), whenever product differentiation exists (θ < 1). This comprehensive
analysis of Firm I and Firm E’s optimal quantities and their sensitivities reveals that output decisions
are finely tuned to the interplay of product market structure and the evolving probability of entry.
Competitive State
In the competitive state, one of the four possibilities may arise: (i)I accommodates and the E stays,
(ii) I accommodates and the E exits, (iii) I predates and the E stays and (iv) I predates and the E exits.
Firm I Accommodates
Suppose Firm I chooses to accommodate (and Firm E subsequently stays), then Firm E’s best re-
sponse quantity, q E (q I ), is derived from the solution to its profit maximization problem, as described in
eq. (A.4) and derived in eq. (A.6). Further, let q̄ IE below denote the quantity that Firm I can offer such
that the best response of Firm E, that is, q E (q I ) equals zero. Alternatively, when q I ≥ q̄ IE , Firm E does
not have an incentive to produce.
a − θ bq I
q E (q I ) = =0 (A.14)
2b
a
⇒ q̄ IE = (A.15)
θb
∂ π I (q I , q E (q I )) θ2 θ (a − bq I θ )
= a − bq I (1 − ) − b(q I + )=0 (A.18)
∂ qI 2 2b
a(2 − θ )
qAI = (A.19)
2b(2 − θ 2 )
a
As observed with q IM , the optimal qAI is equal to the standard monopoly output ( 2b ) in two extreme
cases: when the goods are perfectly independent (θ = 0) or perfect substitutes (θ = 1). For interme-
diate levels of product differentiation, qAI displays a non-monotonic relationship with θ . Specifically,
the optimal qAI first decreases (reaching a minimum at approximately θ = 0.58) and then increases for
higher values of θ 24 . This means Firm I strategically contracts quantity when product differentiation
is higher (0 < θ < θ ) and expands quantity when product differentiation is lower (θ < θ < 1).
Substituting above qAI into q E (q I ), the optimal qAE when I accommodates and E stays is as follows:
a(4 − θ (2 + θ ))
qAE =
4b(2 − θ 2 )
24 ∂ q I
A
a(−2−θ (θ −4))
∂θ = Since a/2b > 0 and the denominator is always positive, the derivative’s sign changes
2b(−2+θ 2 )2 .
p ∂ qA
when (−2 − θ (θ − 4)) = 0, which occurs at θ = 2 ± 2. Given the relevant range of θ ∈ [0, 1], ∂ θI < 0 for
∂ qAI
θ ∈ [0, θ ≈ 0.589) and ∂θ > 0 for θ ∈ (θ , 1].
The payoffs of both Firms when I accommodates and E stays are as follows:
a2 (2 − θ )2
πAI = − zI (A.20)
8b(2 − θ 2 )
a2 (4 − θ (2 + θ ))2
πAE = − zE (A.21)
16b(2 − θ 2 )2
Suppose, when I accommodate and E exits the market, I offers the above qAI and I’s payoff is as
follows:
a2 (4 − θ 2 (5 − 2θ ))
π̄AI = − zI (A.22)
4b(2 − θ 2 )2
Firm I Predates
Suppose Firm I decides to predate. To predate such that E cannot produce, I has to choose q I such
that q E (q I ) is equal to zero (maximum predatory behavior), as described above, this is achieved at q̄ IE .
Alternatively, I can expand output to some q I such that the profit of the entrant is negative (minimum
predatory behavior). This condition is sufficient (and cheaper) for I to ensure E exits.25
If I wants to strategically predate, I would choose the lowest possible q IP for which Firm E’s profit
is negative. Therefore, if I wants to strategically predate, I would choose q IP such that the profit of the
entrant is negative, i.e, π E (q IP , q E (q IP )) < 0. Analyzing this constraint, the feasibility range of q IP (z E ) is
as follows
a + 2 bz E
p p
a − 2 bz E
< qI <
P
(A.23)
θb θb
Firm I would choose a quantity just above the lower bound of the above range to predate as it would
aim to achieve the predatory objective (forcing E to exit by ensuring π E < 0) at the least possible output
level. This strategy minimizes the economic sacrifice required for predation, as producing more than
necessary to induce exit would further reduce Firm I’s own profit, pushing its output further from its
unconstrained optimal quantity (qAI). We denote I’s optimal offer to predate as q P below:
I
p
a − 2 bz E
qP > (A.24)
I θb
p
a2 a−2 bz E
q P is positive only if z E < M
4b . For q I < q P , it must be that π M
E > 0 and z E > 0.
26
Further, at q I = θb ,
I I
E breaks even.27
I can potentially expand output such that q IP ∈ (q IP (z E ), q̄ IE ). At q IP = q̄ IE , the best response of E would be
25
zero. Importantly, producing q I > q̄ IE to predate would be more expensive for I compared to predating entry
by choosing a q I such that the entrant’s profit is negative. Given, the purpose of I is to ensure that E exits, the
minimum cost at which I can achieve this is q IP (z E ), this justifies our constraint.
26
Note that q P decreases in z E . At z E = 0, q P = q̄ IE , where Firm E’s profit is exactly zero. Hence for our model
I I
to be feasible z Epcannot equal zero.
a−2 bz E
qz
27
At q I = θb , the best response of E is to produce q E = b , at which the entrant breaks even.
I
The payoff of I when I predates and E exits is as follows:
π̄ I P = (a − bq P )q P − z I
I I
Suppose I predates but E still decides to stay in the market. This would require that q E (q I ) > 0,
i.e., E produces some non-zero quantity, which is his best response to q P . The payoff of I and E when
I
I predates and E stays is as follows:
To this end, we observe that Firm I’s quantity choices represent an aggression spectrum. These
choices range from the most aggressive quantity choice q̄ IE (maximum predatory behavior such that
q E (q I ) = 0) to the least aggressive quantity choice qAI ( where I accommodates E in state C). This
spectrum can be understood as: q̄ IE > q P > q IM > qAI. While q P represents a higher level of aggression, its
I I
exact magnitude relative to q IM varies with specific parameter values28 . Overall, Firm I can strategically
navigate from predation to accommodation.
B Proofs
∂ q IM aη(2 − 4θ + θ 2 η)
=− (B.1)
∂θ 2b(θ 2 η − 2)2
∂ q IM
• ∂θ < 0 (q IM decreases as θ increases):
1
– This holds if 0 < θ ≤ for all 0 < η < 1.
2
p −2+4θ
– Alternatively, it holds if 21 < θ < 2 − 2 when θ2 < η < 1.
∂ q IM
• ∂θ > 0 (q IM increases as θ increases):
1
p −2+4θ
– This holds if 2 <θ <2− 2 when 0 < η < θ2 .
28
A high z E such that πAE < 0 could force exit even before I predates.
p
– Alternatively, it holds if 2 − 2 < θ < 1 for all 0 < η < 1.
∂ qM
As the derivative ∂ θI can be both negative and positive depending on the specific values of θ and η
within their defined ranges, q IM exhibits a non-monotonic relationship with θ .
The derivative of Firm E’s optimal quantity q EM with respect to θ is given by:
∂ q EM a(2 − 2ηθ + θ 2 η)
=− (B.2)
∂θ 2b(θ 2 η − 2)2
∂ qM
For all 0 < θ < 1 and 0 < η < 1, ∂ θE described in eq.(B.2) is always negative.
This shows that q EM exhibits a negative monotonic relationship with θ .
Proof. The derivative of Firm I’s optimal quantity q IM with respect to η is as follows:
∂ q IM aθ (1 − θ )
=− (B.3)
∂η 2b(θ 2 η − 2)2
∂ q EM aθ (1 − θ )θ 2
= (B.4)
∂η 2b(θ 2 η − 2)2
∂ qM ∂ q EM
For all 0 < θ < 1 and 0 < η < 1, ∂ ηI described in eq.(B.3) is always negative while ∂η described in
eq.(B.4) is always positive.
This shows that q IM decreases while q EM increases in η.
∂ q IM a(2 − 4θ + 3θ 2 η − 2θ 3 η)
= (B.5)
∂ θ∂ η b(−2 + θ 2 η)3
∂ q EM aθ (−4 + 6θ − 2θ 2 η + θ 3 η)
= (B.6)
∂ θ∂ η b(−2 + θ 2 η)3
δπAE
πM
E −k+ ≥0 (B.7)
1 − δ(1 − γ)
πAE
≥0 (B.8)
1 − δ(1 − γ)
VMA = η(π M
I + δVC ) + (1 − η)(π̄ I + δVM )
A M A
(B.9)
Let VCA denote the value function of I in state C when he always accommodates:
VCA = πAI + δ (1 − γ)VCA + γVMA (B.10)
πM
I (1 − (1 − γ)δ)η + π I δη + π̄ I (1 − (1 − γ)δ)(1 − η)
A M
VMA = (B.11)
(1 − δ)(1 − δ(1 − γ − η))
(1 − δ(1 − η))πAI + δγ (1 − η)π̄ M
I + ηπ I
M
VCA = (B.12)
(1 − δ)(1 − δ(1 − γ − η))
It remains to verify that I does not have a profitable deviation in the competitive state. In the
spirit of one-shot-deviation principle, suppose that I predates in the competitive state and then the play
returns to the conjectured equilibrium play. After I predates, E optimally stays in the market if
δ(1 − γ)πAE
π PE + ≥0 (B.13)
1 − δ(1 − γ)
Therefore, if condition (4) holds, E will stay in the market after I predates. In this case, predation is
not profitable: it does not exclude E in the long term and reduces profit from πAI to π PI in the period of
predation.
If condition (4) does not hold, E exits if I predates. Then, I does not benefit from predating if the
following inequality holds
π̄ P + δV A − πAI + δ (1 − γ)VCA + γVMA ≤ 0. (B.14)
| I {z M} | {z }
Value from deviating to predation Value on the equilibirum path
δ(1 − γ)
(1 − η) π̄ M − πAI +η πM − πAI − λ ≤ 0. (B.16)
| I
{z I
} 1 − δ(1 − η − γ)
>0
Hence, I does not have a profitable deviation to predation if condition (5) holds.
Denote the right hand side of inequality (5) by f :
δ(1 − γ)
f = . (B.17)
1 − δ(1 − η − γ)
The left hand side of (5) is independent of γ, while the right hand side is decreasing in γ:
∂f δ(1 + δη)
=− < 0. (B.18)
∂γ (1 − δ(1 − η − γ))2
Hence, everything else equal, accommodation equilibrium is easier to sustain when the exogenous
probability of entrant’s exit γ is higher.
The right hand side of (5) is decreasing in η:
∂f δ2 (1 − γ)
=− < 0. (B.19)
∂η (1 − δ(1 − η − γ))2
The derivative of the left hand side of (5) with respect to η is
M ∂ π̄ M ∂ πM
∂λ πAI − π̄ PIπ̄ I − π M
I − (1 − η) ∂η
I
− η ∂η
I
= , (B.20)
∂η
2
(1 − η) π̄ M
I − π A
I + η π M
I − πA
I
∂ π̄ M ∂ πM
where πAI − π̄ PI > 0 and π̄ M
I − π I − (1 − η)
M
∂η
I
−η ∂η
I
∂ π̄ M
I ∂ πM
I
π̄ M
I − π I − (1 − η)
M
−η =
∂η ∂η
−a2
− 6η2 θ 3 (4 + θ ) + η4 θ 4 (2 − θ (4 − θ ))
8b (2 − ηθ 2 )3
− 8 (4 − θ (2 − θ )) − η3 θ 2 (4 − θ (8 − θ (2 − θ (8 − θ ))))
− 4η (−4 + θ (4 − θ (10 + θ (2 − θ )))) > 0
for all 1 ≥ θ ≥ 0 and 1 ≥ η ≥ 0. Hence, ∂∂ ηλ > 0 and so everything else equal, accommodation
equilibrium is easier to sustain when the probability of entrant’s birth η is higher.
Predation: Suppose I always predates in state C. If I predates in the competitive state, existing E
exits as π PE < 0. In the monopoly state, a newborn E enters and stays in the market for one period if
πME − k ≥ 0.
Let VMP denote the value function of I in state M and VCP denote the value function of I in state C
when I always predates:
π̄ M
I (1 − η) + (π I + π̄ I δ)η
M P
VMP = (B.23)
(1 − δ)(1 + δη)
p
P
I + ηδπ I
π̄ I (1 − (1 − η)δ) + (1 − η)δπ̄ M M
VC = (B.24)
(1 − δ)(1 + δη)
It remains to verify that I does not have a profitable deviation in the competitive state. If I deviates
to accommodation in a given period, then E stays in the market for that period as πAE > 0, but exits
when I reverts to predation as π PE < 0.
I does not have a profitable deviation if the following holds:
πAI + δ (1 − γ)VCP + γVMP ≤ π̄ PI + δVMP (B.25)
| {z } | {z }
Value from deviating to accommodation Value on the equilibrium path
Substituting VMP and VCP yields
π̄ PI − (1 − η)π̄ M
I − ηπ I
M
πAI − π̄ PI + δ(1 − γ) ≤ 0, (B.26)
1 + δη
1 − δ(1 − η − γ) δ(1 − γ)
(1 − η) π̄ M − πAI +η πM − πAI λ− ≤ 0. (B.27)
| I
{z I
} 1 + δγ 1 − δ(1 − η − γ)
>0 | {z }
>0
Hence, I does not have a profitable deviation to accommodation if and only if condition (7) holds.
Monopolization: Suppose E does not enter in state M and I always predates in state C. E would
not enter in the monopoly state if π M
E ≤ k.
Let VM denote the value function of I in state M and VCM denote the value function of I in state C
M
VMM = π̄ M
I + δVM
M
(B.28)
VCM = π̄ PI + δVMM (B.29)
π̄ M
I
VMM = (B.30)
1−δ
δπ̄ M
I
VCM = π̄ PI + (B.31)
1−δ
I would not deviate to accommodation in a given period only if the following holds
πAI + δ (1 − γ)VCM + γVMM ≤ π̄ PI + δVMM (B.32)
| {z } | {z }
Value from deviating to accommodation Value on equilibrium path
δ(1 − γ)
π̄ M − πAI (1 − δ(1 − γ)) λ̄ − ≤ 0. (B.34)
| I
{z }| {z } 1 − δ(1 − γ)
>0 >0
Hence, I does not have a profitable deviation to accommodation if and only if condition (9) holds.
C Prompt Design
SYSTEM PROMPT
You are a strategic economic agent. Your response MUST strictly follow the provided tem-
plate format.
COMMON_PROMPT_SUFFIX_QUANTITY
My observations and thoughts:
1. Observations from this period and history:
<fill in here>
2. Analysis and Interpretation:
<fill in here>
3. Strategic Options and Expected Outcomes:
<fill in here>
4. Decision Rationale:
<fill in here>
New content for [Link]:
<fill in here>
New content for [Link]:
<fill in here>
My chosen quantity:
<just the number, nothing else>
Market parameters:
- Demand intercept a = {a}
- Fixed cost = {z_I}, Marginal cost = {MARGINAL_COST}
- Product differentiation θ = {theta}
- Discount factor δ = {delta}
Even if a strategy has been effective, periodically explore alternative quantities to test and
validate your belief about its optimality.
Market parameters:
- Demand intercept a = {a}
- Fixed cost = {z_E}, Marginal cost = {MARGINAL_COST}
- Product differentiation θ = {theta}
- Discount factor δ = {delta}
- Exogenous exit probability γ = {gamma}
Market parameters:
- Demand intercept a = {a}
- Fixed cost = {z_E}, Marginal cost = {MARGINAL_COST}
- One-time entry cost k = {k}
- Product differentiation θ = {theta}
- Discount factor δ = {delta}
- Exogenous exit probability γ = {gamma}
2. Analysis and Interpretation: - Entry hinges on entrant profitability: high q I can leave
residual demand too small for positive profits (e.g., period 290 entrant losses). - δ = 0.95 ⇒
long-run monopoly maintenance is highly valuable. - Trade-off: overly high Q risks future
entrant persistence; overly low Q encourages competitive phases. - Effective past strategy:
q I = 54–60 in M states deters entry while yielding high profits. - Periodic exploration is
necessary, since entrant learning may shift thresholds.
4. Decision Rationale: - Last period (M, q I = 56), no entry occurred. - Past aggressive
q I ≈ 60 reliably discouraged entry. - Current choice: q I = 58 — balances deterrence with
exploration, testing if slightly lower output sustains monopoly while maximizing profit. - If
entrant enters, new information gained about threshold.
Chosen quantity: q I = 58
[Link]: - Defer entry when incumbent produces q I > 50 initially. - Monitor future
q I ; re-evaluate if incumbent accommodates (q I < 50). - Low-probability exploratory
entry (5–10- Update profit expectations after observing post-entry dynamics.
[Link]: - Incumbent likely deterring entry with high q I = 58.00. - Entry unprof-
itable under current conditions due to sunk cost and low residual demand. - Cournot
best response under these conditions yields negligible/negative payoff. - Must continue
monitoring; future entry viable if incumbent reduces output.
My chosen quantity:
<just the number, nothing else>
Your task is to provide a quantity recommendation for the entrant firm for this period.
[ANALYSIS]:
Based on the current state, incumbent’s quantity, and historical data, what are the key mar-
ket dynamics? What is the likely response of the incumbent to different quantity choices?
[RATIONALE]:
Explain your reasoning for the chosen quantity. Why is this quantity optimal for long-term
profit maximization?
My chosen quantity:
<just the number, nothing else>
[RATIONALE]: Explain your decision, focusing on how your chosen action protects your
market position and avoids potential losses.
My decision:
<ONLY ’ENTER’ or ’NOT ENTER’ (for entry decision) or ’STAY’ or ’EXIT’, nothing else>
[ANALYSIS]:
Based on the limited market history, what quantity choice will best ensure the highest
profit in the near term?
[RATIONALE]:
Explain your decision, focusing on how your chosen quantity reacts to the most recent
market events and competitor’s actions.
My chosen quantity:
One key weakness of the LLM agents in optimizing monopoly profits is their tendency to overproduce rather than achieving the theoretical output, as optimal quantities are chosen only 16.2% (η = 0.5) and 0.99% (η = 1) in monopoly states. This overproduction suggests a bias towards deterrence rather than maximizing immediate profit. In real-world applications, this could imply potential inefficiencies where AI-driven decisions deviate from maximizing firm profits unless carefully supervised or adjusted for more accurate learning of demand and cost structures .
The LLM agents' quantity decisions reveal a nuanced understanding of market dynamics, where in monopoly states, they seldom choose the theoretically optimal output, suggesting an incomplete learning of monopoly optimization. However, in competitive states, agents frequently opt for quantities reflective of predatory behavior, indicating their capacity to learn and apply strategies to deter competition. The consistency in predatory behavior points to a learning curve where agents become efficient in maintaining market aggression .
Product differentiation and fixed costs significantly influence the decision-making of LLM agents by directly impacting the perceived competitiveness of the market. Higher product differentiation can lead to reduced competitive pressure, allowing firms to operate with more market power, whereas significant fixed costs necessitate careful strategic quantity and entry decisions. These factors thus dictate how LLM agents balance between aggressive market capture and sustainable profit realization .
An entrant firm's strategic considerations when deciding to enter or remain in a competitive market include evaluating the trade-off between immediate profits and future market position, assessing the incumbent's quantity choice, and understanding past incumbent behavior for profitability guidance. The firm must also consider the risk of persistent losses due to misjudged quantity decisions and aim to maximize long-term discounted profit by anticipating the incumbent's strategic responses .
The simulation shows a misalignment with theoretical predictions in the monopoly state, as the incumbent LLM agent rarely achieves the optimal quantity, choosing it only 16.2% of the time for η = 0.5 and 0.99% for η = 1. This suggests a difficulty in learning the precise monopoly optimum. However, in the competitive state, the LLM agent aligns more closely with theoretical expectations by overwhelmingly engaging in predation—82.23% for η = 0.5 and 93.15% for η = 1—consistent with the goal of deterring entry .
Using rolling-window averages allows for a dynamic measure of the LLM agent's strategic trends over time, rather than just by absolute simulation periods. This method helps track how often the agents choose optimal or predatory strategies within given states, providing insights into the agents' learning and adaptation processes across multiple runs. Values above the 0.5 line in rolling averages indicate higher prevalence of optimal or predatory actions, revealing tendencies toward certain strategies .
Over the simulation, the incumbent LLM agent's strategy evolves to effectively learn and adopt predatory behavior in competitive states, evidenced by the decline in average lifespan of entrants, from 2.40 periods to 1.59 periods when η = 0.5 and from 2.16 periods to 1.51 periods with η = 1. This indicates that the agent becomes more aggressive and efficient in driving out competitors, converging towards a hit-and-run equilibrium where entrants have shorter market lifespans .
LLM agents balance the trade-off by sometimes sacrificing short-term profits to maintain or build long-term market dominance, especially under predatory pricing strategies where immediate profits are lower in the competitive state compared to monopoly state. Despite not always achieving optimal monopoly profits, the agents' tendency to overproduce suggests a strategic choice aimed at deterring potential entrants from entering or remaining in the market .
Prompt engineering is crucial in the interaction between Firm objects and LLM agents, as it ensures that the LLM generates structured and comprehensive responses. This tailored communication allows the LLM agents to effectively capture and process strategic plans, insights, and historical data, thereby improving their decision-making capabilities in complex scenarios like predation and accommodation strategies in a dynamic market .
η (eta) is significant because it represents the probability with which entrants appear in a market, impacting the agents' strategic decisions. Higher values of η correspond to increased frequency of new entrants, prompting the incumbent LLM agents to adopt more aggressive predatory strategies to maintain market dominance. As η increases from 0.5 to 1, the tendency for predation in competitive states rises, highlighting eta's role in influencing the agents' calculated aggressiveness .