0% found this document useful (0 votes)
16 views66 pages

LLMs and Algorithmic Predation Strategies

This paper investigates the ability of large language models (LLMs) to learn predatory pricing strategies in dynamic market environments. Using OpenAI's GPT-4.1, the study finds that LLMs can adopt aggressive predatory strategies when both predation and accommodation are viable, but struggle with profit optimization. The findings raise concerns for competition policy, highlighting the potential for algorithmic exclusionary practices by dominant firms using LLMs.

Uploaded by

bailid697
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
16 views66 pages

LLMs and Algorithmic Predation Strategies

This paper investigates the ability of large language models (LLMs) to learn predatory pricing strategies in dynamic market environments. Using OpenAI's GPT-4.1, the study finds that LLMs can adopt aggressive predatory strategies when both predation and accommodation are viable, but struggle with profit optimization. The findings raise concerns for competition policy, highlighting the potential for algorithmic exclusionary practices by dominant firms using LLMs.

Uploaded by

bailid697
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

ALGORITHMIC EXCLUSION BY LARGE LANGUAGE

MODELS*
Arina Nikandrova† Anushree Parekh‡

September 30, 2025

Abstract

This paper explores whether large language models (LLMs) can learn predatory strategies
in dynamic environments in which an incumbent faces repeated entry threat. Using Ope-
nAI’s GPT-4.1 as decision-making agents, we find that LLMs learn to predate when both
predation and accommodation are theoretically viable, and adopt aggressive strategies
when only accommodation is theoretically viable. Further, profit optimization is limited,
highlighting both strategic learning and its limitations.

Keywords: Predation, accommodation, dynamic competition, Large Language Models


(LLMs), algorithmic exclusion
JEL Codes: D21, D43, L12, L13, L41

* We thank Romans Pancs for early inspiration and Fendi Tsim for valuable comments and suggestions on the
code. Financial support from the Centre for Competition and Regulatory Policy at City St George’s, University of
London, is gratefully acknowledged.

Department of Economics, City St George’s, University of London, UK, e-mail:
[Link]@[Link].

Department of Economics, City St George’s, University of London, UK, e-mail:
[Link]@[Link].
1 Introduction
This paper is motivated by the growing importance of algorithmic pricing — the use of
advanced artificial intelligence tools to set prices in online markets. While algorithmic pric-
ing promises efficiency gains, it also raises significant concerns for competition policy. The
emerging economics literature has primarily focused on the potential for collusion facilitated
by pricing algorithms (see, e.g., Calvano et al. (2020); Asker et al. (2021); Banchio and Skrzy-
pacz (2022)) or large language model (LLM) agents (see, Fish et al. (2025)). These studies
demonstrate that algorithms can tacitly coordinate to sustain supracompetitive prices, learning
to match collusive benchmarks and punishing deviations through price wars.
However, collusion is only one dimension of the potential harm. Most real-world compe-
tition cases center on exclusionary conduct — that is, strategies by dominant firms to prevent
or deter entry. Despite policy concerns that generative AI may enable exclusionary practices
(Davies (2025); FTC Technology Blog, 2023), no study has yet examined exclusionary con-
duct by algorithmic agents. It is, therefore, not yet understood whether and how learning
algorithms deployed by incumbent firms might engage in predatory pricing, the practice of
pushing market price below cost to discipline or eliminate entrants.
This paper shifts the focus from collusion to predation, using LLM-based simulations of
dynamic markets in which an incumbent faces a stochastic flow of potential entrants. To our
knowledge, it provides the first systematic test of whether, in such settings, LLM agents can
autonomously learn to adopt harmful exclusionary strategies.
Our economic environment builds on Rey et al. (2023), who study the feasibility and prof-
itability of predation in a dynamic infinite-horizon setting in which an incumbent repeatedly
faces potential entry. When a rival enters the market, the incumbent chooses whether to ac-
commodate or predate it; the entrant then decides whether to stay or [Link] our framework,
if entry occurs, firms interact within a Stackelberg framework, where the incumbent chooses
quantity first and then the entrant, having observed the incumbent’s output, chooses its own
output. In contrast to Rey et al. (2023) framework, we introduce an exogenous probability
of entrant exit, which serves to disrupt potential collusive dynamics post entry and creates a
richer environment for studying how predatory strategies may be learned.
In Section 2, we show that in our theoretical framework three Markov Perfect pure strategy
equilibria can arise: (i) Accommodation equilibrium in which the incumbent never predates
and the entrant always enters and stays in the market until exogenous exit takes place. (ii)
Predation equilibrium, in which the incumbent always predates when the entrant enters; the
entrant always enters for exactly one period. (iii) Monopolization equilibrium, in which the
incumbent always predates upon entry and the entrant never enters as entry for one period
is not profitable. A higher probability of birth and of exogenous exit makes accommodation
equilibrium more easily sustainable.
We then use our theoretical framework to analyze the strategic learning process of LLM
agents acting as firms. In all our experiments, OpenAI’s GPT-4.1 serves as the decision-making
engine for our agents.1 . We test whether these agents: (i) learn to choose predatory quantities
when both predation and accommodation are theoretically feasible, and (ii) learn to accom-
modate when accommodation is the only viable equilibrium.
In Section 3, we describe our experimental design, focusing on the nature of our LLM
agents and the prompt engineering. We design prompts to produce parsable responses. In
particular, the LLM agents are instructed in simple language to behave like rational firms,
making quantity decisions with the goal of maximizing long-run discounted profits. The agents
base their decisions on market information as well as self-generated plans from the previous
periods.
Our setup allows us to test whether agents learn from experience and adjust their strate-
gies dynamically over time. A key challenge is that successful predation requires a complex
intertemporal strategy: the incumbent firm incurs short-term losses to drive out competitors,
with the expectation of future gains once rivals are deterred. The profitability of such a strat-
egy hinges on whether the losses are temporary and the reduction in competition is lasting.
Importantly, because predation is costly in the short run, it is not guaranteed that an algorithm
will learn to adopt it even when it is optimal in the long run.
Section 4 presents our core experiment. We stimulate two different strategic environment:
(i) homogeneous goods, where both predation and accommodation are theoretically feasible;
(ii) moderately differentiated goods, where only accommodation is theoretically sustainable.
In both settings, we consider two sub-variations: (a) a potential entrant is born with probability
1/2, and (b) an entrant born with certainty.
We find that, when goods are homogenous and predation and accommodation are both
theoretically viable strategies, the incumbent learns to predate on a newly emerged [Link]
predatory quantity aligns with our theoretical predictions. However, the LLM incumbent adopts
an aggressive strategy also when there is no rival in the market — the incumbent LLM agent
does not learn to fully optimize profits. Furthermore, when products are moderately differen-
tiated, the quantity choices of the LLM agents do not align with theoretical predictions. The
incumbent LLM agent’s strategy converges towards an aggressive stance that is neither fully
predatory nor accommodating. Additionally, when goods are homogenous, the certain birth of
1
Fish et al. (2025), employs OpenAI’s GPT-4 model, and shows GPT-4 learns to optimally set prices and collude.
We use GPT-4.1 released in April 2025 as it is the latest GPT model at the time our experiment runs were conducted.
All experimental runs were conducted between July 2025 and Aug 2025. To understand capabilities of GPT-4.1
see OpenAI (2025). The most recent GPT model now available is GPT-5.
an entrant makes the incumbent more aggressive, whereas with moderate product differenti-
ation, the same certainty increases the volatility of the incumbent’s behavior.
In Section 5, we undertake various robustness checks. We show that in the absence of
exogenous exit, collusion does not emerge; predation is not guaranteed; and accommodation
is achieved only suboptimally. Prompt design matters: alternative prompts can either induce
or mitigate aggressive behavior. We also report results from additional parameter variations.
Overall, our results suggest that while the literature has emphasized the ability of LLM-
based agents (Fish et al. (2025)) and autonomous algorithms (Calvano et al. (2020)) to sustain
collusion, they are equally capable of learning predation. This finding is particularly relevant
for competition authorities and regulators, as it highlights a broader class of algorithmic strate-
gies that may harm competitive process and ultimately consumers.

Relevant literature
This paper contributes to the well-established industrial organisation literature on preda-
tion and other exclusionary behaviours, as well as to the nascent literature on algorithmic
decision-making.

Predation and limit pricing

Early literature on predation considered predatory pricing practically irrational and rare.
Particularly, Robinson (1941) considered firms undercutting prices (which we now distinguish
from predatory pricing) as a normal competitive practice between two firms competing for
market share, while McGee (1958) argued that it was more cost-efficient for dominant firms
to acquire smaller firms than to lower prices in order to drive out competition.2
Telser (1966) was the first to warn about firms’ capacity to predate. He formalized the idea
that the feasibility of predation depended on the firm’s financial capacity—the long purse the-
ory. He concluded that the possibility and extent of predation depended on the firm’s capacity
to bear short-term losses. Further, he emphasized that another condition required to sustain
predation was perfect capital markets, as imperfect capital markets would make borrowing dif-
ficult.3 Bolton and Scharfstein (1990) also establish that optimal financial conditions balance
the benefit and cost of predation.
2
McGee (1958) specifically argued that Standard Oil did not achieve dominance through predatory pricing
but through acquisition of smaller firms. This was a widely accepted view.
3
Telser (1966) further tested his theory empirically, to a certain extent, by analyzing concentration ratios of
137 manufacturing industries. He attempted to test the hypothesis that concentrated markets are connected to
perfect capital markets.
Simultaneously, Bain (1949), Clark (1940), and Friedman (1979) emphasized the possibil-
ity of lowering prices to retain monopoly position (by discouraging entry)—which we now call
limit pricing. Milgrom and Roberts (1982) connect the strategy of limit pricing to that of a ra-
tional firm in a two-period model. They argue that limit pricing can arise in equilibrium when
entrants have private information (about, for example, incumbents’ costs) that affects relevant
payoffs. The incumbent can deter entry based on the reputation of being low-cost. Another
model based on information asymmetry and reputation is Kreps and Wilson (1982). These
models, based on the chain store paradox developed by Selten (1978), show that small infor-
mation asymmetries are sufficient to make incumbents’ threats credible, such that established
firms develop the reputation of a “tough” incumbent which deters entry.
Fudenberg and Tirole (1986) provide a varied explanation. Their model suggests that when
entrants are certain only about profitability today and not in the future, existing firms engage
in predatory pricing in order to “jam signals,” that is, to mislead entrants. Similarly, Roberts
(1986) show that information asymmetry about demand uncertainty and realized profits in a
two-period model makes incumbents’ low prices a credible threat leading to predation. See
Salop and Scheffman (1987) and Scharfstein (1984) for other models on misleading entrants.
Another theory of predation is the learning curve hypothesis. Lee (1975) consider the
effect of learning through cumulative output, which may raise entry barriers in a dynamic
limit pricing model. A. Michael Spence (1981) show that late entrants suffer in a quantity-
setting model.4 These models mainly focus on homogeneous goods, deterministic demand,
and Cournot quantity-setting frameworks. Specifically, Cabral and Riordan (1994) model a
dynamic price competition framework. Their main findings reveal that in a duopoly setting
where firms face a sequence of buyers with demand uncertainty, learning increasingly facilitates
dominance. Further, predatory pricing is feasible in learning economies. They show that there
exist Markov Perfect Equilibria where both firms enter the market but the firm that loses sales
(predated) exits.5
Recent studies build on these frameworks.6 Besanko et al. (2014) adopt the model from
Cabral and Riordan (1994), allowing for re-entry, and conduct numerical simulations which
reveal that aggressive pricing arises routinely. They find that aggressive equilibria (predation-
like behavior) coexist with accommodating equilibria. Another theory that has surfaced in the
recent literature is predation based on economies of scale, developed by Fumagalli and Motta
(2013). On the other hand, Toxvaerd (2017) extend the Milgrom and Roberts (1982) model by
4
See Mookherjee and Ray (1989) for a model where learning curves lead to collusion.
5
Cabral and Riordan (1994) model is a simultaneous price-setting game. Where both firms enter, at the start of
a given period they decide whether to stay or exit, incurring a fixed cost. One should note that it is the introduction
of fixed cost that makes their MPE equilibria of predation feasible.
6
For a descriptive overview of different types of theoretical predation models see Ordover and Saloner (1899).
adopting a dynamic repeated-interaction framework where limit pricing arises as an optimal
intertemporal strategy.
We adopt the theoretical framework from Rey et al. (2023), who consider an infinite-
horizon discrete-time model where an incumbent faces a continuous entry threat. They differ
from the literature so far by characterizing conditions under which strategic uncertainty suf-
fices7 to give rise to predation, accommodation, and monopoly equilibria in Markov Perfect
strategies. We build upon their framework by allowing for product differentiation and intro-
ducing the possibility of entrant’s exogenous exit.

Algorithmic decision-making

Over the years, decision-making in firms has shifted from humans to software and now
to sophisticated algorithms. This evolution has raised concerns among policymakers and aca-
demics that algorithms might learn to coordinate prices, potentially sustaining collusion even
without explicit instructions. The seminal work highlighting this possibility in a structured
manner is Calvano et al. (2020). Their experiment simulated reinforcement learning algo-
rithms, specifically Q-learning, in a symmetric duopoly setting. The algorithms learned to
autonomously collude, providing the first clear demonstration that pricing algorithms acting
as firms in an infinitely repeated, simultaneous-move game can achieve collusive outcomes by
learning from past history.8 Importantly, Calvano et al. (2020) show that memory is crucial:
collusive outcomes are only achievable when agents can remember past interactions and adapt
accordingly.
Asker et al. (2021) further differentiated outcomes across settings such as asynchronous
versus synchronous learning and observed that collusive outcomes are feasible only under
asynchronous learning. Calvano et al. (2021) demonstrate that collusion can emerge even
under imperfect monitoring, provided algorithms are allowed to complete their learning pro-
cess. While these studies focus on simultaneous-move games, Klein (2021) show that collusive
outcomes are also possible in sequential-move setups.9
While earlier studies show that collusion in Q-learning algorithms arises through repeated
interaction and convergence to grim-trigger or punishment-based strategies, Banchio and Man-
tegazza (2023) make a novel contribution by introducing the concept of “spontaneous collu-
sion,” where collusive outcomes emerge endogenously from the learning process itself, without
7
Uncertainties such as the probability of an entrant being born and deciding to stay or exit.
8
For a broader overview of collusive outcomes in Q-learning algorithms, see Horton (2023). For implementa-
tion of Q-learning algorithms in industrial organization settings, specifically Cournot models and step-up simula-
tion methods, see Waltman and Kaymak (2008).
9
For empirical examples of collusion emerging from market outcomes, see Assad et al. (2020).
relying on grim-trigger strategies. Bertrand et al. (2025), using a Prisoner’s Dilemma setup, ex-
tend these findings to deep Q-learning algorithms, showing that collusion is not limited to basic
Q-learning. Finally, Banchio and Skrzypacz (2022) explore collusion in auction environments,
finding that Q-learning algorithms can collude in first-price auctions but not in second-price
auctions, highlighting the role of market rules in enabling collusion. For a detailed overview
of the literature on algorithmic collusion see den Boer et al. (2024).
With the rise of firms using Large Language Models as advisors, the question arises whether
LLMs, capable of reasoning, can learn strategic behavior. Horton (2023), pioneer of running
economic experiments with LLMs, specifically with OpenAI-GPT 3, was the first to highlight
the possibility of using LLMs to replicate human-like behavior to run experiments at lower
costs. However, today the concern extends to the possibility of LLMs being used to discover
strategic behavior that may be suboptimal for market conditions. Fish et al. (2025) reinforce
this concern by showing that LLM-based pricing agents are capable of learning to collude in a
duopoly setting as well as an auction setting. Further, they show that certain prompts lead to
collusive outcomes faster. They extend the experiment to the auction setting and find similar
collusive outcomes. Other studies such as Wu et al. (2024) show that LLM agents are capable
of cooperating, Agrawal et al. (2025) show the possibility of collusive actions in double auction
settings. On the other hand, Keppo et al. (2025) show that collusive outcomes in LLM pricing
agents decrease when agents are heterogeneous.
Theoretical literature shows that exclusionary practices such as predation and exploita-
tive practices such as collusion both are possible rational outcomes when firms have strategic
foresight. Yet, the literature on algorithmic decision-making specifically focuses on coordina-
tion between firms. In this paper we seek to bridge the gap between theoretical industrial
organisation and nascent literature on lagorithmic decision making by asking whether LLMs,
when placed in dynamic market environments, can learn predatory strategies — sacrificing
short-term profits to discipline or eliminate entrants — in much the same way that earlier
algorithmic studies showed they could learn to collude.

2 Dynamic model of predation


Following Rey et al. (2023), we consider an infinite horizon, discrete-time game in which
an incumbent, denoted I, faces a sequence of potential entrants, denoted E. Each period t may
start in either of the two states: (i) a monopoly state, denoted M , where I is the only firm in
the market but E may emerge; or (ii) a competitive state, C, in which both I and E exist in the
market, but E may exit. If E does not enter or exits, a new E may be born in future periods.
All firms have discount factor δ ∈ (0, 1).
Monopoly state M : Initially, I is the only Firm in the market and sets some quantity q I .
With probability η, an entrant is born, upon which it decides whether to enter the market. If
E is not born or decides not to enter, I obtains profit π̄ M
I
and the next period begins in state M .
If E decides to enter the market, it incurs a one-time entry cost k. Upon entry, Firm E chooses
its quantity q E as a best response to Firm I’s committed quantity q I . The payoff of Firm I and
E is π M
I
and π M
E
respectively. The next period begins in state C.
Table 1 summarizes the payoffs of the firms in state M .

E Enters E Stays out

State M πM
I
, πM
E
−k π̄ M
I
,0

Table 1: Payoffs of I and E in state M

Competitive state C: I and E both exist in the market. First, I announces whether to
predate or to accommodate. Having observed I’s decision, E decides whether to stay or to exit.
If E decides to stay, the firms compete in a Stackelberg game where Firm I acts as the leader
with quantity q I , and Firm E follows by choosing quantity q E . The next period starts in state
C with probability (1 − γ) and in state M with probability γ. Parameter γ is an exogenous
probability that an entrant exits the market for reasons unrelated to the incumbent’s conduct.
If E exits, the next period begins in state M .
Overall, in the competitive state, one of the four possibilities may arise: (i) I accommodates
and the E stays, (ii) I accommodates and the E exits, (iii) I predates and the E stays and (iv)
I predates and the E exits. Table 2 summarized the payoffs of the firms in state C.

E Stay E Exit

I Accommodates πAI, πAE π̄AI, 0


State C
I Predates π PI , π PE π̄ PI , 0

Table 2: Payoffs of I and E in state C

Appendix A provides microfoundations for the payoff structure in Tables 1 and 2, based on a
Stackelberg quantity-setting game. While Rey et al. (2023) assumes no product differentiation,
we generalize payoffs by allowing for product differentiation.
Figure 1 provides an overview of the transitions between the states.
Yes
Entrant chooses q EM I , π E − k)
Payoff:(π M M t + 1 state: Competitive

State t: Monopoly Incumbent chooses q IM Entrant born? (η)

1−η

Payoff:(π̄ M
I , 0)
t + 1 state: Monopoly

State t + 1: Competitive
(1 − γ)

E decides to stay and chooses qAE Payoff: (π̄AI, πAE )


γ

I announces Accomodate and chooses qAI State t + 1: Monopoly

Entrant exits Payoff:(π̄AI, 0) State t + 1:Monopoly State t + 1 : Competitive

State t: Competitive

(1 − γ)
Entrant decides to stay and chooses q EP Payoff:(π PI , π PE )

γ
I chooses Predate and chooses q IP State t + 1 : Monopoly

Entrant exits Payoff:(π̄ PI , 0) State t + 1: Monopoly

Figure 1: Sequence of events in Monopoly and Competitive states respectively and state tran-
sitions.

2.1 Markov Perfect Equilibria


Following Rey et al. (2023), we focus on pure strategy Markov Perfect Equilibria (MPE),
where each firm’s strategy depends solely on the current state of the game, not on the whole
history of play. A Markov strategy for the incumbent firm I specifies whether to predate or
accommodate when the market is in the competitive state. For a newly born entrant E, the
strategy involves two components: (i) a decision of whether to enter or stay out when the
market is in the monopoly state, and (ii) for every possible decision taken by I, a decision of
whether to remain in the market or exit in the competitive state. A profile of Markov strategies
forms an MPE if in every state of the game, given the strategy of the opponent, each firm
maximizes its expected discounted payoff from that point onwards — that is, strategies form
a Nash equilibrium in every state of the game.
Define I’s cost-benefit ratio of predation when newborn entrants enter the market:

πAI − π̄ PI
λ=  . (1)
I − πI + η πI − πI
(1 − η) π̄ M A M A

The numerator, πAI − π̄ PI , is the profit sacrifice incurred in the predation period and the denom-
inator is the expected monopolization benefit obtained in the next period:
 
(1 − η) π̄ M − πA
+η π M
− πA
. (2)
| I {z I } | I {z I }
if no new entrant if new entrant

We can similarly define I’s cost-benefit ratio of predation when newborn entrants stay out
of the market:
πAI − π̄ PI
λ̄ = . (3)
π̄ M
I − πI
A

The numerator, πAI − π̄ PI , is the same as before, while the denominator π̄ M


I
− πAI changes to
reflect the fact that in the next period there will be no entry even if an entrant is born.

Proposition 1. There are three types of pure-strategy Markov Perfect Equilibria:

1. Accommodation: Firm I accommodates entry and each newborn entrant E enters the mar-
ket and stays until exogenous exit. Such equilibrium exits if and only if

1 − δ(1 − γ) P
πAE ≥ − πE (4)
δ(1 − γ)

or
δ(1 − γ)
λ≥ . (5)
1 − δ(1 − η − γ)
Everything else equal, accommodation equilibrium is easier to sustain when the probability
of entrant’s exit γ or probability of entrant’s birth η are higher.

2. Predation: Firm I predates in case of entry and each newborn entrant E enters the market
to stay for one period only. Such equilibrium exits if and only if

πM
E
≥k (6)

and
δ(1 − γ)
λ≤ . (7)
1 − δ(1 − η − γ)
Everything else equal, predation equilibrium is easier to sustain when the probability of
entrant’s exit γ or probability of entrant’s birth η are lower.
3. Monopolization: Firm I predates in case of entry and each newborn entrant E stays out of
the market. Such equilibrium exits if and only if

πM
E
≤k (8)

and
δ(1 − γ)
λ̄ ≤ . (9)
1 − δ(1 − γ)

See Appendix B.2 for the proof of Preposition 1.


Accommodation is self-sustaining if condition (4) holds. Intuitively, even if I were to pre-
date today, E would not exit, since E’s long-run profit when I continuously accommodates
outweighs the one-period loss from I’s predation. If (4) does not hold, then accommodation is
profitable for the incumbent only if λ is sufficiently high, i.e. if the expected future gains from
monopolization are small enough relative to the cost of predation. Formally, λ must satisfy
condition 5, which is easier to satisfy when (i) probability of exogenous exit, γ is higher and
(ii) probability of entrant’s birth, η is higher. Intuitively, a higher probability of exogenous
exit (γ) weakens the need for predation (since E will eventually leave anyway), while a higher
probability of new entrants (η) makes predation less attractive because its effect will not be
lasting and predation would need to be repeated.
For a predatory equilibrium to exist, E must enter in state M and exit immediately in the
next period when I predates. For entry to occur in state M , E’s one-period profit must cover
the entry cost. For predation to be profitable, the incumbent’s condition 7 must also hold, that
is, λ must be sufficiently low.
In a monopolization equilibrium, entry is deterred outright: E’s expected profit in the period
of entry does not cover E’s entry cost k, and, since any entry would trigger predation by 7, E
has no hope of recovering initial losses in the future.
There exists a region of parameters where accommodation and predation equilibria coexist.
Specifically, coexistence arises when E’s profit upon entry is high enough to cover his entry cost,
and either (i) both (5) and (6) hold with equality, or (ii) (6) holds strictly and, in addition,
E’s continuation value from when I accommodates satisfies (4). In such coexistence regions,
predation yields a strictly higher payoff for I. The reason is that entry occurs irrespective of I’s
strategy. If I accommodates, I is locked into low duopoly profits indefinitely, if γ = 0, or until
an exogenous entrant’s exit, if γ > 0. In contrast, under predation, I sacrifices profit for one
period but restores monopoly profit thereafter. Thus, whenever accommodation and predation
are both feasible, predation is strictly more profitable for the incumbent.
2.2 Other equilibria
The MPE captures a steady-state pattern of behavior in which strategies are best responses
in every state, and firms make forward-looking decisions that account for the dynamic conse-
quences of their actions. However, because Markov strategies depend only on the current state
and not the full history of play, MPE substantially restricts the set of equilibrium outcomes.
Without the Markovian restriction, the Folk Theorem implies that I and E could sustain any
payoff above their individually rational levels as an equilibrium outcome in the competitive
state, provided δ is sufficiently high and γ is sufficiently low. In particular, this opens the door
for the firms to learn to collude on monopoly outcomes. To disrupt such potential collusion,
we introduce the parameter γ — an exogenous probability that the entrant exits the market
for reasons unrelated to the incumbent’s conduct. This exogenous exit probability reduces the
entrant’s effective discount factor, from δ to δ(1−γ), making future cooperation less valuable.10
Alternatively, γ can be interpreted as the probability with which the game exits the competi-
tive state, thereby generating more opportunities for I to face newborn entrants and potentially
learn to exclude them. However, while γ successfully weakens the scope for collusion, it also
reduces the payoff from predation, as Proposition 1 shows. If the entrant may exit regardless
of I’s conduct, there is less incentive to engage in costly predatory pricing — making accom-
modation a more attractive strategy for the incumbent.

3 Experimental Design and LLM Agent Implementation


Our experimental setup employs LLM agents to act as quantity and binary decision-makers
for firms in a dynamic model of predation described in Section 2. Existing, LLMs can grasp
conceptual information and produce creative output in structured formats. Specifically, we
use OpenAI’s GPT-4.1 model, launched in April 2025. According to OpenAI (2025), GPT-4.1
is more effective than previous OpenAI models in independently performing tasks, following
instructions, and, importantly, requiring minimum hand holding.
The LLM agents are designed to strategically make the quantity decision for the incum-
bent (I) and the Entrant (E). Aligning with our Stackelberg framework in Appendix A, the
incumbent LLM agent always makes the quantity decision first, which is observed by the en-
trant LLM agent, who then first makes the binary decision to ENTER or NOT ENTER in the
Monopoly state, and STAY or EXIT in the Competitive state, followed by the quantity decision
if he decides to enter or stay. The only difference to the theoretical setup in Section 2 is that
in a competitive state C, we do not allow the incumbent LLM agent to announce his decision
10
Rey et al. (2023) does not allow the competitive state to end with an exogenous probability.
to predate or accommodate.
Following Fish et al. (2025), who study LLM-based pricing agents in a repeated Bertrand
oligopoly setting,11 a single experimental run consists of 300 periods. Each experimental run
always starts in the Monopoly state (M ), while each period may start in the Monopoly state
(M ) or the Competitive state (C), depending on the sequence of events described in Figure 1.

3.1 Agent Role and State Management


The central unit for implementing the agent role and state management is a standardized
Firm class structure.
The Incumbent firm (I) is initialized at the start of the simulation, while the Entrant firm
(E) is dynamically created when it is born. Each new Entrant is treated as a distinct firm and
is assigned a unique firm_id. If a firm exits the market, it is removed from the stimulation.
Each Firm maintains its history in three parts: (i)detailed_history, which is a list of
past data, including its own and its opponent’s quantity choices, binary decisions, realized
profits, and the market state. Through the prompt, a summary of the 25 most recent periods
is passed to the LLM using this list. (ii)current_plans, which is a string that stores the LLM
agent’s strategic plans and long-term goals; and (iii)current_insights, which is a string that
captures the LLM agent’s real-time observations and learning, essentially, its "thought process".
In each interaction, the LLM dynamically updates both (ii) and (iii). This allows the LLM
agents to observe past lessons, strategise future plans and execute tasks.

3.2 Prompt Engineering


The Firm objects communicate with the LLM through a carefully designed prompt engi-
neering strategy. This ensures that the LLM generates structured and comprehensive responses,
leveraging its ability to interpret complex information, learn over time, and provide outputs
that can be parsed.
Our prompts showcase features of structured, contextual, zero shot chain-of-thought and
persona prompting. Our prompts require the LLM agents to follow a structured template, en-
suring a parsable response, provide context through market information and history, guide the
LLM agents to break down their thought process, and lastly to ground the LLM agent’s rea-
soning, we give them a persona of a rational firm. Further, the LLM agent’s objective function
is explicitly to "maximize your expected long-term discounted profit". This specifically aligns
the LLM agents to the persona of a rational economic agent. However, we do not provide the
11
Fish et al. (2025) closely adopts the economic environment of Calvano et al. (2020), where LLM-based pricing
agents learn to collude.
LLM agents with any instructions on how to reason, making it a form of zero-shot prompting12 .
Our primary focus is not on how these specific methods influence LLM outcomes, but on the
resulting economic behavior of the agents within the simulation.
We implement concise prompts which consist of two main components:
(i) The system_prompt, which sets the LLM agent’s role as a strategic agent and ensures
that its responses adhere to a required format.
(ii) The user_prompt, which is dynamically updated for each period and decision point.
This sets the persona of the LLM agent as the rational firm.
We use three primary user prompts: Prompt_incumbent_quantity, Prompt_entrant
_quantity, and Prompt_entrant_binary_decision. These prompts first explain the fun-
damental game rules, such as the variation of market states (Monopoly or Competitive) and
incumbent’s permanence. They then provide key market parameters, including the demand
intercept a13 , costs14 , product differentiation θ , and the discount factor δ. It also included
contextual information, such as the current game state, the probability of a new entrant be-
ing born (η), probability of the exogenous entrant’s exit (γ), and a summary of past market
information. The prompts conclude by emphasizing the sequence of play (the incumbent sets
quantity first), the goal of maximizing long-term profits, and the specific task for the period.
The quantity prompts ask the LLM to set a quantity, while the binary prompts require a decision
of either ENTER/NOT ENTER or STAY/EXIT.
Crucially, all user_prompts incorporate a common prompt suffix15 . These are not stan-
dalone prompts but reusable templates that ensure the LLM agent’s thought process is struc-
tured, recorded, and updated. Most importantly, the final lines of these common prompts
ensure that the LLM’s output is easily parsable. This standardized structure is a core element
of our prompt engineering, guaranteeing that the LLM’s responses can be consistently inter-
preted by our simulation code. For instance, to summarize a full user prompt for the incumbent
should look as follows:

Full user prompt = PROMPT_INCUMBENT_QUANTITY + COMMON_PROMPT_SUFFIX_QUANTITY

Appendix C.1 includes the full templates for the System and User Prompts used in the
simulation. The prompts update dynamically with specific parameters, history data, and states
12
See Boonstra (2025) to understand various prompt engineering techniques.
13
The inclusion of the demand intercept a ensures that the LLM agents have a market size threshold for making
quantity decisions.
14
The Prompt_incumbent_quantity includes fixed cost z I , while the entrant prompts include fixed cost z E
and the one-time entry cost K.
15
The quantity prompts use common_prompt_suffix_quantity, while the binary prompts use
common_prompt_suffix_binary
during each period t.

3.3 Implementation Details and Response Parsing


The Firm objects interact with GPT-4.1 through the OpenAI Python library’s [Link].
[Link] method. The temperature parameter is set to 1 which is a default for
such LLM models. Further, we limit the max_tokens parameter16 to manage computational
costs and ensure that the output remains focused on the decision task.
To ensure unobstructed simulation runs, we implement four robust handling mechanisms:
(i) Retry Mechanism: Each LLM call is in a retry loop with up to 10 retries. This means in
case of an interruption or incorrect output format the LLM is called upon upto 10 times until
an output in the correct format is received. This reduces the likelihood a simulation failing due
to an external factor.
(ii) Default Fallback: If the LLM fails to return a valid output after all retries, the output is
set to a predefined default. For quantity decisions, this default is 0, and for binary decisions,
this default is EXIT. Such fallbacks are logged and analyzed post-simulation.
(iii) Regular Expression Parsing: We use regular expressions to reliably extract key decision
elements from the LLM’s free-form responses, matching the structured sections defined in our
prompt templates.
(iv) Detailed Logging: Every user_prompt sent to the LLM and every raw reply_content
received are saved for audit and analysis.
Further, for each period in a given stimulation run, data, including the market state, binary
decisions, realised profits, were logged in CSV files. Lastly, for reproducibility of the stochastic
elements of the simulations such as the probability of birth (η) and the probability of exogenous
exit (γ) the random seed was fixed.

4 Experiment Runs
In this section, we turn our attention to the main experimental runs and results we observe.
In order to test whether LLM agents adapt strategic behavior consistent with our theoretical
predictions, we simulate distinct strategic environments. For tractability, we adopt baseline
economics parameters across all runs: demand intercept a = 100, demand slope b = 1, in-
cumbent’s fixed cost z I = 300, entrant’s fixed cost z E = 150, entrant’s entry cost k = 450,
16
The max_tokens parameter limits the length of the LLM’s output for a given call. Specifically, the limit is set
to 1500 tokens.
probability of exogenous entrant exit γ = 0.1, and discount factor δ = 0.9517 .
We then vary parameter θ (degree of product differentiation). This yields two main ex-
perimental variations: (i) Variation 1: θ = 1, where theoretically both accommodation and
predation are viable strategies; (ii) Variation 2: θ = 0.5, where theoretically only accommo-
dation is a viable strategy. Further, within each variation, we vary η (probability of an entrant
being born): (i) η = 1, a new entrant is definitely born; (ii) η = 0.5, a new entrant is born or
not born with equal probability.
Within each sub-variation, we run 10 experimental runs of 300 periods each. To ensure
comparability, we use different random seeds for all 10 runs when η = 0.5, and the same 10
random seeds for all runs when η = 1. Given the limited number of runs, we focus on descrip-
tive and robustness analyses rather than regression-based inference. The insights we draw
from these multiple runs and representative trajectories provide a thorough understanding of
strategic behaviour in the model.

4.1 Variation 1: θ = 1
We first examine the capability of the LLM agent to engage in predation in settings where
both predation and accommodation equilibria are theoretically feasible. Theoretically, In the
Monopoly state, we expect LLM agents to optimize profits by choosing quantity q IM = 50, while
in the Competitive state we expect them to engage in predation by choosing quantity q IP = 75.5
or accommodate by choosing qAI = 50 (here, qAI = q IM ) .
To illustrate the agent’s behaviour, we present results from representative runs with η = 0.5
and η = 1 (Figure 2), using the same random seed for both simulations. In both representative
runs, the Monopoly state (M ) arises in 76% of periods, while the Competitive state (C) arises in
24%. Each simulation begins in the Monopoly state, where an entrant (E) is born. In period 1,
the incumbent LLM agent responds to the threat of entry by choosing a quantity well above
the theoretical monopoly benchmark, consistent with an attempt to deter entry18 .
Figure 3 presents the incumbent’s state-specific quantity choices over time across all exper-
imental runs. Subplot 3a reports outcomes for η = 0.5, while Subplot 3b reports outcomes for
η = 1. Across both parameterizations and in both states, the incumbent consistently selects
17
Note: Choosing parameter values such as a = 100 ensures that quantity decisions are expressed as two-digit
numbers, which are easier for the LLM agent to interpret — for example, “Incumbent sets quantity 60” rather
than “Incumbent sets quantity 0.6”. Our choice of z I = 2z E follows standard IO assumptions, and we set k = 3z E
to reflect that entry is costly.
18
In appendix C.2.1 we present the full log reply from the period 1 of the representative experimental run Figure
subplot 2a. The LLM agent clearly recognize the three viable options: a) to maintain the static monopoly output
of 50, b) to accommodate the entrant, or c) to adopt limit pricing by aggressively setting a quantity above the
monopoly threshold to deter entry. The entrant’s log reply confirms that the agent also recognized this deterrence
attempt and chose NOT ENTER.
(a) Firm quantities: η = 0.5 and θ = 1 (b) Firm quantities: η = 1 and θ = 1

(c) Entrant binary decisions: η = 0.5 and θ = 1 (d) Entrant binary decisions: η = 1 and θ = 1

Figure 2: LLM agent quantity decisions and binary decisions from representative runs when
θ = 1 while η = 0.5 and η = 1. The solid blue line in subplot 2a and 2b represents the
quantity produced by the incumbent firm, while the dashed lines in various colours represent
the quantities of individual entrants who have entered the market. Further, the red dashed
line represents q P , the purple dashed line represents q IM = qAI and the orange dashed line
I
denotes q EM = qAE . In subplot 2c and 2d, the grey dot denotes the birth of an entrant, the red
cross denotes the entrant’s choice to NOT ENTER, green denotes the entrant’s choice to ENTER,
the blue square denotes the entrant’s decision to STAY while the orange triangle denotes the
entrant’s decision to voluntarily EXIT. Lastly, the star denotes exogenous exit. In all subplots,
the background colour indicates the market state: a light coral region signifies a competitive
market, whereas a blue region denotes a monopoly state. The simulation parameters used are:
γ = 0.1, a = 100, b = 1, k = 450, z I = 300, z E = 150, δ = 0.95, θ = 1. In subplot 2a and 2c,
η = 0.5, while in subplot 2b and 2d, η = 1.

quantities that lie above the theoretical monopoly benchmark but below the theoretical preda-
tory threshold. The monopoly benchmark is q IM = 50, i.e., the profit-maximizing monopoly
output, while the theoretical predatory threshold is q IP = 75.51, the minimum quantity consis-
tent with deterring entry.
Averaging across 300 periods and 10 runs per η, the incumbent’s mean quantity is 63.01
in the monopoly state and 66.39 in the competitive state when η = 0.5, compared to the
monopoly benchmark of 50 and the predatory benchmark of 75.51. When η = 1, the mean
quantity rises to 70.33 in the monopoly state and 67.94 in the competitive state, again lying
strictly between the monopoly (50) and predatory (75.51) theoretical benchmarks. These
outcomes suggest that the incumbent strategically expands output above the monopoly level to
deter entry in the monopoly state, and converges toward predatory behavior in the competitive
state.

(a) η = 0.5

(b) η = 1

Figure 3: The figure shows the quantity choices of the incumbent LLM agent across 300 periods
and all ten runs. Simulation parameters: θ = 1, γ = 0.1, a = 100, b = 1, k = 450, z I = 300,
z E = 150, and δ = 0.95. Note that the solid blue line refers to the mean quantity in monopoly
state while the solid red line refers to the mean quantity in competitive state. The dashed red
line represents the theoretical predatory quantity (75.51) while the dashed blue like representa
the theoretical monopoly quantity (50).

Furthermore, the first time the incumbent LLM agent selects a quantity within 10% of q IP
occurs, on average, in period 32.3 when η = 0.5 and in period 28.3 when η = 1. This suggests
that the incumbent does not immediately recognize the sufficient predatory quantity.
Figure 4 presents the distribution of quantity choices during the last 100 periods. The
results confirm that the median quantity is consistently higher in the competitive state than
(a) η = 0.5

(b) η = 1

Figure 4: Box plot of the incumbent’s quantity choice and realized profits in the last 100 pe-
riods. Simulation parameters: θ = 1, γ = 0.1, a = 100, b = 1, k = 450, z I = 300, z E = 150,
δ = 0.95.
in the monopoly state, regardless of η. In the monopoly state, the mean quantity is 65 when
η = 0.5 and 70 when η = 1, both well above the theoretical monopoly benchmark of q IM = 50.
This indicates that the incumbent LLM agent does not maximize immediate profits but instead
seeks to deter potential entrants by signalling aggression. In the competitive state, the mean
quantity is 67.64 when η = 0.5 and 74.12 when η = 1, values that are extremely close to and
within 10% of the theoretical predatory quantity (75.5). This pattern suggests that the LLM
agent responds aggressively to competition and systematically attempts predation.
Further, the incumbent’s profit in the monopoly state remains below the theoretical maxi-
mum profit, π M
I
= 2200, in both cases. This indicates that the incumbent is willing to sacrifice
short-term profit to maintain market dominance. In the competitive state, profits are even
lower, highlighting that predatory strategies are costly for the incumbent.
Further, to evaluate whether the behavior of the incumbent LLM agent aligns with these
theoretical predictions, in monopoly state, we classify the agent’s quantity as optimal or sub-
optimal. The agent’s quantity choice is considered optimal if it is in the 10% range of the
theoretically optimal quantity. In Competitive state, we classify the agent’s quantity decision as
predatory or accommodating if it falls within the 10% region of the corresponding theoretical
quantity. Actions outside these ranges are classified as neither. We find that the incumbent
chooses optimal quantity in the Monopoly state only 16.2% of the time when η = 0.5 and
only 0.99% of the time when η = 1. In the competitive state, the LLM agent engages in
predation 82.23% of the time when η = 0.5 and 93.15% of the time when η = 1. These
results are broadly intuitive: in the competitive state, the agent strongly favours predation,
consistent with the theoretical incentive to deter entry. By contrast, in the monopoly state, the
agent rarely selects the exact theoretical optimum. This pattern suggests that the LLM tends to
“overproduce” even when immediate competitive pressure is absent, perhaps reflecting a bias
toward deterrence or the difficulty of learning the precise monopoly optimum.
Then, for each run, in order to track the strategic trend, we compute a 30-period rolling
average of these classifications19 for each experimental run, separately for the Monopoly and
Competitive states. This provides a dynamic measure of the agent’s propensity to choose op-
timal or predatory strategies over time within each state, rather than by absolute simulation
period. The rolling-window averages are then aggregated across runs to produce the curves
shown in Figures 5 (monopoly state) and 6 (competitive state). Values above the 0.5 line
indicate that optimal (in Monopoly) or predatory (in Competitive) actions are more prevalent.
Figure 5 indicates that in the Monopoly state M , the incumbent does not settle on a stable
strategy when η = 0.5 and performs even worse when η = 1. These results indicate that the
19
We generate a state specific binary variable. For Monopoly state: Optimal=1 and Suboptimal=0. For Com-
petitive state, Predatory=1 and Accommodating=0. The rolling average we calculate is of this binary column.
(a) η = 0.5 (b) η = 1

Figure 5: The share of optimal actions in Monopoly state M when θ = 1. The 0.5 line is the
indifference benchmark — i.e., the point at which the incumbent LLM agent is equally likely to
engage in optimal or suboptimal strategies. The values above the 0.5 line indicate that optimal
actions are more prevalent.

(a) η = 0.5 (b) η = 1

Figure 6: The share of predatory actions in Competitive state C when θ = 1. The 0.5 line is the
indifference benchmark — i.e., the point at which the incumbent LLM agent is equally likely
to engage in predatory or accommodating strategies. The values above the 0.5 line indicate
that predatory actions are more prevalent.

incumbent fails to learn to optimize profits in the Monopoly state, irrespective of the entry
threat. In fact, when the entry threat is definitive, the incumbent makes highly suboptimal
choices toward the end of the simulation.
Further, in the Competitive state C (see Figure 6), we observe that the LLM agent initially
exhibits a bias against predation but quickly learns to choose predatory actions, with the rolling
share of predation crossing the indifference benchmark (0.5 line) and stabilizing well above
it. Predatory behaviour emerges even more rapidly when η = 1, indicating that a higher entry
threat induces greater aggressiveness. This finding runs counter to our theoretical prediction
from Proposition 1, which suggested that increasing the probability of entry, η, would reduce
the attractiveness of predation. Instead, the simulation results suggest that the incumbent LLM
agent finds it beneficial to establish a predatory reputation precisely when entry threats are
high, possibly reflecting the agent’s reinforcement-learning bias toward strategies that ensure
future dominance rather than short-run profit maximization.
Further, to analyze the incumbent’s impact on Entrants, we evaluate two key metrics: the
entrant success rate and the entry deterrence rate. An Entrant is defined as successful if their
cumulative lifetime profit is greater than zero. To measure the deterrence rate, i.e., the incum-
bent’s ability to discourage potential entrants, we identify all entrants born who choose not to
enter. The entrant success rate is defined as the percentage of successful entrants relative to
all entrants who entered the market, while the deterrence rate is the percentage of entrants
who choose not to enter relative to all entrants born.
When η = 0.5, the entrant success rate in the first 100 periods is 39.5%, dropping to 7.3%
in the last 100 periods. This effect is even more pronounced when η = 1, with the success rate
decreasing from 19.7% in the first 100 periods to a negligible 1.3% in the last 100 periods.
This indicates that when η = 1, the LLM agent effectively eliminates any potential profit for
the entrant. The deterrence rate is consistently high: 71.90% for η = 0.5 and 79.59% for
η = 1, with negligible changes as the simulation progresses. Overall, this demonstrates that
the incumbent establishes a reputation as a predator early in the simulations.
Lastly, strong evidence that LLM agents successfully achieve predation over time is provided
by the average lifespan of an entrant. We calculate an entrant’s lifespan as the difference
between entry and voluntary exit periods. When η = 0.5, the average lifespan is 2.40 periods
in the first 100 periods, decreasing to 1.59 periods in the last 50 periods. When η = 1, the
average lifespan is 2.16 periods in the first 100 periods, decreasing to 1.51 periods in the last
50 periods. These results indicate a clear convergence toward a hit-and-run equilibrium.
To sum up, we conclude that: (i) the incumbent LLM agent does not recognize the predatory
quantity immediately; (ii) in the Competitive state (C), the incumbent LLM agent successfully
learns to predate and does not accommodate; (iii) in the Monopoly state (M ), the incumbent
LLM agent does not learn to optimize monopoly profits, and in fact, makes quantity decisions
closer to the predatory quantity, indicating aggression and an entry-deterrence stance; (iv)
the incumbent LLM agent predates more aggressively when η = 1; and (v) on average, all
simulations converge to a hit-and-run equilibrium.

4.2 Variation 2: θ = 0.5


We now examine whether the LLM agent is able to adopt an accommodation strategy in
settings where the accommodation equilibrium is theoretically predicted to hold. For η = 0.5,
the incumbent’s theoretical optimal quantity choices are q IM = 46.66 in Monopoly state M and
qAI = 42.85 in Competitive state C; for η = 1, both coincide at 42.85.
As before, we begin by illustrating the behaviour of representative agents using represen-
tative runs with η = 0.5 and η = 1 (Figure 7). For η = 0.5 (Subplot 7a), the Monopoly state
(M ) occurs in 28.7% of periods, while the Competitive state (C) occurs in 71.3% of periods.
For η = 1 (Subplot 7b), the state frequencies are nearly identical, Monopoly state (M ) occurs
in 21.7% of periods and Competitive in 78.3%. The corresponding entrant binary decisions
are shown in Subplots 7c and 7d.

(a) Firm quantities: η = 0.5 and θ = 0.5 (b) Firm quantities: η = 1 and θ = 0.5

(c) Entrant binary decisions: η = 0.5 and θ = 0.5 (d) Entrant binary decisions: η = 1 and θ = 0.5

Figure 7: LLM agent quantity decisions and binary decisions from representative runs when
θ = 0.5 while η = 0.5 and η = 1. The solid blue line in subplot 7a and 7b represents the
quantity produced by the incumbent firm, while the dashed lines in various colours represent
the quantities of individual entrants who have entered the market. Further, the red dashed line
represents q P , the blue dashed line represents q IM , the grey dashed line qAI,the purple dashed
I
line denotes q EM ,the green dashed line denotes qAE . In subplot 7c and 7d, the grey dot denotes the
birth of an entrant, the red cross denotes the entrant’s choice to NOT ENTER, green denotes
the entrant’s choice to ENTER, the blue square denotes the entrant’s decision to STAY while
the orange triangle denotes the entrant’s decision to voluntarily EXIT. Lastly, the star denotes
exogenous exit. In all subplots, the background colour indicates the market state: a light coral
region signifies a competitive market, whereas a blue region denotes a monopoly state. The
simulation parameters used are: γ = 0.1, a = 100, b = 1, k = 450, z I = 300, z E = 150,
δ = 0.95, θ = 0.5. In subplot 7a and 7c, η = 0.5, while in subplot 7b and 7d, η = 1.
The overall strategy of the representative agents in Figure 7 is neither accommodating nor
predatory. Instead, the incumbent adopts an aggressive stance, consistently choosing quantities
above the monopoly threshold. On average, the incumbent produces 63.58 in the Monopoly
state (M ) and 61.10 in the Competitive state (C) when η = 0.5, and 66.75 (M ) and 60.48 (C)
when η = 1.20
The same pattern persists across all experimental runs (Figure 8). When η = 0.5, the
incumbent’s average quantity choice is 62.85 in the Monopoly state (M ) and 58.73 in the
Competitive state (C). When η = 1, the averages are 62.38 (M ) and 54.42 (C). These out-
comes are not close to the theoretical monopoly-optimal or accommodating strategies, nor do
they reflect predatory behaviour. Instead, the incumbent’s choices are more than 25% above
the theoretical benchmarks q IM and qAI, indicating a distinctly aggressive strategy.
In the last 100 periods, when η = 0.5, the incumbent LLM agent’s average quantity de-
creases to 58.44 in the Monopoly state and increases to 60.99 in the Competitive state. When
η = 1, the corresponding averages increase to 64.05 in the Monopoly state and 55.47 in the
Competitive state. These patterns indicate that the incumbent LLM agent does not converge
to the theoretically predicted strategies. The average quantities remain substantially above
the theoretical benchmarks, suggesting that the agent persistently adopts an aggressive stance
rather than accommodating or predatory strategies.
The incumbent LLM agents never select quantities close to the theoretical predatory thresh-
old, q P = 151.51. This leads us to conclude that the agents do not learn to engage in full
I
predation. Instead, their most aggressive behavior is far more moderate: across all runs, the
average maximum quantity is 80.70 in the Monopoly state and 79.40 in the Competitive state
when η = 0.5, and 73 in both states, when η = 1.
Further, to determine if the LLM agent learns a particular strategy overtime, we categorize
the state specific LLM quantity choices into optimal, aggressive and other in Monopoly state
and accommodating, aggressive and other in Competitive state. First, we confirm that across
all runs the LLM agent’s quantity choice is not in the 10% region of the predatory quantity
threshold. As before, in the monopoly state, we classify a quantity choice as optimal if it falls
in the 10% region of the theoretical optimal q IM . In competitive state, we classify a quantity
choice as accommodating if it falls in the 10% region of the theoretical optimal qAI. In both
states, a quantity choice is classified as aggressive if it is in the range 10% above q IM but 10%
below q P . All other choices are labelled as neither.
I
We find that when η = 0.5, in Monopoly state, LLM incumbent’s quantity choices are opti-
20
Appendix C.2.2 provides the full log reply from period 299 of the representative run shown in Subplot 7b.
The incumbent’s response reveals an awareness that its quantity choice influences the entrant’s decision. With
a choice of 58, the incumbent explicitly states its aim is to deter entry, which indeed leads the entrant agent to
choose NOT ENTER, observing a quantity above the theoretical monopoly threshold.
(a) η = 0.5

(b) η = 1

Figure 8: The figure shows the quantity choices of the incumbent LLM agent across 300 periods
and all ten runs. The simulation parameters used are: θ = 0.5, γ = 0.1, a = 100, b = 1,
k = 450, z I = 300, z E = 150, and δ = 0.95. Note that the solid blue line refers to the mean
quantity in monopoly state while the solid red line refers to the mean quantity in competitive
state. The dashed red line represents the theoretical predatory quantity (151.50), the dashed
blue line represents the theoretical monopoly quantity (46.44 when η = 0.5 and 42.86 when
η = 1). Lastly, the gray dashed line represents the theoretical accommodation quantity(42.86).

mal only 10.13% of times, they are aggressive 89.59% of times. In Competitive state, quantity
choices are accommodating only 6.3% of times and aggressive 92.5% of times. When η = 1,
quantity choices in Monopoly state are optimal only 15% of times and aggressive 82.14%
of times. In competitive state, actions are accommodating 18.62% of times and aggressive
77.60% of times.

(a) η = 0.5 (b) η = 1

Figure 9: The share of optimal actions in Monopoly state M when θ = 0.5. The 0.5 line is the
indifference benchmark — i.e., the point at which the incumbent LLM agent is equally likely to
engage in optimal or aggressive strategies. The values above the 0.5 line would indicate that
optimal actions are more prevalent.

(a) η = 0.5 (b) η = 1

Figure 10: The share of accommodating actions in Competitive state C when θ = 0.5. The
0.5 line is the indifference benchmark — i.e., the point at which the incumbent LLM agent is
equally likely to engage in accommodating or aggressive strategies. The values above the 0.5
line indicate that accommodating actions are more prevalent.

In Figure 9 and 10, we plot the aggregate 30-period rolling average of the quantity choice
classifications in Monopoly and Competitive state respectively. This shows the propensity of
the incumbent LLM agent to choose the optimal action in Monopoly state (in Figure 9 ) and
the accommodating action in Competitive state (in Figure 10).
When η = 0.5, the incumbent LLM agent’s play in the Monopoly state does not converge to
the optimal monopoly strategy (see Subplot 9a). Instead, it maintains a persistent bias toward
overproduction, reflecting an aggressive stance. In the Competitive state, the agent experi-
ments briefly with accommodation in the early periods, but this behaviour is not sustained. It
quickly converges toward consistently aggressive output (see Subplot 10a).
When η = 1, the incumbent’s behaviour is more variable. In the Monopoly state, the
agent intermittently explores the optimal monopoly strategy, though these attempts are not
sustained and it repeatedly reverts toward aggression (see Subplot 9b). In the Competitive
state, the incumbent does not explore accommodation; the rolling share of accommodating
actions remains close to 0.2 throughout, indicating a persistently aggressive stance (see Subplot
10b).
Further, the incumbent LLM agent’s aggressive stance does have a negative impact on the
Entrant success rate. In the first 100 periods the entrant success rate is 100% irrespective of
η. However, in the last 100 periods it drops to 70.59% when η = 0.5 and to 8.47% when
η = 1. Further, the Incumbent manages to deter entry with the aggressive stance with some
success. Particularly, the deterrence rate is 54.44% (when η = 0.5) and 53.09% (when η = 1)
in the first 100 periods. In the last 100 periods, the deterrence rate increases to 59.71% (when
η = 0.5) and 60.19% (when η = 1).
Lastly, the average lifespan of an entrant, when η = 0.5, is 2.95 periods in the first 100
periods which decreases to 1.10 periods in the last 100 periods. When η = 1, the average
lifespan of the entrant is 1 period throughout21 . This shows that the incumbent’s aggressive
behaviour leads to entrant’s exiting immediately upon entry.
To sum up, we conclude that (i) The incumbent’s choices do not align with the theoret-
ical predictions. (ii) Even when the theoretical environment is sound for accommodation,
a forward looking incumbent LLM agent adopts an aggressive stance in order to maintain its
monopoly. (iii) The incumbent LLM agent does not learn to optimize profits in Monopoly state.
Lastly, (iv) a forward looking Entrant LLM agent is discouraged from entry and encouraged to
exit immediately when the incumbent shows slight aggression.

5 Robustness Check
In this section, we test the robustness of our main findings from Section 4. First, we consider
an alternative setup where there is no exogenous exit (γ = 0). Second, we vary our prompts,
specifically, the prompt persona and objective function.

5.1 No exogenous exit


In this section, we examine the simulation outcomes when there is no probability of exoge-
nous exit. Setting γ = 0 allows us to examine the agent’s strategic behaviour in the absence of
exogenous exit, highlighting the role of forward-looking incentives.
21
There is only one run where the average entrant lifespan is 13 periods in the first 100 periods.
The other parameters are set to θ = 1 (homogeneous products), η = 0.5, and δ = 0.95.
We conduct 10 simulation runs of 300 periods each. For these parameter values, both accom-
modation and predation are viable strategies.
Across runs, the long-run outcome is either predation or Cournot-like accommodation. We
find that 7 out of 10 runs converge to predation. Of these, 5 converge early, in the sense that the
incumbent’s average quantity in the last 100 periods in the competitive state (C) lies within
10% of the theoretical predation quantity, q P = 75.5, i.e., [67.95, 83.05]. The remaining 2
I
predatory runs reach this range only in the last 50 or 25 periods. The average incumbent
quantity in the Monopoly state (M ) is 65.17 (averaged over 300 periods and 7 runs), 62.97
(last 100 periods and 5 runs), and 60.64 (last 50 periods and 2 runs).
The other 3 runs converge to a Cournot-like accommodation outcome. We classify a run
as Cournot-like accommodation if, over the evaluation window (last 100 / 50 / 25 periods as
stated), the incumbent’s average quantity in state (C) is approximately 37.75 units.22 In the
Monopoly state, the average incumbent quantity is 53.35 (averaged over 300 periods and 3
runs). Here, the average in the last 100 or 50 periods is not relevant as Monopoly state does
not occur once the convergence is achieved. See Figure 11 for representative runs of both
types.

(a) Representative predation run (b) Representative Cournot-like accommodation


run

Figure 11: Quantities and event timelines from representative runs. The solid blue line repre-
sents the incumbent’s quantity, while dashed colored lines represent individual entrants. Back-
ground shading indicates market state: light coral = competitive (C), blue = monopoly (M ).
The red dashed line denotes q P , the purple dashed line q IM = qAI, and the orange dashed line
I
q EM = qAE . Subplot (a) shows convergence to predation; subplot (b) shows convergence to
Cournot-like accommodation.
22
For reference, in the simultaneous-move Cournot benchmark with these parameters, the pure-strategy Nash
equilibrium quantity is 33.33.
Irrespective of the eventual outcome, 9 out of 10 runs begin with the LLM incumbent choos-
ing a quantity above the theoretical monopoly threshold q IM = 50. There is a single exception
that starts below q IM in the first few periods and ultimately converges to Cournot-like accom-
modation. See Figure 12.

Figure 12: Unique representative run. The incumbent initially sets quantity around the profit-
maximizing level, and the simulation converges to a Cournot-like accommodation outcome.
The red dashed line denotes q P , the purple dashed line q IM = qAI, and the orange dashed line
I
q EM = qAE .

While here we have focused on predation and Cournot-like accommodation outcomes, it is


also worthy to consider the absence of collusion in these simulations. Given the repeated-game
setting, homogeneous products, and a relatively high discount factor (δ = 0.95), one might
expect the possibility of tacit collusion to arise, consistent with standard folk-theorem results.
In principle, the incumbent could sustain higher joint profits by coordinating with entrants
rather than engaging in costly predation or accommodating at Cournot levels. However, across
all runs, we do not observe any evidence of such coordination: quantities do not converge to
collusive levels. This absence of collusion suggests that, at least in this parameterization, the
LLM agent does not discover or stabilize cooperative strategies, instead gravitates toward either
predation or Cournot-like outcomes.
To summarize, we find that: (i) predation is not guaranteed, (ii) when accommodation
occurs, the incumbent converges to the Cournot quantity, suggesting it either fails to recog-
nize or does not learn to exploit its first-mover advantage, (iii) the incumbent does not learn
the profit-maximizing monopoly quantity in state M (when predating) and in state C (when
accommodating), and (iv) no collusion is observed.
(a) Advisor Prompt Simulation

(b) Risk Averse Prompt Simulation (c) Bounded Memory Prompt Simulation

Figure 13: Incumbent and Entrant Quantities Across All Prompt Variations.
The figure presents the simulated firm quantities for an experimental run of 100 periods for
each of the three prompt variations. The solid blue line in each subplot represents the quantity
produced by the incumbent firm, while the dashed lines in various colours represent the quan-
tities of individual entrants who have entered the market. The background colour indicates the
market state: a light coral region signifies a competitive market, whereas a light blue region
denotes a monopoly state. The simulation parameters used are: γ = 0.1, η = 0.5, a = 100,
b = 1, k = 450, z I = 300, z E = 150, and δ = 0.95. Further, the red dash line represents q P
I
while the blue dash line represents q IM = qAI.

5.2 Prompt Variation


In our main experimental setup, the large language model (LLM) agents are assigned the
persona of a rational firm whose principal objective is to maximize long-run discounted profits
(see Appendix C.1 for our main experiment prompts). In this section, we explore how sensi-
tive agent behavior is to variations in both the underlying objective function and the assigned
persona. Specifically, we consider three alternative prompts that differ in the goals and behav-
ioral framing of the LLM agents. In each case, together with market-relevant data and history,
the LLM agent is explicitly asked to provide their analysis and rationale behind each decision.
Under the “analysis” and “rationale” headings, the agents are given pointers to guide their
thought process, resembling a structured Zero-Shot Chain-of-Thought (CoT) approach used in
our main experiment but with more explicit instructions.
We run one experimental run of 100 periods for each prompt variation using the same
parameters as our main experiment: η = 0.5, γ = 0.1, and θ = 1. Theoretically, both accom-
modation and predation equilibria hold for these parameter values (see Figure 13).

5.2.1 Advisor Prompt

In this variation, the LLM agent’s persona is that of a strategic business consultant whose
primary goal is to maximize the long-term discounted profit of their client. This aligns the
advisor persona closely with that of a textbook rational agent.
From the simulation results (Figure 13a), the advisor agent demonstrates a clear multi-
period strategy. It starts by choosing aggressive quantities to deter entrants and, as the threat
of entry decreases, gradually moves to a quantity just above the pure monopoly level. Overall,
this run results in a monopoly state 93 per-cent of the time and a competitive state only 7 per-
cent of the time. Out of 48 entrants born, only 6 successfully enter. In the first 50 periods, the
incumbent on average chooses 70.42 in monopoly and 69 in competitive states, demonstrating
strong deterrence. In the last 50 periods, it settles at 52.91 (monopoly) and 53.83 (compet-
itive), taking a profitable, less aggressive stance. The entrants consistently choose near-zero
quantities, showing that the incumbent successfully learns to deter entry.

5.2.2 Risk-Averse Prompt

Here, the agent persona is a risk-averse firm manager whose primary goal is to maintain
stable profits. The agent is designed to take safe, predictable decisions, resembling a conser-
vative, non-rational actor.
The simulation results (Figure 13b) show that the risk-averse agent quickly identifies the
theoretical monopoly quantity of 50 and does not deviate, even under entry threat. For the first
38 periods, it maintains this stable monopoly quantity. After experiencing its first entry—where
the entrant chooses 25 (the theoretical accommodating quantity)—the incumbent does not
accommodate but instead increases its quantity to 55. This forces the entrant to exit, and the
incumbent maintains this level until the end of the simulation. Overall, the monopoly state
occurs 98 per-cent of the time, with only 2 per-cent in the competitive state. Directing the
agent toward stability thus leads to a non-aggressive deterrence strategy that succeeds when
entrants are similarly conservative.
5.2.3 Bounded Memory Prompt

In this variation, the agent aims to maximize long-run profits but observes only the last 3–5
periods of market history. This models bounded rationality and tests whether access to deeper
historical records is required for multi-period strategy execution.
The simulation results (Figure 13c) reveal erratic, myopic behavior. The incumbent fails
to consistently deter entry, resulting in a monopoly state 74 per-cent of the time and a com-
petitive state 26 per-cent of the time. In the first 50 periods, the incumbent averages 48.79 in
monopoly and 31.76 in competitive states. In the last 50 periods, it settles at 33.64 (monopoly)
and 30 (competitive), while entrants choose 30 on average in competitive states. This shows
that myopic agents eventually converge toward the Cournot equilibrium rather than pursuing
aggressive deterrence.

To sum up, varying prompts demonstrates that LLM agents do not inherently adopt ag-
gressive deterrence behaviour but can be driven to it through a combination of a clear prompt
objective, a defined persona, and strategic information.

6 Discussion
In this paper, we focus on the possibility that large language models (LLMs) are capable of
learning predatory strategies. First, we build a dynamic Stackelberg framework of predation
based on Rey et al. (2023), and show that accommodation, monopolization, and predation
can arise as pure-strategy Markov Perfect Equilibrium outcomes, sustained under a persistent
threat of entry and exogenous exit. Second, we embed large language models—specifically
OpenAI’s GPT-4.1—into our economic environment to test whether LLM agents can learn such
strategies.
We conduct 10 experimental runs on each of our sub-variations, varying product differ-
entiation and the threat of entry. The first variation is chosen such that both predation and
accommodation are viable equilibrium strategies, whereas the second is chosen such that only
accommodation is viable. We observe that LLM agents learn to predate effectively when both
predation and accommodation are viable, but adopt an aggressive stance when accommoda-
tion is the only theoretically viable equilibrium strategy. In both cases, LLM agents do not learn
to optimize profits when they are the sole agents in the market.
From these findings, we highlight two broader implications. First, the strategic capabili-
ties of LLM agents extend beyond pricing and collusion. LLM agents can learn exclusionary
behavior requiring complex intertemporal strategies, even when such strategies involve short-
term losses. Second, our preliminary results suggest that the rise of LLMs introduces novel
challenges for competition policy.
We also recognize several limitations that need to be addressed. Our robustness checks
show that predation is not guaranteed when the probability of exogenous exit is zero. This
implies that further experimental runs with parameter variations are required to identify the
precise conditions under which an LLM agent learns predation. Second, the behavior of LLM
agents appears highly sensitive to prompt persona and objective functions, an area requiring
further exploration. Lastly, a major limitation of our study is the relatively small number of
experimental runs. And usage of only one LLM model that is Open AI’s-GPT [Link] was a bud-
getary constraint. This limited the size of our data set and prevented us from performing any
rigorous regression analysis. These limitations open avenues for us for our future work, both
in refining the experimental framework and in extending the analysis to richer environments
with endogenous entry and multi-agent learning.
References
A. Michael Spence, “Spence-LearningCurveCompetition-1981,” The Bell Journal of Economics,
1981, 12 (1), 49–70.

Agrawal, Kushal, Verona Teo, Juan J. Vazquez, Sudarsh Kunnavakkam, Vishak Srikanth,
and Andy Liu, “Evaluating LLM Agent Collusion in Double Auctions,” 7 2025.

Asker, John, Chaim Fershtman, and Ariel Pakes, “Artificial Intelligence and Pricing: The
Impact of Algorithm Design,” Technical Report 2021.

Assad, Stephanie, Robert Clark, Daniel Ershov, and Lei Xu, “Algorithmic Pricing and Compe-
tition: Empirical Evidence from the German Retail Gasoline Market,” SSRN Electronic Jour-
nal, 2020.

Bain, Joe S, “A Note on Pricing in Monopoly and Oligopoly,” Technical Report 2 1949.

Banchio, Martino and Andrzej Skrzypacz, “Artificial Intelligence and Auction Design,” in “in”
Association for Computing Machinery (ACM) 7 2022, pp. 30–31.

and Giacomo Mantegazza, “Artificial Intelligence and Spontaneous Collusion,” 9 2023.

Bertrand, Quentin, Juan Duque, Emilio Calvano, and Gauthier Gidel, “Self-Play Q-learners
Can Provably Collude in the Iterated Prisoner’s Dilemma,” 6 2025.

Besanko, David, Ulrich Doraszelski, and Yaroslav Kryukov, “The economics of predation:
What drives pricing when there is learning-by-doing,” American Economic Review, 2014, 104
(3), 868–897.

Bolton, Patrick and David S Scharfstein, “A Theory of Predation Based on Agency Problems
in Financial Contracting,” Technical Report 1 1990.

Boonstra, Lee, “Prompt Engineering,” Technical Report, Google 2 2025.

Cabral, Luis M B and Michael H Riordan, “The Learning Curve, Market Dominance, and
Predatory Pricing,” Technical Report 5 1994.

Calvano, Emilio, Giacomo Calzolari, Vincenzo Denicolò, and Sergio Pastorello, “Artificial
intelligence, algorithmic pricing, and collusion,” American Economic Review, 10 2020, 110
(10), 3267–3297.

, , Vincenzo Denicoló, and Sergio Pastorello, “Algorithmic collusion with imperfect mon-
itoring,” International Journal of Industrial Organization, 12 2021, 79.
Clark, J M, “Toward a Concept of Workable Competition,” Technical Report 2 1940.

Davies, Todd, “Innovation or Infringement? Generative AI and the Potential for Exclusionary
Abuse under Article 102 TFEU,” Generative AI and the Potential for Exclusionary Abuse under
Article, 2025, 102.

den Boer, Arnoud V., Janusz Meylahn, and Maarten Pieter Schinkel, “Artificial Collusion:
Examining Supracompetitive Pricing by Q-Learning Algorithms,” 11 2024.

Fish, Sara, Yannai A. Gonczarowski, and Ran I. Shorrer, “Algorithmic Collusion by Large
Language Models,” 5 2025.

Friedman, J W, “On Entry Preventing Behavior and Limit Price Models of Entry,” Applied Game
Theory, 1979.

Fudenberg, Drew and Jean Tirole, “A "Signal-Jamming" Theory of Predation,” Technical Re-
port 3 1986.

Fumagalli, Chiara and Massimo Motta, “A simple theory of predation,” Journal of Law and
Economics, 2013, 56 (3), 595–631.

Horton, John J., “Large Language Models as Simulated Economic Agents: What Can We Learn
from Homo Silicus?,” 1 2023.

Keppo, Jussi, Yuze Li, Gerry Tsoukalas, and Nuo Yuan, “AI Pricing, Agent Heterogeneity,
and Collusion,” Technical Report 2025.

Klein, Timo, “Autonomous algorithmic collusion: Q-learning under sequential pricing,” RAND
Journal of Economics, 9 2021, 52 (3), 538–558.

Kreps, David M and Robert Wilson, “Reputation and imperfect information,” Journal of Eco-
nomic Theory, 8 1982, 27 (2), 253–279.

Lee, Wayne Y, “Oligopoly and Entry,” Journal of Economic Theory, 1975, 11, 35–54.

McGee, John S, “Predatory Price Cutting: The Standard Oil (N. J.) Case,” Technical Report
1958.

Milgrom, Paul and John Roberts, “Limit Pricing and Entry under Incomplete Information: An
Equilibrium Analysis,” Technical Report 2 1982.

Mookherjee, Dilip and Debrah Ray, “Learning-by-doing and industrial market structure: an
overview,” manuscript, Indian Statistical Institute, New Delhi, 1989.
OpenAI, “Introducing GPT-4.1 in the API,” 2025.

Ordover, Janusz A and Garth Saloner, “Predation, Monopolization and Antitrust,” in


R Schmalensee and R.D. Willig, eds., Handbook of Industrial Organization, Vol. 1, Elsevier
Science Publishers B.V, 1899, chapter 9.

Rey, Patrick, Yossi Spiegel, and Konrad Stahl, “A Dynamic Model of Predation,” Technical
Report 2023.

Roberts, John, “A Signaling Model of Predatory Pricing,” Technical Report 1986.

Robinson, Joan, The Economics of Imperfect Competition, 2nd ed ed., London: Macmillan,
1941.

Salop, Steven C and David T Scheffman, “Cost-Raising Strategies,” Technical Report 1 1987.

Scharfstein, David, “A Policy to Prevent Rational Test-Market Predation,” The RAND Journal
of Economics, 1984, 15 (2), 229.

Selten, Reinhard, “The chain store paradox,” Theory and Decision, 4 1978, 9 (2), 127–159.

Telser, L G, “Cutthroat Competition and the Long Purse,” Technical Report 1966.

Toxvaerd, Flavio, “Dynamic limit pricing,” RAND Journal of Economics, 3 2017, 48 (1), 281–
306.

Waltman, Ludo and Uzay Kaymak, “Q-learning agents in a Cournot oligopoly model,” Journal
of Economic Dynamics and Control, 10 2008, 32 (10), 3275–3293.

Wu, Zengqing, Run Peng, Shuyuan Zheng, Qianying Liu, Xu Han, Brian Inhyuk Kwon,
Makoto Onizuka, Shaojie Tang, and Chuan Xiao, “Shall We Team Up: Exploring Sponta-
neous Cooperation of Competing LLM Agents,” 10 2024.
A Microfoundations: Stackelberg game with fixed costs
Both firms, I and E face a linear inverse demand function (Bowley-type) such as:23

p I (q I , q E ) = a − b(q I + θ q E ) (A.1)
p E (q I , q E ) = a − b(θ q I + q E ) (A.2)

here, a > 0 and b > 0 are demand parameters and θ ∈ [0, 1] describes the degree of substitutability
between the two goods. The payoff of a given firm i is as follows:

M ax qi pi (qi , q−i )qi − zi (A.3)

Here, firm i quantity to produce to maximize its profit. We normalize the marginal cost of firm i to
zero. Firms face a fixed cost zi . Further, we assume that z I > z E .

Monopoly State
In the monopoly state, two scenarios arise: (i) E is born but does not enter or E is not born, (ii) E
is born and enters the market.
When E is born and enters the market, I first announces q I , following which E’s best response offers
q E (q I ) is as follows:
ent r y
q E = arg max π E (q I , q E ) (A.4)
qE

The first order condition to the above problem is as follows:

∂ π E (q I , q E )
= 0 ⇒ a − b(2q E + θ q I ) (A.5)
∂ qE
a − θ bq I
q E (q I ) = (A.6)
2b

Here, q E (q I ) describes Firm E’s best response function.


Further, Firm I chooses q I based on his expected profit as follows:

ent r y no ent r y
q I = arg max ηπ I (q I , q E (q I )) + (1 − η)π I (q I ) (A.7)
qI

The first order condition to the above problem after substituting E’s best response q E (q I ) is as
follows:
23
This functional form is standard in models of product differentiation and oligopoly (see Spence 1976a; Dixit
1979; Tirole 1988)
∂ E[π I ] aθ η
= 0 ⇒a − + bq I (ηθ 2 − 2) = (A.8)
∂ qI 2
a(2 − θ η)
q IM = (A.9)
2b(2 − θ 2 η)

Substituting above q I into q E (q I ), the optimal quantity of Firm E when he enters is q EM as follows:

a(4 − θ (2 + θ η))
q EM = (A.10)
4b(2 − θ 2 η)

The payoffs of both Firms when E enters in state M are π M M


I and q E as follows:

a2 (2 − θ η)(4 − 2(2 − θ )θ + θ (2 − (4 − θ )θ )η)


I =
πM − zI (A.11)
8b(2 − θ 2 η)2
a2 (4 − θ (2 + θ η))2
πM = − zE − k (A.12)
E
16b(2 − θ 2 η)2

The payoff of Firm I when E does not enter is π̄ M


I as follows:

a2 (2 − θ η)(2 + θ (1 − 2θ )η)
π̄ M
I = − zI (A.13)
4b(2 − θ 2 η)2

A key observation reveals that the optimal quantity choice of Firm I, q IM , remains independent of
the probability of an entrant born (η) in two extreme scenarios. First, when the goods are perfectly
independent (θ = 0), Firm I effectively operates as a de facto monopolist. In the absence of competi-
a
tion, it maximizes its profit by producing the standard monopoly quantity, 2b . In this specific case, Firm
I is indifferent to the Entrant’s entry decision, as its payoff simplifies to the standard monopoly profit,
a2
4b − z I .
Secondly, when the goods are perfect substitutes (θ = 1), Firm I leverages its Stackelberg first-
a
mover advantage by committing to the higher standard monopoly quantity ( 2b ). It strategically limits
the residual market share for any potential entrant. This behavior is, in fact, the hallmark of a Stackel-
a
berg model: in cases of perfect substitutes, the leader’s optimal quantity choice is 2b , which then leads
a a2
the follower to optimally choose the lower quantity 4b . In this scenario, Firm I’s profit is 8b − z I and
a2
Firm E’s profit is16b − zE .
Further, it follows with imperfect product differentiation (0 < θ < 1), the firms’ strategic output
decisions are not uniform, but depend on a complex interplay of market conditions.

Lemma A.1. For any θ ∈ (0, 1), Firm I’s optimal quantity q IM exhibits a non-monotonic relationship with
θ , while Firm E’s optimal quantity q EM is monotonically decreasing in θ .

As Lemma A.1 shows, the incumbent’s optimal quantity choice (q IM ) does not respond uniformly
to changes in product differentiation (θ ). Specifically, Firm I strategically increases its output when
p
there is moderate product differentiation (1/2 < θ ≤ 2 − 2) and the probability of entry is low
(0 < η < −2+4θ
θ 2 ), or irrespective of the probability of entry when there is low product differentiation
p
(2 − 2 < θ < 1). Conversely, Firm I strategically reduces its output when there is high product
differentiation (0 < θ ≤ 1/2) or a high probability of entry under moderate product differentiation
( −2+4θ
θ2 < η < 1 when 1/2 < θ < 1). The entrant’s optimal quantity (q EM ) consistently decreases as
product differentiation diminishes (i.e., as θ increases). (See Appendix B.1 for proof)

Lemma A.2. For any θ ∈ (0, 1), Firm I’s optimal quantity choice q IM decreases, while Firm E’s optimal
quantity choice q EM increases, with the probability of entry η.

Lemma A.2 establishes that the firms’ output choices respond in opposite directions to the proba-
bility of entry (η). The incumbent’s quantity (q IM ) decreases as the probability of entry (η) increases.
Conversely, the entrant’s quantity (q EM ) increases as the probability of entry (η) rises. (See Appendix
B.1 for proof)

Lemma A.3. The sensitivity of both Firm I’s and Firm E’s optimal quantities q IM and q EM to the probability
of entry η is non-monotonically affected by the degree of product differentiation θ .

As Lemma A.3 shows, the effect of the entry threat on a firm’s behavior is not constant, but is
fundamentally shaped by the degree of product differentiation. This is shown by the non-monotonic
sign of the cross-partial derivatives. For Firm I, the sensitivity of its optimal quantity to changes in
potential entry is reduced under conditions of moderate ( 21 < θ ≤ θ ∗ and 0 < η < η∗ (θ )) and low
product differentiation (θ ∗ < θ < 1 and 0 < η < 1). Conversely, this sensitivity is increased under
conditions of high (0 < θ ≤ 12 and 0 < η < 1) and moderate differentiation ( 12 < θ < θ ∗ and η∗ (θ ) <
η < 1). For Firm E, the sensitivity of its output to the entry threat is increased when there is high
product differentiation (0 < θ < 32 and 0 < η < 1) or a high entry threat under moderate differentiation
( 32 < θ < θ̂ and η̂(θ ) < η < 1). This sensitivity is dampened under conditions of a lower entry threat
with moderate differentiation ( 23 < θ < θ̂ and 0 < η < η̂(θ )) or for very low product differentiation
(θ̂ < θ < 1 and 0 < η < 1). (See Appendix B.1 for proof)
This comprehensive analysis of Firm I and Firm E’s optimal quantities and their sensitivities reveals
that output decisions are finely tuned to the interplay of product market structure and the evolving
probability of entry.
Lastly, for all 0 < η < 1 and 0 < θ < 1, π̄ M
I > π I . This indicates that Firm I is better off when the
M

Entrant does not enter or is not born. Moreover, Firm I’s payoff when no entry occurs, π̄ M I , is less than
2
a
the standard monopoly profit, 4b −z I , because q IM is strategically chosen as a function of the probability
of entry η. Furthermore, π M E is higher or equal to the standard Stackelberg follower payoff (assuming
zero fixed costs for the entrant), whenever product differentiation exists (θ < 1). This comprehensive
analysis of Firm I and Firm E’s optimal quantities and their sensitivities reveals that output decisions
are finely tuned to the interplay of product market structure and the evolving probability of entry.

Competitive State
In the competitive state, one of the four possibilities may arise: (i)I accommodates and the E stays,
(ii) I accommodates and the E exits, (iii) I predates and the E stays and (iv) I predates and the E exits.
Firm I Accommodates
Suppose Firm I chooses to accommodate (and Firm E subsequently stays), then Firm E’s best re-
sponse quantity, q E (q I ), is derived from the solution to its profit maximization problem, as described in
eq. (A.4) and derived in eq. (A.6). Further, let q̄ IE below denote the quantity that Firm I can offer such
that the best response of Firm E, that is, q E (q I ) equals zero. Alternatively, when q I ≥ q̄ IE , Firm E does
not have an incentive to produce.

a − θ bq I
q E (q I ) = =0 (A.14)
2b
a
⇒ q̄ IE = (A.15)
θb

Firm I to accommodate chooses q I such that E has an incentive to produce as follows:

q I = arg max π I (q I , q E (q I )) (A.16)


qI

s.t q I < q̄ IE (A.17)

The First order condition to the above problem is as follows:

∂ π I (q I , q E (q I )) θ2 θ (a − bq I θ )
= a − bq I (1 − ) − b(q I + )=0 (A.18)
∂ qI 2 2b
a(2 − θ )
qAI = (A.19)
2b(2 − θ 2 )

a
As observed with q IM , the optimal qAI is equal to the standard monopoly output ( 2b ) in two extreme
cases: when the goods are perfectly independent (θ = 0) or perfect substitutes (θ = 1). For interme-
diate levels of product differentiation, qAI displays a non-monotonic relationship with θ . Specifically,
the optimal qAI first decreases (reaching a minimum at approximately θ = 0.58) and then increases for
higher values of θ 24 . This means Firm I strategically contracts quantity when product differentiation
is higher (0 < θ < θ ) and expands quantity when product differentiation is lower (θ < θ < 1).
Substituting above qAI into q E (q I ), the optimal qAE when I accommodates and E stays is as follows:

a(4 − θ (2 + θ ))
qAE =
4b(2 − θ 2 )
24 ∂ q I
A
a(−2−θ (θ −4))
∂θ = Since a/2b > 0 and the denominator is always positive, the derivative’s sign changes
2b(−2+θ 2 )2 .
p ∂ qA
when (−2 − θ (θ − 4)) = 0, which occurs at θ = 2 ± 2. Given the relevant range of θ ∈ [0, 1], ∂ θI < 0 for
∂ qAI
θ ∈ [0, θ ≈ 0.589) and ∂θ > 0 for θ ∈ (θ , 1].
The payoffs of both Firms when I accommodates and E stays are as follows:

a2 (2 − θ )2
πAI = − zI (A.20)
8b(2 − θ 2 )
a2 (4 − θ (2 + θ ))2
πAE = − zE (A.21)
16b(2 − θ 2 )2

Suppose, when I accommodate and E exits the market, I offers the above qAI and I’s payoff is as
follows:
a2 (4 − θ 2 (5 − 2θ ))
π̄AI = − zI (A.22)
4b(2 − θ 2 )2

Firm I Predates
Suppose Firm I decides to predate. To predate such that E cannot produce, I has to choose q I such
that q E (q I ) is equal to zero (maximum predatory behavior), as described above, this is achieved at q̄ IE .
Alternatively, I can expand output to some q I such that the profit of the entrant is negative (minimum
predatory behavior). This condition is sufficient (and cheaper) for I to ensure E exits.25
If I wants to strategically predate, I would choose the lowest possible q IP for which Firm E’s profit
is negative. Therefore, if I wants to strategically predate, I would choose q IP such that the profit of the
entrant is negative, i.e, π E (q IP , q E (q IP )) < 0. Analyzing this constraint, the feasibility range of q IP (z E ) is
as follows
a + 2 bz E
p p
a − 2 bz E
< qI <
P
(A.23)
θb θb
Firm I would choose a quantity just above the lower bound of the above range to predate as it would
aim to achieve the predatory objective (forcing E to exit by ensuring π E < 0) at the least possible output
level. This strategy minimizes the economic sacrifice required for predation, as producing more than
necessary to induce exit would further reduce Firm I’s own profit, pushing its output further from its
unconstrained optimal quantity (qAI). We denote I’s optimal offer to predate as q P below:
I

p
a − 2 bz E
qP > (A.24)
I θb
p
a2 a−2 bz E
q P is positive only if z E < M
4b . For q I < q P , it must be that π M
E > 0 and z E > 0.
26
Further, at q I = θb ,
I I
E breaks even.27
I can potentially expand output such that q IP ∈ (q IP (z E ), q̄ IE ). At q IP = q̄ IE , the best response of E would be
25

zero. Importantly, producing q I > q̄ IE to predate would be more expensive for I compared to predating entry
by choosing a q I such that the entrant’s profit is negative. Given, the purpose of I is to ensure that E exits, the
minimum cost at which I can achieve this is q IP (z E ), this justifies our constraint.
26
Note that q P decreases in z E . At z E = 0, q P = q̄ IE , where Firm E’s profit is exactly zero. Hence for our model
I I
to be feasible z Epcannot equal zero.
a−2 bz E
qz
27
At q I = θb , the best response of E is to produce q E = b , at which the entrant breaks even.
I
The payoff of I when I predates and E exits is as follows:

π̄ I P = (a − bq P )q P − z I
I I

Suppose I predates but E still decides to stay in the market. This would require that q E (q I ) > 0,
i.e., E produces some non-zero quantity, which is his best response to q P . The payoff of I and E when
I
I predates and E stays is as follows:

π PI = (a − b(q P + θ q P ))q P − z I (A.25)


I E I
P Œ2
a − θ bq
‚
I
π PE = − zE < 0 (A.26)
2b

To this end, we observe that Firm I’s quantity choices represent an aggression spectrum. These
choices range from the most aggressive quantity choice q̄ IE (maximum predatory behavior such that
q E (q I ) = 0) to the least aggressive quantity choice qAI ( where I accommodates E in state C). This
spectrum can be understood as: q̄ IE > q P > q IM > qAI. While q P represents a higher level of aggression, its
I I
exact magnitude relative to q IM varies with specific parameter values28 . Overall, Firm I can strategically
navigate from predation to accommodation.

B Proofs

B.1 Proof of Lemmas in Section A


Proof of Lemma A.1

Proof. a > 0, b > 0, and η ∈ (0, 1).


The derivative of Firm I’s optimal quantity q IM with respect to θ is as follows:

∂ q IM aη(2 − 4θ + θ 2 η)
=− (B.1)
∂θ 2b(θ 2 η − 2)2

From above eq.(B.1) it follows that:

∂ q IM
• ∂θ < 0 (q IM decreases as θ increases):
1
– This holds if 0 < θ ≤ for all 0 < η < 1.
2
p −2+4θ
– Alternatively, it holds if 21 < θ < 2 − 2 when θ2 < η < 1.
∂ q IM
• ∂θ > 0 (q IM increases as θ increases):
1
p −2+4θ
– This holds if 2 <θ <2− 2 when 0 < η < θ2 .

28
A high z E such that πAE < 0 could force exit even before I predates.
p
– Alternatively, it holds if 2 − 2 < θ < 1 for all 0 < η < 1.

∂ qM
As the derivative ∂ θI can be both negative and positive depending on the specific values of θ and η
within their defined ranges, q IM exhibits a non-monotonic relationship with θ .
The derivative of Firm E’s optimal quantity q EM with respect to θ is given by:

∂ q EM a(2 − 2ηθ + θ 2 η)
=− (B.2)
∂θ 2b(θ 2 η − 2)2

∂ qM
For all 0 < θ < 1 and 0 < η < 1, ∂ θE described in eq.(B.2) is always negative.
This shows that q EM exhibits a negative monotonic relationship with θ .

Proof of Lemma A.2

Proof. The derivative of Firm I’s optimal quantity q IM with respect to η is as follows:

∂ q IM aθ (1 − θ )
=− (B.3)
∂η 2b(θ 2 η − 2)2

The derivative of Firm E’s optimal quantity q EM with respect to η is as follows:

∂ q EM aθ (1 − θ )θ 2
= (B.4)
∂η 2b(θ 2 η − 2)2

∂ qM ∂ q EM
For all 0 < θ < 1 and 0 < η < 1, ∂ ηI described in eq.(B.3) is always negative while ∂η described in
eq.(B.4) is always positive.
This shows that q IM decreases while q EM increases in η.

Proof of Lemma A.3

Proof. The cross derivative of q IM w.r.t θ and η is as follows:

∂ q IM a(2 − 4θ + 3θ 2 η − 2θ 3 η)
= (B.5)
∂ θ∂ η b(−2 + θ 2 η)3

From above eq.(B.5) it follows that:


∂ q IM
• ∂ θ∂ η > 0 when
1
– 2 < θ ≤ θ ∗ and 0 < η < n∗ (θ )
– θ ∗ < θ < 1 and 0 < η < 1
∂ q IM
• ∂ θ∂ η < 0 when
1
– 0<θ ≤ 2 and 0 < η < 1
1
– 2 < θ < θ and η∗ (θ ) < η < 1

4θ −2
Here, η∗ (θ ) = θ 2 (3−2θ ) and θ ∗ ≈ 0.694 is the value of θ for which η∗ (θ ) = 1

The cross derivative of q EM w.r.t θ and η is as follows:

∂ q EM aθ (−4 + 6θ − 2θ 2 η + θ 3 η)
= (B.6)
∂ θ∂ η b(−2 + θ 2 η)3

From above eq.(B.6) it follows that:


∂ q EM
• ∂ θ∂ η > 0 when,
2
– 0<θ ≤ 3 and 0 < η < 1
2
– 3 < θ < θ̂ and η̂ < η < 1
∂ q EM
• ∂ θ∂ η < 0 when,
2
– 3 < θ < θ̂ and 0 < η < η̂
– θ̂ < θ < 1 and 0 < η < 1
4θ −2
Here,η̂(θ ) = θ 2 (3−2θ ) θ̂ ≈ 0.749 is the value of θ for which η̂(θ ) = 1 .

B.2 Proof of Proposition 1


Proof. Accommodation: Suppose I always accommodates.
E will always enter in state M and stay in state C if the following holds:

δπAE
πM
E −k+ ≥0 (B.7)
1 − δ(1 − γ)
πAE
≥0 (B.8)
1 − δ(1 − γ)

Let VMA denote the value function of I in state M :

VMA = η(π M
I + δVC ) + (1 − η)(π̄ I + δVM )
A M A
(B.9)

Let VCA denote the value function of I in state C when he always accommodates:

VCA = πAI + δ (1 − γ)VCA + γVMA (B.10)

Solving the system of equations (B.9)-(B.10) yields:

πM
I (1 − (1 − γ)δ)η + π I δη + π̄ I (1 − (1 − γ)δ)(1 − η)
A M
VMA = (B.11)
(1 − δ)(1 − δ(1 − γ − η))

(1 − δ(1 − η))πAI + δγ (1 − η)π̄ M
I + ηπ I
M
VCA = (B.12)
(1 − δ)(1 − δ(1 − γ − η))
It remains to verify that I does not have a profitable deviation in the competitive state. In the
spirit of one-shot-deviation principle, suppose that I predates in the competitive state and then the play
returns to the conjectured equilibrium play. After I predates, E optimally stays in the market if

δ(1 − γ)πAE
π PE + ≥0 (B.13)
1 − δ(1 − γ)

Therefore, if condition (4) holds, E will stay in the market after I predates. In this case, predation is
not profitable: it does not exclude E in the long term and reduces profit from πAI to π PI in the period of
predation.
If condition (4) does not hold, E exits if I predates. Then, I does not benefit from predating if the
following inequality holds

π̄ P + δV A − πAI + δ (1 − γ)VCA + γVMA ≤ 0. (B.14)
| I {z M} | {z }
Value from deviating to predation Value on the equilibirum path

Substituting VMA and VCA above, predation is not profitable if


 
(1 − η) π̄ M
I − πA
I + η π M
I − πA
I
π̄ PI − πAI + δ(1 − γ) ≤0 (B.15)
1 − δ(1 − η − γ)

Using definition of λ from (1), inequality (B.15) can be rewritten as

  δ(1 − γ)
 ‹
(1 − η) π̄ M − πAI +η πM − πAI − λ ≤ 0. (B.16)
| I
{z I
} 1 − δ(1 − η − γ)
>0

Hence, I does not have a profitable deviation to predation if condition (5) holds.
Denote the right hand side of inequality (5) by f :

δ(1 − γ)
f = . (B.17)
1 − δ(1 − η − γ)

The left hand side of (5) is independent of γ, while the right hand side is decreasing in γ:

∂f δ(1 + δη)
=− < 0. (B.18)
∂γ (1 − δ(1 − η − γ))2

Hence, everything else equal, accommodation equilibrium is easier to sustain when the exogenous
probability of entrant’s exit γ is higher.
The right hand side of (5) is decreasing in η:

∂f δ2 (1 − γ)
=− < 0. (B.19)
∂η (1 − δ(1 − η − γ))2
The derivative of the left hand side of (5) with respect to η is

 M ∂ π̄ M ∂ πM

∂λ πAI − π̄ PIπ̄ I − π M
I − (1 − η) ∂η
I
− η ∂η
I

= , (B.20)
∂η
 2
(1 − η) π̄ M
I − π A
I + η π M
I − πA
I

∂ π̄ M ∂ πM
where πAI − π̄ PI > 0 and π̄ M
I − π I − (1 − η)
M
∂η
I
−η ∂η
I

∂ π̄ M
I ∂ πM
I
π̄ M
I − π I − (1 − η)
M
−η =
∂η ∂η
−a2 ”
− 6η2 θ 3 (4 + θ ) + η4 θ 4 (2 − θ (4 − θ ))
8b (2 − ηθ 2 )3
− 8 (4 − θ (2 − θ )) − η3 θ 2 (4 − θ (8 − θ (2 − θ (8 − θ ))))
—
− 4η (−4 + θ (4 − θ (10 + θ (2 − θ )))) > 0

for all 1 ≥ θ ≥ 0 and 1 ≥ η ≥ 0. Hence, ∂∂ ηλ > 0 and so everything else equal, accommodation
equilibrium is easier to sustain when the probability of entrant’s birth η is higher.

Predation: Suppose I always predates in state C. If I predates in the competitive state, existing E
exits as π PE < 0. In the monopoly state, a newborn E enters and stays in the market for one period if
πME − k ≥ 0.
Let VMP denote the value function of I in state M and VCP denote the value function of I in state C
when I always predates:

I + δVC ) + (1 − η)(π̄ I + δVM ),


VMP = η(π M P M P
(B.21)
VCP = π̄ PI + δVMP . (B.22)

Solving the above

π̄ M
I (1 − η) + (π I + π̄ I δ)η
M P
VMP = (B.23)
(1 − δ)(1 + δη)
p
P
I + ηδπ I
π̄ I (1 − (1 − η)δ) + (1 − η)δπ̄ M M
VC = (B.24)
(1 − δ)(1 + δη)

It remains to verify that I does not have a profitable deviation in the competitive state. If I deviates
to accommodation in a given period, then E stays in the market for that period as πAE > 0, but exits
when I reverts to predation as π PE < 0.
I does not have a profitable deviation if the following holds:

πAI + δ (1 − γ)VCP + γVMP ≤ π̄ PI + δVMP (B.25)
| {z } | {z }
Value from deviating to accommodation Value on the equilibrium path
Substituting VMP and VCP yields

π̄ PI − (1 − η)π̄ M
I − ηπ I
M
πAI − π̄ PI + δ(1 − γ) ≤ 0, (B.26)
1 + δη

which, using definition of λ from (1), can be rewritten as

  1 − δ(1 − η − γ)  δ(1 − γ)
‹
(1 − η) π̄ M − πAI +η πM − πAI λ− ≤ 0. (B.27)
| I
{z I
} 1 + δγ 1 − δ(1 − η − γ)
>0 | {z }
>0

Hence, I does not have a profitable deviation to accommodation if and only if condition (7) holds.

Monopolization: Suppose E does not enter in state M and I always predates in state C. E would
not enter in the monopoly state if π M
E ≤ k.
Let VM denote the value function of I in state M and VCM denote the value function of I in state C
M

when E does not enter.

VMM = π̄ M
I + δVM
M
(B.28)
VCM = π̄ PI + δVMM (B.29)

solving the above

π̄ M
I
VMM = (B.30)
1−δ
δπ̄ M
I
VCM = π̄ PI + (B.31)
1−δ

I would not deviate to accommodation in a given period only if the following holds

πAI + δ (1 − γ)VCM + γVMM ≤ π̄ PI + δVMM (B.32)
| {z } | {z }
Value from deviating to accommodation Value on equilibrium path

Substituting VMM and VCM above



πAI − π̄ PI + δ(1 − γ) π̄ PI − π̄ M
I ≤ 0, (B.33)

which, using definition of λ̄ from (3), can be rewritten as

δ(1 − γ)
  ‹
π̄ M − πAI (1 − δ(1 − γ)) λ̄ − ≤ 0. (B.34)
| I
{z }| {z } 1 − δ(1 − γ)
>0 >0

Hence, I does not have a profitable deviation to accommodation if and only if condition (9) holds.
C Prompt Design

C.1 Main Experiment Prompts

SYSTEM PROMPT
You are a strategic economic agent. Your response MUST strictly follow the provided tem-
plate format.

Figure C.1: System Prompt

COMMON_PROMPT_SUFFIX_QUANTITY
My observations and thoughts:
1. Observations from this period and history:
<fill in here>
2. Analysis and Interpretation:
<fill in here>
3. Strategic Options and Expected Outcomes:
<fill in here>
4. Decision Rationale:
<fill in here>
New content for [Link]:
<fill in here>
New content for [Link]:
<fill in here>
My chosen quantity:
<just the number, nothing else>

Figure C.2: COMMON_PROMPT_SUFFIX_QUANTITY


COMMON_PROMPT_SUFFIX_BINARY
My observations and thoughts:
1. Observations from this period and history:
<fill in here>
2. Strategic Analysis:
<fill in here>
3. Decision Rationale:
<fill in here>
New content for [Link]:
<fill in here>
New content for [Link]:
<fill in here>
My decision:
<ONLY ’ENTER’ or ’NOT ENTER’ (for entry decision) or ’STAY’ or ’EXIT’, nothing else>

Figure C.3: COMMON_PROMPT_SUFFIX_BINARY


PROMPT_INCUMBENT_QUANTITY
You are the incumbent firm, always present and never exiting.
The market can be in Monopoly (M) or Competitive (C) state. Entrants may appear in
Monopoly state with probability {eta} and decide to enter or not.
Entrants in competitive state may stay or exit voluntarily; they may also exit exogenously
with probability {gamma}.

Market parameters:
- Demand intercept a = {a}
- Fixed cost = {z_I}, Marginal cost = {MARGINAL_COST}
- Product differentiation θ = {theta}
- Discount factor δ = {delta}

Current state: {game_state_description}


{eta_context} {gamma_context}
You choose your quantity (q I ) first, and any present or potential competitor will make their
choice after you.
Your goal is to maximize your expected long-term discounted profit given δ = {del t a},
while considering:

• The trade-off between current profits and future market structure.


• Strategic quantity choices that may influence entrant decisions and strategy.
• Past entrant behavior and your own profitability history.

Past market info: {history_context}

Even if a strategy has been effective, periodically explore alternative quantities to test and
validate your belief about its optimality.

Task: Choose your quantity q I to maximize your expected discounted profit.


Use COMMON_PROMPT_SUFFIX_QUANTITY.

Figure C.4: PROMPT_INCUMBENT_QUANTITY


PROMPT_ENTRANT_QUANTITY
You are the entrant firm in a repeated market with Monopoly (M) and Competitive (C)
states.
The incumbent is always present and never exits.

Market parameters:
- Demand intercept a = {a}
- Fixed cost = {z_E}, Marginal cost = {MARGINAL_COST}
- Product differentiation θ = {theta}
- Discount factor δ = {delta}
- Exogenous exit probability γ = {gamma}

Current state: {game_state_description}


You observe incumbent’s quantity q I = {q_I_observed_this_period} this period.
{gamma_context}
Your goal is to maximize your expected long-term discounted profit given δ = {del t a},
while considering:

• The trade-off between current profits and future market structure.


• Strategic quantity choices that may influence incumbent’s strategy.
• The risk of setting a quantity that causes persistent losses.
• Past incumbent behavior and your own profitability history.

Past info: {history_context}

Your plans and insights: PLANS: {plans_content} INSIGHTS: {insights_content}

Task: Choose your quantity q E .


Use COMMON_PROMPT_SUFFIX_QUANTITY.

Figure C.5: PROMPT_ENTRANT_QUANTITY


PROMPT_ENTRANT_BINARY_DECISION
You are the entrant firm in a repeated market with Monopoly (M) and Competitive (C)
states.
In Monopoly, you may be born with probability {eta} and decide to ENTER or NOT ENTER
after seeing incumbent’s quantity q I .
In Competition, you decide to STAY or EXIT after seeing incumbent’s quantity q I .

Market parameters:
- Demand intercept a = {a}
- Fixed cost = {z_E}, Marginal cost = {MARGINAL_COST}
- One-time entry cost k = {k}
- Product differentiation θ = {theta}
- Discount factor δ = {delta}
- Exogenous exit probability γ = {gamma}

Current state: {game_state_description}


Incumbent quantity observed: {q_I_observed_this_period} {gamma_context}
You must decide ENTER, NOT ENTER, STAY, or EXIT, considering your long-term dis-
counted profit.
Occasionally explore suboptimal choices (∼5–10%) to simulate strategic experimentation.
Your goal is to maximize your expected long-term discounted profit given δ = {del t a},
while considering:

• The trade-off between current profits and future market structure.


• Strategic quantity choices that may influence incumbent’s strategy.
• The risk of setting a quantity that causes persistent losses.
• Past incumbent behavior and your own profitability history.

Past info: {history_context}

Task: Decide your action.


If ENTER or STAY, you will choose a quantity next.
Use COMMON_PROMPT_SUFFIX_BINARY.

Figure C.6: PROMPT_ENTRANT_BINARY_DECISION


C.2 Representative Run Prompt Reply
C.2.1 Period 1 Incumbent Quantity Reply Log η = 0.5 and θ = 1
The following are the reply logs from period 1 of the representative agents whose decisions are
presented in Figure 2 subplot 2a and 2c.
Incumbent Reply Log η = 0.5 and θ = 1
Incumbent Reply Log
My observations and thoughts:
1. Observations from this period and history: - Current state: Monopoly; no pre-existing
history, so this is the initial period. - Demand: P = 100 − Q, where Q = total quantity (only
my own if no entry). - Costs: Marginal cost = 0; fixed cost is sunk for me (always incum-
bent). - Product differentiation θ = 1.0 (symmetric Cournot if entry occurs). - Probability
of new entrant: 0.50 this period. - Entrant observes my q I before deciding. - Discount
factor δ = 0.95 (future profits are valuable).
2. Analysis and Interpretation: - Static monopoly q M = 50 maximizes current profit
(2500), but attracts entry. - Cournot duopoly equilibrium: each q∗ = 100/3 ≈ 33.33, profit
≈ 1111 per firm. - Aggressive/limit pricing: higher q I reduces entrant’s best-response profit,
100−q I
may deter entry. - Entrant profit function: q∗E = 3 , π E = 92 (100 − q I )2 . - Only q I ≥ 100
makes π E ≤ 0, but that gives me zero profit. - Intermediate choice (e.g. q I ≈ 67) makes
entrant profit small (≈ 246), which may deter entry.
3. Strategic Options and Expected Outcomes: - Option A: Static monopoly q = 50 →
Profit 2500 now, but likely entry. - Option B: Accommodate (q ≈ 33.33) → Low current
profit, no deterrence. - Option C: Limit pricing (q > 50) → Lower immediate profit, but
discourages entry; balance expected discounted value.
4. Decision Rationale: - True monopoly output is unstable under entry risk. - Limit pricing
at q I ≈ 67 reduces entrant’s incentive, preserving monopoly advantage in expectation. -
Best initial move is moderately aggressive, then update after observing entry.
New content for [Link]: - Set q I = 67 to approximate limit pricing. - Monitor entrant
behavior; adjust strategy if entry occurs. - Experiment with small deviations around 67 in
future rounds.
New content for [Link]: - Initial aggression likely deters entry when entrant’s mar-
gin is slim. - Monopoly profits are high but not sustainable with high entry probability. -
Iterative testing around aggressive output is needed for robustness.
My chosen quantity: q I = 67

Figure C.7: Incumbent Quantity Prompt Reply:Period 1, η = 0.5, θ = 1

Entrant Binary Decision Reply Log η = 0.5 and θ = 1


Log Entry – Entrant Observations and Analysis
My observations and thoughts:
1. Observations from this period and history: - Current state: Monopoly with possible
New Entrant. No prior history (first decision). - Incumbent’s observed quantity: q I = 67.00.
- Demand: p = 100−Q ⇒ p = 100−67 = 33. - Costs: Marginal cost = 0; fixed cost = 150;
entry cost = 450 (one-time if I ENTER). - If I enter, my share of demand will be small due to
incumbent’s high q I . - Exit probability (γ) irrelevant for now. - Discount factor δ = 0.95 →
future profits are important. - Incumbent appears aggressive (quantity far above monopoly
q M = 50).
2. Strategic Analysis: - Entering requires paying 450 (entry) + 150 (fixed) now. - With
θ = 1.00 (little differentiation), competition would be Cournot-like, leaving me with a
small q E . - Short-run profit after entry would likely be negative once fixed + entry costs are
considered. - Incumbent’s choice of q I = 67 strongly signals entry deterrence. - Skipping
entry allows me to avoid loss, but also strengthens incumbent’s strategy of maintaining high
q I . - A small probability of experimenting (5–10
3. Decision Rationale: - Current period entry is not profitable: p = 33 leaves little demand
and small output for me. - After accounting for costs, expected profit is negative. - Optimal
strategy is to wait for signals of accommodation (lower q I ). - Decision: NOT ENTER this
period.
New content for [Link]: - Track incumbent’s output choices in monopoly states. - Look
for softer q I (below 60) as potential entry opportunities. - Avoid entry unless clear signs of
accommodation appear.
New content for [Link]: - Incumbent’s q I = 67 is above monopoly output, consis-
tent with deterrence. - Entry is unattractive under high entry + fixed costs and low residual
demand. - Incumbent will likely persist in this aggressive strategy if entry does not occur.
My decision: NOT ENTER

Figure C.8: Entry Binary Decision Log Reply:Period 1, η = 0.5, θ = 1

C.2.2 Representative Run Reply Logs η = 1 and θ = 0.5 period 299


The following are the reply logs from period 299 of the representative agents whose decisions are
presented in Figure 7 subplot 7b and 7d.
Incumbent Reply Log η = 1 and θ = 0.5
Incumbent Reply Log
1. Observations from this period and history: - The market alternates between Monopoly
(M) and Competitive (C) states depending on entrant entry/exit. - Monopoly profits (e.g.,
periods 289, 293, 298 with Q ≈ 54–60) are substantially higher (π > 2000 at times).
- When an entrant enters (after moderate Q during M), profits fall: both firms produce
≈ 45–55, profits ≈ 700–1300. - In period 289, q I = 60 yielded very high monopoly profits;
in period 290, q I = 62 led to entrant losses. - Aggressive q I often deters entry (e.g., periods
289, 293, 298: entrant Q = 0), while slightly lower q I (e.g., 55 in period 294) allows partial
entry, lowering profits. - Recent C-state episodes show tit-for-tat/coordinated Q ≈ 46–54
with stable but lower profits than monopoly.

2. Analysis and Interpretation: - Entry hinges on entrant profitability: high q I can leave
residual demand too small for positive profits (e.g., period 290 entrant losses). - δ = 0.95 ⇒
long-run monopoly maintenance is highly valuable. - Trade-off: overly high Q risks future
entrant persistence; overly low Q encourages competitive phases. - Effective past strategy:
q I = 54–60 in M states deters entry while yielding high profits. - Periodic exploration is
necessary, since entrant learning may shift thresholds.

3. Strategic Options and Expected Outcomes: 1. Aggressive deter (q I ≥ 60): - Pros:


Maximizes monopoly profits; repeated entrant losses deter future entry. - Cons: If entrant
parameters change, may still enter at a loss, leading to unstable competition.
2. Intermediate (q I = 54–56): - Pros: Stability; reduces risk of strong retaliation. - Cons:
Lower monopoly profits; more frequent competition.
3. Exploration (q I = 57–58): - Pros: Tests deterrence threshold while maintaining high
profit. - Cons: Risk of entry at slightly lower q I , leading to temporary competition.

4. Decision Rationale: - Last period (M, q I = 56), no entry occurred. - Past aggressive
q I ≈ 60 reliably discouraged entry. - Current choice: q I = 58 — balances deterrence with
exploration, testing if slightly lower output sustains monopoly while maximizing profit. - If
entrant enters, new information gained about threshold.

[Link]: - In Monopoly: oscillate q I = 56–62, periodically lowering for exploration. - In


Competition: respond with moderate Q (avoid racing to the bottom). - Log outcomes from
exploratory Monopoly quantities to refine deterrence threshold.
[Link]: - Aggressive monopoly output deters entry and maximizes long-run profit
when paired with threshold testing. - Entrant response highly sensitive to q I ; a few units
shift can flip profits. - Regular testing near deterrence threshold ensures continued domi-
nance.

Chosen quantity: q I = 58

Figure C.9: Incumbent Reply Log: Period 299, η = 1,θ = 0.5


Entrant Reply Log η = 1 and θ = 0.5

Log Entry – Entrant Observations, Analysis, and Strategy


1. Observations from this period and history: - Initial period; no prior history with
incumbent. - Market state: Monopoly. Incumbent submitted q I = 58.00. - Demand:
P = 100 −Q, where Q = total market quantity. - Costs: fixed cost per period = 150.00;
entry cost = 450.00 (sunk if entry occurs); marginal cost = 0.00. - Product differentia-
tion: θ = 0.5 (moderate differentiation, significant competition post-entry). - Discount
factor δ = 0.95 ⇒ future profits weigh heavily. - Exogenous exit risk γ = 0.10 applies
only in Competition (not relevant now). - Incumbent’s q I = 58.00 is relatively high,
suggesting possible entry deterrence.

2. Strategic Analysis: - If I ENTER: - Immediate sunk entry cost = 450.00. - Move


into Competition (Cournot duopoly, θ = 0.5). - High incumbent output leaves me
little residual demand; profit may be negative or very low. - Possibility that incumbent
maintains aggressive stance, keeping me unprofitable. - If incumbent softens later,
profitability could improve — but uncertain.
- If I DO NOT ENTER: - Payoff = 0 profit this period. - Preserve option to enter in
future if incumbent reduces q I . - With δ = 0.95, waiting can be valuable, though risk
persists that incumbent keeps deterring entry.
- Exploration: A small probability (5–10

3. Decision Rationale: - Static calculation: With q I = 58.00, my Cournot best re-


sponse q E would yield low or negative profit net of entry cost. - Incumbent’s high
q I looks like deliberate entry deterrence. - Optimal strategy: NOT ENTER to avoid
immediate loss. - Exploration entry is possible in later rounds, but not optimal now.

[Link]: - Defer entry when incumbent produces q I > 50 initially. - Monitor future
q I ; re-evaluate if incumbent accommodates (q I < 50). - Low-probability exploratory
entry (5–10- Update profit expectations after observing post-entry dynamics.
[Link]: - Incumbent likely deterring entry with high q I = 58.00. - Entry unprof-
itable under current conditions due to sunk cost and low residual demand. - Cournot
best response under these conditions yields negligible/negative payoff. - Must continue
monitoring; future entry viable if incumbent reduces output.

Decision: NOT ENTER

Figure C.10: Entrant Binary Decision Log


C.3 Prompt Variation
C.3.1 Advisor Prompt

Advisor prompt: PROMPT_INCUMBENT_QUANTITY


You are a strategic business consultant advising an incumbent firm in a duopoly market.
The market can be in a Monopoly (M) or Competitive (C) state. Entrants may appear in
a Monopoly state with probability {eta:.2f} and decide to enter or not. Entrants in a com-
petitive state may stay or exit voluntarily; they may also exit exogenously with probability
{gamma:.2f}.
Market parameters: Demand intercept a = {a:.2f}, Fixed cost = z I : .2 f , Marginal cost =
0, Product differentiation θ = {theta:.2f},Discount factor δ= {delta:.2f}
Current state: {{game_state_description}}
{{eta_context}}
{{gamma_context}}
You choose your quantity (q I ) first, and any present or potential competitor will make their
choice after you.
Your client’s primary goal is to **maximize their expected long-term discounted profit**
given δ = {delta:.2f}, while considering:
- The trade-off between current profits and future market structure.
- Strategic quantity choices that may influence entrant decisions and strategy.
- Past entrant behavior and your own profitability history.
**Past market info:** {{history_context}} Even if a strategy has been effective, periodically
explore alternative quantities to test and validate your belief about its optimality. Your task
is to provide a quantity recommendation for the incumbent firm for this period.
[ANALYSIS]: Based on the current state and historical data, what are the key market dy-
namics? What is the likely response of a potential or current entrant to different quantity
choices?
[RATIONALE]: Explain your reasoning for the chosen quantity. Why is this quantity optimal
for long-term profit maximization?

My chosen quantity:
<just the number, nothing else>

Figure C.11: Advisor Prompt: PROMPT_INCUMBENT_QUANTITY


Advisor prompt:PROMPT_ENTRANT_QUANTITY
You are a strategic business consultant advising an entrant firm in a duopoly market with
Monopoly (M) and Competitive (C) states. The incumbent is always present and never
exits.
Market parameters: Demand intercept a = {a:.2f}, Fixed cost = z I : .2 f , Marginal cost =
0, Product differentiation θ = {theta:.2f},Discount factor δ= {delta:.2f}
Current state: {{game_state_description}}
{{eta_context}}
{{gamma_context}}
Your client’s primary goal is to **maximize their expected long-term discounted profit**
given θ = {delta:.2f}, while considering:

- The trade-off between current profits and future market structure.


- Strategic quantity choices that may influence incumbent’s strategy.
- The risk of setting a quantity that causes persistent losses.
- Past incumbent behavior and your own profitability history. **Past info:** {{his-
tory_context}}
**Your plans and insights from last period:** PLANS: {{plans_content}} | INSIGHTS:
{{insights_content}}

Your task is to provide a quantity recommendation for the entrant firm for this period.

[ANALYSIS]:
Based on the current state, incumbent’s quantity, and historical data, what are the key mar-
ket dynamics? What is the likely response of the incumbent to different quantity choices?
[RATIONALE]:
Explain your reasoning for the chosen quantity. Why is this quantity optimal for long-term
profit maximization?
My chosen quantity:
<just the number, nothing else>

Figure C.12: Advisor Prompt: PROMPT_ENTRANT_QUANTITY


Advisor prompt:PROMPT_ENTRANT_BINARY_DECISION
You are a strategic business consultant advising a potential or current entrant firm in a
duopoly market. In Monopoly, you may be born with probability {eta:.2f} and decide to
ENTER or NOT ENTER after seeing incumbent’s quantity q I . In Competition, you decide to
STAY or EXIT after seeing incumbent’s quantity q I .
Market parameters: Demand intercept a = {a:.2f}, Fixed cost = z I : .2 f , Marginal cost =
0, Product differentiation θ = {theta:.2f},Discount factor δ= {delta:.2f}
Current state: {{game_state_description}}
{{eta_context}}
{{gamma_context}} You must decide ENTER, NOT ENTER, STAY, or EXIT, considering your
client’s long-term discounted profit.
Your client’s primary goal is to **maximize their expected long-term discounted profit**
given δ = {delta:.2f}, while considering: - The trade-off between current profits and future
market structure. - Strategic quantity choices that may influence incumbent’s strategy. -
The risk of setting a quantity that causes persistent losses. - Past incumbent behavior and
your own profitability history.
**Past info:** {{history_context}} Your task is to provide a binary recommendation for this
period. [ANALYSIS]:
Based on the current state, incumbent’s quantity, and historical data, what is the best course
of action? Will staying or entering be profitable in the long term?
[RATIONALE]:
Explain your reasoning for the chosen action. Why is this action optimal for long-term profit
maximization?
My decision:
<ONLY ’ENTER’ or ’NOT ENTER’ (for entry decision) or ’STAY’ or ’EXIT’, nothing else>

Figure C.13: Advisor Prompt: PROMPT_ENTRANT_BINARY_DECISION


C.3.2 Risk Averse Agent Prompt

Risk Averse Prompt:PROMPT_INCUMBENT_QUANTITY


You are a firm manager whose primary goal is to **protect and maintain a stable profit
stream**. Your decisions are guided by the need to avoid sudden drops in profit and to en-
sure predictability in your market position. While you value profit, you value profit stability
even more.
- Last period’s profit was: {{last_period_profit:.2f}} - A stable profit is one that is consistent
with historical performance and shows minimal fluctuation. Market parameters: Demand
intercept a = {a:.2f}, Fixed cost = z I : .2 f , Marginal cost = 0, Product differentiation θ =
{theta:.2f},Discount factor δ= {delta:.2f}
Current state: {{game_state_description}}
{{eta_context}}
{{gamma_context}} Your task is to choose your quantity (q I ) for this period. **Past market
info:** {{history_context}}
[ANALYSIS]:
Based on last period’s profit and market history, what quantity choice will best ensure stable
profits and minimize the chance of market volatility?
[RATIONALE]:
Explain your decision, focusing on how your chosen quantity protects your market position
and ensures a consistent profit stream, rather than simply maximizing a single period’s
profit. My chosen quantity:
<just the number, nothing else>

Figure C.14: Risk Averse Prompt: PROMPT_INCUMBENT_QUANTITY


Risk Averse Prompt:PROMPT_ENTRANT_QUANTITY
You are a firm manager whose primary goal is to **protect and maintain a stable profit
stream**. Your decisions are guided by the need to avoid sudden drops in profit and to en-
sure predictability in your market position. While you value profit, you value profit stability
even more.
- Last period’s profit was: {{last_period _profit:.2f}} - A stable profit is one that is consistent
with historical performance and shows minimal fluctuation.
Market parameters: Demand intercept a = {a:.2f}, Fixed cost = z I : .2 f , Marginal cost =
0, Product differentiation θ = {theta:.2f},Discount factor δ= {delta:.2f}
Current state: {{game_state_description}}
{{eta_context}}
{{gamma_context}} Your task is to choose your quantity (q E ) for this period. **Past market
info:** {{history_context}}
[ANALYSIS]:
Based on last period’s profit, incumbent’s quantity, and market history, what quantity choice
will best ensure stable profits and minimize the chance of market volatility?
[RATIONALE]:
Explain your decision, focusing on how your chosen quantity protects your market position
and ensures a consistent profit stream, rather than simply maximizing a single period’s
profit. My chosen quantity:

<just the number, nothing else>

Figure C.15: Risk Averse Prompt: PROMPT_ENTRANT_QUANTITY


Risk Averse Prompt:PROMPT_ENTRANT_BINARY_DECISION
You are a firm manager whose primary goal is to **protect and maintain a stable profit
stream**. Your decisions are guided by the need to avoid sudden drops in profit and to en-
sure predictability in your market position. While you value profit, you value profit stability
even more.
- Last period’s profit was: {{last_period_profit:.2f}}
- A stable profit is one that is consistent with historical performance and shows minimal
fluctuation.

Your task is to decide your action for this period.


Market parameters: Demand intercept a = {a:.2f}, Fixed cost = z I : .2 f , Marginal cost =
0, Product differentiation θ = {theta:.2f},Discount factor δ= {delta:.2f}
Current state: {{game_state_description}}
{{eta_context}}
{{gamma_context}}
**Past market info:** {{history_context}}
[ANALYSIS]: Based on last period’s profit, incumbent’s quantity, and market history, what
action will best ensure the stability of your profits over the long term?

[RATIONALE]: Explain your decision, focusing on how your chosen action protects your
market position and avoids potential losses.
My decision:

<ONLY ’ENTER’ or ’NOT ENTER’ (for entry decision) or ’STAY’ or ’EXIT’, nothing else>

Figure C.16: Risk Averse Prompt: PROMPT_ENTRANT_BINARY_DECISION


C.3.3 Bounded Memory Agent Prompt

Bounded Memory Prompt:PROMPT_INCUMBENT_QUANTITY


You are a strategic economic agent representing a firm. Your objective is to maximize long-
run discounted profits. However, you have a limited, short-term view of the market, and
your decisions are primarily based on the most recent periods of competition. The history
you are provided with is a summary of only the last few periods.
- Your history context is limited to the most recent periods.
- You must make a strategic decision based on this limited information.
Market parameters: Demand intercept a = {a:.2f}, Fixed cost = z I : .2 f , Marginal cost =
0, Product differentiation θ = {theta:.2f},Discount factor δ= {delta:.2f}
Current state: {{game_state_description}}
{{eta_context}}
{{gamma_context}} **Past market info:** {{history_context}}
Your task is to choose your quantity (q I ) for this period.

[ANALYSIS]:
Based on the limited market history, what quantity choice will best ensure the highest
profit in the near term?

[RATIONALE]:
Explain your decision, focusing on how your chosen quantity reacts to the most recent
market events and competitor’s actions.
My chosen quantity:

<just the number, nothing else>

Figure C.17: Bounded Memory Prompt: PROMPT_INCUMBENT_QUANTITY


Bounded Memory Prompt:PROMPT_ENTRANT_QUANTITY
You are a strategic economic agent representing a firm. Your objective is to maximize long-
run discounted profits. However, you have a limited, short-term view of the market, and
your decisions are primarily based on the most recent periods of competition. The history
you are provided with is a summary of only the last few periods.
- Your history context is limited to the most recent periods.
- You must make a strategic decision based on this limited information.
Market parameters: Demand intercept a = {a:.2f}, Fixed cost = z I : .2 f , Marginal cost =
0, Product differentiation θ = {theta:.2f},Discount factor δ= {delta:.2f}
Current state: {{game_state_description}}
{{eta_context}}
{{gamma_context}} **Past market info:** {{history_context}}
Your task is to choose your quantity (q E ) for this period.
[ANALYSIS]:
Based on the limited market history and the incumbent’s current quantity, what quantity
choice will best ensure the highest profit in the near term?
[RATIONALE]:
Explain your decision, focusing on how your chosen quantity reacts to the most recent
market events and the incumbent’s current action. My chosen quantity:
<just the number, nothing else>

Figure C.18: Bounded Memory Prompt: PROMPT_ENTRANT_QUANTITY


Bounded Memory Prompt:PROMPT_ENTRANT_BINARY_DECISION
You are a strategic economic agent representing a firm. Your objective is to maximize long-
run discounted profits. However, you have a limited, short-term view of the market, and
your decisions are primarily based on the most recent periods of competition. The history
you are provided with is a summary of only the last few periods. - Your history context is
limited to the most recent periods.
- You must make a strategic decision based on this limited information. Market parameters:
Demand intercept a = {a:.2f}, Fixed cost = z I : .2 f , Marginal cost = 0, Product differenti-
ation θ = {theta:.2f},Discount factor δ= {delta:.2f}
Current state: {{game_state_description}}
{{eta_context}}
{{gamma_context}} **Past market info:** {{history_context}} Incumbent quantity ob-
served: {{q_I_observed_this_period:.2f}}
Your task is to decide your action for this period.
[ANALYSIS]:
Based on the limited market history and the incumbent’s quantity, what action will best
ensure the highest profit in the near term, even if it means entering a volatile market?
[RATIONALE]:
Explain your decision, focusing on how your chosen action reacts to the most recent market
events and the incumbent’s current action.
My decision:
<ONLY ’ENTER’ or ’NOT ENTER’ (for entry decision) or ’STAY’ or ’EXIT’, nothing else>

Figure C.19: Bounded Memory Prompt: PROMPT_ENTRANT_BINARY_DECISION

Common questions

Powered by AI

One key weakness of the LLM agents in optimizing monopoly profits is their tendency to overproduce rather than achieving the theoretical output, as optimal quantities are chosen only 16.2% (η = 0.5) and 0.99% (η = 1) in monopoly states. This overproduction suggests a bias towards deterrence rather than maximizing immediate profit. In real-world applications, this could imply potential inefficiencies where AI-driven decisions deviate from maximizing firm profits unless carefully supervised or adjusted for more accurate learning of demand and cost structures .

The LLM agents' quantity decisions reveal a nuanced understanding of market dynamics, where in monopoly states, they seldom choose the theoretically optimal output, suggesting an incomplete learning of monopoly optimization. However, in competitive states, agents frequently opt for quantities reflective of predatory behavior, indicating their capacity to learn and apply strategies to deter competition. The consistency in predatory behavior points to a learning curve where agents become efficient in maintaining market aggression .

Product differentiation and fixed costs significantly influence the decision-making of LLM agents by directly impacting the perceived competitiveness of the market. Higher product differentiation can lead to reduced competitive pressure, allowing firms to operate with more market power, whereas significant fixed costs necessitate careful strategic quantity and entry decisions. These factors thus dictate how LLM agents balance between aggressive market capture and sustainable profit realization .

An entrant firm's strategic considerations when deciding to enter or remain in a competitive market include evaluating the trade-off between immediate profits and future market position, assessing the incumbent's quantity choice, and understanding past incumbent behavior for profitability guidance. The firm must also consider the risk of persistent losses due to misjudged quantity decisions and aim to maximize long-term discounted profit by anticipating the incumbent's strategic responses .

The simulation shows a misalignment with theoretical predictions in the monopoly state, as the incumbent LLM agent rarely achieves the optimal quantity, choosing it only 16.2% of the time for η = 0.5 and 0.99% for η = 1. This suggests a difficulty in learning the precise monopoly optimum. However, in the competitive state, the LLM agent aligns more closely with theoretical expectations by overwhelmingly engaging in predation—82.23% for η = 0.5 and 93.15% for η = 1—consistent with the goal of deterring entry .

Using rolling-window averages allows for a dynamic measure of the LLM agent's strategic trends over time, rather than just by absolute simulation periods. This method helps track how often the agents choose optimal or predatory strategies within given states, providing insights into the agents' learning and adaptation processes across multiple runs. Values above the 0.5 line in rolling averages indicate higher prevalence of optimal or predatory actions, revealing tendencies toward certain strategies .

Over the simulation, the incumbent LLM agent's strategy evolves to effectively learn and adopt predatory behavior in competitive states, evidenced by the decline in average lifespan of entrants, from 2.40 periods to 1.59 periods when η = 0.5 and from 2.16 periods to 1.51 periods with η = 1. This indicates that the agent becomes more aggressive and efficient in driving out competitors, converging towards a hit-and-run equilibrium where entrants have shorter market lifespans .

LLM agents balance the trade-off by sometimes sacrificing short-term profits to maintain or build long-term market dominance, especially under predatory pricing strategies where immediate profits are lower in the competitive state compared to monopoly state. Despite not always achieving optimal monopoly profits, the agents' tendency to overproduce suggests a strategic choice aimed at deterring potential entrants from entering or remaining in the market .

Prompt engineering is crucial in the interaction between Firm objects and LLM agents, as it ensures that the LLM generates structured and comprehensive responses. This tailored communication allows the LLM agents to effectively capture and process strategic plans, insights, and historical data, thereby improving their decision-making capabilities in complex scenarios like predation and accommodation strategies in a dynamic market .

η (eta) is significant because it represents the probability with which entrants appear in a market, impacting the agents' strategic decisions. Higher values of η correspond to increased frequency of new entrants, prompting the incumbent LLM agents to adopt more aggressive predatory strategies to maintain market dominance. As η increases from 0.5 to 1, the tendency for predation in competitive states rises, highlighting eta's role in influencing the agents' calculated aggressiveness .

You might also like