Market-based Architectures in RL and Beyond
Blue Sky Ideas Track
Abhimanyu Pallavi Sudhir Long Tran-Thanh
University of Warwick, UK University of Warwick, UK
[Link]-sudhir@[Link] [Link]-thanh@[Link]
ABSTRACT “Market-based architectures” can be made concrete in the case
Market-based agents refer to reinforcement learning agents which of reinforcement learning (RL), where the majority of work in this
determine their actions based on an internal market of sub-agents. area lies (see 1.1 for a brief summary). For example in the Hayek
We introduce a new type of market-based algorithm where the machine [2] and its derivatives, the setting is a Markov Decision
arXiv:2503.05828v1 [[Link]] 5 Mar 2025
state itself is factored into several axes called “goods”, which allows Problem (MDP) and there is a single resource, the “right to act and
for greater specialization and parallelism than existing market- collect reward”, that is traded between sub-agents. At each time
based RL algorithms. Furthermore, we argue that market-based step, this resource is sold to the highest-bidding sub-agent, who
algorithms have the potential to address many current challenges performs some action that modifies the state and collects reward.
in AI, such as search, dynamic scaling and complete feedback, and In this paper, we argue that market-based agents represent an
demonstrate that they may be seen to generalize neural networks; underexplored and promising niche, especially in context of recent
finally, we list some novel ways that market algorithms may be advancements in language models (LLMs), and have potential to
applied in conjunction with Large Language Models for immediate address a range of present challenges in contemporary AI research.
practical applicability. Specifically, we make the following claims and contributions:
Theoretical framework for market-based agents. We present
KEYWORDS two general frameworks for market-based RL agents: (1) the “deep
market” (Def 2.1), where a single good, the state, is passed through a
markets; prediction markets; alignment; RL
sequence of transacting agents, and (2) the “wide market” (Def 2.2),
ACM Reference Format: in which the state space itself is partitioned into factors called goods.
Abhimanyu Pallavi Sudhir and Long Tran-Thanh. 2025. Market-based Ar- The deep framework is not much of a departure from existing al-
chitectures in RL and Beyond: Blue Sky Ideas Track. In Proc. of the 24th gorithms, and can be applied to any Partially Observed Markov
International Conference on Autonomous Agents and Multiagent Systems (AA-
Decision Process (POMDP); to our knowledge the wide framework
MAS 2025), Detroit, Michigan, USA, May 19 – 23, 2025, IFAAMAS, 5 pages.
is original to us, and is a generalization of the deep framework
which better mirrors the success of real-world markets allowing for
1 INTRODUCTION greater specialization and parallelism. A Python library for creating
Before neural networks won the mandate of heaven, an AI research market-based algorithms will be released upon publication.
paradigm that had shown considerable promise was that of market- Markets, neural networks and backpropagation. We demon-
based architectures, i.e. AI agents that determine their output based strate that these market-based agents can in principle be applied
on some internal market-based mechanism [3, 18]. even to basic supervised learning tasks such as classification, and
There are general intuitive arguments that motivate such a line that neural networks (though not backpropagation or gradient
of research. Philosophers and psychologists have long pondered descent) emerge as a special case of them. Furthermore we general-
multi-agent models of the mind [23]; moreover, one may imagine ize the result in [33] to wide markets, demonstrating a suggestive
that “any” machine learning task could in principle be solved by relationship between backpropagation and markets at equilibrium.
a market of agents solving sub-tasks with their individual reward Search, complete feedback and alignment. We claim that
set by the sale value of their output. There is work suggesting that markets can address several present problems in AI research, specif-
markets can capture some notion of bounded rationality, e.g. the ically: they are a natural framework for search, their scale or depth
Boundedly Rational Inductive Agent (BRIA) [25] and Algorithmic can be dynamic rather than fixed, and they allow complete feedback
Bayesian Epistemology [24]. In some sense, markets “aggregate” the [8], a property widely regarded as valuable in AI alignment.
intelligence or capacities of their individual participants.1 Markets and LLMs. We present novel ways in which market
algorithms might be applied in conjunction with LLMs to address
1 See e.g. Hayek on the role of markets in aggregating information [32] – or the famous
their limitations: they can be used for developing reasoning models
parable “I, Pencil” [19]: “... no one person, no matter how smart, could create from
scratch a small, everyday pencil [yet the market makes over a billion of them each like o1, and LLMs can facilitate “information markets” that can in
year] ...”. There is also some empirical work on emergent intelligent behaviour in turn improve human feedback mechanisms in AI training.
markets comprised of zero-intelligence traders [11, 17, 31]
This work is licensed under a Creative Commons Attribution Inter-
national 4.0 License.
1.1 Related Work
Market-based RL. The majority of early work in this area has fo-
Proc. of the 24th International Conference on Autonomous Agents and Multiagent Systems cused on market-based reinforcement learning (RL) algorithms. The
(AAMAS 2025), Y. Vorobeychik, S. Das, A. Nowé (eds.), May 19 – 23, 2025, Detroit, Michigan,
USA. © 2025 International Foundation for Autonomous Agents and Multiagent Systems pioneering work in this domain consists of Holland’s Learning Clas-
([Link]). sifier Systems or “bucket brigade” [15] in rule-based systems, where
condition-action agents (“classifiers”) bid to post messages (which Algorithm 1 Deep market
can be actions described in some language) onto a global message procedure Market ⊲ Forward pass
board. Improvements to this paradigm were made by Schmidhuber parameters: 𝑤 [𝛼] ∈ R ⊲ wealths of each 𝛼 ∈ A
[29, 30] who allowed agents to determine their own bids and im- input: 𝜔 ∈ Ω
posed credit conservation, and by Baum [2] who further strictly 𝑏 [𝛼] ← min(𝛼𝑏 (𝜔), 𝑤 [𝛼]) for 𝛼 ∈ A ⊲ Cap bids by wealth
enforced property rights, resulting in a much more familiar set-up 𝛼 ∗ ← arg max 𝑏 [𝛼] ⊲ Choose winning agent
called the “Hayek Machine”. The Hayek machine, which can be 𝑥 ← 𝛼ˆ ∗ (𝜔) ⊲ Determine action
applied to any Markov Decision Process (MDP), was subsequently return 𝑥, 𝛼 ∗
extended to POMDPs by [18] by adding external memory, and more end procedure
recently [3] modified the framework to use Vickrey auctions and procedure Capitalism ⊲ Training loop
prove that a Nash equilibrium of the market produces a globally Initialize agent wealths 𝑤 [𝛼] ∈ R for each 𝛼 ∈ A
optimal policy. Initialize original owner of the world 𝛼 ∗
Bounded rationality and markets. Other, more recent work Initialize state 𝑠 ∈ S
in this area includes: Logical Induction [9], an algorithm that as- while 𝑡 ∈ N do
signs probabilities to mathematical sentences based on their prices 𝜔 ← 𝜔 (𝑠) ⊲ Generate observation
in a prediction market that (roughly speaking) pays off when a ∗
𝛼 prev ← 𝛼∗
sentence is proven; and the Boundedly Rational Inductive Agent 𝑥, 𝛼 ∗ ← Market(𝜔)
(BRIA) [25] which solves finite decision problems by assigning it 𝑤 [𝛼 ∗ ] ← 𝑤 [𝛼 ∗ ] − 𝑏 [𝛼 ∗ ] ⊲ Pay bid
to the highest-bidding trader, similar to in market-based RL. The ∗ ] ← 𝑤 [𝛼 ∗ ] + 𝑏 [𝛼 ∗ ]
𝑤 [𝛼 prev ⊲ to previous owner
prev
key insight of these works is that markets are useful for modeling 𝑠 prev ← 𝑠
boundedly rational agents. The main results of each work – the fact 𝑠 ← 𝑥 (𝑠) ⊲ Transition state
that the logical inductor cannot by dominated by any polynomial- 𝑤 [𝛼 ∗ ] ← 𝑤 [𝛼 ∗ ] + R(𝑠 prev, 𝑥, 𝑠) ⊲ add reward to wealth
time trader, and the “boundedly rational inductive agent criterion” end while
in the latter paper – are specific and precise formulations of the end procedure
Efficient Market Hypothesis [24].
Markets and neural networks. A specific equivalence between
classifier systems and neural networks has been studied in the Neu-
ral bucket brigade [7, 30], although this does not consider backpropa- Definition 2.1 (Deep market). Assume a POMDP setup, and let
gation. A suggestive analogy between markets and backpropagation A be a collection of “agents”, which are (stochastic) maps 𝛼 : Ω →
is discussed in [33], though only for the case of a strictly sequential, X × R. The first component 𝛼ˆ : Ω → X of an agent is called its
unit-width market like that in Def 2.1. In our work we make this action, the second component 𝛼𝑏 : Ω → R is called its bid. The
more precise and generalize it to 2.2 market algorithm then proceeds as in Algorithm 1.2
Some details have been ignored. A will usually be infinite and so
2 MARKET ALGORITHMS tables like 𝑤 [𝛼] cannot simply be indexed on it: instead, A must be
Setting (POMDP). We assume a ususal POMDP setting, with a countably enumerated and added to the economy one-by-one in the
state space S, action space X, transition probability P(𝑠 ′ | 𝑠, 𝑥) training loop with each agent being endowed with some allowance.
(which allows us to treat actions as stochastic functions i.e. 𝑥 (𝑠) ∼ To prevent holdout problems, one may impose a small fixed “rent”
P(𝑠 ′ | 𝑠, 𝑥)), reward function R(𝑠, 𝑥, 𝑠 ′ ), and observation distribution on 𝛼 ∗ at each training step i.e. 𝑤 [𝛼 ∗ ] ← (1 − 𝜀)𝑤 [𝛼 ∗ ]. A simple
𝜔 (𝑠) ∼ O(𝜔 | 𝑠) over a set of observations Ω. A policy is a map first-price auction is shown for simplicity, and may be replaced
𝛼 : Ω → X, and the process proceeds as per usual. with a Vickrey auction in line with [3]. One may also replace the
explicit reward function R with a class of “consumers” C who place
Mirroring [18] and similar to Belief-MDP formulations, we can
bids upon desirable states, which may be a useful formulation for
extend the state and message spaces by taking the cartesian product
reinforcement learning from diverse human feedback3 .
space with a message space str which is always preserved by 𝜔;
Definition 2.1, which subsumes existing market-based RL, al-
this gives the policy a “memory”, or in terms of markets, creates
ready illustrates one of the key defining features of markets rec-
informational goods.
ognizable to any student of economics: markets serve not only to
The first algorithm we describe is Def 2.1: here, agents bid at each
select (via market competition) the best process to achieve a task,
time step for the right to act and collect reward, and the highest-
but also to distribute a complex task among agents which are in-
bidding agent is chosen to act. As in previous work, e.g. [2, 3],
dividually much too weak or uninformed to complete the entire
these agents are not utility-maximizers but programs out of a pos-
task. This means that the collection of actions A can be a class
sibly infinite collection of agents A (which we leave abstract). The
of “simple” agents, so that enumerating A can quickly find many
parameters of this algorithm are the wealths of each agent 𝑤 [𝛼].
valuable agents.
trained by the training loop Capitalism: at each step, the agent
pays its bid to the previous agent, collects the reward generated
2 For training, this may be executed in multiple episodes with different initial state 𝑠 ,
by its actions and receives the bid of the next. This means that at
either with finite episodes or in parallel (with shared wealth variables across running
equilibrium, each agent is incentivized to bid the value function, instances).
and perform the action with maximum Q-value. 3 see [6, 10] for a primer on this area
Specifically, Def 2.1 exploits modularity of action, where the Algorithm 2 Wide market
state can be transformed one step at a time. There is however procedure Market ⊲ Forward pass
another form of modularity, missed by all existing market-based RL parameters: 𝑤 [𝛼] ∈ R ⊲ wealths of each 𝛼 ∈ A
algorithms, which we may call modularity of state, and is crucial to input: 𝜔 ∈ Ω
the success of real-world markets: here, agents are not constantly ⊲ Cap bids by budget
transacting the whole “state of the world”: instead, the state of 𝑏 [𝛼] ← 𝜆𝑠 : min(𝛼𝑏 (𝑠), 𝑤 [𝛼]) for all 𝛼 ∈ A
the world is decomposed into several components, called goods4 : ⊲ Compute equilibrium prices and allocations
S = S1 ⊕ · · · ⊕ S𝑛 . For instance, 𝑠 1 ∈ S1 might represent the p, 𝜔 ′ [𝛼 1 ], . . . 𝜔 ′ [𝛼 𝑛 ] ← Equ( 𝜔 [𝛼], 𝑏 [. . . ])
Í
quantity of iron ore in the world. Agents bid for small quantities of 𝑥 [𝛼] ← 𝛼ˆ (𝜔 ′ [𝛼]) for all 𝛼 ∈ A ⊲ Determine actions
each good; no agent owns the whole world, and does not have to return 𝑥 [𝛼 1 ], . . . 𝑥 [𝛼 𝑛 ], p, 𝜔 ′ [𝛼 1 ], . . . 𝜔 ′ [𝛼 𝑛 ]
bother performing a valuation of the whole world. The “state of the end procedure
world” may be recovered as the vector sum of all agents’ holdings. procedure Capitalism ⊲ Training loop
At least two new difficulties are introduced by considering mar- Initialize agent wealths 𝑤 [𝛼] ∈ R for each 𝛼 ∈ A
kets of multiple divisible goods: Initialize agent properties 𝜔 [𝛼] ∈ Ω for each 𝛼 ∈ A
General equilibrium theory. Allocating goods is no longer as Initialize state 𝑠 ∈ S
easy as an auction, because agents might have joint demand sched- while 𝑡 ∈ N do
ules for goods that are complementary or substitute to each other. 𝜔 ← 𝜔 (𝑠) ⊲ Generate observation
The problem of matching buyers and sellers in this setting is the · · · ← Market(𝜔) ⊲ get all outputs
domain of “General Equilibrium Theory” in economics, where there 𝑤 [𝛼] ← 𝑤 [𝛼] − p · 𝜔 ′ [𝛼] for all 𝛼 ⊲ Charge buyers
are models such as the Fisher market and the Arrow-Debreu ex- 𝑤 [𝛼] ← 𝑤 [𝛼] + p · 𝜔 [𝛼] for all 𝛼 ⊲ Pay sellers
change market [1, 20]. Computing the equilibrium in these models 𝑠 [𝛼] ← 𝑠𝑥 [𝛼 ] (𝜔 [𝛼]) for all 𝛼⊲ Calculate property rights
is non-trivial and often intractable [4, 5] 𝑠 ′ [𝛼] ← 𝑥 [𝛼] (𝑠 [𝛼]) for all 𝛼 ⊲ Transform goods
Property rights in POMDPs. A more subtle difficulty lies in 𝑤 [𝛼] ← 𝑤 [𝛼] + R(𝑠 [𝛼], 𝑥 [𝛼], 𝑠 ′ [𝛼]) for all 𝛼
the fact that we want to divide the state 𝑠 ∈ S, which is not directly 𝑠 ← 𝑠 ′ [𝛼]
Í
⊲ Update state
observed, among bidding agents (so that each of their actions only end while
transform their respective portions of the state, i.e. their properties), end procedure
but the agents only submit demand schedules over 𝜔 ∈ Ω. It is not
obvious how to map a decomposition of a vector 𝜔 (𝑠) back onto 𝑠.
Both of these have to do with specific questions of how buy-
ers and sellers meet and match in real markets, i.e. having to do 𝑠𝑥𝜔 2 (𝜔 2 )) (this is used to combine actions by different agents). The
with institutions such as property rights and mechanism design. agents now have dependent type signatures 𝛼 : (𝜔 : Ω) → X𝜔 × R,
These questions are out of scope for us, and we abstract them away and the transition probability 𝑥 (𝑠) ∼ P(𝑠 ′ | 𝑠, 𝑥), reward function
by postulating some effective equilibrium computation algorithm5 R(𝑠, 𝑥, 𝑠 ′ ) and observation distribution O(𝜔 | 𝑠) are now interpreted
Equ(𝜔, 𝛼𝑏1 , . . . 𝛼𝑏𝑚 ) = (p, 𝜔 [𝛼 1 ], . . . 𝜔 [𝛼 𝑚 ]) i.e. which takes the to- as applying to “private property”, i.e. to any goods bundle in their
tal perceived quantity of goods in the world Ω and each agent’s respective domains, rather than to the whole state, e.g. each action
valuation function 𝛼𝑏𝑖 : Ω → R, and returns a price vector p ∈ Ω 𝑥 (𝑠) defines a production function that transforms one goods bundle
and allocations to each agent 𝜔 [𝛼 𝑖 ] ∈ Ω, such that (in line with a into another, and 𝛼𝑏 is an agent’s valuation function over all possible
Walrasian equilibrium with quasilinear utilities [22]): bundles, i.e. how much it is willing to pay for a particular perceived
Í
• 𝜔 = 𝜔 [𝛼 𝑖 ] (the full quantity is allocated) bundle (if it’s differentiable, then ∇𝛼𝑏 (𝜔) can be interpreted as
• p · 𝜔 [𝛼 𝑖 ] ≤ 𝛼𝑏𝑖 (𝜔 [𝛼 𝑖 ]) for all 𝛼 𝑖 (no agent pays for what it the price vector it offers). The market algorithm proceeds as in
doesn’t value), and Algorithm 2 .
• 𝜔 [𝛼 𝑖 ] = arg max𝜔 ′ ∈Ω 𝛼𝑏𝑖 (𝜔 ′ ) − p · 𝜔 ′ for all 𝛼 𝑖 (each agent
Computing prices via backpropagation. Though it remains
gets a utility-maximizing bundle at the given price).
to be seen how standard RL problems might be cast in this setting,
Definition 2.2 (Wide market). Everything from the POMDP setup we expect implementations of this algorithm to be much more
and the agent type in Def 2.1 remains the same; except that S and effective than of Def 2.1, as it allows us to use simpler and more
Ω are now vector spaces with each vector called a goods bundle. specialized agents in the collection A. In particular, these agents
Further, we have action spaces X𝜔 indexed by 𝜔 ∈ Ω such that do not need to estimate the valuations of the whole world, but only
(1) for any 𝑥 ∈ X𝜔 , there is an “exercised property right” denoted of their particular input goods.
𝑠𝑥 (𝜔) ∈ S such that 𝜔 (𝑠𝑥 (𝜔)) = 𝜔 and 𝑥 (𝑠) = 𝑥 (𝑠𝑥 (𝜔)) + (𝑠 − This last point can be illustrated particularly nicely when the
𝑠𝑥 (𝜔)) (i.e. each agent’s actions transform only the goods they own) setup is an MDP, and rewards are replaced by consumers – here,
and (2) there is an injective map 𝜉 : X𝜔 1 × X𝜔 2 → X𝜔 1 +𝜔 2 such Ω = S and 𝜔 (𝑠) = 𝑠, so 𝛼ˆ : S → X can directly be interpreted as a
that 𝜉 (𝑥𝜔 1 , 𝑥𝜔 2 ) = 𝑥𝜔 1 (𝑠𝑥𝜔 1 (𝜔 1 )) +𝑥𝜔 2 (𝑠𝑥𝜔 2 (𝜔 2 )) + (𝑠 −𝑠𝑥𝜔 1 (𝜔 1 ) − production function 𝛼ˆ : S → S := 𝛼ˆ (𝑠)(𝑠). Then if the agent can
estimate what the market prices of its output goods will be (e.g.
4 ⊕ denotes the direct sum of vector spaces, which is a Cartesian product equipped
if prices are sufficiently stable that it makes sense to speak of a
with a pointwise vector addition operator
5 e.g. there are results demonstrating that simple tâtonnement converges to a Walrasian “prevailing price” p), then it can compute its offered prices via the
equilibrium when the agents’ valuations are gross subtitutes [13]. chain rule – where 𝐷 𝛼ˆ denotes the Jacobian:
4 PRACTICALITY AND FUTURE WORK
∇𝛼𝑏 = 𝐷 𝛼ˆ · p (1) We have presented two general frameworks for market-based RL
i.e. once the market “graph” is fixed, prices can be computed by agents, and illustrated that they may be seen to generalize neural
simply backpropagating consumer bids through the graph. This networks in a supervised learning setting, albeit with a more flexible
generalizes the result in [33], which demonstrated this relationship training mechanism that holds promise to address the limitations of
for deep markets only. current-day AIs with respect to reasoning and alignment properties.
Despite these theoretical strengths, our algorithm as described
3 MOTIVATION FOR MARKET-BASED AI faces practical challenges to implement in real-world machine learn-
In this section, we describe how markets could potentially general- ing tasks: blindly enumerating large classes of even simple agents
ize neural networks and provide a more “flexible training mecha- is inefficient (compared to backpropagation, where the search is
nism”. Although we have presented our algorithms in an RL setting, guided by gradients), and we have to store many more agents in
they can even be applied to supervised learning tasks by treating memory than the “size” of the network (the exact number depending
internal representations as “states”. To see this, it is illustrative to on the rule we use to prune low-wealth agents). Some potentially
see how a simple neural network can be recast as a market. promising approaches include:
• “integrated” models which perform backpropagation by de-
Theorem 3.1 (Neural networks as markets). Consider a fully-
fault but intelligently resort to markets when it expects
connected neural network 𝑓 : X → Y := 𝑓𝑛 ◦ . . . 𝑓1 where each
changing the network structure to be worthwhile
𝑓𝑖 : R𝑚𝑖 −1 → R𝑚𝑖 is a layer, i.e. a function of the form 𝑓𝑖 (x) =
• having each agent simultaneously learn its parameters via
𝜎 (𝑊𝑖 x + b𝑖 ) where 𝜎 is a ReLU activation. Then there is a deep market
backpropagation
whose forward pass performs the same operation as 𝑓 .
• decentralized set-ups, perhaps using frameworks such as
Proof. The
É construction is straightforward. Define the state BitTensor [28], allowing traders to be shared across machine
space S := 𝑚𝑖
1≤𝑖 ≤𝑛 S𝑖 ⊕ Y (with each S𝑖 := R ), with 𝜔 : S → learning applications.
Ω discarding only the last component Y which represents the true Markets of LLMs. A more immediately feasible application is
label which is unchanged under all actions. Each X𝜔 = {(𝑊𝑖 , b𝑖 ) : to develop markets comprised of LLMs, i.e. where A is a collection
𝑊𝑖 ∈ R𝑚𝑖 ×𝑚𝑖 −1 , b𝑖 ∈ R𝑚𝑖 } if 𝜔 ∈ S𝑖 −1 and empty if no such 𝑖 exists, of LLM agents. For instance, one may let S = Ω be a message space,
and an action 𝑥 = (𝑊𝑖 , i) acts on 𝑠 ∈ S𝑖 −𝑖 as 𝑥 (𝑠) = 𝜎 (𝑊𝑖 𝑠 +b𝑖 ). The and let actions act on 𝑠 by appending some “chain-of-thought item”
reward R(𝑠, 𝑥, 𝑠 ′ ) = −ℓ (𝑠 ′, 𝑠 Y
′ ) for some loss function ℓ if 𝑠 ′ ∈ S
𝑛 to the current message. The final reward is determined by human
and 0 otherwise. Finally, let A consist of all constant maps to X𝜔 feedback, and intermediate rewards by bids. Such a market would
and endow non-zero wealth to only those agents whose actions’ function as a “reasoning model” analogous to o1.
parameters are the same as some 𝑓𝑖 . □ The extension to a wide market is also immediate: agents may
While the market model, i.e. the forward pass, in Theorem 3.1 bid for the right to read only a portion of the message space6 –
is the same as the neural network, the training mechanism is Cap- this allows for more precise credit assignment to contributions by
italism (as defined in Algorithm 1) rather than backpropagation. different agents, and may be understood as to trees-of-thought [35]
Detailed below are some strengths of this we anticipate: what o1 is to chain-of-thought.
Search and dynamic scale. Reasoning is widely touted as a Theoretical work. The most pressing need at present is for
key limitation of current-day LLMs [16, 21]. A view held by some precise theoretical results on the effectiveness of market-based algo-
researchers including Yann LeCun [34], is that this is due to the rithms. An immediate research agenda includes the following:
fact that “[neural networks] produce their answers with a con- • Determining convergence and optimality conditions of
stant number of computational steps between input and output”, market algorithms; in particular, generalizing the “coverage”
independent of the complexity required by the problem. Some results of BRIA [25] and logical induction [9], i.e. demon-
proposed architectures that avoid this limitation include dynamic strating that the market will give a fair chance to the best
neural networks [14], adaptive computation time [12] as well as policy, conditional on some suitable wealth endowments.
chain-of-thought based methods such as o1 [26]. Markets provide • A Learning Theory perspective on markets and the wealth
a principled alternative, as here the structure of the computational update mechanism. In particular, (real-world) markets ap-
graph is itself learned, and different agents and structures may be pear to have many useful features from an alignment stand-
active for different inputs. point, such as their inherent capacity for online learning and
Complete feedback. Informally speaking, markets allow any generalization even from imperfect reward signals.
aspect of the system to be optimized. Formal results are needed to • A thorough translation of economic terminology into
make this statement precise, but intuitively: any aspect of a learner, our model – especially concepts like perfect competition,
such as any hyperparameter, or meta-learning, can be changed economies of scale, growth and welfare.
by adding a trader to the market who will profit if his changes Finally, to accelerate empirical work with market-based algo-
are beneficial and the incentives are correctly designed. This is rithms, we plan to release a Python library for efficiently creating
suggestive of the notion of “complete feedback” in AI alignment and applying market-based algorithms.
research, which refers to the property that “the trainer can enact 6 As for how to enable the agent to “inspect” the message to make an informed bid
any modification they’d like to make to the system” [8], and is without it stealing the entire message, [27] is relevant: the agent can subcontract
viewed as a desirable characteristic of an AI system for alignment. another LLM to inspect the message and place the bid, then have its context deleted.
REFERENCES [16] Jie Huang and Kevin Chen-Chuan Chang. 2023. Towards Reasoning in Large
[1] Kenneth J. Arrow and Gerard Debreu. 1954. Existence of an Equilibrium for a Language Models: A Survey. In Findings of the Association for Computational
Competitive Economy. Econometrica 22, 3 (1954), 265–290. [Link] Linguistics: ACL 2023, Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki
2307/1907353 arXiv:1907353 (Eds.). Association for Computational Linguistics, Toronto, Canada, 1049–1065.
[2] Eric B. Baum. 1999. Toward a Model of Intelligence as an Economy of [Link]
Agents. Machine Learning 35, 2 (May 1999), 155–185. [Link] [17] Karim Jamal, Michael S. Maier, and Shyam Sunder. 2015. Simple Agents, Intel-
1007593124513 ligent Markets. SSRN Electronic Journal (2015). [Link]
[3] Michael Chang, Sid Kaushik, S. Matthew Weinberg, Tom Griffiths, and Sergey 2478665
Levine. 2020. Decentralized Reinforcement Learning: Global Decision-Making via [18] Ivo Kwee, Marcus Hutter, and Juergen Schmidhuber. 2001. Market-Based Re-
Local Economic Transactions. In Proceedings of the 37th International Conference inforcement Learning in Partially Observable Worlds. In Proceedings of the In-
on Machine Learning. PMLR, 1437–1447. ternational Conference on Artificial Neural Networks. arXiv, 865–873. https:
[4] Xi Chen, Decheng Dai, Ye Du, and Shang-Hua Teng. 2009. Settling the Complexity //[Link]/10.48550/[Link]/0105025 arXiv:cs/0105025
of Arrow-Debreu Equilibria in Markets with Additively Separable Utilities. In [19] Leonard E Read. 1958. I, Pencil: My Family Tree as Told to Leonard E. Read. The
2009 50th Annual IEEE Symposium on Foundations of Computer Science. 273–282. Freeman 8 (December 1958), 32–37.
[Link] [20] Lionel W. McKenzie. 1959. On the Existence of General Equilibrium for a Compet-
[5] Xi Chen and Shang-Hua Teng. 2009. Spending Is Not Easier Than Trading: On the itive Market. Econometrica 27, 1 (1959), 54–71. [Link]
Computational Equivalence of Fisher and Arrow-Debreu Equilibria. In Algorithms arXiv:1907777
and Computation, Yingfei Dong, Ding-Zhu Du, and Oscar Ibarra (Eds.). Springer [21] Aidan McLau. 2024. AI Search: The Bitter-er Lesson.
Berlin Heidelberg, Berlin, Heidelberg, 647–656. [Link]
[6] Vincent Conitzer, Rachel Freedman, Jobst Heitzig, Wesley H. Holliday, Bob M. Ja- [22] Nolan Miller. 2006. Notes on Microeconomic Theory. (August 2006). (Lecture
cobs, Nathan Lambert, Milan Mosse, Eric Pacuit, Stuart Russell, Hailey Schoelkopf, Notes).
Emanuel Tewolde, and William S. Zwicker. 2024. Position: Social Choice Should [23] Marvin Minsky. 1988. Society Of Mind. Simon and Schuster.
Guide AI Alignment in Dealing with Diverse Human Feedback. In Proceedings of [24] Eric Neyman. 2024. Algorithmic Bayesian Epistemology. [Link]
the 41st International Conference on Machine Learning. PMLR, 9346–9360. 48550/arXiv.2403.07949 arXiv:2403.07949 [cs]
[25] Caspar Oesterheld, Abram Demski, and Vincent Conitzer. 2023. A Theory
[7] Lawrence Davis. 1988. Mapping Classifier Systems into Neural Networks. In
of Bounded Inductive Rationality. Electronic Proceedings in Theoretical Com-
Proceedings of the 1st International Conference on Neural Information Processing
puter Science 379 (July 2023), 421–440. [Link]
Systems (NIPS’88). MIT Press, Cambridge, MA, USA, 49–56.
arXiv:2307.05068 [cs]
[8] Abram Demski. 2024. Complete Feedback.
[26] OpenAI. 2024. Learning to Reason with LLMs.
[Link]
[27] Nasim Rahaman, Martin Weiss, Manuel Wüthrich, Yoshua Bengio, Li Erran
feedback.
Li, Chris Pal, and Bernhard Schölkopf. 2024. Language Models Can Reduce
[9] Scott Garrabrant, Tsvi Benson-Tilsen, Andrew Critch, Nate Soares, and Jessica
Asymmetry in Information Markets. [Link]
Taylor. 2020. Logical Induction. [Link]
arXiv:2403.14443 [cs]
arXiv:1609.03543 [cs, math]
[28] Yuma Rao, Jacob Steeves, Ala Shaabana, Daniel Attevelt, and Matthew McAteer.
[10] Luise Ge, Daniel Halpern, Evi Micha, Ariel D. Procaccia, Itai Shapira, Yevgeniy
2021. BitTensor: A Peer-to-Peer Intelligence Market. [Link]
Vorobeychik, and Junlin Wu. 2024. Axioms for AI Alignment from Human Feed-
arXiv.2003.03917 arXiv:2003.03917
back. In The Thirty-eighth Annual Conference on Neural Information Processing
[29] Jürgen Schmidhuber. 1987. Evolutionary Principles in Self-Referential Learning,
Systems.
or on Learning How to Learn: The Meta-Meta-. Hook.
[11] Dhananjay K. Gode and Shyam Sunder. 1993. Allocative Efficiency of Markets
[30] Jurgen Schmidhuber. 1989. A Local Learning Algorithm for Dynamic Feedforward
with Zero-Intelligence Traders: Market as a Partial Substitute for Individual
and Recurrent Networks. Connection Science 1, 4 (January 1989), 403–412. https:
Rationality. Journal of Political Economy 101, 1 (February 1993), 119–137. https:
//[Link]/10.1080/09540098908915650
//[Link]/10.1086/261868
[31] Alan Schwartz. 2008. How Much Irrationality Does the Market Permit? The
[12] Alex Graves. 2017. Adaptive Computation Time for Recurrent Neural Networks.
Journal of Legal Studies 37, 1 (January 2008), 131–159. [Link]
[Link] arXiv:1603.08983 [cs]
519963
[13] Faruk Gul and Ennio Stacchetti. 1999. Walrasian Equilibrium with Gross Substi-
[32] F. A. von Hayek. 1937. Economics and Knowledge. Economica 4, 13 (1937), 33–54.
tutes. Journal of Economic Theory 87, 1 (1999), 95–124. [Link]
arXiv:2548786
jeth.1999.2531
[33] John Wentworth. 2018. Competitive Markets as Distributed Backprop.
[14] Yizeng Han, Gao Huang, Shiji Song, Le Yang, Honghui Wang, and Yulin Wang.
[34] Yann LeCun. 2023. Towards Machines That Can Learn, Reason, and Plan. In AI
2021. Dynamic Neural Networks: A Survey. [Link]
and Barrier of Meaning Workshop. Santa Fe Institute.
04906 arXiv:2102.04906
[35] Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao,
[15] John H. Holland. 1985. Properties of the Bucket Brigade. In Proceedings of the 1st
and Karthik Narasimhan. 2024. Tree of Thoughts: Deliberate Problem Solving
International Conference on Genetic Algorithms. L. Erlbaum Associates Inc., USA,
with Large Language Models. In Proceedings of the 37th International Conference
1–7.
on Neural Information Processing Systems (NIPS ’23). Curran Associates Inc., Red
Hook, NY, USA, 11809–11822.