0% found this document useful (0 votes)
8 views27 pages

Dynamic Learning in Strategic Communication

This paper explores a dynamic model of strategic communication between a principal and an expert with conflicting preferences, demonstrating that the principal can achieve perfect information extraction in just two periods if the expert's bias is not excessively large. The study introduces a learning protocol that allows the expert to provide truthful information through a series of cheap-talk conversations, even in scenarios where traditional assumptions of the Crawford and Sobel model do not hold. The findings suggest that effective communication can be maintained without the principal knowing the expert's bias, and full information revelation is possible even with significant bias in an unbounded state space.

Uploaded by

shughao
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
8 views27 pages

Dynamic Learning in Strategic Communication

This paper explores a dynamic model of strategic communication between a principal and an expert with conflicting preferences, demonstrating that the principal can achieve perfect information extraction in just two periods if the expert's bias is not excessively large. The study introduces a learning protocol that allows the expert to provide truthful information through a series of cheap-talk conversations, even in scenarios where traditional assumptions of the Crawford and Sobel model do not hold. The findings suggest that effective communication can be maintained without the principal knowing the expert's bias, and full information revelation is possible even with significant bias in an unbounded state space.

Uploaded by

shughao
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Int J Game Theory (2016) 45:627–653

DOI 10.1007/s00182-015-0474-x

Dynamic learning and strategic communication

Maxim Ivanov1

Accepted: 19 March 2015 / Published online: 27 March 2015


© Springer-Verlag Berlin Heidelberg 2015

Abstract This paper investigates a dynamic model of strategic communication


between a principal and an expert with conflicting preferences. In each period, the
uninformed principal selects an experiment which privately reveals information about
an unknown state to the expert. The expert then sends a cheap talk message to the
principal. We show that the principal can elicit perfect information from the expert
about the state and achieve the first-best outcome in only two periods if the expert’s
preference bias is not too large. If the state space is unbounded, full information revela-
tion is possible for an arbitrarily large bias. Moreover, full revelation of information is
feasible in more general frameworks than those considered in the literature, including
frameworks which include non-quasiconcave and non-supermodular payoff functions
and those with a privately known bias of the expert.

Keywords Communication · Information acquisition · Cheap talk

JEL Classification C72 · D81 · D82 · D83

1 Introduction

This paper focuses on the classic problem of inefficient cheap talk communication
between parties with conflicting interests who are asymmetrically informed and have

The paper has been previously circulated under the title “Dynamic Informational Control”.

B Maxim Ivanov
mivanov@[Link]

1 Department of Economics, McMaster University, 1280 Main Street West, Hamilton, ON L8S 4M4,
Canada

123
628 M. Ivanov

different decision-making powers. It adds to the literature by investigating the follow-


ing question. Consider the principal (he) who does not possess important information
about the economic, political, or military consequences of his decisions. As a result, he
consults the expert (she), who has either more expertise in particular areas or signifi-
cantly lower costs of acquiring and processing new information. However, the expert’s
preferences are biased in a sense that her optimal decisions differ from those of the
principal, and thus does not report information truthfully (Crawford and Sobel 1982,
hereafter CS). The principal, however, can determine the quality of the expert’s private
information without being able to observe its content, and then request a report from
the expert about her observations. For instance, the principal may select the type of an
experiment performed by the expert, but he cannot see its outcome. The main ques-
tion of this paper is how precisely can the principal elicit information from the expert
through cheap-talk communication by determining the expert’s learning process?
This paper adds to the literature in two ways. First, it introduces a dynamic protocol
of learning information by the expert which allows the principal to extract perfect infor-
mation from the expert in only two rounds of cheap talk conversation.1 Second, the
paper emphasizes that perfect information extraction is feasible without maintaining
several fundamental assumptions of the CS model: the concavity and the supermod-
ularity of the players’ payoff functions, and the principal’s knowledge of the expert’s
bias.2 The only factors essential for effective work of the protocol are the maximal
intensity of the expert’s bias and the dependence of the expert’s payoff on the princi-
pal’s decisions in different states. Moreover, if the state space is unbounded, perfectly
informative communication is feasible for an arbitrarily large bias of the expert.
In order to explain these findings in a more detailed way, we start with an observation
that the failure of perfectly informative communication in the CS model is driven by the
possibility of the expert to mimic any information arbitrarily close to the true state. To
get around this difficulty, we introduce a dynamic protocol for acquiring information
(hereafter, a learning protocol) that (1) allows the expert to learn the state perfectly
in the second period, and (2) sustains truthful communication in both periods. The
learning protocol is not a mechanism, since it does not require any commitment of
the principal to the quality of expert’s information or decisions. It is deterministic,
that is, each state is mapped into a single signal in any period. The key feature of
the protocol is that it decomposes communication about the continuous state into a
continuum of cheap-talk conversations about binary posteriors, and then separates the
true state from the irrelevant one. In particular, it includes two experiments such that
the precision of the second experiment is contingent on the expert’s report on the
first experiment. The first experiment returns information about two isolated points
in the state space—the true state and some irrelevant complement state, which is

1 A single-stage protocol of learning information cannot implement the first-best outcome (Ivanov 2010a).
2 All or some of these assumptions are used in various communication models, e.g., communication with
multiple experts (Krishna and Morgan 2001a, b; Battaglini 2002), delegation (Dessein 2002; Alonso and
Matouschek 2008; Kovác and Mylovanov 2009), mediated communication (Goltsman et al. 2009), dynamic
communication (Krishna and Morgan 2004; Golosov et al. 2014), contracting for information (Krishna and
Morgan 2008), and communication via a noisy channel (Blume et al. 2007). Also, Morgan and Stocken
(2003) and Li and Madaràsz (2008) show that an outcome of communication can change drastically if the
expert’s bias is her private information.

123
Dynamic learning and strategic communication 629

sufficiently distinct from the true one—but does not reveal which state is true. Such
a pair forms a subset of posterior states. The second experiment allows the expert
to distinguish between reported posterior states only. Otherwise, the outcome of the
experiment is completely uninformative. Thus, the expert’s information is updated
in the second period if and only if she reported the truth in the previous period,
and even local distortions of information result in severe informational losses.3 In
that period, however, the principal knows the binary distribution of posteriors, which
prevents the expert from manipulating the second-period information. If the preference
conflict is not large, then the informational benefits of learning the true state outweigh
the benefits of manipulating the first-period imprecise information. Also, because
discrete second-period posterior beliefs sustain the fully informative equilibrium and
the expert is interested in learning the state in the second period if the absolute value of
the sender’s bias is below some cut-off, the learning protocol is robust to the principal’s
knowledge of the sender’s bias.
The main technical innovation of our protocol is that it employs the non-convexity
of posterior distributions for expert’s information acquisition and, as a consequence,
for influencing the principal’s beliefs through communication. This contrasts with
related works by Krishna and Morgan (2004) and Golosov et al. (2014) which estab-
lish that the principal can benefit from multi-stage communication in which the expert
reveals the non-convex set containing the state in the early period(s) and separates
the set into convex subsets afterwards (given the principal’s proper behavior). How-
ever, these authors consider the case of the perfectly informed expert and thus utilize
non-convex sets only for manipulating the principal’s beliefs. Such a communication
scheme, however, does not achieve the first-best outcome of the principal. In contrast,
generating the expert’s posterior beliefs with non-convex supports can qualitatively
improve the quality of transmitted information.
As a potential application of our results, consider interaction between an auto
mechanic (expert) and a car manufacturer (principal). The mechanic repairs cars on
the manufacturer warranty, and the manufacturer covers the costs of the repair. Before
servicing a car, the mechanic gathers information about it by using the testing equip-
ment that isolates the source of the problem and estimates its severity. This information
determines the total cost of the repair. Though both parties are concerned about repair-
ing the car, the interests of the mechanic are likely to be biased toward repairs with
higher costs. This naturally leads to the problem of manipulating information about
the severity of the problem.4
In this work, we offer a potential solution to the problem of communication about
car issues with high repair costs. In the context of our model, the car manufacturer

3 Ambrus and Lu (2013) apply discontinuous punishment of the expert for local distortions of information
via the principal’s actions the model of communication with multiple experts. As a result, the principal can
elicit almost full information from the experts.
4 Roberts (2004) provides a clear example of manipulating such information by auto mechanics. In his
words, “Sears Roebuck in 1992 sought to motivate the mechanics in its auto repair business by setting
targets for the amount of work they did. The mechanics responded by telling customers that they needed
steering and suspension repairs that were in fact unnecessary. Customers could not easily verify the need
for the repairs themselves, and many paid for the unnecessary work. When the fraud was uncovered, Sears
not only paid large fines, it lost much of the precious trust it had once enjoyed among its customers.”

123
630 M. Ivanov

may benefit from selecting a specific testing equipment which is used by mechanics
to estimate the problems. Many typical tests are performed exclusively by electronic
devices which are sufficiently accurate and which include the aforementioned quan-
tizers. For example, electronic car scanners return the code of the car’s sub-system not
operating properly. (This example is a sufficiently precise approximation of a contin-
uous state because the set of different outcomes is sufficiently large and may exceed
100 codes for a single car model.) Thus, instead of providing a single code, the first test
would return two codes, one of which is correct. Upon observing the outcome of this
test, the mechanic reports it to the manufacturer and uses it as an input signal for the
second test. The second test is conducted upon an approval from the manufacturer and
allows the mechanic to distinguish between the reported values only. The mechanic
then reports the outcome of the second test to the manufacturer who covers the costs
of the repair.
In the context of the above example, the principal can often acquire information
directly. However, this possibility is often restricted by the large volume, diverse
range, and complexity of other responsibilities (Radner 1993; Krishna and Morgan
2001a; Austen-Smith 1994). Therefore, high costs of information acquisition and
opportunity costs may require the principal to delegate information acquisition to
the expert. Moreover, Argenziano et al. (2013) show that the benefits of delegating
information acquisition to the biased expert (with further cheap-talk communication)
can be higher than those in the case of acquiring information by the principal even if
the players’ costs of information acquisition are identical. Another argument in favor
of using the expert is the decreasing returns to scale of the information acquisition
technology with respect to the number of performed tests.5
In a related paper, Ivanov (2015) considers a dynamic learning protocol in which
the sets of posterior states in each period are subintervals, and the precision of the
expert’s information is independent of her past reports. In this case, the expert also
faces a trade-off between future informational benefits and the set of feasible actions,
however, the natures of the trade-offs in the two papers are different. In the setup by
Ivanov (2015), the convexity of subsets of posterior states allows the expert to locally
distort her information. As a result, there is a subinterval in which states cannot be
separated. Outside of that subinterval, full information extraction is feasible in the
limit, i.e., as the number of periods increases without a bound. This is because the
learning protocol is independent of expert’s messages, so that the number of periods
cannot be reduced by refining the reported information only.
In addition to the literature on multi-period communication, our paper adds to works
on communication with the imperfectly informed expert(s). In this area, Green and
Stokey (2007) first demonstrated that the less informed expert can be beneficial to
the principal. Fischer and Stocken (2001) show this result for the CS model. Ivanov
(2010a) extends this result by demonstrating that communication with the imperfectly

5 If a single testing procedure must be applied to multiple independent variables, then obtaining information
about all variables by the principal is costly if, e.g., there are capacity constraints on the number of tests
performed. For instance, evaluating problems of cars covered by the manufacturer warranty is delegated
to car dealers, who then submit reports and bills to the car producer. In this case, designing the standards
on testing procedures for estimating car conditions and auditing the dealers is a less costly task for the
manufacturer than testing cars directly.

123
Dynamic learning and strategic communication 631

informed expert can benefit the principal more than optimally delegating authority
to the perfectly informed expert. Besides cheap-talk communication, the situations
in which the uninformed party(s) can influence the quality of private information
acquired by the informed agent(s) have attracted attention in various areas. Applica-
tions include auctions (Bergemann and Pesendorfer 2007; Board 2009), monopoly
(Lewis and Sappington 1994; Johnson and Myatt 2006), and markets of differentiated
products (Damiano and Li 2007; Ivanov 2013).6
Our work also complements a few papers on (almost) full information extraction in
communication games. These include cheap talk with multiple experts investigated by
Krishna and Morgan (2001a, b), Battaglini (2002), Esö and Fong (2008) and Ambrus
and Lu (2013). Kartik et al. (2007) consider a setting where the expert incurs lying
costs. The main distinction of our paper is the endogenous and dynamic quality of
expert’s information, whereas the information structure of the expert(s) in the afore-
mentioned works is exogenously given at the beginning of the game.
The remainder of the paper is structured as follows. Section 2 presents the formal
model. Section 3 highlights an illustrative example. The general analysis of the model
is performed in Sect. 4. Section 5 concludes the paper.

2 The model

Consider a two-period communication game with two players, an expert and a princi-
pal. The state θ is distributed on the interval  = [0, 1] according to prior distribution
F (θ ) with a positive and continuous density f (θ ). The expert can privately observe
some information about θ , whereas the principal makes a decision a ∈ R that affects the
payoffs of both players. The payoff functions are of the form U (a, θ, bi ) , i ∈ {E, P}
where the inherent bias parameter bi ∈ R reflects the divergence in players’ interests.
The principal’s bias is normalized to be 0, whereas the expert’s bias is b ≥ 0.
We consider a more general class of players’ preferences than those in the CS set-
ting. In particular, the function U (a, θ, b) is continuous in (a, θ, b) and has a unique
ideal decision a ∗ (θ, b) = arg maxa U (a, θ, b), which is continuous in θ and b. Here-
after, we write U (a, θ ) = U (a, θ, 0) and V (a, θ, b) = U (a, θ, b) as the principal’s
and the expert’s payoff functions, respectively. We also write a p (θ ) = a ∗ (θ, 0) and
y (θ, b) = a ∗ (θ, b) as the ideal decisions of the principal and the expert, respectively.
Denote the set of the principal’s ideal decisionsA = {a p (θ ) |θ ∈ }. We call deci-
sion a rationalizable if a ∈ A. Suppose that V a p (1) , 1, b ≥ V (a, 1, b) , a ∈ A,
i.e., the expert of the highest type would (weakly) prefer to separate her type
if the principal interpreted expert’s messages as truthful and optimally reacted to
them.
Also, we assume there exists s̄ ∈ (0, 1) and a continuously differentiable bijective
function ϕ : [0, s̄] → [s̄, 1], such that ϕ  (s) > 0 and a p (θ ) = a p (ϕ (θ )) , θ ∈ [0, s̄].
This condition reflects two facts. First, the subset [0, 1) of  can be split into two-point

6 Klein and Mylovanov (2011) investigate the conflict between the expert’s reputational concerns about
appearing competent and the incentives to report information truthfully in the dynamic environment in
which the expert can accumulate information about her competence over time.

123
632 M. Ivanov

 
sets θ, θ  , such that θ < s̄ ≤ θ  and the principal’sidealdecisions for any pair of
states θ and θ  are distinct. Second, all pairs of states θ, θ  with distinct principal’s
ideal decisions are not arbitrarily close. Hereafter, we refer to ϕ (θ ) as the complement
state function.7

2.1 Actions

At the beginning of period t = 1, 2, the principal selects a publicly observable expert’s


information structure It = {Ft (st |θ )}θ∈ ∈ S, where the signal space S ⊃  is
compact. That is, an information structure determines conditional distributions of the
signal for all θ in each period. Denote I = S ×  the space of all distributions of
signals conditional on the state. Then, the expert privately observes a signal st ∈ S
drawn from an associated distribution Ft (st |θ ). At the end of period t, the expert sends
a message m t ∈ M to the principal. Finally, upon receiving messages {m 1 , m 2 }, the
principal makes a decision a ∈ R.8

2.2 Strategies

We call a pair of first-period and second-period information structures a learning


protocol {I1 , I2 (m 1 , I1 )}, where I2 : M × I → I is a function of the first-period
principal’s history {m 1 , I1 }. Note the principal does not commit to the second-period
information structure I2 (m 1 , I1 ) from the beginning. Instead, he takes into account his
history {m 1 , I1 } while selecting I2 in the second period. A behavioral strategy of the
principal consists of the learning protocol and a decision rule a : M2 ×I 2 → R, which
is a function of the second-period principal’s history {m 1 , m 2 , I1 , I2 }. The behavioral
strategy of the expert is a pair of functions {σ1 , σ2 }, where σ1 : I × S → M
and σ2 : I 2 × S 2 × M → M, which map the expert’s private histories h 1 =
{I1 , s1 } and h 2 = {h 1 , I2 , s2 , m 1 } in the first and the second periods, respectively,
into the spaces of probability distributions on measurable message sets M in these
periods.

2.3 Beliefs

The principal’s belief is a pair μ = {μ1 (m 1 , I1 ) , μ2 (m 1 , m 2 , I1 , I2 )}, where μ1 :


M × I →  and μ2 : M × M × I 2 →  determine the principal’s beliefs in
the end of the first and the second periods, respectively. The belief is consistent if it is
derived from the players’ strategies on the basis of Bayes’ rule where applicable.9

7 Thus, compared to the CS setup, our model does not require: (1) the concavity of U (a, θ, b) in a for each
θ ; (2) the supermodularity of U (a, θ, b) in (a, θ ) or, equivalently, the monotonicity of a ∗ (θ, b) in θ ; and
(3) the single-sided bias y (θ, b)  = a (θ ) for all b and θ .
8 For simplicity, we restrict attention to pure strategies of the principal as functions of messages {m , m }.
1 2
9 All messages m ∈ 0
t / M, t = 1, 2 are interpreted by the principal as some m t ∈ M.

123
Dynamic learning and strategic communication 633

s1 (θ) s2 (θ, m1 )
s0 .. ..
1
... ...
m1 + 1 .. ..
... ...
2

.. ..
... ...
.. ..
... ...
1 .. ..
... ...
2

m1 .. ..
... ...
.. ..
θ . . . . . . . .... . . . . . . . . . . . . . . . . . . ... ... ...
.. .. .. ..
... ... ... ...
. . θ . . θ, m1
0 θ 1 θ =θ+ 1
2 1 0 θ m1 = θ θ m1 +
1
2 1
2

Fig. 1 Information structures in two periods

2.4 Equilibrium

A perfect
 ∗ ∗Bayesian equilibrium
 (hereafter, equilibrium) consists of the learning proto-
col I ,
 ∗ 1 2∗ I (m 1 , I 1 ) , the decision rule a ∗ (.), the belief μ∗ (.) and the expert’s strategy

σ1 (.) , σ2 (.) , such that μ (.) is consistent with the players’ strategies and the strate-
gies are optimal given the principal’s beliefs and any sender’s history at each moment.
The formal definition of an equilibrium is given in the Appendix.

3 Example: the CS uniform-quadratic case

We start with an illustrative example outlining the learning protocol under which
the expert truthfully reveals the state upon perfectly learning it in the second period.
This example highlights the key ideas that we use below in the more general economic
environments than the standard CS settings. For our example, we employ the uniform-
quadratic setup in which the prior distribution of θ is uniform on  = [0, 1] and the
preferences are quadratic:10
U (a, θ, bi ) = − (a − θ − bi )2 , i ∈ {E, P} .

Because b P = 0 and b E = b, the ideal decisions of the principal and the expert are
a p (θ ) = θ and y (θ, b) = θ + b, respectively. Suppose that b ≤ b̄ = 41 , where b̄ is
the largest bias, which sustains informative communication in the CS model.
Define the first-period information structure as follows (see Fig. 1). Each pair of state
θ < 21 and the complement state θ  = ϕ (θ ) = θ + 21 map into the same signal s = θ .
Such a partitioning of  into two-point sets generates a binary
 posterior
 distribution
of θ conditional on s with probabilities Pr {θ = s|s} = Pr θ = s + 21 |s = 21 , s < 21 .

10 The uniform-quadratic setup is known for the tractability in various modifications of the basic CS model.
See, for example, Blume et al. (2007), Goltsman et al. (2009), Krishna and Morgan (2001a, b, 2004), Krishna
and Morgan (2008), Melumad and Shibano (1991), Ottaviani and Squintani (2006), and Esö and Szalay
(2010).

123
634 M. Ivanov

The second-period information structure updates the expert’s reported information


only. In particular, given message m 1 < 21 , the expert can separate state m 1 from the
complement state ϕ (m 1 ). For θ ∈ / {m 1 , ϕ (m 1 )}, the information structure returns
an uninformative signal s0 ∈ / [0, 1). Thus, the expert learns the true state in the
second period if and only if she truthfully reports her information in the first period.
Otherwise, reporting m 1 = s does not update the expert’s information. Upon receiving
message m 2 ∈ {m 1 , ϕ (m 1 )}, the principal interprets it as truthful and makes a decision
a (m 2 ) = m 2 .
Intuitively, the expert in the first period faces a trade-off between informational
benefits and flexibility over actions. On one hand, she can induce any rationalizable
action y ∈ A = [0, 1] by sending either m 1 = y if y < 21 , or m 1 = y − 21 if y ≥ 21 ,
and then sending m 2 = y in the second period. In this case, however, the decisions are
based on the expert’s interim information only. On the other hand, truthfully reporting
m 1 = s allows the expert to learn θ , but shrinks the set of feasible actions to two. These
are the principal’s best-responses to posterior realizations s and ϕ (s) , a p (s) = s and
 
a p s + 21 = s + 21 , respectively.
From the expert’s perspective, inducing decision y ∈ A results in the payoff

1 1
E [V (y, θ, b) |s] = V (y, s, b) + V (y, ϕ (s) , b)
2 2
 2
1 1 1
= − (y − s − b)2 − y−s− −b ,
2 2 2

which is maximized at y1 (s, b) = s + 41 + b < 1, since b ≤ 1


4 and s < 21 . Inducing
y1 (s, b) provides the payoff to the expert:

1 1 1
E [V (y1 (s, b) , θ ) |s] = − (ϕ (s) − s)2 = − , s < .
4 16 2
Truthful reporting in both periods results in the payoff:

   1 1 1
E V a p (θ ) , θ, b |s = − (s − s − b)2 − (ϕ (s) − ϕ (s) − b)2 = −b2 , s < .
2 2 2
Finally, the second-period incentive compatibility constraints must prevent the pos-
itively biased expert from distorting information in period two. That is, inducing
a p (θ ) = θ upon learning θ < 21 must be more beneficial than inducing a p (ϕ (θ )) =
ϕ (θ ) > a p (θ ), or

    1
V a p (s) , s, b ≥ V a p (ϕ (s)) , s, b , s < . (1)
2

These constraints are satisfied if a p (ϕ (θ )) − a p (θ ) = ϕ (s) − s = 21 ≥ 2b. Thus, if


b ≤ 41 , then the inequalities (1) and the “ learning incentive constraints”

   1
E V a p (θ ) , θ, b |s = −b2 ≥ − = E θ [V (y1 (s, b) , θ ) |s] ,
16

123
Dynamic learning and strategic communication 635

imply that the expert prefers to convey her information truthfully in both periods and,
hence, reveals the true state in the second period.
Since the principal achieves his first-best outcome, he does not have profitable
deviations in any period. However, we need to specify the players’ reaction to the
principal’s deviations from the prescribed information structures. In this case, the
expert sends an uninformative message, or babbles, until the end of the game, whereas
the principal ignores all of the expert’s messages after his deviation. Because there
is a babbling equilibrium given any information structure, such out-of-equilibrium
behavior is optimal for both players.
Intuitively, the efficiency of the constructed learning protocol is driven by a combi-
nation of two factors. First, the quality of the expert’s information in the second period,
and thus the benefits of updating the information are contingent on her first-period
report. If the expert conveys her imprecise information truthfully, the second-period
information structure allows her to perfectly learn the state. Second, truthtelling in
the first period reveals the particular binary distribution of posteriors to the princi-
pal, who can use it in order to (partially) verify the expert’s report in the next period.
This significantly restricts the expert’s possibilities of manipulating her information in
round two, because the principal makes only those decisions that are consistent with
received information in both periods. Together, these factors imply that the expert
faces a trade-off between the future informational benefits (received in the case of
truthful reporting) and the set of feasible actions (induced by distorting her interim
information). If the preference bias is not large, the former effect outweighs the latter
one. In addition, a smaller set of feasible actions in round two helps sustain truthful
communication in that period. In fact, the expert can induce only the principal’s best
responses to either true or complement states. If the latter option is unfavorable, the
expert prefers to tell the truth. In short, the learning protocol gives a lot of freedom
over decisions to the expert because the principal can be easily manipulated by the
expert’s messages. However, the value of such freedom is low if the expert is not well
informed about the true state.11
Technically, the learning protocol decomposes the prior distribution into a contin-
uum of discrete posterior distributions. This decomposition plays a dual role. First, it
introduces substantial uncertainty in the first-period expert’s information which moti-
vates her to acquire information in the second period. Second, the discrete set of
posterior states prevents the expert from manipulating her information in the second
period.
An important feature of the introduced protocol is its high sensitivity to reported
information. For implementation, however, this feature is less critical than it may
seem. The sensitivity is dictated by the extreme efficiency of the learning protocol
within the very limited time of interaction. If the principal does not aim to achieve
the first-best outcome, he can use simpler information structures that are less sensi-
tive to conveyed information while preserving sufficient efficiency. This is because

11 In this light, it is important to note that the efficiency of interaction relies on the expert’s being unable to
get outside information. Otherwise, the expert could run the first test to find out the two possible states and
then privately eliminate one of them with an outside investigation. Having then discovered the true state,
he could mislead the principal afterwards.

123
636 M. Ivanov

the construction is equally applicable to arbitrary non-convex subsets as to pairs of


states. Consider a learning protocol in which the first-period information structure
Nk
reveals a finite family of non-convex measurable subsets Q k = kj that contain
j=1
the state. In the second period, the expert can identify the subset kjθ ∈ Q k contain-
ing the state upon truthtelling in the previous period. Consider, for example,
 the bias
b = 15 and the simplest non-trivial learning protocol with the posterior sets 11 , 12 =
 1  1 3   2 2   1 1  3  12
0, 4 , 2 , 4 and 1 , 2 = 4 , 2 , 4 , 1 . This results in an ex-ante payoff
to the principal of − 192
1
, which exceeds the ex-ante payoff in all known non-first-best
13
incentive schemes. In general, for quadratic preferences and an arbitrary distri-
bution, a learning protocol that achieves the principal’s ex-ante losses ε relative
to the first-best payoff must contain approximately (12ε)−1/2 sub-intervals kj in
total.14

4 Dynamic information extraction

In this section, we extend the example above to general prior distributions and players’
payoff functions. In particular, we construct learning protocols that sustain the fully
informative equilibria in which the expert reveals the state upon learning it in the
second period.

4.1 Learning protocols

Consider the first-period information structure which generates the expert’s signal

θ if θ < s̄ or θ = 1,
s1 (θ ) = (2)
ϕ −1 (θ ) if s̄ ≤ θ < 1.

Because the first-period information structure maps state θ < s̄ and a complement state
θ  = ϕ (θ ) ≥ s̄ into the same signal, the expert cannot distinguish between states θ and
ϕ (θ ) upon observing a signal s = θ . Thus, the information structure (2) generates the
family of the first-period posterior distributions {F (θ |s) , s ∈ [0, s̄), θ ∈ {s, ϕ (s)}}

12 The simplest non-trivial learning protocol is the protocol with the smallest number of Q and k in
k j
which the expert acquires private information in the first period (that is, it contains at least two families Q k )
and can update her information in the second period (i.e., each Q k contains at least two subsets).
13 To the best of our knowledge, these schemes are restricted delegation with monetary transfers for reported
information (Krishna and Morgan 2008) and restricted delegation to an imperfectly informed expert (Ivanov
1 and − 1 , respectively.
2010a). They provide the ex-ante payoffs to the principal approximately − 38 98
14 The principal’s ex-ante payoff in the most informative equilibrium is equal to the average residual
variance across sub-intervals kj . Suppose that the total number of sub-intervals in the learning protocol is
N and all sub-intervals are of the equal length  = N1 . Since the density function can be approximated by
2
the piecewise constant density on each sub-interval for large N , this gives EU −  1
12 = − 12N 2 . Setting
EU = −ε, where ε > 0 is sufficiently small, determines the necessary number of sub-intervals.

123
Dynamic learning and strategic communication 637

and the degenerate distribution at θ = 1 such that each F (θ |s) is supported on


{s, ϕ (s)} with the probabilities:15

f (s)
ps = Pr {θ = s|s} = , and
f (s) + f (ϕ (s)) ϕ  (s)
psc = Pr {θ = ϕ (s) |s} = 1 − ps , if s ∈ [0, s̄).

Hereafter, we call the support {s, ϕ (s)} of a posterior distribution F (θ |s) the posterior
set, and each pair of states s and s  = ϕ (s) posterior states.
Given message m 1 ∈ [0, s̄), the expert’s second-period information structure gen-
erates a signal which separates posterior states m 1 and ϕ (m 1 ), but keeps other states
unseparated:

θ if θ ∈ {m 1 , ϕ (m 1 )} and m 1 ∈ [0, s̄),


s2 (θ, m 1 ) = (3)
s0 ∈
/ [0, 1) otherwise.

The learning protocol (2) and (3) is characterized by a few key properties that influ-
ence the expert’s incentives to convey information. First, the binary set of posteriors
makes the expert’s information in the first period sufficiently imprecise. Because the
principal’s ideal decisions for states s and ϕ (s) are different, he would be interested in
learning the state upon learning that θ ∈ {s, ϕ (s)}. If the expert’s bias is not large, her
ideal decisions are close to those of the principal. Therefore, the expert is also inter-
ested in learning the true state in the second period.16 At the same time, the principal
wants to limit the set of posteriors for each signal s. This is necessary in order to restrict
the expert’s possibilities of mimicking another posterior after perfectly learning the
state in round two.
Also, the distance between the posterior states ϕ (s) − s affects the expert’s incen-
tives due to two effects. First, a smaller distance reduces the value of learning θ in the
second period. This is because it increases the expert’s interim payoff from inducing
her optimal decision while the interim payoff from learning the state and inducing
an ideal principal’s decision does not change significantly. Second, it provides the
incentive to the expert to distort information in the second period upon learning θ .
For example, the positively biased expert may benefit from inducing a p (ϕ (θ )) if it
is closer to her ideal decision than a p (θ ). However, the principal is limited in maxi-
mizing the distance ϕ (s) − s for all s by choosing a different function ϕ (s). This is

15 The derivation of the formulas for p and p c is in the Appendix.


s s
16 The effect of the binary first-period information structure on the expert’s uncertainty is clearly seen
in the case of the risk-averse expert, i.e., if V (a, θ, b) is concave in a for all (θ, b). For a fixed interval
[θ1 , θ2 ], consider all distributions of a random variable X with the support L ⊂ [θ1 , θ2 ] and a given
mean value E [X ] ∈ [θ1 , θ2 ]. In this class, the binary distribution on {θ1 , θ2 } with the probabilities pθ1 =
θ −E[X ]
Pr {X = θ1 } = 2θ −θ and pθ2 = 1 − pθ1 , respectively, dominates any other distribution by the convex
2 1
order. The proof follows from (3.A.8) in Shaked and Shanthikumar (2007). In our model, this means that the
first-period information structure with binary distributions of posterior values of θ maximizes the expert’s
uncertainty. This increases her benefits from learning θ in the second period and motivates her to reveal her
first-period signal.

123
638 M. Ivanov

because for any partition of [0, 1) consisting of the two-point subsets, the minimum
distance between pairs of points is bounded by 21 .

4.2 Information extraction in bounded state space

We show now how the constructed learning protocol allows the principal to elicit full
information from the expert. Suppose that the expert acquires information according
to the information structures (2)–(3) and sends the sequence of messages {m 1 , m 2 }.
For the information structures (2)–(3), define the principal’s decision rule a ∗ (m 1 , m 2 )
as ⎧
⎪ a p (m 2 ) if m 1 ∈ [0, s̄) and m 2 ∈ {m 1 , ϕ (m 1 )} ,


∗ a (1) if m 1 = 1, ∀m 2 ,
a (m 1 , m 2 ) = (4)

⎪ a p (0) / [0, s̄) ∪ {1} , ∀m 2 , and
if m 1 ∈

a p (m 1 ) if m 1 ∈ [0, s̄) and m 2 ∈/ {m 1 , ϕ (m 1 )} .
If the expert sends equilibrium messages {m 1 , m 2 }, i.e., m 1 ∈ [0, s̄) and m 2 ∈
{m 1 , ϕ (m 1 )} or m 1 = 1, then the principal interprets them as truthful and implements
decision a p (m 2 ). Any first-period out-of-equilibrium message m 1 is interpreted by the
principal as m 01 = 0 and results in decision a p (0) for any m 2 .17 Finally, any second-
period out-of-equilibrium message m 2 is disregarded by the principal who makes a
decision on the basis of m 1 .
Let Y1 (s, b) be the set of maximizers of the expert’s interim payoff:

Y1 (s, b) = arg max ps V (a, s, b) + (1 − ps ) V (a, ϕ (s) , b) , s ∈ [0, s̄).


a∈A

That is, each y1 (s, b) ∈ Y1 (s, b) is an optimal interim decision based on the first-
period information only. The expert can induce y = y1 (s, b) by sending the proper
messages in both periods.18 Because her information will not be updated in the second
period, inducing y1 (s, b) results in the payoff:

E [V (y1 (s, b) , θ, b) |s] = max ps V (a, s, b) + (1 − ps ) V (a, ϕ (s) , b)


a∈A
= ps V (y1 (s, b) , s, b) + (1 − ps ) V (y1 (s, b) , ϕ (s) , b) , s ∈ [0, s̄). (5)

If the expert reveals her information truthfully by sending m 1 = s, her interim infor-
mation
 will be updated in the second period, but the set of feasible actions shrinks
to a p (s) , a p (ϕ (s)) . In this case, reporting the truth in both periods results in the
payoff:
      
E V a p (θ ) , θ, b |s = ps V a p (s) , s, b + (1 − ps ) V a p (ϕ (s)) , ϕ (s) , b .
(6)

17 The choice of m 0 = 0 is not important. Equivalently, it can be replaced by any m  ∈ [0, s̄).
1 1
18 The messages {m , m } that induce y = y (s, b) can be constructed as follows. Denote A =
  1 2  1  1
a p (θ ) |θ ∈ [0, s̄) and A2 = a p (θ ) |θ ∈ [s̄, 1) . Then, define m 1 = m 2 ∈ [0, s̄) ∩ a −1 p (y) if
y ∈ A1 , m 1 = ϕ −1 (m 2 ) , m 2 ∈ [s̄, 1) ∩ a −1
p (y) if y ∈ A2 , and m 1 = m 2 = 1 if y = a p (1).

123
Dynamic learning and strategic communication 639

Since the expert can induce only a p (s) or a p (ϕ (s)) upon reporting s < s̄, her second-
period incentive-compatibility constraints are given by:
   
V a p (s) , s, b ≥ V a p (ϕ (s)) , s, b , s ∈ [0, s̄), and (7)
   
V a p (ϕ (s)) , ϕ (s) , b ≥ V a p (s) , ϕ (s) , b , s ∈ [0, s̄). (8)

Also, the learning incentive constraints in the first period require:


  
E V a p (θ ) , θ, b |s ≥ E [V (y1 (s, b) , θ, b) |s] , s ∈ [0, s̄). (9)

Finally, we model the out-of-equilibrium behavior of the players in the case of prin-
cipal’s deviation to a different information structure at any period as in the example
above. Then, fully informative communication is sustainable if (7)–(9) hold.
The following theorem demonstrates that the trade-off between the informational
benefits and the flexibility over available actions is in favor of the former if the bias
in preferences is not large. All proofs are collected in the Appendix.

Theorem 1 There exists b̄ such that if b ≤ b̄, there is a fully informative equilibrium
with the learning protocol determined by (2) and (3).

The proof is constructive and follows from the structure of the learning protocol. If
the players’ interests are close enough, the ideal decisions of the players are sufficiently
close for each state. Since the expert’s information in the first period is imprecise, she
is more interested in learning the state perfectly at the cost of inducing the principal’s
ideal decision than inducing an optimal interim decision. Because the expert can learn
the state only by revealing her first-period information, the principal knows the binary
distribution of posteriors at the beginning of the second period. For this distribution,
truthtelling communication is sustainable if the bias is not large.
A few comments are necessary. First, effective communication in our model does
not rely on the (quasi)-concavity of U (a, θ, b) in a and the supermodularity in (a, θ ),
because these properties of the payoff function influence the expert’s incentives to
manipulate her information locally by mimicking a slightly different state. Because
the learning protocol precludes such a possibility for the expert, these assumptions are
unnecessary.
Second, dynamic learning is different from mediated communication introduced by
Myerson (1991) and studied by Blume et al. (2007), Goltsman et al. (2009) and Ivanov
(2010b). In mediated communication, the expert faces the message-contingent lottery
over induced actions chosen by the mediator, who privately requests information from
the expert and gives recommendations to the principal. In our model, the expert in
the first period faces the message-contingent lottery over her own posterior types
generated by the second-period information structure. The bigger effectiveness of a
lottery over the expert’s types relative to a lottery over the principal’s decisions is
driven by two factors. First, in order to screen a continuum of types, the mediation
rule has to be delicate in punishing the expert for lying. Suppose that the mediation
rule punishes the expert’s type θ for claiming type θ  close to θ by inducing some
unfavorable distribution over actions. By the continuity of the expert’s payoff function,

123
640 M. Ivanov

such a lottery also punishes type θ  for truthtelling and motivates this type to distort
information. In contrast, the lottery over posterior types is able to separate the intensity
of expert’s punishment for distorting information across first-period types (signals).
As a result, each expert’s type can be punished for slight distortions of information
without damaging the incentives of arbitrarily close types to convey their information.
Sending a message s  different from the observed signal s in the first period does not
provide new information only to type s. These informational losses, however, do not
affect type s  who can still benefit from truthtelling.
Also, mediated communication relies on the concavity of expert’s preferences. The
mediator explores this property by offering a message-contingent lottery in which the
variance over recommended actions depends on the reported type. A side effect of
this randomization is that each state maps into several actions, which cannot be the
ideal principal’s decisions together. In dynamic learning, the expert in the first period
is interested in acquiring more information unconditionally on the concavity of her
payoff function. That is, instead of imposing the (quasi)-concavity condition on the
sender’s payoff function, we only assume that the sender’s ideal decisions for posterior
states are unique and distinct. Also, there is no need to randomize over actions after the
expert learns θ in the second period. By that time, the principal knows the distribution
of posterior types, which are distant enough to separate themselves without additional
incentives.
Finally, if the expert’s bias is her private information, but all values of b are below
b̄, Theorem 1 still holds. In contrast, if b is sufficiently large, then the fully informative
equilibrium might not exist for any function ϕ (θ ) in the learning protocol given
by (2) and (3). In order to demonstrate this, assume that a p (ϕ (s)) > a p (s) for
some s and and y (s,  b) > sup A for a sufficiently  large b. Then, it follows that
V a p (ϕ (s)) , s, b∗ = V (y(s, b∗ ), s, b∗ ) > V a p (s) , s, b∗ = V (y (s, 0) , s, b∗ )
for some b∗ . Thus, the expert strictly prefers to induce action a p (ϕ (s)) in the second
period upon learning that the state is s if b ≥ b∗ .19

4.3 Comparative statics

In the light of Theorem 1, an important question to examine is whether full information


extraction is monotone in the magnitude of the expert’s bias. In other words, what are
the conditions on the primitives, i.e., the players’ payoff functions and the distribution
of states, which sustain the full information extraction if and only if b is below some
cut-off? To address this issue and extend the results to the case of the unbounded ,
we impose the standard CS conditions on the players’ preferences in this subsection:

 < 0, U  > 0, and U  > 0.


A 1 U (a, θ, b) is twice-differentiable and satisfies Uaa aθ ab

 
19 In the uniform-quadratic example above, the distance s − s   ≤ 1 for s  = 1 . Hence, if b > 1 ,
2 2 4
then the second-period incentive-compatibility constraints (1) are violated for some s. In addition,
sup 1 > E  V a (θ ) , θ, b |s = −b2 means that the expert’s
E [V (y1 (s, b) , θ ) |s] ≥ − 16 p
s,ϕ(s):∪{s,ϕ(s)}=
s
learning incentive constraints (9) are also violated.

123
Dynamic learning and strategic communication 641

Given these conditions, a ∗ (θ, b) is strictly increasing in θ and b. Through this


subsection, we also suppose that Ub (a ∗ (θ, b) , θ, b) ≤ 0.20 Consider a differentiable
and strictly increasing function ϕ : [0, 1/2] → [1/2, 1] and the difference between
the expert’s interim payoffs in the cases of truthful communication and inducing an
arbitrary decision y ∈ R:
  
V (y, s, b) = E V a p (θ ) , θ, b |s − E [V (y, θ, b) |s]
   
= ps V a p (s) , s, b + (1 − ps ) V a p (ϕ (s)) , ϕ (s) , b
− ps V (y, s, b) − (1 − ps ) V (y, ϕ (s) , b) .

Suppose that y1∗ (s, b) ∈ R is the expert’s optimal (possibly unfeasible) interim deci-
sion:

y1∗ (s, b) = arg max E [V (a, θ, b) |s]


a∈R
= arg max ps V (a, s, b) + (1 − ps ) V (a, ϕ (s) , b) , s ∈ [0, 1/2).
a∈R

Then, if the difference in the payoffs in the cases of truthful communication and
inducing the (non-feasible) decision y1∗ (s, b) is not increasing in b, then the fully
informative equilibrium exists if and only if the bias is below some cut-off.
 
Lemma 1 If Vb y = y1∗ (s, b) , s, b ≤ 0, s < 1/2, then there is a fully informative
equilibrium with the learning protocol determined by (2) and (3) if and only if b ≤ b̄.
The main reason for using y1∗ (s, b) instead of y1 (s, b) in Lemma 1 is that it is not
necessary to check whether the boundary condition
 y1 (s, b) ≤ a p (1) is binding for
all s and b. Thus, verifying the sign of Vb y = y1∗ (s, b) , s, b is a less complicated
exercise than checking it for Vb (y = y1 (s, b) , s, b). The following corollary is an
implication of Lemma 1 for a class of preferences broadly used in cheap-talk and
mechanism design literature.21
Corollary 1 If V (a, θ, b) = V (a − β (b, θ ) , θ ), where β (0, θ ) = 0, βb ≥ 0, and
 ≤ 0, then there is a fully informative equilibrium with the learning protocol
βbθ
determined by (2) and (3) if and only if b ≤ b̄.

4.4 Information extraction in unbounded state space

If the state space is unbounded, the expert’s first-period information structure can
be partitioned into posterior sets such that the distance between any posterior states
θ and θ  = ϕ (θ ) is arbitrarily large. This might suppress the expert’s incentive to

20 That is, there are no private benefits to the expert for simply being more biased.
21 Besides the quadratic payoff function with a constant bias, this class includes the generalized quadratic
functions V (a, θ, b) = − (a − y (θ, b))2 used by Alonso and Matouschek (2008) and Kovác and Mylo-
 
vanov (2009), and the symmetric functions V (a, θ, b) = V |a − (θ + b)| exploited by Dessein (2002).
Krishna and Morgan (2004) use a special case of the symmetric functions, V (a, θ, b) = − |a − (θ + b)|ρ .

123
642 M. Ivanov

distort her information due to two factors. First, a large distance between the posterior
states implies that the optimal decision y1 in the case of observing the signal s and
deviating from truthtelling in period one would differ significantly from at least one
of the ideal decisions y (s, b) and y (ϕ (s) , b). As a result, the expected losses from
distorting information should increase. Second, upon communicating truthfully   in the
first period and observing s < s  = ϕ (s) in the next period, inducing a p s  results in
high losses of the expert if s  is sufficiently distant from s. These observations suggest
that the principal can potentially extract all information from the expert for an arbitrary
magnitude of the expert’s bias.
The intuition above, however, misses an essential detail. The value of second-period
information is high if the posterior states are sufficiently distinct and approximately
equally likely. For a fixed state θ , if the distance between the posterior states θ and
θ  > θ increases, then the density f θ  eventually converges to 0 and the posterior
probability of θ converges to 1. Thus, the expert infers that the posterior realization
θ = s is more likely than θ  = ϕ (s). This decreases her informational benefits in
the second period and, as a result, the incentives to truthfully convey information in
the first period. These arguments
  imply that truthtelling can be sustained only if the
magnitudes of f (θ ) and f θ  do not differ significantly.
In order to generalize the logic above, consider the distribution of states with a
positive density f (θ ) > 0 on  = R+ . In addition to the assumptions on the payoff
functions maintained for the case of bounded  , suppose thatthe set of
 rationalizable
decisions of U (a, θ ) is unbounded, A = a p (0) , ∞ . Let y1∗ θ, θ  , b be the expert’s
optimal action given that the state is either θ or θ  with probabilities pθ = pθ  = 1/2:
   
y1∗ θ, θ  , b = arg max V (a, θ, b) + V a, θ  , b . (10)
a∈A

Suppose V (a, θ, b) is strictly pseudo-concave in a and satisfies the following condi-


tions.
 
A 2 For any b, there exists δ (b) > 0, such that  y (θ, b) − a p (θ ) ≤ δ (b) , ∀θ .

A 3 For any b and δ > 0, there exists δ̄ (b, δ) > 0, such that
 
V (y (θ, b) ∓ δ, θ, b) ≥ V y (θ, b) ± δ̄ (b, δ) , θ, b , for all θ.

A 4 For any b and δ > 0, there exists d (b, δ) > 0 such that y (θ, b) + δ <
y1∗ (θ, θ + d (b, δ) , b) < y (θ + d (b, δ) , b) − δ, ∀θ .

Condition (A2) states that there is a uniform bound on the expert’s bias. Condition
(A3) requires that the expert is not infinitely more sensitive to the actions below her
ideal decision than to the actions above and vice versa.22 Finally, (A4) establishes that
(1) moving the posterior states θ  and θ sufficiently apart from each other guarantees
that the expert’s ideal decisions will be sufficiently apart as well, and (2) the optimal
decision given information that the true state is either θ or θ  is not too close to the ideal

22 Ambrus and Lu (2013) use this condition in the context of communication with multiple experts.

123
Dynamic learning and strategic communication 643

decisions for either realization. These conditions do not impose strong restrictions on
the shape of the expert’s payoff function, and imposing only (A2) might be sufficient
for satisfying (A3) and (A4).23
Given these preliminaries, the theorem below provides sufficient conditions on
f (θ ) under which full information extraction is feasible for the expert with an arbitrary
bias.

Theorem 2 If V (a, θ, b) satisfies conditions A2–A4, then for any b, there exist
d (b) > 0 and ε (θ, b) ∈ (0, 1), such that if ε (θ, b) ≤ f (θ+d(b))
f (θ) ≤ ε(θ,b)
1
, ∀θ , then
there is a learning protocol that sustains the fully informative equilibrium.

The learning protocol in Theorem 2 works as follows. First, the state space is
divided into the sequence of half-open intervals of a fixed length, which depends on
the expert’s bias. Then, each interval is partitioned into two-state posterior sets such
that the distance between the posterior states is equal to a half of the interval’s length.
The conditions (A2)–(A4) imply that if the probabilities of posterior realizations in
each posterior set are equal, then moving the posterior states sufficiently apart from
each other by increasing the length of intervals forces the expert to communicate
truthfully. Because it is impossible to keep the identical posterior probabilities for any
distribution with an unbounded support, the condition in Theorem 2 means that the
likelihood ratio between any states θ and θ + d must not vary significantly, where
d depends on the magnitude of the bias. In this case, the signal in the first period
is sufficiently uninformative. This guarantees that the expert’s benefits more from
learning the state in the second period than from manipulating her interim information.
As a special case of Theorem 2, consider the family of quadratic payoff func-
tions V (a, θ, b) = − (a − θ − β (b, θ ))2 and suppose F (θ ) satisfies the following
condition.

f (θ+d)
A 5 There exist α > 0 and d > 0 such that e−αd ≤ f (θ) ≤ eαd , θ ≥ 0.

This condition states that the likelihood ratio f (θ+d)


f (θ) varies less rapidly than expo-
nentially with the rate α, i.e., the prior distribution varies more smoothly than the
exponential distribution with the parameter α. The following corollary establishes the
sufficient conditions for full information extraction under the above assumptions.

Corollary 2 If V (a, θ, b) = − (a − θ − β (b, θ ))2 , β (b, θ ) ≤ δ, and F (θ ) satisfies


(A5) for d = 2δ (2/z ∗ + 1) and α = z ∗ /δ, where z ∗ = exp (−z ∗ /2 − 1) 0.314,
then there exists the learning protocol that sustains the fully informative equilibrium.

23 Consider the symmetric payoff function V (a, θ, b) = V (|a − θ − β (b, θ )|), where V  (x) ≤
0, V  (x) < 0, x ≥ 0. Then, a p (θ ) = θ, y (θ, b) = θ + β (b, θ ), and y1∗ (θ, θ + d, b) = θ + d2 +
β(b,θ)+β(b,θ+d)
2   . For a fixed b, (A2) requires β (b, θ ) ≤ δ (b) , ∀θ . This implies V (y (θ, b) ∓ δ, θ, b) =
 
V (δ) ≥ V δ̄ = V y (θ, b) ± δ̄, θ, b if δ̄ ≥ δ, so (A3) holds. Finally, y (θ, b) + δ < y1∗ (θ, θ + d, b) <
y (θ + d, b) − δ if d > δ (b) + 2δ, so (A4) holds.

123
644 M. Ivanov

5 Conclusion and discussion

This paper demonstrates how dynamic management of information assessed by a pri-


vately informed expert can completely eliminate the informational asymmetry between
the expert an uninformed decision maker. Though the parties have conflicting inter-
ests and communicate via cheap-talk messages, the proper design of the protocol
of learning information by the expert allows the principal to elicit perfect infor-
mation about an unknown state and achieve his first-best decision after only two
rounds.
Our results raise several natural questions. First, what is the class of all learning
protocols under which the principal can learn the state perfectly? Though the full
characterization of this class seems to be a difficult problem, a partial characteri-
zation is possible by noting that the learning protocol offered in this paper belongs
to a more general class of protocols whose information structures efficiently relax
the expert’s incentives to distort information while allowing her to acquire perfectly
precise information on the equilibrium path. In this class, an information structure
in each period is deterministic, that is, each state maps into a single signal. As a
result, posterior sets indexed by signals form a partition of the posterior set in the
previous period. The expert’s report, however, determines how her information is
updated in the next period. In particular, the second-period information structure
separates the states in the posterior distribution associated with a reported signal
only. Because the set of posterior states indexed by signals forms a partition, non-
truthful reporting of a signal induces the same posterior distribution of the expert as
that in the previous period, i.e., it does not update expert’s information. In contrast,
truthful reporting allows the expert to learn the state perfectly. Thus, the second-
period information structure implements the extreme reward/punishment in terms
of informational benefits for truthtelling/misreporting, respectively. In addition, the
expert’s posterior distribution in the first period (and thus in the second period) is
discrete. This provides the expert with the incentives to learn the state in the sec-
ond period (by truthful reporting in the first period) and sustains truthful reporting
in the second period.24 Therefore, modifying any characteristics of the learning
protocol—the partitional character of information structures, the extreme sensitiv-
ity of the quality of the expert’s updated information to her reports, or the discreteness
of posterior distributions—can only tighten the expert’s incentive-compatibility con-
straints.
In this light, it is important to note that full information extraction sometimes
demands first-period information structures which generate non-binary sets of poste-
rior states or coupling posterior states in a different way. This is especially effective in
the case of distributions with zero density at some states. Consider, for example, a trian-

24 Denote t =supp F (θ |h ) , t = 1, 2 the support of the posterior distribution F (θ |h ) in period


ht t t t t
t given a learning protocol I1 , I2 (m 1 , I1 ) and expert’s private histories h t , where h 1 = {I1 , s1 } and
h 2 = {h 1 , I2 , s2 , m 1 }. In a described class of learning protocols, a first-period information structure I1 is
partitional, i.e., 1I ,s ∩1  = ∅for s1  = s1 and ∪ 1I ,s = . Also, the second-period information
1 1 I1 ,s1 s1 ∈S 1 1
structure I2 (m 1 , I1 ) separates states in the posterior distribution associated with a reported signal only, i.e.,
2h = s2 ∈ 1h if m 1 = s1 , and 2h = 1h if m 1  = s1 .
2 1 2 1

123
Dynamic learning and strategic communication 645

gle distribution supported on the unit interval. If first-period posterior distributions are
binary, then there is no fully informative equilibrium for any positive bias. The reason
is that states around the point with zero density have to be linked to the states with pos-
itive densities. This creates an imbalance between posterior probabilities of posterior
states which eliminates the expert’s uncertainty in the first period and hence forces her
to distort information. The problem can be relaxed by specifying a discrete first-period
distribution which generates multiple posterior states such that at least two posterior
states are (approximately) equally likely. In this case, the expert in the first period
remains sufficiently uncertain, so that the fully informative equilibrium is restored.
Also, our model does not impose restrictions on information structures feasible to
the principal. As a result, the first-best outcome is an upper bound that the principal can
achieve by having access to any learning protocol. Thus, another question for further
research is to investigate benefits of dynamic learning under natural restrictions on
information structures which are feasible to the principal. For addressing this question,
our paper suggests that in addition to the precision of information structures, the key
factor for efficient communication is the structure of posterior sets, in particular, its
cardinality. Also, as shown by Ivanov (2010a), if the expert’s information is imperfect,
the possibility of generating discrete posterior distributions improves communication
in a single-period model compared to the CS setup with the perfectly informed expert.

Acknowledgments I am grateful to Ricardo Alonso, Dirk Bergemann, Archishman Chakraborty, Kalyan


Chatterjee, Ettore Damiano, Péter Esö, Maria Goltsman, Seungjin Han, Johannes Hörner, Sergei Izmalkov,
Andrei Karavaev, Vijay Krishna, Wei Li, Tymofiy Mylovanov, Gregory Pavlov, Vasiliki Skreta, Joel Sobel,
Dezsö Szalay, and participants of CETC 2011, CEA 2011, 22nd Game Theory Conference in Stony Brook,
WZB conference on Cheap Talk and Signaling in Berlin, and a seminar at the New Economic School for
valuable suggestions and discussions on different versions of the paper. I am also indebted to the associate
editor and two anonymous referees for helpful and constructive comments. Misty Ann Stone provided
invaluable help with copy editing the manuscript.

Appendix
 ∗ ∗ 
A perfect Bayesian equilibrium includes the learning protocol  I1 , I2 (m 1, I1 ) , the
decision rule a ∗ (.), the belief μ∗ (.), and the expert’s strategy σ1∗ (.) , σ2∗ (.) , such that
μ∗ (.) is consistent with the players’ strategies and the strategies satisfy the following
conditions.
(1) Given any {I1 , I2 , m 1 , m 2 } and μ∗2 (.) , a ∗ maximizes the principal’s payoff:

a ∗ ∈ arg max E θ U (y, θ ) |μ∗2 (m 1 , m 2 , I1 , I2 ) .
y∈R

(2) Given a ∗ (.) and any {I1 , I2 , s1 , s2 , m 1 } , σ2∗ maximizes the expert’s second-period
payoff:25
  
m 2 ∈ arg max E θ V a ∗ (m 1 , x, I1 , I2 ) , θ, b |I1 , I2 , s1 , s2 .
x∈M

25 For simplicity, this definition implies that the expert does not randomize across messages. None of our
results changes if we allow for mixed expert’s strategies.

123
646 M. Ivanov

 
(3) Given σ1∗ (.) , σ2∗ (.) , a ∗ (.) , μ∗1 (.), and any {m 1 , I1 }, the information structure
I2∗ (m 1 , I1 ) maximizes the principal’s interim payoff:

    
I2∗ ∈ arg max E θ,s1 ,s2 U a ∗ m 1 , σ2∗ (I1 , I2 , s1 , s2 , m 1 ) , I1 , I2 , θ |I1 , I2 , m 1 .
I2 ∈I

 
(4) Given a ∗ (.) , σ2∗ (.) , I2∗ (.) , and any {I1 , s1 } , σ1∗ maximizes the expert’s interim
payoff:

   
m 1 ∈ arg max E θ,s2 V a ∗ x, σ2∗ (I1 , I2 (x, I1 ) , s1 , s2 , x) , I1 , I2 (x, I1 ) ,
x∈M
θ, b) |I1 , I2 (x, I1 ) , s1 ] .

 
(5) Given σ1∗ (.) , σ2∗ (.) , a ∗ (.) , I2∗ (.) , the information structure I1∗ ∈ I maximizes
the principal’s ex-ante payoff.

Now, we derive the expressions for ps and psc . Split  into two subinter-
vals, [0, s̄] and [s̄, 1]. Then, split the subinterval [0, s̄] into N subintervals of the
equal length  = Ns̄ , and consider two-interval sets Ss, = Ss, 1 ∪ Ss,
2 =
[s, s + ] ∪ [ϕ (s) , ϕ (s + )], where s = 0, N , ..., s̄ − N , Ss, = [s, s + ], and
s̄ s̄ 1
2 = [ϕ (s) , ϕ (s + )]. Then, we have
Ss,

1 ]
Pr[θ ∈ Ss, 1
 (F (s +)− F (s))
ps () =  = ,
Pr θ ∈ Ss, 1
 (F (s + )− F (s))+ 1 (F (ϕ (s + ))− F (ϕ (s)))

f (s)
which gives ps = lim ps () = f (s)+ f (ϕ(s))ϕ  (s) and ps = →0
c lim (1 − ps ()) =
→0
1 − ps . Since f (s) > 0, f (ϕ (s)) 
> 0, and ϕ (s) > 0 for all s ∈ [0, s̄], then ps > 0
and psc > 0.

Proof of Theorem 1 Consider the principal’s decision rule defined by (4). Upon
observing
 s = 1, the expert prefers to reveal it truthfully in both periods, since
V a p (1) , 1, b ≥ V (a, 1, b) , a ∈ A. After observing a signal s ∈ [0, s̄) in the
first period, the expert’s optimal strategy is either to learn θ in the second
 period by
reporting s truthfully and inducing an action a ∈ a p (s) , a p (ϕ (s)) or to induce the
optimal interim decision y1 (s, b) ∈ Y1 (s, b).
Consider the expert’s second-period incentive-compatibility constraints (7)–(8)
given truthful reporting in t = 1. If b = 0, then a p (s) = y (s, 0) , ∀s. Since
a p (s) = a p (ϕ (s)) , s ∈ [0, s̄], then

   
V a p (s) , s, 0 − V a p (ϕ (s)) , s, 0 = V (y (s, 0) , s, 0) − V (y (ϕ (s) , 0) , s, 0) > 0,
s ∈ [0, s̄] , and
   
V a p (ϕ (s)) , ϕ (s) , 0 − V a p (s) , ϕ (s) , 0 > 0, s ∈ [0, s̄] .

123
Dynamic learning and strategic communication 647

By the continuity of V (a, s, b) , a p (s), and ϕ (s) in (a, s, b) and the compactness of
[0, s̄], there are bs > 0 and bϕ > 0, such that26
   
V a p (s) , s, b ≥ V a p (ϕ (s)) , s, b , s ∈ [0, s̄] , b ≤ bs , and
   
V a p (ϕ (s)) , ϕ (s) , b ≥ V a p (s) , ϕ (s) , b , s ∈ [0, s̄] , b ≤ bϕ .
  
The expert’s interim payoffs E V a p (θ ) , θ, b |s and E [V (y1 (s, b) , θ, b) |s] from
truthtelling in both periods and inducing y1 (s, b), respectively, are
      
E V a p (θ ) , θ, b |s = ps V a p (s) , s, b + (1 − ps ) V a p (ϕ (s)) , ϕ (s) , b ,
and E [V (y1 (s, b) , θ, b) |s] = max ps V (a, s, b) + (1 − ps ) V (a, ϕ (s) , b) .
a∈A

Suppose b = 0. Then, y (ϕ (s) , 0) = a p (ϕ (s)) = a p (s) = y (s, 0) implies


y1 (s, 0) = y (s, 0) and/or y1 (s, 0) = y (ϕ (s) , 0). Because 0 < ps < 1, we have

E [V (y1 (s, 0) , θ, 0) |s] = ps V (y1 (s, 0) , s, 0) + (1 − ps ) V (y1 (s, 0) , ϕ (s) , 0)


< ps V (y (s, 0) , s, 0) + (1 − ps ) V (y (ϕ (s) , 0) , ϕ (s) , 0)
   
= ps V a p (s) , s, 0 + (1 − ps ) V a p (ϕ (s) , 0) , ϕ (s) , 0
  
= E V a p (θ ) , θ, 0 |s , s ∈ [0, s̄] .

Because A is an image of a continuous function a p (θ ) on the compact , it is compact


(Arkhangel’skii and Pontrjagin 2011).  By
 the continuity of ϕ (s) , ps , V (a, s, b), and
a p (s) in (a, s, b), it follows that E V a p (θ ) , θ, b |s and E [V (y1 (s, b) , θ, b) |s]
are continuous in (s, b) (Sydsæter et al. 2005). Hence, there is b y > 0 such that
  
E V a p (θ ) , θ, b |s ≥ E [V (y1 (s, b) , θ, b) |s] , s ∈ [0, s̄] , b ≤ b y .
 
Thus, there is the fully informative equilibrium if b ≤ b̄, where b̄ = min bs , bϕ , b y >
0. 

Proof of Lemma 1 Suppose there is the fully informative equilibrium for b̄ > 0, that
is,    
V a p (s) , s, b̄ − V a p (ϕ (s)) , s, b̄ ≥ 0, (11)
and
           
V y1 s, b̄ , s, b̄ = E V a p (θ ) , θ, b̄ |s − E V y1 s, b̄ , θ, b̄ |s
   
= ps V a p (s) , s, b̄ + (1 − ps ) V a p (ϕ (s)) , ϕ (s) , b̄
       
− ps V y1 s, b̄ , s, b̄ −(1− ps ) V y1 s, b̄ , ϕ (s) , b̄ ≥ 0,

26 The argument behind b > 0 and b > 0 is as follows. Consider a function H : S × R → R,


s ϕ +
which is continuous in (s, b) and H (s, 0) > 0 for all s ∈ S, where S is compact. Define β (s) =
inf {b ∈ R+ |H (s, b) = 0} and b H = inf s∈S β (s). We claim that b H > 0. By contradiction, suppose
∞ , such that β (s ) → 0. Because S is compact, then exists a
b H = 0. Then, there is a sequence {si }i=1 i
∞ → s̃, and β (s̃ ) → 0. The continuity of H (s, b) and H (s̃ , β (s̃ )) = 0
converging subsequence {s̃i }i=1 i i i
mean H (s̃, 0) = 0 that contradicts H (s̃, 0) > 0.

123
648 M. Ivanov

for all s ∈ [0, 1/2). We will prove that there exists the fully informative equilibrium
for any b < b̄.  
Because a p (ϕ (s)) > a p (s) and Vab  > 0, then Vb a p (s) , s, b − Vb
 
a p (ϕ (s)) , s, b < 0, s ∈ [0, 1/2). This implies that (11) holds for b < b̄:

     
V a p (s) , s, b − V a p (ϕ (s)) , s, b > V a p (s) , s, b̄
 
−V a p (ϕ (s)) , s, b̄ ≥ 0, b < b̄. (12)


Also, a p (ϕ (s)) > y (s, b) , s ∈ [0, 1/2), b < b̄. Otherwise, if a p (ϕ (s)) ≤ y (s, b) ,

 a p (s) < a p (ϕ (s)) ≤ y (s, b)  and
then  Va (a, θ, b) > 0, a < y (θ, b) result in
V a p (s) , s, b < V a p (ϕ (s)) , s, b .
Since Vab  > 0, then Vb (a, s, b) is increasing in a. Because a p (s) <
y (s, b) , ∀s, b > 0, we have:

    
Vb a p (s) , s, b < Vb (y (s, b) , s, b) ≤ 0, andE Vb a p (θ ) , θ, b |s
   
= ps Vb a p (s) , s, b + (1 − ps ) Vb a p (ϕ (s)) , ϕ (s) , b < 0.

Because y1 (s, b) > a p (0), we need to consider two cases: y1 (s, b) < a p (1) and
y1 (s, b) = a p (1). First, let y1(s, b) < a p (1). ∗

 In this case, y1 (s, b) = y1 (s, b)
and V (y1 (s, b) , s,
 b) = V y1 (s, b) , s, b . This implies Vb (y1 (s, b) , s, b) =
Vb y1∗ (s, b) , s, b ≤ 0. Second, let y1 (s, b) = a p (1), which
 means  y1 (s, b) >
a p (ϕ (s)) > a (s). Since V  > 0, then V  a (s) , s, b < V  a (1) , s, b
 p  ab  b p  b p
and Vb a p (ϕ (s)) , ϕ (s) , b < Vb a p (1) , ϕ (s) , b . By the Envelope theorem,
 
db E [V (y1 (s, b) , θ, b) |s] = E Vb (y1 (s, b) , θ, b) |s results in:
d

dV (y1 (s, b) , s, b)    


= ps Vb a p (s) , s, b + (1 − ps ) Vb a p (ϕ (s)) , ϕ (s) , b
db
− ps Vb (y1 (s, b) , s, b) − (1 − ps ) Vb (y1 (s, b) , ϕ (s) , b)
    
= ps Vb a p (s) , s, b − Vb a p (1) , s, b + (1 − ps )
    
× Vb a p (ϕ (s)) , ϕ (s) , b − Vb a p (1) , ϕ (s) , b < 0.

   
Thus, Vb (y1 (s, b) , s, b) ≤ 0 and V (y1 (s, b) , s, b) ≥ V y1 s, b̄ , s, b̄ ≥
0, b < b̄. 


Proof of Corollary 1 The decision y1∗ (s, b) is given by the first-order condition:

   
ps Va y1∗ (s, b) − β (b, s) , s + (1 − ps ) Va y1∗ (s, b) − β (b, ϕ (s)) , ϕ (s) = 0,

   
where Va y1∗ (s, b) − β (b, s) , s < 0 < Va y1∗ (s, b) − β (b, ϕ (s)) , ϕ (s) due
 > 0. Because V  (a − β (b, θ ) , θ ) = −V  (a − β (b, θ ) , θ ) β  (b, θ ) and
to Vaθ b  a  b
βb (b, s) ≥ 0, this implies −Va y1∗ (s, b) − β (b, s) , s βb (b, s) ≥ 0, and


123
Dynamic learning and strategic communication 649

    
E Vb y1∗ (s, b) , θ, b |s = − ps Va y1∗ (s, b) − β (b, s) , s βb (b, s) − (1 − ps )
 
×Va y1∗ (s, b) − β (b, ϕ (s)) , ϕ (s) βb (b, ϕ (s))
  
≥ −βb (b, ϕ (s)) ps Va y1∗ (s, b) − β (b, s) , s
 
+ (1 − ps ) Va y1∗ (s, b) − β (b, ϕ (s)) , ϕ (s) = 0,

  
where the
 inequality

  from βb (b,s)  ≥ βb (b, ϕ (s)) due to βbθ (b, θ ) ≤ 0.
follows
Thus, E Vb a p (θ ) , θ, b |s < 0 and E Vb (y1 (s, b) , θ, b) |s ≥ 0, which gives
     
Vb (y1 (s, b) , s, b) = E Vb a p (θ ) , θ, b |s − E Vb y1∗ (s, b) , θ, b |s < 0.




Proof of Theorem 2 According to (A2), there exists δ (b) > 0 such that  y (s, b) − a p
(s)| ≤ δ (b) , ∀s. By (A3), there exists δ̄ = δ̄ (b, δ (b)) > 0 such that
V (y (s, b) − δ (b) , s, b) ≥ V y (s, b) + δ̄, s, b and V (y (s, b) + δ (b) , s, b) ≥
 
V y (s, b) − δ̄, s, b , ∀s. Then, δ̄ ≥ δ (b). (Otherwise, if δ̄ < δ (b), then
 
V (y (s, b) − δ (b) , s, b) ≥ V y (s, b) + δ̄, s, b > V (y (s, b) + δ (b) , s, b) ≥
 
V y (s, b) − δ̄, s, b , where the first and the last inequalities follow from (A3), and
the second one follows from y (s, b) + δ̄ < y (s, b) + δ (b) and Va (a, s, b) < 0, a >
y (s, b) due to the strict pseudo-concavity of V (a, s, b). By the same property, we
have the contradiction y (s, b) − δ (b) > y (s, b) −δ̄, or δ (b) < δ̄.)
Put ϕ (θ ) = θ + d2 . By (A4), there exists d = d b, δ̄ > 0, such that

y (s, b) + δ̄ < y1∗ (s, ϕ (s) , b) < y (ϕ (s) , b) − δ̄, ∀s. (13)




Construct the learning protocol as follows:

θ if θ ∈ [kd, kd + d2 ), k = 0, 1, . . . ,
s1 (θ ) =
ϕ −1 (θ ) otherwise, and

⎨θ if θ = m 1 or θ = ϕ (m 1 ) , and m 1 ∈ [kd, kd + d2 ),
s2 (θ, m 1 ) = k = 0, 1, . . . ,

s0 ∈
/ R+ otherwise.

Consider the expert’s second-period incentive-compatibility constraints given truth-


ful reporting of s in the first period and observing θ = s < ϕ (s) in the second period.
First, we have

a p (ϕ (s)) ≥ y (ϕ (s) , b) − δ (b) ≥ y (ϕ (s) , b) − δ̄ > y (s, b) + δ̄, ∀s,

where the first inequality follows from (A2) and the third inequality follows from (13).
This gives
     
V a p (s) , s, b ≥ V y (s, b) + δ̄, s, b > V a p (ϕ (s)) , s, b , ∀s, (14)

123
650 M. Ivanov

where the first inequality follows from (A3) if a p (s) ≤ y (s, b) and from Va (a, s, b) <
0, a > y (s, b) if a p (s) > y (s, b). This property also implies the second inequality.
Similarly, if the expert observes θ = ϕ (s) in the second period, then

a p (s) ≤ y (s, b) + δ (b) ≤ y (s, b) + δ̄ < y (ϕ (s) , b) − δ̄, ∀s,

where the first inequality follows from (A2) and the third inequality follows from (13).
This gives
     
V a p (ϕ (s)) , ϕ (s) , b ≥ V y (ϕ (s) , b)− δ̄, ϕ (s) , b > V a p (s) , ϕ (s) , b , ∀s,
(15)
where the first inequality follows from (A3) if a p (ϕ (s)) ≥ y (ϕ (s) , b) and from
Va (a, ϕ (s) , b) > 0, a < y (ϕ (s) , b) if a p (ϕ (s)) < y (ϕ (s) , b). This property also
implies the second inequality.
Consider now the expert’s interim payoff in the first period. First, note that
     
V a p (s) , s, b ≥ V y (s, b) + δ̄, s, b > V y1∗ (s, ϕ (s) , b) , s, b , ∀s,

where the first inequality follows from (14), and the second inequality follows from
(A4) and Va (a, s, b) < 0, a > y (s, b). Similarly, we obtain
     
V a p (ϕ (s)) , ϕ (s) , b ≥ V y (ϕ (s) , b)− δ̄, ϕ (s) , b >V y1∗ (s, ϕ (s), b), s, b , ∀s,

where the first inequality follows from (15), and the second inequality follows from
(A4) and Va (a, ϕ (s) , b) > 0, a < y (ϕ (s) , b). As a result,
   
V a p (s) , s, b + V a p (ϕ (s)) , ϕ (s) , b
   
> V y1∗ (s, ϕ (s) , b) , s, b + V y1∗ (s, ϕ (s) , b) , ϕ (s) , b . (16)

Note that (14)–(15) do not depend on ps . Also, the optimal decision in the case of
deviation from truthtelling in the first period,
   
d 1 − ps
y1 s, s + , b = arg max ps V (y, s, b) + V (y, ϕ (s) , b) ,
2 a∈A ps

is continuous in ps > 0 and y1 (s, ϕ (s) , b) = y1∗ (s, ϕ (s) , b) if ps = 21 . Because


(16) holds strictly, there exists an ε–neighborhood of 1 for each (s, b), such that if it
contains 1−psps = f (ϕ(s))
f (s) , then

  
E V a p (θ ) , θ, b |s
 
  f (ϕ (s))  
= ps V a p (s) , s, b + V a p (ϕ (s)) , ϕ (s) , b
f (s)
 
f (ϕ (s))
> ps V (y1 (s, ϕ (s) , b) , s, b) + V (y1 (s, ϕ (s) , b) , ϕ (s) , b)
f (s)
= E [V (y1 (s, ϕ (s) , b) , θ, b) |s] .

123
Dynamic learning and strategic communication 651

Proof of Corollary 2 We construct the learning protocol as that in the proof of


Theorem 2 with the distance d2 between the posterior states, which is determined
below. Given a signal s ∈ [kd, kd + d2 ), k ∈ N0 in the first period, truthtelling in both
periods provides the payoff to the expert:
 
   d
E V a p (θ ) , θ, b |s = − ps β 2 (b, s) − psc β 2 b, s + ≥ −δ 2 .
2
 
Inducing the optimal interim decision y1 s, s + d2 , b = ps (s + β (b, s)) +
  
psc s + d2 + β b, s + d2 provides the payoff:
         2
d c d d
E V y1 s, s + , b , θ, b |s = − ps ps + β b, s + − β (b, s) .
2 2 2
 
Because a p s + d2 − a p (s) = s + d2 − s = d2 and V (a, s, b) is symmetric in a
around y (s, b) = s + β (s, b), the second-period incentive compatibility constraints
are given by:  
d d
ap s + − a p (s) = ≥ 2β (b, s) . (17)
2 2
Since β (b, s) ≤ δ, (17) are satisfied if d ≥ 4δ. Also, (17) implies d
2 − β (b, s) ≥
β (b, s) > 0 and
       2
d c d
E V y1 s, s + , b , θ, b |s < − ps ps −δ
2 2
   2
f (s) f s + d2 d
= −  2 2 − δ .
f (s) + f s + d2
   
Suppose f (s) ≥ f s + d2 . (The case of f (s) < f s + d2 is symmetric.) If f (θ )
satisfies condition (A5) for some α > 0 and d > 0, then
    d d
f (s) f s + d2 f s + d2 1 e−α 2 e−α 2
  2 =    2 ≥    2 ≥ ,
f (s) + f s + d2 f (s) f s+ d2 f s+ d2 4
1 + f (s) 1 + f (s)

      d  2
d e−α 2 d
and E V y1 s, s + , b , θ, b |s < − − δ = φ (d, α) ,
2 4 2
−2−αδ
where φ (d, α) attains the minimum φ (d ∗ (α) , α) = − e α 2 at d ∗ (α) = α4 + 2δ
for any α > 0. Equivalently, each d > 2δ is the minimizer of φ (x, α) over x for
α ∗ (d) = d−2δ
4
> 0.

 
Because (A5) holds for α = zδ and d = d ∗ (α) = 2δ z2∗ + 1 , where z ∗ is a
z z∗ αδ
unique solution to the equation z = e−1− 2 , we have αδ = z ∗ = e−1− 2 = e−1− 2
and e−2−αδ = α 2 δ 2 . This results in

123
652 M. Ivanov

   e−2−αδ  
E V a p (θ ) , θ, b |s ≥ −δ 2 = − = φ d ∗ (α) , α
  α
2
  
d
> E V y1 s, s + , b , θ, b |s .
2
 
Finally, (17) holds because z ∗ 0.314 < 1 implies d = 2δ 2
z∗ + 1 > 4δ. 


References
Alonso R, Matouschek N (2008) Optimal delegation. Rev Econ Stud 75:259–293
Ambrus A, Lu S (2013) Robust almost fully revealing equilibria in multi-sender cheap talk. Working paper
Argenziano R, Severinov S, Squintani F (2013) Strategic information acquisition and transmission. Working
paper
Arkhangel’skii A, Pontrjagin L (2011) General topology I: basic concepts and constructions dimension
theory. Springer, Berlin
Austen-Smith D (1994) Strategic transmission of costly information. Econometrica 62:955–963
Battaglini M (2002) Multiple referrals and multidimensional cheap talk. Econometrica 70:1379–1401
Bergemann D, Pesendorfer M (2007) Information structures in optimal auctions. J Econ Theory 137:580–
609
Blume A, Board O, Kawamura K (2007) Noisy talk. Theor Econ 2:395–440
Board S (2009) Revealing information in auctions: the allocation effect. Econ Theory 38:125–135
Crawford V, Sobel J (1982) Strategic information transmission. Econometrica 50:1431–1451
Damiano E, Li H, (2007) Information provision and price competition. Working paper
Dessein W (2002) Authority and communication in organizations. Rev Econ Stud 69:811–838
Esö P, Fong Y (2008) Wait and see: a theory of communication over time. Working paper
Esö P, Szalay D (2010) Incomplete language as an incentive device. Working paper
Fischer P, Stocken P (2001) Imperfect information and credible communication. J Account Res 39:119–134
Golosov M, Skreta V, Tsyvinski A, Wilson A (2014) Dynamic strategic information transmission. J Econ
Theory 151:304–341
Goltsman M, Hörner J, Pavlov G, Squintani F (2009) Mediation, arbitration and negotiation. J Econ Theory
144:1397–1420
Green J, Stokey N (2007) A two-person game of information transmission. J Econ Theory 135:90–104
Ivanov M (2010a) Informational control and organizational design. J Econ Theory 145:721–751
Ivanov M (2010b) Communication via a strategic mediator. J Econ Theory 145:869–884
Ivanov M (2013) Information revelation in competitive markets. Econ Theory 52:337–365
Ivanov M (2015) Dynamic information revelation in cheap talk. BE J Theor Econ (forthcoming)
Johnson J, Myatt D (2006) On the simple economics of advertising, marketing, and product design. Am
Econ Rev 93:756–784
Kartik N, Ottaviani M, Squintani F (2007) Credulity, lies, and costly talk. J Econ Theory 134:93–116
Klein N, Mylovanov T (2011) Will truth out? An advisor’s quest to appear competent. Working paper
Kovác E, Mylovanov T (2009) Stochastic mechanisms in settings without monetary transfers: the regular
case. J Econ Theory 144:1373–1395
Krishna V, Morgan J (2001a) A model of expertise. Q J Econ 116:747–775
Krishna V, Morgan J (2001b) Asymmetric information and legislative rules: some amendments. Am Polit
Sci Rev 95:435–452
Krishna V, Morgan J (2004) The art of conversation: eliciting information from experts through multi-stage
communication. J Econ Theory 117:147–179
Krishna V, Morgan J (2008) Contracting for information under imperfect commitment. RAND J Econ
39:905–925
Lewis T, Sappington D (1994) Supplying information to facilitate price discrimination. Int Econ Rev
35:309–327
Li M, Madaràsz K (2008) When mandatory disclosure hurts: expert advice and conflicting interests. J Econ
Theory 139:47–74
Melumad N, Shibano T (1991) Communication in settings with no transfers. RAND J Econ 22:173–198

123
Dynamic learning and strategic communication 653

Morgan J, Stocken P (2003) An analysis of stock recommendations. RAND J Econ 34:183–203


Myerson R (1991) Game theory: analysis of conflict. Harvard University Press, Cambridge, MA
Ottaviani M, Squintani F (2006) Naive audience and communication bias. Int J Game Theory 35:129–150
Radner R (1993) The organization of decentralized information processing. Econometrica 62:1109–1146
Roberts J (2004) The modern firm: organizational design for performance and growth. Oxford University
Press, New York
Shaked M, Shanthikumar JG (2007) Stochastic orders. Springer, New York
Sydsæter K, Strøm A, Berck P (2005) Economists’ mathematical manual. Springer, Berlin

123

Common questions

Powered by AI

The protocol accommodates the expert's private bias by implementing a structure in which the precision of acquired information and the available sets of actions are affected by bias intensity. It uses a dynamics of binary posterior states, which are sensitive to deviations in reporting and prevent full access to biaised manipulations. The protocol indicates that unbounded communication efficiency depends on the expert's truthful reporting and precisely partitioned information sets, hence mitigating the bias effects through the mechanism of trade-offs in actions .

The learning protocol restricts manipulation by decomposing communication into a continuum of cheap-talk rounds about binary posteriors and structuring two key experiments. The first experiment isolates the true state with an irrelevant complement, ensuring that only truthfully reported information in the first period influences the second period's learning. Hence, any local distortions in information lead to significant informational losses for the expert . In a babbling equilibrium, the principal disregards the expert's messages if deviations are suspected, which further discourages manipulation .

Achieving a fully informative equilibrium implies that the expert reveals the true state through truthful communication across all periods. The implications include promoting strategic decisions by the principal that align with accurate information, thus maximizing payoff efficiency without deviations. However, the feasibility hinges on maintaining structured posterior states and balancing informational benefits against bias-induced distortions, leading to enhanced negotiation and decision-making accuracy within bounded rational conditions .

Truthtelling in the first period is crucial because it establishes a particular binary distribution of posteriors, which the principal can use to verify the expert's subsequent reports. This limits the expert’s ability to manipulate information in the second period and balances the trade-off between informational benefits from truthful reporting and possible outcomes from distorting reports. Therefore, it sustains efficient communication and helps in achieving a fully informative second period .

The document suggests that the essential factors for the effective functioning of the learning protocol are the maximal intensity of the expert’s bias and the dependence of the expert’s payoff on the principal’s decisions in different states. It shows that perfectly informative communication is feasible even for an arbitrarily large bias, provided these factors are accounted for, without necessarily relying on other assumptions like concavity and supermodularity of payoff functions or the principal's knowledge of the expert's bias .

The structure and cardinality of posterior sets are crucial because they determine the degree of uncertainty and thus the accuracy of the expert's reports. The model shows that the first-period structure can generate non-binary posterior states that are effective for distributions with zero density at some points, thereby restoring a fully informative equilibrium. The cardinality of these sets prevents absolute certainty that could lead to report distortion. Hence, these elements are key factors alongside information structure precision in achieving efficient communication .

The trade-offs involve balancing future informational benefits from truthful reporting against potential immediate gains from distorting information. The expert risks losing the value of accurate updates and optimal decision flexibility if reporting is inconsistent. Truthtelling assures improved second-period insights while manipulation may lead to babbling equilibria or ignored communication, thus affecting strategic outcomes negatively .

Modifying information structures, such as making them partitional or adjusting the quality of updates, affects communication outcomes by either constraining or enhancing the expert's incentive-compatible strategies. An increase in structural sensitivity may tighten constraints and improve clarity in communication, thus achieving more truthful insights. Conversely, a mismatch or imbalance could lead to distortions, particularly at distributions with zero density, thereby subverting full information extraction .

The principal's ability to interpret binary posterior distributions is significant because it acts as a verification mechanism for the expert’s reports. By accurately understanding the binary outcomes, the principal can cross-verifying the consistency of the expert's claims across periods, thereby deterring the expert from false reporting. This verification minimizes the risk of misrepresentation and manipulation by aligning the expert's actions with verifiable outcomes .

The dynamic protocol differs from traditional mechanisms in that it does not require the principal's commitment to the expert's information quality. It is deterministic and decomposes communication into binary posteriors and distinct state experiments, unlike mechanisms that maintain strict assumptions on payoff structures and information fidelity. This protocol emphasizes a more flexible structure based on the expert's bias intensity and subsequent information updating .

You might also like