Selling Data
Selling Data
Carlos Segura-Rodriguez†
Abstract
* I am grateful to the coeditor and two anonymous referees for their constructive suggestions.
I also would like to express my gratitude to George Mailath, Annie Liang, Rakesh Vohra, Andrew
Postlewaite, Aislinn Bohren and Rohit Lamba for their insightful comments and support through-
out this project. I also thank Dirk Bergemann and Nima Haghpanah for their valuable input.
Thanks as well to my fellow classmates, especially Ashwin Kambhampati, Youngsoo Heo, Joao
Granja, and Paolo Martellini for listening and providing help during the project. I also gratefully
acknowledge all the participants in the UPenn Micro Theory Lunch, the UPenn Micro Theory
Seminar, NASMES 2019 and SEA 2019 for their feedback. All errors are my own.
† Central Bank of Costa Rica, Department of Economic Research, Av 0 and 1, St 2 and 4, San
1 Introduction 1
1.1 Literature Review . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3
2 Model 6
2.1 Elliptical Distributions . . . . . . . . . . . . . . . . . . . . . . . . . . . 8
3 Value of Information 10
3.1 Linear Mechanisms . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
B Proofs 33
Markets for data are developing quickly worldwide due to the ease of gathering
information from electronic commerce and on-line surfing. In these markets, data
brokers (sellers) are facing new challenges since data’s core properties are differ-
ent from most goods. First, data is intrinsically multidimensional: data brokers
observe multiple characteristics of each individual. Second, data is easy to trans-
form and summarize in statistics, and these statistics are highly substitutable. In
this context, it is important to understand the optimal way for a data monopo-
list seller to present and price data when buyers have private information about
which unknown consumer characteristic (willingness to pay for a product, credit-
worthiness, political preferences,...) the buyer (firm, financial institution, political
party,...) wants to forecast and the value that she assigns to this knowledge.1 I
introduce a novel model in which a seller sells multi-attribute data and chooses a
menu of statistics and prices to differentiate among buyers in both dimensions.
As an example, consider a firm that plans to introduce a new product to the
market and wants to identify which consumers will react positively to its mar-
keting efforts, in order to better target its marketing campaign. The data broker
believes the firm is either a casual or a fine-dining restaurant, but this is private
information of the buyer. Suppose the consumer’s income and the distance from
the consumer’s home to the restaurant are observable. Further, assume that the
restaurant is interested in knowing how much the consumer is willing to pay for
one of its meals, and the consumer’s willingness to pay is, in both cases, posi-
tively related with his income and negatively related with the distance he lives
from the restaurant. Due to the substitutability of information, any statistic that
is increasing in income and decreasing in distance will be valuable for both types
of firms, so that if any offered statistic is sufficiently cheap both of them will buy
it. To avoid this issue, the data broker needs to carefully design statistics and
1I use male pronouns for the seller (data broker) and female pronouns for the buyer (firm).
My paper contributes to a growing body of literature that studies the optimal way
in which a monopolist should sell information. The work that comes closest to my
approach is Bergemann, Bonatti, and Smolin (2018). They study an environment
5 Some examples in the literature show that, in the absence of conflict between players, using
only linear statistics is not always optimal (see, for instance, Witsenhausen (1968)).
6 Since the optimal statistics are a linear combination of the buyers’ preferred statistics, the set
that the buyer can take after buying information is sufficiently rich.
2 Model
Data brokers (sellers) sell consumer data to firms (buyers). I study how a monop-
olist seller, who wants to maximize profits, should sell the multi-attribute data
he can access when he is uncertain about the buyer’s motives for purchasing it.
By assuming that consumer’s characteristics come from the same distribution, I
focus on the problem of selling data of a representative consumer.
Buyer The seller knows there is a buyer that wants to forecast the value of an
unknown consumer characteristic in the vector θ = (θ1 , θ2 , . . . , θn ). For simplicity,
I assume that each buyer wants to forecast a unique characteristic θi ∈ θ. In the ex-
ample in the introduction, the vector θ corresponds to the consumer’s willingness
to pay for a meal in a casual or in a fine-dining restaurant.
The buyer’s preferences are represented by a two-dimensional type (i, v) ∈
{1, . . . , n} × R++ where i is her “information type” and v is her “valuation type”.
Type (i, v) takes an action a ∈ R to minimize
that is, type (i, v) is interested in making the most accurate forecast, a, of con-
sumer’s characteristic θi according to a quadratic loss function.9 Finally, the
buyer’s payoff is quasilinear in money; it is equal to the forecast loss minus the
9 Appendix A presents a microfoundation for this assumption. It shows that a monopolist
who is uncertain about the intercept of a linear demand, faces a profit loss that is quadratic with
respect to the forecast of the intercept that the monopolist uses in his decision-making process.
Prior information It is assumed that the firm’s prior over the tuple ( X, θ, ϵ)
corresponds to an elliptical distribution with finite first and second moments,
which is denoted by F. Further, the marginal distributions over ( X, θ ) and θ,
F(X,θ ) and Fθ exist; and that the conditional distribution of X given θ, FX |θ , exists.
Moreover, without loss of generality I assume that all variables have zero mean
with Var ( X ) = Ik and and Var (ϵ) = [Link] information environment is common
knowledge to both the data broker and the firm.
10 Babaioff, Kleinberg, and Paes Leme (2012) show that the seller might be able to use his private
information to extract surplus from the buyer beyond what he is able to capture by using direct
mechanisms. I thank an anonymous referee for this observation.
Timing The game between the buyer and seller proceeds as follows
3. The buyer reports her type to the seller and receives the promised statistic
at the promised price.
In this section, I briefly introduce the family of elliptical distributions and describe
its main properties. A random variable with an elliptical distribution whose first
two moments are finite and that assumes a density f has a density given by
Claim 1 Suppose the joint distribution of (θ, X, ϵ) is jointly elliptical. Then for any
non-zero vector L and/or non-zero constant ℓ,
This property is key. Claim 1 implies that the buyer’s optimal forecast corre-
sponds to the conditional expectation of the variable she is interested in given the
statistic she buys. Therefore, this property implies that if the buyer buys a linear
statistic his optimal action is a linear transformation of this statistic. Theorem 2
uses this fact to show that it is optimal for the seller to only offer linear statistics,
simplifying the analysis.
12 Because of these properties, the family of elliptical distributions have been used in multivari-
ate analysis (Anderson, 2003), portfolio choice (Chamberlain, 1983; Owen and Rabinovitch, 1983)
and information economics (Mailath and Nöldeke, 2008; Deimen and Szalay, 2019; Ball, 2019).
I start the analysis by calculating the maximum price that type (i, v) is willing
to pay for any statistic ψ. Since the firm utility is quadratic, type (i, v)’s optimal
forecast equals the conditional expectation of the characteristic she is interested
in, θi , given the statistic ψ. With this forecast, type (i, v)’s expected loss is equal
to v times the negative of the expected conditional variance of θi given ψ.13
Claim 2 If the data broker offers to type (i, v) the statistic ψ at the price p, her optimal
forecast is a∗ = EX,ϵ (θi | ψ), and her ex-ante expected utility is equal to
This result allows me to define the incentive compatibility and individual ra-
tionality constraints. The mechanism (ψ, p) is incentive compatible (IC) if
−vEX,ϵ [Var (θi | ψ(i, v))] − p(i, v) ≥ −vEX,ϵ [Var (θi | ψ( j, v′ ))] − p( j, v′ ) ∀(i, v), ( j, v′ ),
−vEX,ϵ [Var (θi | ψ(i, v))] − p(i, v) ≥ −vVar Fθ (θi ) ∀(i, v).
These constraints have two distinct properties that make them slightly differ-
ent from those commonly found in the literature. First, the IC constraints embed
two possible deviations: As usual the buyer can misreport her type. But when
the buyer lies, she can use the statistic she receives to make a forecast about the
variable she is interested in, which might not coincide with the variable that the
statistic was targeted to.
Second, the right-hand side of the IR constraint is not the same across types,
since when the buyer does not buy any information, her optimal forecast is her
prior expectation, which results in a payoff equal to the negative of her prior vari-
13 The proof of Claim 2 and all others that are not in the main text can be found in Appendix B.
10
Claim 2 implies that the buyer’s optimal forecast is equal to the conditional ex-
pectation given the statistic she buys. But, in this elliptical environment, if the
buyer buys all data X her conditional expectation is linear in X. Further, when
she buys a statistic that is a linear combination of the data and noise, her condi-
tional expectation is still linear in X and ϵ. Therefore, it is sensible to pay special
attention to the set of mechanisms in which the seller only offers linear statistics.
Definition 1 A mechanism is linear if all the statistics take the form ψ(i, v) = L(i, v) T X +
ℓ(i, v)ϵ, where L(i, v) ∈ Rk and ℓ(i, v) ∈ R.
Lemma 1 Type (i, v)’s willingness to pay for a linear statistic L T X + ℓϵ with L ̸= 0
and/or ℓ ̸= 0, corresponds to
( L T Cov(θi , X ))2
v .
L T L + ℓ2
There are two clarifications worth mentioning at this point. First, if ϵ where
multidimensional, the value of a linear statistic L T X + ℓ̂ T ϵ would be
( L T Cov(θi , X ))2
v ,
L T L + ℓ̂ T Σϵ ℓ̂
11
L2
γi
L0
L1
Figure 1: For a fixed direction, the distance between the origin and the blue locus
represents type (i, v)’s willingness to pay for a noiseless statistic L T X with vector
L pointing in that direction.
Further, when the statistic does not contain independence noise, the square of
the correlation has a nice geometrical interpretation: It corresponds to the square
of the cosine of the angle β between γi and L. Figure 1 plots type (i, v)’s will-
ingness to pay for any noiseless linear statistic. For a fixed direction, the distance
between the origin and the blue locus represents the function, vγiT γi cos2 ( β), i.e.,
type i’s willingness to pay for the statistic L T X with vector L pointing in that di-
14 This is analogous to the coefficient of determination R2 , which in regression analysis is used
as a goodness-of-fit measure and corresponds to the square of the correlation between the ob-
served and the regression predicted values.
12
The optimal mechanism has two properties that generalize those that have been
found in unidimensional environments that satisfy the single crossing condition.15
First, there is no rent at the bottom: at least one type pays his willingness to pay
for the statistic she receives. If this were not the case, the data broker can increase
his profits, without affecting the buyers’ incentives, by uniformly increasing all
prices. Second, there is no distortion at the top: at least one type receives an
statistic that for her is equivalent to receiving all data. The reasoning is subtler;
if there is no type that receives an undistorted statistic, the seller can offer a new
package containing all data at a price high enough that only one type is willing to
pay. Since this new product has a price that is too high for other types, all other
IC constraints are satisfied.
13
14
To gain intuition, suppose as in Bergemann, Bonatti, and Smolin (2018) that all
buyers are interested in forecasting the same variable, i.e., there is a unique in-
formation type. The critical difference with their environment is that in here all
buyer types share the same prior.
17 This is a result of the fact that in any strictly competitive game, if there is an equilibrium, then
max min = min max. This property, further, implies that the equilibrium payoffs are independent
of the timing in the game.
18 The argument relies on the existence of conflict between the two players. In fact, Witsen-
hausen (1968) introduced an example with normal distributions in which, in the absence of con-
flict, is not always optimal to use linear statistics and linear updating rules.
15
Proposition 1 Suppose there is a unique information type t who wants to forecast the
1− G ( v | t )
random variable θ and that v − g(v|t)
is non-decreasing in v. Then, there exists v∗
such that in the optimal mechanism all types v ≥ v∗ receive the statistic E[θ | X ] and
pay v∗ Var Fθ − EX,ϵ [Var (θ | X )] , and all types v < v∗ do not receive any information
Suppose now that there are multiple information types that share a common
valuation type. This case is more interesting than the previous one because now
the perception of the quality of a statistic by a buyer depends on which consumer
characteristic she wants to forecast. To simplify the analysis, I normalize, without
loss of generality, the valuation type to 1, and assume that γ1T γ1 ≥ γ2T γ2 ≥ . . . ≥
19 The proof of this proposition is omitted since it follows from standard arguments.
16
Assumption 1 Let β ij be the angle between γi and γ j . For any i ̸= j, β ij ∈ (0, π ).21
Theorem 3 Let I ⊆ {1, . . . , n} be the set of types to whom, in the optimal mechanism,
the data broker sells data. To each i ∈ I, the seller offers the statistic LiT X with
!
λ ji c ji
Li = Var ( X )−1 Cov(θi , X ) − ∑ Cov(θ j , X ) ,
j∈ I,j̸=i ∑ j∈ I,j̸=i ij
λ + µi
where λkℓ is the Lagrange multiplier associated with the constraint ICkℓ , µi is the La-
LkT Cov(θk ,X )
grange multiplier associated with the constraint IRi , and ckℓ = LℓT Cov(θℓ ,X )
.
The optimal statistics satisfy the following main property: When an IC con-
straint binds, the data broker decreases the correlation between the statistics of-
fered to those two types by modifying the statistic given to the mimicked type
20 If none of the inequalities are strict, the data broker can reveal all the information to all types
and charge a uniform price equal to γ1T γ1 .
21 If there are two types i and j with angle β = 0, they would only differ in the value they
ij
are willing to pay for any statistic, but they would have the same preferred statistic. If β ij = π
both types’ preferred statistics point in opposite directions, so buying any of their two preferred
statistics provide them exactly the same information.
17
incentive to mimic, which might not be those with the lowest willingness to pay for all data.
24 Uncorrelated noise has a proportional effect as time delay does in Stokey (1979). Fortunately
for the seller, in here, changing the direction of the statistic has a non-proportional effect.
18
Proposition 2 Suppose there exists a collection of small positive numbers {ϵij } such
2
that cos2 ( β ij ) ∥γi ∥2 ≤ γj + ϵij ∀i < j. Then the solution to the relaxed problem
2
that considers only the constraints with i < j such that cos2 ( β ij ) ∥γi ∥2 > γ j is the
solution to the seller’s problem.25
The previous proposition only imposes restrictions on the value that types
assign to each others’ preferred statistics. However, in this environment, it is very
important how types evaluate any statistic that is offered by the seller. To give
further structure, I define an order concept for types that depends only on the
angle between the types’ preferred statistics.
−1
Definition 2 Types are ordered if for any i < j < k, cos2 ( β ik ) < ∏kj= 2
i cos ( β j,j+1 ).
25 A straightforward corollary of the result is that in this case type 1 receives an undistorted
statistic and type n always receives zero rent.
19
Proposition 3 Suppose types are ordered. There exists ϵ > 0 such that if ∥γi ∥2 cos2 ( β i,i+1 )
< (1 + ϵ) ∥γi+1 ∥2 ∀i, then the solution to the relaxed problem that considers only the
local downward constraints is the solution to the seller’s problem.
Further, if the constraint from type i to type i + 1 binds, there exists ci,i+1 ∈ (−c̄i,i+1 , c̄i,i+1 )
with c̄i,i+1 < 1 such that the seller offers to type i + 1 the statistic γi+1 − ci,i+1 γi .26,27
To gain more intuition I consider the case with only two information types,
a special case in which types are ordered. In this case it cannot be that both IC
constraints bind simultaneously, since Theorem 1 implies that at least one type
must receive her optimal statistic. As type 1 is willing to pay more than type 2 for
all data it has to be that the constraint IC21 does not bind. Further, by Theorem 1
26 The direction of the distortion which is measured by the parameter ci,i+1 depends on the sign
ofγiT γi+1 . If this sign is positive ci,i+1 is positive and if it is negative ci,i+1 is negative.
27 The parameter c the multiplier for the constraint ICi,i+1 . And the value of
i,i +1 depends on
this multiplier might depend on the value of the multipliers for the local downward constraints
IC12 , . . . , ICi−1,i .
20
21
Corollary 1 Suppose the constraint IC12 binds. Then, there exist α̃ ∈ (0, 1) and c ∈
(0, α̃) such that the optimal mechanism is a linear mechanism with L1 = γ1 and
In the first case the constraint IR1 does not bind, but it does in the second one.
Figure 2: Optimal mechanism when γ1 = (6, 2.4) and γ2 = (4, 3) for different
values of α1 .
22
Finally , it is possible to analyze which data set the data broker would acquire
if he can choose from two data sets that have the same cost and for which both
types are willing to pay the same. As our intuition will suggest, Corollary 3 shows
that the seller prefers to acquire the data set that generates the smallest correlation
between the conditional expectations of θ1 and θ2 given X.
Corollary 3 Suppose there are two data sets X and X ′ such that for i ∈ {1, 2}, type i’s
willingness to pay for the statistic E[θi | X ] and for the statistic E[θi | X ′ ] is the same. If
ρ2 (E[θ1 | X ], E[θ2 | X ]) < ρ2 (E[θ1 | X ′ ], E[θ2 | X ′ ]), then the seller’ profits are larger
by acquiring X than by acquiring X ′ .
This nice intuition is true as a result of types being ordered. In the opposite case,
the problem is more difficult since the set of binding constraints is endogenous
and cannot be characterized. In particular, by decreasing the price of a statistic or
by distorting its direction, the seller might create an incentive for another type to
want to buy this statistic. This makes it difficult to obtain more general results.
In the following example, I consider a very simple case, in which data is two-
dimensional (X ∈ R2 ) and there are only three information types, to show how
general these difficulties are and some of their implications.
23
γ3 γ2
γ3 γ3
Example 1 Suppose data X is two-dimensional and there are three information types.
Figure 3 represents the cases that might occur under these assumptions. In the figure on
the left, types are ordered: Type 2 is to the right of 1, and 3 is to the right of 2. In the other
two figures, types are not ordered: In the one in the center, type 3 is located in the middle
and in the one on the right, type 1 is located in the middle. By modifying the magnitude
of γ2 and γ3 , we obtain all possible configurations for the seller’s problem.
Figure 4 presents which IC constraints bind for each of the cases above and for all
possible magnitudes of γ2 (v2 ) and γ3 (v3 ). I vary both v2 and v3 between 0 and 1, with
v3 ≤ v2 , and fix the magnitude of γ1 to be equal to 1.29
The figure has many interesting implications. First, the set of IC constraints that bind
varies a lot with v2 and v3 , and the limits of the regions defining those sets are non-linear.
Therefore, it is difficult to obtain closed form characterizations for those regions, but it
is easy to solve for them numerically due to the optimality of linear contracts. Second,
in many cases, some upward IC constraints bind, that is, constraints from types that are
willing to pay less for all data to those who are willing to pay more. This results from
the seller offering a price below the monopolist price for the mimicked type and/or from
modifying a statistic on the direction towards the type with lower willingness to pay’s
preferred statistic.
Furthermore, some sets of binding IC constraints generate contracts that exhibit some
properties we do not obtain in other environments. First, when both IC12 and IC32 bind
29 For this figure, I assume that the angle between any two consecutive vectors is π/6. The
patterns are very similar when considering other angles.
24
(IC13 and IC23 ) the seller, in some occasions, finds it optimal no to target an informative
statistic to type 2 (3, respectively), i.e., exclude this type from the mechanism. Interest-
ingly, the seller might not exclude the type with the lowest willingness to pay for all data,
but he excludes the type that other types have a higher incentive to mimic.
Second, which types are ”at the top” or ”at the bottom” is not completely determined
by types willingness to pay for all data. Consider the case when constraints IC13 and
IC21 bind. These constraints bind simultaneously when type 1 is in the middle, type 3’s
willingness to pay is low, and type 2’s willingness to pay is high. In this case, the seller
wants to reduce the price charged to type 1, so that she does not have an incentive to mimic
type 3. However, this reduction in p(1) creates an incentive for type 2 to mimic type 1.
To reduce this incentive, the seller modifies the direction of the statistic targeted to type 1
away from type 2’s preferred statistic. Therefore, in this case, type 2 is the only type that
receives an undistorted statistic, converting her into the type ”at the top”, even when she
does not have the largest willingness to pay for all data. Similarly, when only constraints
IC12 and IC32 bind and type 2 is not excluded, the only type who receives zero rent is
type 2, making her the type ”at the bottom”.
25
is nondecreasing in v.
I restrict the analysis by assuming that the seller can only offer linear mecha-
nisms, as defined before. To simplify notation, I let q jiv be the reduction of θ j ’s
forecast error variance after observing the linear statistic targeted to type (i, v);
T γ )2
( Liv
that is, q jiv = T
j
Liv Liv +ℓ2iv
≤ 1, since γiT γi = 1. This variable measures the quality,
as perceived by type j, of the statistic targeted to type i. This environment differs
from other multidimensional environments in that different types perceive the
30 In
a related environment, Smolin (2019) uses a direct approach to find the optimal mecha-
nism. However, the steps he needs to implement and the optimal mechanism are different from
the ones I use. Furthermore, the optimal mechanisms differ as well.
26
√
Theorem 4 Let h(q22v ) := cos2 ( β + cos−1 ( q22v )). There exist cutoffs v̂1 , v̂2 , and
31 As pointed out before in footnote 28, offering all data to each information type is never
optimal.
27
Figure 5: Optimal Mechanism when the angle between γ1 and γ2 is equal to 0.07̄π,
α1 = 0.85, v1 = v2 = 0, v̄1 = 1, v̄2 = 2.4, g(v | 1) = 15v14 and g(v | 2) = 0.83̄ for
v < 0.8 and g(v | 2) ∝ 4(2.4 − v)3 for v > 0.8.
In the optimal mechanism, a buyer with type (1, v) with v > v1∗ , still receives
a fully informative statistic. However, to reduce the incentive of type 1 to hori-
zontally mimic type 2, the seller decreases the quality of the statistic targeted to
types (2, v) by increasing the minimum valuation type v to whom he sells above
Myerson’s threshold, v2∗ , and by never offering an undistorted statistic to these
types. Further, the seller vertically differentiates between types (1, v) by offering
those types with v < v1∗ a distorted statistic, and might vertically differentiate
among types (2, v) by offering them a continuum of statistics.
A comment is in order. The theorem describes the quality of the statistic that
is targeted to each type, but it does not specify the actual statistics that are of-
32 Thisexpression is well defined, since the function h is infinitely differentiable and the inverse
of any of its derivatives exists.
28
29
Anderson, S. P., and L. Celik (2015): “Product line design,” Journal of Economic
Theory, 157, 517–526.
Balestrieri, F., S. Izmalkov, and J. Leao (2015): “The market for surprises: sell-
ing substitute goods through lotteries,” .
Bergemann, D., A. Bonatti, and A. Smolin (2018): “The design and price of
information,” American Economic Review, 108(1), 1–48.
Birbil, Ş. İ., J. B. G. Frenk, and G. J. Still (2007): “An elementary proof of the
Fritz-John and Karush–Kuhn–Tucker conditions in nonlinear programming,”
European journal of operational research, 180(1), 479–484.
30
Deimen, I., and D. Szalay (2019): “Delegated expertise, authority, and communi-
cation,” American Economic Review, 109(4), 1349–74.
Fang, K., S. Kotz, and K. Ng (1990): Symmetric multivariate and related distributions.
Chapman and Hall.
Federal Trade Commission (2014): “Data brokers: A call for transparency and
accountability,” Washington, DC.
Mailath, G. J., and G. Nöldeke (2008): “Does competitive pricing cause mar-
ket breakdown under extreme adverse selection?,” Journal of Economic Theory,
140(1), 97–125.
Maskin, E., and J. Riley (1984): “Monopoly with incomplete information,” The
RAND Journal of Economics, 15(2), 171–196.
Mussa, M., and S. Rosen (1978): “Monopoly and product quality,” Journal of
Economic theory, 18(2), 301–317.
31
Osborne, M. J., and A. Rubinstein (1994): A course in game theory. MIT press.
Owen, J., and R. Rabinovitch (1983): “On the class of elliptical distributions and
their applications to the theory of portfolio choice,” The Journal of Finance, 38(3),
745–752.
Pavlov, G. (2011): “Optimal mechanism for selling two goods,” The BE Journal of
Theoretical Economics, 11(1).
32
that is, the monopolist’s losses due to its uncertainty about the parameter ā cor-
respond to a quadratic loss function.
B Proofs
Proof of Claim 2
Let S be the set containing all realizations of a statistic ψ. For any functional
H : S → R,
E[θ |ψ], and the equality follows from the Law of Iterated Expectations and the
definition of conditional variance.
Proof of Lemma 1
Consider the family of functions k( L T X + ℓϵ) for k ∈ R. Since all variables
33
= Eθi [θi2 ] − 2kLT EX,ϵ [Eθi [θi X | LT X + ℓϵ]] + k2 LT EX,ϵ [Eθi [ XX T | LT X + ℓϵ]] L + k2 ℓ2 Eϵ [ϵ2 ]
where the first equality follows from ϵ being independent of X and θ, and he sec-
ond equality follows the Law of Iterated Expectations, the definition of variance
and covariance, and that Var ( X ) = Ik and Var (ϵ) = 1. By Claim 2, the conditional
expectation minimizes this expression. The First Order Condition with respect to
k implies that
L T Cov(θi , X )
k̂ = .
L T L + ℓ2
Plugging in this value of k, we obtain that the forecast variance is equal to
( L T Cov(θi , X ))2
Var (θi ) − .
L T L + ℓ2
The result follows from type (i, v)’s willingness to pay being equal to the differ-
ence between the prior and the posterior forecast variance times v.
Proof of Theorem 1
To prove part 1. suppose that all types receive a positive surplus bounded
away from zero and let P = infi,v {vEX,ϵ [Var (θi | (ψ(i, v)))] − p(i, v)} > 0. Then
the data broker can charge new prices p̃(i, v) = p(i, v) + P, so that all IC con-
straints are unaffected and the IR constraints are satisfied. This modification of
the mechanism clearly increases the seller’s profits.
To prove part 2. suppose by contradiction that in the optimal mechanism no
type receives all data or a statistic that for her is equivalent to receiving all data.
Let R(i, v) be type (i, v)’s information rent in this mechanism and let V (i, v) be
type (i, v)’s willingness to pay for all data. Define (i∗ , v∗ ) ∈ arg maxi,v V (i, v) −
34
35
If only the constraints IR1 and IR2 bind, the problem is straightforward be-
cause the seller can offer the optimal statistics to both types at the monopolist
prices. These optimal statistics correspond to the conditional expectations, which
are linear since the joint distribution of the ( X, θ, ϵ) is elliptical.
In Step 1-4, I assume that only constraints IC12 and IR2 bind, and show that
the optimal mechanism is still linear. In Step 5, I extend the analysis to include
the case in which constraints IC12 , IR1 and IR2 bind simultaneously.
Step 1: It is without loss of generality only considering recommendation
mechanisms.
In a recommendation mechanism, buyer i’s forecast, ai , is equal to the statistic
sold by the seller to her, ψ(i ). As Claim 2 establishes that ai = E[θi | ψ(i )( X = x )],
in a recommendation mechanism we would have ψ(i )( x ) = E[θi | ψ(i )( X = x )].
Considering only this kind of mechanisms is without loss of generality by an
argument similar to the one in Kamenica and Gentzkow (2011). First, Claim 2
implies that type i learns the same either by observing ψ(i ) or E[θi | ψ(i )( X )].
Moreover, type j cannot learn more about θ j by observing E[θi | ψ(i )( X )] than
by observing ψ(i ). To see this, let a( j, E[θi | ψ(i )( X )]) be type j’s forecast when
observing E[θi | ψ(i )( X )], and notice that type j can recover a( j, E[θi | ψ(i )( X )])
from ψ(i ) in a two-step estimation: first calculate E[θi | ψ(i )( X )] and them, from
it, create a( j, E[θi | ψ(i )( X )]).
Therefore, if the seller offers E[θi | ψ(i )] to type i ratter than ψ(i ) and charges
her the same price as before, type i’s incentives are unchanged and all ICji con-
straints are still satisfied.
Step 2: The mechanism design problem is equivalent to a sequential strictly
competitive game.
The problem is a sequential game in which the seller first chooses a mechanism
that satisfies all constraints and then each buyer type decides how to use the
36
which generates a payoff equal to zero for type 2. Further, as there is not IC
constraint pointing towards type 1, the seller offers all information to her, by
setting ψ(1) = E[θ1 | X ]. Thus, and as constraint IC12 binds, the seller charges to
type 1 the price
= E[(θ1 − a(1, 2, ψ(2)))2 ] − E[(θ1 − a(1, 1, ψ(1)))2 ] + Var (θ2 ) − E[(θ2 − a(2, 2, ψ(2)))2 ].
=Var (θ1 ) − E[(θ1 − a(1, 2, ψ(2)))2 ] − Var (θ2 ) + E[(θ2 − a(2, 2, ψ(2)))2 ]
Type 1 chooses her forecasts a(1, ·, ·) to maximize her payoff. However, the
only component of her forecasts that appears in her payoff is a(1, 2, ψ(2)). Notice
that whenever type 1 (2) buys the statistic ψ(1) (ψ(2)), by Claim 2, it is sequen-
tially rational for her to choose the optimal updating rule a(1, 1, ψ(1)) = E[θ1 |
ψ(1)] = E[θ1 | X ] ( a(2, 2, ψ(2)) = E[θ2 | ψ(2)]). With this mechanism and buyer’s
forecasts, the seller’s profits are equal to
α1 −E[Var (θ1 | X )] + E[(θ1 − a(1, 2, ψ(2)))2 ] + Var (θ2 ) − E[Var (θ2 | ψ(2))], 34
where α1 is the probability that the seller assigns to type 1. The chooses the
34 Notice that as a(1, 1, ψ(1)) = E[θ1 | X ] and a(2, 2, ψ(2)) = E[θ2 | ψ(2)] we have E[(θ1 −
a1 (1, 1, ψ(1)))2 ] = E[Var (θ1 | X )] and E[(θ2 − a2 (2, 2, ψ(2)))2 ] = E[Var (θ2 | ψ(2))].
37
Notice that for the buyer it is equivalent to choose a(1, 2, ψ(2)) to minimize G or
to choose a(1, 2, ψ(2)) to maximize her payoff, and for the seller it is equivalent to
choose ψ(2) to maximize his profits or to maximize G. Therefore, the functional G
defines a sequential stochastic strictly competitive game equivalent to the mech-
anism design problem, in which type 1 chooses a(1, 2, ψ(2)) to minimize G and
the seller chooses ψ(2) to maximize G.
Step 3: The sequential strictly competitive game has an equilibrium in lin-
ear strategies.
I divide the proof for this step in three parts. First, I show that best responses
to linear strategies, whenever they exist, are linear. Then, I show that in the set of
linear mechanisms, it is without loss to focus on linear mechanisms whose vector
of coefficients magnitude is equal to 1. As this set is compact, an application of
the Maximum Theorem implies that the seller’s best response exists when he is
restricted to choose a linear mechanism. Finally, an application of the Minmax
Theorem implies that this linear solution is actually the equilibrium of the strictly
competitive game, i.e., it is the solution to the mechanism design problem.
Step 3.1: Set of linear mechanisms is closed with respect to best responses.
The following definition states what I mean by a linear mechanism to be closed
with respect to best responses.
Definition 3 The set of linear mechanisms is closed with respect to best responses if
2. the seller’s best response, whenever it exists, to a linear forecast is a linear statistic.
38
Since the objective function is quadratic, the point-wise necessary First Order
Condition delivers a linear function of the conditional expectations:
E[ θ2 | X ] − α1 c1 E[ θ1 | X ] − α1 c0 c1
ψ ∗ (2) = . (1)
1 − α1 c21
This statistic is linear because Claim 1 implies that these conditional expectations
are linear. Note that I have only proved that if there is a statistic that solves
the seller’s problem this statistic has to be linear, i.e., this is a necessary, but no a
sufficient condition. This is enough for my purposes since I only want to conclude
that if there is a mechanism that maximizes the seller’s profits when the buyer
uses a linear forecast that maximizing mechanism is a linear mechanism.
Step 3.2: In the set of linear mechanisms, it is without loss to only consider
linear statistics whose coefficient vectors have magnitude equal to one.
This result is a direct application of Lemma 1. Suppose that the seller offers
the statistic L T X + ℓϵ, with ( L, ℓ) ̸= 0. Lemma 1 implies that the buyer obtains the
L ℓ
same value from the alternative statistic L̃ T X + ℓ̃ϵ with L̃ = ∥( L,ℓ)∥
and ℓ̃ = ∥( L,ℓ)∥
,
independently of which variable θ she wants to forecast. Therefore, from the point
of view of the seller both statistics are equivalent as well.
If the seller offers the uninformative statistic L T X + ℓϵ with ( L, ℓ) = 0 he
can equivalently offer the uninformative statistic the only contains noise, i.e., the
statistic L̃ T X + ℓ̃ϵ with L̃ = 0 and ℓ̃ = 1.
39
Lemma 2 Suppose ( a∗ (1, 2, ·), ψ∗ (2)) is an equilibrium of the strictly competitive game
defined by G. Then mina(1,2,·) maxψ(2) G ( a(1, 2, ·), ψ(2)) = maxψ(2) mina(1,2,·) G ( a(1, 2, ·), ψ2 ),
and thus all the Nash equilibria of the game defined by G yield the same payoffs.
As a consequence, the Nash equilibria of a strictly competitive game are interchange-
able: if ( a∗ (1, 2, ·), ψ∗ (2)) and ( â(1, 2, ·), ψ̂(2)) are equilibria then so are ( a∗ (1, 2, ·), ψ̂(2))
40
Proof See the proof of Theorem 22.2 in Osborne and Rubinstein (1994) and the
discussion that follows after.
In Step 3, I showed that there is a minmax point in linear strategies of the zero-
sum game defined by G. Then the lemma guarantees that the linear mechanism
at this minmax point generates to the seller at least the same profits as any other
mechanism.35,36
I complete the argument by showing that there no exists a non-linear mech-
anism that generates the same profits. Let ( a∗ (1, 2, ·), ψ∗ (2)) be an equilibrium
in linear strategies and suppose there is another equilibrium ( â(1, 2, ·), ψ̂(2)).
Lemma 2 implies that ( a∗ (1, 2, ·), ψ̂(2)) is an equilibrium as well. Since ψ̂(2) is
a best response to a∗ (1, 2, ·), Equation 1, in Step 3.1, implies that ψ̂(2) is a lin-
ear statistic. Therefore, in any equilibrium of the game, the seller offers a linear
statistic, completing the argument.
Step 5: Generalization of the argument when both IR constraints bind.
Notice that only Steps 2 and 3.1 depend on which constraints do bind. I now
argue that these steps are still satisfied when both IR constraints and constraint
IC12 bind. In this case, it is still optimal for the seller to offer a full information
statistic to type 1. Further, the seller will still want to charge to type 2 the price
p(2) = Var (θ2 ) − E[(θ2 − a(2, ψ(2)))2 ]. However, now, when considering the
constraints IR1 and IC12 , the price p(1) can be expressed in two alternative ways:
p(1) = E[(θ1 − a(1, 2, ψ(2)))2 ] − Var (θ1 | X ) + Var (θ2 ) − E[(θ2 − a(2, 2, ψ(2)))2 ].
35 Strictly speaking the right equilibrium concept in the sequential game is sequential rationality.
However, in this case, beyond the requirement that the profile of strategies is a Nash equilibrium,
this concept only requires that the seller uses the statistic she buys in an optimal way, which has
been imposed already in Steps 2 and 3. I thank an anonymous referee for this observation.
36 In discrete games, the MinMax Theorem applies only if the players can randomize. In here,
that is not a problem, since both the seller’s and buyer’s action space is a dense and connected
space. I thank an anonymous referee for pointing this out.
41
Var (θ1 ) − E[(θ1 − a(1, 2, ψ(2)))2 ] = Var (θ2 ) − E[(θ2 − a(2, 2, ψ(2)))2 ].
This constraint set is non-empty. Actually, it is easy to check, by using the result
T
γ2T (γ1 −γ2 )
in Lemma 1, that the statistic ψ(2)( X ) = γ2 − γT (γ −γ ) γ1 X belongs to this
1 1 2
set. As the objective function is still quadratic and the constraint is quadratic, the
Fritz conditions,37 imply that it is necessary that the seller’s best response to a
linear forecast is a linear statistic. Further, type 1’s problem is unaffected, so that
her best response to a linear mechanism is still a linear forecast. Therefore, the
argument generalizes to this case.
Steps 1-5 show that when there are only two information types the optimal
mechanism is linear. Step 6 argues why this argument can be generalized to the
case with more information types (but still only one valuation type).
Step 6: Generalization to more than two information types.
Now, consider the case with a finite number of information types. By sequen-
tial rationality each type updates optimally when buying the statistic targeted
to her. If type i buys the statistic targeted to type j, she chooses the forecast
a(i, j, ψ( j)) .
37Fritz (1948) showed that for any solution x to a constrained optimization problem there exists
a vector of multipliers µ = (µ0 , µ1 , . . . , µn ), µ ̸= 0 and µi ≥ 0 such that
n
µ0 ∇ f ( x ) + ∑ µi ∇ gi ( x ) = 0,
i =1
where f corresponds to the objective function and ( gi )in=1 to the set of constraints. These condi-
tions generalize the Kuhn-Tucker conditions by allowing the coefficient assigned to the gradient
of the objective function to be equal to 0, and do not require a constraint qualification. Though
they might generate many candidates for the optimal point(s), they are useful in here because it
allows me to show that any critical point is linear. For a simple proof of the Fritz conditions that
apply to this environment see Birbil, Frenk, and Still (2007).
42
If one or more prices are pin down by more than one constraint, then the set of
feasible mechanisms is restricted as in Step 5. Let ψH denote the set of feasible
mechanisms.38
Notice that all prices are linear combinations of constants and quadratic terms:
E[(θi − a(i, j, ψ( j)))2 ] and −E[(θi − ψ(i ))2 ], for different i and j. Since the seller
prefers larger prices, he wants to maximize E[(θi − a(i, j, ψ( j)))2 ] and −E[(θi −
ψ(i ))2 ]. Each buyer type, however, wants to minimize E[(θi − a(i, j, ψ( j)))2 ] be-
cause the way she pays a lower price without affecting the quality of the statistic
she buys. Further, if the constraint ICij binds and type j is able to reduce the price
she pays, this reduction benefits type i as well.
Therefore, the problem reduces to a strictly competitive game between the
seller who wants to maximize prices by choosing statistics ψ(i ) for each buyer
type in the set of feasible mechanisms ψH and a fictitious player who chooses
forecasts a(i, j, ψ( j)) to minimize prices, which favors all buyer types.
Claim 2 implies that, if the seller offers linear statistics, the fictitious player
wants to choose linear forecasts. At the same time, if the fictitious player chooses
linear forecasts, the seller’s best response is to choose linear statistics. This re-
sults from applying the Fritz’ conditions to the seller’s problem, which objective
function and constraints are quadratic. Therefore, as before, the set of linear
mechanisms is closed with respect to best responses.
As Lemma 2 still applies in this strictly competitive game, the rest of the argu-
ment follows and the result is still true for this extension.
38 The set of restrictions is analogous to the restriction in Step 5. Therefore, the same class of
linear statistics satisfy these restrictions, so that the set ψH is non-empty.
43
V ( I, δ) = max ∑i ∈ I αi pi
{ pi ,Li ,ℓi }i∈ I
( LiT γi )2 ( L Tj γi )2
s.t. LiT Li +ℓ2i σ2
− pi ≥ L Tj L j +ℓ2j σ2
− pj ∀i, j ∈ I
( LiT γi )2
T
Li Li +ℓ2i σ2
− pi ≥ 0 ∀i ∈ I
pi ≥ δ ∀i ∈ I
and let V (δ) = max V ( I, δ) and I ∗ (δ) = arg max V ( I, δ).39 By Lemma 1,
I ⊆{1,...,n} I ⊆{1,...,n}
the mechanism with L̃i = γi , ℓ̃i = 0 and pi = min{γiT γi : i ∈ {1, . . . , n}} satisfy
all the constraints of the auxiliary problem. Then, this problem has a solution
and V ( I, δ) > 0. As the number of types is finite, in the optimal mechanism
min{ pi : pi > 0} > 0. Therefore, as δ → 0, the sequences V (δ) and I ∗ (δ) are
finally constant, and equal to the seller’s profits and the set of types to which the
seller sells in the optimal mechanism, respectively.
Let I ∗ be the set of types to which the data broker sells in the optimal mecha-
nism. Consider the auxiliary problem with I ∗ and small δ, such that the solution
of the auxiliary problem corresponds to the optimal mechanism.
We use the Fritz’s conditions to characterize the optimal solution.40 Let λ0
be the multiplier corresponding to the objective function, λij be the Lagrange
multiplier corresponding to constraint ICij and µi be the Lagrange multiplier cor-
LiT γ j
responding to the constraint IRi for i ̸= j ∈ I ∗ . Defining a ji ≡ T
Li Li +ℓ2i
, the Fritz’
39 Ifthere are multiple subsets of individuals that are optimal, pick any of them.
40 The Fritz’ conditions provide a necessary condition that the optimal solution must satisfy.
Then, they allow to conclude what functional form the optimal solution have, though no to to
solve for it. See Footnote 37 for more details and for a reference for a proof that applies to this
environment.
44
Li
Lemma 1 implies that the buyer’s willingness to pay for the statistic Li and L̃i = k
with k ∈ R++ is the same. Then it is without loss to normalize the denominator
of Li to 1. Further, the problem is bang-bang with respect to ℓi , so that ℓi = 0 for
a ji
all i ∈ I ∗ . By defining c ji ≡ aii we obtain the desired expression.
It remains to show that the mapping from the vector (c ji )i,j ∈ R(n−1)×(n−1)
into itself has a least one fixed point. The mapping is clearly continuous. I ar-
gue that thecodomain and the domain of the mapping are bounded. Define
T
( L i γi ) 2 ( LiT γi )2
ζ = mini∈ I ∗ L LT > 0 since by construction LT L
> 0 for each i ∈ I ∗ . Fur-
i i i i
( LiT γ j )2
ther, define M = maxi∈ I ∗ γiT γi . Lemma 1 implies that ≤ γ Tj γ j ≤ M.
LiT Li
M
Therefore, for each i and j, c2ji ≤ ζ .
Therefore, the mapping that define cij can be restricted to go from the compact
h q q i(n−1)×(n−1)
convex subset − M ζ ,
M
ζ to itself. Brouwer Fixed-Point Theorem
implies that this mapping has at least one fixed-point. Since the proof of Theo-
rem 2 implies that this problem has a solution, its solution needs to satisfy the
Fritz necessary conditions, and one of the fixed-points corresponds to the optimal
mechanism.
Proof of Proposition 2
None of the IC constraints bind if and only if the full-information mechanism
that targets to each type i the linear statistic with coefficients γi , and charges each
type the monopolistic price γiT γi is feasible. Denote by β ij the angle between γi
and γ j . In the full-information mechanism the constraint ICij with i < j is satisfied
iff
(γiT γ j )2 2
T
− γ Tj γ j < 0 ⇔ (γiT γ j )2 > (γ Tj γ j )2 ⇔ cos2 ( β ij ) ∥γi ∥2 < γ j .
γj γj
45
This inequality is always satisfied since cos( β lk ) < 1 and by assumption, for k > l,
∥ γl ∥ 2 ≥ ∥ γk ∥ 2 .
Therefore, if for all i < j, cos2 ( β ij ) ∥γi ∥2 ≤ γ j , the full-information mecha-
nism with monopolistic prices is feasible.
Consider now the case in which some i < j, cos2 ( β ij ) ∥γi ∥2 = γ j + ϵij , with
ϵij > 0 small and for all other k and ℓ cos2 ( β kℓ ) ∥γk ∥2 < ∥γℓ ∥. By Theorem
3 and continuity, the seller offers to type j a statistic with coefficients L j that
are distorted a bit away from γ j in the direction opposite to type i’s preferred
direction and charges to type i a price slightly below ∥γi ∥2 . This deviation in
price and direction reduces as ϵij decreases.
These changes relax the constraints ICiℓ for ℓ ̸= j since now the utility of type
i is positive, and any constraint ICni with n ̸= i are still satisfied as long as the
reduction in the price p(i ) is small enough. Further, any constraint ICkj with k ̸= i
is still satisfied as long as the change in the direction of L j is sufficiently small.
Finally, since type j is obtaining the same rent, any constraint ICjm is also satisfied.
Therefore, when all ϵij are small enough, the relaxed problem the considers
2
only the constraints ICij with i < j such that cos2 ( β ij ) γ j > ∥γi ∥2 solves the
seller’s problem.
Proof of Proposition 3
I consider the relaxed problem that contains only the adjacent downward con-
straints. From Theorem 3 we have that if the constraint ICi,i+1 binds then
L T γi λi,i+1
with ci,i+1 = α̃i+1 LTi+γ1 and α̃i+1 = λi+1,i+2 +µi+1 > 0. Fixing the Lagrange multi-
i +1 i +1
46
By using that (γit γi+1 )2 = ∥γ1 ∥2 ∥γ2 ∥2 cos2 ( β i,i+1 ), where β i,i+1 is the angle
between γi and γi+1 , it is easy to check that the positive root of this quadratic
equation is always larger than or equal to c̄i,i+1 and that the negative root is
always always positive and smaller than or equal to c̄i,i+1 . As ci,i+1 = 0 and
ci,i+1 = c̄i,i+1 are clearly sub-optimal for the seller, the solution to the problem is
given by the negative root of this quadratic equation.
If γiT γi+1 < 0, let γ̃i+1 = −γi+1 . With this change of variables γ̃iT γ̃i+1 > 0.
The procedure above delivers the optimal c̃i,i+1 . By substituting back, ci,i+1 =
−c̃i,i+1 ∈ (c̄i,i+1 , 0).
Now, I check that by picking ϵ small enough the constraint ICi,k for any k >
47
By substituting from the first expression into the second one and so on, we obtain
the following inequality
k
∥γk ∥2 (1 + ϵ)(k−i) > ∏ cos2 ( β j,j+1 ) ∥γi ∥2 .
j =i
48
∂Rent
∂c ∝ −γ1T γ2 + c(γ1T γ1 + γ2T γ2 ) − c2 γ1T γ2 .
I argue that this derivative is negative. This results from this expression being
increasing in c for c < c̄ and being negative when evaluated at c̄. To see the
first, notice that the derivative with respect to c of this expression is equal to
γ1T γ1 + γ2T γ2 − 2cγ1T γ2 and it is larger than zero since c < c̄ < 1 and 2γ1T γ2 <
γ1T γ1 + γ2T γ2 . To see the second, notice that when evaluated at c̄ it is equal to:
−(1 + c̄2 )γ1T γ2 + c̄(γ1T γ1 + γ2T γ2 ) ∝ − ∥γ1 − γ2 ∥2 (γ1T γ1 γ2T γ2 − (γ1T γ2 )2 ) < 0.
49
The only difference in this term between comes from γ1T γ2 . Besides,
∂r (c′ )
∂γ1T γ2
∝ −(γ1T γ2 − c′ γ1T γ1 )(γ2T γ2 − c′ γ1T γ2 )(1 − c′2 ) < 0.
It is easy to check that the first two factors are positive. The third factor
T
is positive since by Corollary 1, c′ < 1. As γ1T γ2 < γ′ 1 γ2′ , r (c′ ) is larger for
data X. Then, for fixed c′ , p(1) > p′ (1), which increases the seller’s profits. By
choosing optimally c the seller can further increase his profits after buying data
X . Therefore, the seller would prefer to buy X rather than X ′ .
Proof of Theorem 4
Consider the relaxed problem without the constraints IC2v−1v′ and the mono-
tonicity constraints. In the next lemmas I present some properties of its solution.
50
√
where I use that sin(cos−1 ( x )) = 1 − x2 in the second equality.
The alternative statistic relaxes the constraints IC1v′ ,2v , which allows the seller
to increase the price p(1v′ ) and his profits, since at least one of them binds.
Proof It follows directly from the fact that for any v ≥ v1∗ the virtual value is
positive, so that increasing q11v relaxes all constraints IC1v′ ,2w , and this allows the
seller to increase his profits.
Lemma 5 There exists v1∗ > v̂2 ≥ v2∗ such that q22v > 0 iff v ≥ v̂2 .
Proof For v < v2∗ the seller optimally picks q22v = 0. This allocation trivially
satisfies all constraints IC1v,2v and maximizes profits because for any v < v2∗ type
(2, v)’s virtual value is negative.
Further, v̂2 < v1∗ because the allocation q11v = q22v = 1v≥v1∗ and constant
prices satisfy all IC constraints with a slack. Then, there exists a neighborhood
V = (v1∗ − ϵ, v1∗ ) such that making q22v = 1 for v ∈ V satisfies all the IC constraints,
and it increases the seller’s profits, since type (2, v)’s virtual value is positive.
v̂2 q22v̂2
Lemma 6 Only the constraints IC1v,2v̂2 for v ∈ (v̂1 , v1∗ ) with v̂1 = h(q22v̂2 )
∈ (v̂2 , v1∗ ),
and IC1v1∗ ,2v′ for v′ ∈ [v̂2 , v̄2 ] bind. Further, v̂2 > v2∗ .
51
where the first inequality follows from constraints IC1v1∗ ,2v′ and q11v = 1 for v >
v1∗ , and the second inequality follows from h(q22v′ ) < 1 by definition of h.
Moreover, the constraint IC1v1∗ ,2v̂2 binds; if it does not, the seller can set v̂2 = v2∗ ,
q22v̂2 = 1 and v̂1 = v1∗ . However, this allocation is not optimal when some crossed
IC constraints bind.
Additionally, it has to be that v̂2 > v2∗ and q22v̂2 < 1: A small reduction of
q22v̂2 or an increase of v̂2 almost does not affect the profits accrue from type 2v̂2 ,
but it reduces the information rent given to all types 1v with v ≥ v1∗ .Then, by
continuity of the RHS of the IC constraints, for v′ in a neighborhood to the right
of v̂2 , q22v′ < 1. This implies that the constraints IC1v1∗ ,2v′ bind: if they did not bind
the seller prefers to set q22v′ = 1. Therefore, in this neighborhood, the derivative
of the RHS of the IC constraints with respect to v′ is equal to zero:
∂q22v′ ∂q ′
v1∗ h′ (q22v′ ) |v′ =v̂2 −v′ 22v | ′ = 0.
∂v ′ ∂v′ v =v̂2
v′
0 or h′ (q22v′ ) = As limx→1 h′ ( x ) =
∂q22v′
Thus, in this neighborhood either dv′ |v′ =v̂2 = v1∗ .
∞, q22v′ is bounded away from 1 for all v′ in this neighborhood. But this neigh-
borhood must contain all values v > v̂2 : If for some type v̊ > v̂2 the constraint
IC1v1∗ −2v̊ does not bind, it would be optimal to set q22v̊ = 1. However, this con-
52
where the inequality follows from the constraint IC1v̂,2v′ and q11v = 0 for v < v̂.
I conclude by arguing that constraints IC1v,2v̂2 for all v ∈ [v̂1 , v1∗ ] have to bind.
Suppose it does not bind for some v ∈ [v̂1 , v1∗ ]. Then the seller would like to
pick q11v = 0 since type (1, v)’s virtual value is negative. But, then type (1, v)
can obtain a positive utility by mimicking type (2, v̂2 ). By taking the derivative in
both sides of these IC constraints I conclude that q11v = h(q22v̂2 ).
We only need to find the allocation for q22v with v ≥ v̂2 . Since, these types’
virtual value is positive, it is optimal to pick the largest q22v′ that, as shown in the
v′
0 or solves h′ (q22v′ ) =
∂q22w
previous lemma, is either the one that solves dw |w=v′ = v1∗ .
It is easy to show that the second derivative of h(q22v ) is strictly positive,
−1
so that h′ (q22v ) is strictly increasing, its inverse exists, and h′ is increasing.
Therefore, there exists ṽ2 ≥ v̂2 such that q22v = q22v̂2 for v < ṽ2 and q22v =
−1 v
h′ v∗ for v > ṽ2 , so that q22v is non-decreasing. To complete the argument
1
h(q22v̂ )
I show that ṽ2 > v̂2 . They are equal only if h′ (q22v̂2 ) = q22v̂ 2 , or equivalently,
r 2
1 − q
if tan(cos−1 (q22v̂2 ) + β) −
22 v̂ 2
q22v̂ = 0. The LHS of this condition is strictly
2
decreasing with lim LHS = tan( β). Since β > 0, there is not value of q22v̂2 < 1
q22v̂2 →1
that satisfies this condition, and ṽ2 > v̂2 .41
Finally, it is straightforward to check that this solution satisfies the constraints
that were omitted at the beginning of the proof.
41 Note that q22v might be a constant function since it is not necessary that ṽ2 ≤ v̄2 .
53