0% found this document useful (0 votes)
3 views55 pages

Selling Data

The document presents a model where a monopolist seller offers multi-attribute consumer data to buyers with uncertain preferences, aiming to maximize profits through optimal pricing and statistics. It demonstrates that the seller should provide linear combinations of data statistics and may need to offer a continuum of these statistics to effectively differentiate between buyers. The findings contribute to the literature on information design and mechanism design in data markets.

Uploaded by

bailid697
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views55 pages

Selling Data

The document presents a model where a monopolist seller offers multi-attribute consumer data to buyers with uncertain preferences, aiming to maximize profits through optimal pricing and statistics. It demonstrates that the seller should provide linear combinations of data statistics and may need to offer a continuum of these statistics to effectively differentiate between buyers. The findings contribute to the literature on information design and mechanism design in data markets.

Uploaded by

bailid697
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Selling Data*

Carlos Segura-Rodriguez†

August 23, 2022

Abstract

A profit-maximizing monopolist (seller) sells multi-attribute consumer


data to a firm (buyer). The seller is uncertain about which unknown con-
sumer characteristic the buyer is interested in forecasting and how much the
buyer values information. In order to screen among potential buyers along
both margins, the seller chooses a menu of statistics of the data to offer and
the price of each statistic. Assuming that the data and unknown character-
istics follow an elliptical distribution, I obtain two results. First, I show that
the seller optimally offers statistics that are linear combinations of the data.
Second, I show that the seller might need to offer a continuum of statistics,
and that they are less correlated than they would be if the seller could per-
fectly discriminate. Every optimal statistic contains information about every
variable in the data, and does not include uncorrelated noise.
Keywords: Information Design, Mechanism Design, Multidimensional Screen-
ing, Product Design, Data
JEL Classification: D42, D82, D83, D86.

* I am grateful to the coeditor and two anonymous referees for their constructive suggestions.
I also would like to express my gratitude to George Mailath, Annie Liang, Rakesh Vohra, Andrew
Postlewaite, Aislinn Bohren and Rohit Lamba for their insightful comments and support through-
out this project. I also thank Dirk Bergemann and Nima Haghpanah for their valuable input.
Thanks as well to my fellow classmates, especially Ashwin Kambhampati, Youngsoo Heo, Joao
Granja, and Paolo Martellini for listening and providing help during the project. I also gratefully
acknowledge all the participants in the UPenn Micro Theory Lunch, the UPenn Micro Theory
Seminar, NASMES 2019 and SEA 2019 for their feedback. All errors are my own.
† Central Bank of Costa Rica, Department of Economic Research, Av 0 and 1, St 2 and 4, San

Jose, 10101, segurarc@[Link].

Electronic copy available at: [Link]


Contents

1 Introduction 1
1.1 Literature Review . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3

2 Model 6
2.1 Elliptical Distributions . . . . . . . . . . . . . . . . . . . . . . . . . . . 8

3 Value of Information 10
3.1 Linear Mechanisms . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11

4 General Properties of the Optimal Mechanism 13

5 Optimal Mechanism with a Unique Information Type 15

6 Optimal Mechanism with a Unique Valuation Type 16

7 Mix model: Two Information Types and a Continuum of Valuation Types 26

A A Micro Foundation for Quadratic Loss 33

B Proofs 33

Electronic copy available at: [Link]


1 Introduction

Markets for data are developing quickly worldwide due to the ease of gathering
information from electronic commerce and on-line surfing. In these markets, data
brokers (sellers) are facing new challenges since data’s core properties are differ-
ent from most goods. First, data is intrinsically multidimensional: data brokers
observe multiple characteristics of each individual. Second, data is easy to trans-
form and summarize in statistics, and these statistics are highly substitutable. In
this context, it is important to understand the optimal way for a data monopo-
list seller to present and price data when buyers have private information about
which unknown consumer characteristic (willingness to pay for a product, credit-
worthiness, political preferences,...) the buyer (firm, financial institution, political
party,...) wants to forecast and the value that she assigns to this knowledge.1 I
introduce a novel model in which a seller sells multi-attribute data and chooses a
menu of statistics and prices to differentiate among buyers in both dimensions.
As an example, consider a firm that plans to introduce a new product to the
market and wants to identify which consumers will react positively to its mar-
keting efforts, in order to better target its marketing campaign. The data broker
believes the firm is either a casual or a fine-dining restaurant, but this is private
information of the buyer. Suppose the consumer’s income and the distance from
the consumer’s home to the restaurant are observable. Further, assume that the
restaurant is interested in knowing how much the consumer is willing to pay for
one of its meals, and the consumer’s willingness to pay is, in both cases, posi-
tively related with his income and negatively related with the distance he lives
from the restaurant. Due to the substitutability of information, any statistic that
is increasing in income and decreasing in distance will be valuable for both types
of firms, so that if any offered statistic is sufficiently cheap both of them will buy
it. To avoid this issue, the data broker needs to carefully design statistics and
1I use male pronouns for the seller (data broker) and female pronouns for the buyer (firm).

Electronic copy available at: [Link]


set prices that allow him to differentiate both vertically (in terms of the buyer’s
valuation) and horizontally (in terms of the buyer’s underlying needs).
In my model, a random vector θ represents the set of consumer characteristics
the buyer might want to forecast (in the example, willingness to pay for a meal in
a casual or in fine-dining restaurant). Each buyer wants to forecast only one of the
elements of θ and faces a quadratic loss function that is weighted by how valuable
this information is for her. There is a seller who can observe some consumer data
(in the example, consumer’s income and the distance he lives from the restaurant).
The seller chooses a mechanism to maximize the profits of selling this data by
differentiating among buyers. I assume the seller can commit to this mechanism
before observing the realization of the data.2 Further, I assume that the seller
and the buyer share a common prior about the joint distribution of all consumer
characteristics, observable or not, and that its joint distribution is elliptical.3
The problem the data broker faces is difficult: It is a multidimensional and
multiproduct mechanism design problem. Besides the intrinsic complications that
such an environment possesses (see, for example Daskalakis, Deckelbaum, and
Tzamos (2017)), the payoff functions in this information environment are non-
linear. Therefore, the commonly used first-order approach does not work (see
Rochet and Choné (1998)). And, as far as I know, there is no general solution
known for this class of problems. By using a direct approach, however, I am able
to characterize the optimal mechanism.
I show, first, that the seller should optimally offer only mechanisms in which
the statistics are linear combinations of the data and uncorrelated noise. Hence,
the seller will only offer linear regressions of the data.4 The argument depends
crucially on the existence of conflict between the seller and the buyer in the de-
2 The commitment assumption is satisfied if, for example, the seller can publish menus only in
discrete periods of time, but new consumer data arrives continuously.
3 Section 2.1 presents more details about this family of distributions.
4 This regression does not need to coincide with the Least Squares regression.

Electronic copy available at: [Link]


signing problem.5 By focusing on linear mechanisms, I solve for different spec-
ifications of the model with increasing levels of complexity, and show that the
solution must satisfy three general properties.
First, a generalization of the no distortion at the top and no rent at the bottom
properties hold: There is a type that receives an undistorted statistic of the data in
the sense that the statistic she receives is equivalent for her to receiving all data,
and there is a type that does not receive any information rent. Second, the offered
statistics are less correlated than they would be if the seller could perfectly dis-
criminate among the buyers. Without introducing uncorrelated noise, the seller
degrades the statistic targeted to a mimicked type by relatively decreasing the
weights assigned to the variables that are more informative for the mimicking
type(s), affecting the mimicking type(s) more than the mimicked type. However,
generically, all variables in the data are used to create every statistic that is of-
fered.6 This implies that the seller does not create differentiation by selling a set
of variables to one type and a different set to other types. Third, if there are a con-
tinuum of valuation types, the seller might need to offer a continuum of statistics.
Using a continuum of degraded statistics allows the seller to continuously differ-
entiate between buyers, as randomization does in multidimensional mechanism
design (see Daskalakis, Deckelbaum, and Tzamos (2017), for example).7

1.1 Literature Review

My paper contributes to a growing body of literature that studies the optimal way
in which a monopolist should sell information. The work that comes closest to my
approach is Bergemann, Bonatti, and Smolin (2018). They study an environment
5 Some examples in the literature show that, in the absence of conflict between players, using
only linear statistics is not always optimal (see, for instance, Witsenhausen (1968)).
6 Since the optimal statistics are a linear combination of the buyers’ preferred statistics, the set

of coefficients that generate a zero entry has measure zero.


7 Zhong (2018) has shown that this occurs in information environments when the set of actions

that the buyer can take after buying information is sufficiently rich.

Electronic copy available at: [Link]


in which all buyer types want to match the realization of a single variable that
is observed by the seller. Since their priors are different, their valuations depend
on the realization of the variable. In contrast, I consider the problem in which
the seller sells multidimensional data, and in which buyers’ private information
has two dimensions: which unobservable variable they want to forecast and how
much they value this information. Then, the model, further, allows me to analyze
how horizontal and vertical information that is privately held by the buyer interact
and shape the optimal mechanism.
Thus, both models focus on two different strategies that data brokers use to
provide information. According to the Federal Trade Commission (2014), data
brokers usually provide information in one of two forms: Either the data broker
appends lost information to consumer’s lists that are already held by the buyer, or
it provides the firms with consumer marketing scores. While Bergemann, Bonatti,
and Smolin (2018)’s model is better suited to understand how the data brokers
should append consumer lists, my model is better suited to study how the data
brokers should create and sell these marketing scores.
In the easy-to-work with framework that I introduce, it is possible, opposite to
Bergemann, Bonatti, and Smolin (2018)’s environment, to study how the data bro-
ker can aggregate different variables in the data into marketing scores (statistics
or data summaries) that allow data buyers identify which consumers are more
likely to buy a product.8
A few other papers have studied how information is sold in different markets.
In an environment without private information, Bergemann and Bonatti (2015)
study the problem of selling information in a competitive market when the seller
can only sell signals that reveal perfectly one of the states. Yang (2018) analyzes
an environment in which a profit-maximizing monopolist reveals information to
8 Though the model’s scope is limited, it is general enough to include the generalized OLS,
multinomial probit and multinomial logit models, which are extensively used in applied work to
understand consumer behavior.

Electronic copy available at: [Link]


consumers about their valuations for a good, but it charges the sellers of goods
for this information. Babaioff, Kleinberg, and Paes Leme (2012) consider a prob-
lem in which a seller and a buyer have private information about one state, but
the buyer’s payoff function depends on both of them. They focus on finding
conditions under which the revelation principle holds.
Smolin (2019) studies the related problem of a seller that provides information
about some product’s attributes. His main result is that it is without loss for the
seller to provide a binary signal that recommends to buy the product if a linear
combination of the attributes is above a threshold and not to buy if it is below
it. While in his environment a linear mechanism being optimal is a result of the
consumer’s decision being binary, in my environment it is a result of the joint
distribution of the data and consumer’s unknown characteristics being elliptical.
The problem I study is technically related to the literature on multidimen-
sional mechanism design (Armstrong, 1996; Thanassoulis, 2004; Manelli and Vin-
cent, 2006; Daskalakis, Deckelbaum, and Tzamos, 2017), and the literature on
quality degradation (Mussa and Rosen, 1978; Maskin and Riley, 1984) and prod-
uct design (Anderson and Celik, 2015). While in the first branch of literature it is
commonly assumed that the products are fixed, in the second it is assumed that
the preferences are unidimensional and that the ranking of buyer’s valuations
are product-independent. In contrast, in my multidimensional environment, the
seller decides the characteristics and quality of the products he offers, and the
ranking of types according to their valuations is product-dependent, in the sense
that statistics that are considered to be of high quality by some types are consid-
ered to be of low quality by others.
Finally, this paper relates to the literature on selling substitutable goods. In
these environments, due to substitution effects, the optimal mechanism may in-
volve lotteries, as shown by Pavlov (2011) in a linear two-goods problem, and
Balestrieri, Izmalkov, and Leao (2015) in a Hotelling-type model. The substitution

Electronic copy available at: [Link]


patterns in my environment are are similar to those found in the discrete version
of the model of Wilson (1993) (see Vohra (2011)). However, payoffs are nonlinear
in mine, while in his they are linear. Hence, the techniques necessary to solve the
problem and its solution are different.

2 Model

Data brokers (sellers) sell consumer data to firms (buyers). I study how a monop-
olist seller, who wants to maximize profits, should sell the multi-attribute data
he can access when he is uncertain about the buyer’s motives for purchasing it.
By assuming that consumer’s characteristics come from the same distribution, I
focus on the problem of selling data of a representative consumer.

Buyer The seller knows there is a buyer that wants to forecast the value of an
unknown consumer characteristic in the vector θ = (θ1 , θ2 , . . . , θn ). For simplicity,
I assume that each buyer wants to forecast a unique characteristic θi ∈ θ. In the ex-
ample in the introduction, the vector θ corresponds to the consumer’s willingness
to pay for a meal in a casual or in a fine-dining restaurant.
The buyer’s preferences are represented by a two-dimensional type (i, v) ∈
{1, . . . , n} × R++ where i is her “information type” and v is her “valuation type”.
Type (i, v) takes an action a ∈ R to minimize

−vEθi [(θi − a)2 ];

that is, type (i, v) is interested in making the most accurate forecast, a, of con-
sumer’s characteristic θi according to a quadratic loss function.9 Finally, the
buyer’s payoff is quasilinear in money; it is equal to the forecast loss minus the
9 Appendix A presents a microfoundation for this assumption. It shows that a monopolist
who is uncertain about the intercept of a linear demand, faces a profit loss that is quadratic with
respect to the forecast of the intercept that the monopolist uses in his decision-making process.

Electronic copy available at: [Link]


price she pays for information.

Seller The seller has access to consumer data x = ( x1 , . . . , xk ) ∈ Rk . In the ex-


ample in the introduction, the vector x corresponds to the consumer’s income and
the distance form the consumer’s house to the restaurant. The seller can append
to this data a random noise variable, ϵ ∈ R, that is assumed to be independently
distributed from X and θ.
The seller does not know the buyer’s type, so he cannot perfectly price dis-
criminate. Instead, he assigns probability αi to the buyer being information type
i and believes that the valuation type is distributed, conditional on the informa-
tion type being i, according to an absolutely continuous distribution G (v | i ) that
admits a density g(v | i ).
Further, assume that the seller, before knowing the data realization x, commits
to a selling mechanism. I look for the incentive-compatible and ex-ante indi-
vidually rational direct mechanism that maximizes the seller’s profits.10 In this
environment, a direct mechanism is a pair (ψ, p), where ψ(i, v) : Rk+1 → Rm is
the statistic (garbling) of the data and noise that the seller offers to type (i, v), and
p(i, v) ∈ R is the monetary transfer from type (i, v) to the seller.

Prior information It is assumed that the firm’s prior over the tuple ( X, θ, ϵ)
corresponds to an elliptical distribution with finite first and second moments,
which is denoted by F. Further, the marginal distributions over ( X, θ ) and θ,
F(X,θ ) and Fθ exist; and that the conditional distribution of X given θ, FX |θ , exists.
Moreover, without loss of generality I assume that all variables have zero mean
with Var ( X ) = Ik and and Var (ϵ) = [Link] information environment is common
knowledge to both the data broker and the firm.
10 Babaioff, Kleinberg, and Paes Leme (2012) show that the seller might be able to use his private

information to extract surplus from the buyer beyond what he is able to capture by using direct
mechanisms. I thank an anonymous referee for this observation.

Electronic copy available at: [Link]


Notice that these assumptions on the information environment have two im-
plications. First, by restricting the joint distribution of X and θ, I restrict the
environment I study. Second, by restricting the joint distribution of the tuple
( X, θ, ϵ), I restrict the type of noise that the seller can introduce into the statistics,
that is, I restrict the scope of mechanisms that the seller can offer.11

Timing The game between the buyer and seller proceeds as follows

1. The data broker commits to a direct mechanism.

2. Nature jointly draws the tuple ( X, θ, ϵ) according to the joint distribution F.

3. The buyer reports her type to the seller and receives the promised statistic
at the promised price.

4. The buyer takes an action and her forecasting loss is realized.

2.1 Elliptical Distributions

In this section, I briefly introduce the family of elliptical distributions and describe
its main properties. A random variable with an elliptical distribution whose first
two moments are finite and that assumes a density f has a density given by

f ( x ) = kg(( x − µ)′ (λΣ)−1 ( x − µ)),

where k is a normalization constant, µ is the mean, λ is a constant and Σ is the


covariance matrix. This formulation makes clear that this family of distributions
generalizes the normal distribution, which corresponds to the special case with
11 That ϵ is one-dimensional is without loss of generality in this elliptical environment, as it will
be clear later on. However, that the joint distribution of ( X, θ, ϵ) is elliptical restricts the garbles
that the seller can obtain by introducing random noise. This assumption is necessary to use the
properties of elliptical distributions on Claim 1, which are key for the proof of Theorem 2. It is
beyond the scope of this paper to investigate whether allowing the seller to use a non-elliptical
random noise can help him increase his profits.

Electronic copy available at: [Link]


g = exp and λ = 1. This family includes distributions that are skewed or lep-
tokurtic, distributions for which zero covariance is different from independence,
and distributions with fatter tails than the normal distribution. Among those
distributions with finite first and second moments, this family includes the multi-
variate normal distribution, the multivariate t-student distribution, the multivari-
ate symmetric Laplace distribution, the multivariate logistic distribution, and the
multivariate symmetric hyperbolic distribution.
In spite of its generality, this family of distributions satisfies two properties
that simplify the Bayes updating process, and for which the normal distribution
is commonly used. First, any conditional expectation is linear, and, second, any
linear combination of elliptical distributions is an elliptical distribution.12 The
proof of the following claim follows directly from the general proposition stated
in Fang, Kotz, and Ng (1990).

Claim 1 Suppose the joint distribution of (θ, X, ϵ) is jointly elliptical. Then for any
non-zero vector L and/or non-zero constant ℓ,

1. (θ, L T X + ϵℓ) is elliptical; and

2. E[θ | L T X + ϵℓ] is linear in L T X + ϵℓ.

This property is key. Claim 1 implies that the buyer’s optimal forecast corre-
sponds to the conditional expectation of the variable she is interested in given the
statistic she buys. Therefore, this property implies that if the buyer buys a linear
statistic his optimal action is a linear transformation of this statistic. Theorem 2
uses this fact to show that it is optimal for the seller to only offer linear statistics,
simplifying the analysis.
12 Because of these properties, the family of elliptical distributions have been used in multivari-
ate analysis (Anderson, 2003), portfolio choice (Chamberlain, 1983; Owen and Rabinovitch, 1983)
and information economics (Mailath and Nöldeke, 2008; Deimen and Szalay, 2019; Ball, 2019).

Electronic copy available at: [Link]


3 Value of Information

I start the analysis by calculating the maximum price that type (i, v) is willing
to pay for any statistic ψ. Since the firm utility is quadratic, type (i, v)’s optimal
forecast equals the conditional expectation of the characteristic she is interested
in, θi , given the statistic ψ. With this forecast, type (i, v)’s expected loss is equal
to v times the negative of the expected conditional variance of θi given ψ.13

Claim 2 If the data broker offers to type (i, v) the statistic ψ at the price p, her optimal
forecast is a∗ = EX,ϵ (θi | ψ), and her ex-ante expected utility is equal to

−vEX,ϵ [Var (θi | ψ)] − p.

This result allows me to define the incentive compatibility and individual ra-
tionality constraints. The mechanism (ψ, p) is incentive compatible (IC) if

−vEX,ϵ [Var (θi | ψ(i, v))] − p(i, v) ≥ −vEX,ϵ [Var (θi | ψ( j, v′ ))] − p( j, v′ ) ∀(i, v), ( j, v′ ),

and it is individually rational (IR) if

−vEX,ϵ [Var (θi | ψ(i, v))] − p(i, v) ≥ −vVar Fθ (θi ) ∀(i, v).

These constraints have two distinct properties that make them slightly differ-
ent from those commonly found in the literature. First, the IC constraints embed
two possible deviations: As usual the buyer can misreport her type. But when
the buyer lies, she can use the statistic she receives to make a forecast about the
variable she is interested in, which might not coincide with the variable that the
statistic was targeted to.
Second, the right-hand side of the IR constraint is not the same across types,
since when the buyer does not buy any information, her optimal forecast is her
prior expectation, which results in a payoff equal to the negative of her prior vari-
13 The proof of Claim 2 and all others that are not in the main text can be found in Appendix B.

10

Electronic copy available at: [Link]


ance times her valuation type. Further, by reordering the IR constraint, I conclude
that type (i, v)’s willingness to pay for a statistic ψ is v times the reduction in the
variance that she can achieve by buying it.

3.1 Linear Mechanisms

Claim 2 implies that the buyer’s optimal forecast is equal to the conditional ex-
pectation given the statistic she buys. But, in this elliptical environment, if the
buyer buys all data X her conditional expectation is linear in X. Further, when
she buys a statistic that is a linear combination of the data and noise, her condi-
tional expectation is still linear in X and ϵ. Therefore, it is sensible to pay special
attention to the set of mechanisms in which the seller only offers linear statistics.

Definition 1 A mechanism is linear if all the statistics take the form ψ(i, v) = L(i, v) T X +
ℓ(i, v)ϵ, where L(i, v) ∈ Rk and ℓ(i, v) ∈ R.

Considering this class of mechanisms simplifies the analysis. Lemma 1 presents


a close-formed expression for the buyer’s willingness to pay for a linear statistic.

Lemma 1 Type (i, v)’s willingness to pay for a linear statistic L T X + ℓϵ with L ̸= 0
and/or ℓ ̸= 0, corresponds to
( L T Cov(θi , X ))2
v .
L T L + ℓ2
There are two clarifications worth mentioning at this point. First, if ϵ where
multidimensional, the value of a linear statistic L T X + ℓ̂ T ϵ would be
( L T Cov(θi , X ))2
v ,
L T L + ℓ̂ T Σϵ ℓ̂

where Σϵ corresponds to ϵ’s covariance matrix. Therefore, considering a uni-


dimensional noise component is without loss of generality. Second, for this result
it is critical that the joint distribution of the tuple ( X, θ, ϵ) is elliptical. In other

11

Electronic copy available at: [Link]


case, the conditional expectation might be no linear. Then, the buyer’s willingness
to pay might contain an interaction term between X and ϵ.
To facilitate notation, I denote by γi the covariance between θi and X. As it
is standard in linear regressions, it is easy to show that E[θi | X ] = γiT X, that
is, γi is the coefficient vector that generates the optimal forecast for θi given the
available data X. Then, type (i, v)’s willingness pay for a statistic L T X + ℓϵ can
be written as vρ2 γiT γi , where ρ represents the correlation between E(θi | X ) and
L T X + ℓϵ and γiT γi = Var ( E[θi | X ]). That is, fixing type (i, v)’s willingness to
pay for all the data, vγiT γi , the buyer is willing to pay more when the square of
the correlation between her optimal forecast, E[θi | X ], and the statistic she buys
is large.14

L2

γi

L0

L1

Figure 1: For a fixed direction, the distance between the origin and the blue locus
represents type (i, v)’s willingness to pay for a noiseless statistic L T X with vector
L pointing in that direction.

Further, when the statistic does not contain independence noise, the square of
the correlation has a nice geometrical interpretation: It corresponds to the square
of the cosine of the angle β between γi and L. Figure 1 plots type (i, v)’s will-
ingness to pay for any noiseless linear statistic. For a fixed direction, the distance
between the origin and the blue locus represents the function, vγiT γi cos2 ( β), i.e.,
type i’s willingness to pay for the statistic L T X with vector L pointing in that di-
14 This is analogous to the coefficient of determination R2 , which in regression analysis is used

as a goodness-of-fit measure and corresponds to the square of the correlation between the ob-
served and the regression predicted values.

12

Electronic copy available at: [Link]


rection—for example, in the same direction as L0 . As the angle between γi and L
varies between 0 and π/2, type (i, v)’s willingness to pay decreases continuously
from vγiT γi to 0.
A similar comparison holds if the linear statistic contains noise. In this case,
the correlation between L T X + ℓϵ and γi X decreases either as ℓ increases or as
the angle between L and γi increases. A key difference between modifying L
or ℓ in a statistic is that the willingness to pay of different buyer types might
change in different directions when modifying L, but the willingness to pay of all
types reduces proportionally as ℓ increases. Further, increasing ℓ makes all types
willingness to pay for a statistic less sensitive to a change in L.

4 General Properties of the Optimal Mechanism

The optimal mechanism has two properties that generalize those that have been
found in unidimensional environments that satisfy the single crossing condition.15
First, there is no rent at the bottom: at least one type pays his willingness to pay
for the statistic she receives. If this were not the case, the data broker can increase
his profits, without affecting the buyers’ incentives, by uniformly increasing all
prices. Second, there is no distortion at the top: at least one type receives an
statistic that for her is equivalent to receiving all data. The reasoning is subtler;
if there is no type that receives an undistorted statistic, the seller can offer a new
package containing all data at a price high enough that only one type is willing to
pay. Since this new product has a price that is too high for other types, all other
IC constraints are satisfied.

Theorem 1 In the optimal mechanism,

i. at least one type receives zero information rent; and


15 Interestingly,
these properties are satisfied even without making any assumption about the
joint distribution of (θ, X, ϵ).

13

Electronic copy available at: [Link]


ii. at least one type receives an undistorted statistic.

Contrary to unidimensional environments, which type is at the bottom or at


the top is not completely determined by the types’ willingness to pay for all data.
Rather, a type is at the top if there is no binding IC constraint pointing towards
her, and a type is at the bottom if there is no binding IC constraint going from this
type to another type. Example 1, in Section 6, provides cases in which the types
at the top and at the bottom do not coincide with the types with the highest and
lowest willingness to pay for all data, respectively.
The other important characteristic of the optimal mechanism is what class of
statistics the seller will offer. Since a buyer’s incentive to mimic another type de-
pends on how much she can learn from the statistic that is targeted to that other
type, the seller can increase his profits by selling statistics that allow one type to
reduce her forecast variance significantly, but it impedes other types to learn that
much. One might expect that to attain this objective the seller needs to offer statis-
tics that are complicated transformations of the data and noise. Theorem 2 shows
this is not the case if there is a unique valuation type v: In the optimal mechanism,
the seller only need to offer linear statistics, which are rather simple.16

Theorem 2 Suppose G (· | i ) := δv for all i, where δv is the Dirac measure on some


valuation v ∈ R++ . The optimal mechanism is a linear mechanism.

If none of the IC constraints bind, the argument is trivial: it is optimal for


the seller to target to each type the conditional expectation of the variable she is
interested in, and by Claim 1 this conditional expectation is linear. The significant
part of the result is that even when some IC constraints bind, the seller cannot
improve upon linear mechanisms. This occurs despite the fact that when at least
one IC constraint binds, the seller will offer linear statistics different from the
conditional expectations.
16 I
suspect that this property is still true for more general cases. The proof of that statement is
beyond the scope of this paper.

14

Electronic copy available at: [Link]


The proof of the theorem, which was inspired by the argument in one example
in Basar (2008), consists of two steps. In the first step, analogously to Kamenica
and Gentzkow (2011), I show that it is without loss for the seller to offer only
deterministic action recommendation mechanisms, that is, mechanisms in which
the statistics offered and the buyers’ actions (forecasts) coincide: a(i, v) = ψ(i, v).
In the second step, I consider a stochastic strictly competitive game between
the seller and the buyer, in which the buyer’s payoff function moves in a opposite
direction to the seller’s profits. The conflict between the seller and the player
arises from two opposite forces. On one hand, the seller’s profits increase when
choosing statistics that reduce what a type can learn from buying the statistic
targeted to another one. On the other hand, the buyer’s utility increases when
choosing a forecasting procedure that minimize the posterior variance of θi after
observing any statistic ψ, including those statistics targeted to other types. I argue
that, in this game, the seller’s best response to the buyer’s best response to a linear
mechanism is a linear mechanism, which generates an equilibrium of the game.
Therefore, since all equilibria of a strictly competitive game generate the same
payoffs to both players,17 it is optimal for the seller to offer a linear mechanism.18

5 Optimal Mechanism with a Unique Information Type

To gain intuition, suppose as in Bergemann, Bonatti, and Smolin (2018) that all
buyers are interested in forecasting the same variable, i.e., there is a unique in-
formation type. The critical difference with their environment is that in here all
buyer types share the same prior.
17 This is a result of the fact that in any strictly competitive game, if there is an equilibrium, then

max min = min max. This property, further, implies that the equilibrium payoffs are independent
of the timing in the game.
18 The argument relies on the existence of conflict between the two players. In fact, Witsen-

hausen (1968) introduced an example with normal distributions in which, in the absence of con-
flict, is not always optimal to use linear statistics and linear updating rules.

15

Electronic copy available at: [Link]


Therefore, when buying any statistic, all of the buyers will choose the same
optimal action and obtain the same reduction in ex-post variance. In this case, one
can define a static’s quality q as how much this ex-post variance can be reduced
by using this statistic. This interpretation relies on the fact that all types rank all
the statistics according to the same variance reduction.
Therefore, with this interpretation, the problem is equivalent to the problem
in Myerson (1981), in which the buyer’s willingness to pay for a statistic (product)
is proportional to the buyer’s valuation and the quality of the statistic: vq. Since
it is cost-less to produce any statistic, and if I suppose that the distribution of
the valuation types satisfies Myerson (1981)’s Regularity Condition, then, as in
Myerson (1981), the optimal mechanism is simple: The seller offers all data to
buyers with valuation above a threshold v∗ , and no information to buyers with
valuation below that threshold. Proposition 1 formalizes this statement.19

Proposition 1 Suppose there is a unique information type t who wants to forecast the
1− G ( v | t )
random variable θ and that v − g(v|t)
is non-decreasing in v. Then, there exists v∗
such that in the optimal mechanism all types v ≥ v∗ receive the statistic E[θ | X ] and
pay v∗ Var Fθ − EX,ϵ [Var (θ | X )] , and all types v < v∗ do not receive any information


and pay nothing.

6 Optimal Mechanism with a Unique Valuation Type

Suppose now that there are multiple information types that share a common
valuation type. This case is more interesting than the previous one because now
the perception of the quality of a statistic by a buyer depends on which consumer
characteristic she wants to forecast. To simplify the analysis, I normalize, without
loss of generality, the valuation type to 1, and assume that γ1T γ1 ≥ γ2T γ2 ≥ . . . ≥
19 The proof of this proposition is omitted since it follows from standard arguments.

16

Electronic copy available at: [Link]


γnT γn , with at least one strict inequality.20 That is, types are ordered according
to their willingness to pay for all data. Further, I assume that each buyer type
wants to forecast a different consumer characteristic, that is, the angle between
any vectors γi and γ j , which I called β ij cannot be either 0 or π. Assumption 1
formally states this. I relegate the study of the mixed model in which different
buyer types might want to forecast the same variable to Section 7.

Assumption 1 Let β ij be the angle between γi and γ j . For any i ̸= j, β ij ∈ (0, π ).21

I start the analysis by presenting, in Theorem 3, a general characterization of


the class of statistics the seller offers in the optimal mechanism. Since it is not
possible to know a priori which IC constraints bind, this characterization depends
on the endogenous value of the Lagrange multipliers. For simplicity of language,
from now on, I refer to the IC constraint that prevents type i of mimicking type j
as ICij and refer to the IR constraint for type i as IRi .

Theorem 3 Let I ⊆ {1, . . . , n} be the set of types to whom, in the optimal mechanism,
the data broker sells data. To each i ∈ I, the seller offers the statistic LiT X with
!
λ ji c ji
Li = Var ( X )−1 Cov(θi , X ) − ∑ Cov(θ j , X ) ,
j∈ I,j̸=i ∑ j∈ I,j̸=i ij
λ + µi

where λkℓ is the Lagrange multiplier associated with the constraint ICkℓ , µi is the La-
LkT Cov(θk ,X )
grange multiplier associated with the constraint IRi , and ckℓ = LℓT Cov(θℓ ,X )
.

The optimal statistics satisfy the following main property: When an IC con-
straint binds, the data broker decreases the correlation between the statistics of-
fered to those two types by modifying the statistic given to the mimicked type
20 If none of the inequalities are strict, the data broker can reveal all the information to all types
and charge a uniform price equal to γ1T γ1 .
21 If there are two types i and j with angle β = 0, they would only differ in the value they
ij
are willing to pay for any statistic, but they would have the same preferred statistic. If β ij = π
both types’ preferred statistics point in opposite directions, so buying any of their two preferred
statistics provide them exactly the same information.

17

Electronic copy available at: [Link]


in the direction opposite to the mimicking types’ preferred statistics. Because
there might be multiple types that want to mimic the same type, the data broker
weights (through the Lagrange multipliers) in which direction to distort the infor-
mation. This generalizes the result obtained by Bergemann, Bonatti, and Smolin
(2018) to multi-attribute data, and makes clear that the optimal way for the seller
to degrade quality is to offer statistics that are less correlated than the buyers’
preferred statistics.22
Further, in the optimal mechanism, some types might not obtain any valuable
statistic. The reason is that selling information to them might force the seller to
create excessive distortions in the statistics offered to the other types.23 Finally,
notice that, since the problem is bang-bang with respect to the uncorrelated noise,
none of the offered statics contains noise. The reason is that adding uncorrelated
noise is an inefficient instrument: While uncorrelated noise reduces proportion-
ally the willingness to pay of all types for a statistic, modifying the direction of
the statistic targeted to type i away from type j’s preferred direction has a non-
proportional effect, affecting more type j’s willingness to pay.24
Though Theorem 3 is very general, one important issue is that it does not
restrict the set of IC and IR constraints the seller might need to consider, and this
set might be large. Intuitively, the IC constraint that prevents type i of mimicking
type j, constraint ICij , will bind only if type i can make an accurate guess about
E[θi | X ] from observing E[θ j | X ], while buying the statistic targeted to type j at
a relatively low price. How much can learn i from the statistic targeted to type j
depends on the correlation between these two conditional expectations: If they are
22 The optimal mechanism in Bergemann, Bonatti, and Smolin (2018)’s environment implicitly
shares a similar property: By introducing noise only to the statisitc offered in one of the states
to one of the buyers, the seller simultaneously increases its variance and reduces the correlation
between this statistic and the full-information statistic.
23 Example 1 shows that the seller might exclude those types that other types have a higher

incentive to mimic, which might not be those with the lowest willingness to pay for all data.
24 Uncorrelated noise has a proportional effect as time delay does in Stokey (1979). Fortunately

for the seller, in here, changing the direction of the statistic has a non-proportional effect.

18

Electronic copy available at: [Link]


perfectly correlated, type i can learn the same from both; if they are orthogonal,
type i cannot learn anything useful from observing E[θ j | X ]. As seen before,
this correlation can be measured in terms of the angle between the vectors γi and
γ j , β ij . Further, the price that the seller charges to each type depends on how
valuable information is for them, which is measured by the magnitude of γi and
γj .
Proposition 2 uses this insight to show that if all types have similar willingness
to pay for all data, the only relevant IC constraints are the downward constraints
for types with vectors γi and γ j pointing in a similar direction, while γi ’s magni-
tude is large relative to γ j ’s magnitude. In other words, the downward ICij binds
whenever both types are interested in similar features of the data, but one type is
willing to pay more than the other one. The idea behind the proof is that as long
as the IC constraints are not too tight the seller only modifies the statistics a little
away from the buyer’s optimal ones and only slightly reduces their prices. In this
case, by continuity, all the upward IC constraints are still satisfied.

Proposition 2 Suppose there exists a collection of small positive numbers {ϵij } such
2
that cos2 ( β ij ) ∥γi ∥2 ≤ γj + ϵij ∀i < j. Then the solution to the relaxed problem
2
that considers only the constraints with i < j such that cos2 ( β ij ) ∥γi ∥2 > γ j is the
solution to the seller’s problem.25

The previous proposition only imposes restrictions on the value that types
assign to each others’ preferred statistics. However, in this environment, it is very
important how types evaluate any statistic that is offered by the seller. To give
further structure, I define an order concept for types that depends only on the
angle between the types’ preferred statistics.

−1
Definition 2 Types are ordered if for any i < j < k, cos2 ( β ik ) < ∏kj= 2
i cos ( β j,j+1 ).

25 A straightforward corollary of the result is that in this case type 1 receives an undistorted
statistic and type n always receives zero rent.

19

Electronic copy available at: [Link]


Remember that type i’s willingness to pay for type j’s preferred statistic is
exactly equal to cos2 ( β ij ) ∥γi ∥2 . This property generates an order for type i’s will-
ingness to pay for other types’ preferred statistic. Type i’s incentive to mimic a
type k with lower valuation is smaller than the composition of all types between
type i and k incentive to mimic the closest type with lower valuation. This condi-
tion is a generalization of a Hotelling model in R2 with three types, in which these
types’ γ vectors are in the same quadrant, and their magnitudes are decreasing
from left to right (right to left).
The next proposition shows that if types are ordered and all types’ willingness
to pay for all data is not too different, it is sufficient to consider only adjacent
downward IC constraints. In this case, it is possible to completely characterize
the optimal mechanism. When the IC constraint from type i to type i + 1 binds,
without adding independent noise, the seller distorts the statistic targeted to type
i + 1 away from type i’s preferred statistic.

Proposition 3 Suppose types are ordered. There exists ϵ > 0 such that if ∥γi ∥2 cos2 ( β i,i+1 )
< (1 + ϵ) ∥γi+1 ∥2 ∀i, then the solution to the relaxed problem that considers only the
local downward constraints is the solution to the seller’s problem.
Further, if the constraint from type i to type i + 1 binds, there exists ci,i+1 ∈ (−c̄i,i+1 , c̄i,i+1 )
with c̄i,i+1 < 1 such that the seller offers to type i + 1 the statistic γi+1 − ci,i+1 γi .26,27

To gain more intuition I consider the case with only two information types,
a special case in which types are ordered. In this case it cannot be that both IC
constraints bind simultaneously, since Theorem 1 implies that at least one type
must receive her optimal statistic. As type 1 is willing to pay more than type 2 for
all data it has to be that the constraint IC21 does not bind. Further, by Theorem 1
26 The direction of the distortion which is measured by the parameter ci,i+1 depends on the sign
ofγiT γi+1 . If this sign is positive ci,i+1 is positive and if it is negative ci,i+1 is negative.
27 The parameter c the multiplier for the constraint ICi,i+1 . And the value of
i,i +1 depends on
this multiplier might depend on the value of the multipliers for the local downward constraints
IC12 , . . . , ICi−1,i .

20

Electronic copy available at: [Link]


at least one type receives zero rent, so that it has to be that the IR constraint for
type 2 always bind.
In case the IC constraint for type 1 reporting type 2 does not bind, which hap-
pens when cos2 ( β) ∥γ1 ∥2 < ∥γ2 ∥2 , the seller can attain the first best: He offers
to each buyer the conditional expectation of the variable she is interested in and
charges her a price equal to her willingness to pay for it.28 In the opposite case,
such a simple mechanism is not implementable. Since adding uncorrelated noise
is never optimal, the data broker optimally weights between two possibilities: Re-
duce the price targeted to type 1 or deteriorate the quality of the statistic targeted
to type 2 by modifying the coefficients vector L2 .
Corollary 1 presents a closed-form solution for the optimal mechanism. The
seller deteriorates the informativeness of the statistic targeted to type 2, so that the
statistics that he offers are less correlated than they would be if he could perfectly
discriminate among the buyers. To attain this objective, the seller obfuscates the
statistic targeted to type 2 in the direction opposite to type 1’s preferred statistic.
The argument in the proof shows that there exists α̃ ∈ (0, 1) such that the
constraint IR1 does not bind if and only if α < α̃. When type 1’s IR constraint
does bind, the data broker cannot increase the price targeted to type 1. Therefore,
the seller does not have an incentive to distort further the statistic targeted to
type 2, resulting in type 2 always receiving a statistic that provides her some
information. This contrasts with the literature in quality degradation, where there
are always instances in which the low type is offered the worst feasible quality
(Mussa and Rosen, 1978; Maskin and Riley, 1984). This difference results from
the ranking of types according to their valuations being statistic-dependent in
28 Offeringthe rough data to both types is never optimal since it allows both players to obtain
their optimal statistics, but the seller cannot differentiate prices, so the highest willingness to pay
buyer is going to mimic the other type and buys the cheapest package. If instead the seller offers
the optimal statistics E[θ1 | X ] and E[θ2 | X ], he can differentiate among buyers and charge each
of them their willingness to pay for this optimal statistic, which is a clear improvement for the
seller.

21

Electronic copy available at: [Link]


my environment, while in the literature it has been normally assumed that it is
product-independent.

Corollary 1 Suppose the constraint IC12 binds. Then, there exist α̃ ∈ (0, 1) and c ∈
(0, α̃) such that the optimal mechanism is a linear mechanism with L1 = γ1 and

i. L2 = γ2 − cγ1 if α < α̃; and

ii. L2 = γ2 − α̃γ1 if α ≥ α̃.

In the first case the constraint IR1 does not bind, but it does in the second one.

Figure 2 illustrates the optimal mechanism. The vectors γ1 and γ2 ’s directions


represent type 1 and 2’s optimal statistic, respectively; and their magnitude rep-
resents their willingness to pay for them. The vectors L̃1 and L̃2 represent the
statistics targeted to type 1 and 2, respectively; and their magnitude represents
the price that the seller charges for these statistics. The figure shows the main
properties of the optimal mechanism: Type 1 always receives the optimal statistic,
the statistic targeted to type 2 is distorted away from type 1’s preferred statis-
tic, and, generically, each statistic contains information form each variable that is
available, since, generically, all entries in the vectors L̃1 and L̃2 are non-zero.

Figure 2: Optimal mechanism when γ1 = (6, 2.4) and γ2 = (4, 3) for different
values of α1 .

22

Electronic copy available at: [Link]


The same figure shows that the quality degradation of the statistic targeted to
type 2 and type 1’s information rent (the difference between type 1’ willingness
to pay and p(1)) depends on the probability of facing type 1, α1 . Corollary 2
shows that as α1 increases, the data broker reduces type 1’s information rent, and
degrades the quality of the statistic targeted to type 2 further.

Corollary 2 The optimal mechanism satisfies the following properties:

i. the quality of the statistic targeted to type 2 is decreasing in α1 , and

ii. type 1’s information rent is nonincreasing in α1 .

Finally , it is possible to analyze which data set the data broker would acquire
if he can choose from two data sets that have the same cost and for which both
types are willing to pay the same. As our intuition will suggest, Corollary 3 shows
that the seller prefers to acquire the data set that generates the smallest correlation
between the conditional expectations of θ1 and θ2 given X.

Corollary 3 Suppose there are two data sets X and X ′ such that for i ∈ {1, 2}, type i’s
willingness to pay for the statistic E[θi | X ] and for the statistic E[θi | X ′ ] is the same. If
ρ2 (E[θ1 | X ], E[θ2 | X ]) < ρ2 (E[θ1 | X ′ ], E[θ2 | X ′ ]), then the seller’ profits are larger
by acquiring X than by acquiring X ′ .

This nice intuition is true as a result of types being ordered. In the opposite case,
the problem is more difficult since the set of binding constraints is endogenous
and cannot be characterized. In particular, by decreasing the price of a statistic or
by distorting its direction, the seller might create an incentive for another type to
want to buy this statistic. This makes it difficult to obtain more general results.
In the following example, I consider a very simple case, in which data is two-
dimensional (X ∈ R2 ) and there are only three information types, to show how
general these difficulties are and some of their implications.

23

Electronic copy available at: [Link]


γ1 γ1 γ1
γ2 γ2

γ3 γ2
γ3 γ3

Figure 3: Possible configurations with three information types on R2 .

Example 1 Suppose data X is two-dimensional and there are three information types.
Figure 3 represents the cases that might occur under these assumptions. In the figure on
the left, types are ordered: Type 2 is to the right of 1, and 3 is to the right of 2. In the other
two figures, types are not ordered: In the one in the center, type 3 is located in the middle
and in the one on the right, type 1 is located in the middle. By modifying the magnitude
of γ2 and γ3 , we obtain all possible configurations for the seller’s problem.
Figure 4 presents which IC constraints bind for each of the cases above and for all
possible magnitudes of γ2 (v2 ) and γ3 (v3 ). I vary both v2 and v3 between 0 and 1, with
v3 ≤ v2 , and fix the magnitude of γ1 to be equal to 1.29
The figure has many interesting implications. First, the set of IC constraints that bind
varies a lot with v2 and v3 , and the limits of the regions defining those sets are non-linear.
Therefore, it is difficult to obtain closed form characterizations for those regions, but it
is easy to solve for them numerically due to the optimality of linear contracts. Second,
in many cases, some upward IC constraints bind, that is, constraints from types that are
willing to pay less for all data to those who are willing to pay more. This results from
the seller offering a price below the monopolist price for the mimicked type and/or from
modifying a statistic on the direction towards the type with lower willingness to pay’s
preferred statistic.
Furthermore, some sets of binding IC constraints generate contracts that exhibit some
properties we do not obtain in other environments. First, when both IC12 and IC32 bind
29 For this figure, I assume that the angle between any two consecutive vectors is π/6. The
patterns are very similar when considering other angles.

24

Electronic copy available at: [Link]


(a) (b) (c)

Figure 4: Pattern of binding IC constraints with three information types.

(IC13 and IC23 ) the seller, in some occasions, finds it optimal no to target an informative
statistic to type 2 (3, respectively), i.e., exclude this type from the mechanism. Interest-
ingly, the seller might not exclude the type with the lowest willingness to pay for all data,
but he excludes the type that other types have a higher incentive to mimic.
Second, which types are ”at the top” or ”at the bottom” is not completely determined
by types willingness to pay for all data. Consider the case when constraints IC13 and
IC21 bind. These constraints bind simultaneously when type 1 is in the middle, type 3’s
willingness to pay is low, and type 2’s willingness to pay is high. In this case, the seller
wants to reduce the price charged to type 1, so that she does not have an incentive to mimic
type 3. However, this reduction in p(1) creates an incentive for type 2 to mimic type 1.
To reduce this incentive, the seller modifies the direction of the statistic targeted to type 1
away from type 2’s preferred statistic. Therefore, in this case, type 2 is the only type that
receives an undistorted statistic, converting her into the type ”at the top”, even when she
does not have the largest willingness to pay for all data. Similarly, when only constraints
IC12 and IC32 bind and type 2 is not excluded, the only type who receives zero rent is
type 2, making her the type ”at the bottom”.

25

Electronic copy available at: [Link]


7 Mix model: Two Information Types and a Contin-
uum of Valuation Types

I introduce heterogeneity in the buyers’ valuations to study vertical differentia-


tion. As buyers’ payoffs are non-linear, I cannot use the common tools that are
implemented in mechanism design (see Rochet and Choné (1998)) to solve for
the optimal mechanism. Instead, I use a novel direct approach that allows me to
identify the set of sufficient and necessary IC constraints.30
Suppose there are two information types {1, 2}, and that for each information
type there is a continuum of valuation types which are conditionally distributed
according to absolutely continuous distributions G (v | 1) and G (v | 2), with
densities g(v | 1) and g(v | 2), and supports [v1 , v̄1 ] and [v2 , v̄2 ], respectively. Since
¯ ¯
I do not impose any restriction in the supports of these distributions, without
loss of generality, I normalize γ1T γ1 = γ2T γ2 = 1. Further, I suppose that these
distributions satisfy the Myerson (1981)’s Regularity Condition.

Assumption 2 For each i ∈ {1, 2},


1 − G (v | i )
v−
g(v | i )

is nondecreasing in v.

I restrict the analysis by assuming that the seller can only offer linear mecha-
nisms, as defined before. To simplify notation, I let q jiv be the reduction of θ j ’s
forecast error variance after observing the linear statistic targeted to type (i, v);
T γ )2
( Liv
that is, q jiv = T
j
Liv Liv +ℓ2iv
≤ 1, since γiT γi = 1. This variable measures the quality,
as perceived by type j, of the statistic targeted to type i. This environment differs
from other multidimensional environments in that different types perceive the
30 In
a related environment, Smolin (2019) uses a direct approach to find the optimal mecha-
nism. However, the steps he needs to implement and the optimal mechanism are different from
the ones I use. Furthermore, the optimal mechanisms differ as well.

26

Electronic copy available at: [Link]


quality of a statistic differently.
The problem that the seller faces has three different sets of constraints: the
traditional IR constraints, the vertical IC constraints between two types with dif-
ferent valuations, but the same information type, and the crossed IC constraints
between different information types. The first two sets of constraints are stan-
dard. In particular, analogous to the result in Section 5, if none of the crossed IC
constraints binds, the optimal mechanism is Myerson’s solution: The seller offers
a menu with two packages. Each package is targeted to a specific information
type and contains a take-it-or-leave-it offer for a distinct statistic that allows her
to recover the conditional expectation she is interested in, i.e., qiiv = 1.31 I denote
by v1∗ and v2∗ the take-it-or-leave-it prices, and, without loss of generality, assume
that v1∗ > v2∗ . If we let β be the angle between γ1 and γ2 , this will be optimal
as long as v1∗ cos2 ( β) ≤ v2∗ , i.e., when the angle between both information types’
optimal statistics is large and/or the difference between v1∗ and v2∗ is small.
In the opposite case, when some of the crossed IC constraints bind, it is not
possible to know a priori which of them actually bind. I use a direct approach
that allows me to identify the set of relevant IC constraints. For this, it is key to
reconsider the intuition that the angle between the vectors L and γi completely
determines how informative the statistic L T X is for type i. This allows me to
concisely express how informative the statistic targeted to type (2, v) is for type 1,
as a function of how informative it is for type 2, q22v , and the angle β between their
preferred vectors, γ1 and γ2 . This yields a simple expression that summarizes
how much each type is willing to pay for any linear statistic the seller offers.
With this is mind, Theorem 4 provides a complete characterization of the allo-
cation q in the optimal mechanism. Figure 5 presents an example.


Theorem 4 Let h(q22v ) := cos2 ( β + cos−1 ( q22v )). There exist cutoffs v̂1 , v̂2 , and
31 As pointed out before in footnote 28, offering all data to each information type is never

optimal.

27

Electronic copy available at: [Link]


v̂2 q22v̂2
ṽ2 with v2∗ < v̂2 < v̂1 = h(qv̂2 )
< v1∗ , and ṽ2 > v̂2 defined by the implicit equation
v1∗ h′ (q22ṽ2 ) = ṽ2 , such that in the optimal mechanism
 


 0 if v < v̂1 

 0 if v < v̂2
 
q11v = h(q22v̂2 ) if v ∈ [v̂1 , v1∗ ) and q22v = q22v̂2 if v ∈ [v̂2 , ṽ2 ) .32
   
 h′ −1 v∗
 
if v ≥ v1∗ if v ≥ ṽ2
 1
 
v 1

Figure 5: Optimal Mechanism when the angle between γ1 and γ2 is equal to 0.07̄π,
α1 = 0.85, v1 = v2 = 0, v̄1 = 1, v̄2 = 2.4, g(v | 1) = 15v14 and g(v | 2) = 0.83̄ for
v < 0.8 and g(v | 2) ∝ 4(2.4 − v)3 for v > 0.8.

In the optimal mechanism, a buyer with type (1, v) with v > v1∗ , still receives
a fully informative statistic. However, to reduce the incentive of type 1 to hori-
zontally mimic type 2, the seller decreases the quality of the statistic targeted to
types (2, v) by increasing the minimum valuation type v to whom he sells above
Myerson’s threshold, v2∗ , and by never offering an undistorted statistic to these
types. Further, the seller vertically differentiates between types (1, v) by offering
those types with v < v1∗ a distorted statistic, and might vertically differentiate
among types (2, v) by offering them a continuum of statistics.
A comment is in order. The theorem describes the quality of the statistic that
is targeted to each type, but it does not specify the actual statistics that are of-
32 Thisexpression is well defined, since the function h is infinitely differentiable and the inverse
of any of its derivatives exists.

28

Electronic copy available at: [Link]


fered. Types (1, v) with v > v1∗ receive the optimal statistic, γ1T X, and to types
(2, v) with v > v̂2 the seller targets, without adding uncorrelated noise, a linear
combination that distorts the coefficients on the data away from information type
1’s preferred direction: γ2 − cγ1 for some c > 0. There are multiple statistics the
seller might offer to types (1, v) with v ∈ [v̂1 , v1∗ ]: modify the statistic in the op-
posite direction of information type 2’s preferred direction, include uncorrelated
noise to information type 1’s preferred linear combination, or offer a statistic that
includes both distortions.
It is worthy to notice how this result differs from the one in Bergemann, Bon-
atti, and Smolin (2018). In their environment, all buyer types want to learn about
the same variable, but they have different valuations according to their priors.
In contrast, in my environment all buyers have the same prior, but differ in two
dimensions: which variable they want to learn and how valuable, once we fixed
what they want to learn, is any information that is provided to them.
This difference affects the optimal mechanism in multiple ways. First, there
is an implementation of the optimal mechanism in which some of the statistics
include independent noise, while including uncorrelated noise is strictly sub-
optimal in their environment. Second, the seller needs to provide information
to some valuation types that would not acquire any information if all buyers
would want to forecast the same variable. Third, in here, the seller might need
to offer a continuum of linear statics. Each linear statistic can be thought of as a
randomization over each dimension of the data, so that in those cases the seller
offers a continuous randomization over the space of statistics.33 And fourth, the
set of binding IC constraints is more involved. In particular, some IC constraints
across information types might bind. Therefore, there is not pattern of local IC
constraints that is sufficient for the problem.
33 In other multidimensional and substitutable goods environments such randomization is
needed (Daskalakis, Deckelbaum, and Tzamos, 2017; Thanassoulis, 2004; Pavlov, 2011; Balestrieri,
Izmalkov, and Leao, 2015). Further, Zhong (2018) has shown that in information environments
this is a result of the richness of the set forecasts that the buyer can make.

29

Electronic copy available at: [Link]


References

Anderson, S. P., and L. Celik (2015): “Product line design,” Journal of Economic
Theory, 157, 517–526.

Anderson, T. (2003): “An Introduction to Multivariate Statistical Analysis (Wiley


Series in Probability and Statistics),” July 11.

Armstrong, M. (1996): “Multiproduct nonlinear pricing,” Econometrica: Journal of


the Econometric Society, pp. 51–75.

Babaioff, M., R. Kleinberg, and R. Paes Leme (2012): “Optimal mechanisms


for selling information,” in Proceedings of the 13th ACM Conference on Electronic
Commerce, pp. 92–109. ACM.

Balestrieri, F., S. Izmalkov, and J. Leao (2015): “The market for surprises: sell-
ing substitute goods through lotteries,” .

Ball, I. (2019): “Scoring Strategic Agents,” arXiv preprint arXiv:1909.01888.

Basar, T. (2008): “Variations on the theme of the Witsenhausen counterexample,”


in 2008 47th IEEE Conference on Decision and Control, pp. 1614–1619. IEEE.

Bergemann, D., and A. Bonatti (2015): “Selling cookies,” American Economic


Journal: Microeconomics, 7(3), 259–94.

Bergemann, D., A. Bonatti, and A. Smolin (2018): “The design and price of
information,” American Economic Review, 108(1), 1–48.

Birbil, Ş. İ., J. B. G. Frenk, and G. J. Still (2007): “An elementary proof of the
Fritz-John and Karush–Kuhn–Tucker conditions in nonlinear programming,”
European journal of operational research, 180(1), 479–484.

Chamberlain, G. (1983): “A characterization of the distributions that imply


mean-Variance utility functions,” Journal of Economic Theory, 29(1), 185–201.

30

Electronic copy available at: [Link]


Daskalakis, C., A. Deckelbaum, and C. Tzamos (2017): “Strong Duality for a
Multiple-Good Monopolist,” Econometrica, 85(3), 735–767.

Deimen, I., and D. Szalay (2019): “Delegated expertise, authority, and communi-
cation,” American Economic Review, 109(4), 1349–74.

Fang, K., S. Kotz, and K. Ng (1990): Symmetric multivariate and related distributions.
Chapman and Hall.

Federal Trade Commission (2014): “Data brokers: A call for transparency and
accountability,” Washington, DC.

Fritz, J. (1948): “Extremum problems with inequalities as subsidiary conditions,


Studies and Essays: Courant Anniversary Volume, KO Friedrichs, OE Neuge-
bauer, and JJ Stoker, eds,” .

Kamenica, E., and M. Gentzkow (2011): “Bayesian persuasion,” American Eco-


nomic Review, 101(6), 2590–2615.

Mailath, G. J., and G. Nöldeke (2008): “Does competitive pricing cause mar-
ket breakdown under extreme adverse selection?,” Journal of Economic Theory,
140(1), 97–125.

Manelli, A. M., and D. R. Vincent (2006): “Bundling as an optimal selling


mechanism for a multiple-good monopolist,” Journal of Economic Theory, 127(1),
1–35.

Maskin, E., and J. Riley (1984): “Monopoly with incomplete information,” The
RAND Journal of Economics, 15(2), 171–196.

Mussa, M., and S. Rosen (1978): “Monopoly and product quality,” Journal of
Economic theory, 18(2), 301–317.

31

Electronic copy available at: [Link]


Myerson, R. B. (1981): “Optimal auction design,” Mathematics of operations re-
search, 6(1), 58–73.

Osborne, M. J., and A. Rubinstein (1994): A course in game theory. MIT press.

Owen, J., and R. Rabinovitch (1983): “On the class of elliptical distributions and
their applications to the theory of portfolio choice,” The Journal of Finance, 38(3),
745–752.

Pavlov, G. (2011): “Optimal mechanism for selling two goods,” The BE Journal of
Theoretical Economics, 11(1).

Rochet, J.-C., and P. Choné (1998): “Ironing, sweeping, and multidimensional


screening,” Econometrica, pp. 783–826.

Smolin, A. (2019): “Disclosure and pricing of attributes,” Available at SSRN


3318957.

Stokey, N. L. (1979): “Intertemporal price discrimination,” The Quarterly Journal


of Economics, pp. 355–371.

Thanassoulis, J. (2004): “Haggling over substitutes,” Journal of Economic theory,


117(2), 217–245.

Vohra, R. V. (2011): Mechanism design: a linear programming approach, vol. 47.


Cambridge University Press.

Wilson, R. B. (1993): Nonlinear pricing. Oxford University Press on Demand.

Witsenhausen, H. S. (1968): “A counterexample in stochastic optimum control,”


SIAM Journal on Control, 6(1), 131–147.

Yang, K. H. (2018): “Non-linear Pricing of Information Structure: An Equivalence


Result,” .

Zhong, W. (2018): “Selling Information,” arXiv preprint arXiv:1809.06770.

32

Electronic copy available at: [Link]


A A Micro Foundation for Quadratic Loss

Suppose there is a monopolist that faces a linear demand q( p) = ā − [Link]


monopolist knows b, but he is uncertain about the value of ā. If the monopolist

knew that the value of the intercept is ā, he would produce a quantity q̄ = 2 and
ā2
his profits would be equal to π̄ = 4b .
Suppose that when the monopolist is uncertain about ā he chooses some quan-
tity q. This quantity would be optimal if a = 2q, that is, the monopolist chooses
the quantity q as if he thinks that ā = a. By choosing this quantity the monopolist
q a2
obtains profits equal to π = q bā − b = 2a bā − 2b
a a ā

= 2b − 4b .
Then, the monopolist is leaving on the table the quantity
ā2 a ā a2 ( a − ā)2
π̄ − π = − + = ,
4b 2b 4b 4b

that is, the monopolist’s losses due to its uncertainty about the parameter ā cor-
respond to a quadratic loss function.

B Proofs

Proof of Claim 2
Let S be the set containing all realizations of a statistic ψ. For any functional
H : S → R,

−EX,ϵ,θ ( H (ψ) − θ )2 | ψ ≤ −EX,ϵ,θ (E(θ | ψ) − θ )2 | ψ = −EX,ϵ [Var (θ | ψ)].


   

The inequality follows from E ( H (ψ) − θ )2 |ψ being minimized by setting H (ψ) =


 

E[θ |ψ], and the equality follows from the Law of Iterated Expectations and the
definition of conditional variance.
Proof of Lemma 1
Consider the family of functions k( L T X + ℓϵ) for k ∈ R. Since all variables

33

Electronic copy available at: [Link]


have mean equal to 0, Claim 1 implies that E[θi | L T X + ℓϵ] belongs to this family.
When using k( L T X + ℓϵ) as forecast, θi ’s forecast variance error is equal to

EX,ϵ [Eθi [(θi − kL T X − k ℓϵ)2 | L T X + ℓϵ]]

= Eθi [θi2 ] − 2kLT EX,ϵ [Eθi [θi X | LT X + ℓϵ]] + k2 LT EX,ϵ [Eθi [ XX T | LT X + ℓϵ]] L + k2 ℓ2 Eϵ [ϵ2 ]

= Var (θi ) − 2kLT Cov(θi , X ) + k2 LT L + k2 ℓ2 ,

where the first equality follows from ϵ being independent of X and θ, and he sec-
ond equality follows the Law of Iterated Expectations, the definition of variance
and covariance, and that Var ( X ) = Ik and Var (ϵ) = 1. By Claim 2, the conditional
expectation minimizes this expression. The First Order Condition with respect to
k implies that
L T Cov(θi , X )
k̂ = .
L T L + ℓ2
Plugging in this value of k, we obtain that the forecast variance is equal to
( L T Cov(θi , X ))2
Var (θi ) − .
L T L + ℓ2

The result follows from type (i, v)’s willingness to pay being equal to the differ-
ence between the prior and the posterior forecast variance times v.
Proof of Theorem 1
To prove part 1. suppose that all types receive a positive surplus bounded
away from zero and let P = infi,v {vEX,ϵ [Var (θi | (ψ(i, v)))] − p(i, v)} > 0. Then
the data broker can charge new prices p̃(i, v) = p(i, v) + P, so that all IC con-
straints are unaffected and the IR constraints are satisfied. This modification of
the mechanism clearly increases the seller’s profits.
To prove part 2. suppose by contradiction that in the optimal mechanism no
type receives all data or a statistic that for her is equivalent to receiving all data.
Let R(i, v) be type (i, v)’s information rent in this mechanism and let V (i, v) be
type (i, v)’s willingness to pay for all data. Define (i∗ , v∗ ) ∈ arg maxi,v V (i, v) −

34

Electronic copy available at: [Link]


R(i, v) be the type that is willing to pay a larger extra amount for all the data.
Consider the alternative mechanism that offers to type (i∗ , v∗ ) all data at a
price p′ (i∗ , v∗ ) = V (i∗ , v∗ ) − R(i∗ , v∗ ) > p(i∗ , v∗ ), and keep unchanged all other
prices and statistics. Since all types receive the same rents and p′ (i∗ , v∗ ) is set
to be high enough such that type ( j, v) ̸= (i∗ , v∗ ) does not have an incentive to
report type (i∗ , v∗ ) all constraints are still satisfied. Therefore, the new mechanism
is feasible and gives higher profits to the seller than the original mechanism, a
contradiction.
Proof of Theorem 2
Without loss of generality, I normalize the valuation type v to 1. To simplify
notation, I do not indicate any dependence on v. To facilitate comprehension, in
Steps 0-5, I first present a detailed proof for the case with only two information
types. Then, in Step 6, I argue why this argument generalizes.
Step 0: Reducing the set of constraints to consider
I argue that the IR constraint for type 2, IR2 , always binds and that the IC
constraint for type 2 mimicking type 1, IC21 , never binds.
First, it cannot be that both IC12 and IC21 bind because then both types would
receive a distorted statistic, which contradicts Theorem 1.
Then, suppose by contradiction that only IC21 binds. Therefore, type 2 receives
the statistic E[θ2 | X ] and pays p(2) = p(1) − Var (θ2 | X ) + E(X,ϵ) [Var (θ2 | ψ(1))].
As the statistic ψ(1) can at most provide the same information to type 2 as the
statistic E[θ2 | X ] does, −Var (θ2 | X ) + E(X,ϵ) [Var (θ2 | ψ(1))] > 0. Therefore,
p(1) < p(2). Further, IR2 implies that p(2) ≤ Var (θ2 ) − Var (θ2 | X ). But, then
the seller can offer all data to both types and charge the price Var (θ2 ) − Var (θ2 |
ψ(2)) to each type, which trivially satisfies all constraints and increases profits.
Therefore, the constraint IC21 never binds.
This immediately implies that the constraint IR2 always binds; if this were not
the case, the seller can slightly increase p(2) so that IC21 is still satisfied. This

35

Electronic copy available at: [Link]


increases the seller’s profits and relaxes the constraint IC12 .

If only the constraints IR1 and IR2 bind, the problem is straightforward be-
cause the seller can offer the optimal statistics to both types at the monopolist
prices. These optimal statistics correspond to the conditional expectations, which
are linear since the joint distribution of the ( X, θ, ϵ) is elliptical.
In Step 1-4, I assume that only constraints IC12 and IR2 bind, and show that
the optimal mechanism is still linear. In Step 5, I extend the analysis to include
the case in which constraints IC12 , IR1 and IR2 bind simultaneously.
Step 1: It is without loss of generality only considering recommendation
mechanisms.
In a recommendation mechanism, buyer i’s forecast, ai , is equal to the statistic
sold by the seller to her, ψ(i ). As Claim 2 establishes that ai = E[θi | ψ(i )( X = x )],
in a recommendation mechanism we would have ψ(i )( x ) = E[θi | ψ(i )( X = x )].
Considering only this kind of mechanisms is without loss of generality by an
argument similar to the one in Kamenica and Gentzkow (2011). First, Claim 2
implies that type i learns the same either by observing ψ(i ) or E[θi | ψ(i )( X )].
Moreover, type j cannot learn more about θ j by observing E[θi | ψ(i )( X )] than
by observing ψ(i ). To see this, let a( j, E[θi | ψ(i )( X )]) be type j’s forecast when
observing E[θi | ψ(i )( X )], and notice that type j can recover a( j, E[θi | ψ(i )( X )])
from ψ(i ) in a two-step estimation: first calculate E[θi | ψ(i )( X )] and them, from
it, create a( j, E[θi | ψ(i )( X )]).
Therefore, if the seller offers E[θi | ψ(i )] to type i ratter than ψ(i ) and charges
her the same price as before, type i’s incentives are unchanged and all ICji con-
straints are still satisfied.
Step 2: The mechanism design problem is equivalent to a sequential strictly
competitive game.
The problem is a sequential game in which the seller first chooses a mechanism
that satisfies all constraints and then each buyer type decides how to use the

36

Electronic copy available at: [Link]


statistic she has acquired. Denote by a(i, j, ψ( j)) type i’s forecast when he mimics
type j and acquires the statistic ψ( j). Notice that these forecasts are random
variables that, indirectly, depend on the realization of X through the statistic ψ( j).
Remember I am assuming that only the constraints IC12 and IR2 bind. As
constraint IR2 binds the seller charges to type 2 the price

p(2) = Var (θ2 ) − E[(θ2 − a(2, 2, ψ(2)))2 ],

which generates a payoff equal to zero for type 2. Further, as there is not IC
constraint pointing towards type 1, the seller offers all information to her, by
setting ψ(1) = E[θ1 | X ]. Thus, and as constraint IC12 binds, the seller charges to
type 1 the price

p(1) = E[(θ1 − a(1, 2, ψ(2)))2 ] − E[(θ1 − a(1, 1, ψ(1)))2 ] + p(2)

= E[(θ1 − a(1, 2, ψ(2)))2 ] − E[(θ1 − a(1, 1, ψ(1)))2 ] + Var (θ2 ) − E[(θ2 − a(2, 2, ψ(2)))2 ].

Therefore, given such a mechanism, type 1’s payoffs are equal to

Var (θ1 ) − E[(θ1 − a(1, 1, ψ(1)))2 ] − p(1)

=Var (θ1 ) − E[(θ1 − a(1, 2, ψ(2)))2 ] − Var (θ2 ) + E[(θ2 − a(2, 2, ψ(2)))2 ]

Type 1 chooses her forecasts a(1, ·, ·) to maximize her payoff. However, the
only component of her forecasts that appears in her payoff is a(1, 2, ψ(2)). Notice
that whenever type 1 (2) buys the statistic ψ(1) (ψ(2)), by Claim 2, it is sequen-
tially rational for her to choose the optimal updating rule a(1, 1, ψ(1)) = E[θ1 |
ψ(1)] = E[θ1 | X ] ( a(2, 2, ψ(2)) = E[θ2 | ψ(2)]). With this mechanism and buyer’s
forecasts, the seller’s profits are equal to
 
α1 −E[Var (θ1 | X )] + E[(θ1 − a(1, 2, ψ(2)))2 ] + Var (θ2 ) − E[Var (θ2 | ψ(2))], 34

where α1 is the probability that the seller assigns to type 1. The chooses the
34 Notice that as a(1, 1, ψ(1)) = E[θ1 | X ] and a(2, 2, ψ(2)) = E[θ2 | ψ(2)] we have E[(θ1 −
a1 (1, 1, ψ(1)))2 ] = E[Var (θ1 | X )] and E[(θ2 − a2 (2, 2, ψ(2)))2 ] = E[Var (θ2 | ψ(2))].

37

Electronic copy available at: [Link]


statistic ψ(2) to maximize his profits.
Consider the functional

G ( a(1, ψ(2)), ψ(2)) = α1 E[(θ1 − a(1, 2, ψ(2)))2 ] − E[Var (θ2 | ψ(2))].

Notice that for the buyer it is equivalent to choose a(1, 2, ψ(2)) to minimize G or
to choose a(1, 2, ψ(2)) to maximize her payoff, and for the seller it is equivalent to
choose ψ(2) to maximize his profits or to maximize G. Therefore, the functional G
defines a sequential stochastic strictly competitive game equivalent to the mech-
anism design problem, in which type 1 chooses a(1, 2, ψ(2)) to minimize G and
the seller chooses ψ(2) to maximize G.
Step 3: The sequential strictly competitive game has an equilibrium in lin-
ear strategies.
I divide the proof for this step in three parts. First, I show that best responses
to linear strategies, whenever they exist, are linear. Then, I show that in the set of
linear mechanisms, it is without loss to focus on linear mechanisms whose vector
of coefficients magnitude is equal to 1. As this set is compact, an application of
the Maximum Theorem implies that the seller’s best response exists when he is
restricted to choose a linear mechanism. Finally, an application of the Minmax
Theorem implies that this linear solution is actually the equilibrium of the strictly
competitive game, i.e., it is the solution to the mechanism design problem.
Step 3.1: Set of linear mechanisms is closed with respect to best responses.
The following definition states what I mean by a linear mechanism to be closed
with respect to best responses.

Definition 3 The set of linear mechanisms is closed with respect to best responses if

1. type 1’s best response to a linear statistic is a linear forecast, and

2. the seller’s best response, whenever it exists, to a linear forecast is a linear statistic.

It is straightforward to argue that type 1’s best response to a linear statistic

38

Electronic copy available at: [Link]


exists, is unique and it is a linear updating rule. Suppose ψ(2) is a linear com-
bination of X. Claim 2 implies that a∗ (1, 2, ψ(2)) = E(θ1 | ψ(2)) and Claim 1
implies that E(θ1 | ψ(2)) is linear. Therefore, a∗ (1, 2, ψ(2)) is a linear function of
ψ (2).
Now, suppose that a(1, 2, ψ(2)) is linear, that is, a(1, ψ(2)) = c1 ψ(2) + c0 . Then
the seller wants to maximize
 
E α1 (θ1 − c1 ψ(2) − c0 )2 − (θ2 − ψ(2))2 | X .

Since the objective function is quadratic, the point-wise necessary First Order
Condition delivers a linear function of the conditional expectations:
E[ θ2 | X ] − α1 c1 E[ θ1 | X ] − α1 c0 c1
ψ ∗ (2) = . (1)
1 − α1 c21

This statistic is linear because Claim 1 implies that these conditional expectations
are linear. Note that I have only proved that if there is a statistic that solves
the seller’s problem this statistic has to be linear, i.e., this is a necessary, but no a
sufficient condition. This is enough for my purposes since I only want to conclude
that if there is a mechanism that maximizes the seller’s profits when the buyer
uses a linear forecast that maximizing mechanism is a linear mechanism.
Step 3.2: In the set of linear mechanisms, it is without loss to only consider
linear statistics whose coefficient vectors have magnitude equal to one.
This result is a direct application of Lemma 1. Suppose that the seller offers
the statistic L T X + ℓϵ, with ( L, ℓ) ̸= 0. Lemma 1 implies that the buyer obtains the
L ℓ
same value from the alternative statistic L̃ T X + ℓ̃ϵ with L̃ = ∥( L,ℓ)∥
and ℓ̃ = ∥( L,ℓ)∥
,
independently of which variable θ she wants to forecast. Therefore, from the point
of view of the seller both statistics are equivalent as well.
If the seller offers the uninformative statistic L T X + ℓϵ with ( L, ℓ) = 0 he
can equivalently offer the uninformative statistic the only contains noise, i.e., the
statistic L̃ T X + ℓ̃ϵ with L̃ = 0 and ℓ̃ = 1.

39

Electronic copy available at: [Link]


In both cases, the new vector ( L̃, ℓ̃) has magnitude equal to 1.
Step 3.3: There exists a local strict minmax point in linear strategies of the
strictly competitive game.
From Step 3.1 we know that the buyer’s best response to any mechanism al-
ways exist and it is unique. Further, if the seller offers a linear mechanism then
the buyer’s best response is a linear forecast.
Further, Step 3.1 guarantees that for the seller it is without loss to restrict to
linear mechanisms whenever the buyer is using a linear strategy. Besides, Step
3.2 implies that in the set of linear mechanisms it is without loss to consider
only statistics whose coefficients vectors have magnitude equal to 1. This set
is equivalent to the set of vectors in Rk+1 with magnitude equal to 1, and it
is straightforward to show that this set is compact. Therefore, the Maximum
Theorem implies that, when restricting to linear mechanisms, the seller’s best
response to a linear updating rule exists. This and Step 3.1 imply that the unique
seller’s best response to a buyer’s linear forecast is a linear mechanism.
Therefore, the strictly competitive game has a local strict minmax point in
linear strategies.
Step 4: The seller’s optimal payoff can only be accrued by offering a linear
mechanism.
Strictly competitive games, in which there exist a minmax point, satisfy two
important properties. First, the order of play is irrelevant since each minmax
solution is a maxmin solution. Second, all Nash equilibria yield the same payoffs
to all players. The following lemma formalizes these statements.

Lemma 2 Suppose ( a∗ (1, 2, ·), ψ∗ (2)) is an equilibrium of the strictly competitive game
defined by G. Then mina(1,2,·) maxψ(2) G ( a(1, 2, ·), ψ(2)) = maxψ(2) mina(1,2,·) G ( a(1, 2, ·), ψ2 ),
and thus all the Nash equilibria of the game defined by G yield the same payoffs.
As a consequence, the Nash equilibria of a strictly competitive game are interchange-
able: if ( a∗ (1, 2, ·), ψ∗ (2)) and ( â(1, 2, ·), ψ̂(2)) are equilibria then so are ( a∗ (1, 2, ·), ψ̂(2))

40

Electronic copy available at: [Link]


and ( â(1, 2, ·), ψ∗ (2)).

Proof See the proof of Theorem 22.2 in Osborne and Rubinstein (1994) and the
discussion that follows after.
In Step 3, I showed that there is a minmax point in linear strategies of the zero-
sum game defined by G. Then the lemma guarantees that the linear mechanism
at this minmax point generates to the seller at least the same profits as any other
mechanism.35,36
I complete the argument by showing that there no exists a non-linear mech-
anism that generates the same profits. Let ( a∗ (1, 2, ·), ψ∗ (2)) be an equilibrium
in linear strategies and suppose there is another equilibrium ( â(1, 2, ·), ψ̂(2)).
Lemma 2 implies that ( a∗ (1, 2, ·), ψ̂(2)) is an equilibrium as well. Since ψ̂(2) is
a best response to a∗ (1, 2, ·), Equation 1, in Step 3.1, implies that ψ̂(2) is a lin-
ear statistic. Therefore, in any equilibrium of the game, the seller offers a linear
statistic, completing the argument.
Step 5: Generalization of the argument when both IR constraints bind.
Notice that only Steps 2 and 3.1 depend on which constraints do bind. I now
argue that these steps are still satisfied when both IR constraints and constraint
IC12 bind. In this case, it is still optimal for the seller to offer a full information
statistic to type 1. Further, the seller will still want to charge to type 2 the price
p(2) = Var (θ2 ) − E[(θ2 − a(2, ψ(2)))2 ]. However, now, when considering the
constraints IR1 and IC12 , the price p(1) can be expressed in two alternative ways:

p(1) = Var (θ1 ) − Var (θ1 | X ), and

p(1) = E[(θ1 − a(1, 2, ψ(2)))2 ] − Var (θ1 | X ) + Var (θ2 ) − E[(θ2 − a(2, 2, ψ(2)))2 ].
35 Strictly speaking the right equilibrium concept in the sequential game is sequential rationality.

However, in this case, beyond the requirement that the profile of strategies is a Nash equilibrium,
this concept only requires that the seller uses the statistic she buys in an optimal way, which has
been imposed already in Steps 2 and 3. I thank an anonymous referee for this observation.
36 In discrete games, the MinMax Theorem applies only if the players can randomize. In here,

that is not a problem, since both the seller’s and buyer’s action space is a dense and connected
space. I thank an anonymous referee for pointing this out.

41

Electronic copy available at: [Link]


Type 1’s payoffs remain the same when we plug in the price as in the second
expression. But, having two expressions for p(1) restricts the set of mechanisms
from which the seller can choose. In fact, by making equal both expressions for
p(1), it has to be satisfied that

Var (θ1 ) − E[(θ1 − a(1, 2, ψ(2)))2 ] = Var (θ2 ) − E[(θ2 − a(2, 2, ψ(2)))2 ].

This constraint set is non-empty. Actually, it is easy to check, by using the result
T
γ2T (γ1 −γ2 )

in Lemma 1, that the statistic ψ(2)( X ) = γ2 − γT (γ −γ ) γ1 X belongs to this
1 1 2

set. As the objective function is still quadratic and the constraint is quadratic, the
Fritz conditions,37 imply that it is necessary that the seller’s best response to a
linear forecast is a linear statistic. Further, type 1’s problem is unaffected, so that
her best response to a linear mechanism is still a linear forecast. Therefore, the
argument generalizes to this case.

Steps 1-5 show that when there are only two information types the optimal
mechanism is linear. Step 6 argues why this argument can be generalized to the
case with more information types (but still only one valuation type).
Step 6: Generalization to more than two information types.
Now, consider the case with a finite number of information types. By sequen-
tial rationality each type updates optimally when buying the statistic targeted
to her. If type i buys the statistic targeted to type j, she chooses the forecast
a(i, j, ψ( j)) .
37Fritz (1948) showed that for any solution x to a constrained optimization problem there exists
a vector of multipliers µ = (µ0 , µ1 , . . . , µn ), µ ̸= 0 and µi ≥ 0 such that
n
µ0 ∇ f ( x ) + ∑ µi ∇ gi ( x ) = 0,
i =1

where f corresponds to the objective function and ( gi )in=1 to the set of constraints. These condi-
tions generalize the Kuhn-Tucker conditions by allowing the coefficient assigned to the gradient
of the objective function to be equal to 0, and do not require a constraint qualification. Though
they might generate many candidates for the optimal point(s), they are useful in here because it
allows me to show that any critical point is linear. For a simple proof of the Fritz conditions that
apply to this environment see Birbil, Frenk, and Still (2007).

42

Electronic copy available at: [Link]


Type i’s price is pin down by the IRi constraint or/and by some ICij con-
straints, that is,

p(i ) = Var (θi ) − E[(θi − ψ(i ))2 ] or

p(i ) = −E[(θi − ψ(i ))2 ] + E[(θi − a(i, ψ( j)))2 ] + p( j) for some j ̸= i.

If one or more prices are pin down by more than one constraint, then the set of
feasible mechanisms is restricted as in Step 5. Let ψH denote the set of feasible
mechanisms.38
Notice that all prices are linear combinations of constants and quadratic terms:
E[(θi − a(i, j, ψ( j)))2 ] and −E[(θi − ψ(i ))2 ], for different i and j. Since the seller
prefers larger prices, he wants to maximize E[(θi − a(i, j, ψ( j)))2 ] and −E[(θi −
ψ(i ))2 ]. Each buyer type, however, wants to minimize E[(θi − a(i, j, ψ( j)))2 ] be-
cause the way she pays a lower price without affecting the quality of the statistic
she buys. Further, if the constraint ICij binds and type j is able to reduce the price
she pays, this reduction benefits type i as well.
Therefore, the problem reduces to a strictly competitive game between the
seller who wants to maximize prices by choosing statistics ψ(i ) for each buyer
type in the set of feasible mechanisms ψH and a fictitious player who chooses
forecasts a(i, j, ψ( j)) to minimize prices, which favors all buyer types.
Claim 2 implies that, if the seller offers linear statistics, the fictitious player
wants to choose linear forecasts. At the same time, if the fictitious player chooses
linear forecasts, the seller’s best response is to choose linear statistics. This re-
sults from applying the Fritz’ conditions to the seller’s problem, which objective
function and constraints are quadratic. Therefore, as before, the set of linear
mechanisms is closed with respect to best responses.
As Lemma 2 still applies in this strictly competitive game, the rest of the argu-
ment follows and the result is still true for this extension.
38 The set of restrictions is analogous to the restriction in Step 5. Therefore, the same class of
linear statistics satisfy these restrictions, so that the set ψH is non-empty.

43

Electronic copy available at: [Link]


Proof of Theorem 3
Fix δ > 0 with δ < min{γiT γi : i ∈ {1, . . . , n}}. For any subset of types
I ⊆ {1, . . . , n}, consider the auxiliary problem

V ( I, δ) = max ∑i ∈ I αi pi
{ pi ,Li ,ℓi }i∈ I

( LiT γi )2 ( L Tj γi )2
s.t. LiT Li +ℓ2i σ2
− pi ≥ L Tj L j +ℓ2j σ2
− pj ∀i, j ∈ I

( LiT γi )2
T
Li Li +ℓ2i σ2
− pi ≥ 0 ∀i ∈ I

pi ≥ δ ∀i ∈ I

and let V (δ) = max V ( I, δ) and I ∗ (δ) = arg max V ( I, δ).39 By Lemma 1,
I ⊆{1,...,n} I ⊆{1,...,n}
the mechanism with L̃i = γi , ℓ̃i = 0 and pi = min{γiT γi : i ∈ {1, . . . , n}} satisfy
all the constraints of the auxiliary problem. Then, this problem has a solution
and V ( I, δ) > 0. As the number of types is finite, in the optimal mechanism
min{ pi : pi > 0} > 0. Therefore, as δ → 0, the sequences V (δ) and I ∗ (δ) are
finally constant, and equal to the seller’s profits and the set of types to which the
seller sells in the optimal mechanism, respectively.
Let I ∗ be the set of types to which the data broker sells in the optimal mecha-
nism. Consider the auxiliary problem with I ∗ and small δ, such that the solution
of the auxiliary problem corresponds to the optimal mechanism.
We use the Fritz’s conditions to characterize the optimal solution.40 Let λ0
be the multiplier corresponding to the objective function, λij be the Lagrange
multiplier corresponding to constraint ICij and µi be the Lagrange multiplier cor-
LiT γ j
responding to the constraint IRi for i ̸= j ∈ I ∗ . Defining a ji ≡ T
Li Li +ℓ2i
, the Fritz’

39 Ifthere are multiple subsets of individuals that are optimal, pick any of them.
40 The Fritz’ conditions provide a necessary condition that the optimal solution must satisfy.
Then, they allow to conclude what functional form the optimal solution have, though no to to
solve for it. See Footnote 37 for more details and for a reference for a proof that applies to this
environment.

44

Electronic copy available at: [Link]


necessary conditions with respect to Li imply that
a ji
!
1 ∑ j̸=i λ ji a
Li = ∑ j̸=i λ ji a2ji
γi − ii
∑ j̸=i λij +µi j
γ .
aii
2 − 2(∑ j̸=i λij +µi ) aii

Li
Lemma 1 implies that the buyer’s willingness to pay for the statistic Li and L̃i = k
with k ∈ R++ is the same. Then it is without loss to normalize the denominator
of Li to 1. Further, the problem is bang-bang with respect to ℓi , so that ℓi = 0 for
a ji
all i ∈ I ∗ . By defining c ji ≡ aii we obtain the desired expression.
It remains to show that the mapping from the vector (c ji )i,j ∈ R(n−1)×(n−1)
into itself has a least one fixed point. The mapping is clearly continuous. I ar-
gue that thecodomain  and the domain of the mapping are bounded. Define
T
( L i γi ) 2 ( LiT γi )2
ζ = mini∈ I ∗ L LT > 0 since by construction LT L
> 0 for each i ∈ I ∗ . Fur-
i i i i
( LiT γ j )2
ther, define M = maxi∈ I ∗ γiT γi . Lemma 1 implies that ≤ γ Tj γ j ≤ M.

LiT Li
M
Therefore, for each i and j, c2ji ≤ ζ .
Therefore, the mapping that define cij can be restricted to go from the compact
h q q i(n−1)×(n−1)
convex subset − M ζ ,
M
ζ to itself. Brouwer Fixed-Point Theorem
implies that this mapping has at least one fixed-point. Since the proof of Theo-
rem 2 implies that this problem has a solution, its solution needs to satisfy the
Fritz necessary conditions, and one of the fixed-points corresponds to the optimal
mechanism.
Proof of Proposition 2
None of the IC constraints bind if and only if the full-information mechanism
that targets to each type i the linear statistic with coefficients γi , and charges each
type the monopolistic price γiT γi is feasible. Denote by β ij the angle between γi
and γ j . In the full-information mechanism the constraint ICij with i < j is satisfied
iff
(γiT γ j )2 2
T
− γ Tj γ j < 0 ⇔ (γiT γ j )2 > (γ Tj γ j )2 ⇔ cos2 ( β ij ) ∥γi ∥2 < γ j .
γj γj

45

Electronic copy available at: [Link]


Further, the constraint ICkl with k > l is satisfied iff
(γlT γk )2
− γlT γl < 0 ⇔ (γkT γl )2 < (γlT γl )2 ⇔ ∥γl ∥2 ∥γk ∥2 cos2 ( β lk ) < ∥γl ∥4 .
γlT γl

This inequality is always satisfied since cos( β lk ) < 1 and by assumption, for k > l,
∥ γl ∥ 2 ≥ ∥ γk ∥ 2 .
Therefore, if for all i < j, cos2 ( β ij ) ∥γi ∥2 ≤ γ j , the full-information mecha-
nism with monopolistic prices is feasible.
Consider now the case in which some i < j, cos2 ( β ij ) ∥γi ∥2 = γ j + ϵij , with
ϵij > 0 small and for all other k and ℓ cos2 ( β kℓ ) ∥γk ∥2 < ∥γℓ ∥. By Theorem
3 and continuity, the seller offers to type j a statistic with coefficients L j that
are distorted a bit away from γ j in the direction opposite to type i’s preferred
direction and charges to type i a price slightly below ∥γi ∥2 . This deviation in
price and direction reduces as ϵij decreases.
These changes relax the constraints ICiℓ for ℓ ̸= j since now the utility of type
i is positive, and any constraint ICni with n ̸= i are still satisfied as long as the
reduction in the price p(i ) is small enough. Further, any constraint ICkj with k ̸= i
is still satisfied as long as the change in the direction of L j is sufficiently small.
Finally, since type j is obtaining the same rent, any constraint ICjm is also satisfied.
Therefore, when all ϵij are small enough, the relaxed problem the considers
2
only the constraints ICij with i < j such that cos2 ( β ij ) γ j > ∥γi ∥2 solves the
seller’s problem.
Proof of Proposition 3
I consider the relaxed problem that contains only the adjacent downward con-
straints. From Theorem 3 we have that if the constraint ICi,i+1 binds then

Li+1 = γi+1 − ci,i+1 γi ,

L T γi λi,i+1
with ci,i+1 = α̃i+1 LTi+γ1 and α̃i+1 = λi+1,i+2 +µi+1 > 0. Fixing the Lagrange multi-
i +1 i +1

46

Electronic copy available at: [Link]


pliers, these two expressions define the following fixed-point problem in ci,i+1 :
(γi+1 − ci,i+1 γi )T γi
ci,i+1 = α̃i+1 (2)
(γi+1 − ci,i+1 γi )T γi+1

I show there is a unique ci,i +1 that solves this problem.
Suppose first that γiT γi+1 > 0. First, notice that the constraint IRi implies that
LiT γi LiT γi LiT+1 γi
p (i ) ≤ LiT Li
. As the constraint ICi,i+1 binds then p(i ) = LiT Li
− LiT+1 Li+1
+ p ( i + 1).
LiT+1 γi
To satisfy both constraints it is necessary that LiT+1 Li+1
≥ 0. Therefore, it has to
∗ γiT γi+1
be that ci,i +1 ≤ c̄i,i +1 := γiT γi
< 1. Further, it is easy to check that the RHS of
Equation 2 is strictly decreasing with respect to ci,i+1 , that when ci,i+1 = 0 the RHS
of Equation 2 is larger than 0 and that when ci,i+1 = c̄i,i+1 the RHS of Equation
2 is equal to 0. Therefore, as both sides of of Equation 2 are continuous with
respect to ci,i+1 , there exists a unique ci,i+1 ∈ (0, c̄i,i+1 ) that solves the fixed-point
problem.
At the same time, the definition of ci,i+1 generates a quadratic equation with
solutions
q
2 2
∥γi+1 ∥ + α̃i,i+1 ∥γi ∥ ± (∥γi+1 ∥2 + α̃i,i+1 ∥γi ∥2 )2 − 4α̃i,i+1 (γiT γi+1 )2
.
2γiT γi+1

By using that (γit γi+1 )2 = ∥γ1 ∥2 ∥γ2 ∥2 cos2 ( β i,i+1 ), where β i,i+1 is the angle
between γi and γi+1 , it is easy to check that the positive root of this quadratic
equation is always larger than or equal to c̄i,i+1 and that the negative root is
always always positive and smaller than or equal to c̄i,i+1 . As ci,i+1 = 0 and
ci,i+1 = c̄i,i+1 are clearly sub-optimal for the seller, the solution to the problem is
given by the negative root of this quadratic equation.
If γiT γi+1 < 0, let γ̃i+1 = −γi+1 . With this change of variables γ̃iT γ̃i+1 > 0.
The procedure above delivers the optimal c̃i,i+1 . By substituting back, ci,i+1 =
−c̃i,i+1 ∈ (c̄i,i+1 , 0).
Now, I check that by picking ϵ small enough the constraint ICi,k for any k >

47

Electronic copy available at: [Link]


i + 1 is satisfied. By assumption we have that

∥γi+1 ∥2 (1 + ϵ) > cos2 ( β i,i+1 ) ∥γi ∥2


∥γi+2 ∥2 (1 + ϵ) > cos2 ( β i+1,i+2 ) ∥γi+1 ∥2
....
..

∥γk ∥2 (1 + ϵ) > cos2 ( β k−1,k ) ∥γk ∥2 .

By substituting from the first expression into the second one and so on, we obtain
the following inequality
k
∥γk ∥2 (1 + ϵ)(k−i) > ∏ cos2 ( β j,j+1 ) ∥γi ∥2 .
j =i

∏kj=i cos2 ( β j,j+1 )


Define η = mink,i (1+ ϵ ) ( k −i )
. By the definition of ordered types, if ϵ is
smalle enough, η > 1. Therefore, the inequality above implies that for ϵ small
enough, ∥γk ∥2 > cos2 ( β i,k ) ∥γi ∥2 .
At the same time, when ϵ is small, by continuity, the Lagrange multipliers are
small. Therefore, pk ≊ ∥γk ∥2 and the distortion in Lk is small, that is the angle
between Lk and γi is approximately equal to β ik . At this price and statistic Lk , the
inequality ∥γk ∥2 > cos2 ( β i,k ) ∥γi ∥2 implies that if type i buys the statistic targeted
to type k he would get a negative information rent. Therefore the constraint ICik
is satisfied.
Finally, if ϵ is small enough, the prices in the mechanism are very close to the
monopolist prices, so that all the upward IC constraints are satisfied, and such
that the seller does not want to exclude any buyer type.
Proof of Corollary 1
From Theorem 1 and the discussion in Step 0 of the proof of Theorem 2, it
follows directly that the constraint IC21 does not bind and that the constraint IR2
always binds. Therefore, the equality L1 = γ1 follows directly, and from Lemma
( L2T γ2 )2
1 we obtain that p(2) = L2T L2
.

48

Electronic copy available at: [Link]


Proposition 3 implies that L2 = γ2 − c∗ γ1 . Then, notice that the constraint IR1
does not bind, i.e., p(1) < γ1T γ1 iff ( L̃2T γ1 )2 > ( L̃2T γ2 )2 . A quick calculation shows
γ2T (γ1 −γ2 )
that it is equivalent to c∗ < c̄ := γ1T (γ1 −γ2 )
.
Checking the FOC conditions it is straightforward to check that the multiplier
corresponding to the constraint IC12 , λ1,2 , is equal to α1 and the multiplier cor-
responding to the constraint IR2 , µ2 , is equal to 1. Therefore, , as long as the
constraint IR1 does not bind, from Proposition 3 we conclude that
q

2
∥γ2 ∥ + α1 ∥γ1 ∥ − (∥γ2 ∥2 + α1 ∥γ1 ∥2 )2 − 4α1 (γ1T γ2 )2
2
c = .
2γ1T γ2

Further, by definition c̄ is such that (γ2 − c̄γ1 ) T γ1 = (γ2 − c̄γ1 ) T γ2 . This


equality, the definition of L̃2 and the definition c12 in Proposition 3 imply that c̄
∂c∗
is a solution when α1 = c̄. As ∂α1 > 0, for α1 > c̄ the optimum is reached at this
boundary and I obtain the desired characterization.
Proof of Corollary 2
Part 1. follows directly from the proof of Corollary 1.
For part 2., write type 1’s information rent as
((γ2 − cγ1 )T γ1 )2 − ((γ2 − cγ1 )T γ2 )2
Rent = ∥γ1 ∥2 − p(1) = .
(γ2 − cγ1 )T (γ2 − cγ1 )

A couple of algebra steps show that

∂Rent
∂c ∝ −γ1T γ2 + c(γ1T γ1 + γ2T γ2 ) − c2 γ1T γ2 .

I argue that this derivative is negative. This results from this expression being
increasing in c for c < c̄ and being negative when evaluated at c̄. To see the
first, notice that the derivative with respect to c of this expression is equal to
γ1T γ1 + γ2T γ2 − 2cγ1T γ2 and it is larger than zero since c < c̄ < 1 and 2γ1T γ2 <
γ1T γ1 + γ2T γ2 . To see the second, notice that when evaluated at c̄ it is equal to:

−(1 + c̄2 )γ1T γ2 + c̄(γ1T γ1 + γ2T γ2 ) ∝ − ∥γ1 − γ2 ∥2 (γ1T γ1 γ2T γ2 − (γ1T γ2 )2 ) < 0.

49

Electronic copy available at: [Link]


Therefore, ∂Rent
∂c < 0. As ∂c
α1 ≥ 0, I conclude that ∂Rent
∂α1 ≤ 0.
Proof of Corollary 3
T
The condition in the statement is equivalent to γ1T γ2 < γ′ 1 γ2′ . It is easy to
check that if the constraint IC12 binds with data X then it binds with data X ′ .
Fix the optimal statistic for type 2 with dataset X ′ L2′ = γ2 − c′ γ1 and assume
that the seller sells this statistic independently of which of the two datasets he
buys. Therefore, by Corollary 1, p(2) is independent of which of the two datasets
the seller observes. Further, p(1) = γ1T γ1 + r (c′ ) with
2 2
′ γ2T γ2 − c′ γ1T γ2 − γ1T γ2 − c′ γ1T γ1
r (c ) = .
γ2T γ2 − 2c′ γ1T γ2 + c′ 2 γ1T γ1

The only difference in this term between comes from γ1T γ2 . Besides,

∂r (c′ )
∂γ1T γ2
∝ −(γ1T γ2 − c′ γ1T γ1 )(γ2T γ2 − c′ γ1T γ2 )(1 − c′2 ) < 0.

It is easy to check that the first two factors are positive. The third factor
T
is positive since by Corollary 1, c′ < 1. As γ1T γ2 < γ′ 1 γ2′ , r (c′ ) is larger for
data X. Then, for fixed c′ , p(1) > p′ (1), which increases the seller’s profits. By
choosing optimally c the seller can further increase his profits after buying data
X . Therefore, the seller would prefer to buy X rather than X ′ .
Proof of Theorem 4
Consider the relaxed problem without the constraints IC2v−1v′ and the mono-
tonicity constraints. In the next lemmas I present some properties of its solution.

Lemma 3 If q22v ̸= 0 then ℓ2v = 0.

Proof Suppose by contradiction that in the optimal mechanism q22v ̸= 0 and


ℓ2v ̸= 0. Then, at least one of the constraints IC1v′ ,2v binds. Let δ22v be the
angle between L2v and γ2 , and β 12 be the angle between γ1 and γ2 . Lemma
cos2 ( β 22v )∥ L2v ∥2 ′ , L′ v ) with ℓ′ = 0
1 implies that q22v = . Take the statistic (ℓ2v 2 2v
∥ L2v ∥2 +ℓ22v
′ the vector with magnitude one that generates an angle with γ equal to
and L2v 2

50

Electronic copy available at: [Link]



β′2v = cos−1 ′ = cos2 ( β′2v ) = q22v . Furthermore,

q22v . By construction q22v
q

q12v = cos( β′2v + β) = cos( β′2v )cos( β) − sin( β′2v )sin( β)
r
cos( β 2v )∥ L2v ∥ cos2 ( β 2v )∥ L2v ∥2
=q cos ( β ) − 1 − 2 2
sin( β)
2
∥ L2v ∥ +ℓ22v ∥ L2v ∥ +ℓ2v

cos( β 2v )∥ L2v ∥ 1−cos2 ( β 2v )∥ L2v ∥
< q cos( β) − q sin( β)
∥ L2v ∥2 +ℓ22v ∥ L2v ∥2 +ℓ22v

cos( β 2v + β)∥ L2v ∥ √


= q = q12v ,
∥ L2v ∥2 +ℓ̂22v


where I use that sin(cos−1 ( x )) = 1 − x2 in the second equality.
The alternative statistic relaxes the constraints IC1v′ ,2v , which allows the seller
to increase the price p(1v′ ) and his profits, since at least one of them binds.

Lemma 4 For all v ≥ v1∗ , q11v = 1.

Proof It follows directly from the fact that for any v ≥ v1∗ the virtual value is
positive, so that increasing q11v relaxes all constraints IC1v′ ,2w , and this allows the
seller to increase his profits.

Lemma 5 There exists v1∗ > v̂2 ≥ v2∗ such that q22v > 0 iff v ≥ v̂2 .

Proof For v < v2∗ the seller optimally picks q22v = 0. This allocation trivially
satisfies all constraints IC1v,2v and maximizes profits because for any v < v2∗ type
(2, v)’s virtual value is negative.
Further, v̂2 < v1∗ because the allocation q11v = q22v = 1v≥v1∗ and constant
prices satisfy all IC constraints with a slack. Then, there exists a neighborhood
V = (v1∗ − ϵ, v1∗ ) such that making q22v = 1 for v ∈ V satisfies all the IC constraints,
and it increases the seller’s profits, since type (2, v)’s virtual value is positive.

v̂2 q22v̂2
Lemma 6 Only the constraints IC1v,2v̂2 for v ∈ (v̂1 , v1∗ ) with v̂1 = h(q22v̂2 )
∈ (v̂2 , v1∗ ),
and IC1v1∗ ,2v′ for v′ ∈ [v̂2 , v̄2 ] bind. Further, v̂2 > v2∗ .

51

Electronic copy available at: [Link]


Proof Consider the relaxed problem with allocations q1 and q2 that satisfy the
conditions in the previous lemmas, and only impose the constraints IC1v,2v̂2 for
v ∈ (v̂1 , v1∗ ) and IC1v1∗ ,2v′ with v′ ≥ v̂2 . For this problem, it is trivially optimal to
offer zero information rent to types 1v1 and 2v2 , and set q11v = 0 for types v < v̂1
¯ ¯
since these types’ virtual value is negative.
Notice that the constraints IC1v,2v′ with v′ < v̂2 are satisfied since q22v′ = 0.
Further, for v ≥ v1∗ and v′ ≥ v̂2 the constraint IC1v,2v′ is satisfied since
Rv R v∗ Rv R v′
v q11w dw = v
1 q11w dw + v1∗ q11w dw ≥ v1∗ h(q22v′ ) − v′ q22v′ + v q22w dw + (v − v1∗ )
¯ ¯ ¯
R v′
≥ vh(q22v′ ) − v′ q22v′ + v q22w dw,
¯

where the first inequality follows from constraints IC1v1∗ ,2v′ and q11v = 1 for v >
v1∗ , and the second inequality follows from h(q22v′ ) < 1 by definition of h.
Moreover, the constraint IC1v1∗ ,2v̂2 binds; if it does not, the seller can set v̂2 = v2∗ ,
q22v̂2 = 1 and v̂1 = v1∗ . However, this allocation is not optimal when some crossed
IC constraints bind.
Additionally, it has to be that v̂2 > v2∗ and q22v̂2 < 1: A small reduction of
q22v̂2 or an increase of v̂2 almost does not affect the profits accrue from type 2v̂2 ,
but it reduces the information rent given to all types 1v with v ≥ v1∗ .Then, by
continuity of the RHS of the IC constraints, for v′ in a neighborhood to the right
of v̂2 , q22v′ < 1. This implies that the constraints IC1v1∗ ,2v′ bind: if they did not bind
the seller prefers to set q22v′ = 1. Therefore, in this neighborhood, the derivative
of the RHS of the IC constraints with respect to v′ is equal to zero:
∂q22v′ ∂q ′
v1∗ h′ (q22v′ ) |v′ =v̂2 −v′ 22v | ′ = 0.
∂v ′ ∂v′ v =v̂2
v′
0 or h′ (q22v′ ) = As limx→1 h′ ( x ) =
∂q22v′
Thus, in this neighborhood either dv′ |v′ =v̂2 = v1∗ .
∞, q22v′ is bounded away from 1 for all v′ in this neighborhood. But this neigh-
borhood must contain all values v > v̂2 : If for some type v̊ > v̂2 the constraint
IC1v1∗ −2v̊ does not bind, it would be optimal to set q22v̊ = 1. However, this con-

52

Electronic copy available at: [Link]


tradicts the continuity of the RHS of the IC constraints: q2v′ being bounded away
from 1 implies that limv→v̊− q22v < 1.
Further, the constraints IC1v,2v′ for all v < v1∗ and v′ ≥ v̂2 are satisfied since
Rv R v̂1 R v̂1 R v′
v q11w dw = v q11w dw − v q11w dw ≥ v̂1 h(q22v′ ) − v′ q22v′ + v q22w dw
¯ ¯ ¯

where the inequality follows from the constraint IC1v̂,2v′ and q11v = 0 for v < v̂.
I conclude by arguing that constraints IC1v,2v̂2 for all v ∈ [v̂1 , v1∗ ] have to bind.
Suppose it does not bind for some v ∈ [v̂1 , v1∗ ]. Then the seller would like to
pick q11v = 0 since type (1, v)’s virtual value is negative. But, then type (1, v)
can obtain a positive utility by mimicking type (2, v̂2 ). By taking the derivative in
both sides of these IC constraints I conclude that q11v = h(q22v̂2 ).
We only need to find the allocation for q22v with v ≥ v̂2 . Since, these types’
virtual value is positive, it is optimal to pick the largest q22v′ that, as shown in the
v′
0 or solves h′ (q22v′ ) =
∂q22w
previous lemma, is either the one that solves dw |w=v′ = v1∗ .
It is easy to show that the second derivative of h(q22v ) is strictly positive,
−1
so that h′ (q22v ) is strictly increasing, its inverse exists, and h′ is increasing.
Therefore, there exists ṽ2 ≥ v̂2 such that q22v = q22v̂2 for v < ṽ2 and q22v =
 
−1 v
h′ v∗ for v > ṽ2 , so that q22v is non-decreasing. To complete the argument
1
h(q22v̂ )
I show that ṽ2 > v̂2 . They are equal only if h′ (q22v̂2 ) = q22v̂ 2 , or equivalently,
r 2
1 − q
if tan(cos−1 (q22v̂2 ) + β) −
22 v̂ 2
q22v̂ = 0. The LHS of this condition is strictly
2
decreasing with lim LHS = tan( β). Since β > 0, there is not value of q22v̂2 < 1
q22v̂2 →1
that satisfies this condition, and ṽ2 > v̂2 .41
Finally, it is straightforward to check that this solution satisfies the constraints
that were omitted at the beginning of the proof.

41 Note that q22v might be a constant function since it is not necessary that ṽ2 ≤ v̄2 .

53

Electronic copy available at: [Link]

You might also like