0% found this document useful (0 votes)
16 views18 pages

Chapter 1

The document discusses the complexities and methodologies involved in making economic predictions, emphasizing the role of uncertainty and the various forecasting techniques available. It highlights the challenges of accurately predicting economic variables, the influence of expert judgment, and the limitations of traditional statistical methods in the context of evolving data and economic behaviors. Additionally, it touches on the importance of integrating predictions into decision-making processes across various economic agents and the philosophical considerations surrounding the nature of uncertainty.

Uploaded by

shepjoel2
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
16 views18 pages

Chapter 1

The document discusses the complexities and methodologies involved in making economic predictions, emphasizing the role of uncertainty and the various forecasting techniques available. It highlights the challenges of accurately predicting economic variables, the influence of expert judgment, and the limitations of traditional statistical methods in the context of evolving data and economic behaviors. Additionally, it touches on the importance of integrating predictions into decision-making processes across various economic agents and the philosophical considerations surrounding the nature of uncertainty.

Uploaded by

shepjoel2
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Predictions: chapter 1

[Link]
November 2024

1 Introduction to Economic Predictions


1.1 Uncertainty
Almost all human decisions involve some kind of predictions. Humans can
predict by using information and lessons from the past, and their imagination.
Anything can be a prediction: claims from TV news, discussions within social
interactions, politicians’ statements,...
However, the public production of predictions and forecast, separated from
the direct use of these predictions only began in public administrations in the
XX th century, mostly after WWII with the setting up of national statistical
offices.
We are here interested in the predictions made for economics, and partic-
ularly the ones involving leading economic indicators. Examples of variables
subject to predictions and forecasting are: GDP growth, asset prices in stock
exchanges, unobserved characteristics of individuals for administrative or policy
purposes, election results, etc.
Each prediction or forecast problem requires a methodology, often based on
what the available data are. Common sense and statistics are often combined.
However, they are usually based on systematic procedures. They are ’forecasting
rules’ for making statements about future events.
There are many methods of predictions: guessing, astrology, rules of thumb,
time series models, leading indicators, extrapolations, structural econometric
forecasting models...
There are forecasts with econometric time series models (e.g., DGP described
by ARIMA models) and statistical forecasts without them (e.g., exponential
smoothing models). In all these cases, there is no express need of causal ex-
planations as a basis for the forecast calculus. In addition, forecasting experts
exert some subjective judgment about their statistical results. So, there are
black boxes somewhere in the kitchen of forecasters. It is not always clear that
these expert opinions always improve the forecasts, while many professionals
tend to believe that this is the case. This implies an additional analytical step
once the statistical forecasts have been produced, and possibly changing them.
Factors hard to predict are technology innovations, revolutions, fashions and
changes in consumer tastes, natural disasters, etc.

1
In particular, conjuncture forecasts do not respond to the same needs than
long term forecasts. They assist monetary and budgetary government policies,
and for the firms their decisions about prices, stocks and cash flows. Most other
organizations are also concerned about prediction and forecasts in their own
way.
Even basic prediction or forecast can be useful to a firm, if they yield better
results than its competitors.
Prediction seems to be the general term, while forecast rather refer to a
time framework. Cases in which predictions are used without referring to time
or to the future are frequent. For example, when part of a population is not
observed, the policy maker or the administrations must predict what are the
characteristics of the individuals that are not observed.
However, the performance of the typical economic forecast is poor, to say
the least. Some of the stated forecasts by experts and forecasting institutions
are often belied by the occurrence of contrary events. Some of the best time
series models for forecasting are sometimes beaten by a mere random walk.
Other indicators are better forecast such as population projections and foreign
trade deficits, while not always. Prediction errors and forecast errors are often
originated from emotions, with predicting agents hoping or fearing too much.
Blind extrapolations are also common sources of prediction errors.
Moreover, it is usually hard to draw clear conclusions about the economic
mechanisms at play from the published forecasts. Besides, this is consistent with
forecasts not needing to be related to causal relationships. As a consequence,
The notions of exogeneity and endogeneity of variables are somewhat blurred
in predictions. Endogenous variables, in terms of an explanatory models, can
often be used as valid predicting variables, while past exogenous variables are
also massively exploited.
Predictions and structural models are not necessarily related. There are
many examples where causal models do not provide the best forecasts. Some
forecasters even completely neglect economic theoretical bases for design forecast
methods. Still, having in mind an economic model should help listing predicting
factors and having ideas about how they interact. Then, one can search data
about these factors. In that case, it is also necessary to think about other factors
that are not included in the theory, but may be relevant.
Another issue, with the current explosion of data availability and comput-
erized power of calculus, is that traditional asymptotics in statistics does not
handle well the situations where they are more available variables than observa-
tions. Moreover, the usual goodness-of-fit indicators of statistics do not respond
well to evaluating the future. Standard errors of predictions and forecasts are
still useful for dealing with inaccuracies in parameter estimation and specifica-
tion search, but they do not help for other types of uncertainties occurring with
predictions. Besides, overfitting is a scourge of many prediction attempts. The
nature of the data used is also unstable, with macroeconomic indicators being
routinely revised over time, for example. In all these respect, more and different
statistical methods are needed when predicting. However, traditional statistics
must be well understood and mastered first.

2
An issue with macroeconomic forecasts, is the Lucas critique: The way the
economic system respond to changes and to policies evolve over time and is
affected by exogenous variables and policies themselves. This challenges the
idea that a stable model of agents’ behaviour could provide precise predictions.
More specifically, the publication of forecasts can itself influence people, and
made them believe that they will be true, and thus make them self-fulfilling
when people follow these results by supporting them, or self-defeating when
people fight them for example because of protection motives.
Finally, predictions and forecasts should be integrated in the decisions that
they are supposed to guide. This is rarely the case from an analytical point of
view. For example, forecasting can be used for many different decisions: policy
design and evaluation, tax and budget management, control of processes, scien-
tific research and model evaluation, risk and portfolio management, investment
and saving decisions, insurance decisions, planning of operations, etc.
However, any economic agent engaged in decision-making is better try to
know where it goes, that is to predict and forecast. Our challenge in this course is
to try to better understand predictions and forecast, and make progress on their
methods. Statistical decision theory and time series techniques are among the
tools that will assist us. Some ways of summarizing the data, with parsimonious
while flexible models should be developed. Some ways of investigating out-of-
the-sample situations should be used.
Rational decision-making often requires evaluation of risks and the careful
weighting of the potential positive and negative consequences of risky choices.
We may willingly embrace a certain level of risk based on our assessment of
specific contexts.
All calculations about the future involve inductive reasoning, but they do
not involve the future itself. These reasonings are only reasonings. For exam-
ple, strictly speaking, it remains a possibility that it might snow in Marseille
tomorrow, despite inductive reasonings based on past records suggesting other-
wise. Moreover, fundamentally all that we discuss about the uncertainty of the
future can be transposed to any situation of ignorance, not exclusively about
future uncertainties. So, predictions can deal with the future or with current
ignorance.
Uncertainty results from the fact that all the factors affecting an agent’s
actions cannot be determined or predicted. One can distinguish two cases:
(1) The uncertainty that is generated by interferences with the actions of
other persons, and
(2) The uncertainty that directly stems from random events, whether these
are natural or social events.
In contrast, regarding the uncertainty emerging from pure randomness, log-
ical tools that describe the random events must be constructed. When these
events can be associated with probabilities - often probabilities of states of na-
ture that also need be defined, we are able to specify probability spaces, which
are associated with a set of random events.
Another, technical, distinction of how uncertainty is modeled, is whether
it is probabilistic, as for financial risks; measurable, when random events can

3
be specified but not their probabilities; and more evanescent than this. This
distinction is somewhat related to the notions of predictive uncertainty that we
will dealt with, and unpredictive uncertainty that cannot be grasped precisely
but must be accounted for.
Among the error terms that partly account for uncertainty in econometric
models, some are motivated by diverse envisaged sources: innovation shocks,
measurement errors, optimization errors by agents, individual heterogeneity,
environment heterogeneity, efficiency error, changes in behavior, sampling error,
policy-generated randomness, etc.

1.2 Cross Sections, Time Series and Panel Data


   
y1 x11 ··· ··· x1k
 ..   .. .. 
 .   . . 
where Y = 
 ..
, X = 
  .. ..
,

 .   . . 
yn xn1 ··· · · · xnk
   
β1 u1
 ..   .. 
 .   . 
β=
 ..
 and u = 
  ..
.

 .   . 
βk un

In that case, Y and u are random n-vectors and X can be a fixed or random
n × k matrix. The model can also be written separately for each observation i:

yi = β1 xi1 + β2 xi2 + ... + βk xik + ui ,


E(ui |xi ) = 0

and
E(ui uj |xi , xj ) = σ 2 δij
for all i, j = 1, ..., n,
where δij is the Kronecker index (δij = 1 if i = j, and 0 otherwise).
In all these expressions, the expectations are conditional on X because these
variables can be random. In general, there is a constant in the X ′ s.
These notations may also directly denote random variables for Y and u, and
a k-random vector for X that includes k distinct random variables. This is a
perspective before a sample has been drawn, and before the sample size is known.
It has the advantage of focussing on the fundamental notion of randomness
associated to a given distribution for each variable.
Sometimes, in these two cases, the variables in X can be considered as fixed
or chosen by the analyst, as is the case for some experiments. This view has the

4
advantage of making some calculations more concrete, for example the formula
of the OLS estimator β̂ = (X ′ X)−1 X ′ Y can be discussed as a mere algebraic
calculus based on some given observed numbers.
When distinguishing the observations of each individuals i, one can denote
Yi , Xi and ui , respectively.
With time series instead, the indices i and t can be substituted. However,
in that case, the sequentiality of the periods somehow changes the theoretical
perspective. First, serial correlations will play a central role in the econometric
analyses. Second, non-stationarity issues when they arise will invalidate many
methods used in cross sections, and new methods must be used. Third, since
forecasting is the objective here, on can systematically take advantage of the
available information about the past to predict the future, and this in a way
that account for the distinct time distance of the past observed periods to the
future.
Typical equations could read like:

yt = β1 xt1 + β2 xt2 + ... + βk xtk + ut ,


E(ut |xt ) = 0
and appropriate conditions for the autocorrelation matrix of the errors.
These conditions are crucial here. Ideally, one tries to include in the models
enough controls to obtain non-correlated errors, which much facilitates the esti-
mation and the analysis. However, imposing E(ut |xt ) = 0 and V(ut |xt ) = σ 2 IT
with time series from the start is somewhat counterintuitive as the same vari-
ables with different lags should be highly correlated in general.
Let us now define a fixed effects panel data model. Let us consider the
following random variables for each individual i :

(yi1 , ..., yiT , x1i , ..., xiT , ηi ), i = 1, ..., N . Then, one writes the model:

yit = x′it β + ηi + vit , (1)


i = 1, ..., N,
t = 1, ..., T.
where the β is a vector of parameters to estimate, and vit is an unobserved
error term, called idiosyncratic. ηi is another unobserved error that has the
special feature of being constant over time, and is called the fixed effect.
Panel data analysts are mostly interested in the cases ‘N large, T small’,
with asymptotic approximation when N goes towards infinity. On the other
hand, no constraints for the temporal dimension are imposed. In particular, no
assumption of stationarity or stability over time is made. Assume the following
hypotheses A1 and A2:

A1: E (vi |xi , ηi ) = 0 for all t and i,

5
where vi = (vi1 , ..., viT )′ and xi = (xi1 , ..., xiT )′ , with vi et ηi unobserved.

This is called the strict exogeneity hypothesis. One says that xi = (xi1 , ..., xiT )′
are strictly exogenous conditionnally to the unobserved effects ηi .

Alternatively, the strict exogeneity hypothesis in this model is sometimes


expressed by the next condition :
E (yit |xi1 , ..., xiT , ηi ) = E (yit |xit , ηi ), for t = 1, ..., T , and for which similar
results can be obtained. This means that once the xit et ηi are included in the
linear model, the other xis (s ̸= t) have no partial effects on yit .
Hypothesis A1 allows the use of difference and the within-group’ estimators,
which are convergent in this case. Let us note, however, that A1 is not satisfied
in the models with lagged dependent variable. Moreover, in this case the usual
OLS exogenous hypothesis is also violated.

A2: V (vi |xi , ηi ) = σ 2 IT , où σ 2 is a parameter to estimate.

Here again, the conditioning is with respect to all periods!


However, in most prediction problems one rather considers data with large
T and small N. First, this implies to reverse the use of these two indices in the
assumptions, as follows; and second to consider the sequentiality of time, which
can lead to serial correlations. The fixed effect is now realized for each period
instead of each individual, and can be denoted ηt .
A1’: E (vt |xt , ηt ) = 0 for all t and i,
where vt = (v1t , ..., vN t )′ and xt = (x1t , ..., xN t )′ , with vt et ηt unobserved.

A2’: V (vt |xt , ηt ) = σ 2 IN , où σ 2 is a parameter to estimate.


Again the sequentiality of the periods can be exploited to extract infor-
mation from the past to get a grasp on the future. In addition, the issue of
non-stationarity can become important and requires specific estimation and
prediction methods.

1.3 Reminders of Probability and Algebra


Probability theory is a main tool to deal with ignorance, notably about the
future.
In probability theory, the risks are described by random variables. Using
standard probability theory is not mandatory (fuzzy set theory, ambiguity the-
ory, and context dependence could be explicit instead), although it is the most
convenient, useful and popular approach.
A philosophical problem that arises when attempting to define concretely
probability numbers to use in decision processes is: Does uncertainty funda-
mentally lie ‘in’ the things themselves (a kind of materialism à la Plato), or in
human judgements about things and events (a kind of subjectivism)? Accord-
ingly, two distinct families of statistical methodologies are used when discussing
probabilities in economics.

6
(1) The probabilities of events in the sense of observable frequencies (about
our observations of the world, and extrapolations from these observations).
They can be called objective probabilities, notably because they generally do
not vary across individuals. For example, the number of rainy days within a
year is a frequency number that provides a measure of the probability of rain
over the year; Another example is the occurrence of the results of a dice, which
is 1/6 for each face and corresponds approximately to the observed frequencies
when throwing it enough times. Objective probabilities are the ones typically
used in time series models for forecasting.
(2) The probabilities of events in the sense of likelihood (in our judgements
and feelings). They can be called: subjective probability, notably because they
generally vary across individuals. Some degrees of likelihood may become mea-
surable by referring to a fictional story or to personal feelings. One can start
from prior subjective probabilities of events and use implicit or explicit calcu-
lus of probabilities to refine them. For example, by looking at the sky in the
morning, one may develop a feeling about the probability of rain in the after-
noon, then by seeing how the clouds gather overtime in the sky, one may adjust
this probability number intuitively. Subjective probabilities are often used by
experts for predictions when data is poor or missing.
Let s be a vector of possible values given to each of the uncertain elements in
the given context. These values are typically called the states of nature, random
states, or the states of the world. For example, if we are deciding to take an
umbrella with us or not when leaving home, rainy and sunny days can define
two distinct states of nature in which we are interested. Another example is,
for a given financial asset, the infinite set of possible monetary returns on the
next day. Let us denote Ω as the set of all possible states s, or universe of
eventualities, or of possibilities.
Let us now formally define the states of nature and the events. Often,
economic theories are specified in terms of the states of nature. However, the
policies or practical decisions rather refer to the set of events, E, that is a way to
summarize all the useful information for decision-making. These two notions are
related. For example, two events of interest for a financial asset are: its return
tomorrow is positive; and, respectively, negative. Therefore, moving from Ω
to E is not merely a question of information aggregation but also of modeling
choice. When predicting or forecasting, choosing what are the relevant variables,
and the relevant random events, looks like a preliminary stage of the reflexion.
For example, when forecasting GDP for a given year, using the past observed
values of GDP as a basis for the prediction is such a restrictive while inevitable
convention.

Definition:
Axioms for a set of random events E and a measurable space (Ω, E):
E ⊂ P(Ω ), Ω ∈ E;
E is stable by countable intersection ( ∩), countable union ( ∪) and comple-
menting ( ).

7
Ω is the set of random states, while E is denoted the σ-field of events.1
An uncertain or random event is defined as a subset, say E, of Ω, belonging
to a given σ-field E of Ω.

This is the harmonious definition of random events, allowing for their com-
binations, that makes the space ‘measurable’. A useful kind of measure are the
probability functions that will fit well the description of the events. We discuss
these functions below. We note the distinction between random states that be-
long to the universe Ω, for the complete set of events E, and random events
that belong to the set of subsets of Ω, i.e., to P(Ω ). That is, for an event E:
E ⊂ Ω and E ∈ E. Each of the events in E is a set of states from Ω.
Indeed, to define events consistently, a σ- field structure must be imposed on
the set E of the events. This allows, for example, to specify the coincidence of two
events as a new event. For example, if Ω = {rain, snow, sun}, then one can spec-
ify E = {{rain} , {snow} , {sun} , {rain, snow} , {rain, sun} , {snow, sun} , {rain, snow, sun} , }.
This is for the analyst who chooses the most appropriate specification for
E, while Ω may be somewhat imposed on her by a generally admitted theory.
Defining a σ-field E of Ω often amounts to defining the kind of information used.
This is one reason a careful specification of the possible events in E matters.

Definition:
(i) Let be a measurable space (Ω, E). A probability function on this space is
such that

P : E → [0, 1]; P (Ω) = 1;

an (Ai )i∈N all incompatible (i.e., disjoint) events from E, let the
and for P

event A ≡ i=1 Ai [whichP∞is the occurrence of any of the events Ai , and can
be denoted in that case i=1 Ai ], then
P∞ P∞ P∞
P( i=1 Ai ) = P ( i=1 Ai ) = i=1 P (Ai ) .

(ii) The result of a probability function applied to any event in E is a proba-


bility number. The probability function P (.) on the set of states Ω is not neces-
sarily known a priori, which justifies allowing for a set of probability functions,
P.
(iii) A probability space is defined as (Ω, E, P), where P is a set of probability
functions on (Ω, E).

An example is the usual regression model with normal errors for a given
observation i:
yi = x′i β + ui , where u follows a Gaussian distribution N (0, σ 2 ).
This yields a probability density function with density
1 In this context, a σ-algebra is a set of set of states that satisfies the above-described

properties for the operations ∩,∪ and .

8
(1/(2πσ))exp(−(yi − x′i β)2 /(2σ)).
The corresponding random events are generated from intersections and unions
of intervals of values for yi . The set of the states of the world can be anything
imagined that can generate the random events and values of yi . useful to the
analysis.

Because of the axioms, P () = 0. Moreover, the probability function to con-


sider, P (.), may differ across various agents, for example, because it stems from
subjective assessments. This heterogeneity in probability functions can be taken
care of by assuming a set of probabilities P rather than a unique given prob-
ability function P (.). There may also be cases in which one can easily define
events but not their probabilities (such as for rainy days). Then, allowing for
alternative probability functions in a set P is a useful precaution.
The basic intuition behind the notion of a probability function for a discrete
σ-field of a finite number of events {E1 , ..., En } is merely that of an addition
of ‘weights’ (p1 , ..., pn ) that represent the likelihood of each of the elementary
events. Each elementary even can be seen as a singleton for each corresponding
state of the world among Ω = {s1 , ..., sn }. In that case, the distinction between
events and states is moot. Summing these weights to 1 is a convenient convention
to characterize certainty: p1 + ... + pn = 1.

Definition:
(i) In a probability space, a measurable function of the random states is
called a random variable, and it can be endowed with a corresponding ‘image
probability function’.
A real-valued random variable X is a measurable function from a measurable
space (Ω, E) to a real measurable space (R, R), and describes the consequences
in each state s in Ω in terms of real numbers: X(s) ∈ R. That is:
For any event B in R, there exists an event A in E such that X −1 (B) ⊂ A.
The σ-field on real consequence events, R, is typically the ‘ Borel set’ of
events that can be constructed by using intersections and unions of the intervals
in R.
(ii) One can also define real-valued random vectors that are obtained by
stacking several real-valued random variables (e.g., Y = (X1 , ..., Xn )′ ).

In a probability space, a real-valued random variable can be characterized


by its distribution function (cdf ), FX (.), which indicates the cumulative prob-
abilities that the realizations of X are below any chosen threshold. Often, for
convenience, one assumes that the considered random variable has a continuous
cdf (‘cumulative distribution function’), or alternatively a discrete cdf.

1.4 Independence and Conditioning


First, let us recall the definition of independence in probability theory.

9
Definition:
(i) Two events A and B are independent if and only if P (A∩B) = P (A)P (B).
(ii) The events in a set E are independent (‘among themselves’) if the probabil-

ity of the intersection of any subset of events from E is equal to the product of
the probabilities of these events in the subset.

The independence properties are what provides some structure to probabil-


ity theory. They allow for the passage from intersection operations among sets
representing events - that is, from logical conjunction among events - to multi-
plication of probability numbers, which is much simpler to handle. Moreover,
they can be used to decompose complex probability formulae, e.g., mixtures of
probabilities, into simpler ones.

Definition:
The conditional probability of an event B conditional on an event A , which
is such that P (A) ̸= 0, is:2
P (B|A) = P (A ∩ B)/P (A).
In the case that A and B are independent, then the knowledge that A
happens does not influence the perceived probability of B:

P (B|A) = P (A)P (B)/P (A) = P (B).

Moreover, it can be shown that

Proposition
For any event A and any complete finite set of exclusive events Bi , i =
1, ..., I, PI
P (A) = P (A|B)P (B) + P (A|B̄)P (B̄) = i=1 P (A|Bi )P (Bi ),
PI
Proof: Use P (A) = i=1 P (A ∩ Bi ), [since the A ∩ Bi are exclusive events
and sum to A] and substitute P (A|Bi )P (Bi ) for P (A ∩ Bi ).
2 The joint occurrence of two events A and B, in A ∩ B, brings some information about the

likelihood of B happening if A has already been observed. One can take advantage of this
intuition to develop the notion of conditional probability. Let us consider any two events A
and B, which are not independent, and such that A is nonnegligible (i.e., P (A) ̸= 0). One’s
subjective feeling about the likelihood that B occurs should be affected by the knowledge that
A has taken place. If A and B are incompatible (A ∩ B =), this likelihood is necessarily zero.
If A ⊂ B, the probability of occurrence of B is clearly one. In the other cases, one could
assume that this likelihood is proportional to P (A ∩ B). This is supported by noticing that
in the Venn diagram of A and B, the measured area A ∩ B may be seen as depicting the
information of A about B, and vice versa versa. This suggests a definition of the conditional
probability of B conditional on A: P (B|A) = kP (A ∩ B), with k a positive scalar constant to
determine. The normalization of this newly constructed probability function P (.|A), based on
a nonnegligible event A (P (A) ̸= 0) yields the formula. Indeed, given that A has happened,
A is certain and therefore 1 = P (A | A) = kP (A).

10
Defining a conditional probability P (B|A) can be seen as a tentative partial
prediction of event B with event A.
The proposition shows that the conditional probabilities can be used to
decompose the probability number of a given event into a sum of products of
probability numbers. The Bayes’s formula in the next subsection is a specific
use of these decompositions.
Learning about some uncertainty often amounts to learning about the prob-
ability numbers of events of interest. This learning process can be modelled
as updating ex ante probabilities by using information about the occurrence of
some observable events.

Proposition (Bayes) Let A and B be two random events with nonzero


probabilities. Then: (i)

P (A|B).P (B)
P (B|A) = P (A|B).P (B)+P (A|B̄).P (B̄)
.

(ii) Consider a probability space with a finite number S of states of the world:
s = 1, ..., S; and a given probability function P (.) Let E be a random event
whose realization brings some information that can be exploited for updating
the probability function P (.). If all the singletons of the states of the world
{{s = j} , j = 1, ..., S} can be considered as random events, then any ex ante
probability number pj ≡ P ({s = j}) of a state s = j, can be updated as follows
into an ex-post probability number pj∗ ≡ P ({s = j} | E):

P (E | {s=j}).pj
pj∗ = PS = w j pj ,
P (E | {s=k}).pk
k=1

P (E | {s=j})
with the updating factors wj = PS .
P (E | {s=k}).pk
k=1

(iii) More generally, with a sequence of K exclusive events Bi , i = 1, ..., K:

P (E | Bj ).P (Bj )
P(Bj | E) = PK
P (E | Bk ).P (Bk )
k=1

Proof : By definition, P (A|B) = P (A ∩ B)/P (B). Therefore, P (A ∩ B) =


P (A|B)P (B) = P (B|A)P (A) [by symmetry]. By solving the latter equality, we
obtain the classical Bayes equality:
P (B|A) = P (A|B)P
P (A)
(B)
= P (A|B)P P (A|B)P (B)
(B)+P (A|B̄)P (B̄)
(using conditioning with
respect to B and B̄), which is the result (i). Decomposing B̄ further into
S measurable states and identifying B with the j th state at the denominator
yields the result (ii); the proof of (iii) is similar. QED.

11
The intuition behind the formula in (ii) is that the ex post probability num-
ber of a state is a rescaling of its ex ante probability numbers by the likelihood
of the realized event E conditional on the considered state, normalized by the
denominator so that all ex-post probabilities sum to one. This implies that, due
to knowledge updating, some probability numbers of events should increase,
while other ones should diminish.

Proposition
(a) For a discrete distribution of a random variable X:
∞ ∞
X X P ({Xi } ∩ A)
E (X | A) = P ({Xi } | A) Xi = Xi
i=1 i=1
P (A)

(b) Let be the, respectively, joint, conditional and marginal density func-
tions,( f , fY |X and fX , respectively), of two random variables X and Y , with
any respective realizations x and y, then

f (x , y) = f Y |X (y | x ).f X (x ).

(c) If X and Y have continuous distributions that do not cancel, then


Z Z
fX ,Y (x , y))
E (Y | X = x) = yfY |X=x (y | X = x )dy = y dy.
Ω Ω fX (x )

(d) (Generalized Fubini)


Z
P (AX ×AY ) = PX
Y (ω X , AY )dP X (ω X )
AX

where PYX is a transition probability function


(that is: PYX is a mapping from (domainof X) ≍ (eventsbasedonY ) into
[0, 1] , such that for any value ωX of X , PYX (ωX , .) is a probability on the
events based on Y ; and for any event AY based on Y , PYX (., AY ) is measurable
in terms of X)

and, for any mesurable variable Z,


Z Z Z 
ZdP = Z(ωX , ωY ) PYX (ωX , ωY ) dP X (ω X ).
Ω ΩX ΩY

Conditioning in an econometric model

Some specific attention will be devoted to the conditioning operation. This is


justified by several reasons. First, the conditional expectation of the dependent
variable is the core of the explanation of this variable. Second, conditioning

12
allows us to examine the role of the unobserved errors in the models, and the
problems that these may generate for the estimation Unclear. Third, they are
suggestive of modelling approaches based on regressions. Fourth, and not least,
conditoning expectations is often a good way to construct predictors.

One can always describe a model of a random variable Y , explained by a


random vector X, in terms of conditional expectation, as follows:

Y = E(Y | X) + u,

where the realizations of Y and X can be observed for a sample, u is an


unobservable random variable such that E(u | X) = 0 by construction, owing
to the linearity and the idempotence of E(. | X). If all random variables in Y
and X are square integrable (i.e., in L2 ), then conditional expectation is also
called regression. This is therefore the central notion of applied econometrics.
Note first that E(u | X) = 0 implies that E(u) = 0, by integration over x,
even if there is no intercept in the model!
Second, u is independent of any function of x.

Let be a k-dimensional real vector X and a real random variable Y , the con-
ditional expectation of Y given X is a numerical mesurable function E(Y | X)
defined on Rk , such that
   
2 2
E [Y − E(Y | X)] = min E [Y − f (X)]
f

where the functions f are measurable and square integrable.

Proposition:
(i) The conditional expectation E(Y | X) is not unique, in general. However,
it is almost surely unique.

(ii) The FOC that are here necessary and sufficient for a solution is that for
any measurable and square integrable function f ,
′ 
E [Y − E(Y | X)] [f (X)] = 0.

(iii) When X and Y are jointly Gaussian, then the conditional expectation
is equal to the linear regression and one can write:

′
E(Y | X) = β0 +β1 X 1 +...+βk X k , where X = X 1 , ..., X k and β0 , β1 , ..., βk
are real coefficients corresponding to the OLS formula.

Other sets of function f can be used to, and yield slightly different properties
of the conditional expectation.

13
The Law of Iterated Expectations
We will use the law of iterated expectations for calculating conditional ex-
pectation of the dependent variables on independent variables; and on diverse
random events. This property may not be as convenient for models for truncated
or censored data that are highly nonlinear.
The law of iterated expectations is what makes feasible many results of the
linear regression model. In particular, it allows the description of a (dependent)
variable as the mean over the covariates of the regression function (conditional
expectation conditional on covariates).

Proposition:
Let Y and X be two random variables and PX the probability function of
X, then (Law of Iterated Expectations):
Z
EY = E (Y | X = x )dP X (x ) = E PX [E (Y | X = x)] ,

where E(Y | X = x) = y dP Y |X=x (y) = y dPYX (y, x) and EPX denoted


R R

the expectation with respect to PX .

Proposition:
Let W be a random vector that determines X through a functional rela-
tionship, X = f (W ). Typically the dimension of X is smaller than that of W .
Then,

E(Y | X) = E [E (Y | W ) | X] .

On the other hand, we also have:

E(Y | X) = E [E (Y | X) | W ]

Indeed, knowing W implies knowing X, because X is a function of W . The


additional information in W is not useful for determining E(Y | X).
A way to exploit this insight is to note:
E(Y | X) = E (E (Y | X, u) | X), where u is any random vector. Therefore,
error term for a model can be eliminated through averating while considering
E(Y | X). This is also true when the error u are not additive in the model.
Here is an example of the use of the law of iterated expectation. Assume
the usual regression hypothesis E(Y | X) = 0. Then, for any function g(.),
E(Y.g(X)) = 0. Indeed, E(Y.g(X)) = E[E(Y.g(X)|X)] = E[g(X).E(Y |X)] =
E[g(X).0] = 0.

It is often convenient to specify a parametric model for E(Y | X) as an


approximation. For example, E(Y | X) ≃ Xβ, with β a vector of coefficients.
Then, one can write the model:

Y = Xβ + u,

14
where u is an unobserved residual. Since this model is only an approxima-
tion, u ̸= Y − E(Y | X) is likely. As a consequence, and the properties that u
is centered, and of u independent of the X ′ s may not be satisfied. That is: this
linear functional form may suffer from endogeneity issues for its estimation.

1.5 Utility of Information

Since predictions must serve to decisions, they must be based on information


useful to these decisions. This issue will be linked with the statistical decision
theory in chapter 3. However, we can have a first view of it.

Let us show that, under the expected utility hypothesis, information is useful
to decisions, and that its usefulness can be measured in terms of VNM utility
levels.
Let us express the idea of information through the conditioning of the expec-
tation operator on a given informative event. Indeed, a sigma-field of events is a
way to describe relevant and consistent information about the world. Consider
the following example with two random events A and B, inspired from Gollier
(2004) a changer. Let P (A) = 1/2, P (B) = 1/4, P (A ∩ B) = 1/8. Assume the
expected utility hypothesis and that the VNM utility function, u, under the act
X, has a level equal to u = 1 for any random state in A ∩ B̄ or in Ā ∩ B, and
that its level is equal to u = −1 otherwise. Because of the cardinality property
of the expected utility, it is always possible to choose a VNM utility function
that satisfies these normalizations.
example to changeExample 1 : The event A could be ‘rain tomorrow’ and
the event B could be ‘be given an umbrella tomorrow’, while the act X could
consist in deciding to use the umbrella if it is given, whether it rains or not. The
umbrella prevents one from getting wet when it rains, but is uselessly heavy to
carry when it does not rain.
It is easy to calculate: E(u(X)) = 0, E(u(X)|A) = 1/2 and E(u(X)|Ā) =
−1/2.
Indeed, Eu(X) = P (Ā ∩ B̄).u(Ā ∩ B̄) + P (A ∩ B̄).u(A ∩ B̄) + P (A ∩ B).u(A ∩
B) + P (Ā ∩ B).u(Ā ∩ B)
= P (Ā ∩ B̄).(−1) + P (A ∩ B̄).1 + P (A ∩ B).(−1) + P (Ā ∩ B).1.
Moreover, P (Ā∩B) = P (B)−P (A∩B) = 1/8, P (A∩B̄) = P (A)−P (A∩B) =
3/8, P (Ā ∩ B̄) = 1 − P (A) − P (Ā ∩ B) = 1 − 1/2 − 1/8 = 3/8. Therefore:
Eu(X) = −3/8+3/8−1/8+1/8 = 0. The conditional expectations E(u(X)| A)
and E(u(X)| Ā), respectively under the information that event A occurs or
that event A does not occur, can be calculated similarly with first computing
the appropriate conditional probabilities using the formula P (E | A) = P (E ∩
A)/P (A) for any event E.
The words that we use when we read P (E | A), as ‘Probability of E knowing
A’ fit well the interpretation of the conditioning on A as referring to information

15
about the occurrence of A. This can also be related to Bayes formula used to
update some information quantified by using probabilities.
Consider now another act X ′ such that the VNM utility level is now equal
to -1 for any random state in A ∩ B̄ or in Ā ∩ B, and that it is equal to 1 oth-
erwise. That is: exactly the opposite signs to above. In that case: E(u(X ′ )) =
0, E(u(X ′ )|A) = −1/2 and E(u(X ′ )|Ā) = 1/2, with a symmetric calculus as
above.
Therefore, the agent should be indifferent between the two acts X and X ′
since in both cases Eu(X) = 0 and Eu(X ′ ) = 0.
However, if the agent knew whether A is true or not, he would be able to
choose act X if A is true since E(u(X)|A) = 1/2 while E(u(X ′ )|A) = −1/2;
and instead to choose act X ′ if A is not true, since E(u(X ′ )|Ā) = 1/2 (that is:
‘Eu knowing non-A’, while E(u(X)|Ā) = −1/2. Therefore, the information on
the occurrence of A has a utility value of 1/2.
Therefore, the gain from information can not only be manifested by com-
paring the levels of E(u(X)|Ā) and E(u(X)), but also concretely by making
distinct decisions, according to whether the agent is or is not informed, through
assessing the consequences of the selected act vs. those of the non-selected act.
Note however, that this utility gain has only a relative, not an absolute, mean-
ing, because the cardinality of the VNM utility function makes it unsufficient
to identify the gain exactly.
The utility value of information is always nonnegative in this model. This
may not be the case with other ‘non expected utility’ models.
These results, under the expected utility hypothesis, justify the interest of
searching for more information for predicting or forecasting. However, in that
case, new complications may arise because of searching costs that should be
included in the decision model. Moreover, ex ante information search may itself
be risky because the agent does not know which observations he will make during
his search.

1.6 Reminders of Convergences and Derivatives

perhaps to move with probabilities?

Definition:
Let (Xn )n∈N be a sequence of real random variables and X is a real random
variable.
(i) Convergence in Probability:
P
Xn X as n −→ +∞ if and only if
−→

∀ε > 0, P (|Xn − X| > ε) −→ 0 as n −→ +∞.

16
(ii) Almost Sure Convergence:
a.s.
Xn X as n −→ +∞ if and only if
−→
 
∀ε > 0, P sup |Xm − X| > ε −→ 0 as n −→ +∞.
m≥n

The following laws of large numbers and central limit theorems are useful to
prove the consistency and the asymptotic normality of the estimators that will
be considered, including the OLS estimators.

Proposition: (i ) Weak Law of Large Numbers of Khinchine:

If Xi , i = 1, ..., n is iid with E(Xi ) = µ < ∞, then X̄ converges in probability


to µ.

Again here, by abuse of notation, the number µ is assimilated to a constant


random variable that takes the value µ with probability one.

(ii) Weak Law of Large Numbers of Chebishev:


If Xi , i = 1, ..., n follow a law stationary in covariances such that E(Xi ) =
µ < ∞, V ar(X) = σi2 < ∞ Pn
and ∀j, ∃γj < ∞, Cov(Xi , Xi+j ) = γj , and if lim n1 j=1 γj = 0,
n−→∞
then converges in probabilité to µ.

(iii) Strong Law of Large Numbers of Kolmogorov:


If the Xi , i = 1, ..., n are independent with E(Xi ) = µi < ∞, V ar(Xi ) =
σi2 < ∞

P∞ σi2
and i=1 i2 < ∞, Pn
1
then X̄ − µ̄ converges almost surely to zero, where µ̄ ≡ n i=1 µi .

(iv) If Xi , i = 1, ..., n converges to zero in mean of order r, then it converges


to zero in mean of order s, for all s < r.

(v) Central-Limit Theorem of Lindeberg-Lévy:


If Xi , i = 1, ..., n are iid with E(Xi ) = µ < ∞ and V (xi ) = σ 2 < ∞,

√  d
then n X̄ − µ N (0, σ 2 ).

17
(vi) Central-Limit Theorem of Lindeberg-Feller:
If Xi , i = 1, ..., n are independent with E(Xi ) = µi < ∞ and V (Xi ) = σi2 <
∞.

1
Pn 1
Pn
Let µ̄ = n i=1 µi and σ̄n2 = n i=1 σi2 .

max(σi )
If lim nσ̄n = 0 and if lim σ̄n2 = σ 2 < ∞,
n−→∞ n−→∞

√  d
then n X̄ − µ̄ N (0, σ̄ 2 ).

Proposition (derivatives):
(vii) Leibniz formula: Let be a n-times differentiable function f of two vari-
ables x and y,
n
∂ n−k ∂ k

∂ ∂
dn f = dx + dy f =nk=0 f dxn−k dy k .
∂x ∂y ∂xn−k ∂y k

(viii) Matrix derivatives: The matrix at the denominator is the one that
indicates
 ′the orientation of the elements in the obtained matrix of derivatives.
∂f ∂f
∂A = ∂A ′.

∂ (a′ x) ∂ (a′ x)
∂x = a and ∂x′ = a′ .

∂(Ax) ∂(Ax)
∂x = A′ , where A is n × k and x is k × 1; and ∂x′ = A.

∂ (x′ Ax)
∂x = 2Ax, when A is symmetric; = (A + A′ )x when A is squared k × k
but not symmetric.

∂ (x′ Ax) ∂ (x′ Ax)


∂A = xx′ and ∂aij = xi xj , where is squared k × k and aij is the
th
(i, j) general term of matrix A.

References:

Amemiya, T. (1985), ’Advanced Econometrics,’ Harvard University Press.


Gollier, C. (2004), ‘The Economics of Risk and Time,’ MIT Press.
Billingsley, P. (2012), ’Probability and Measure,’ Wiley.

18

You might also like